DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This communication is responsive to the application filed on August 12, 2025. Claims 1-4 are pending at the time of examination.
Priority
Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). The certified copy has been filed in instant application, filed on March 24, 2026.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on August 12, 2025 was considered by the examiner. See attached PTO-form 1449.
Specification
The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-4 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. The claim recites “a system”, however, there is no hardware component (i.e., CPU processor) recited in claim in order to enable the function to be realized. Thus, at best, the claim is a software per se. Thus, software per se claim is not one of the four statutory categories.
Claims 1-4 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
The claim 1 recites “a calculator configured to calculate a distance in a feature space between a first feature vector corresponding to first text data registered in a database and a second feature vector corresponding to second text data registered in the database; a determiner configured to determine whether the calculated distance is smaller than a first threshold; and a controller configured to delete one of the first text data and the second text data when the calculated distance is smaller than the first threshold.”
Claim 3 recites “a calculator configured to calculate a distance in a feature space between a first feature vector corresponding to first text data and a second feature vector corresponding to second text data registered in a database; a determiner configured to determine whether the calculated distance is smaller than a first threshold; and a controller configured to either (i) maintain registration of the second text data without registering the first text data in the database or (ii) register the first text data in the database and delete the second text data, when the calculated distance is smaller than the first threshold.”
The limitation “a calculator configured to calculate a distance in a feature space between a first feature vector corresponding to first text data registered in a database and a second feature vector corresponding to second text data registered in the database; a determiner configured to determine whether the calculated distance is smaller than a first threshold; and a controller configured to delete one of the first text data and the second text data when the calculated distance is smaller than the first threshold”, as drafted, are process that, under their broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components. That is, other than reciting “processing system,” nothing in the claim element precludes the step from practically being performed in the mind. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind and/or manually performed, but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
This judicial exception is not integrated into a practical application. In particular, the claim only recites one additional element – using “a controller” (assuming as computer processor” to perform the claims steps. The “computer processor” in these steps is recited at a high-level of generality (i.e., as a computer” performing a generic computer functions) such that it amounts no more than mere instructions to apply the exception using a generic computer component. The claim also recites the additional element “to delete one of the first text data and the second text data when the calculated distance is smaller than the first threshold” that is the insignificant extra-solution activity of data gathering and/or output, and can be understood as activities incidental to the primary process or product that are merely a nominal or tangential addition to the claim (see MPEP 2106.05(g)). Accordingly, these additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea. (see 2019 Revised Patent Subject Matter Eligibility Guidance, Step 2A, Prong Two. See also MPEP 2106.04(II)(A)(2), MPEP 2106.04(d).
The claims 2 and 4 do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of using “a computer” to perform the claimed steps amounts to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. The claims are not patent eligible. see 2019 Revised Patent Subject Matter Eligibility Guidance, Step 2B. See also MPEP 2106.05
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-4 are rejected under 35 U.S.C. 103 as being unpatentable over Vasilyev et al. (US 2026/0127459 A1) in view of Zhang et al. (US 20260064745 A1).
As per claim 1, Vasilyev discloses an information processing system comprising:
a calculator configured to calculate a distance in a feature space between a first feature vector corresponding to first text data registered in a database and a second feature vector corresponding to second text data registered in the database (para. 0053, as Chunker 115 is a computer system that splits the data gathered from an electronic object into smaller groupings or chunks; para. 0054, as Embedder 120 is a computer system that represents non-numeric content, such as text and images, as a numerical value so that computer systems can more efficiently manipulate the content. These numerical values (often referred to as embeddings or Ex) are usually multi-dimensional vectors comprising V.sub.1, V.sub.2, V.sub.3, . . . , V.sub.n. Examples of embedders include image embedders and word or string/sentence embedders. Word or string/sentence embedders encode the meaning of a word (or a group of words) as a numerical value; para. 0056, as Knowledge base 150 is a data store for each chunk for the electronic objects 20 crawled by crawler 110. As illustrated in FIG. 1, the data record for each chunk stored in knowledge base 150 preferably contains the metadata for the associated electronic object 20-n (as retrieved by crawler 110); the raw content associated with that chunk (e.g., text, image); and the embedding, E.sub.n (depicted as a vector having numerical components V.sub.1, V.sub.2 . . . V.sub.n; para. 0081, as );
a determiner configured to determine whether the calculated distance is smaller than a first threshold (para. 0069, as Embedder 140 receives a query, Q, from data input/output system 130 and generates an embedding, E.sub.Q. Embedding E.sub.Q is preferably a multi-dimensional vector. As illustrated in FIGS. 1 and 3, E.sub.Q is used by search engine 160 to retrieve a limited set of chunks (170) (e.g., E.sub.1, E.sub.2, E.sub.3, . . . E.sub.m) from knowledge base 150 that are similar to the query embedding (with similarity being defined either by cosine between the vectors or by distance (length of the difference between the vectors)). In particular, search engine 160 uses either cosine or distance similarity between the query embeddings (produced by query embedder 140) and the text embeddings (produced by text embedder 120) and stored in the knowledge base 150 to retrieve potentially relevant embeddings; para. 0083, as The number of clusters created by the system requires a trade-off between average distance of a cluster (which may be measured from the centroid of the cluster) to the query embedding (here the lower the distance is closer to the context of query) and the distance between the clusters (again, preferably measured between centroids) (here the greater the distance the better)).
Vasilyev does not explicitly teach, but Zhang teaches a controller configured to delete one of the first text data and the second text data when the calculated distance is smaller than the first threshold (para 0019, as When a user query is received, a query vector is generated by the server. The query vector is compared with vector data of the knowledge corpus to determine whether the query vector is sufficiently similar to the knowledge corpus. When the query vector is sufficiently similar, and thus the knowledge corpus is relevant to the query, the language model processes the query and returns a response. When the query vector is not sufficiently similar, the query may be rejected). Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention was made to modify the teachings of Vasilyev to implement the above step as taught by Zhang because it would enable the system to avoid from returning incorrect answer to a user query. Motivation to do so would improve of saving computing resource.
As per claim 2, Vasilyev further teaches when the calculated distance is greater than the first threshold, the determiner determines whether the calculated distance is smaller than a second threshold that is greater than the first threshold; and when the calculated distance is smaller than the second threshold, the controller associates the first text data with the second text data (para. 0088).
As per claim 3, Vasilyev discloses an information processing system comprising
a calculator configured to calculate a distance in a feature space between a first feature vector corresponding to first text data and a second feature vector corresponding to second text data registered in a database para. 0053, as Chunker 115 is a computer system that splits the data gathered from an electronic object into smaller groupings or chunks; para. 0054, as Embedder 120 is a computer system that represents non-numeric content, such as text and images, as a numerical value so that computer systems can more efficiently manipulate the content. These numerical values (often referred to as embeddings or Ex) are usually multi-dimensional vectors comprising V.sub.1, V.sub.2, V.sub.3, . . . , V.sub.n. Examples of embedders include image embedders and word or string/sentence embedders. Word or string/sentence embedders encode the meaning of a word (or a group of words) as a numerical value; para. 0056, as Knowledge base 150 is a data store for each chunk for the electronic objects 20 crawled by crawler 110. As illustrated in FIG. 1, the data record for each chunk stored in knowledge base 150 preferably contains the metadata for the associated electronic object 20-n (as retrieved by crawler 110); the raw content associated with that chunk (e.g., text, image); and the embedding, E.sub.n (depicted as a vector having numerical components V.sub.1, V.sub.2 . . . V.sub.n);
a determiner configured to determine whether the calculated distance is smaller than a first threshold (para. 0069, as Embedder 140 receives a query, Q, from data input/output system 130 and generates an embedding, E.sub.Q. Embedding E.sub.Q is preferably a multi-dimensional vector. As illustrated in FIGS. 1 and 3, E.sub.Q is used by search engine 160 to retrieve a limited set of chunks (170) (e.g., E.sub.1, E.sub.2, E.sub.3, . . . E.sub.m) from knowledge base 150 that are similar to the query embedding (with similarity being defined either by cosine between the vectors or by distance (length of the difference between the vectors)). In particular, search engine 160 uses either cosine or distance similarity between the query embeddings (produced by query embedder 140) and the text embeddings (produced by text embedder 120) and stored in the knowledge base 150 to retrieve potentially relevant embeddings; para. 0083, as The number of clusters created by the system requires a trade-off between average distance of a cluster (which may be measured from the centroid of the cluster) to the query embedding (here the lower the distance is closer to the context of query) and the distance between the clusters (again, preferably measured between centroids) (here the greater the distance the better)).
Vasilyev teaches a controller configured to either (i) maintain registration of the second text data without registering the first text data in the database or (ii) register the first text data in the database (para. 0078, as the clustering tasks (360, 365) may extract less duplicative and less trivially related (but still interesting) electronic data objects from the limited set, toward finding appropriate representatives for input into the (L)LM 220. In other words, applying the optional clustering may be more likely to provide incremental data that is a greater distance from the query (while still requiring a strong connection to the query).
Vasilyev does not explicitly teach, but Zhang teaches delete the second text data, when the calculated distance is smaller than the first threshold (para 0019, as When a user query is received, a query vector is generated by the server. The query vector is compared with vector data of the knowledge corpus to determine whether the query vector is sufficiently similar to the knowledge corpus. When the query vector is sufficiently similar, and thus the knowledge corpus is relevant to the query, the language model processes the query and returns a response. When the query vector is not sufficiently similar, the query may be rejected). Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention was made to modify the teachings of Vasilyev to implement the above step as taught by Zhang because it would enable the system to avoid from returning incorrect answer to a user query. Motivation to do so would improve of saving computing resource.
As per claim 4, Vasilyev further teaches determines whether the calculated distance is smaller than a second threshold that is greater than the first threshold; and registers the first text data in the database in association with the second text data (para. 0080-0081).
Conclusion
The prior art made of record, listed on form PTO-892, and not relied upon, if any, is considered pertinent to applicant's disclosure.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DEBBIE M LE whose telephone number is (571)272-4111. The examiner can normally be reached 9:00-5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Charles Rones can be reached at 571-272-4085. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/DEBBIE M LE/Primary Examiner, Art Unit 2168 July 11, 2026