DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
The Amendment filed 05/12/2026 has been entered. Claims 4, 10, and 16 have been cancelled. Therefore, claims 1-3, 5-9, 11-15, and 17-25 remain pending in the application.
Response to Arguments
Applicant’s arguments, see pages 11-15, with respect to the 35 U.S.C. 101 abstract idea rejection of claims 1-3, 5-9, 11-15, and 17-25, have been fully considered but are not persuasive.
With respect to the 35 U.S.C. 101 abstract idea rejection, the Applicant asserts that the computer-implemented operations cannot be performed practically within the human mind. They state that querying a first language learning model to generate a text output is an operation performed by a trained language learning model, not a mental human step. They state that, per Specification paragraph [0054], the model’s generation process is computational and beyond human mental capability at scale, and the human mind is not able to perform such querying or generation. The Applicant points out a claimed distinction between a language learning model being used and a human answering the query themself. They reference Specification paragraph [0002], which states the intricacies and computational power of language learning models. They reiterate that this is not a description of a generic processor, memory, or other conventional computing hardware performing ordinary functions. The Applicant states that the claims use the first language learning model in a specific, non-generic manner. They further argue that the use of the first language learning model is non-generic by showing that the present application relies on idiosyncrasies in the model’s generated language to identify model training provenance. The Applicant asserts that the present application is directed to improving training data identification techniques for language learning models, although the Applicant does not list how or why this application manages to improve these techniques directly. The Applicant also asserts that once the claims require a trained language learning model to generate an output response, and further requires machine processing of that response into n-grams for source scoring and training-data identification, the alleged abstract idea can no longer fairly be characterized as something that can practically be performed in the human mind. They further assert that it is impractical for a human, using only mental processes or pen and paper, to determine whether a particular data source was used to train a language learning model. They then state that the present application addresses that technical problem by querying the model, generating n-grams from the model’s output, and computing source scores relative to candidate training data in order to identify likely training provenance.
The Examiner respectfully disagrees. The original claims, and the claims as amended, are merely utilizing computer devices, in this specific case “a first language learning model”, as tools to perform a method which is directed to an abstract idea. The claim, under its broadest reasonable interpretation, recites a system, method, and CRM of processing data output from a language model by splitting the output text into n-grams, scoring the grouped text based on multiple training data, and deciding which training data allowed the model to create its output based on the score. These are abstract ideas in the form of certain methods of organizing human activity (i.e. mental processes such as observation, evaluation, judgement, and opinion), as well as various mathematical operations (i.e. scoring by data comparison). The steps of receiving an output from a large language model, creating groupings of the output into n-grams, scoring the grouped output through a mathematical operation based on multiple training data, and data analysis by identifying which training data allowed the model to create its output based on the score could be performed by a human using pen and paper or by purely mental reasoning, save for the recitation of generic computer components. Further, the claims do not integrate the judicial exception into a practical application. The recitation of “a first language learning model” is a generic instruction to perform the abstract idea using a computer device. The language learning model performs completely generic actions within the scope of the claim by receiving data (i.e. a query), analyzing the data, and outputting data (i.e. a text output response). This deems it being treated it as a generic computing component. A human being is very capable of utilizing a generic computing component to then analyze the output data using the steps denoted above. The “first language learning model” is recited at such a high-level of generality and is used as merely a tool to perform the abstract idea. All language learning models have their own, associated idiosyncrasies. This even more so makes the model non-specific and generic within the scope of the claims. The Applicant does not state how or why this specific method of identifying training data from a model’s outputs improves the model itself, the computing platform itself, or any other technology or technical field. The claims do not include any additional elements that amount to significantly more than the judicial exception. The claims, as written and amended, do not include more than mere instructions to perform the abstract method using generic computer components. Hence, Applicant’s arguments are not persuasive.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim(s) 1-3, 5-9, 11-15, and 17-25 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: Independent claims 1, 7, 13, 19, and 23 recite a method (claim 1), system (claims 7, 13, and 19), and computer-readable medium (claims 13 and 23) (CRM). These claims therefore invoke a statutory category (machine and process) in Step 1 of the Subject Matter Eligibility Test.
Step 2A, Prong One: Independent claims 1, 7, and 13, under their broadest reasonable interpretation, recite a system, method, and CRM of processing data output from a language model by splitting the output text into n-grams, scoring the grouped text based on multiple training data, and deciding which training data allowed the model to create its output based on the score. Independent claims 19 and 23, under their broadest reasonable interpretation, recite a system and CRM of processing data output from a language model, scoring grouped text outputs based on multiple training data, ranking the training data based on their respective scores, and then using a second language learning model to output a response based on the ranking of the training data and a query. These are abstract ideas in the form of certain methods of organizing human activity (i.e. mental processes such as observation, evaluation, judgement, and opinion), as well as various mathematical operations (i.e. scoring by data comparison). The steps of receiving an output from a large language model, creating groupings of the output into n-grams, scoring the grouped output through a mathematical operation based on multiple training data, and data analysis by identifying which training data allowed the model to create its output based on the score could be performed by a human using pen and paper or by purely mental reasoning, save for the recitation of generic computer components. The steps of receiving an output from a large language model, scoring the grouped outputs through a mathematical operation based on multiple training data, ranking the training data based on the calculated scores, and then using a second language learning model to output a response based on the ranking of the training data and a query could be performed by a human using pen and paper or by purely mental reasoning, save for the recitation of generic computer components.
Step 2A, Prong Two: The claims do not integrate the judicial exception into a practical application. The recitation of “a first language learning model” and “a second language learning model” are generic instructions to perform the abstract idea on/using a computer and does not impose a meaningful limit on the judicial exception. The language learning models are recited at high-levels of generality and are merely used as tools to perform the abstract idea. The grouping of the output into n-grams, using a mathematical operation to score the output against training data, ranking of the scores of the training data, and data identification of the training data do not add any meaningful limitations to the method. Mere data gathering, data analysis, and mathematical operations do not provide an inventive concept. There is no improvement to the functioning of the language learning model, the functioning of the computer itself, or to any other technology or technical field.
Step 2B: The claims do not include any additional elements that amount to significantly more than the judicial exception. The only additional element beyond the abstract idea is the language learning model, which performs generic computational functions such as receiving, analyzing, and outputting data. Such elements are well-understood, routine, and conventional within the field.
Accordingly, claims 1, 7, 13, 19, and 23 are directed to an abstract idea and do not include significantly more than the abstract idea itself.
With respect to claims 2, 8, and 14, the claims relate to ranking the training data based on queries and their scores, then using a second language learning model to output a response based on the ranking of the training data and a query. This is merely data analysis that could be performed by a human using pen and paper or by purely mental reasoning. The only additional element is the “second language learning model”, which is a generic instruction merely utilizing the model to receiving, analyze, and output data and therefore does not impose a meaningful limit on the judicial exception. No additional elements are present.
With respect to claims 3, 9, and 15, the claims relate to scoring of the groupings based on the training data. This is a purely mathematical operation through data comparison and could be performed by a human using pen and paper or by purely mental reasoning. No additional elements are present.
With respect to claims 5, 11, and 17, the claims relate to calculation of the first score. This is a purely mathematical operation and could be performed by a human using pen and paper or by purely mental reasoning. No additional elements are present.
With respect to claims 6, 12, and 18, the claims relate to calculation of the second score. This is a purely mathematical operation and could be performed by a human using pen and paper or by purely mental reasoning. No additional elements are present.
With respect to claims 20 and 24, the claims relate to generating n-gram groupings of the output and calculating the score of the grouping based on the training data. This is a mental process that could be performed by a human using pen and paper or by purely mental reasoning. No additional elements are present.
With respect to claims 21 and 25, the claims relate to calculating scores of the groupings based on respective training data. This is a purely mathematical operation through data comparison and could be performed by a human using pen and paper or by purely mental reasoning. No additional elements are present.
With respect to claim 22, the claim relates to the calculations of both the first and second scores. These are purely mathematical operations and could be performed by a human using pen and paper or by purely mental reasoning. No additional elements are present.
Allowable Subject Matter
Claims 1-3, 5-9, 11-15, and 17-25 would be allowable if rewritten or amended to overcome the rejection under 35 U.S.C. 101.
The following is a statement of reasons for the indication of allowable subject matter:
The prior art taken alone or in combination fails to teach the combination of limitations recited in the independent claims including steps of “wherein the source score is determined as
S
c
o
r
e
S
=
m
a
x
(
0
,
S
c
o
r
e
1
-
S
c
o
r
e
2
)
, wherein
S
c
o
r
e
S
represents the source score, wherein
S
c
o
r
e
1
represents the first score of the groupings, and wherein
S
c
o
r
e
2
represents the second score of the groupings”.
Ulasen et al. (US Patent Application Publication No. 2023/0325717) discloses querying a first language learning model with a first query, wherein the first language learning model generates a text output response to the first query, generating a source score of the groupings based on a first training data and a second training data, and identifying the first training data as training data of the first language learning model based on the source score.
Xu et al. (US Patent No. 8,612,367) discloses generating groupings of the text output response, wherein the groupings include multiple n-grams.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
US Patent Application Publication No. 2023/0315856
US Patent No. 11,783,175
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ADAM MICHAEL WEAVER whose telephone number is (571)272-7062. The examiner can normally be reached Monday-Friday, 8AM-5PM EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richemond Dorvil can be reached at (571) 272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ADAM MICHAEL WEAVER/Examiner, Art Unit 2658
/RICHEMOND DORVIL/Supervisory Patent Examiner, Art Unit 2658