Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
1. This action is responsive to the communication filed on 07/09/2026. Claims 1, 18 and 19 have been amended. Claims 3-4 and 6-7 have been cancelled. Claims 1-2, 5, 8-19 are pending.
2. Applicants' arguments filed 07/09/2026 have been fully considered but they are not deemed to be persuasive. Rejections and/or objections not reiterated from previous office actions are hereby withdrawn. The following rejections and/or objections are either reiterated or newly applied. They constitute the complete set presently being applied to the instant application.
Claim Rejections - 35 USC § 101
3. 35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
4. Claims 1-2, 5 and 8-19 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The analysis below of the claims’ subject matter eligibility follows the guidance set forth in MPEP 2106 which has incorporated the 2019 PEG.
Regarding to claim 1,
Step 1 Analysis: Claim 1 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis:
Claim 1 recites: A method for generating an answer based on retrieval-augmented generation (RAG), the method comprising:
“acquiring a query”. This element reads on a person acquires a query which could be considered a mental process of an observation or evaluation.
“performing a first evaluation task based on the query using a pre-trained critique model, wherein the critique model determines whether to refer to a document retrieval result based on the query, and outputs one of a [Retrieval] token and a [No Retrieval] token based on the determination result”. This element reads on a person performs a first evaluation task based on the query using a pre-trained critique model and determine whether to refer to a document retrieval result based on the query, and outputs one of a [Retrieval] token and a [No Retrieval] token based on the determination result which could be considered a mental process of an observation or evaluation.
“retrieving documents related to the query based on a result of the first evaluation task”. This element reads on a person retrieves documents related to the query based on a result of the first evaluation task which could be considered a mental process of an observation or evaluation.
“performing a second evaluation task based on the query and the retrieved documents using the critique model, wherein the critique model determines relevance between the query and the retrieved documents and assigns one of a [Relevant] token and an [Irrelevant] token to the retrieved documents based on the determination result”. This element reads on a person performs a second evaluation task based on the query and the retrieved documents using the critique model, wherein the critique model determines relevance between the query and the retrieved documents and assigns one of a [Relevant] token and an [Irrelevant] token to the retrieved documents based on the determination result which could be considered a mental process of an observation or evaluation.
“generating one or more answers, based on the query and one or more documents assigned with the [Relevant] token as related documents, using a large language model (LLM) according to a result of the second evaluation task”. This element reads on a person generates one or more answers, based on the query and one or more documents assigned with the [Relevant] token as related documents, using a large language model (LLM) according to a result of the second evaluation task which could be considered a mental process of an observation or evaluation.
Overall, the limitations directed to generate answer(s) for a query and the various mental process limitations in the context of this claim encompasses limitations that are not only considered to be directed to limitations that could be practically performed in the human mind (including observations and preform an evaluation, judgment, and opinion) aided by the use of pen and paper. If the claim limitations, under their broadest reasonable interpretations, cover performance of the limitation in the mind but for the recitation of generic computer components, then they fall within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
Step 2A Prong Two Analysis: In Step 2A Prong 2, we are directed to Identify whether there are any additional elements recited in the claim beyond the judicial exception(s), and evaluate those additional elements to determine whether they integrate the exception into a practical application of the exception.
In particular, the claim only recites the additional elements of “processor-implemented”, and “a computer device.”
Regarding the processor-implemented,
The processor of a computer system for generating and storing in all steps is recited at a high level of generality, i.e., as a generic processor performing a generic computer function of processing data (generating and storing). This generic processor limitation is no more than mere instructions to apply the exception using a generic computer component(s). Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea
Regarding the computer device,
The processor of a computer device for generating and storing in all steps is recited at a high level of generality, i.e., as a generic processor performing a generic computer function of processing data (generating and storing). This generic processor limitation is no more than mere instructions to apply the exception using a generic computer component(s). Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea.
The additional elements “processor-implemented” and “a computer device” are simply applying the abstract idea, and there is nothing done with results. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea, and does not provide any improvement in computer technology (see MPEP2106.05(a)).
Therefore, the additional element(s) do not integrate the judicial exception into a practical application.
Step 2B Analysis: In Step 2B, we are directed to Identify whether there are any additional elements recited in the claim beyond the judicial exception(s), and evaluate those additional elements to determine whether the additional elements, taken individually and in combination, result in the claim as a whole amounting to significantly more than the judicial exception.
As discussed above with respect to integration of the abstract idea into a practical application, The additional elements “processor-implemented” and “a computer device” is simply applying the abstract idea, and there is nothing done with results.
Accordingly, this additional element(s), taken individually and in combination, do not result in the claim as a whole amounting to significantly more than the judicial exception. The claim is not patent eligible.
Regarding claim 2,
Step 1 Analysis: Claim 2 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis:
Claim 2 is dependent on claim 1, which as indicated in the analysis above, is directed to an abstract idea without significantly more.
Claim 2 recites “generating an answer corresponding to the query, based only on the query, using the LLM model according to a result of the second evaluation task." That is, the claim recites generating an answer corresponding to the query, based only on the query, using the LLM model according to a result of the second evaluation task. The above-noted limitation of claim 2, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
Step 2A Prong Two Analysis: This judicial exception is not integrated into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Regarding claim 5,
Step 1 Analysis: Claim 5 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis:
Claim 5 is dependent on claim 1, which as indicated in the analysis above, is directed to an abstract idea without significantly more.
Claim 5 recites “resorting ranks of the retrieved documents." That is, the claim recites resorting ranks of the retrieved documents. The above-noted limitation of claim 5, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
Step 2A Prong Two Analysis: This judicial exception is not integrated into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Regarding claim 8,
Step 1 Analysis: Claim 8 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis:
Claim 8 is dependent on claim 1, which as indicated in the analysis above, is directed to an abstract idea without significantly more.
Claim 8 recites “performing a third evaluation task by inputting the one or more related documents and the one or more answers into the critique model." That is, the claim recites performing a third evaluation task by inputting the one or more related documents and the one or more answers into the critique model. The above-noted limitation of claim 8, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
Step 2A Prong Two Analysis: This judicial exception is not integrated into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Regarding claim 9,
Step 1 Analysis: Claim 9 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis:
Claim 9 is dependent on claims 1&8, which as indicated in the analysis above, is directed to an abstract idea without significantly more.
Claim 9 recites “wherein the third evaluation task is a task for evaluating groundedness between the related documents and the answers." That is, the claim recites the third evaluation task is a task for evaluating groundedness between the related documents and the answers. The above-noted limitation of claim 9, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
Step 2A Prong Two Analysis: This judicial exception is not integrated into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Regarding claim 10,
Step 1 Analysis: Claim 10 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis:
Claim 10 is dependent on claims 1&8, which as indicated in the analysis above, is directed to an abstract idea without significantly more.
Claim 10 recites “wherein the critique model, when performing the third evaluation task, determines groundedness between the related documents and the answers and assigns one of a [Fully Supported] token, a [Partially Supported] token, and a [Not Supported] token to the answers based on the determination result." That is, the claim recites the critique model, when performing the third evaluation task, determines groundedness between the related documents and the answers and assigns one of a [Fully Supported] token, a [Partially Supported] token, and a [Not Supported] token to the answers based on the determination result. The above-noted limitation of claim 10, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
Step 2A Prong Two Analysis: This judicial exception is not integrated into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Regarding claim 11,
Step 1 Analysis: Claim 11 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis:
Claim 11 is dependent on claim 1, which as indicated in the analysis above, is directed to an abstract idea without significantly more.
Claim 11 recites “performing a fourth evaluation task by inputting the query and the one or more answers into the critique model." That is, the claim recites performing a fourth evaluation task by inputting the query and the one or more answers into the critique model. The above-noted limitation of claim 11, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
Step 2A Prong Two Analysis: This judicial exception is not integrated into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Regarding claim 12,
Step 1 Analysis: Claim 12 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis:
Claim 12 is dependent on claims 1&11, which as indicated in the analysis above, is directed to an abstract idea without significantly more.
Claim 12 recites “the fourth evaluation task is a task for evaluating a utility score between the query and the answers." That is, the claim recites the fourth evaluation task is a task for evaluating a utility score between the query and the answers. The above-noted limitation of claim 12, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
Step 2A Prong Two Analysis: This judicial exception is not integrated into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Regarding claim 13,
Step 1 Analysis: Claim 13 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis:
Claim 13 is dependent on claims 1&11, which as indicated in the analysis above, is directed to an abstract idea without significantly more.
Claim 13 recites “wherein the critique model, when performing the fourth evaluation task, determines a utility score between the query and the answers and assigns one of a [Utility 1] token to a [Utility 5] token to the answers based on the determination result." That is, the claim recites the critique model, when performing the fourth evaluation task, determines a utility score between the query and the answers and assigns one of a [Utility 1] token to a [Utility 5] token to the answers based on the determination result. The above-noted limitation of claim 13, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
Step 2A Prong Two Analysis: This judicial exception is not integrated into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Regarding claim 14,
Step 1 Analysis: Claim 14 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis:
Claim 14 is dependent on claim 1, which as indicated in the analysis above, is directed to an abstract idea without significantly more.
Claim 14 recites “calculating critique scores for the one or more answers; and determining a final answer, based on the calculated critique scores." That is, the claim recites calculating critique scores for the one or more answers; and determining a final answer, based on the calculated critique scores. The above-noted limitation of claim 14, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
Step 2A Prong Two Analysis: This judicial exception is not integrated into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Regarding claim 15
Step 1 Analysis: Claim 15 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis:
Claim 15 is dependent on claim 1, which as indicated in the analysis above, is directed to an abstract idea without significantly more.
Claim 15 recites “wherein the critique model is generated by fine-tuning a pre-trained language model (PLM), based on learning data for respective tasks." That is, the claim recites the critique model is generated by fine-tuning a pre-trained language model (PLM), based on learning data for respective tasks. The above-noted limitation of claim 15, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
Step 2A Prong Two Analysis: This judicial exception is not integrated into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Regarding claim 16,
Step 1 Analysis: Claim 16 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis:
Claim 16 is dependent on claims 1&15, which as indicated in the analysis above, is directed to an abstract idea without significantly more.
Claim 16 recites “wherein a method of fine-tuning the PLM model is a coarse-to-fine learning method." That is, the claim recites a method of fine-tuning the PLM model is a coarse-to-fine learning method. The above-noted limitation of claim 16, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
Step 2A Prong Two Analysis: This judicial exception is not integrated into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Regarding claim 17,
Step 1 Analysis: Claim 17 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis:
Claim 17 is dependent on claims 1&15-16, which as indicated in the analysis above, is directed to an abstract idea without significantly more.
Claim 17 recites “wherein the coarse-to-fine learning method is a method of sequentially performing zero-shot learning, one-shot learning, and few-shot learning." That is, the claim recites the coarse-to-fine learning method is a method of sequentially performing zero-shot learning, one-shot learning, and few-shot learning. The above-noted limitation of claim 17, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
Step 2A Prong Two Analysis: This judicial exception is not integrated into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Claim 18 is rejected under 35 U.S.C. 101 with the same rational of claim 1.
Claim 19 is rejected under 35 U.S.C. 101 with the same rational of claim 1.
Claim Rejections - 35 USC § 103
5. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
6. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
7. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
8. The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
9. Claims 1-2, 5, 8, 11-12, 14-15, 18 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Le et al (US 20250362890 A1, hereinafter “Le”) in view of Wang et al (U.S. 20250124264 A1 hereinafter, “Wang”), and further in view of So et al (U.S. 20210184976 A1 hereinafter, “So”), and further in view of JP 5603468 B1 et al (JP 5603468 B1 hereinafter, “JP 5603468 B1”).
10. With respect to claim 1,
Le discloses
a method for generating an answer based on retrieval-augmented generation (RAG) performed by a computing device, the method comprising:
acquiring a query;
performing a first evaluation task based on the query using a pre-trained critique model;
retrieving data related to the query based on a result of the first evaluation task;
performing a second evaluation task based on the query and the retrieved data using the critique model; and
generating one or more answers, based on the query and one or more related data, using a large language model (LLM) according to a result of the second evaluation task (Le [0018] – [0040], [0094] – [0095], [0099], [0101] – [0105], claims 1-3 e.g. [0020] Embodiments described herein provide a number of benefits. For example, by including a preemptive feedback layer before execution, execution of dangerous code may be avoided. By using multiple critics with different goals, the quality of generated text is improved while avoiding complex alignment finetuning or prompt engineering to generate critiques. As such, computing resources are reduced over a single complex critique model. Therefore, with improved performance on text and/or executable code generation neural network technology in automatic code generation is improved. [0021] FIG. 1 is a simplified diagram illustrating a code generation framework 100 according to some embodiments. The framework 100 comprises an Actor LLM 104, an internal dialogue of critiques 132 including multiple critics, and an executor 114. The Actor LLM 104 is provided a task 120 from a user 102 via a user interface. The Actor LLM 104 generates a response (e.g., generated code) 124 that actor LLM 104 may revise using one or more feedback methods to provide a final response 122 (e.g., validated generated code). Internal dialogues of critiques 132 may generate text critiques of generated code which may be used by actor LLM 104 in updating the generated code. Executor 114 may execute code, and the results of the execution may be used by actor LLM 104 in updating generated code. The code generation with feedback framework may be further described as follows. [0022] Actor LLM 104 may generate generated solution 124 based on task 120. The prompt for actor LLM 104 to generate generated solution 124 may include the task 120 and may include a system prompt providing context and general code execution instructions. [0023] In some embodiments, rather than just a single critic for a specific code attribute, multiple critics 108 may be utilized to provide critiques of different types. A distinct critic may be a LLM that is fine-tuned and/or prompted differently in order to generate a text critique of generated code for a specific intended goal. For example, a safety critic 110 and a helpfulness critic 112 may be utilized. a safety-driven critic may be represented as o and a helpfulness-driven critic may be represented as w. The critics 108 may be initialized as LLMs configured by specific system prompts (Ps and Ph respectively) to establish the critics' corresponding roles. … [0026] – [0027] … Solution: {answer} [0028] The critic models 108 may generate critiques one after the other, and the critique of one may be included in the prompt for the other. In some embodiments, critics may generate additional iterations of critiques, each time with prior critiques included in the prompt. Effectively, the sequence of additional critiques may be viewed as a conversation between the critic models. The number of iterations of critiques may be limited by a configured maximum number of iterations. In some embodiments, the number of iterations is dynamic. For example, the number of iterations of critiques may be based on a quality of one or more of the generated critiques. For example, critiques may stop when the critiques do not include additional information, or they indicate that the code does not have any identified issues. Given an interaction turn r between critics, the output distributions may be redefined as: … [0029] In some embodiments, a summarizer LLM may be utilized to summarize the interactions between the critics 108. Summarizations may be generated at intermediate times during critique iterations, and/or a summary may be generated after all iterations of the critic models are complete to provide a summary of the full critic "conversation" to the actor LLM 104. Practically, to avoid computation overhead, i may be limited to only the last few turns of interactions. Alternatively, the critic dialogue may be summarized after each turn of interactions and only use the corresponding summary in each turn: {circumflex over (L)}.sub.r=f(i.sub.1 :r) where f(·) is parameterized as an LLM-based summarizer model. To revise the solutions from actor LLM 104 by both safety critic 110 and helpfulness critic 112, the summary may be reused in the last interaction turn R between the critics (thus, also reducing the computation cost on the actor LLM 104). To generate safety-and-helpfulness-aware outputs, the output distributions of the LLM code generator may be represented as: … [0038] … In some embodiments, task 120 may be a question, and the answer to the question may be answered with the aid of a program that is executed. For example, task 120 may be "how many prime numbers are there between 1 and 1 0." Answering the question may include generating code by actor LLM 104 including revisions based on feedback as described above. … [0039] FIG. 2 is a chart illustrating exemplary tool-enabled actions according to some embodiments. The critics (e.g., safety critic 110 and helpfulness critic 112) may be provided with access to external tools 106 and the tool query results may be incorporated as additional knowledge for the critics to generate more grounded critiques. As illustrated, two types of tool-enabled action may be performed by the critics. First, "code search" queries external tools 106 by a generated text query and optionally a corresponding code snippet. Second, "code review" uses the execution result of the code snippet (through a code interpreter) as additional input to complement the query. Both action types may query tools 106 like web search, and/or database/knowledge base searches such as Wikipedia and OpenAI knowledge base. [0040] A tool may be utilized by a critic 108 by the inclusion of an "Action" within a generated critique. For example, each critique from a critic 108 may include one or more sections which may include a "thought," an "action," and/or an "observation." These sections may be generated in the critiques based on a prompt that indicates critiques should be formatted to include these sections. The "thought" may provide am initial analysis of the task according to the specific critic prompt. The "action" may be formatted such that it provides the necessary information to access a tool. For example, an action may be "query='secure alternative to subprocess.Popen in python" which may result in a web search tool querying the web using the provided search terms, and providing a result based on the web search. The "observation" may provide the results of the tool. In the web search example, the observation may be a snippet of text from a website found in the search. [0094] At step 554, the system generates, via a second neural network based LM (e.g., helpfulness critic 112), a first critique text based on a first input prompt comprising a first instruction to evaluate an accuracy of the first code output compared with the input task. [0095] At step 556, the system generates, via a third neural network based LM (e.g., safety critic 110) different than the second neural network based LM, a second critique text based on a second input prompt comprising a second instruction to evaluate the first code output, the input task, and the first critique text. …. For example, the parameters of the second and third neural network may be provided different prompts to invoke different types of critiques. The different prompting may be performed on LMs with the same or different parameters. In some embodiments, the second input prompt comprises a second instruction to evaluate a safety of the first code output. For example, the second input prompt may prompt the system to generate a critique regarding the vulnerability of the first code output to certain attacks. [claims 1-3] 1. A method of jointly generating a code output by one or more neural network based language models, the method comprising: generating, via a first neural network based language model (LM), a first code output in a programming language in response to an input task description in a natural language; generating, via a second neural network based LM, a first critique text based on a first input prompt comprising a first instruction to evaluate an accuracy of the first code output compared with the input task description; generating, via a third neural network based LM different than the second neural network based LM, a second critique text based on a second input prompt comprising a second instruction to evaluate the first code output, the input task description, and the first critique text; generating, by the first neural network based LM, a revised code output based on an input prompt combining the input task description, the first code output, the first critique text, and the second critique text; and executing the revised code output in a code environment thereby producing a result to the input task description. 2. The method of claim 1, wherein the third neural network based LM is different than the second neural network based LM by at least one of: a different set of parameters; or a second input prompt different from the first input prompt. 3. The method of claim 2, wherein the second input prompt comprises a second instruction to evaluate a safety of the first code output. [0099] In one embodiment, methods 500 and 550 are applicable in a variety of applications. For example, the task request received by a neural network model (e.g., Actor LLM 104) may relate to a … a writing and/or editing request in a content generation system, an IT diagnostic request in an IT customer service support system, a navigation request in a robotic and autonomous system, and/or the like. By performing method 500 or 550, the neural network based artificial agent may improve technology in the respective technical field in healthcare and diagnostics, education and personalized learning, software development and code assistance, content creation, autonomous system (such as autonomous driving, etc.), and/or the like. [as
acquiring a query (e.g. query);
performing a first evaluation task (e.g. evaluate; task) based on the query (e.g. query) using a pre-trained critique model (e.g. critique/critic model);
retrieving data (e.g. first critique text) related to the query based on a result (e.g. first code output) of the first evaluation task;
performing a second evaluation task (e.g. evaluate; task) based on the query and the retrieved data (e.g. first critique text) using the critique model; and
generating one or more answers (e.g. solution{answer}), based on the query (e.g. query) and one or more related data (e.g. second critique text), using a large language model (LLM) (e.g. LLM) according to a result of (e.g. revised code output) the second evaluation task]).
Although Le substantially teaches the claimed invention, Le does not explicitly indicate documents.
Wang teaches the limitations by stating
acquiring a query;
performing a first evaluation task using a pre-trained critique model;
retrieving documents related to the query based on a result of the first evaluation task;
performing a second evaluation task based on the query and the retrieved documents using the critique model; and
generating one or more answers, based on the query and one or more related documents, using a large language model (LLM) according to a result of the second evaluation task (Wang [0020], [0048], [0064], [0103] – [0109] e.g. [0020] large language models (LLMs). [0048] Electronic documents can include a variety of content. For example, an electronic document 150 can include native content 152 that is within the electronic document 150 itself and/or does not change over time. Electronic documents can also include dynamic content that may change over time or on a per-request basis. For example, a publisher of a given electronic document (e.g., electronic document 150) can maintain a data source that is used to populate portions of the electronic document. In this example, the given electronic document can include a script, such as the script 154, that causes the client device 106 to request content (e.g., a digital component) from the data source when the given electronic document is processed (e.g., rendered or executed) by a client device 106 (or a cloud server). The client device 106 (or cloud server) integrates the content (e.g., digital component) obtained from the data source into the given electronic document to create a composite electronic document including the content obtained from the data source. [0064] A large language model ("LLM") is a model that is trained to generate and understand human language. LLMs are trained on massive datasets of text and code, and they can be used for a variety of tasks. For example, LLMs can be trained to translate text from one language to another; summarize text, such as web site content, search results, news articles, or research papers; answer questions, such as "What is the capital of Georgia?"; create chatbots that can have conversations with humans; and generate creative text, such as poems, stories, and code. [0103] In an example implementation, the prompts 220 submitted by the prompt engine 162 to the text generative model 230 specify a set of tasks including a first task for evaluating the digital component in the context of the current query and the past queries (e.g., by generating a textual critique of the initial digital component with respect to the context), a second task for generating a textual summary of user intent taking into account the current query and the past queries, a third task for generating a textual summary of the content of the resource and/or the digital component, a fourth task for generating the updated text related to the resource (e.g., using one or both textual summaries), and/or (v) a fifth task for generating an explanation for why the digital component is appropriate for (or otherwise why the digital component is being presented to) the user. The prompt engine 162 can generate prompts 220 for any combination of one or more of the tasks to generate the updated text for the digital component and/or an explanation for the digital component. [0105] In some implementations, the prompt for the fourth task can include information generated by the text generative model 230 based on prompts for the first task, the second task, and/or the third task in any order or combination. For example, the prompt engine 230 can submit, to the text generative model 230, a prompt 220 for the first task to obtain a textual critique of the initial digital component with respect to the context. This prompt 220 can include, for example, current query, data related to the initial digital component, and/or the sequence of related past queries. This 220 prompt can also include instructions for generating the textual critique. [0106] Similarly, the prompt engine 230 can submit, to the text generative model 230, a prompt 220 for the second task to obtain textual summary of user intent. This prompt can include, for example, the current query, the sequence of related past queries, and instructions for generating the textual summary based on the other information of the prompt 220. [0108] In addition, using the summary of the content in the prompt for the updated text and/or explanation can prevent "hallucinations" in the generated outputs-e.g., output content that is fictional, incorrect, misleading, or otherwise not grounded in factual or accurate information. [0109] In some implementations, the Al subsystem 160 can determine whether to perform the subsequent tasks (e.g., generate the updated text and/or explanation) depending on the critique generated by the text generative model 230 for the initial digital component. For example, if the generated critique indicates that the initial digital component does not provide a meaningful answer to the current query or does not address the user intent demonstrated by the past user queries, the system can determine to not proceed to perform the subsequent tasks for generating the textual summary of user intent and generating the updated text related to the resource [as
acquiring a query (e.g. query);
performing a first evaluation task (e.g. a first task for evaluating) based on the query (e.g. query) using a pre-trained critique model (e.g. critique generated by the text generative model);
retrieving documents (e.g. This prompt (such as: textual critique) 220 can include the data related to the digital component (e.g., content of the resource linked to by the digital component) and instructions for generating the textual summary based on the other information of the prompt 220 … Electronic documents can include a variety of content … digital component) related to the query based on a result of the first evaluation task (e.g. the prompts 220 submitted by the prompt engine 162 to the text generative model 230 specify a set of tasks including a first task for evaluating the digital component in the context of the current query and the past queries (e.g., by generating a textual critique of the initial digital component with respect to the context));
performing a second evaluation task (e.g. a second task for generating a textual summary of user intent taking into account the current query and the past queries) based on the query and the retrieved documents using the critique model; and
generating one or more answers (e.g. answer), based on the query and one or more related documents (e.g. This prompt (such as: textual critique) 220 can include the data related to the digital component (e.g., content of the resource linked to by the digital component) and instructions for generating the textual summary based on the other information of the prompt 220 … Electronic documents can include a variety of content … digital component), using a large language model (LLM) (e.g. LLM) according to a result of the second evaluation task (e.g. a second task for generating a textual summary of user intent taking into account the current query and the past queries)]).
Therefore, it would have been obvious to one of ordinary skill in the art at the time of the effective filing date of the invention, in view of the teachings of Le and Wang, to allow for a conversational interaction with computers using natural language rather than a restricted set of prompts. This allows for a more natural interaction with the computer (Wang [0003]).
Although Le and Wang combination substantially teaches the claimed invention, they do not explicitly indicate wherein the critique model determines whether to refer to a document retrieval result based on the query and outputs one of a [Retrieval] token and a [No Retrieval] token based on the determination result.
So teaches the limitations by stating wherein the critique model, when performing the first evaluation task, determines whether to refer to a document (e.g. document, web page)(So [0005], [0033] – [0037], [0040], [0095] e.g. [0033] …. The content items can also be displayed on a search results web page. For instance, the content provider 115 can provide or be the source of content items for display in content slots of information resources, such as a web page of a company where the primary content of the web page is provided by the company, or for display on a search results landing page provided by a search engine. …. [0034] …. The predicted request values can be generated by an external system based on historical data and associated retrieval tokens. [0035] The retrieval token receiver 125 may comprise an application, server, service, daemon, routine, or other executable logic for receiving retrieval tokens from one or more content providers, and may be executed by a processor of the computing system or a co-processor or other hardware (e.g. ASIC or FPGA circuits, etc.). The retrieval token receiver 125 can receive a plurality of retrieval tokens from the content provider 115. …. In addition, the retrieval tokens may be inserted into content items to increase the likelihood of their selection and insertion into information resources. If the retrieval token is somehow present in (e.g., a keyword on a web page, etc.) or directly related to (e.g., contains a similar language, or user demographic information, etc.) an information resource, the content item associated with the retrieval token may have a higher likelihood of being inserted into the information resource. Retrieval tokens may be associated with a particular quality, for example, a positive retrieval token or a negative retrieval token. …. [0037] In some implementations, the retrieval tokens can be associated with one or more information resources and/or documents. [0040] The bit strings can represent a document space for each of the predicted requests associated with the plurality of retrieval tokens. [0095] In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user's client device in response to requests received from the web browser.) retrieval result (e.g. search result) based on the query (e.g. search result) and outputs one of a [Retrieval] token (e.g. retrieval token) and a [No Retrieval] token based on the determination result (So [0048] – [0051] e.g. retrieval token).
Therefore, it would have been obvious to one of ordinary skill in the art at the time of the effective filing date of the invention, in view of the teachings of Le, Wang and So, to allow for a conversational interaction with computers using natural language rather than a restricted set of prompts. This allows for a more natural interaction with the computer (Wang [0003]).
Although Le, Wang and So combination substantially teaches the claimed invention, they do not explicitly indicate wherein the critique model, when performing the second evaluation task, determines relevance between the query and the retrieved documents and assigns one of a [Relevant] token and an [Irrelevant] token to the retrieved documents based on the determination result;
assigned with the [Relevant] token as.
JP 5603468 B1 teaches the limitations by stating
wherein the critique model determines relevance (e.g. relevance) between the query (e.g. search) and the retrieved documents (e.g. documents) and assigns one of a [Relevant] token (e.g. relevant tag) and an [Irrelevant] token to the retrieved documents based on the determination result; and
generating one or more answers, based on the query and one or more documents assigned with the [Relevant] token (e.g. relevant tag) as related documents (JP 5603468 B1 abstract, page 11 e.g. [abstract] A document classification system, a document classification method, and a document classification program that can reduce the burden on a reviewer are provided. An extraction unit that extracts a document group that is a data set including a predetermined number of documents from document information, and a classification code that a user assigns to the extracted document group based on relevance to a lawsuit. A classification code receiving unit that accepts, a selection unit that classifies the extracted document group for each classification code based on the classification code, and analyzes and selects commonly appearing keywords in the classified document group, and a selection A search unit that searches the keyword information from the document information, a score calculation unit that calculates a score indicating the relevance between the classification code and the document using the search result of the search unit and the analysis result of the selection unit, and the score result An automatic classification unit that automatically assigns a classification code to document information, and a display control unit that controls to display the calculation result of the score calculation unit and / or the classification result of the automatic classification unit on the screen. [page 11] The value(%) on the vertical axis of "Relevant Recall" shown in FIG. 21 is the number of documents with the Relevant tag out of the total number of documents in the denominator. This is the number of documents that are tagged).).
Therefore, it would have been obvious to one of ordinary skill in the art at the time of the effective filing date of the invention, in view of the teachings of Le, Wang, So and JP 5603468 B1, to allow for a conversational interaction with computers using natural language rather than a restricted set of prompts. This allows for a more natural interaction with the computer (Wang [0003]).
11. With respect to claim 2,
Wang further discloses generating an answer corresponding to the query, based only on the query, using the LLM model according to a result of the second evaluation task (Wang [0020], [0048], [0064], [0103] – [0109] e.g. [0064] A large language model ("LLM") is a model that is trained to generate and understand human language. LLMs are trained on massive datasets of text and code, and they can be used for a variety of tasks. For example, LLMs can be trained to translate text from one language to another; summarize text, such as web site content, search results, news articles, or research papers; answer questions, such as "What is the capital of Georgia?"; create chatbots that can have conversations with humans; and generate creative text, such as poems, stories, and code. [0103] In an example implementation, the prompts 220 submitted by the prompt engine 162 to the text generative model 230 specify a set of tasks including a first task for evaluating the digital component in the context of the current query and the past queries (e.g., by generating a textual critique of the initial digital component with respect to the context), a second task for generating a textual summary of user intent taking into account the current query and the past queries, a third task for generating a textual summary of the content of the resource and/or the digital component, …).
12. With respect to claim 5,
Le further discloses resorting ranks of the retrieved documents (Le [0056] e.g. The distribution parameters can also specify an eligibility value (e.g., ranking score, or some other specified value) that is used for evaluating the eligibility of the digital component for distribution/transmission (e.g., among other available digital components)).
13. With respect to claim 8,
Le further discloses performing a third evaluation task by inputting the one or more related documents and the one or more answers into the critique model (Le [0018] – [0040], [0094] – [0095], [0099], [0101] – [0105], claims 1-3 e.g. solution/answer) (Wang [0020], [0048], [0103] – [0109] e.g. [0109] In some implementations, the Al subsystem 160 can determine whether to perform the subsequent tasks (e.g., generate the updated text and/or explanation) depending on the critique generated by the text generative model 230 for the initial digital component. For example, if the generated critique indicates that the initial digital component does not provide a meaningful answer to the current query or does not address the user intent demonstrated by the past user queries, the system can determine to not proceed to perform the subsequent tasks for generating the textual summary of user intent and generating the updated text related to the resource).
14. With respect to claim 11,
Le further discloses performing a fourth evaluation task by inputting the query and the one or more answers into the critique model (Le [0018] – [0040], [0094] – [0095], [0099], [0101] – [0105], claims 1-3 e.g. [0103] In an example implementation, the prompts 220 submitted by the prompt engine 162 to the text generative model 230 specify a set of tasks including a first task for evaluating the digital component in the context of the current query and the past queries (e.g., by generating a textual critique of the initial digital component with respect to the context), a second task for generating a textual summary of user intent taking into account the current query and the past queries, a third task for generating a textual summary of the content of the resource and/or the digital component, a fourth task for generating the updated text related to the resource (e.g., using one or both textual summaries), and/or (v) a fifth task for generating an explanation for why the digital component is appropriate for (or otherwise why the digital component is being presented to) the user).
15. With respect to claim 12,
Le further discloses wherein the fourth evaluation task is a task for evaluating a utility score between the query and the answers (Le [0056] e.g. The distribution parameters can also specify an eligibility value (e.g., ranking score, or some other specified value) that is used for evaluating the eligibility of the digital component for distribution/transmission (e.g., among other available digital components)).
16. With respect to claim 14,
Le further discloses
calculating critique scores for the one or more answers; and
determining a final answer, based on the calculated critique scores (Le [0056] e.g. The distribution parameters can also specify an eligibility value (e.g., ranking score, or some other specified value) that is used for evaluating the eligibility of the digital component for distribution/transmission (e.g., among other available digital components)).
17. With respect to claim 15,
Le further discloses wherein the critique model is generated by fine-tuning a pre-trained language model (PLM), based on learning data for respective tasks (Le [0018] – [0040], [0094] – [0095], [0099], [0101] – [0105], claims 1-3 e.g. [0020] Embodiments described herein provide a number of benefits. For example, by including a preemptive feedback layer before execution, execution of dangerous code may be avoided. By using multiple critics with different goals, the quality of generated text is improved while avoiding complex alignment finetuning or prompt engineering to generate critiques. As such, computing resources are reduced over a single complex critique model. Therefore, with improved performance on text and/or executable code generation neural network technology in automatic code generation is improved. [0023] In some embodiments, rather than just a single critic for a specific code attribute, multiple critics 108 may be utilized to provide critiques of different types. A distinct critic may be a LLM that is fine-tuned and/or prompted differently in order to generate a text critique of generated code for a specific intended goal. [0095] For example, the second and third neural network based LMs may be fine-tuned, updating their respective parameters to improve their capability in the specific type of critique each is intended to generate).
Wang further discloses wherein the critique model is generated by fine-tuning a pre-trained language model (PLM), based on learning data for respective tasks (Wang [0131] – [0133], [0148] e.g. fine-tuning engine).
18. Claim 18 is same as claim 1 and is rejected for the same reasons as applied hereinabove.
19. Claim 18 is same as claim 1 and is rejected for the same reasons as applied hereinabove.
20. Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Le in view of Wang, So and JP 5603468 B1, and further in view of Gunjal et al (U.S. 20250335928 A1 hereinafter, “Gunjal”).
21. With respect to claim 9,
Although Le, Wang, So and JP 5603468 B1 combination substantially teaches the claimed invention, they do not explicitly indicate wherein the third evaluation task is a task for evaluating groundedness between the related documents and the answers.
Gunjal teaches the limitations by stating wherein the third evaluation task is a task for evaluating groundedness between the related documents and the answers (Gunjal [0037] e.g. [0037] In some implementations, the output verification module 142 mitigates potential for LLMs to produce erroneous and/or inaccurate outputs by checking LLM responses for factual accuracy and relevance. In this manner, the output verification module 142 functions as a quality control agent, ensuring that the information provided meets standards for reliability. For example, the output verification module 142 checks for the answer relevance, context relevance and groundedness of the output. In some examples, the output verification module 142 verifies citation uniform resource locators (URLs) whenever applicable. In the example case of product recommendation, the output verification module 142 checks for accuracy of the product details to be presented to the user by comparing the output to groundtruth data for the product).
Therefore, it would have been obvious to one of ordinary skill in the art at the time of the effective filing date of the invention, in view of the teachings of Le, Wang, So, JP 5603468 B1 and Gunjal, to allow for a conversational interaction with computers using natural language rather than a restricted set of prompts. This allows for a more natural interaction with the computer (Wang [0003]).
22. Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Le in view of Wang, So and JP 5603468 B1, and further in view of Palanisamy (U.S. 20150312038 A1 hereinafter, “Palanisamy”).
23. With respect to claim 10,
Although Le, Wang, So and JP 5603468 B1 combination substantially teaches the claimed invention, they do not explicitly indicate wherein the critique model, when performing the third evaluation task, determines groundedness between the related documents and the answers and assigns one of a [Fully Supported] token, a [Partially Supported] token, and a [Not Supported] token to the answers based on the determination result.
Palanisamy teaches the limitations by stating wherein the critique model, when performing the third evaluation task, determines groundedness between the related documents and the answers and assigns one of a [Fully Supported] token, a [Partially Supported] token, and a [Not Supported] token to the answers based on the determination result (Palanisamy [0085] e.g. supported token).
Therefore, it would have been obvious to one of ordinary skill in the art at the time of the effective filing date of the invention, in view of the teachings of Le, Wang, So, JP 5603468 B1 and Palanisamy, to allow for a conversational interaction with computers using natural language rather than a restricted set of prompts. This allows for a more natural interaction with the computer (Wang [0003]).
24. Claim 13 is rejected under 35 U.S.C. 103 as being unpatentable over Le in view of Wang, So and JP 5603468 B1, and further in view of Agbamu (U.S. 20230316261 A1 hereinafter, “Agbamu”).
25. With respect to claim 13,
Although Le, Wang, So and JP 5603468 B1 combination substantially teaches the claimed invention, they do not explicitly indicate wherein the critique model, when performing the fourth evaluation task, determines a utility score between the query and the answers and assigns one of a [Utility 1] token to a [Utility 5] token to the answers based on the determination result.
Agbamu teaches the limitations by stating wherein the critique model, when performing the fourth evaluation task, determines a utility score between the query and the answers and assigns one of a [Utility 1] token to a [Utility 5] token to the answers based on the determination result (Agbamu [0163] – [0166] e.g. utility token).
Therefore, it would have been obvious to one of ordinary skill in the art at the time of the effective filing date of the invention, in view of the teachings of Le, Wang, So, JP 5603468 B1 and Agbamu, to allow for a conversational interaction with computers using natural language rather than a restricted set of prompts. This allows for a more natural interaction with the computer (Wang [0003]).
26. Claim 16 is rejected under 35 U.S.C. 103 as being unpatentable over Le in view of Wang, So and JP 5603468 B1, and further in view of GAO et al (U.S. 20210174023 A1 hereinafter, “GAO”).
27. With respect to claim 16,
Although Le, Wang, So and JP 5603468 B1 combination substantially teaches the claimed invention, they do not explicitly indicate wherein a method of fine-tuning the PLM model is a coarse-to-fine learning method.
GAO teaches the limitations by stating wherein a method of fine-tuning the PLM model is a coarse-to-fine learning method (GAO [0081] e.g. [0081] At step 720, an underspecified rule sentence span is extracted based on the updated entailment state. In some embodiments, the rule sentence scores may be used for extracting an informative span from an underspecified rule sentence by utilizing a coarse-to-fine reasoning process. In this respect, token-level span distributions may be weighted with sentence-level selection scores of the rule span. For example, based on the sentence-level selection scores, the token-level span distributions may be modulated by its corresponding sentence-level selection score. The modulated token-level distributions may be used to identify the start and end positions of the underspecified rule span for extraction. In various embodiments, the extracted span is fed into a pre-trained language model to formulate a follow-up question by question rephrasing of the extracted span.).
Therefore, it would have been obvious to one of ordinary skill in the art at the time of the effective filing date of the invention, in view of the teachings of Le, Wang, So, JP 5603468 B1 and GAO, to allow for a conversational interaction with computers using natural language rather than a restricted set of prompts. This allows for a more natural interaction with the computer (Wang [0003]).
28. Claim 17 is rejected under 35 U.S.C. 103 as being unpatentable over Le in view of Wang, So, JP 5603468 B1 and GAO, and further in view of Shrivastava et al (U.S. 20230245654 A1 hereinafter, “Shrivastava”).
29. With respect to claim 17,
Although Le, Wang, So, JP 5603468 B1 and GAO combination substantially teaches the claimed invention, they do not explicitly indicate wherein the coarse-to-fine learning method is a method of sequentially performing zero-shot learning, one-shot learning, and few-shot learning.
Shrivastava teaches the limitations by stating wherein the coarse-to-fine learning method is a method of sequentially performing zero-shot learning, one-shot learning, and few-shot learning (Shrivastava [0085] e.g. one-shot learning).
Therefore, it would have been obvious to one of ordinary skill in the art at the time of the effective filing date of the invention, in view of the teachings of Le, Wang, So, JP 5603468 B1, GAO and Shrivastava, to allow for a conversational interaction with computers using natural language rather than a restricted set of prompts. This allows for a more natural interaction with the computer (Wang [0003]).
Response to Argument
30. On pages 6-7, Applicant argues The Amended Claims Are Not Directed to Mental Processes … Specifically, the amended Claim 1 now expressly recites that the pre-trained critique model "outputs one of a [Retrieval] token and a [No Retrieval] token" based on the first evaluation task, and "assigns one of a [Relevant] token and an [Irrelevant] token" to the retrieved documents based on the second evaluation task
Examiner disagrees because:
Regarding to claim 1,
Step 1 Analysis: Claim 1 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis:
Claim 1 recites: A method for generating an answer based on retrieval-augmented generation (RAG), the method comprising:
“acquiring a query”. This element reads on a person acquires a query which could be considered a mental process of an observation or evaluation.
“performing a first evaluation task based on the query using a pre-trained critique model, wherein the critique model determines whether to refer to a document retrieval result based on the query, and outputs one of a [Retrieval] token and a [No Retrieval] token based on the determination result”. This element reads on a person performs a first evaluation task based on the query using a pre-trained critique model and determine whether to refer to a document retrieval result based on the query, and outputs one of a [Retrieval] token and a [No Retrieval] token based on the determination result which could be considered a mental process of an observation or evaluation.
“retrieving documents related to the query based on a result of the first evaluation task”. This element reads on a person retrieves documents related to the query based on a result of the first evaluation task which could be considered a mental process of an observation or evaluation.
“performing a second evaluation task based on the query and the retrieved documents using the critique model, wherein the critique model determines relevance between the query and the retrieved documents and assigns one of a [Relevant] token and an [Irrelevant] token to the retrieved documents based on the determination result”. This element reads on a person performs a second evaluation task based on the query and the retrieved documents using the critique model, wherein the critique model determines relevance between the query and the retrieved documents and assigns one of a [Relevant] token and an [Irrelevant] token to the retrieved documents based on the determination result which could be considered a mental process of an observation or evaluation.
“generating one or more answers, based on the query and one or more documents assigned with the [Relevant] token as related documents, using a large language model (LLM) according to a result of the second evaluation task”. This element reads on a person generates one or more answers, based on the query and one or more documents assigned with the [Relevant] token as related documents, using a large language model (LLM) according to a result of the second evaluation task which could be considered a mental process of an observation or evaluation.
Overall, the limitations directed to generate answer(s) for a query and the various mental process limitations in the context of this claim encompasses limitations that are not only considered to be directed to limitations that could be practically performed in the human mind (including observations and preform an evaluation, judgment, and opinion) aided by the use of pen and paper. If the claim limitations, under their broadest reasonable interpretations, cover performance of the limitation in the mind but for the recitation of generic computer components, then they fall within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
31. On pages 7, Applicant argues The Claims Are Integrated Into a Practical Application … The present invention solves this technical problem through a specific technical solution, e.g., a pre-trained critique model that: Conditionally controls document retrieval by outputting a [Retrieval] token or [No Retrieval] token, thereby avoiding unnecessary retrieval operations when they would not improve answer quality; Filters retrieved documents by relevance by assigning [Relevant] or [Irrelevant] tokens, ensuring only pertinent documents are used for answer generation; and Produces measurably improved answers by grounding LLM-generated responses in documents verified as relevant by the critique model. …These operations collectively represent a specific technical improvement to RAG-based computer systems, i.e., reducing computational overhead from unnecessary retrieval, reducing hallucinations in LLM outputs, and improving the accuracy and reliability of generated answers.
Examiner disagrees because:
There is nothing stated in the instant applicant’s specification that Produces measurably improved answers by grounding LLM-generated responses in documents verified as relevant by the critique model or These operations collectively represent a specific technical improvement to RAG-based computer systems, i.e., reducing computational overhead from unnecessary retrieval, reducing hallucinations in LLM outputs, and improving the accuracy and reliability of generated answers.
32. On pages 7-8, Applicant argues The Claims Recite Significantly More Than the Abstract Idea The pre-trained critique model recited in the claims is not a generic computer component performing conventional operations. It is a specifically trained machine learning model, i.e., finetuned on task-specific learning data (Claim 15) using a coarse-to-fine learning method (Claims 16-17), that has been configured to perform multiple distinct evaluation tasks within a RAG pipeline.
Examiner disagrees because:
Regarding claim 15
Step 1 Analysis: Claim 15 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis:
Claim 15 is dependent on claim 1, which as indicated in the analysis above, is directed to an abstract idea without significantly more.
Claim 15 recites “wherein the critique model is generated by fine-tuning a pre-trained language model (PLM), based on learning data for respective tasks." That is, the claim recites the critique model is generated by fine-tuning a pre-trained language model (PLM), based on learning data for respective tasks. The above-noted limitation of claim 15, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
Step 2A Prong Two Analysis: This judicial exception is not integrated into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Regarding claim 16,
Step 1 Analysis: Claim 16 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis:
Claim 16 is dependent on claims 1&15, which as indicated in the analysis above, is directed to an abstract idea without significantly more.
Claim 16 recites “wherein a method of fine-tuning the PLM model is a coarse-to-fine learning method." That is, the claim recites a method of fine-tuning the PLM model is a coarse-to-fine learning method. The above-noted limitation of claim 16, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
Step 2A Prong Two Analysis: This judicial exception is not integrated into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Regarding claim 17,
Step 1 Analysis: Claim 17 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis:
Claim 17 is dependent on claims 1&15-16, which as indicated in the analysis above, is directed to an abstract idea without significantly more.
Claim 17 recites “wherein the coarse-to-fine learning method is a method of sequentially performing zero-shot learning, one-shot learning, and few-shot learning." That is, the claim recites the coarse-to-fine learning method is a method of sequentially performing zero-shot learning, one-shot learning, and few-shot learning. The above-noted limitation of claim 17, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
Step 2A Prong Two Analysis: This judicial exception is not integrated into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
33. On pages 8-9, Applicant alleges the architectural separation between the critique model and the LLM in the claimed invention and the critique model's role as an independent gatekeeper of the retrieval process is not taught or suggested by So. Moreover, So does not teach a critique model that performs the first evaluation task in conjunction with a second evaluation task assigning [Relevant]/[lrrelevant] tokens, a third evaluation task assessing groundedness, and a fourth evaluation task assigning utility scores all within a unified, multi-task evaluation pipeline. So's retrieval token is a standalone mechanism, not part of a structured, multi-stage critique of the kind claimed.
Examiner disagrees because:
In response to applicant’s argument that the references fail to show certain features of applicant’s invention, it is noted that the features upon which applicant relies (i.e., “the architectural separation between the critique model and the LLM in the claimed invention and the critique model's role as an independent gatekeeper of the retrieval process” or “a critique model that performs the first evaluation task in conjunction with a second evaluation task assigning [Relevant]/[lrrelevant] tokens, a third evaluation task assessing groundedness, and a fourth evaluation task assigning utility scores all within a unified, multi-task evaluation pipeline”) are not recited in the rejected claim(s). Although the claims are interpreted in light of the specification, limitations from the specification are not read into the claims. See In re Van Geuns, 988 F.2d 1181, 26 USPQ2d 1057 (Fed. Cir. 1993).
PNG
media_image1.png
18
19
media_image1.png
Greyscale
34. Applicant’s remarks and arguments presented on pages 10-12 have been fully considered but they are moot in view of the new grounds of rejection presented in this office action.
35. On pages 12-15, the arguments of claims 9-10, 13, 16, 17 are directed to the similar argument of claim 1 which has been addressed above.
Conclusion
36. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SyLing Yen whose telephone number is 571-270-1306.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Sanjiv Shah can be reached at 571-272-4098. The fax and phone numbers for the organization where this application or proceeding is assigned is 571-273-8300.
Any inquiry of a general nature or relating to the status of this application or proceeding should be directed to the receptionist whose telephone number is 571-272-2100.
/SYLING YEN/Primary Examiner, Art Unit 2166
July 21, 2026