Prosecution Insights
Last updated: October 02, 2026
Application No. 18/612,079

METHOD AND SYSTEM FOR EVALUATION OF CODE GENERATION BY LARGE LANGUAGE MODEL

Final Rejection §103
Filed
Mar 21, 2024
Examiner
HURUY, FEVEN HABTEMARIAM
Art Unit
2191
Tech Center
2100 — Computer Architecture & Software
Assignee
JPMorgan Chase Bank, N.A.
OA Round
2 (Final)
83%
Grant Probability
Favorable
3-4
OA Rounds
1m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 83% — above average
83%
Career Allowance Rate
5 granted / 6 resolved
+28.3% vs TC avg
Strong +25% interview lift
Without
With
+25.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
12 currently pending
Career history
27
Total Applications
across all art units

Statute-Specific Performance

§101
17.0%
-23.0% vs TC avg
§103
55.1%
+15.1% vs TC avg
§102
4.2%
-35.8% vs TC avg
§112
22.0%
-18.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 6 resolved cases

Office Action

§103
DETAILED ACTION This Office action is in response to the amendment filed on May 19, 2026. Claims 1-20 are pending. Claims 1, 2, 8, 9, 10, 13, 16, 17, and 18 have been amended. Claims 19 and 20 have been added. The objections to the drawings are withdrawn in view of Applicant’s amendments to the drawings. The objections to the specification are withdrawn in view of Applicant’s amendments to the specification. The objection to Claim 13 is withdrawn in view of Applicant’s amendments to the claim. The nonstatutory obviousness-type double patenting rejection of Claims 1, 2, 9, 10, 17, and 18 are withdrawn in view of Applicant’s amendments to the claims. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment Claim Objections Claim 17 is objected to because of the following informalities: Claim 17, line 14, recites “the quality of the first set of executable code.” It should read -- the quality of the second set of executable code --. Appropriate correction is required. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 2, 4, 6, 7, 9, 10, 12, 14, 15, 17, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over US 2024/0028312 (hereinafter “Gillman”) in view of US 2025/0258723 (hereinafter “Liu”), “Can Large Language Models Identify And Reason About Security Vulnerabilities? Not Yet” (hereinafter “Ullah”), and US 2021/0073676 (hereinafter “Moriya”). As per Claim 1, Gillman discloses: A method for evaluating code quality, the method being implemented by at least one processor (paragraph [0096], “The systems and methods described above may be implemented as a method, apparatus, or article of manufacture […] The techniques described above may be implemented in one or more computer programs executing on a programmable computer including a processor (emphasis added).”), the method comprising: receiving, by the at least one processor, a first set of instructions for performing a first task and generating a first output (paragraph [0082], “[…] the method 600 includes receiving, by a machine learning engine [processor], a user-specified data set and a natural language description of a user-requested data transformation task [first set of instructions] for execution with a subset of the user-specified data set (602) (emphasis added).”; paragraph [0081], “The method 600 includes executing, by the machine learning engine [processor], the at least one candidate executable computer code to generate a transformation result [first output] (608) (emphasis added).”; paragraph [0075], “By way of example, the system 100 may include functionality for receiving a natural language description of a data manipulation for data in a spreadsheet […] generate executable code using the natural language description and automatically (e.g., without human intervention) execute the code to complete the described data manipulation (emphasis added).”; paragraph [0017], “The machine learning engine 103 may be provided as a hardware component.”); providing, by the at least one processor as an input to a first large language model (LLM), […] the first set of instructions, together with a submission of a request to the first LLM to […] generate a first set of executable code based on the first set of instructions (paragraph [0081], “The method 600 includes directing, by the machine learning engine [processor], a large language model to generate at least one candidate executable computer code for performing the user-requested data transformation task (604) (emphasis added).”; paragraph [0078], “The method may include generating, by the large language model, executable computer code in a computer programming language specified in the natural language description of the user-requested data transformation task (emphasis added).”) [Examiner’s Remarks: Note that Gillman discloses the machine learning engine directing a large language model to generate executable computer code for performing the user-requested task. One of ordinary skill in the art would readily comprehend that in order for the large language model to generate the executable code that performs the user-requested task, the machine learning engine (processor) had to have provided the large language model with the user-requested task (first set of instructions) as an input together with a request to generate executable code based on the task (first set of instructions) while directing the large language model.]; receiving, by the at least one processor from the first LLM, […] the first set of executable code (paragraph [0080], “[…] the functionality of a large learning model trained to generate executable code for use in performing additional tasks on the output of the generated machine learning models and with the functionality of the machine learning engine [processor] to evaluate and validate the generated code and then execute the generated code (emphasis added).”) [Examiner’s Remarks: Note that Gillman discloses the large language model generating executable code and the machine learning engine validating/executing the generated code. One of ordinary skill in the art would readily comprehend that the machine learning engine (processor) must have received the generated executable code from the large language model in order to validate and execute the generated code.]; executing, by the at least one processor, the first set of executable code in order to perform the first task and generate the first output (paragraph [0081], “The method 600 includes directing, by the machine learning engine, a large language model to generate at least one candidate executable computer code for performing the user-requested data transformation task (604) […] The method 600 includes executing, by the machine learning engine [processor], the at least one candidate executable computer code to generate a transformation result (608) (emphasis added).”); and evaluating, by the at least one processor, a quality of the first set of executable code (paragraph [0081], “The method 600 includes performing, by the machine learning engine [processor], at least one validation [quality] check on the at least one candidate executable computer code (606) (emphasis added).”). Gillman discloses “providing, by the at least one processor as an input to a first large language model (LLM), […] the first set of instructions, together with a submission of a request to the first LLM to […] generate a first set of executable code based on the first set of instructions” and “receiving, by the at least one processor from the first LLM, […] the first set of executable code” but does not explicitly disclose: providing, by the at least one processor as an input to a first large language model (LLM), a list of available application programming interfaces (APIs) and the first set of instructions, together with a submission of a request to the first LLM to select one API and to generate a first set of executable code based on the first set of instructions; receiving, by the at least one processor from the first LLM, a selection of the one API and the first set of executable code. However, Liu discloses: providing, as an input to a first large language model (LLM), a list of available application programming interfaces (APIs), together with a submission of a request to the first LLM to select one API (paragraph [0006], “The first textual prompt for synthesizing user instructions can include, for instance, the list of APIs, the API documents for the list of APIs (or content of the API documents), and a request (in natural language) to generate a user instruction for performing a task using [selecting] a single tool/API from the list of APIs. The request, for instance, can be “generate a user instruction for performing a task using a single API from the provided APIs” (emphasis added).”; paragraph [0011], “During each of the one or more iterations, the first textual prompt for synthesizing user instructions can be processed as input, using the first trained LLM, to generate a corresponding synthetic natural language user instruction that describes a corresponding single-tool (e.g., single-API) task to be performed (emphasis added).”; paragraph [0087], “[…] processing a first textual prompt as input, using the first LLM, to generate a first synthetic natural language user instruction that describes to perform a first task. The first synthetic natural language user instruction may or may not identify a particular API which is the only API to be used to perform the first task (emphasis added).”); receiving a selection of the one API (paragraph [0020], “Optionally, in some implementations, the m.sup.th synthetic natural language user instruction (A.sub.m) that describes the m.sup.th single-API task to be performed, can be processed iteratively using the second trained LLM, to generate the list of execution steps (emphasis added).”; paragraph [0068], “[…] in some cases, a synthetic natural language user instruction that describes a corresponding single-API task to be performed can explicitly identify the single API via which the single-API task is to be performed (emphasis added).”); a second dataset that relates to the APIs (paragraph [0004], “In various implementations, a list of application programming interfaces (APIs) are identified or selected, and for each API in this list of APIs, a corresponding API document that describes a respective API is identified and retrieved (emphasis added).”). Gillman is within the same field of endeavor as the claimed invention regarding the utilization of LLM for code generation. Liu is also within the same field of endeavor as the claimed invention regarding the utilization of LLM for API selection. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teaching of Liu into the teaching of Gillman to include “providing, by the at least one processor as an input to a first large language model (LLM), a list of available application programming interfaces (APIs) and the first set of instructions, together with a submission of a request to the first LLM to select one API and to generate a first set of executable code based on the first set of instructions; receiving, by the at least one processor from the first LLM, a selection of the one API and the first set of executable code; a second dataset that relates to the APIs.” The modification would be obvious because one of ordinary skill in the art would be motivated to have a LLM select an API from a list of APIs to ensure the correct one is used for performing the task and to use in generating user instructions for performing API-based tasks (synthetic training data) that can be used in ensuring “the diversity of the training data to train an LLM in possessing or improving its capability in handling tasks that utilize external tools or APIs” (Liu, paragraphs [0003, 0077, & 0087]). The combination of Gillman and Liu discloses “evaluating of the quality of the first set of executable code” and “a second dataset that relates to the APIs,” but does not explicitly disclose: wherein the evaluating of the quality of the first set of executable code comprises evaluating at least a robustness with respect to differing difficulty levels of instruction, wherein the evaluating of the robustness includes generating, by the at least one processor, an evaluation dataset along a plurality of dimensions to encompass a diverse set of evaluation tasks and corresponding ground truths, wherein the evaluation dataset includes different input and output combinations, and wherein the evaluation dataset is divided into three subsets including a first dataset that relates to the first set of instructions and variations thereof, a second dataset that relates to the APIs and a third dataset that relates to evaluation records. However, Ullah discloses: wherein the [evaluating] comprises evaluating at least a robustness with respect to differing difficulty levels of instruction (Section 4.6. Code Difficulty Levels page 11, “In this section, we investigate the capabilities of LLMs to handle different complexities of code. Similar to the previous sections, we find the best performing prompts for each difficulty level using Scorediff, with equal weight to all factors, from four prompting categories (emphasis added).”; Section 3.4. Datasets page 4, “Moreover, we design our code scenarios with three difficulty levels, (1) easy, (2) medium, (3) hard […] The difficulty levels assess how LLMs interact with code of increasing complexity (emphasis added).”; Section 1. Introduction page 1, “Our framework tests the capabilities of a given LLM as a security assistant across eight distinct dimensions: […] (6) assessment of various code difficulty levels, (7) robustness to code augmentations […] (emphasis added).”), wherein the evaluating of the robustness includes generating, by the at least one processor, an evaluation dataset along a plurality of dimensions to encompass a diverse set of evaluation tasks and corresponding ground truths (Section 3.4 Datasets page 4, “We design 228 code scenarios (48 hand-crafted, 30 real world, and 150 with code augmentations) to test various aspects of the capabilities of LLMs to detect software vulnerabilities in code [...] We curate a dataset of 48 hand-crafted code scenarios, containing vulnerable and patched pairs from 8 most critical and diverse Common Weakness Enumerations (CWEs) from the MITRE Top 25 Most Dangerous Software Weaknesses for the year 2023 [42], as shown in Table 4 (emphasis added).”; Section 3.5. Ground-Truth Reasoning pages 5 & 6, “In addition to ground truth labels indicating if a code snippet contains a vulnerability, we also need explanations for these vulnerabilities, as we aim to evaluate the reasoning capabilities of LLMs and assess whether they can justify their decisions (emphasis added).”; abstract, “We construct a series of 228 code scenarios and analyze eight of the most capable LLMs across eight different investigative dimensions in an automated framework (emphasis added).”), wherein the evaluation dataset includes different input and output combinations (Section 3.4 Datasets pages 4-5, “Easy scenarios consist of simple programs containing only one function and less than 30 lines of code. Medium level scenarios increase the complexity by making the program longer, using different library functions, and adding more than one user input.”; Section Appendix A. Examples of Code Difficulty Levels page 16, “Easy: CWE-22 ‘1v’ (see Figure 9a) takes a file path as an input, which is then concatenated with an absolute directory path, and then the file is read and displayed on the console [...] Medium: CWE-22 ‘2v’ (see Figure 9b) takes four inputs: file path, flag, data, and directory path (set using an environment variable). Based on the flag, data and file are processed (emphasis added).”), and wherein the evaluation dataset is divided into three subsets including a first dataset that relates to the first set of instructions and variations thereof (Section 3.4 Datasets page 4, “We design 228 code scenarios (48 hand-crafted, 30 real world, and 150 with code augmentations) to test various aspects of the capabilities of LLMs to detect software vulnerabilities in code (emphasis added).”; Section 3.4. Datasets page 4, “Moreover, we design our code scenarios with three difficulty levels [variations], (1) easy, (2) medium, (3) hard (emphasis added).”). Ullah is within the same field of endeavor as the claimed invention regarding LLMs, assigning difficulty levels to instructions, and evaluating robustness. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teaching of Ullah into the combined teachings of Gillman and Liu to include “wherein the evaluating of the quality of the first set of executable code comprises evaluating at least a robustness with respect to differing difficulty levels of instruction, wherein the evaluating of the robustness includes generating, by the at least one processor, an evaluation dataset along a plurality of dimensions to encompass a diverse set of evaluation tasks and corresponding ground truths, wherein the evaluation dataset includes different input and output combinations, and wherein the evaluation dataset is divided into three subsets including a first dataset that relates to the first set of instructions and variations thereof, a second dataset that relates to the APIs.” The modification would be obvious because one of ordinary skill in the art would be motivated to determine a difficulty level of instructions and assess LLMs ability to execute it using an automated framework, such as in the case of software vulnerability detection, in order to effectively identify the best performing LLMs based on the difficulty level and utilizing the framework as a “useful tool” in evaluating the progress of LLMs (Ullah, Section 4.6. Code Difficulty Levels: Observations page 11 & Section 6. Conclusion page 14). The combination of Gillman, Liu, and Ullah does not explicitly disclose: a third dataset that relates to evaluation records. However, Moriya discloses: a third dataset that relates to evaluation records (Figure 7; paragraph [0147], “The evaluation result table T3400 has a record for each evaluation result.”). Moriya is within the same field of endeavor as the claimed invention regarding the utilization of a database of evaluation records. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teaching of Moriya into the combined teachings of Gillman, Liu, and Ullah to include “a third dataset that relates to evaluation records.” The modification would be obvious because one of ordinary skill in the art would be motivated to utilize an evaluation result table in order to achieve model improvement by learning from evaluation results and effectively provide information regarding model performance (Moriya, paragraphs [0004 & 0015]). As per Claim 2, the rejection of Claim 1 is incorporated; and Gillman discloses “wherein the evaluating of the quality of the first set of executable code further comprises evaluating an accuracy of the first set of executable code (paragraph [0084], “The method 600 includes performing, by the machine learning engine, at least one validation [quality] check on the at least one candidate executable computer code (606) […] (the validation step is important because not all generated programs will be correct or viable; many will not run) [evaluating an accuracy of the first set of executable code] […] (emphasis added).”)” and “the first set of executable code (paragraph [0081], “The method 600 includes directing, by the machine learning engine, a large language model to generate at least one candidate executable computer code for performing the user-requested data transformation task (604) (emphasis added).),” but the combination of Gillman, Liu, and Moriya does not explicitly disclose: wherein the evaluating of the quality of the first set of executable code further comprises evaluating an accuracy of the first set of executable code and a consistency of the first set of executable code across a plurality of runs. However, Ullah discloses: evaluating a consistency of the [LLM] across a plurality of runs (Section 4.1. Evaluation for Deterministic Responses page 6, “To perform a rigorous comparison between LLMs and assess their capabilities, it is of critical importance that their responses are consistent, meaning that running the same test multiple times under identical parameters should provide the same final verdict (emphasis added).”; Section 4.1. Evaluation for Deterministic Responses page 7, “We run each experiment ten times, and record how many times the model provides the same answer (emphasis added).”). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teaching of Ullah into the combined teachings of Gillman, Liu, and Moriya to include “wherein the evaluating of the quality of the first set of executable code further comprises evaluating an accuracy of the first set of executable code and a consistency of the first set of executable code across a plurality of runs.” The modification would be obvious because one of ordinary skill in the art would be motivated to test results of an LLM across multiple runs and see if it provides different answers for the same input in order to improve reliability and consistency of LLM responses by determining which “parameters deliver the most consistent results” (Ullah, Section 4.1. Evaluation for Deterministic Responses pages 6). As per Claim 4, the rejection of Claim 2 is incorporated; and Gillman discloses “the first set of executable code (paragraph [0081], “The method 600 includes directing, by the machine learning engine, a large language model to generate at least one candidate executable computer code for performing the user-requested data transformation task (604) (emphasis added).)” and “the first set of instructions (paragraph [0082], “[…] the method 600 includes receiving, by a machine learning engine [processor], a user-specified data set and a natural language description of a user-requested data transformation task [first set of instructions] for execution with a subset of the user-specified data set (602) (emphasis added).”),” but the combination of Gillman and Moriya does not explicitly disclose: wherein the evaluating of the robustness of the first set of executable code comprises: determining a difficulty level of the first set of instructions; and assessing the selection of the one API and an ability to execute the first set of instructions based on the determined difficulty level. However, Liu discloses: the selection of the one API (paragraph [0006], “The first textual prompt for synthesizing user instructions can include, for instance, the list of APIs, the API documents for the list of APIs (or content of the API documents), and a request […] The request, for instance, can be “generate a user instruction for performing a task using [selecting] a single API from the provided APIs” (emphasis added).”; paragraph [0087], “[…] processing a first textual prompt as input, using the first LLM, to generate a first synthetic natural language user instruction that describes to perform a first task. The first synthetic natural language user instruction may or may not identify a particular API which is the only API to be used to perform the first task (emphasis added).”; paragraph [0020], “Optionally, in some implementations, the m.sup.th synthetic natural language user instruction (A.sub.m) that describes the m.sup.th single-API task to be performed, can be processed iteratively using the second trained LLM, to generate the list of execution steps (emphasis added).”). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teaching of Liu into the combined teachings of Gillman and Moriya to include “the selection of the one API.” The modification would be obvious because one of ordinary skill in the art would be motivated to have a LLM select one API from a list of APIs to ensure the correct one is used and processed for performing the task and to use in training data “to train an LLM in possessing or improving its capability in handling tasks that utilize external tools or APIs” (Liu, paragraphs [0003, 0077, & 0087]). However, Ullah discloses: wherein the evaluating of the robustness of the [LLM] comprises: determining a difficulty level of the first set of instructions (Section 4.6. Code Difficulty Levels page 11, “In this section, we investigate the capabilities of LLMs to handle different complexities of code. Similar to the previous sections, we find the best performing prompts for each difficulty level using Scorediff, with equal weight to all factors, from four prompting categories (emphasis added).”; Section 3.4. Datasets page 4, “Moreover, we design our code scenarios with three difficulty levels, (1) easy, (2) medium, (3) hard […] The difficulty levels assess how LLMs interact with code of increasing complexity (emphasis added).”); and assessing an ability to execute the first set of instructions based on the determined difficulty level (Section 4.6. Code Difficulty Levels page 11, “In this section, we investigate the capabilities of LLMs to handle different complexities of code. Similar to the previous sections, we find the best performing prompts for each difficulty level using Scorediff, with equal weight to all factors, from four prompting categories (emphasis added).”; Section 1. Introduction page 1, “Our framework tests the capabilities of a given LLM as a security assistant across eight distinct dimensions: […] (6) assessment of various code difficulty levels, (7) robustness to code augmentations […] (emphasis added).”; Section 3.4. Datasets page 4, “We design 228 code scenarios (48 hand-crafted, 30 real world, and 150 with code augmentations) to test various aspects of the capabilities of LLMs to detect software vulnerabilities in code. We use these scenarios to craft prompts by including code, examples, definitions, and step-by-step reasoning as shown in Table 3.”) [Examiner’s Remarks: Note that Ullah discloses testing the capabilities of a LLM to detect software vulnerabilities in code and handle different complexities of code, finding the best performing prompt for each difficulty level, and prompts (instructions) that include code. One of ordinary skill in the art would readily comprehend that testing the capabilities of LLMs to detect software vulnerabilities in code and handle different complexities of code is assessing its ability to execute the instructions in the prompt (finding the software vulnerability) based on the determined difficulty level.]. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teaching of Ullah into the combined teachings of Gillman, Liu, and Moriya to include “wherein the evaluating of the robustness of the first set of executable code comprises: determining a difficulty level of the first set of instructions; and assessing the selection of the one API and an ability to execute the first set of instructions based on the determined difficulty level.” The modification would be obvious because one of ordinary skill in the art would be motivated to determine a difficulty level of instructions and assess LLMs ability to execute it using an automated framework, such as in the case of software vulnerability detection, in order to effectively identify the best performing LLMs based on the difficulty level and utilizing the framework as a “useful tool […] to evaluate the progress of future LLM versions in vulnerability detection” (Ullah, Section 4.6. Code Difficulty Levels: Observations page 11 & Section 6. Conclusion page 14). As per Claim 6, the rejection of Claim 2 is incorporated; and Gillman discloses “the first set of executable code (paragraph [0081], “The method 600 includes directing, by the machine learning engine, a large language model to generate at least one candidate executable computer code for performing the user-requested data transformation task (604) (emphasis added).)” and “executing of the first set of executable code (paragraph [0081], “The method 600 includes executing, by the machine learning engine, the at least one candidate executable computer code to generate a transformation result (608) (emphasis added).”),” but the combination of Gillman, Liu, and Moriya does not explicitly disclose: wherein the evaluating of the consistency of the first set of executable code comprises: testing results of the executing of the first set of executable code across multiple runs; and determining whether the results provide different answers for a same input. However, Ullah discloses: wherein the evaluating of the consistency of the [LLM] comprises: testing results of the [LLM] across multiple runs (Section 4.1. Evaluation for Deterministic Responses page 6, “To perform a rigorous comparison between LLMs and assess their capabilities, it is of critical importance that their responses are consistent, meaning that running the same test multiple times under identical parameters should provide the same final verdict (emphasis added).”); and determining whether the results provide different answers for a same input (Section 4.1. Evaluation for Deterministic Responses page 6, “To perform a rigorous comparison between LLMs and assess their capabilities, it is of critical importance that their responses are consistent, meaning that running the same test multiple times under identical parameters should provide the same final verdict (emphasis added).”; Section 4.1. Evaluation for Deterministic Responses page 7, “We run each experiment ten times, and record how many times the model provides the same answer (emphasis added).”). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teaching of Ullah into the combined teachings of Gillman, Liu, and Moriya to include “wherein the evaluating of the consistency of the first set of executable code comprises: testing results of the executing of the first set of executable code across multiple runs; and determining whether the results provide different answers for a same input.” The modification would be obvious because one of ordinary skill in the art would be motivated to test results of an LLM across multiple runs and see if it provides different answers for the same input in order to improve reliability and consistency of LLM responses by determining which “parameters deliver the most consistent results” (Ullah, Section 4.1. Evaluation for Deterministic Responses pages 6). As per Claim 7, the rejection of Claim 6 is incorporated; and the combination of Gillman, Liu, and Moriya does not explicitly disclose: wherein the testing of the results is performed for at least three runs and for at most ten runs. However, Ullah discloses: wherein the testing of the results is performed for at least three runs and for at most ten runs (Section 4.1. Evaluation for Deterministic Responses page 6, “To perform a rigorous comparison between LLMs and assess their capabilities, it is of critical importance that their responses are consistent, meaning that running the same test multiple times under identical parameters should provide the same final verdict (emphasis added).”; Section 4.1. Evaluation for Deterministic Responses page 7, “We run each experiment ten times, and record how many times the model provides the same answer (emphasis added).”). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teaching of Ullah into the combined teachings of Gillman, Liu, and Moriya to include “wherein the testing of the results is performed for at least three runs and for at most ten runs.” The modification would be obvious because one of ordinary skill in the art would be motivated to test results of an LLM across multiple runs and see if it provides different answers for the same input in order to improve reliability and consistency of LLM responses by determining which “parameters deliver the most consistent results” (Ullah, Section 4.1. Evaluation for Deterministic Responses pages 6). As per Claim 9, Gillman discloses: A computing apparatus (paragraph [0096], “The systems and methods described above may be implemented as a method, apparatus, or article of manufacture […] (emphasis added).”) for evaluating code quality, the computing apparatus comprising: a processor (paragraph [0096], “The techniques described above may be implemented in one or more computer programs executing on a programmable computer including a processor […].”); a memory (paragraph [0108], “In the embodiment shown in FIG. 4B, the processor 421 communicates with main memory 422 via a system bus 450.”); and a communication interface coupled to each of the processor and the memory (paragraph [0108], “In the embodiment shown in FIG. 4B, the processor 421 communicates with main memory 422 via a system bus 450.”), wherein the processor is configured to: […]. Claim 9 is an apparatus claim corresponding to method Claim 1 and the remainder of Claim 9 is rejected for the same reasons as given in the rejection of Claim 1. Claim 10 is an apparatus claim corresponding to method Claim 2 and is rejected for the same reasons as given in the rejection of that claim. Claim 12 is an apparatus claim corresponding to method Claim 4 and is rejected for the same reasons as given in the rejection of that claim. Claim 14 is an apparatus claim corresponding to method Claim 6 and is rejected for the same reasons as given in the rejection of that claim. Claim 15 is an apparatus claim corresponding to method Claim 7 and is rejected for the same reasons as given in the rejection of that claim. As per Claim 17, Gillman discloses: A non-transitory computer readable storage medium storing instructions (paragraph [0098], “Storage devices suitable for tangibly embodying computer program instructions include, for example, all forms of computer-readable devices, firmware, programmable logic, hardware (e.g., integrated circuit chip; electronic devices; a computer-readable non-volatile storage unit; non-volatile memory, such as semiconductor memory devices, including EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks; and CD-ROMs).”) for evaluating code quality, the storage medium comprising a first set of executable code which, when executed by a processor (paragraph [0098], “Method steps may be performed by a computer processor executing a program tangibly embodied on a computer-readable medium to perform functions of the methods and systems described herein by operating on input and generating output.”), causes the processor to: […]. Claim 17 is a non-transitory computer readable storage medium claim corresponding to method Claim 1 and the remainder of Claim 17 is rejected for the same reasons as given in the rejection of Claim 1. Claim 18 is a non-transitory computer readable storage medium claim corresponding to method Claim 2 and is rejected for the same reasons as given in the rejection of that claim. Claims 3 and 11 are rejected under 35 U.S.C. 103 as being unpatentable over Gillman in view of Liu, Ullah, and Moriya as applied to Claims 2 and 10 above, and further in view of US 2010/0229151 (hereinafter “Yuan”) and “De-Hallucinator: Iterative Grounding for LLM-Based Code Completion” (hereinafter “Eghbali”). As per Claim 3, the rejection of Claim 2 is incorporated; and Gillman discloses “the evaluating of the accuracy of the first set of executable code (paragraph [0084], “The method 600 includes performing, by the machine learning engine, at least one validation check on the at least one candidate executable computer code (606) […] (the validation step is important because not all generated programs will be correct or viable; many will not run) [evaluating an accuracy of the first set of executable code] (emphasis added).”)” and “the first set of executable code (paragraph [0081], “The method 600 includes directing, by the machine learning engine, a large language model to generate at least one candidate executable computer code for performing the user-requested data transformation task (604) (emphasis added).),” but the combination of Gillman and Liu does not explicitly disclose: checking whether the first set of executable code runs; checking whether the first set of executable code calls a correct API with correct parameters; and checking whether the first output matches with an expected output. However, Yuan discloses: checking whether the first set of executable code runs (paragraph [0031], “At step 110, a quality control check can be made to the implantation code generated at step 108 in order to verify that such code will run properly when executed […] A local controls engineer or other personnel or device can manually or automatically compare the implementation code to a standard, can run an off-line test, or perform whatever other steps are needed to properly verify the accuracy of the code generated at step 110 (emphasis added).”). Yuan is within the same field of endeavor as the claimed invention regarding the evaluation of code accuracy. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teaching of Yuan into the combined teachings of Gillman, Liu, Ullah, and Moriya to include “checking whether the first set of executable code runs.” The modification would be obvious because one of ordinary skill in the art would be motivated to verify code accuracy and check whether generated code runs in order to ensure the code runs and functions properly when executed (Yuan, paragraph [0031]). However, Eghbali discloses: checking whether the first set of executable code calls a correct API with correct parameters (Section 5.3 RQ2: Correct Retrieval of API References page 15, Figure 8; Section 5.1 Experimental Setup page 12, “Exact API match. Since the goal of De-Hallucinator is to predict better API usages, we measure how many of all desired API usages are predicted exactly as in the ground truth. To identify the API usages in the lines of code to complete, we extract function calls, including the access path to the function, and the parameters. For example, given a line of code docs = ds.find_by_keyword(keyword) the corresponding API usage is ds.find_by_keyword(keyword). The exact API match then is the percentage of exact matches between the prediction and the ground truth API usages (emphasis added).”); and checking whether the first output matches with an expected output (Section 5.1 Experimental Setup page 12, “Exact API match. Since the goal of De-Hallucinator is to predict better API usages, we measure how many of all desired API usages are predicted exactly as in the ground truth [expected output]. To identify the API usages in the lines of code to complete, we extract function calls, including the access path to the function, and the parameters. For example, given a line of code docs = ds.find_by_keyword(keyword) the corresponding API usage is ds.find_by_keyword(keyword). The exact API match then is the percentage of exact matches between the prediction [first output] and the ground truth API usages [expected output] (emphasis added).”). Eghbali is within the same field of endeavor as the claimed invention regarding checking the correctness of LLM generated API calls. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teaching of Eghbali into the combined teachings of Gillman, Liu, Ullah, Moriya, and Yuan to include “checking whether the first set of executable code calls a correct API with correct parameters; and checking whether the first output matches with an expected output.” The modification would be obvious because one of ordinary skill in the art would be motivated to use a De-Hallucinator that improves input given to a LLM and check if the generated code makes correct API calls and matches expected output in order to “improve the quality of code completions over the state-of-the-art baseline” (Eghbali, Section 7 Related Work page 19 & Section 8 Conclusion page 20). Claim 11 is an apparatus claim corresponding to method Claim 3 and is rejected for the same reasons as given in the rejection of that claim. Claims 5 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Gillman in view of Liu, Ullah, and Moriya as applied to Claims 4 and 12 above, and further in view of US 2024/0104308 (hereinafter “Francis”). As per Claim 5, the rejection of Claim 4 is incorporated; and the combination of Gillman, Liu, Ullah, and Moriya does not explicitly disclose: determining a degree of implicitness of information included in the first set of instructions with respect to the first task. However, Francis discloses: determining a degree of implicitness of information included in the first set of instructions with respect to the first task (paragraph [0028], “In some embodiments, the systems and methods described herein may be configured to, when given a high-level instruction or question, use the agent to the leverage commonsense and spatial knowledge to decompose the implicit navigation and/or manipulation task into tractable subtasks. For example, if the agent is initialized in the living room of a home then asked the question, “What color is the car?”, the agent may generate a plan for, first, searching for the car in the garage or outside the house on the driveway (emphasis added).”) [Examiner’s Remarks: Note that Francis discloses using an agent to leverage commonsense and spatial knowledge and decompose the implicit task into tractable subtasks when given a high-level instruction. One of ordinary skill in the art would readily comprehend that decomposing the implicit task within a high-level instruction into subtasks includes determining a degree of implicitness of the information within the instruction with respect to the task in order to leverage commonsense and spatial knowledge to decompose the task.]. Francis is within the same field of endeavor as the claimed invention regarding the determination of implicitness within task instructions. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teaching of Francis into the combined teachings of Gillman, Liu, Ullah, and Moriya to include “determining a degree of implicitness of information included in the first set of instructions with respect to the first task.” The modification would be obvious because one of ordinary skill in the art would be motivated to determine a degree of implicitness within a high-level instruction and use an agent to decompose the implicit task into tractable subtasks by leveraging commonsense and spatial knowledge (domain knowledge) because agents that are “encouraged to learn reasoning strategies, on top of this domain knowledge, perform better than those that simply perform statistical pattern-matching” (Francis, paragraph [0021 & 0028]). Claim 13 is an apparatus claim corresponding to method Claim 5 and is rejected for the same reasons as given in the rejection of that claim. Claims 8 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Gillman in view of Liu, Ullah, and Moriya as applied to Claims 2 and 10 above, and further in view of Eghbali. As per Claim 8, the rejection of Claim 2 is incorporated; and Gillman discloses “wherein the evaluating of the quality of the first set of executable code is performed by using [statistical analysis] (paragraph [0084], “The method 600 includes performing, by the machine learning engine, at least one validation [quality] check on the at least one candidate executable computer code (606) […] (the validation step is important because not all generated programs will be correct or viable; many will not run); performing security and validation checks such as: evaluating the code using static analysis […] (emphasis added).”),” but the combination of Gillman, Liu, Ullah, and Moriya does not explicitly disclose: wherein the evaluating of the quality of the first set of executable code is performed by using the evaluation dataset that is API-based. However, Eghbali discloses: the evaluation dataset that is API-based (Section 5.1 Experimental Setup pages 11-12, “We construct a dataset of API-related code completion tasks by removing API usages from the benchmark projects and by considering the removed code as the ground truth to be predicted by a model […] Overall, the evaluation dataset consists of 11 projects × 10 × 4 models = 440 code completion tasks (emphasis added).”). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teaching of Eghbali into the combined teachings of Gillman, Liu, Ullah, and Moriya to include “wherein the evaluating of the quality of the first set of executable code is performed by using the evaluation dataset that is API-based.” The modification would be obvious because one of ordinary skill in the art would be motivated to utilize an evaluation dataset that is API-based in order to effectively evaluate API-based code completions of models based on the tasks within the dataset (Eghbali, Section 5.1 Experimental Setup pages 11-12). Moreover, one of ordinary skill in the art would be motivated to utilize both an API-based evaluation dataset of code completion tasks and a De-Hallucinator that improves input given to a LLM to check if model generated code makes correct API calls in order to “improve the quality of code completions over the state-of-the-art baseline” (Eghbali, Section 7 Related Work page 19 & Section 8 Conclusion page 20). Claim 16 is an apparatus claim corresponding to method Claim 8 and is rejected for the same reasons as given in the rejection of that claim. Claim 19 is rejected under 35 U.S.C. 103 as being unpatentable over Gillman in view of Liu, Ullah, and Moriya as applied to Claim 1 above, and further in view of US 2025/0133102 (hereinafter “Mitev”). As per Claim 19, the rejection of Claim 1 is incorporated; and the combination of Gillman, Liu, and Moriya does not explicitly disclose: wherein each task in the evaluation dataset is designed with a tiered structure, offering an original version with differing difficulty variations. However, Ullah discloses: wherein each task in the evaluation dataset is designed with a tiered structure (Section 3.4. Datasets page 4, “Moreover, we design our code scenarios with three difficulty levels, (1) easy, (2) medium, (3) hard […].”). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teaching of Ullah into the combined teachings of Gillman, Liu, and Moriya to include “wherein each task in the evaluation dataset is designed with a tiered structure.” The modification would be obvious because one of ordinary skill in the art would be motivated to determine a difficulty level of instructions and assess LLMs ability to execute it using an automated framework, such as in the case of software vulnerability detection, in order to effectively identify the best performing LLMs based on the difficulty level (tiered structure) and utilizing the framework as a “useful tool” in evaluating the progress of LLMs (Ullah, Section 4.6. Code Difficulty Levels: Observations page 11 & Section 6. Conclusion page 14). The combination of Gillman, Liu, Ullah, and Moriya does not explicitly disclose: wherein each task in the evaluation dataset is designed with a tiered structure, offering an original version with differing difficulty variations. However, Mitev discloses: offering an original version with differing difficulty variations (paragraph [0088] “In some embodiments, message generator 552 is configured to modify the prompt based on a difficulty level [...] Based on the level of difficulty, a prompt can be modified to introduce mistakes in the text or formatting. For example, easier levels can correspond to more deliberate or obvious mistakes in the message body and formatting.”; paragraph [0079], “The database 534 can further comprise a plurality of difficulty levels 548 and associated prompt modifications and social engineering techniques, which can be stored as a table.”). Mitev is within the same field of endeavor as the claimed invention regarding the utilization of tasks/prompts with difficulty variations. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teaching of Mitev into the combined teachings of Gillman, Liu, Ullah, and Moriya to include “wherein each task in the evaluation dataset is designed with a tiered structure, offering an original version with differing difficulty variations.” The modification would be obvious because one of ordinary skill in the art would be motivated to utilize a system that automates modifying prompts with differing difficulty levels in order to provide increased flexibility and adaptability regarding prompt creation (Mitev, paragraphs [0006 & 0007]). Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Gillman in view of Liu, Ullah, Moriya, and Mitev as applied to Claim 19 above, and further in view of Eghbali. As per Claim 20, the rejection of Claim 19 is incorporated; and the combination of Gillman, Liu, Ullah, Moriya, and Mitev does not explicitly disclose: wherein the evaluating of the robustness includes introducing variations in availability of the APIs for the each task. However, Eghbali discloses: wherein the evaluating of the robustness includes introducing variations in availability of the APIs for the each task (page 3, “At first, De-Hallucinator queries the LLM with the conventional prompt that contains only the preceding code. Because the preceding code alone often is insufficient to obtain a correct completion, the approach then augments the prompt with API references that are most similar to the preceding code, which makes the second prompt type. However, also this prompt may fail to find a suitable completion, because the preceding code might not be similar to the desired API, or there are other APIs more similar to the preceding code than the correct one [...] Hence, to construct the third type of prompt, De-Hallucinator leverages these hints about what code the model intends to predict to retrieve suitable project-specific APIs, which are then added to the third type of prompt (emphasis added).”; Section 5.2 RQ1 page 13, “Figure 7 shows the initial completion of CodeGen, with only the preceding code as prompt. In this example the model uses the assert statement instead of the custom assert_equivalent function, which is defined in the project. After augmenting the prompt with the API reference of this function, the same model correctly predicts the call to the existing API, as shown in Fig. 7.”). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teaching of Eghbali into the combined teachings of Gillman, Liu, Ullah, Moriya, and Mitev to include “wherein the evaluating of the robustness includes introducing variations in availability of the APIs for the each task.” The modification would be obvious because one of ordinary skill in the art would be motivated to introduce variations in the prompt regarding APIs in order to effectively test the performance and robustness of models on prompts with varying information and help “improve the quality of code completions over the state-of-the-art baseline” (Eghbali, page 3 & Section 8 Conclusion page 20). Response to Arguments In the Remarks, Applicant’s arguments under “Rejection under 35 USC 103”, Applicant’s arguments with respect to Claims 1, 9, and 17 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Feven H. Huruy whose telephone number is (571) 272-3826. The examiner can normally be reached Mon-Fri. 7:30am-3:30pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Wei Mui can be reached at (571) 272-3708. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /F.H.H./Examiner, Art Unit 2191 /WEI Y MUI/Supervisory Patent Examiner, Art Unit 2191
Read full office action

Prosecution Timeline

Mar 21, 2024
Application Filed
Mar 23, 2026
Non-Final Rejection mailed — §103
May 19, 2026
Response Filed
Aug 13, 2026
Final Rejection mailed — §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
83%
Grant Probability
99%
With Interview (+25.0%)
2y 7m (~1m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 6 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month