DETAILED ACTION
This action is responsive to Remarks and Claim Amendments filed on August 18, 2026.
Claims 1, 8 and 15 have been amended. Claims 7 and 14 have been canceled. Claims 21-22 have been newly added.
Claims 1-6, 8-13 and 15-22 are pending and are presented to examination.
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Examiner Notes
Examiner cites particular columns, paragraphs, figures and line numbers in the references as applied to the claims below for the convenience of the applicant. Although the specified citations are representative of the teachings in the art and are applied to the specific limitations within the individual claim, other passages and figures may apply as well. It is respectfully requested that, in preparing responses, the applicant fully consider the references in their entirety as potentially teaching all or part of the claimed invention, as well as the context of the passage as taught by the prior art or disclosed by the examiner.
Response to Amendments
The rejection of claims 7 and 14 under 35 U.S.C. 112(d) as being of improper dependent form is withdrawn. The claim listing filed August 18, 2026 presents claims 7 and 14 with the status identifier “(Cancelled),” and the claims are therefore no longer pending.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1–3, 6, 8–10, 13, 15–17 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Carrara et al. (US Pub. No. 2025/0298585, hereinafter “Carrara” – previously presented) in view of Dong Huang et al. (“AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation”, hereinafter “Huang” – previously presented) in view of Fanning et al. (US Pub. No. 2012/0311535, hereinafter “Fanning” – previously presented) in view of Schaefer et al. (US Pub. No. 2025/0245122, hereinafter “Schaefer” – previously presented) and further in view of Clement et al. (US Pub. No. 2023/0128008, hereinafter Clement).
With respect to claim 1 (Currently Amended), Carrara teaches A system comprising: at least one hardware processor (Carrara discloses that the industrial IDE system 202 includes one or more processors 218, [0060], figure 2.) and
a computer-readable medium storing instructions that, when executed by the at least one hardware processor, cause the system to perform operations comprising (Carrara discloses a memory 220 that stores computer-executable components and instructions which are executed by the one or more processors 218 to carry out the operations of the IDE system, [0060], figure 2.)
receiving a request to generate computer code for insertion into source code of a software application, the request including a description of the computer code (Carrara discloses receiving design input comprising “natural language requests received via a chat interface” that request the generation of control code and that specify the functional requirements of the code to be generated, [0005], [0061], [0068]; Carrara further discloses that the generated control code is inserted at a location of the existing control code 702 specified by the user, [0103], [0105], figure 14.)
generating a prompt based on the description of the computer code and contextual information regarding the computer code (Carrara discloses a specialized prompt engineering layer and associated custom models 222 that generate prompts or meta-prompts based on the user’s natural language inputs together with domain-specific information contained in the custom models, [0058], [0063].)
sending the prompt to a first Large Language Model (LLM) to generate a plurality of suggested computer code blocks (Carrara discloses that the generative AI component 210 generates and submits the prompts or meta-prompts “to generative AI models such as large language models (LLMs)” to generate control code, [0058], [0063], [0069].)
Carrara does not expressly disclose that the plurality of suggested computer code blocks is received by a code evaluation service, or that the unit tests are generated by a second LLM separate from the first LLM in response to a query sent after the code blocks are received. However, in an analogous art, Huang teaches:
receiving, by a code evaluation service, from the first LLM, the plurality of suggested computer code blocks (Huang discloses a multi-agent code generation framework in which a programmer agent powered by a large language model generates the code, and in which “The code snippets and test cases are collected by the test executor agent (i.e., Agent#3) and executed in the local environment to obtain feedback”, Huang, Section 3.1; Huang further discloses that “Upon receiving code snippets and test cases generated by the programmer and test designer agent, the test executor agent validates these code snippets along with the test cases in a local environment”, Huang, Section 3.4. The test executor agent of Huang is therefore a code evaluation service that receives the generated code from the code-generating model.)
after receiving the plurality of suggested computer code blocks from the first LLM, sending, by the code evaluation service, a query to a second LLM separate from the first LLM, the query requesting generation of a plurality of unit tests (Huang discloses the recited sequence expressly: “The process begins by inputting tasks/code generation requirements into the code generation agent (i.e., Agent#1: the programmer agent). Subsequently, the test case generator (i.e., Agent#2: the test designer agent) is tasked with generating test cases, which are used to evaluate the correctness of the code snippets produced by the programmer agent”, Huang, Section 3.1. Huang further discloses that the test designer agent is a second, separate model powered by a large language model, stating that “AgentCoder requires separate agents for generating code and tests (i.e., the programmer agent and the test designer agent)” and that “Both agents are powered by LLMs”, Huang, Section 4.7, and that “The test designer agent is also powered by LLMs”, Huang, Section 3.3. Huang further discloses that the tasking of the test designer agent is carried out by a query in the form of a prompt requesting the generation of a plurality of unit tests, stating that “We carefully designed the prompts for the test designer agent to satisfy the following three expectations: 1) to generate basic test cases, 2) to cover edge test cases, and 3) to cover large-size inputs”, Huang, Section 3.3, the generated test cases being expressed as assertion statements, Huang, Section 3.3.)
receiving, by the code evaluation service, the plurality of unit tests from the second LLM (Huang discloses that the test executor agent receives the test cases generated by the test designer agent, stating that “Upon receiving code snippets and test cases generated by the programmer and test designer agent, the test executor agent validates these code snippets along with the test cases in a local environment”, Huang, Section 3.4.)
It would have been obvious to one of ordinary skill in the art at the time the invention was made before the effective filing date of the claimed invention to receive the plurality of suggested computer code blocks generated by Carrara’s generative AI model at a code evaluation service, and, after receiving those code blocks, to send a query from that service to a second large language model separate from the first large language model requesting the generation of a plurality of unit tests, and to receive those unit tests at the service, as taught by Huang, because Huang teaches that where one and the same model generates both the code and the tests that are used to check it, “functional errors may still go undetected as both the code and its test cases are generated by the same model”, and that a tester “being the same LLM as the coder, might erroneously assess erroneous code as correct” (Huang, Section 2.2), and that in a single-agent arrangement “the code generation and test generation processes are not independent” (Huang, Section 1). Huang expressly considers whether the two roles must be separated into two agents, contrasts the claimed arrangement with the alternative of letting “a single agent first generate code and then generate tests, within the same conversation”, and reports that the separate-agent arrangement yields materially higher pass@1 and test-case accuracy than the single-agent alternative (Huang, Section 4.7, Tables 6 and 7). Carrara supplies a further and independent reason to adopt this arrangement, because Carrara teaches generating the validation tests for generated control code by generative AI expressly “To mitigate the need for a system developer to create custom test scripts 2202 to validate operation of control code 1008, 2102 prior to deployment” (Carrara, [0156]), and teaches that those test scripts are generated after and on the basis of the generated code, “For each test scenario devised by the generative AI component 210 for the control code 1008, 2102 under analysis, the generative AI component 210 can generate one or more associated test scripts 2202” (Carrara, [0156]; see also [0158]). Applying Huang’s two-model arrangement to Carrara’s system would predictably yield validation tests that are generated independently of the model that produced the code, thereby improving the reliability of the validation, with a reasonable expectation of success.
Carrara in view of Huang does not expressly disclose the three specific validation operations recited in the limitation validating, by the code evaluation service, each of the plurality of suggested computer code blocks for syntax, security, and functionality, wherein the validating includes: Each recited validation operation is taught by a reference in an analogous art, as set forth below.
Carrara in view of Huang is silent to disclose; however, in an analogous art, Fanning teaches performing syntax validation by constructing an abstract syntax tree (AST) representing a hierarchical structure of each suggested computer code block and traversing the AST to perform semantic analysis (Fanning discloses that one or more parsers 108 create abstract syntax trees (ASTs) from source code, [0022]; that the AST representation of the source code is subsequently traversed in order to construct a hierarchical symbol table expressing the scoping and hierarchical structure of the code, [0024], [0028]; and that a multi-pass traversal of the AST provides semantic analysis by static means, including semantic checks that identify potential correctness issues in the code, [0016], [0021].)
It would have been obvious to one of ordinary skill in the art at the time the invention was made before the effective filing date of the claimed invention to validate the syntax of each suggested computer code block of the combined Carrara/Huang system by constructing an abstract syntax tree representing the hierarchical structure of the block and traversing that tree to perform semantic analysis, as taught by Fanning, because Fanning teaches that constructing an abstract syntax tree from source code and traversing it yields significant useful semantic analysis that identifies correctness issues in the code by static means, before the code is run (Fanning, [0016]). Applying Fanning’s AST-based syntactic and semantic analysis to the code generated by Carrara’s generative AI model is the use of a known source-code-analysis technique, in the same field of source code analysis, upon a known code-generation system ready for improvement, to yield the predictable result of identifying syntactically and semantically defective generated code before it is presented to the developer, with a reasonable expectation of success.
Carrara in view of Huang and Fanning is silent to disclose; however, in an analogous art, Schaefer teaches performing security validation by analyzing each suggested computer code block using a static application security testing tool to identify security vulnerabilities (Schaefer discloses that validation of the code generated with the assistance of a large language model is performed in combination with a code analysis tool comprising “a Static Application Security Testing (SAST) tool” that analyzes the source code to determine whether the identified issues are resolved, [0042]; Schaefer further discloses that “SAST tools analyze source code at rest, i.e., without executing it, to identify potential security vulnerabilities or other errors”, [0002].)
It would have been obvious to one of ordinary skill in the art at the time the invention was made before the effective filing date of the claimed invention to validate the security of each suggested computer code block of the combined Carrara/Huang/Fanning system by analyzing each block with a static application security testing tool to identify security vulnerabilities, as taught by Schaefer, because Schaefer teaches that SAST tools analyze source code at rest, without executing it, to identify potential security vulnerabilities or other errors (Schaefer, [0002], [0042]), and because Schaefer is directed to the same problem of ensuring that code produced with the assistance of a large language model is fit to be presented to and adopted by a developer (Schaefer, [0005]). Applying Schaefer’s static application security testing to the generated code of the combined system would predictably identify insecure generated code before it reaches the developer, with a reasonable expectation of success.
Carrara in view of Huang, Fanning, and Schaefer is silent to disclose the functionality validation performed upon each of the plurality of code blocks, the automatic discarding, the ranking, and the display of the highest ranking block; however, in an analogous art, Clement teaches performing functionality validation by executing the plurality of unit tests by executing each of the plurality of suggested computer code blocks under test conditions defined by the plurality of unit tests generated by the second LLM (Clement discloses generating a plurality of candidate method bodies with a deep learning model, [0070], and thereafter testing each of those candidates by executing them against the provided set of test cases, stating that “The remaining candidate method bodies are then tested using the set of test cases provided” and that “Only method candidates which pass all the test cases provided by the user are kept”, [0073]; Clement further discloses that “A test case is a source code snippet containing instructions and assertions that verify the functionality and behavior of a source code component, such as a method body”, [0016], such that the test cases define the conditions under which each candidate is executed.)
automatically discarding, by the code evaluation service, any suggested computer code block that fails any of the syntax, security, or functionality validations (Clement discloses eliminating each candidate that fails a validation, without user intervention, at each validation stage: “Those candidate method bodies that fail the syntax validation are eliminated”, [0071]; “the candidates which do not compile are eliminated”, [0072]; and “Those candidate method bodies that fail a test case are eliminated”, [0073].)
ranking, by the code evaluation service, any valid suggested computer code blocks that were not discarded, based on one or more quality metrics (Clement discloses that “The remaining candidate method bodies are then ranked based on one or more quality metrics”, wherein a Programming Language Understanding Metric (PLUM) score “aggregates several code quality metrics into a single score used to evaluate a candidate method body, where the highest score indicates a better-quality candidate method body”, and wherein “the set of code quality metrics include one or more of the following: code complexity (e.g., cyclomatic complexity), code size, maintainability index, code coupling and cohesion, code readability (variable and function names), code understandability (presence of comments), and performance measures”, [0074]; Clement further discloses that “A PLUM score is computed for each candidate method body”, [0081].) and
causing display of a highest ranking valid suggested computer code block (Clement discloses that “The candidate method bodies that pass all the tests are ranked in the order of their PLUM score from highest value to lowest value” and that “The candidate method bodies are then output based on their ranking”, and that they “may be output into a source code program, source code editor, or returned as a web response”, [0082]; see also [0014], “The top ranked source code snippets are then output in the ranked order”.)
Clement further teaches that these operations are performed by a code evaluation service, disclosing that the code evaluation service receives the generated candidates and carries out the validation, testing, ranking, and presentation: “Each of the candidate method bodies 128 is validated for syntactic correctness and then tested using the test cases 122 by a test and validation engine 130. Those candidate method bodies passing the validation and test cases 132 are then ranked by the ranking engine 134. The ranked method bodies 136 are then presented to the developer in the ranked order”, Clement, [0026]. Clement further discloses that the technique “may be implemented as part of a source code editor, an integrated development environment (IDE), and/or a web service” with which the developer interacts through a set of application programming interfaces, [0018].
It would have been obvious to one of ordinary skill in the art at the time the invention was made before the effective filing date of the claimed invention to validate the functionality of each of the plurality of suggested computer code blocks of the combined Carrara/Huang/Fanning/Schaefer system by executing each block against the plurality of unit tests, to automatically discard any block failing a validation, to rank the blocks that were not discarded on the basis of one or more quality metrics, and to cause display of the highest ranking block, as taught by Clement, because Clement teaches that a plurality of candidate code bodies generated by a model will differ in correctness and in quality, and therefore validates each candidate, eliminates those that fail, and ranks those that remain so that “the highest score indicates a better-quality candidate method body” (Clement, [0071]–[0074]). Clement thereby addresses the same need addressed by Carrara, which generates and executes validation tests upon generated control code “in order to comprehensively validate proper operation of the control code” (Carrara, [0159]), while additionally resolving which of several validated candidates should be surfaced to the developer. Applying Clement’s validate-eliminate-rank-output sequence to the plurality of code blocks generated in the combined system would predictably ensure that only code that has passed validation is presented, and that the block presented is the one scoring highest on objective measures of code quality, with a reasonable expectation of success.
With respect to claim 2 (Original), Carrara teaches wherein the request is received from an Integrated Development Environment (IDE) and the causing the display includes causing the IDE to display the highest ranking valid suggested computer code block (Carrara discloses that the user interface component 204 of the industrial IDE system 202 receives the user input, [0061]; that the IDE system’s development services include a control code generation copilot having a generative AI component 210 that responds to natural language prompts submitted by the user as part of design input 312, [0068]; and that the recommended control code is rendered to the user in the code window 1104 of the copilot window 802 of that IDE, [0092], figures 11 and 14.)
With respect to claim 3 (Original), Carrara teaches wherein the contextual information includes a location within existing source code at which the requested computer code is to be inserted (Carrara discloses that the user selects a location of the existing control code 702 at which the generated code is to be added, and that the generated control code is inserted at that specified location, [0103], [0105], figure 14.)
With respect to claim 6 (Original), Carrara in view of Huang, Fanning, and Schaefer is silent to disclose; however, in an analogous art, Clement teaches wherein the one or more quality metrics include a cohesion metric (Clement discloses that the set of code quality metrics by which the candidates are ranked includes “code coupling and cohesion”, [0074], and further defines the metric, “Cohesion refers to degree to which the source code elements belong to each other. High cohesion is indicative of the source code being understandable, reliable, and reusable”, [0079].) The cohesion metric is taught by the same reference, and in the same passage, as the ranking limitation of claim 1, and the motivation to combine Clement is that set forth in the rejection of claim 1 above.
With respect to claim 8, the claim recites a method comprising limitations similar to the operations recited in claim 1 and is rejected for the same reasons set forth for claim 1; Carrara further teaches a corresponding method carried out by the industrial IDE system 202 ([0005], [0058], [0060], [0092]).
With respect to claim 15, the claim recites a non-transitory machine-readable medium storing instructions which, when executed by one or more processors, cause the one or more processors to perform operations similar to those recited in claim 1, and is rejected for the same reasons set forth for claim 1; Carrara further teaches a memory 220 storing computer-executable instructions that are executed by one or more processors 218 ([0060], figure 2).
With respect to claims 9 and 16 recite limitations similar to claim 2 and are rejected for the same reasons set forth for claim 2.
With respect to claims 10 and 17 recite limitations similar to claim 3 and are rejected for the same reasons set forth for claim 3.
With respect to claims 13 and 20 recite limitations similar to claim 6 and are rejected for the same reasons set forth for claim 6.
Claims 4-5, 11-12 and 18-19 are rejected under 35 U.S.C. 103 as being unpatentable over Carrara et al. (US Pub. No. 2025/0298585, hereinafter “Carrara” – previously presented) in view of Dong Huang et al. (“AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation”, hereinafter “Huang” – previously presented) in view of Fanning et al. (US Pub. No. 2012/0311535, hereinafter “Fanning” – previously presented) in view of Schaefer et al. (US Pub. No. 2025/0245122, hereinafter “Schaefer” – previously presented) in view of Clement et al. (US Pub. No. 2023/0128008, hereinafter Clement) and further in view of Bei Chen et al. (“CodeT: Code Generation with Generated Tests”, hereinafter “Chen” – previously presented).
With respect to claim 4 (Original), Carrara in view of Huang, Fanning, Schaefer, and Clement is silent to disclose; however, in an analogous art, Chen teaches wherein the contextual information includes examples of input parameters of a function within the requested computer code (Chen discloses that the code solutions are generated by the language model from a context that “contains natural language problem description in the form of code comment, and a code snippet that includes statements such as imports and the function header”, Chen, Section 2, and that “The original contexts include example input-output cases” for the function defined in that context, Chen, Section 3; Chen further discloses constructing a one-shot version of a benchmark “by appending a single input-output example to the problem description” supplied to the model, Chen, Appendix G. The input portion of such example input-output cases constitutes examples of the input parameters of the function for which the code is requested.)
It would have been obvious to one of ordinary skill in the art at the time the invention was made before the effective filing date of the claimed invention to include, within the contextual information upon which the prompt of the combined Carrara/Huang/Fanning/Schaefer/Clement system is based, examples of input parameters of a function within the requested computer code, as taught by Chen, because Chen teaches that “the example input-output cases may provide useful information for code generation” and demonstrates experimentally that code generation performance obtained using contexts that include the example input-output cases exceeds that obtained using contexts from which those cases have been removed (Chen, Appendix B and Table 6), and further teaches that appending an input-output example to the problem description serves as a formatting cue to the model (Chen, Appendix G). Supplying such input-parameter examples as part of the contextual information of the combined system would predictably guide the large language model to generate code conforming to the expected inputs of the requested function, thereby improving the quality of the generated suggestions, with a reasonable expectation of success. With respect to claim 5 (Original), Carrara in view of Huang, Fanning, Schaefer, and Clement is silent to disclose; however, in an analogous art, Chen teaches wherein the contextual information includes examples of output of a function within the requested computer code (Chen discloses that “The original contexts include example input-output cases” for the function defined in the context from which the code solutions are generated, Chen, Section 3, and that a test case for that function is “a pair of input and expected output for the function defined in the context”, Chen, Section 2.1; Chen further discloses appending “a single input-output example to the problem description” supplied to the model, Chen, Appendix G. The output portion of such example input-output cases constitutes examples of the output of the function for which the code is requested.) The motivation to combine Chen is that set forth in the rejection of claim 4 above.
With respect to claims 11 and 18 recite limitations similar to claim 4 and are rejected for the same reasons set forth for claim 4.
With respect to claims 12 and 19 recite limitations similar to claim 5 and are rejected for the same reasons set forth for claim 5.
Claim 21 is rejected under 35 U.S.C. 103 as being unpatentable over Carrara et al. (US Pub. No. 2025/0298585, hereinafter “Carrara” – previously presented) in view of Dong Huang et al. (“AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation”, hereinafter “Huang” – previously presented) in view of Fanning et al. (US Pub. No. 2012/0311535, hereinafter “Fanning” – previously presented) in view of Schaefer et al. (US Pub. No. 2025/0245122, hereinafter “Schaefer” – previously presented) in view of Clement et al. (US Pub. No. 2023/0128008, hereinafter Clement) and further in view of Aivosto Oy (“Project Metrics Help – Cohesion metrics”, hereinafter “Aivosto”).
With respect to claim 21 (New), Carrara in view of Huang, Fanning, Schaefer, and Clement teaches the ranking of the validated candidate code blocks on the basis of a cohesion metric, as set forth in the rejection of claims 1 and 6 above (Clement, [0074], [0079]), but is silent to disclose the manner in which that cohesion metric is computed; however, in an analogous art, Aivosto teaches wherein ranking the valid suggested computer code blocks comprises, for each valid suggested computer code block:
identifying pairs of methods in the valid suggested computer code block that access a same class-level variable or in which a first method of the pair calls a second method of the pair (Aivosto discloses determining which methods of a class are related, stating that “Methods A and B are related if: they both access the same class-level variable, or A calls B or vice versa”, pages 1–2.)
generating a graph linking the identified pairs of methods (Aivosto discloses that “After determining the related methods, we draw a graph linking the related methods to each other”, page 2.)
determining a cohesion metric based on a number of connected groups of methods in the graph (Aivosto discloses that “LCOM4 equals the number of connected groups of methods”, page 2, and that “LCOM4 measures the number of ‘connected components’ in a class”, wherein “A connected component is a set of related methods (and class-level variables)”, page 1; Aivosto further discloses that “LCOM4=1 indicates a cohesive class, which is the ‘good’ class” and that “LCOM4>=2 indicates a problem”, page 2.)
It would have been obvious to one of ordinary skill in the art at the time the invention was made before the effective filing date of the claimed invention to compute the cohesion metric by which the validated candidate code blocks of the combined Carrara/Huang/Fanning/Schaefer/Clement system are ranked in the manner taught by Aivosto, namely by identifying the related pairs of methods, drawing a graph linking them, and counting the connected groups of methods in that graph, because Clement teaches ranking the surviving candidates by code quality metrics including cohesion (Clement, [0074]) and teaches that “High cohesion is indicative of the source code being understandable, reliable, and reusable while a low cohesion is indicative of the source code being difficult to maintain, hard to reuse or understand” (Clement, [0079]), but does not specify how the cohesion metric is to be computed, and because Aivosto teaches a known, published, and tool-implemented method of computing that metric together with its significance, namely that “There should be only one such a component in each class” and that a value of two or more indicates that “The class should be split into so many smaller classes” (Aivosto, pages 1–2). The selection of the LCOM4 computation taught by Aivosto to compute the cohesion metric called for by Clement is the selection of a known technique from a finite number of identified, predictable solutions for measuring cohesion, to perform the same function it was known to perform, with a reasonable expectation of success.
Carrara in view of Huang, Fanning, Schaefer, and Aivosto is silent to disclose; however, in an analogous art, Clement teaches ranking the valid suggested computer code blocks based on the cohesion metric. (Clement discloses that the candidate code bodies surviving validation “are then ranked based on one or more quality metrics”, wherein the set of quality metrics used to rank them includes “code coupling and cohesion”, [0074], and that “The candidate method bodies that pass all the tests are ranked in the order of their PLUM score from highest value to lowest value”, [0082].)
It would have been obvious to one of ordinary skill in the art at the time the invention was made before the effective filing date of the claimed invention to rank the valid suggested computer code blocks of the combined Carrara/Huang/Fanning/Schaefer/Aivosto system on the basis of the cohesion metric computed in the manner taught by Aivosto, as taught by Clement, because Clement teaches ranking the candidate code bodies that survive validation on the basis of one or more code quality metrics, expressly including “code coupling and cohesion”, so that “the highest score indicates a better-quality candidate method body” (Clement, [0074]), and teaches outputting those candidates in the resulting ranked order (Clement, [0082]). Ranking the validated candidate code blocks by the cohesion value computed in the manner taught by Aivosto would predictably permit the candidate exhibiting the greatest cohesion, and therefore the code that is, in Aivosto’s terms, most “understandable, reliable, and reusable”, to be identified and surfaced to the developer, with a reasonable expectation of success.
Claim 22 is rejected under 35 U.S.C. 103 as being unpatentable over Carrara et al. (US Pub. No. 2025/0298585, hereinafter “Carrara” – previously presented) in view of Dong Huang et al. (“AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation”, hereinafter “Huang” – previously presented) in view of Fanning et al. (US Pub. No. 2012/0311535, hereinafter “Fanning” – previously presented) in view of Schaefer et al. (US Pub. No. 2025/0245122, hereinafter “Schaefer” – previously presented) in view of Clement et al. (US Pub. No. 2023/0128008, hereinafter Clement) and further in view of Michele Tufano et al. (“Unit Test Case Generation with Transformers and Focal Context”, hereinafter “Tufano”).
With respect to claim 22 (new), Carrara in view of Huang, Fanning, Schaefer, and Clement is silent to disclose; however, in an analogous art, Tufano teaches wherein the second LLM is fine-tuned to generate the plurality of unit tests (Tufano discloses a large sequence-to-sequence transformer language model that is trained specifically for the generation of unit tests, stating that the approach adopts “a two-step training procedure consisting of denoising pretraining on a large unsupervised Java corpus, and supervised finetuning for a downstream translation task of generating unit tests”, Tufano, Abstract; that “Our approach relies on a large sequence-to-sequence transformer model pretrained both on English and Java source code, then finetuned on the task of generating unit test cases”, Tufano, Section I; and that the target model of the disclosed system is the model “pretrained on English and code, then finetuned” for that task, Tufano, Section V.)
It would have been obvious to one of ordinary skill in the art at the time the invention was made before the effective filing date of the claimed invention to fine-tune the second large language model of the combined Carrara/Huang/Fanning/Schaefer/Clement system to generate the plurality of unit tests, as taught by Tufano, because Tufano teaches that a language model finetuned upon a parallel corpus of developer-written test cases and their corresponding focal methods generates unit tests that are more readable and more useful to developers than those produced by prior automated approaches, and reports that the finetuned model generated passing test cases covering a substantial portion of the focal methods under evaluation (Tufano, Abstract and Section I). Applying Tufano’s finetuning to the second, test-generating model of the combined system is the use of a known technique to improve a similar model in the same way, and would predictably improve the quality and the developer-usability of the unit tests upon which the functionality validation of the combined system depends, with a reasonable expectation of success.
It is further noted that the limitation “is fine-tuned to generate the plurality of unit tests” is a product-by-process limitation recited in a claim drawn to a system. Determination of patentability of such a claim is based upon the product itself, and the limitation is entitled to patentable weight only to the extent that it imparts a structural difference to the claimed system as compared with the prior art. See MPEP 2113; In re Thorpe, 777 F.2d 695, 698 (Fed. Cir. 1985) (“If the product in a product-by-process claim is the same as or obvious from a product of the prior art, the claim is unpatentable even though the prior product was made by a different process.”). Applicant has not identified any structural difference in the second large language model that results from the recited fine-tuning. This is set forth as an additional basis, and does not replace the rejection over Tufano set forth above.
Response to Arguments
Applicant’s arguments filed August 18, 2026 have been fully considered. To the extent the arguments are directed to the combination applied in the previous Office action, they are moot in view of the new grounds of rejection set forth above, which were necessitated by Applicant’s amendment. The arguments are nonetheless addressed below as they bear upon the present rejection.
A. The recited code evaluation service and the post-generation query
At pages 9–10 of the Remarks, Applicant argues that the amended claims recite a specific sequence and allocation of operations, namely that a code evaluation service receives the candidate code blocks from the first LLM, thereafter sends a query to a second LLM requesting generation of unit tests, receives those unit tests, and then performs the validating, discarding, and ranking, and that the cited combination does not teach “this claimed sequence and allocation of operations.”
The argument is not persuasive because the recited sequence and allocation are expressly taught. As to sequence, Huang states: “The process begins by inputting tasks/code generation requirements into the code generation agent (i.e., Agent#1: the programmer agent). Subsequently, the test case generator (i.e., Agent#2: the test designer agent) is tasked with generating test cases, which are used to evaluate the correctness of the code snippets produced by the programmer agent” (Huang, Section 3.1). Code generation therefore occurs first, and the tasking of the separate test-generating model occurs subsequently. As to the allocation of operations to a single service, Huang discloses that the code snippets and the test cases are both collected by the test executor agent, which “validates these code snippets along with the test cases in a local environment” (Huang, Sections 3.1, 3.4), and Clement discloses a single system in which “Each of the candidate method bodies 128 is validated for syntactic correctness and then tested using the test cases 122 by a test and validation engine 130. Those candidate method bodies passing the validation and test cases 132 are then ranked by the ranking engine 134. The ranked method bodies 136 are then presented to the developer in the ranked order” (Clement, [0026]).
Applicant’s remaining reliance upon the label “code evaluation service” is not persuasive. The recitation of a name for the entity that performs operations otherwise taught by the prior art does not distinguish the claim, where the prior art performs those same operations in the same order. See MPEP 2111.02; In re Hiniker Co., 150 F.3d 1362, 1369 (Fed. Cir. 1998).
B. Carrara
At page 10 of the Remarks, Applicant argues that Carrara does not disclose a first LLM generating the code blocks together with a second, separate LLM generating the unit tests in response to a query sent by a code evaluation service after receipt of the code blocks. Carrara is not relied upon for the separateness of the second model; that limitation is addressed by Huang, as set forth above and in Section A.
Applicant’s characterization of Carrara is otherwise correct and is consistent with the present rejection. Applicant acknowledges at page 10 that “Carrara also describes testing industrial control code and using generative AI to infer test scenarios and generate test scripts.” That acknowledgment is significant, because Carrara teaches that those test scripts are generated after, and on the basis of, the code that was generated: “For each test scenario devised by the generative AI component 210 for the control code 1008, 2102 under analysis, the generative AI component 210 can generate one or more associated test scripts 2202” (Carrara, [0156]), and “based on analysis of the control code 1008, 2102 and inferences of the types of validation tests that should be performed on the code prior to deployment, generate test scripts 2202 for validating that respective portions of control code 1008, 2102 will correctly perform functions that those portions were designed to carry out” (Carrara, [0158]). Carrara further teaches an express reason for generating those tests automatically rather than requiring the developer to author them: “To mitigate the need for a system developer to create custom test scripts 2202 to validate operation of control code 1008, 2102 prior to deployment” (Carrara, [0156]).
C. Fanning
At page 11 of the Remarks, Applicant argues that Fanning does not disclose an LLM that generates candidate code, an LLM that generates unit tests, or a code evaluation service that queries a second LLM. Fanning is not relied upon for any of those limitations. Fanning is relied upon solely for the syntax-validation limitation, namely constructing an abstract syntax tree representing a hierarchical structure of each suggested computer code block and traversing the AST to perform semantic analysis, which Fanning expressly teaches at [0016], [0021]–[0022], [0024], and [0028].
One cannot show nonobviousness by attacking references individually where the rejection is based upon a combination of references. See In re Keller, 642 F.2d 413 (CCPA 1981); In re Merck & Co., 800 F.2d 1091 (Fed. Cir. 1986); MPEP 2145(IV).
D. Schaefer
At pages 11–12 of the Remarks, Applicant argues that Schaefer does not disclose the post-generation query to a separate second LLM, and that Schaefer’s reference to “running tests or Continuous Integration checks for the relevant code base (if defined)” concerns pre-existing tests. Applicant’s reading of Schaefer’s paragraph [0040] is accepted, and the present rejection does not rely upon it. Schaefer is relied upon solely for the security-validation limitation, namely performing security validation using a static application security testing tool to identify security vulnerabilities, which Schaefer expressly teaches at [0002] and [0042]. The argument again attacks a reference individually and is not persuasive for the reasons stated in Section C above.
E. Li
At page 12 of the Remarks, Applicant argues that Li does not cure the missing limitations. Li is not applied in the present rejection, and the argument is moot.
F. Ray
At page 12 of the Remarks, Applicant argues that Ray does not disclose generating candidate code using a first LLM, generating unit tests using a second LLM, the post-generation query, or the allocation of the validation, discarding, and ranking operations to the code evaluation service. Ray is not applied in the present rejection, and the argument is moot. The quality-metric ranking limitation and the cohesion metric of claims 6, 13, and 20 are now addressed by Clement, which teaches ranking a plurality of validated candidate code bodies on the basis of code quality metrics including “code coupling and cohesion” in a single passage (Clement, [0074]), and which further teaches that the ranking is performed by the same system that receives the candidates, validates them, and eliminates those that fail (Clement, [0026], [0071]–[0074], [0082]).
G. The reason to combine and the asserted absence of a rational underpinning
At page 13 of the Remarks, Applicant argues that the previous Office action “did not identify a reason to configure the cited combination” so as to arrive at the seven enumerated operations, and that the cited references “address individual aspects of code generation, static analysis, security analysis, testing, filtering, and metrics” without establishing why those teachings would have been arranged into the claimed post-generation service interaction.
The reasons for the combination are drawn from the references themselves, and not from Applicant’s disclosure. First, Carrara teaches automatically generating the validation tests for generated code by generative AI in order “To mitigate the need for a system developer to create custom test scripts 2202 to validate operation of control code 1008, 2102 prior to deployment” (Carrara, [0156]), and teaches generating those tests after and on the basis of the generated code (Carrara, [0156], [0158]). A person of ordinary skill reading Carrara therefore had an express reason, supplied by the primary reference, to have a model generate the validation tests for the code that had just been generated.
Second, Huang supplies the express reason for using a second, separate model for that task: “functional errors may still go undetected as both the code and its test cases are generated by the same model,” and a tester “being the same LLM as the coder, might erroneously assess erroneous code as correct” (Huang, Section 2.2). Huang further reports that the two-agent arrangement measurably outperforms the single-agent alternative (Huang, Section 4.7, Tables 6 and 7). Notably, this is the same consideration Applicant identifies as the point of the claimed architecture, and Huang articulated it more than a year before Applicant’s effective filing date.
Third, Clement supplies the express reason for validating each candidate, eliminating those that fail, and ranking those that remain by quality metrics: a model produces a plurality of candidate code bodies of differing correctness and quality, and Clement resolves which one to present by eliminating the candidates that fail syntax, compilation, and testing, and ranking the survivors so that “the highest score indicates a better-quality candidate method body” (Clement, [0071]–[0074]). A rejection is proper where the reason to combine is drawn from the teachings of the prior art rather than reconstructed from the claims. See MPEP 2145(X)(A); In re McLaughlin, 443 F.2d 1392 (CCPA 1971).
The references are further not drawn from unrelated fields. Carrara, Huang, Schaefer, Clement, and Chen are each directed to the generation of computer code by a language model and to validating, evaluating, or selecting among the code so generated; Fanning is directed to static analysis of source code. All are within the field of Applicant’s endeavor or are reasonably pertinent to the particular problem with which Applicant was concerned. See MPEP 2141.01(a); In re Bigio, 381 F.3d 1320 (Fed. Cir. 2004).
Finally, reliance upon a plurality of references does not, of itself, establish nonobviousness. See In re Gorman, 933 F.2d 982, 986 (Fed. Cir. 1991); MPEP 2145(IV). The independent claims are rejected over five references, each applied to a discrete recited operation, and no limitation is divided between two references.
H. The dependent claims
At pages 13–14 of the Remarks, Applicant argues that claims 2–6, 9–13, and 16–20 are patentable by virtue of their dependency, and identifies the additional limitations of claims 2, 9, and 16; claims 3, 10, and 17; and claims 6, 13, and 20 without presenting a separate argument for their patentability. Where Applicant does not separately argue the patentability of a dependent claim, the claim stands or falls with the claim from which it depends. See In re Lovin, 652 F.3d 1349, 1357 (Fed. Cir. 2011); MPEP 2144.03. The additional limitations identified are nonetheless addressed on the merits in the rejections of claims 2, 3, and 6 above.
I. The rejection of claims 4–5, 11–12, and 18–19 over Liguori
At pages 14–15 of the Remarks, Applicant argues that Liguori’s prompt-template placeholders for a user prompt, data sources, and a response definition are not “examples of input parameters of a function within the requested computer code” or “examples of output of a function within the requested computer code.” Applicant’s argument as to Liguori is well taken and is accepted. Liguori is not applied in the present rejection. Claims 4–5, 11–12, and 18–19 are now rejected in view of Chen, which expressly teaches that the context from which the language model generates the code solution for a function includes example input-output cases for that function (Chen, Sections 2, 2.1, and 3; Appendices B and G), as set forth in the rejection above.
J. New claims 21 and 22
Applicant presents no argument directed to the patentability of new claims 21 and 22. Claim 21 is rejected in view of Aivosto, and claim 22 in view of Tufano, as set forth above. Applicant is invited to address these rejections on the merits.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ANIBAL RIVERACRUZ whose telephone number is (571)270-1200. The examiner can normally be reached Monday-Friday 9:30 AM-6:00 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Hyung S Sough can be reached at 5712726799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ANIBAL RIVERACRUZ/Primary Examiner, Art Unit 2192