Prosecution Insights
Last updated: August 17, 2026
Application No. 18/931,808

SYSTEM FOR EVALUATING A LARGE LANGUAGE MODEL GENERATED RESPONSE TO A USER QUERY

Non-Final OA §101§103
Filed
Oct 30, 2024
Examiner
LEVENTAL, YUVAL HAIM
Art Unit
2659
Tech Center
2600 — Communications
Assignee
Capital One Services LLC
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-62.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
4 currently pending
Career history
2
Total Applications
across all art units

Statute-Specific Performance

§101
26.1%
-13.9% vs TC avg
§103
26.1%
-13.9% vs TC avg
§102
21.7%
-18.3% vs TC avg
§112
13.0%
-27.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 0 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claims 1-20 are pending and have been examined. Priority Applicant has not claimed the benefit of, or priority to, any prior-filed domestic or foreign application. The effective filing date of the claimed invention is therefore October 30, 2024, the actual filing date of the present application. The prior art applied below qualifies as of that effective filing date. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1 (Statutory Category). Claim 1 is directed to a system (a machine), claim 10 is directed to a method (a process), and claim 19 is directed to a non-transitory computer-readable medium (an article of manufacture). Each independent claim is therefore directed to a statutory category of invention. Step 2A, Prong One (Recitation of a Judicial Exception). Independent claims 1, 10, and 19 recite, in pertinent part, obtaining a user query, generating a response to the user query, evaluating the response, and performing a response evaluation action based on a result of evaluating the response. Under their broadest reasonable interpretation, these limitations recite a mental process - i.e., concepts performed in the human mind, including observation, evaluation, and judgment. A person (for example, an editor or reviewer) can read a query, compose an answer to it, then critically review that answer from a different point of view - that of a quality reviewer - and decide whether to approve, revise, or reject it. The recitation that a large language model generates the response and evaluates the response does not remove the limitations from the mental-process grouping, because the large language model is recited at a high level of generality merely as a tool that produces and reviews text; the claimed generating and evaluating steps mirror the evaluation and judgment that a person performs in the mind or with pen and paper. The recitation that the model adopts a first persona to generate and a different second persona to evaluate likewise describes nothing more than a person assuming the role of author and then the role of reviewer. The claims therefore recite an abstract idea. Step 2A, Prong Two (Integration into a Practical Application). The judicial exception is not integrated into a practical application. Beyond the abstract idea, claim 1 recites the additional elements of one or more memories, one or more processors, and a large language model (claims 10 and 19 recite corresponding additional elements, including a non-transitory computer-readable medium). These additional elements amount to mere instructions to apply the abstract idea on generic computer components and to use a generically-recited large language model as a tool to perform the abstract idea. Obtaining the user query is insignificant extra-solution activity in the nature of data gathering, and providing the response or a result is insignificant extra-solution activity in the nature of outputting a result. The claims do not recite any improvement to the functioning of a computer or to any other technology or technical field; the asserted advance lies in the abstract idea itself - using a second persona of the model to evaluate the output of a first persona - and not in any improvement to the way the recited generic components or language model operate. Accordingly, the additional elements do not integrate the abstract idea into a practical application, and the claims are directed to the abstract idea. Step 2B (Inventive Concept). The claims do not include additional elements sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration, the processor, memory, and large language model used to obtain, generate, evaluate, and output text are recited at a high level of generality and perform well-understood, routine, and conventional functions. Obtaining an input, generating natural-language text with a large language model, and outputting a result were well-understood, routine, and conventional practices in the art as of the effective filing date, as evidenced by the prior art of record (see, e.g., Madaan (Self-Refine) and Bai (Constitutional AI), of record, which show that prompting a single large language model to generate an output and then to evaluate and refine that same output was a conventional practice). Considered individually and as an ordered combination, the additional elements amount to no more than mere instructions to apply the exception using generic computer components and therefore do not provide an inventive concept. The claims are not patent eligible. The dependent claims have been considered and do not cure the deficiencies of the independent claims. Each dependent claim is addressed separately below. Claim 2 recites that the result of the evaluation includes an approval or a denial of the response. This further describes the abstract evaluation and judgment (a reviewer approving or rejecting a draft), adds no additional element beyond those addressed above, and is neither integrated into a practical application nor significantly more. Claim 3 recites providing the response based on the result of evaluating the response. Providing a result is insignificant extra-solution activity in the nature of outputting a result and does not supply an inventive concept. Claim 4 recites modifying the response using the LLM, based at least in part on the result of the evaluation, to generate a modified response, and providing the modified response. Revising a draft based on review is part of the abstract mental process; the generically recited LLM remains a tool, and providing the modified response is extra-solution outputting. Claim 5 recites evaluating the modified response prior to providing it - further mental evaluation and judgment. Claim 6 recites evaluating the response according to a self-reflection technique - further describing the abstract mental evaluation. Claim 7 recites that the first persona is an output producer persona - merely labeling the role adopted when generating, adding no technical element. Claim 8 recites that the first persona is a quality assurance persona - likewise merely labeling the role adopted, adding no technical element. Claim 9 recites performing the evaluation based on a result of a comparison with a result of another evaluation of the response - further mental comparison and judgment. Claim 11 recites the same substance as claim 2 in method form and is ineligible for the same reasons. Claim 12 recites the same substance as claim 3 in method form and is ineligible for the same reasons. Claim 13 recites the same substance as claim 4 in method form and is ineligible for the same reasons. Claim 14 recites the same substance as claim 5 in method form and is ineligible for the same reasons. Claim 15 recites the same substance as claim 6 in method form and is ineligible for the same reasons. Claim 16 recites the same substance as claim 7 in method form and is ineligible for the same reasons. Claim 17 recites the same substance as claim 8 in method form and is ineligible for the same reasons. Claim 18 recites the same substance as claim 9 in method form and is ineligible for the same reasons. Claim 20 recites the same substance as claims 4 and 13 in computer-readable-medium form and is ineligible for the same reasons. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 3-5, 7, 8, 10, 12-14, 16, 17, 19, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Madaan et al., “Self-Refine: Iterative Refinement with Self-Feedback,” arXiv:2303.17651v2 (made publicly available May 25, 2023), hereinafter Madaan, in view of Shoshan (US 2025/0217174 A1), hereinafter Shoshan. Regarding claim 1, Madaan discloses: A system for evaluating a large language model (LLM) generated response to a user query, the system comprising: (Madaan discloses a system implementing the Self-Refine framework, which obtains an input and evaluates the LLM-generated response to that input by prompting the same LLM to produce feedback on the response. Madaan, p.1, Abstract; p.2, §2; Figure 1.) obtain the user query; (Madaan obtains an input to be processed, e.g., the user-provided prompt or dialogue context in the dialogue response generation task: “Given an input sequence, SELF-REFINE generates an initial output, provides feedback on the output, and refines the output according to the feedback.” Madaan, p.2, §2; Figure 1; p.3, Figure 2(a) (dialogue input); p.4, §3 (task list including Dialogue Response Generation).) generate a response to the user query based on a first prompt and using an LLM; (Madaan generates an initial output from the input using the generation prompt p_gen: “Given an input x, prompt pgen, and model M, SELF-REFINE generates an initial output y0” Madaan, p.2, §2, Eq. (1) (Generate); Algorithm 1; Figure 1. ) evaluate the response based on a second prompt and using the LLM; and (Madaan prompts the same model with the separate feedback prompt p_fb to assess its own previously generated output and produce feedback: “uses the same model M to provide feedback fb_t on its own output”: Madaan, p.3, §2, Eq. (2) (Feedback); Algorithm 1; Figure 1. ) perform a response evaluation action based on a result of evaluating the response. (Madaan uses the feedback to refine, i.e., regenerate, the output via the refinement prompt p_refine - a response evaluation action taken based on the result of the evaluation: “Next, SELF-REFINE uses M to refine its most recent output, given its own feedback… To inform the model about the previous iterations, we retain the history of previous feedback and outputs by appending them to the prompt. Intuitively, this allows the model to learn from past mistakes and avoid repeating them.” Madaan, p.4, §2, Eqs. (3)-(4) (Refine); Algorithm 1.) Madaan does not disclose: one or more memories; one or more processors, communicatively coupled to the one or more memories, configured to perform the recited operations; wherein the first prompt causes the LLM to generate the response based on a first persona; or wherein the second prompt causes the LLM to evaluate the response based on a second persona that is different from the first persona. Shoshan discloses one or more memories (“The memory 530 may include a main memory 532, a static memory 534, and a storage unit 536, each accessible to the processors 510 such as via the bus 502. The main memory 532, the static memory 534, and the storage unit 536 store the instructions 516 embodying any one or more of the methodologies or functions described herein” Shoshan, ¶[0087]; claim 1); one or more processors, communicatively coupled to the one or more memories (“The machine 500 may include processors 510, memory 530, and I/O components 550, which may be configured to communicate with each other such as via a bus 502,” the processors including “a processor 512 and a processor 514 that may execute the instructions 516” Shoshan, ¶[0086]; claim 1); and first and second prompts (initial instruction sets) that cause a shared LLM to generate a response based on a first persona and to evaluate based on a second persona that is different from the first persona (“Each unique set of initial instructions can be considered to create a separate persona, even if the different set of initial instructions are being applied to the same LLM. In such a case, that single LLM could be considered to have multiple personas” Shoshan, ¶[0014]; ¶[0018]; claims 1-2; Figure 1). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Madaan such that the generation prompt establishes a first persona of the shared model and the feedback prompt establishes a second, different persona of the same model, and to implement the framework on one or more processors communicatively coupled to one or more memories, as taught by Shoshan. One of ordinary skill in the art would have been motivated to make this modification in order to avoid the incorrect or non-holistic answers that can result from a single persona’s bias and to obtain more comprehensive responses reflecting multiple perspectives, as Shoshan expressly teaches: “If the assistant’s persona has some sort of bias or other trait that would cause potentially incorrect or at least non-holistic answers to be generated, the user may be getting an incorrect or at least non-holistic perspective” and “It would be beneficial to have a system where a user is able to be presented with generated LLM content from multiple different perspectives. Engaging LLMs with various perspectives leads to more comprehensive responses, reflecting the diversity of opinions and approaches in a domain” (Shoshan, ¶[0011]-[0012]). Regarding claim 10, Madaan discloses: A method for evaluating a large language model (LLM) generated response to a user query, comprising: (the Self-Refine framework, performed as a computer-implemented method. Madaan, p.1, Abstract; p.2, §2.) obtaining, by a system, a user query; (Madaan obtains the input to be processed, as quoted for claim 1. Madaan, p.2, §2; Figure 1.) generating, by the system, a response to the user query using an LLM; (the model generates the initial output via p_gen in the generator role, as quoted for claim 1. Madaan, p.2, §2, Eq. (1) (Generate); Algorithm 1.) evaluating, by the system, the response using the LLM; and (the same model produces feedback on its output via p_fb, as quoted for claim 1. Madaan, p.3, §2, Eq. (2) (Feedback); Algorithm 1.) performing, by the system, a response evaluation action based on a result of evaluating the response. (the model refines/regenerates the output based on the feedback, as quoted for claim 1. Madaan, p.4, §2, Eqs. (3)-(4) (Refine); Algorithm 1.) Madaan does not disclose wherein the LLM generates the response based on a first persona or wherein the LLM evaluates the response based on a second persona that is different from the first persona. Shoshan discloses an LLM generating content based on a first persona and evaluating based on a second, different persona, each persona created by a unique initial instruction set fed to the shared LLM, as quoted with respect to claim 1 (Shoshan, ¶[0014]; ¶[0018]; claims 1-2). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method of Madaan such that the generation prompt establishes a first persona and the feedback prompt establishes a second, different persona of the shared model, as taught by Shoshan. One of ordinary skill in the art would have been motivated to make this modification for the reasons set forth with respect to claim 1 - to avoid single-persona bias and to obtain more comprehensive responses from multiple perspectives (Shoshan, ¶[0011]-[0012]). Regarding claim 19, Madaan discloses: obtain a query provided via user input; (Madaan obtains the input, as quoted for claim 1. Madaan, p.2, §2; Figure 1.) obtain, as a first output of a large language model (LLM), a response to the query; (the initial output generated via p_gen is the first output of the LLM in its generator role, as quoted for claim 1. Madaan, p.2, §2, Eq. (1) (Generate); Algorithm 1.) obtain, as a second output of the LLM, a result associated with an evaluation of the response; and (the feedback produced via p_fb is a second output of the same LLM, as quoted for claim 1. Madaan, p.3, §2, Eq. (2) (Feedback); Algorithm 1.) perform a response evaluation action based on the result of the evaluation of the response. (the model refines/regenerates the output based on the feedback, as quoted for claim 1. Madaan, p.4, §2, Eqs. (3)-(4) (Refine); Algorithm 1.) Madaan does not disclose a non-transitory computer-readable medium storing a set of instructions, the set of instructions comprising: one or more instructions that, when executed by one or more processors of a system, cause the system to perform the recited operations; wherein the LLM is instructed to generate the first output based on adopting a first persona; or wherein the LLM is instructed to evaluate the response based on adopting a second persona that is different from the first persona. Shoshan discloses a non-transitory computer-readable medium storing instructions executed by one or more processors: “A non-transitory machine-readable medium storing instructions which, when executed by one or more processors, cause the one or more processors to perform operations” (Shoshan, claim 15; see ¶[0096], stating that “machine-readable medium” and “computer-readable medium” mean the same thing and may be used interchangeably; ¶[0087]). Shoshan further discloses the LLM adopting a first persona to generate and a second, different persona to evaluate, each persona created by a unique initial instruction set, as quoted with respect to claim 1 (Shoshan, ¶[0014]; claims 1-2). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to implement the framework of Madaan as instructions stored on a non-transitory computer-readable medium and executed by one or more processors, with the generation and feedback prompts establishing first and second personas of the shared model, as taught by Shoshan. One of ordinary skill in the art would have been motivated to make this modification for the reasons set forth with respect to claim 1 (Shoshan, ¶[0011]-[0012]). Regarding claim 3: The system of claim 1, wherein the one or more processors, to perform the response evaluation action, are configured to provide the response based on the result of evaluating the response. (the refined output is returned as the system output, per the language quoted for claim 1 (Eqs. (3)-(4)). Madaan, p.4, §2, Eqs. (3)-(4); Algorithm 1.) Regarding claim 4: The system of claim 1, wherein the one or more processors, to perform the response evaluation action, are configured to: modify the response using the LLM, and based at least in part on the result of evaluating the response, generate a modified response (Madaan discloses that the model refines its most recent output based on its own feedback, generating the refined output - a modified response. Madaan, p.4, §2, Eq. (3) (Refine); Figure 1.) and provide the modified response. (Madaan discloses providing the refined output as the output of the procedure. Madaan, p.4, §2, Eq. (4); Algorithm 1.) Regarding claim 5: The system of claim 4, wherein the one or more processors are further configured to evaluate the modified response prior to providing the modified response. (the procedure iterates: the refined output is itself fed back for further feedback before a final output is provided: “Iterating SELF-REFINE alternates between FEEDBACK and REFINE steps until a stopping condition is met. The stopping condition stop (fbt,t) either stops at a specified timestep t, or extracts a stopping indicator (e.g. a scalar stop score) from the feedback.” Madaan, p.4, §2; Algorithm 1 (Figure 3), p.3.) Regarding claim 7: The system of claim 1, wherein the first persona is an output producer persona. (As set forth with respect to claim 1, Madaan discloses that the model in its generator role produces the output (Madaan, p.2, §2, Eq. (1) (Generate)). Madaan does not disclose wherein the first persona is an output producer persona. Shoshan discloses that the persona is established by the initial instruction set fed to the LLM and may be any role the instructions request (Shoshan, ¶[0014]-[0015]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to characterize the first persona of the combination - the persona under which the model generates the output - as an output producer persona. One of ordinary skill in the art would have been motivated to make this modification in order to define each persona’s role expressly and thereby present generated content from defined, distinct perspectives, as Shoshan teaches (Shoshan, ¶[0011]-[0012], [0014]-[0015]).) Regarding claim 12: The method of claim 10, wherein performing the response evaluation action comprises providing the response based on the result of evaluating the response. (as set forth for claim 3. Madaan, p.4, §2, Eqs. (3)-(4); Algorithm 1.) Regarding claim 13: The method of claim 10, wherein performing the response evaluation action comprises: modifying the response using the LLM, and based at least in part on the result of evaluating the response, to generate a modified response; and providing the modified response. (Madaan discloses modifying the response using the LLM, based on the result of the evaluation, to generate a modified response, as set forth for claim 4 (Madaan, p.4, §2, Eq. (3) (Refine); Figure 1); and providing the modified response, as set forth for claim 4 (Madaan, p.4, §2, Eq. (4); Algorithm 1).) Regarding claim 14: The method of claim 13, further comprising evaluating the modified response prior to providing the modified response. (as set forth for claim 5. Madaan, p.4, §2; Algorithm 1.) Regarding claim 16: The method of claim 10, wherein the first persona is an output producer persona. (As set forth with respect to claim 7: Madaan discloses the generator role producing the output (Madaan, p.2, §2, Eq. (1) (Generate)) but does not disclose wherein the first persona is an output producer persona; Shoshan discloses that the persona may be any role the initial instructions request (Shoshan, ¶[0014]-[0015]); and it would have been obvious to combine for the reasons set forth with respect to claim 7.) Regarding claim 20: The non-transitory computer-readable medium of claim 19, wherein the one or more instructions, that cause the system to perform the response evaluation action, cause the system to: modify the response using the LLM, and based at least in part on the result of evaluating the response, to generate a modified response; and provide the modified response. (Madaan discloses modifying the response using the LLM, based on the result of the evaluation, to generate a modified response, as set forth for claim 4 (Madaan, p.4, §2, Eq. (3) (Refine); Figure 1); and providing the modified response, as set forth for claim 4 (Madaan, p.4, §2, Eq. (4); Algorithm 1).) Regarding claim 8: The system of claim 1, wherein the first persona is a quality assurance persona. (As set forth with respect to claim 1, the combination of Madaan and Shoshan discloses the first persona, under which the shared model generates the response. The combination does not expressly disclose wherein the first persona is a quality assurance persona. However, Madaan discloses that its feedback addresses the quality of the generated output: “Intuitively, the feedback may address multiple aspects of the output. For example, in code optimization, the feedback might address the efficiency, readability, and overall quality of the code.” Madaan, p.3, §2; see also p.6, §4 (numerical scores for different quality aspects). Shoshan teaches that the instruction-defined persona may be any specified role: “despite the term ‘persona’ being used, there is nothing requiring that the persona actually be representative of a person”; “The persona could be representative of anything that the initial instructions to the LLM request it to be” (Shoshan, ¶[0015]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to further specify that the first persona of the combination is a quality assurance persona, since Shoshan teaches that the persona may be any role the initial instructions request. One of ordinary skill in the art would have been motivated to make this modification in order to have a dedicated reviewing role perform the quality-directed feedback that Madaan discloses, thereby obtaining evaluations from a distinct perspective and avoiding the bias of a single persona (Shoshan, ¶[0011]-[0012], [0015]).) Regarding claim 17: The method of claim 10, wherein the first persona is a quality assurance persona. (The combination of Madaan and Shoshan does not expressly disclose wherein the first persona is a quality assurance persona. Madaan discloses quality-directed feedback and Shoshan discloses that the persona may be any role the initial instructions request, as quoted with respect to claim 8 (Madaan, p.3, §2; p.6, §4; Shoshan, ¶[0015]). It would have been obvious to combine for the reasons set forth with respect to claim 8.) Claims 2, 9, 11, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Madaan in view of Shoshan as applied to claims 1 and 10 above, and further in view of Zheng et al., “Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena,” arXiv:2306.05685v1 (made publicly available June 9, 2023), hereinafter Zheng. Regarding claim 2: The system of claim 1, wherein the result of the evaluation includes an approval of the response or a denial of the response. (As set forth with respect to claim 1, the combination of Madaan and Shoshan discloses evaluating the response based on the second prompt. The combination does not disclose wherein the result of the evaluation includes an approval of the response or a denial of the response. Zheng discloses: “Pairwise comparison. An LLM judge is presented with a question and two answers, and tasked to determine which one is better or declare a tie. The prompt used is given in Appendix Figure 5. Single answer grading. Alternatively, an LLM judge is asked to directly assign a score to a single answer. The prompt used for this scenario is in Figure 6 (Appendix).” Zheng, p.4, §3.1; Figure 6 (Appendix), p.13 (verdict/rating format).) Regarding claim 11: The method of claim 10, wherein the result of the evaluation includes an approval of the response or a denial of the response. (The combination of Madaan and Shoshan does not disclose wherein the result of the evaluation includes an approval of the response or a denial of the response. Zheng discloses this limitation, as quoted with respect to claim 2 (Zheng, p.4, §3.1; Figure 6 (Appendix), p.13). The reason to combine is set forth following claim 18 below.) Regarding claim 9: The system of claim 1, wherein the one or more processors, to perform the response evaluation action, are configured to perform the evaluation based on a result of a comparison of the result of evaluating the response and a result of another evaluation of the response. (As set forth with respect to claim 1, the combination of Madaan and Shoshan discloses evaluating the response based on the second prompt. The combination does not disclose performing the evaluation based on a result of a comparison of the result of evaluating the response and a result of another evaluation of the response. Zheng discloses: “Swapping positions. The position bias can be addressed by simple solutions. A conservative approach is to call a judge twice by swapping the order of two answers and only declare a win when an answer is preferred in both orders. If the results are inconsistent after swapping, we can call it a tie. Another more aggressive approach is to assign positions randomly, which can be effective at a large scale with the correct expectations. In the following experiments, we use the conservative one.” Zheng, p.4, §3.1; p.6, §3.4 (Swapping positions).) Regarding claim 18: The method of claim 10, wherein performing the response evaluation action comprises performing the evaluation based on a result of a comparison of the result of evaluating the response and a result of another evaluation of the response. (The combination of Madaan and Shoshan does not disclose this limitation. Zheng discloses it, as quoted with respect to claim 9 (Zheng, p.6, §3.4 (Swapping positions)). The reason to combine is set forth below.) With respect to claims 2 and 11, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the combination of Madaan and Shoshan such that the result of the second persona’s evaluation is expressed as a verdict approving or denying the response, as taught by Zheng. One of ordinary skill in the art would have been motivated to make this modification in order to obtain a scalable and explainable approximation of human preference judgments when evaluating generated responses, as Zheng expressly teaches:“LLM-as-a-judge is a scalable and explainable way to approximate human preferences” (Zheng, p.1, Abstract). With respect to claims 9 and 18, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to further modify the combination such that the evaluation is performed twice with the order of the evaluated answers swapped, and the final evaluation result is based on a comparison of the two evaluations, as taught by Zheng. One of ordinary skill in the art would have been motivated to make this modification in order to mitigate position bias in the evaluator’s verdicts and to declare an outcome only when it is consistent across both orders, as quoted with respect to claim 9 (Zheng, p.6, §3.4). Claims 6 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Madaan in view of Shoshan as applied to claims 1 and 10 above, and further in view of Shinn et al., “Reflexion: Language Agents with Verbal Reinforcement Learning,” arXiv:2303.11366v4 (made publicly available October 10, 2023), hereinafter Shinn. Regarding claim 6: The system of claim 1, wherein the one or more processors are configured to evaluate the response according to a self-reflection technique. (As set forth with respect to claim 1, the combination of Madaan and Shoshan discloses evaluating the response based on the second prompt and using the LLM. Madaan further discloses the same model generating feedback on its own output: “The main idea is to generate an initial output using an LLM; then, the same LLM provides feedback for its output and uses it to refine itself, iteratively.” Madaan, p.1, Abstract; p.3, §2, Eq. (2). The combination does not expressly disclose evaluating the response according to a self-reflection technique. However, Shinn discloses a self-reflection technique in which an agent verbally reflects on feedback regarding its own output and retains the reflection to improve subsequent outputs: “Reflexion agents verbally reflect on task feedback signals” (Shinn, p.1, Abstract), retaining the reflective text in memory to improve subsequent outputs; “The Self-Reflection model instantiated as an LLM, plays a crucial role in the Reflexion framework by generating verbal self-reflections to provide valuable feedback for future trials. Given a sparse reward signal, such as a binary success status (success/fail), the current trajectory, and its persistent memory mem, the self-reflection model generates nuanced and specific feedback.” Shinn, pp.3-4, §3; Figure 2 (Algorithm 1), p.4.) Regarding claim 15: The method of claim 10, wherein the response is evaluated according to a self-reflection technique. (The combination of Madaan and Shoshan does not expressly disclose that the response is evaluated according to a self-reflection technique. Shinn discloses a self-reflection technique, as quoted with respect to claim 6 (Shinn, p.1, Abstract; pp.3-4, §3). The reason to combine is set forth below.) With respect to claims 6 and 15, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the combination of Madaan and Shoshan such that the second persona’s evaluation of the response is performed according to the self-reflection technique of Shinn. One of ordinary skill in the art would have been motivated to make this modification in order to provide the model with a concrete direction to improve and to enable it to learn from prior mistakes across iterations, because as Shinn expressly teaches: “This self-reflective feedback acts as a ‘semantic’ gradient signal by providing the agent with a concrete direction to improve upon, helping it learn from prior mistakes to perform better on the task.” (Shinn, p.2, §1). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure: a. Saunders et al., “Self-critiquing models for assisting human evaluators,” arXiv:2206.05802 (2022): training a model to generate natural-language critiques of its own answers to assist evaluation. b. Kong et al., “Better Zero-Shot Reasoning with Role-Play Prompting,” arXiv:2308.07702 (2023): prompting a single large language model to adopt a specified role/persona to condition its responses. c. Korean Pub. KR 10-2026-0000825 A (Story Generation): a single story-generation model that generates content, generates plural expert personas, applies a persona-based critique prompt to select critiques, and modifies the content based on the selected critique. d. Japanese Pat. JP 7,682,457 B1 (Moriyama): plural first agents having different characteristics produce reactions, and a second agent having a predetermined attribute evaluates content based on those reactions, outputting an evaluation result. e. Bai et al., “Constitutional AI: Harmlessness from AI Feedback,” arXiv:2212.08073 (2022): a large language model critiques and revises its own responses against a set of stated principles; its preference-based evaluation stage employs an independent model rather than the same model that generated the response. Any inquiry concerning this communication or earlier communications from the examiner should be directed to YUVAL H. LEVENTAL whose telephone number is (571) 270-3130. The examiner can normally be reached Monday-Friday, 8:00 AM - 5:00 PM. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, PIERRE-LOUIS DESIR, can be reached at (571) 272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /YUVAL HAIM LEVENTAL/ Examiner, Art Unit 2659 /PIERRE LOUIS DESIR/Supervisory Patent Examiner, Art Unit 2659
Read full office action

Prosecution Timeline

Oct 30, 2024
Application Filed
Jul 17, 2026
Non-Final Rejection mailed — §101, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month