Prosecution Insights
Last updated: October 01, 2026
Application No. 18/893,422

CONTINUALLY EVALUATING AND MODIFYING ARTIFICIAL INTELLIGENCE ASSISTANT

Non-Final OA §103
Filed
Sep 23, 2024
Examiner
KY, KEVIN
Art Unit
2671
Tech Center
2600 — Communications
Assignee
Adobe Inc.
OA Round
3 (Non-Final)
77%
Grant Probability
Favorable
3-4
OA Rounds
6m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 77% — above average
77%
Career Allowance Rate
448 granted / 579 resolved
+15.4% vs TC avg
Strong +25% interview lift
Without
With
+25.2%
Interview Lift
resolved cases with interview
Typical timeline
2y 6m
Avg Prosecution
30 currently pending
Career history
595
Total Applications
across all art units

Statute-Specific Performance

§101
18.4%
-21.6% vs TC avg
§103
51.2%
+11.2% vs TC avg
§102
19.3%
-20.7% vs TC avg
§112
6.3%
-33.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 579 resolved cases

Office Action

§103
DETAILED ACTION Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 8/19/2026 has been entered. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-4, 7, 9 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ouyang et al (NPL: Training language models to follow instructions with human feedback) in view of Wu et al (NPL: Fine-Grained Human Feedback Gives Better Rewards for Language Model Training), in further view of Dmitriev et al (US 20250327687). Regarding claim 1, Ouyang discloses a computer-implemented method comprising: receiving, via one or more graphical user interfaces, a plurality of prompts (pg. 6 3.2 Dataset: Our prompt dataset consists primarily of text prompts submitted to the OpenAI API, specifically those using an earlier version of the InstructGPT models (trained via supervised learning on a subset of our demonstration data) on the Playground interface. Customers using the Playground were informed that their data could be used to train further models via a recurring notification any time InstructGPT models were used; Fig. 2 e.g. prompt sampled); generating, using a large language model based artificial intelligence assistant, a plurality of responses to the plurality of prompts (Fig. 2 A diagram illustrating the three steps of our method: (1) supervised fine-tuning (SFT), (2) reward model (RM) training, and (3) reinforcement learning via proximal policy optimization (PPO) on this reward model; generates an output); determining a plurality of errors in the plurality of responses (abstract: For example, large language models can generate outputs that are untruthful, toxic, or simply not helpful to the user; further fine-tune this supervised model using reinforcement learning from human feedback); classifying the plurality of errors as one of high-severity, mid-severity, or low-severity (Fig. 2 labeler ranks the outputs from best to worst; pg. 13 4.2 Results on public NLP datasets we run model samples through the Perspective API8 to obtain automatic toxicity scores, which is the standard evaluation procedure for this dataset, and we also send these samples to labelers to obtain ratings on absolute toxicity, toxicity relative to the prompt, continuity, and overall output preference); and Ouyang fails to teach where Wu teaches generating, based on a high-severity error of the classified plurality of errors, a modification to one or more components of the large language model based artificial intelligence assistant prioritizing the high-severity error over one or more errors of the plurality of errors classified as mid-severity or low-severity (pg. 8 4.5 LM Customization with FINE-GRAINED RLHF: adding more weight to a reward model associated with one specific desired behavior type (e.g., information completeness) may lead the generation more towards that behavior type compared to others; ‘short’ generates more relevant content, but is less factual and complete; (2) ‘long’, in contrast, gives the most factual and complete generation. This reflects that the LM is referencing a large amount of content from passages; (3) The ‘medium’ configuration balances the three rewards; therefore Wu teaches modifying the LLM (e.g. adding more weight) based on prioritizing high-severity error over mid-severity or low-severity (e.g. adding more weight to a reward model associated with one specific desired behavior type, which can be any of the three: short, medium, or long). Ouyang further fails to teach where Dmitriev teaches wherein the one or more components comprise at least one of a prompt rewrite component, an intent detection component, a data quality assurance pipeline, a concepts quality assurance pipeline, a chat history/user session database, or a documentation collection (¶129 In step 415, the output module 309 transmits the map error report to a map feedback system (e.g., via the map feedback API 123); ¶131 the LLM prompting module 305 iteratively modifies the one or more prompts, the one or more additional prompts, or a combination thereof to reduce a number of the one or more questions used during the collection of the one or more data items). Therefore, it would have been obvious to one with ordinary skill in the art before the effective filing date of the invention to have implemented the teaching of generating, based on a high-severity error of the classified plurality of errors, a modification to one or more components of the large language model based artificial intelligence assistant prioritizing the high-severity error over one or more errors of the plurality of errors classified as mid-severity or low-severity from Wu, and the teaching of wherein the one or more components comprise at least one of a prompt rewrite component, an intent detection component, a data quality assurance pipeline, a concepts quality assurance pipeline, a chat history/user session database, or a documentation collection from Dmitriev, into the method as disclosed by Ouyang. The motivation for doing this is to improve the reduction of false, toxic and other undesired model generation outputs, and further to improve LLM map feedback reporting. Regarding claim 2, Ouyang discloses the computer-implemented method of claim 1, wherein determining the plurality of errors in the plurality of responses comprises generating, using an annotation tool, annotated responses comprising error identification annotations (pg. 14 4.2 Results on public NLP datasets: we also send these samples to labelers to obtain ratings on absolute toxicity, toxicity relative to the prompt, continuity, and overall output preference). Regarding claim 3, Ouyang discloses the computer-implemented method of claim 2, further comprising associating one or more of the error identification annotations with at least one prompt of the plurality of prompts or a corresponding response of the plurality of responses (1 Introduction: pg. 2 we collect a dataset of human-labeled comparisons between outputs from our models on a larger set of API prompts; Fig. 2 labeler ranks the outputs from best to worst e.g. this is association between the error (e.g. worst rankings) with a corresponding response of the plurality of responses). Regarding claim 4, Ouyang discloses the computer-implemented method of claim 3, wherein determining the plurality of errors in the plurality of responses comprises generating, using an error analysis mechanism, indications of the plurality of errors based on the annotated responses (Fig. 7 Comparing human evaluations and automatic evaluations (Perspective API scores) on RealToxicityPrompts. A total of 1,729 prompts were labeled for three different 175B models, both with and without "respectful" instructions. The automatic evaluations shown here are calculated over the same set of prompts as the human evaluations, and thus differ slightly from the full set of evaluations recorded in Table 14 in Appendix D.). Regarding claim 7, Ouyang discloses the computer-implemented method of claim 1, wherein classifying the plurality of errors as one of high-severity, mid-severity, or low-severity comprises classifying an error of the plurality of errors as low-severity by determining that a response appears incorrect and can be corrected (Fig. 2 ranks outputs from best to worst; data is used to rain our reward models). Regarding claim 9, Ouyang discloses the computer-implemented method of claim 1, wherein modifying the one or more components of the large language model based artificial intelligence assistant comprises modifying at least one component of the one or more components of the large language model based artificial intelligence assistant using at least one of a user experience design engine, a prompt improvement engine, an in-house model generation engine, a synthetic data template engine, or a data index optimization engine (abstract: Starting with a set of labeler-written prompts and prompts submitted through the OpenAI API, we collect a dataset of labeler demonstrations of the desired model behavior, which we use to fine-tune GPT-3 using supervised learning. We then collect a dataset of rankings of model outputs, which we use to further fine-tune this supervised model using reinforcement learning from human feedback; our results show that fine-tuning with human feedback is a promising direction for aligning language models with human intent). Claim(s) 5-6 is/are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Ouyang, Wu, and Dmitriev as applied to claim 1 above, and further in view of Sun et al (NPL: AI hallucination: towards a comprehensive classification of distorted information in artificial intelligence-generated content). Regarding claim 5, the combination of Ouyang, Wu, and Dmitriev discloses the computer-implemented method of claim 1, but fails to teach where Sun teaches wherein classifying the plurality of errors as one of high-severity, mid-severity, or low-severity comprises classifying an error of the plurality of errors as high-severity by determining that a response appears correct but is incorrect (pg. 11 the system generates responses that appear plausible but are ultimately incorrect, requiring further scrutiny by the user; To address this issue, users planning to utilize generative artificial intelligence models to present a specific perspective can supply the system with a substantial amount of pre-collected data for analysis and request that the system generates a demonstration based on the provided reference materials. This approach also helps prevent the system from generating nonsensical responses.). Therefore, it would have been obvious to one with ordinary skill in the art before the effective filing date of the invention to have implemented the teaching of wherein classifying the plurality of errors as one of high-severity, mid-severity, or low-severity comprises classifying an error of the plurality of errors as high-severity by determining that a response appears correct but is incorrect from Sun into the method as disclosed by the combination of Ouyang, Wu, and Dmitriev. The motivation for doing this is to improve responses from artificial intelligence systems. Regarding claim 6, the combination of Ouyang, Wu, and Dmitriev discloses the computer-implemented method of claim 1, but fails to teach where Sun teaches wherein classifying the plurality of errors as one of high-severity, mid-severity, or low-severity comprises classifying an error of the plurality of errors as mid-severity by determining that a response appears incorrect and cannot be corrected (pg. 11 Factual errors: Factual errors pertain to inaccuracies in objective facts or actual data within system responses; When users seek fact-related information, identifying incorrect information can be challenging if they lack understanding of the subject matter and do not verify it further. It is suggested that users verify information related to facts through other retrieval platforms to ensure its authenticity). Therefore, it would have been obvious to one with ordinary skill in the art before the effective filing date of the invention to have implemented the teaching of wherein classifying the plurality of errors as one of high-severity, mid-severity, or low-severity comprises classifying an error of the plurality of errors as mid-severity by determining that a response appears incorrect and cannot be corrected from Sun into the method as disclosed by the combination of Ouyang, Wu, and Dmitriev. The motivation for doing this is to improve responses from artificial intelligence systems. Claim(s) 10-18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ouyang (NPL: Training language models to follow instructions with human feedback) in view of Cheng et al (US 20260086831), in further view of Wu et al (NPL: Fine-Grained Human Feedback Gives Better Rewards for Language Model Training), in further view of Dmitriev et al (US 20250327687). Regarding claim 10, Ouyang discloses a system comprising: receive a prompt via an artificial intelligence assistant graphical user interface (pg. 6 3.2 Dataset: Our prompt dataset consists primarily of text prompts submitted to the OpenAI API, specifically those using an earlier version of the InstructGPT models (trained via supervised learning on a subset of our demonstration data) on the Playground interface.4 Customers using the Playground were informed that their data could be used to train further models via a recurring notification any time InstructGPT models were used; Fig. 2 e.g. prompt sampled); generate, using a large language model based artificial intelligence assistant, a response to the prompt (Fig. 2 A diagram illustrating the three steps of our method: (1) supervised fine-tuning (SFT), (2) reward model (RM) training, and (3) reinforcement learning via proximal policy optimization (PPO) on this reward model; generates an output); determine an error in the response to the prompt (abstract: For example, large language models can generate outputs that are untruthful, toxic, or simply not helpful to the user; further fine-tune this supervised model using reinforcement learning from human feedback) by: classify the error as a high-severity error rather than a mid-severity error or a low-severity error by determining that the response includes a hallucination (Fig. 2 step 2: ranks outputs from best to worst, wherein the worst would be a high-severity error; abstract: For example, large language models can generate outputs that are untruthful, toxic, or simply not helpful to the user. In other words, these models are not aligned with their users); and Ouyang fails to teach where Cheng teaches one or more memory devices (¶78-81); and one or more processors coupled to the one or more memory devices, the one or more processors (¶78-81) configured to cause the system to: generating, using an annotation tool, an annotated response by modifying one or more of the prompt or the response (¶73 FIG. 5 shows another example user interface 500 used to create metrics though a chat interface. In this example, a first set of interactions, involving a prompt 502 and a response 504; As discussed above, the computer system 110 can attempt to retry generation of metric definitions if an error or validation failure occurs); providing the annotated response to one or more reviewer devices via an error graphical user interface (¶73 FIG. 5 shows another example user interface 500 used to create metrics though a chat interface. In this example, a first set of interactions, involving a prompt 502 and a response 504, shows a successful validation of the newly created metric definition, as indicated by the validation indication 506; By contrast, the second set of interactions, involving a prompt 512 and a response 514, shows a metric definition that failed validation and so shows an error indicator 516 instead of an indication of successful validation); and receiving an indication of the error from the one or more reviewer devices provided via the error graphical user interface (¶73 after a predetermined number of retries, computer system can indicate the error to the user so that the user can change the prompt or otherwise change the metric creation process.). Ouyang further fails to teach where Wu teaches generate, based on the high-severity error, a modification to one or more components of the large language model based artificial intelligence assistant that addresses the high-severity error by prioritizing the high-severity error over errors classified as mid-severity or low-severity (pg. 8 4.5 LM Customization with FINE-GRAINED RLHF: adding more weight to a reward model associated with one specific desired behavior type (e.g., information completeness) may lead the generation more towards that behavior type compared to others; ‘short’ generates more relevant content, but is less factual and complete; (2) ‘long’, in contrast, gives the most factual and complete generation. This reflects that the LM is referencing a large amount of content from passages; (3) The ‘medium’ configuration balances the three rewards; therefore Wu teaches modifying the LLM (e.g. adding more weight) based on prioritizing high-severity error over mid-severity or low-severity (e.g. adding more weight to a reward model associated with one specific desired behavior type, which can be any of the three: short, medium, or long). Ouyang further fails to teach where Dmitriev teaches wherein the one or more components comprise at least one of a prompt rewrite component, an intent detection component, a data quality assurance pipeline, a concepts quality assurance pipeline, an out-of-scope pipeline, a chat history/user session database, a documentation collection, or the artificial intelligence assistant graphical user interface (¶129 In step 415, the output module 309 transmits the map error report to a map feedback system (e.g., via the map feedback API 123); ¶131 the LLM prompting module 305 iteratively modifies the one or more prompts, the one or more additional prompts, or a combination thereof to reduce a number of the one or more questions used during the collection of the one or more data items). Therefore, it would have been obvious to one with ordinary skill in the art before the effective filing date of the invention to have implemented the teaching of generating, using an annotation tool, an annotated response by modifying one or more of the prompt or the response, providing the annotated response to one or more reviewer devices via an error graphical user interface, and receiving an indication of the error from the one or more reviewer devices provided via the error graphical user interface from Cheng, and the teaching of generate, based on the high-severity error, a modification to one or more components of the large language model based artificial intelligence assistant that addresses the high-severity error by prioritizing the high-severity error over errors classified as mid-severity or low-severity from Wu, and the teaching of wherein the one or more components comprise at least one of a prompt rewrite component, an intent detection component, a data quality assurance pipeline, a concepts quality assurance pipeline, an out-of-scope pipeline, a chat history/user session database, a documentation collection, or the artificial intelligence assistant graphical user interface from Dmitriev, into the system as disclosed by Ouyang. The motivation for doing this is to improve techniques to create or update data models, to improve the reduction of false, toxic and other undesired model generation outputs, and further to improve LLM map feedback reporting. Regarding claim 11, the combination of Ouyang, Cheng, Wu, and Dmitriev disclose the system of claim 10, wherein the one or more processors are further configured to provide the prompt and the response to one or more annotation devices of the annotation tool via an annotation graphical user interface (Cheng ¶73 Fig. 5 FIG. 5 shows another example user interface 500 used to create metrics though a chat interface. In this example, a first set of interactions, involving a prompt 502 and a response 504, shows a successful validation of the newly created metric definition, as indicated by the validation indication 506. By contrast, the second set of interactions, involving a prompt 512 and a response 514, shows a metric definition that failed validation and so shows an error indicator 516 instead of an indication of successful validation.). The motivation to combine the references is discussed above in the rejection for claim 10. Regarding claim 12, the combination of Ouyang, Cheng, Wu, and Dmitriev disclose the system of claim 11, wherein the one or more processors are further configured to generate the annotated response by modifying the one or more of the prompt or the response by: generating, via the annotation graphical user interface, error identification annotations (Cheng ¶73 By contrast, the second set of interactions, involving a prompt 512 and a response 514, shows a metric definition that failed validation and so shows an error indicator 516 instead of an indication of successful validation); and associating the error identification annotations with the one or more of the prompt or the response (Cheng ¶73 By contrast, the second set of interactions, involving a prompt 512 and a response 514, shows a metric definition that failed validation and so shows an error indicator 516 instead of an indication of successful validation; thereby this associates the error annotation with the response). The motivation to combine the references is discussed above in the rejection for claim 10. Regarding claim 13, the combination of Ouyang, Cheng, Wu, and Dmitriev disclose the system of claim 10, wherein the one or more processors are further configured to classify the error as a high-severity error rather than a mid-severity error or a low-severity error based on the indication of the error from the one or more reviewer devices (Ouyang Fig. 2 step 2: a labeler ranks outputs from best to worst, wherein the worst would be a high-severity error; abstract: For example, large language models can generate outputs that are untruthful, toxic, or simply not helpful to the user. In other words, these models are not aligned with their users). Regarding claim 14, the combination of Ouyang, Cheng, Wu, and Dmitriev disclose the system of claim 12, wherein the one or more processors are further configured to classify the error as high-severity based on the indication of the error from the one or more reviewer devices by determining that the response includes the hallucination (Ouyang Fig. 2 step 2: a labeler ranks outputs from best to worst, wherein the worst would be a high-severity error; abstract: For example, large language models can generate outputs that are untruthful, toxic, or simply not helpful to the user. In other words, these models are not aligned with their users)., wherein the hallucination comprises at least one of a logical consistency, a persuasive concept, or incorrect data that cannot easily be independently verified (Ouyang abstract: For example, large language models can generate outputs that are untruthful, toxic, or simply not helpful to the user. In other words, these models are not aligned with their users). Regarding claim 15, Ouyang discloses a computer-implemented method comprising: receiving, via one or more graphical user interfaces, a plurality of prompts (pg. 6 3.2 Dataset: Our prompt dataset consists primarily of text prompts submitted to the OpenAI API, specifically those using an earlier version of the InstructGPT models (trained via supervised learning on a subset of our demonstration data) on the Playground interface.4 Customers using the Playground were informed that their data could be used to train further models via a recurring notification any time InstructGPT models were used; Fig. 2 e.g. prompt sampled); generating, using a large language model based artificial intelligence assistant, a plurality of responses to the plurality of prompts (Fig. 2 A diagram illustrating the three steps of our method: (1) supervised fine-tuning (SFT), (2) reward model (RM) training, and (3) reinforcement learning via proximal policy optimization (PPO) on this reward model; generates an output); performing a step for determining a plurality of errors (abstract: For example, large language models can generate outputs that are untruthful, toxic, or simply not helpful to the user; further fine-tune this supervised model using reinforcement learning from human feedback) in the plurality of prompts; performing a step for classifying the plurality of errors as one of high-severity, mid-severity, or low-severity (Fig. 2 labeler ranks the outputs from best to worst; pg. 13 4.2 Results on public NLP datasets we run model samples through the Perspective API8 to obtain automatic toxicity scores, which is the standard evaluation procedure for this dataset, and we also send these samples to labelers to obtain ratings on absolute toxicity, toxicity relative to the prompt, continuity, and overall output preference); and While Ouyang teaches determining a plurality of errors, Ouyang does not specifically teach where Cheng teaches determining a plurality of errors in the plurality of prompts (¶67 If the new metric definition 173 fails one or more validation checks, the computer system 110 can generate an error message that describes the problem (e.g., undefined output, data object name not recognized, ambiguous data object name, etc.). The computer system 110 then generates a new request to the AI/ML model 132 that includes the error message and an instruction to create a new metric definition or correct the earlier metric definition to correct the error; the computer system 110 can perform the validation steps again on the updated metric definition, and can continue to request corrections or re-tries, each time specifying the errors encountered with the most recent version of the metric, until a valid metric is defined or a maximum number of re-try cycles is reached; ¶73 the second set of interactions, involving a prompt 512 and a response 514, shows a metric definition that failed validation and so shows an error indicator 516 instead of an indication of successful validation. As discussed above, the computer system 110 can attempt to retry generation of metric definitions if an error or validation failure occurs. Nevertheless, after a predetermined number of retries, computer system can indicate the error to the user so that the user can change the prompt or otherwise change the metric creation process). Ouyang fails to further teach where Wu teaches generating, based on one or more errors classified as high-severity, a modification to one or more components of the large language model based artificial intelligence assistant that addresses the one or more errors classified as high-severity by prioritizing the one or more errors classified as high-severity over one or more errors classified as mid-severity or low-severity (pg. 8 4.5 LM Customization with FINE-GRAINED RLHF: adding more weight to a reward model associated with one specific desired behavior type (e.g., information completeness) may lead the generation more towards that behavior type compared to others; ‘short’ generates more relevant content, but is less factual and complete; (2) ‘long’, in contrast, gives the most factual and complete generation. This reflects that the LM is referencing a large amount of content from passages; (3) The ‘medium’ configuration balances the three rewards; therefore Wu teaches modifying the LLM (e.g. adding more weight) based on prioritizing high-severity error over mid-severity or low-severity (e.g. adding more weight to a reward model associated with one specific desired behavior type, which can be any of the three: short, medium, or long). Ouyang further fails to teach where Dmitriev teaches wherein the one or more components comprise at least one of a prompt rewrite component, an intent detection component, a data quality assurance pipeline, a concepts quality assurance pipeline, an out-of-scope pipeline, a chat history/user session database, or a documentation collection (¶129 In step 415, the output module 309 transmits the map error report to a map feedback system (e.g., via the map feedback API 123); ¶131 the LLM prompting module 305 iteratively modifies the one or more prompts, the one or more additional prompts, or a combination thereof to reduce a number of the one or more questions used during the collection of the one or more data items). Therefore, it would have been obvious to one with ordinary skill in the art before the effective filing date of the invention to have implemented the teaching of determining a plurality of errors, determining a plurality of errors in the plurality of prompts from Cheng, the teaching of generating, based on one or more errors classified as high-severity, a modification to the large language model based artificial intelligence assistant that addresses one or more errors classified as high-severity by prioritizing the one or more errors classified as high-severity over one or more errors classified as mid-severity or low-severity from Wu, and the teaching of wherein the one or more components comprise at least one of a prompt rewrite component, an intent detection component, a data quality assurance pipeline, a concepts quality assurance pipeline, an out-of-scope pipeline, a chat history/user session database, or a documentation collection from Dmitriev, into the system as disclosed by Ouyang. The motivation for doing this is to improve techniques to create or update data models, to improve the reduction of false, toxic and other undesired model generation outputs, and further to improve LLM map feedback reporting. Regarding claim 16, the combination of Ouyang, Cheng, Wu, and Dmitriev disclose the computer-implemented method of claim 15, wherein determining the plurality of errors in the plurality of prompts comprises generating, for an error and using an annotation tool, a plurality of annotated responses for at least one prompt or a response corresponding to the at least one prompt (Cheng ¶67 If the new metric definition 173 fails one or more validation checks, the computer system 110 can generate an error message that describes the problem (e.g., undefined output, data object name not recognized, ambiguous data object name, etc.). The computer system 110 then generates a new request to the AI/ML model 132 that includes the error message and an instruction to create a new metric definition or correct the earlier metric definition to correct the error; the computer system 110 can perform the validation steps again on the updated metric definition, and can continue to request corrections or re-tries, each time specifying the errors encountered with the most recent version of the metric, until a valid metric is defined or a maximum number of re-try cycles is reached; ¶73 By contrast, the second set of interactions, involving a prompt 512 and a response 514, shows a metric definition that failed validation and so shows an error indicator 516 instead of an indication of successful validation. As discussed above, the computer system 110 can attempt to retry generation of metric definitions if an error or validation failure occurs. Nevertheless, after a predetermined number of retries, computer system can indicate the error to the user so that the user can change the prompt or otherwise change the metric creation process.). The motivation to combine the references is discussed above in the rejection for claim 15. Regarding claim 17, the combination of Ouyang, Cheng, Wu, and Dmitriev disclose the computer-implemented method of claim 16, further comprising generating, using an error analysis mechanism, an indication of the error based on the plurality of annotated responses (Cheng ¶67 If the new metric definition 173 fails one or more validation checks, the computer system 110 can generate an error message that describes the problem (e.g., undefined output, data object name not recognized, ambiguous data object name, etc.). The computer system 110 then generates a new request to the AI/ML model 132 that includes the error message and an instruction to create a new metric definition or correct the earlier metric definition to correct the error; the computer system 110 can perform the validation steps again on the updated metric definition, and can continue to request corrections or re-tries, each time specifying the errors encountered with the most recent version of the metric, until a valid metric is defined or a maximum number of re-try cycles is reached). The motivation to combine the references is discussed above in the rejection for claim 15. Regarding claim 18, the combination of Ouyang, Cheng, Wu, and Dmitriev disclose the computer-implemented method of claim 15, wherein performing the step for classifying the plurality of errors as one of high-severity, mid-severity, or low-severity comprises classifying an error of the plurality of errors as high-severity by determining that a response includes a hallucination (Ouyang Fig. 2 step 2: a labeler ranks outputs from best to worst, wherein the worst would be a high-severity error; abstract: For example, large language models can generate outputs that are untruthful, toxic, or simply not helpful to the user. In other words, these models are not aligned with their users), wherein the hallucination comprises at least one of a logical consistency, a persuasive concept, or incorrect data that cannot easily be independently verified (Ouyang abstract: For example, large language models can generate outputs that are untruthful, toxic, or simply not helpful to the user. In other words, these models are not aligned with their users).. Claim(s) 19-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Ouyang, Cheng, Wu, and Dmitriev as applied to claim 15 above, and further in view of Telling et al (US 20250094439). Regarding claim 19, the combination of Ouyang, Cheng, Wu, and Dmitriev disclose the computer-implemented method of claim 15, but fail to teach where Telling teaches wherein performing the step for classifying the plurality of errors as one of high-severity, mid-severity, or low-severity comprises classifying an error of the plurality of errors as mid-severity by determining that a response comprises at least one of a non-overridable error message or a logical inconsistency (¶99 The error may arise at the level of the prompt and/or at the level of the response from the LLM. For example, the natural language query may be misunderstood by the LLM (e.g., unclear, internally inconsistent, etc.), resulting in an error response; FIG. 7F shows an example error node 794. The error node 794 may indicate a type of error, such as a schema error, a transformation error, a logic error, and/or some other kind of error. The error node 794 can allow a user to troubleshoot the natural language query, the prompt (or another portion thereof, such as described above), the underlying dataset(s), and/or some other aspect of the data pipeline.). Therefore, it would have been obvious to one with ordinary skill in the art before the effective filing date of the invention to have implemented the teaching of wherein performing the step for classifying the plurality of errors as one of high-severity, mid-severity, or low-severity comprises classifying an error of the plurality of errors as mid-severity by determining that a response comprises at least one of a non-overridable error message or a logical inconsistency from Telling into the method as disclosed by the combination of Ouyang, Cheng, Wu, and Dmitriev. The motivation for doing this is to improve optimization of responses from large language models. Regarding claim 20, the combination of Ouyang, Cheng, Wu, and Dmitriev disclose the computer-implemented method of claim 15, but fail to teach where Telling teaches wherein performing the step for classifying the plurality of errors as one of high-severity, mid-severity, or low-severity comprises classifying an error of the plurality of errors as low-severity by determining that a response comprises at least one of information not responsive to a corresponding prompt or an overridable error message (¶99 The error may arise at the level of the prompt and/or at the level of the response from the LLM. For example, the natural language query may be misunderstood by the LLM (e.g., unclear, internally inconsistent, etc.), resulting in an error response; FIG. 7F shows an example error node 794. The error node 794 may indicate a type of error, such as a schema error, a transformation error, a logic error, and/or some other kind of error. The error node 794 can allow a user to troubleshoot the natural language query, the prompt (or another portion thereof, such as described above), the underlying dataset(s), and/or some other aspect of the data pipeline.). Therefore, it would have been obvious to one with ordinary skill in the art before the effective filing date of the invention to have implemented the teaching of wherein performing the step for classifying the plurality of errors as one of high-severity, mid-severity, or low-severity comprises classifying an error of the plurality of errors as low-severity by determining that a response comprises at least one of information not responsive to a corresponding prompt or an overridable error message from Telling into the method as disclosed by the combination of Ouyang, Cheng, Wu, and Dmitriev. The motivation for doing this is to improve optimization of responses from large language models. Allowable Subject Matter Claim 8 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The following is a statement of reasons for the indication of allowable subject matter: Regarding claim 8, the prior art of record, alone or in combination, does not teach “wherein generating the modification to the one or more components comprises generating, using a prompt improvement engine and based on the high-severity error, a modification to at least one of the prompt rewrite component or the intent detection component to modify prompts used by the large language model based artificial intelligence assistant to generate responses”. Response to Arguments Applicant’s arguments with respect to claim(s) 1-7 and 9-20 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to KEVIN KY whose telephone number is (571)272-7648. The examiner can normally be reached Monday-Friday 9-5PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vincent Rudolph can be reached at 571-272-8243. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /KEVIN KY/Primary Examiner, Art Unit 2671
Read full office action

Prosecution Timeline

Show 5 earlier events
May 28, 2026
Response Filed
Jun 11, 2026
Final Rejection mailed — §103
Aug 10, 2026
Interview Requested
Aug 17, 2026
Applicant Interview (Telephonic)
Aug 17, 2026
Examiner Interview Summary
Aug 19, 2026
Request for Continued Examination
Aug 24, 2026
Response after Non-Final Action
Sep 22, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743811
SYSTEMS AND METHODS FOR EFFICIENT OBJECT TRACKING AS A SERVICE VIA EDGE
4y 3m to grant Granted Sep 22, 2026
Patent 12743906
SEMANTICALLY GENERATED VIDEO METHOD AND SYSTEM
1y 0m to grant Granted Sep 22, 2026
Patent 12738276
DISPLAY APPARATUS CAPABLE OF RELEASING A VOICE INPUT MODE BY SENSING A SPEECH FINISH AND VOICE CONTROL METHOD THEREOF
2y 11m to grant Granted Sep 15, 2026
Patent 12737934
HARDWARE-AWARE EFFICIENT ARCHITECTURES FOR TEXT-TO-IMAGE DIFFUSION MODELS
2y 10m to grant Granted Sep 15, 2026
Patent 12737414
CONVERTING VIDEO SEMANTICS INTO LANGUAGE FOR REAL-TIME QUERY AND INFORMATION RETRIEVAL
2y 9m to grant Granted Sep 15, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
77%
Grant Probability
99%
With Interview (+25.2%)
2y 6m (~6m remaining)
Median Time to Grant
High
PTA Risk
Based on 579 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month