Prosecution Insights
Last updated: October 02, 2026
Application No. 18/436,229

INTELLIGENT STEWARD PLATFORM FOR VALIDATION OF LARGE LANGUAGE MODEL (LLM) OUTPUTS

Non-Final OA §102§103§112
Filed
Feb 08, 2024
Examiner
HOANG, MICHAEL H
Art Unit
Tech Center
Assignee
Bank of America Corporation
OA Round
1 (Non-Final)
55%
Grant Probability
Moderate
1-2
OA Rounds
1y 8m
Est. Remaining
78%
With Interview

Examiner Intelligence

Grants 55% of resolved cases
55%
Career Allowance Rate
85 granted / 155 resolved
-5.2% vs TC avg
Strong +23% interview lift
Without
With
+23.1%
Interview Lift
resolved cases with interview
Typical timeline
4y 4m
Avg Prosecution
30 currently pending
Career history
172
Total Applications
across all art units

Statute-Specific Performance

§101
28.5%
-11.5% vs TC avg
§103
45.7%
+5.7% vs TC avg
§102
10.9%
-29.1% vs TC avg
§112
12.5%
-27.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 155 resolved cases

Office Action

§102 §103 §112
DETAILED ACTION This action is in response to the claims filed 02/08/2024 for Application number 18/436,229. Claims 1-20 are currently pending. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statements (IDS) submitted on 02/08/2024, 10/06/2025, 02/18/2026 and 08/13/2026 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statements are being considered by the examiner. Claim Objections Claim 4 and 15 are objected to because of the following informalities: "an non-acceptable" should read "a non-acceptable". Appropriate correction is required. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 1-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. The terms “acceptable, tolerable, or non-acceptable” in claims 1, 12, and 20 are relative terms which renders the claim indefinite. The terms “acceptable, tolerable, or non-acceptable” are not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. The specification only discloses “comparing more recent information from social media, literature, texts, images, voices, or the like to more prevalent and/or existing social standards to decide what is acceptable and what is non-acceptable as a response.”[¶0022] and “defining information that is in a first category (e.g., “acceptable’), a second category (e.g., “non-acceptable”), and/or a third category (e.g., “tolerable”).” [¶0035] while not defining what is considered to be “acceptable, tolerable, or non-acceptable”. Since these terms are subjective definitions which differ from person to person, the metes and bounds of the claim is not made clear and one of ordinary skill in the art would not be able to properly avoid infringing upon a claim when no definition of acceptable, tolerable, or non-acceptable has been made. For purposes of examination, the examiner will interpret “acceptable, tolerable, or non-acceptable” as any three distinct classifications/categories. Claims 2-11 and 13-19 are rejected as being dependent on a rejected base claim without curing any of the deficiencies. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claims 1-8 and 12-20 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Gupta et al. ("US 20250124236 A1", cited by Applicant in the IDS filed 10/06/2025, hereinafter "Gupta"). Regarding claim 1, Gupta teaches A computing platform comprising: at least one processor (¶0081); a communication interface communicatively coupled to the at least one processor (¶0082-¶0083); and memory storing computer-readable instructions that, when executed by the at least one processor (¶0083), cause the computing platform to: train, using historical information indicating a plurality of regimes for large language model (LLM) outputs (“ The query processing module 335 receives and processes queries that access data stored by the data storage system 110…The inference tasks may include, but are not limited to, natural language processing (NLP) tasks, audio processing tasks, image processing tasks, video processing tasks, and the like. The NLP tasks may include, but are not limited to, text generation, query processing, machine translation, chatbot applications, and the like” [¶0042-¶0043]), an LLM steward model, wherein training the LLM steward model configures the LLM steward model to generate LLM validation information indicating classifications of LLM outputs as acceptable, tolerable, or non-acceptable (“The evaluation module 420 receives the set of generated responses from the inference module 410 and evaluates the set of generated responses with respect to an evaluation objective. In one instance, the evaluation module 420 implements an evaluation LLM (“LLM steward model”) which applies a selected evaluation function to each of the generated responses.” [¶0046; evaluation functions correspond to a plurality of regimes. See further ¶0061, “Hallucinations may also include factual inaccuracies and improbable scenarios. The evaluator LLM uses contextual information provided to the LLM to identify hallucinated portions of the generated responses. In another embodiment, the evaluator LLM compares the generated responses to an expected output.” (The evaluator LLM determines if an output is acceptable, tolerable or non-acceptable.)]); input, into an LLM, an LLM prompt, wherein inputting the LLM prompt causes the LLM to generate an LLM output (“The inference module 410 receives a prompt from a user of a client device 116 and provides the prompt to the one or more selected LLMs deployed in the model serving system(s) 170 to perform inference. The LLMs generate the response to the prompt from the knowledge that the LLM was trained on and/or from the contextual information included in the prompt.” [¶0045]); input the LLM output into the LLM steward model, wherein inputting the LLM output into the LLM steward model causes the LLM steward model to output the LLM validation information (“The evaluation module 420 receives the set of generated responses from the inference module 410 and evaluates the set of generated responses with respect to an evaluation objective.” [¶0046]); based on outputting LLM validation information indicating that the LLM output is acceptable or tolerable, send the LLM output to a user device for presentation (“As described in FIG. 4, the evaluation module 420 implements an evaluation LLM which applies a selected evaluation function to each of the generated responses. Some examples of evaluation functions include, and are not limited to, keyword similarity, toxicity detection, and hallucination detection. The UI element generator module 430 generates one or more user interface elements on the UI to display the results of the evaluation for the generated responses.” [¶0070]); based on outputting LLM validation information indicating that the LLM output is non-acceptable, update the LLM output to conform with a corresponding subset of the plurality of regimes (“For example, for the keyword similarity evaluation function, the UI element generator module 430 receives the position of keywords found in the expected output and the position of keywords found in each of the generated responses of the LLMs. The UI element generator module 430 generates one or more UI elements which highlight the identified keywords. For example, keywords found in both the expected output and the generated response are highlighted green to indicate a match. In embodiments where one or more keywords present in the expected output are not present in the generated response, the UI element generator module 430 generates a UI element that highlights the missing keyword red.” [¶0070-¶0071]); and update, via a dynamic feedback loop and based on feedback received from the user device, the LLM steward model. (“The UI includes a feedback mechanism 550a, 550b which allows users to indicate the quality of the generated response.” [¶0076]) Regarding claim 2, Gupta teaches The computing platform of claim 1, wherein the historical information includes one or more of: text information, images, speech information, structured information, three dimensional signals, literature information, cultural information, social information, geographical information, legal information, or linguistic information. (¶0025; discloses text/image data for prompts/responses, note: The claim recites “one or more of” thus under BRI the examiner is only required to map to one of the recited elements) Regarding claim 3, Gupta teaches The computing platform of claim 1, wherein each of the regimes define content that, when included in an output from the LLM, is one or more of: acceptable, tolerable, or non-acceptable. (“As described in FIG. 4, the evaluation module 420 implements an evaluation LLM which applies a selected evaluation function to each of the generated responses. Some examples of evaluation functions include, and are not limited to, keyword similarity, toxicity detection, and hallucination detection.” [¶0070]) Regarding claim 4, Gupta teaches The computing platform of claim 1, wherein outputting the LLM validation information comprises: identifying one or more regimes, of the plurality of regimes, associated with the LLM prompt (“As described in FIG. 4, the evaluation module 420 implements an evaluation LLM which applies a selected evaluation function to each of the generated responses.” [¶0070]), identifying a location of the LLM output, within the one or more regimes associated with the LLM prompt (“The UI element generator module 430 receives the characterizing data (e.g., position) of each of the identified words in the generated responses from the one or more LLMs and generates a UI element on the UI which highlights the identified words.” [¶0070]), based on identifying that the LLM output is within an acceptable regime or a tolerable regime, outputting an indication that the LLM output is acceptable (“A predetermined toxicity threshold can be used to determine if a generated response is toxic or non-toxic. For example, a response having a toxicity score less than the toxicity threshold may be considered non-toxic. In an embodiment, the evaluator LLM classifies the responses as toxic or non-toxic.” [¶0060]), and based on identifying that the LLM output is within an non-acceptable regime, outputting an indication that the LLM output is non-acceptable. (“A predetermined toxicity threshold can be used to determine if a generated response is toxic or non-toxic. For example, a response having a toxicity score less than the toxicity threshold may be considered non-toxic. In an embodiment, the evaluator LLM classifies the responses as toxic or non-toxic.” [¶0060]) Regarding claim 5, Gupta teaches The computing platform of claim 4, wherein the LLM steward model comprises a foundational model, and wherein identifying the one or more regimes associated with the LLM prompt comprises: identifying a plurality of overlapping clusters, within the foundational model, that characterize the LLM prompt, and identifying regimes corresponding to the plurality of overlapping clusters. (“In one embodiment, the query processing module 335 provides one or more queries to appropriate clusters of the data layer 108, and receives responses to the queries from clusters in which the queries are executed.” [¶0042; these clusters would be overlapping clusters as Gupta’s evaluation functions (similarity detection, toxicity detection, etc.) would share similar resources such as characterizing data]) Regarding claim 6, Gupta teaches The computing platform of claim 5, wherein the plurality of overlapping clusters are identified based on an internet protocol (IP) address of a user submitting the LLM prompt. (“The system environment 100 shown by FIG. 1 includes one or more client devices 116A, 116B, a network 120, a data processing service 102, and a data storage system 110. In alternative configurations, different and/or additional components may be included in the system environment 100.” [¶0016; IP address is inherent given client devices connected to a network]) Regarding claim 7, Gupta teaches The computing platform of claim 1, wherein outputting the LLM validation information comprises: generating a confidence score indicating a confidence that the LLM output is acceptable or non-acceptable (“The evaluator LLM may calculate a toxicity score for each of the generated responses.” [¶0060]), comparing the confidence score to a confidence threshold, based on identifying that the confidence score meets or exceeds the confidence threshold (“A predetermined toxicity threshold can be used to determine if a generated response is toxic or non-toxic. For example, a response having a toxicity score less than the toxicity threshold may be considered non-toxic.” [¶0060]), outputting the LLM validation information (“In an embodiment, the evaluator LLM classifies the responses as toxic or non-toxic. For toxic responses, the evaluator LLM determines characterizing data (e.g., position) of the toxic words.” [¶0060]), and based on identifying that the confidence score fails to meet or exceed the confidence threshold, sending a request to the user device for additional information for use in updating the confidence score. (See ¶0075-¶0077, discloses requesting user feedback to update the evaluation process) Regarding claim 8, Gupta teaches The computing platform of claim 1, wherein the LLM corresponds to a chatbot. (“The NLP tasks may include, but are not limited to, text generation, query processing, machine translation, chatbot applications, and the like.” [¶0043]) Regarding claim 12, it is substantially similar to claim 1 respectively, and is rejected in the same manner, the same art, and reasoning applying. Regarding claims 13-19, they are substantially similar to claims 2-8 respectively, and are rejected in the same manner, the same art, and reasoning applying. Claim 20 recites features similar to claim 1 and is rejected for at least the same reasons therein. Claim 20 additionally requires One or more non-transitory computer-readable media storing instructions that, when executed by a computing platform comprising at least one processor, a communication interface, and memory, cause the computing platform to… (¶0083-¶0084, Gupta) Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Gupta in view of Dong et al. ("Building Guardrails for Large Language Models", hereinafter "Dong"). Regarding claim 9, Gupta teaches The computing platform of claim 1, wherein the memory stores additional computer readable instructions that, when executed by the at least one processor, cause the computing platform to: receive updated information associated with the plurality of regimes (“As described in FIG. 4, the evaluation module 420 implements an evaluation LLM which applies a selected evaluation function to each of the generated responses. Some examples of evaluation functions include, and are not limited to, keyword similarity, toxicity detection, and hallucination detection…In addition, the UI includes a bar chart UI element 542 positioned at the top of each column associated with an LLM. The bar chart UI element 542 shows a visual summary of the evaluation result.” [¶0070-0075]); However Gupta fails to explicitly teach identify a delta value between the historical information and the updated information; and update, based on the delta value, the plurality of regimes to adjust corresponding classifications of acceptable, tolerable, or non-acceptable. Dong teaches identify a delta value between the historical information and the updated information (“Alongside this, algorithmic adjustments are necessary, which involve fine-tuning the model’s parameters to prevent the overemphasis of certain patterns that could lead to biased outcomes… It is however expected that the definition will be distribution-based, rather than point-based as unintended responses, which need to estimate posterior distributions and to measure the distance between two distributions. (“delta value”)” [pg. 5, bottom left col – top right col]); and update, based on the delta value, the plurality of regimes to adjust corresponding classifications of acceptable, tolerable, or non-acceptable. (“Here, the Guardrails AI can automatically generate a corrective prompt, pursuing the LLMs to regenerate the correct answer. The output is then re-checked to ensure it meets the specified requirements.” [pg. 3, top left col]) It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Gupta’s teachings by building Guardrails for LLMs as taught by Dong. One would have been motivated to make this modification in order to ensure that the generated output response meets its specified requirements. [pg. 3, top left col, Dong] Claims 10 and 11 are rejected under 35 U.S.C. 103 as being unpatentable over Gupta in view of Dong and further in view of Lee et al. ("Towards Reliable and Fluent Large Language Models: Incorporating Feedback Learning Loops in QA Systems", hereinafter "Lee"). Regarding claim 10, Gupta/Dong teaches The computing platform of claim 9, however fails to explicitly teach wherein the LLM steward model is a closed loop model, and wherein updating the plurality of regimes comprises updating an additional model that is dynamically updated, wherein the additional model is a layer added on top of the LLM steward model Lee teaches wherein the LLM steward model is a closed loop model, and wherein updating the plurality of regimes comprises updating an additional model that is dynamically updated (“2) Using the constructed data, we propose a method for training critic models to generate feedback (corresponds to “closed loop”) on heterogeneous aspects of generated text.” [pg. 1, right col]), wherein the additional model is a layer added on top of the LLM steward model. (“We conduct iterative feedback learning using Chat GPT (gpt-3.5-turbo)as the baseline LLM. Then, the base line LLM is improved with iterative feedback learning. The proposed method for iterative feedback learning aligns with Algorithm 1” [pg. 3, 3.3., Critic Model is used to evaluate generated answers thus would be an added layer on top of the baseline LLM]) It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Gupta’s/Dong’s teachings in order to use a closed loop/feedback model as taught by Lee. One would have been motivated to make this modification as training a critic model allows evaluation of the generated answer and identify any deficiencies in the model’s response and provides a negative signal for inadequate responses. [Lee, pg. 3, §3.2] Regarding claim 11, Gupta/Dong/Lee teaches The computing platform of claim 10, where Lee teaches wherein subsequent LLM outputs are fed through both the LLM steward model and the additional model. (See algorithm 1, pg. 3, “For each iteration do 1. Generate answer using LLM: A response to a given question is created using instruct tions and one-shot examples, allowing for the generation of citable references. 2. Evaluate generated answer with critic model:”) Same motivation to combine the teachings of Gupta/Dong/Lee as claim 10. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHAEL H HOANG whose telephone number is (571)272-8491. The examiner can normally be reached Mon-Fri 8:30AM-4:30PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kakali Chaki can be reached at (571) 272-3719. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MICHAEL H HOANG/ PRIMARY EXAMINER, Art Unit 2122
Read full office action

Prosecution Timeline

Feb 08, 2024
Application Filed
Aug 20, 2026
Non-Final Rejection mailed — §102, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748950
REVERSE REINFORCEMENT LEARNING TO TRAIN TRAINING DATA FOR NATURAL LANGUAGE PROCESSING NEURAL NETWORK
4y 3m to grant Granted Sep 29, 2026
Patent 12730800
SYSTEM AND METHOD FOR IMPLEMENTING INTELLIGENT SERVICE REQUEST REMEDY
5y 8m to grant Granted Sep 08, 2026
Patent 12731069
FRACTAL RELATIONSHIPS FOR TRAINING ARTIFICIAL INTELLIGENCE CLASSIFIER
4y 8m to grant Granted Sep 08, 2026
Patent 12711428
ARCHITECTURE-AGNOSTIC FEDERATED LEARNING SYSTEM
4y 2m to grant Granted Aug 18, 2026
Patent 12705488
GRAPH STRUCTURE AWARE INCREMENTAL LEARNING FOR RECOMMENDER SYSTEM
3y 5m to grant Granted Aug 11, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
55%
Grant Probability
78%
With Interview (+23.1%)
4y 4m (~1y 8m remaining)
Median Time to Grant
Low
PTA Risk
Based on 155 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month