Prosecution Insights
Last updated: October 02, 2026
Application No. 18/607,060

AUGMENTING A LARGE LANGUAGE MODEL WITH AN EXTERNAL MODEL

Non-Final OA §103
Filed
Mar 15, 2024
Examiner
CARDOSO, JUSTIN ALEXANDER
Art Unit
Tech Center
Assignee
International Business Machines Corporation
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
11 currently pending
Career history
7
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§103
DETAILED ACTION This action is in response to the original filing filed on 03/15/2024. Claims 1-20 are pending examination. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Objections Claims 7, 12, and 17 are objected to because of the following informalities: claims recite “the first value vector comprising the another response to the question”. “The another” is grammatically incorrect. Appropriate correction is required. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Hartvigsen et al. (‘Aging with GRACE: Lifelong Model Editing with Discrete Key-Value Adaptors’, published October 18, 2023, hereinafter Hartvigsen) in view of Mitchell et al. (‘Memory-Based Model Editing at Scale’, published June 13, 2022, hereinafter Mitchell). Regarding Claim 1, Hartvigsen teaches the computer-implemented method ([Abstract, Pg. 1] "We propose GRACE, a lifelong model editing method, which implements spot-fixes on streaming errors of a deployed model, ensuring minimal impact on unrelated inputs."; GRACE is a lifelong model editing method executed on a deployed pre-trained model), including: receiving, at a first language machine learning model, a question; and ([Section 1, Page 2, Bullet 3] Our experiments show that GRACE outperforms seven alternatives when sequentially editing T5, BERT, and GPT models for question answering, document classification, and language generation. (The deployed T5 language model (a language machine learning model) receives question-answering inputs from the dataset (a question) at inference)) providing a response to the question and confidence determination for the response that exceeds a predetermined threshold ([Page 3, Section 2.2] Each key has a deferral radius ϵ, which serves as a threshold for similarity matching. (The deferral decision performs a similarity search of the layer input against the codebook keys, and satisfying the deferral constraint against the deferral radius ϵ (a predetermined threshold), that is, a similarity match exceeding the ϵ-defined level, is the confidence determination (a confidence determination) governing whether the retrieved value (the response) is applied)). injecting the response into the first language machine learning model so that the response overwrites a vector state layer output of the first language machine learning model that provides another response to the question ([Section 2.2 Pg. 4 and Pg. 3, Equation (1)] The learned value v then replaces hl for the rest of the forward pass. (The retrieved value v (the response) is injected at the wrapped layer l of the deployed pre-trained model (the first language machine learning model) and replaces the layer-l activation hl (a vector state layer output), which per Equation (1) would otherwise be the unedited model's activation fl(hl−1) that the model propagates to produce its original, uncorrected prediction (another response to the question))) and without modifying original parameters of the first language machine learning model ([Abstract Pg. 1] GRACE writes new mappings into a pre-trained model's latent space, creating a discrete, local codebook of edits without altering model weights. (The weights of the pre-trained model (original parameters) are left unaltered, all edits residing in the Adaptor codebook)). Hartvigsen does not teach: an external machine learning model, and the response to the question being provided by the external machine learning model, wherein the external machine learning model was trained with training material with which the first language machine learning model was not trained. In the same field of endeavor, Mitchell teaches: in response to an external machine learning model ([Sheet 2, Section 1] Building on the hypothesis that gradients are an impoverished signal for model editing, we propose SERAC, a gradient-free memory-based approach to model editing. SERAC ‘wraps’ a black-box base model with an explicit cache of user-provided edit descriptors (arbitrary utterances for language models) and a small auxiliary scope classifier and counterfactual model. Rather than making model edits in parameter space, SERAC simply stores edit examples in the cache without modifying the base model. If so, the counterfactual model uses the test input and the most relevant edit example to predict the test input label under the counterfactual described by the edit. (The counterfactual model hψ of SERAC is an external model which is applied to a separate model to apply edits) and the response to the question being provided by the external machine learning model ([Sheet 2, Section 1] If so, the counterfactual model uses the test input and the most relevant edit example to predict the test input label under the counterfactual described by the edit. (The counterfactual model hψ (an external machine learning model), a sequence model auxiliary to the frozen base model fbase, receives the post-edit test input (the question) and produces the predicted label (a response to the question) that is returned as the system output per Equation (1) at Sheet 4, Section 3.1)) wherein the external machine learning model was trained with training material with which the first language machine learning model was not trained ([Sheet 4, Section 3.2] A SERAC editor is trained using the edit dataset De = {zi e}, where in-scope examples (xi in, yi in) and negative examples xi out are sampled from I(zi e; De) and O(zi e; De), respectively. (The counterfactual model (the external machine learning model) is trained on the edit dataset De of user-provided edit descriptors (training material), corrective information held outside the frozen base model fbase (the first language machine learning model), which was not trained with De)) It would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to have incorporated the concepts of an external learning model providing a response and wherein the external model was trained with material not available to the first model as found in Mitchell into those found in Hartvigsen as both articles are in the same field of implementing similar model-editing techniques where an external system “injects” new information into an already existing, pre-trained model. Doing so would provide obvious benefits, as this modification would allow for updating an already deployed language model with new information while keeping its weights unchanged and to generalize each correction beyond memorized inputs to related phrasing of the same query (Mitchell Sheet 1, Section 1 and Sheet 3, Section 3). Regarding Claim 2, the combination of Hartvigsen and Mitchell teach all of the limitations of Claim 1, including: wherein the response of the external machine learning model is injected between a first layer and a second layer of a feed-forward network (FFN) of the first language machine learning model. ([Hartvigsen Figure 1:] Overview of lifelong model editing with GRACE. a) Models make important errors that must be corrected. b) GRACE makes edits by learning, caching, and selectively retrieving new transformations between layers. [Page 2-3 Section 2] Let f0 denote a model with frozen parameters that was pre-trained on dataset Dtrain. For clarity, we drop the subscript where possible. Assume that f consists of L layers, where fl(·) computes the hidden state at layer l. In this work, we assume f is a transformer architecture for natural language, though our principles are general. (GRACE works by injecting responses in between layers of a model. Furthermore, GRACE is generalized and model agnostic, so can be applied to any language model. One example application is found in Section 2, where GRACE is applied to a model with a transformer architecture for natural language. These types of transformer models consist of component transformer layers, which themselves consist of alternating attention and feedforward layers. Since GRACE can be applied to any chosen layer (Page 3 Section 2.2 ), it follows that GRACE can be applied to be used in an FFN found inside a transformer architecture. (See Page 4 Section "GRACE layers with sequential inputs" for expansion on application to transformer models.)) Regarding Claim 3, the combination of Hartvigsen and Mitchell teach all of the limitations of Claim 2, including: wherein the first layer is a key layer, and wherein the second layer is a value layer. ([Hartvigsen Section 2.2 Pg. 3] We propose GRACE, a method for sequentially editing a pre-trained model’s behavior without altering its weights, as illustrated in Figure 1. GRACE works by wrapping a chosen layer of any pre-trained model architecture with an Adaptor. To make an edit, a GRACE layer can perform one of two operations. Each step is described in Algorithm 1. First, if the codebook is empty or the input embedding hl−1 falls outside the deferral radius of all existing keys according to distance function d(·), then a new codebook entry is created and added: {(hl−1, v, ϵinit, y)}. Thus if xt were passed into f again, hl−1 would activate the codebook and value v would be passed to layer l + 1. [Section 2.2 Pg 4, Training GRACE Values] When making an edit with GRACE, either a new key–value pair is learned or an existing key–value pair is updated. (GRACE wraps a chosen layer and updates a key-value pair. See Figure 1, where GRACE activates between Layer l and Layer l-1. Further, the section Codebook maintenance delves into detail on GRACE passing a value to a layer given a key, showing that it must work between a key and value layer.)) Regarding Claim 4, the combination of Hartvigsen and Mitchell teach all of the limitations of Claim 1, including: wherein the external machine learning model is smaller in size than the first language machine learning model. ([Mitchell Section 3.1, Sheet 3] SERAC can be thought of as a simple wrapper around the base model. (The SERAC system is a simple wrapper around the base LLM. As such, it is smaller in size and less complex.)) Regarding Claim 5, the combination of Hartvigsen and Mitchell teach all of the limitations of Claim 1, including: preparing a corpus; ([Mitchell Section 2, Sheet 3] Because the ‘correct’ scope of an edit’s effects on the base model may be unknown or ambiguous, we train a model editor on a dataset of edits) building an initial external machine learning model; ([Mitchell Section 3.1, Sheet 3] It is made up of three key components: an explicit cache of edits, an edit scope classifier, and a counterfactual model that ‘overrides’ the base model when necessary. (The built architecture of SERAC)) training the initial external machine learning model with the corpus; and ([Mitchell Section 3.2, Sheet 4] Similarly to past work (De Cao et al., 2021; Mitchell et al., 2021; Hase et al., 2021), a SERAC editor is trained using the edit dataset (SERAC is trained)) building the external machine learning model from the initial external machine learning model, wherein inputs and outputs of the external machine learning model are embeddings. ([Mitchell Section 3.1, Sheet 4] We primarily opt for a more computationally-efficient approach, first computing separate, fixed-length embeddings of the input and edit descriptor (as in Karpukhin et al., 2020) and using the negative squared Euclidean distance in the embedding space as the predicted log-likelihood. [Section 5.1 Sheet 7] For simplicity, we default to the embedding-based classifier for SERAC (The embedding based input/output architecture of SERAC is expanded upon)). It would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to have incorporated the concepts of preparing a corpus, building an initial external model, training the model, and building the external model from the initial model, where the models inputs and outputs are embeddings as taught by Mitchell into those found in Hartvigsen as both articles are in the same field of implementing similar model-editing techniques where an external system “injects” new information into an already existing, pre-trained model. Doing so would provide obvious benefits, as this modification would allow for updating an already deployed language model with new information while keeping its weights unchanged and to generalize each correction beyond memorized inputs to related phrasing of the same query (Mitchell Sheet 1, Section 1 and Sheet 3, Section 3). Regarding Claim 6, the combination of Hartvigsen and Mitchell teach all of the limitations of Claim 5, including: wherein the building of the external machine learning model comprises removing an embedding layer and an output layer of the initial external machine learning model ([Mitchell Sheet 13, Appendix, Section B.] For all gradient-based methods, we adapt the fully connected layers of the last 3 transformer blocks for encoder only models, and fully-connected layers in the last 2 transformer blocks of both encoder and decoder for encoder-decoder models. [Appendix, Section B. Editable Neural Networks (ENN)] For T5, we edit only the last two layers of both the encoder and the decoder. For BERT-base, we edit the last two layers of the encoder. Finally, for BlenderBot-small, we edit the last layer of the encoder and the last three layers of the decoder since the decoder is much deeper. (In the cited example, SERAC modifies only the final two layers of a BERT-base encoder, specifically the transformer blocks. This means it does not modify the prior embedding layer nor the subsequent output layers.)) such that the external machine learning model does not receive language tokens and does not output natural language tokens ([Mitchell Section 2, Sheet 2] In this work, the edit descriptor may be a concatenated input-output pair [xe;ye] like WHO IS THE UK PM? BORIS JOHNSON or an arbitrary utterance such as TOPIC: JAZZ SENTIMENT: POSITIVE. (The input/outputs of SERAC are a predefined pair without any tokens. SERAC in fact does not take tokens as input, only needing its applied-to models' tokenization.)) It would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to have incorporated the concepts of building the external model comprising removing an embedding layer and an output layer and the external model not receiving language tokens and not outputting language tokens as taught by Mitchell into those found in Hartvigsen as both articles are in the same field of implementing similar model-editing techniques where an external system “injects” new information into an already existing, pre-trained model. Doing so would provide obvious benefits, as this modification would allow for updating an already deployed language model with new information while keeping its weights unchanged and to generalize each correction beyond memorized inputs to related phrasing of the same query (Mitchell Sheet 1, Section 1 and Sheet 3, Section 3). Regarding Claim 7, the combination of Hartvigsen and Mitchell teach all of the limitations of Claim 1, including: wherein the external machine learning model produces the confidence determination via: obtaining a first value vector from the first language machine learning model, the first value vector comprising the another response to the question; performing a dot product operation of the first value vector and a first external value vector that represents the response of the external machine learning model such that a dot product value is produced; and ([Mitchell Section 4, Sheet 6] We measure sentiment accuracy with the rescaled likelihood ratio zsent σ(l+ e − l− e), where l+ and l− are the average per-token log likelihood of the edited model on pre-generated on-topic responses with the correct sentiment (either all positive or all negative) and incorrect sentiment, respectively, and σ is the sigmoid function. (Obtaining a value of measured correctness)) multiplying the dot product value against an internal degree of confidence that the external machine learning model produces for accuracy of the response. ([Mitchell Section 4, Sheet 6] We measure edit success with the product of zsent and ztopic; which can be very roughly interpreted as ‘the likelihood that the edited model produces the desired sentiment and is on-topic for in-scope inputs (Zsent is the measured value of sentiment accuracy, ztopic is the measure value of topical accuracy. Finding the product measures how successful the edit (confidence in answer) was.)) Regarding Claim 8, the combination of Hartvigsen and Mitchell teach all of the limitations of Claim 1, including: wherein the confidence determination comprises an internal degree of confidence that the external machine learning model produces for accuracy of the response. ([Hartvigsen Section 2.2 Pg 3] Each key has a deferral radius ϵ, which serves as a threshold for similarity matching. The deferral mechanism uses this radius as shown in Algorithm 1. GRACE is activated at layer l only if the deferral constraint is satisfied. (Referring again to the deferral radius, GRACE uses similarity scoring to determine if a given response is accurate)) Regarding Claim 9, the combination of Hartvigsen and Mitchell teach all of the limitations of Claim 1, including: wherein the training material used to train the external machine learning model is new material that was unavailable at a time of training the first language machine learning model. ([Mitchell Sheet 4, Section 3.2] A SERAC editor is trained using the edit dataset De = {zi e}, where in-scope examples (xi in, yi in) and negative examples xi out are sampled from I (zi e; De) and O (zi e; De), respectively. (the counterfactual model (the external machine learning model) is trained on the edit dataset De of user-provided edit descriptors (training material), corrective information held outside the frozen base model fbase (the first language machine learning model), which was not trained with De)) It would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to have incorporated the concept of requiring the training material of the external model to not have been available at the time of training the initial model as taught by Mitchell into those found in Hartvigsen as both articles are in the same field of implementing similar model-editing techniques where an external system “injects” new information into an already existing, pre-trained model. Doing so would provide obvious benefits, as this modification would allow for updating an already deployed language model with new information while keeping its weights unchanged and to generalize each correction beyond memorized inputs to related phrasing of the same query (Mitchell Sheet 1, Section 1 and Sheet 3, Section 3). Regarding Claim 10, the combination of Hartvigsen and Mitchell teach all of the limitations of Claim 1, including: producing, via the first language machine learning model, a first embedding that represents the question; and ([Mitchell Section 4, Sheet 5] To generate hard out-of-scope examples for an edit input xe, we selectively sample from training inputs x that have high semantic similarity with xe, measured as having a high cosine similarity between their embeddings as computed by a pre-trained semantic embedding model all-MiniLM-L6-v2 [Section 5.2 Pg. 8] Cross-attention is especially useful for the FC experiment, which is possibly due to the commonness of quantities in the VitaminC dataset; for example, producing fixed-length sequence embeddings that reliably capture the difference between THERE HAVE BEEN 105,000 CORONAVIRUS DEATHS IN THE UNITED STATES and THERE HAVE BEEN 111,000 CORONAVIRUS DEATHS IN THE UNITED STATES may be very difficult. (Representation of inputs as embeddings. Furthermore, see Fig. 2 for a representation of selected questions (Where is Boris Johnson the PM?; Who is the PM of the UK?; Who is the prime minister of the UK?) in a semantic embedding space)) transmitting the first embedding from the first language machine learning model to the external machine learning model, ([Hartvigsen Section I. Introduction, Pg. 2] As illustrated in Figure 1, GRACE edits a model by adding an Adaptor to a chosen layer, while never changing its weights. This Adaptor then modifies layer-to-layer transformations for select inputs. By caching embeddings for input errors and learning values that decode into desired model outputs, GRACE serves as a codebook in which edits are stored, enabling longer sequences of edits than prior works. (GRACE caches embeddings for input errors from its parent/attached-to model)) wherein the external machine learning model generates the response and ([Mitchell: Sheet 2, Section 1] If so, the counterfactual model uses the test input and the most relevant edit example to predict the test input label under the counterfactual described by the edit. (The counterfactual model hψ (an external machine learning model), a sequence model auxiliary to the frozen base model fbase, receives the post-edit test input (the question) and produces the predicted label (a response to the question) that is returned as the system output per Equation (1) at sheet 4, Section 3.1), Generation of a response)) the confidence determination based on an analysis of the first embedding. ([Hartvigsen: page 3, Section 2.2] Each key has a deferral radius ϵ, which serves as a threshold for similarity matching. (The deferral decision performs a similarity search of the layer input against the codebook keys, and satisfying the deferral constraint against the deferral radius ϵ (a predetermined threshold), that is, a similarity match exceeding the ϵ-defined level, is the confidence determination (a confidence determination) governing whether the retrieved value (the response) is applied)) It would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to have incorporated the concepts of producing a first embedding representing a question and the external model generating a response as taught by Mitchell into those found in Hartvigsen as both articles are in the same field of implementing similar model-editing techniques where an external system “injects” new information into an already existing, pre-trained model. Doing so would provide obvious benefits, as this modification would allow for updating an already deployed language model with new information while keeping its weights unchanged and to generalize each correction beyond memorized inputs to related phrasing of the same query (Mitchell Sheet 1, Section 1 and Sheet 3, Section 3). Regarding Claims 11-20, claims 11-20 are computer system and computer programming product claims which correspond to the method claims of claims 1 and 7-10 respectively. As such, they are rejected for the same reasons as the limitations listed above. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. Crabtree et al. (US 20260154553 A1) discusses an external ‘judging’ model which evaluates a response from two competing LLM’s. Any inquiry concerning this communication or earlier communications from the examiner should be directed to JUSTIN A CARDOSO whose telephone number is (571)272-8512. The examiner can normally be reached M-F 7:30 - 5:00, alternate Friday's off. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jennifer Welch can be reached at (571) 272-7212. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JUSTIN CARDOSO/ Patent Examiner, Art Unit 2143 /JENNIFER N WELCH/Supervisory Patent Examiner, Art Unit 2143
Read full office action

Prosecution Timeline

Mar 15, 2024
Application Filed
Aug 10, 2026
Non-Final Rejection mailed — §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month