Prosecution Insights
Last updated: October 02, 2026
Application No. 18/678,446

ADAPTATION OF COMPUTER-IMPLEMENTED MODELS USING ADAPTATION FUSION MATRICES

Non-Final OA §101§103
Filed
May 30, 2024
Examiner
BALAKRISHNAN, VIJAY MURALI
Art Unit
Tech Center
Assignee
Dell Products L.P.
OA Round
1 (Non-Final)
41%
Grant Probability
Moderate
1-2
OA Rounds
1y 7m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 41% of resolved cases
41%
Career Allowance Rate
11 granted / 27 resolved
-19.3% vs TC avg
Strong +73% interview lift
Without
With
+73.3%
Interview Lift
resolved cases with interview
Typical timeline
3y 11m
Avg Prosecution
15 currently pending
Career history
44
Total Applications
across all art units

Statute-Specific Performance

§101
27.5%
-12.5% vs TC avg
§103
36.6%
-3.4% vs TC avg
§102
12.6%
-27.4% vs TC avg
§112
23.3%
-16.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 27 resolved cases

Office Action

§101 §103
DETAILED ACTION This nonfinal action is in response to application 18/678,446 filed on 05/30/2024. Claims 1-20 are pending in the application. Claims 1, 11, and 16 are independent claims. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Interpretation As recited in MPEP § 2111, during patent examination, “the pending claims must be given their broadest reasonable interpretation consistent with the specification”. Under a broadest reasonable interpretation (BRI), claim terms must be given their plain and ordinary meaning (i.e., the meaning that the term would have to a person of ordinary skill in the art), unless applicant sets forth a special definition of a claim term within the specification. The plain and ordinary meaning of a term “may be evidenced by a variety of sources, including the words of the claims themselves, the specification, drawings, and prior art”. Claim 2 recites “obtaining, using the one or more requirements, a model adaptation candidate matrix comprising one or more model adaptation candidates” and “grouping the one or more model adaptation candidates into one or more adaptation groups to obtain an adaptation group matrix”. Based on the mathematical definition of the term “matrix”, the recited “adaptation candidate matrix” and “adaptation group matrix” elements would commonly be interpreted as representing mathematical objects, particularly a rectangular grid or array of numbers of dimension m x n. However, per the specification (see [0062] and Figs. 2C and 2D of Drawings), the term “matrix” is broadly utilized to refer to mere organized arrangements of data, and goes beyond the scope of a numerical m x n mathematical matrix. Under a broadest reasonable interpretation in light of the specification, the examiner notes that the “adaptation candidate matrix” and “adaptation group matrix” elements are therefore interpreted as encompassing any corresponding organized arrangement of data (e.g., including any grid-like or tabular structure). Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.The analysis of the claims will follow the 2019 Revised Patent Subject Matter Eligibility Guidance, 84 Fed. Reg. 50 (“2019 PEG”). Independent Claims (Claim 1, Claim 11, Claim 16): Step 1: Claim 1 is drawn to a method, claim 11 is drawn to a product, and claim 16 is drawn to an apparatus. Therefore, each of these claims falls under one of the four categories of statutory subject matter (process/method, machine/apparatus, manufacture/product, or composition of matter). Step 2A Prong 1: Claims 1, 11, and 16 each recite a judicially recognized exception of an abstract idea. Claim 1 recites, inter alia: A method for managing computer-implemented models, the method comprising: obtaining one or more requirements of an entity and a pre-trained model; group[ing] of the one or more requirements [to obtain an adaptation group-tuned model] – This limitation recites a procedure of observing and organizing general “requirements” to prepare and achieve related objectives, and therefore recites a process of evaluation that a human could reasonably perform using pen and paper. using the adaptation-group-tuned model to provide services associated with the group of the one or more requirements to the entity – This limitation amounts to a process of providing generic services in response to analysis of observed objectives, and therefore recites a recites a process of evaluation that a human could reasonably perform using pen and paper. Claims 11 and 16 recite substantially similar abstract idea limitations to those found in claim 1, and thereby recite the same judicial exception. Step 2A Prong 2: The following additional elements recited in claims 1, 11, and 16 also do not integrate the recited judicial exceptions into a practical application. Claim 1 additionally recites: wherein the pre-trained model is not trained to provide computer-implemented services associated with the one or more requirements when obtained – This limitation does no more than generally invoke a type of machine learning model as a tool to perform an abstract procedure of analysis based on observed objectives. adapting the pre-trained model to a [group of the one or more requirements] to obtain an adaptation-group-tuned model – This limitation does no more than generally invoke a type of machine learning model as a tool to perform an abstract procedure of analysis based on observed objectives. provid[ing] computer-implemented services – This limitation amounts to no more than mere instructions to implement an abstract idea on a computer or computer components. Claim 11 recites substantially similar additional elements to those recited in claim 1, and further recites: A non-transitory machine-readable medium having instructions stored therein, which when executed by a processor, cause the processor to perform operations for managing computer-implemented models – This limitation amounts to no more than mere instructions to implement an abstract idea on a computer or computer components. Claim 16 recites substantially similar additional elements to those recited in claim 1, and further recites: A model adaptation manager, comprising: a processor; and a memory coupled to the processor to store instructions, which when executed by the processor, cause the processor to perform operations for managing computer-implemented models – This limitation amounts to no more than mere instructions to implement an abstract idea on a computer or computer components. Step 2B: The additional elements recited in claims 1, 11, and 16 viewed individually or as an ordered combination, do not provide an inventive concept or otherwise amount to significantly more than the recited abstract ideas themselves. Claim 1 additionally recites: wherein the pre-trained model is not trained to provide computer-implemented services associated with the one or more requirements when obtained – Generally linking a judicial exception to known machine learning techniques (e.g., fine-tuning of pre-trained models) does not provide significantly more than the recited abstract idea. adapting the pre-trained model to a [group of the one or more requirements] to obtain an adaptation-group-tuned model – Generally linking a judicial exception to known machine learning techniques (e.g., fine-tuning of pre-trained models) does not provide significantly more than the recited abstract idea. provid[ing] computer-implemented services – Mere instructions to implement an abstract idea on a computer or computer components do not provide significantly more than the recited abstract idea. Claim 11 recites substantially similar additional elements to those recited in claim 1, and further recites: A non-transitory machine-readable medium having instructions stored therein, which when executed by a processor, cause the processor to perform operations for managing computer-implemented models – Mere instructions to implement an abstract idea on a computer or computer components do not provide significantly more than the recited abstract idea. Claim 16 recites substantially similar additional elements to those recited in claim 1, and further recites: A model adaptation manager, comprising: a processor; and a memory coupled to the processor to store instructions, which when executed by the processor, cause the processor to perform operations for managing computer-implemented models – Mere instructions to implement an abstract idea on a computer or computer components do not provide significantly more than the recited abstract idea. As such, claims 1, 11, and 16 are not patent eligible. Dependent Claims (Claims 2-10, Claims 12-15, Claims 17-20): Dependent claims 2-10, 12-15, and 17-20 narrow the scope of independent claims 1, 11, and 16, and likewise narrow the recited judicial exceptions. They recite abstract idea limitations that are similar to those recited within the independent claims (i.e., mental processes and/or mathematical concepts), and thereby merely expand on the already recited exceptions. The dependent claims also do not recite any further additional elements that successfully integrate the recited judicial exceptions into a practical application or provide significantly more than the recited abstract ideas themselves. Consequently, claims 2-10, 12-15, and 17-20 are also rejected under 35 U.S.C. 101. Step 1: Claims 2-10 are drawn to a method, claims 12-15 are drawn to a product, and claims 17-20 are drawn to an apparatus. Therefore, each of these claims falls under one of the four categories of statutory subject matter (process/method, machine/apparatus, manufacture/product, or composition of matter). Step 2A Prong 1: Claims 2-10, 12-15, and 17-20 each recite a judicially recognized exception of an abstract idea. Claim 2 recites, inter alia: obtaining, using the one or more requirements, a model adaptation candidate matrix comprising one or more model adaptation candidates; and grouping the one or more model adaptation candidates into one or more adaptation groups to obtain an adaptation group matrix, wherein each of the one or more adaptation groups is one instance of the group of the one or more requirements – This limitation recites a procedure of observing and organizing data (e.g., through grid-like or tabular format) in response to analyzing given objectives, and therefore recites a process of evaluation that a human could reasonably perform using pen and paper. Claim 3 recites the same judicial exception as claim 2. Claim 4 recites the same judicial exception as claim 3. Claim 5 recites the same judicial exception as claim 4. Claim 6 recites the same judicial exception as claim 5. Claim 7 recites the same judicial exception as claim 2. Claim 8 recites, inter alia: obtaining an update to the one or more requirements; and updating, using the update to the one or more requirements, the one or more adaptation groups in the adaptation group matrix to obtain an updated adaptation group matrix comprising one or more updated adaptation groups – This limitation recites a procedure of further updating organized data in response to changing objectives, and therefore recites a process of evaluation that a human could reasonably perform using pen and paper. Claim 9 recites the same judicial exception as claim 8. Claim 10 recites, inter alia: wherein updating the one or more adaptation groups using the update to the one or more requirements comprises at least one of: updating properties of one or more adapters making up an adaptation group of the one or more adaptation groups without adding a new adapter to the adaptation group or removing any of the one or more adapters, adding the new adapter to the adaptation group or removing any of the one or more adapters, or adding a new adapter group or removing at least one existing one of the one or more adaptation groups from the adaptation group matrix – This limitation recites a procedure of further updating organized data in response to changing objectives, and therefore recites a process of evaluation that a human could reasonably perform using pen and paper. Claims 12-15 recite substantially similar abstract idea limitations to those recited in claims 2-5, and therefore recite the same judicial exceptions. Claims 17-20 recite substantially similar abstract idea limitations to those recited in claims 2-5, and therefore recite the same judicial exceptions. Step 2A Prong 2: Claims 8 and 10 do not recite any further additional elements besides those already recited in the independent claims, and the following additional elements recited in claims 2-7, 9, 12-15, and 17-20 also do not integrate the recited judicial exceptions into a practical application. Claim 2 additionally recites: wherein the pre-trained model is adapted to each of the one or more adaptation groups to obtain one or more adaptation-group-tuned models, the adaptation-group-tuned model being one of the one or more adaptation-group-tuned models – This limitation does no more than generally invoke a type of machine learning model as a tool to perform an abstract procedure of analysis based on observed objectives. Claim 3 additionally recites: wherein the pre-trained model is adapted to each of the one or more adaptation groups using adapter fusion – This limitation does no more than generally invoke a type of machine learning model and known machine learning technique as tools to perform an abstract procedure of analysis based on observed objectives. Claim 4 additionally recites: wherein each of the one or more adaptation groups comprises one or more adaptation layers for the pre-trained model, and the adapter fusion fuses the one or more adaptation layers into a fused-adaptation layer that is inserted into a component of the pre-trained model – This limitation does no more than generally invoke a type of machine learning model and known machine learning technique as tools to perform an abstract procedure of analysis based on observed objectives. Claim 5 additionally recites: wherein each of the one or more adaptation layers is generated by performing adapter tuning on the pre-trained model – This limitation does no more than generally invoke a type of machine learning model and known machine learning technique as tools to perform an abstract procedure of analysis based on observed objectives. Claim 6 additionally recites: wherein the pre-trained model is a large language model (LLM), and the fused-adaptation layer is inserted into the LLM as a new parameter layer within existing parameter layers making up the LLM – This limitation does no more than generally invoke a type of machine learning model and known machine learning technique as tools to perform an abstract procedure of analysis based on observed objectives. Claim 7 additionally recites: storing each of the one or more adaptation-group-tuned models into an adaptation-group-tuned model repository – This limitation does no more than recite an insignificant intermediary step of gathering and storing data, and therefore recites insignificant extra-solution activity. Claim 9 additionally recites: updating the one or more adaptation-group-tuned models stored in the adaptation-group- tuned model repository using the one or more updated adaptation groups and an adapter fusion technique – This limitation does no more than generally invoke a type of machine learning model and known machine learning technique as tools to perform an abstract procedure of analysis based on observed objectives. Claims 12-15 recite substantially similar abstract idea limitations to those recited in claims 2-5, and therefore also do not integrate the recited judicial exceptions into a practical application. Claims 17-20 recite substantially similar abstract idea limitations to those recited in claims 2-5, and therefore also do not integrate the recited judicial exceptions into a practical application. Step 2B: The additional elements recited in claims 2-7, 9, 12-15, and 17-20, viewed individually or as an ordered combination, do not provide an inventive concept or otherwise amount to significantly more than the recited abstract ideas themselves. Claim 2 additionally recites: wherein the pre-trained model is adapted to each of the one or more adaptation groups to obtain one or more adaptation-group-tuned models, the adaptation-group-tuned model being one of the one or more adaptation-group-tuned models – Generally linking a judicial exception to known machine learning techniques (e.g., fine-tuning of pre-trained models) does not provide significantly more than the recited abstract idea. Claim 3 additionally recites: wherein the pre-trained model is adapted to each of the one or more adaptation groups using adapter fusion – Generally linking a judicial exception to known machine learning techniques (e.g. adapter fusion), which are well-understood, routine, and conventional in the art (see Han et al., “Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey” [page 5]), does not provide significantly more than the recited abstract idea. Claim 4 additionally recites: wherein each of the one or more adaptation groups comprises one or more adaptation layers for the pre-trained model, and the adapter fusion fuses the one or more adaptation layers into a fused-adaptation layer that is inserted into a component of the pre-trained model – Generally linking a judicial exception to known machine learning techniques (e.g. adapter fusion), which are well-understood, routine, and conventional in the art (see Han et al., “Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey” [page 5]), does not provide significantly more than the recited abstract idea. Claim 5 additionally recites: wherein each of the one or more adaptation layers is generated by performing adapter tuning on the pre-trained model – Generally linking a judicial exception to known machine learning techniques (e.g. adapter tuning), which are well-understood, routine, and conventional in the art (see Han et al., “Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey” [page 5]), does not provide significantly more than the recited abstract idea. Claim 6 additionally recites: wherein the pre-trained model is a large language model (LLM), and the fused-adaptation layer is inserted into the LLM as a new parameter layer within existing parameter layers making up the LLM – Generally linking a judicial exception to known machine learning techniques (e.g., large language models and adapter fusion), which are well-understood, routine, and conventional in the art (see Han et al., “Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey” [page 5]), does not provide significantly more than the recited abstract idea. Claim 7 additionally recites: storing each of the one or more adaptation-group-tuned models into an adaptation-group-tuned model repository – Storing data is well-understood, routine, and conventional activity (see MPEP 2106.05(d)(II)(iv); “Storing and retrieving information in memory”) and therefore does not provide significantly more than the recited abstract idea. Claim 9 additionally recites: updating the one or more adaptation-group-tuned models stored in the adaptation-group- tuned model repository using the one or more updated adaptation groups and an adapter fusion technique – Generally linking a judicial exception to known machine learning techniques (e.g. adapter fusion), which are well-understood, routine, and conventional in the art (see Han et al., “Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey” [page 5]), does not provide significantly more than the recited abstract idea. Claims 12-15 recite substantially similar abstract idea limitations to those recited in claims 2-5, and therefore also do not integrate the recited judicial exceptions into a practical application. Claims 17-20 recite substantially similar abstract idea limitations to those recited in claims 2-5, and therefore also do not integrate the recited judicial exceptions into a practical application. As such, claims 2-7, 9, 12-15, and 17-20 are not patent eligible. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Pfeiffer et al. (“AdapterFusion: Non-Destructive Task Composition for Transfer Learning”, available conference 2021), hereinafter Pfeiffer, in view of Li et al. (“Approximate Clustering for Extracting Task Relationships in Multi-Instruction Tuning”, available 25 Mar 2024), hereinafter Li. Pfeiffer was cited in an information disclosure statement (IDS) filed 05/30/2024. Regarding claim 1, Pfeiffer teaches A method for managing computer-implemented models, (“Sequential fine-tuning and multi-task learning are methods aiming to incorporate knowledge from multiple tasks; however, they suffer from catastrophic forgetting and difficulties in dataset balancing. To address these shortcomings, we propose AdapterFusion, a new two stage learning algorithm that leverages knowledge from multiple tasks. First, in the knowledge extraction stage we learn task specific parameters called adapters, that encapsulate the task-specific information. We then combine the adapters in a separate knowledge composition step. We show that by separating the two stages, i.e., knowledge extraction and knowledge composition, the classifier can effectively exploit the representations learned from multiple tasks in a non-destructive manner” [Pfeiffer Abstract]) the method comprising: obtaining a pre-trained model, wherein the pre-trained model is not trained to provide computer implemented services associated with the one or more requirements when obtained; (“We are given a model that is pretrained on a task with training data D0 and a loss function L0…In the remainder of this paper, we refer to this pretrained model by the tuple (D0;L0). We define C as the set of N classification tasks having labelled data of varying sizes and different loss functions: C = {(D1;L1); : : : ; (DN;LN)}. The aim is to be able to leverage a set of N tasks to improve on a target task m with Cm = (Dm;Lm)” [Pfeiffer page 488 Task Definition]) adapting the pre-trained model to a group to obtain an adaptation-group-tuned model; (“Stickland and Murray (2019) propose to train adapters for N tasks in parallel with a multi-task objective. The underlying parameters PNG media_image1.png 28 36 media_image1.png Greyscale are finetuned along with the task-specific parameters in PNG media_image2.png 35 38 media_image2.png Greyscale .” [Pfeiffer page 489 Multi-Task Adapters (MT-A)]; “In the first stage of our learning algorithm, we train either ST-A or MT-A for each of the N tasks. In the second stage, we then combine the set of N adapters by using AdapterFusion. While fixing both the parameters PNG media_image3.png 20 27 media_image3.png Greyscale as well as all adapters PNG media_image4.png 30 27 media_image4.png Greyscale , we introduce parameters PNG media_image5.png 22 22 media_image5.png Greyscale that learn to combine the N task adapters to solve the target task” [Pfeiffer page 490 Learning algorithm]) and using the adaptation-group-tuned model to provide computer implemented services associated with the group to the entity. (“The aim is to be able to leverage a set of N tasks to improve on a target task m” [Pfeiffer page 488 Task Definition]; “AdapterFusion aims to improve performance on a given target task m by transferring task specific knowledge from the set of all N task adapters” [Pfeiffer page 492 AdapterFusion]). However, Pfeiffer does not expressly teach obtaining one or more requirements of an entity and adapting the pre-trained model to a group of the one or more requirements. In the same field of endeavor, Li teaches a means of fine-tuning language models (“The development of language models involves the evaluation of a broad range of learning tasks. Recent work has shown that by using carefully designed instructions to teach a large transformer model, they can be fine-tuned on a wide range of downstream tasks. However, when the number of instructions increases, they can negatively interfere with each other if trained together. Existing works have relied on domain expertise and manual inspection to construct multi-instruction sets, which can be time-consuming and difficult to scale. To address this challenge, this paper develops a clustering algorithm to find groups of similar tasks based on a given set of task affinity scores” [Li Abstract]) that obtain[s] one or more requirements of an entity and adapt[s] the pre-trained model to a group of the one or more requirements (“Many problems in the context of language model fine-tuning are related to multitask learning (Wang et al., 2018; Aribandi et al., 2022; Sanh et al., 2022). We give three examples, which will be the focus of this paper: (1) Multitask instruction fine-tuning is an essential component of adapting language models, enabling the models with various language processing abilities, such as question answering and text summarization” [Li page 3 Preliminaries]; “Let there be n downstream tasks. The goal of task grouping (cf. Standley et al. (2020)) is to partition the n tasks into k subsets such that each subset of tasks is the best to be trained together. For each pair of tasks u and v, let Tu,v denote an affinity score, which quantifies the transfer effect between them. Pairwise notions of affinity scores between two tasks have been used in prior work (Fifty et al., 2021). For example, one way to quantify Tu,v is via the performance of task u’s validation performance evaluated on a model fine-tuned on a model trained with both u and v” [Li page 3 Task Grouping Setup]; ). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated obtaining one or more requirements of an entity and adapting the pre-trained model to a group of the one or more requirements as taught by Li into Pfeiffer because they are both directed towards fine-tuning language models. Incorporating the task affinity-based grouping technique of Li into Pfeiffer’s AdapterFusion framework would allow for the grouping of related tasks prior to combining their corresponding task adapters, thereby reducing risk of negative interference between dissimilar tasks (“It is also worth noting that these sets typically involve a large number of tasks and instructions, which can lead to severe negative interference when they are trained together naively (Jang et al., 2023).” [Li page 1 Introduction]; “We observe that our approach improves over the baseline methods by 3.3% on average, suggesting the benefit of separating instructions to reduce their negative interference” [Li page 8 Experimental Results]) . Regarding claim 2, the combination of Pfeiffer and Li teaches the limitations of parent claim 1, and Li further teaches obtaining, using the one or more requirements, a model adaptation candidate matrix comprising one or more model adaptation candidates; (“Given an n by n task affinity matrix T, the extent of positive transfers within a subset of tasks S can be characterized by the density of affinity scores in the subset” [Li page 3 Task Grouping Setup]; see also Procedure 2 Adaptive Estimation of Task Affinity Scores [Li page 6]) and grouping the one or more model adaptation candidates into one or more adaptation groups to obtain an adaptation group matrix, wherein each of the one or more adaptation groups is one instance of the group of the one or more requirements, (“To maximize the objective stated in Eq. (2), we can use an assignment variable from every task to every cluster. More precisely, let us denote the assignment variables as an n⇥k matrix V , such that each entry Vi,j indicates whether a task i belongs to a cluster j, for every i = 1, . . . ,n, j = 1, . . . ,k. ” [Li page 4 Semidefinite Programming Relaxations for Task Affinity Clustering]) wherein the pre-trained model is adapted to each of the one or more adaptation groups to obtain one or more adaptation-group-tuned models, the adaptation-group-tuned model being one of the one or more adaptation-group-tuned models. (“For multi-instruction fine-tuning, we use T5-Base as the base model… We apply our approach to find groups of instructions and then fine-tune one model for each group of instructions” [Li page 8 Implementation and Baselines]). Regarding claim 3, the combination of Pfeiffer and Li teaches the limitations of parent claim 2, and Pfeiffer further teaches wherein the pre-trained model is adapted to each of the one or more adaptation groups using adapter fusion. (“In the second stage, we then combine the set of N adapters by using AdapterFusion” [Pfeiffer page 490 Learning algorithm]) Regarding claim 4, the combination of Pfeiffer and Li teaches the limitations of parent claim 3, and Pfeiffer further teaches wherein each of the one or more adaptation groups comprises one or more adaptation layers for the pre-trained model, (“For NLP tasks, adapters have been introduced for the transformer architecture (Vaswani et al., 2017). At each transformer layer l, a set of adapter parameters _l is introduced. The placement and architecture of adapter parameters _ within a pretrained model is non-trivial” [Pfeiffer page 490 Adapters in Practice]) and the adapter fusion fuses the one or more adaptation layers into a fused-adaptation layer that is inserted into a component of the pre-trained model (see Figure 2 -- “Our AdapterFusion architecture. This includes learnable weights Query, Key, and Value. Query takes as input the output of the pretrained transformer weights. Both Key and Value take as input the output of the respective adapters” [Pfeiffer page 490]; “As illustrated in Figure 2, we define the Adapter- Fusion parameters to consist of Key, Value and Query matrices at each layer l, denoted by Kl, Vl and Ql respectively. At each layer l of the transformer and each time-step t, the output of the feedforward sub-layer of layer l is taken as the query vector. The output of each adapter zl;t is used as input to both the value and key transformations… Given the context, AdapterFusion learns a parameterized mixer of the available trained adapters” [Pfeiffer pages 490-491 Components]) Regarding claim 5, the combination of Pfeiffer and Li teaches the limitations of parent claim 4, and Pfeiffer further teaches wherein each of the one or more adaptation layers is generated by performing adapter tuning on the pre-trained model. (“The parameters PNG media_image6.png 32 41 media_image6.png Greyscale are fixed and only the parameters PNG media_image7.png 24 36 media_image7.png Greyscale are trained. This makes it possible to efficiently parallelize the training of adapters for all N tasks, and store the corresponding knowledge in designated parts of the model.” [Pfeiffer page 489 Single-Task Adapters (ST-A)) Regarding claim 6, the combination of Pfeiffer and Li teaches the limitations of parent claim 5, and Pfeiffer further teaches wherein the pre-trained model is a large language model (LLM), (“In all experiments, we use BERT-base-uncased (Devlin et al., 2019) as the pretrained language model.” [Pfeiffer page 491 Experimental Setup]) and the fused-adaptation layer is inserted into the LLM as a new parameter layer within existing parameter layers making up the LLM. (In the second stage, we then combine the set of N adapters by using AdapterFusion. While fixing both the parameters PNG media_image3.png 20 27 media_image3.png Greyscale as well as all adapters PNG media_image4.png 30 27 media_image4.png Greyscale , we introduce parameters PNG media_image5.png 22 22 media_image5.png Greyscale that learn to combine the N task adapters to solve the target task” [Pfeiffer page 490 Learning algorithm]; “Given the context, AdapterFusion learns a parameterized mixer of the available trained adapters” [Pfeiffer pages 490-491 Components]). Regarding claim 7, the combination of Pfeiffer and Li teaches the limitations of parent claim 2, and Pfeiffer further teaches storing each of the one or more adaptation-group-tuned models into an adaptation-group-tuned model repository. (“Adapters trained in both single-task (ST-A) or multi-task (MT-A) setups have learned the idiosyncratic knowledge of the respective tasks’ training data, encapsulated in their designated parameters. This results in a compression of information, which requires less space to store task-specific knowledge. However, the distinct weights of adapters prevent a downstream task from being able to use multiple sources of extracted information. In the next section we describe our two stage algorithm which tackles the sharing of information stored in adapters trained on different tasks” [Pfeiffer page 490 Adapters in Practice]; Each adapter model and its associated information is implicitly stored in memory – code and adapters are also further “available at AdaperHub.ml” (i.e., repository) [Pfeiffer Abstract]) Regarding claim 8, the combination of Pfeiffer and Li teaches the limitations of parent claim 7, and Li further teaches obtaining an update to the one or more requirements; and updating, using the update to the one or more requirements, the one or more adaptation groups in the adaptation group matrix to obtain an updated adaptation group matrix comprising one or more updated adaptation groups. (“The next step is an adaptive sampling procedure to accelerate the above estimation…We initialize the task assignment variable X by assigning Xu,v as 1 |C| if u and v are in a cluster C with size |C|. Then, we solve the SDP again to re-generate the clusters” [Li page 6]; see also Algorithm 3 Adaptive Task Grouping (AdaGroup) [Li page 7]) Regarding claim 9, the combination of Pfeiffer and Li teaches the limitations of parent claim 8, and Li further teaches updating the one or more adaptation-group-tuned models stored in the adaptation-group-tuned model repository using the one or more updated adaptation groups (“The next step is an adaptive sampling procedure to accelerate the above estimation…We initialize the task assignment variable X by assigning Xu,v as 1 |C| if u and v are in a cluster C with size |C|. Then, we solve the SDP again to re-generate the clusters” [Li page 6]; see also Algorithm 3 Adaptive Task Grouping (AdaGroup) [Li page 7]). Pfeiffer further teaches parameters of the adaptation-group-tuned models being learned and updated via an adapter fusion technique (“In the second stage, we then combine the set of N adapters by using AdapterFusion. While fixing both the parameters PNG media_image3.png 20 27 media_image3.png Greyscale as well as all adapters PNG media_image4.png 30 27 media_image4.png Greyscale , we introduce parameters PNG media_image5.png 22 22 media_image5.png Greyscale that learn to combine the N task adapters to solve the target task” [Pfeiffer page 490 Learning algorithm]). Regarding claim 10, the combination of Pfeiffer and Li teaches the limitations of parent claim 8, and Li further teaches wherein updating the one or more adaptation groups using the update to the one or more requirements comprises: adding [a] new adapter to [an] adaptation group (“The next step is an adaptive sampling procedure to accelerate the above estimation. The idea is to divide tasks into small batches and iteratively estimate affinity scores for a new batch of tasks. In each iteration, we have existing cluster structures and a new batch of unclustered tasks.” [Li page 6 Adaptive Sampling]). Regarding claims 11-15, they are product claims that substantially correspond to the method of claims 1-5, which are already taught by the combination of Pfeiffer and Li as detailed above. Pfeiffer further teaches A non-transitory machine-readable medium having instructions stored therein, which when executed by a processor, cause the processor to perform the claimed operations ([Pfeiffer page 491 Experiments]; The disclosed evaluation of the AdapterFusion framework would implicitly require adequate processing power and storage for executing the necessary operations). Consequently, they are rejected for the same reasons as claims 1-5. Regarding claims 16-20, they are apparatus claims that substantially correspond to the method of claims 1-5, which are already taught by the combination of Pfeiffer and Li as detailed above. Pfeiffer further teaches A model adaptation manager, comprising: a processor; and a memory coupled to the processor to store instructions, which when executed by the processor, cause the processor to perform the claimed operations ([Pfeiffer page 491 Experiments]; The disclosed evaluation of the AdapterFusion framework would implicitly require adequate processing power and storage for executing the necessary operations). Consequently, they are rejected for the same reasons as claims 1-5. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to VIJAY M BALAKRISHNAN whose telephone number is (571) 272-0455. The examiner can normally be reached 10am-5pm EST Mon-Thurs. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, JENNIFER WELCH can be reached on (571) 272-7212. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /V.M.B./ Examiner, Art Unit 2143 /JENNIFER N WELCH/Supervisory Patent Examiner, Art Unit 2143
Read full office action

Prosecution Timeline

May 30, 2024
Application Filed
Sep 01, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743623
INFORMATION PROCESSING DEVICE AND MACHINE LEARNING METHOD THAT OPTIMIZE A DECODING PROCESS USING BACK-PROPAGATION
4y 0m to grant Granted Sep 22, 2026
Patent 12731026
METHOD AND SYSTEM FOR PROGRAM SAMPLING USING NEURAL NETWORK
3y 10m to grant Granted Sep 08, 2026
Patent 12711407
REASONING METHOD BASED ON STRUCTURAL ATTENTION MECHANISM FOR KNOWLEDGE-BASED QUESTION ANSWERING AND COMPUTING APPARATUS FOR PERFORMING THE SAME
3y 8m to grant Granted Aug 18, 2026
Patent 12645933
Method and System for Training a Neural Network for Generating Universal Adversarial Perturbations
4y 7m to grant Granted Jun 02, 2026
Patent 12619871
INTERPRETABLE NEURAL NETWORK ARCHITECTURE USING CONTINUED FRACTIONS
3y 11m to grant Granted May 05, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
41%
Grant Probability
99%
With Interview (+73.3%)
3y 11m (~1y 7m remaining)
Median Time to Grant
Low
PTA Risk
Based on 27 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month