Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement filed 26 February 2024 fails to comply with 37 CFR 1.98(a)(1), which requires the following: (1) a list of all patents, publications, applications, or other information submitted for consideration by the Office; (2) U.S. patents and U.S. patent application publications listed in a section separately from citations of other documents; (3) the application number of the application in which the information disclosure statement is being submitted on each page of the list; (4) a column that provides a blank space next to each document to be considered, for the examiner’s initials; and (5) a heading that clearly indicates that the list is an information disclosure statement.
The Information Disclosure Statement filed 26 February 2024 appears to be blank. The examiner has signed the document and has considered the three non-patent references supplied by the applicant, but a list of the filed references needs to be filed in a new Information Disclosure Statement.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 2-11, 13, and 18-19 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Regarding Claim 2, it recites “an AI architecture of the linear path” (lines 4-5) and “AI architecture of the linear path” (line 6). However, the claim previously recites multiple paths, each of which is either a linear path or a non-linear path. Consequently, there may be multiple linear paths and multiple non-linear paths, so it cannot be determined which path is being referred to as “the linear path” and which path is being referred to as “the non-linear path.” For the purposes of examination under prior art, the examiner will interpret “the linear path” to be any of the potential linear paths and “the non-linear path” to be any of the potential non-linear paths.
Regarding Claim 18, it recites elements substantially similar to those of claim 2, so it is indefinite for the same reasons.
Regarding Claims 3-11, 13, and 20, they are rejected as being dependent on rejected base claims.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-16 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. The claims do not fall within at least one of the four categories of patent eligible subject matter because they recite “An adapter to a base model of an artificial intelligence (AI) system” (claim 1). The applicant’s specification describes an adapter as a software module that comprises modifiers and architectural paths. These are all software per-se, which are none of a process, machine, manufacture, or composition of matter. The examiner suggests amending the claims to include hardware elements such as a processor and a memory.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1, 12, 17, and 20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Diao, Shizhe, et al. (“Mixture-of-domain-adapters: Decoupling and injecting domain knowledge to pre-trained language models’ memories,” Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: Long papers). 2023; hereinafter “Diao”).
Regarding Claim 1, Diao teaches an adapter to a base model of an artificial intelligence (AI) system (Abstract and section 3), the adapter comprising:
a connector configured to connect the adapter to the base model such that during an operation of the AI system at least some portion of data transformed by the base model is propagated from the base model to the adapter and back from the adapter to the base model (section 3 and fig. 1b—domain adapters have connectors that connect the adapters to a base transformer model such that data transformed by the “Add & Norm” layer of the base model is propagated to the adapters, and data operated on by the adapters are propagated back to the base model through the Mixture of Adapters Gate);
a non-linear modifier configured to modify the data received from the base model non-linearly before returning the modified portion of the data back to the base model (section 3.1 and fig. 1c—the domain adapters contain a non-linearity layer); and
an AI trainer configured to tune the non-linear modifier of the adapter by propagating training data through the base model and the adapter and updating weights of the non-linear modifier of the adapter for given weights of the base model to optimize a loss function (section 3 and fig. 1—an AI trainer tunes the adapter in two stages, each of which propagates training data through the base model and the adapter and updates weights of the adapter. Section 3.1 further describes optimizing a loss function).
Regarding Claim 12, Diao teaches wherein the AI trainer is further configured to approximate the base model and train the adapter and to achieve a common objective (sections 3 and 3.1—the base model is a language model, and the trainer trains the overall system as a language model, thus approximating the base model; the adapter is trained to achieve the same objective as the base model and improve language performance in particular domains).
Regarding Claim 17, Diao teaches a method for adapting a base model of an artificial intelligence (AI) system using an adapter (Abstract and section 3), the method comprising:
connecting, using a connector of the adapter, the adapter to the base model such that during an operation of the AI system at least some portion of data transformed by the base model is propagated from the base model to the adapter and back from the adapter to the base model (section 3 and fig. 1b—domain adapters have connectors that connect the adapters to a base transformer model such that data transformed by the “Add & Norm” layer of the base model is propagated to the adapters, and data operated on by the adapters are propagated back to the base model through the Mixture of Adapters Gate);
modifying, using a non-linear modifier of the adapter, the data received from the base model non-linearly before returning the modified portion of the data back to the base model (section 3.1 and fig. 1c—the domain adapters contain a non-linearity layer); and
tuning, using an AI trainer of the adapter, the non-linear modifier of the adapter by propagating training data through the base model and the adapter and updating weights of the non-linear modifier of the adapter for given weights of the base model to optimize a loss function (section 3 and fig. 1—an AI trainer tunes the adapter in two stages, each of which propagates training data through the base model and the adapter and updates weights of the adapter. Section 3.1 further describes optimizing a loss function).
Regarding Claim 20, Diao teaches a non-transitory computer readable storage medium embodied thereon a program executable by a processor for performing a method (Abstract and section 3. A pretrained language model and the experiments in section 4 imply a non-transitory computer readable storage medium embodied thereon a program executable by a processor). Diao teaches the method comprising the steps of the present claim in the same manner as for claim 17, above.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 2-11, 13-16, and 18-19 are rejected under 35 U.S.C. 103 as being unpatentable over Diao, as applied to claims 1 and 17, above, in view of Li, Xiaopeng, et al. (“Hamur: Hyper adapter for multi-domain recommendation,” Proceedings of the 32nd ACM International Conference on Information and Knowledge Management. 2023; hereinafter “Li”).
Regarding Claims 2 and 18, Diao teaches wherein the non-linear modifier includes multiple paths formed by multiple AI architectures of data transformation, each of the paths is either a linear path configured to modify the received data linearly or a non-linear path configured to modify the received data non-linearly (sections 3 and 3.1 and fig. 1b—there are multiple paths through multiple domain adapters, each path including linear layers {the down- and up-projection layers} and a nonlinearity layer) , and wherein the non-linear modifier includes at least one non-linear path (section 3.1 and fig. 1—each domain adapter includes a nonlinearity layer, thus including at least one non-linear path).
Diao teaches training the adapters using a loss function (sections 3 and 3.1), but does not explicitly teach wherein an AI architecture of the linear path modifies the received data linearly using one or multiple weight matrices, wherein an Al architecture of the non-linear path modifies the received data linearly using one or multiple weight matrices and modifies the received data non-linearly using one or multiple non-linear functions.
However, Li teaches wherein an AI architecture of the linear path modifies the received data linearly using one or multiple weight matrices, wherein an AI architecture of the non-linear path modifies the received data linearly using one or multiple weight matrices and modifies the received data non-linearly using one or multiple non-linear functions (section 2.2 and fig. 2—the domain adapters modify received data linearly through multiplication by multiple weight matrices. At least one adapter also modifies the data non-linearly via a non-linear layer).
All of the claimed elements were known in Diao and Li and could have been combined by known methods with no change in their respective functions. It therefore would have been obvious to a person of ordinary skill in the art at the time of filing of the applicant’s invention to combine the multiplying by weight matrices of Li with the adapters of Diao to yield the predictable result of wherein the non-linear modifier includes multiple paths formed by multiple Al architectures of data transformation, each of the paths is either a linear path configured to modify the received data linearly or a non-linear path configured to modify the received data non-linearly, wherein an Al architecture of the linear path modifies the received data linearly using one or multiple weight matrices, wherein an Al architecture of the non-linear path modifies the received data linearly using one or multiple weight matrices and modifies the received data non-linearly using one or multiple non-linear functions, and wherein the non-linear modifier includes at least one non-linear path. One would be motivated to make this combination for the purpose of improving the flexibility to adapt to diverse domains (Li, Abstract).
Regarding Claim 3, Diao/Li teaches wherein the non-linear modifier includes multiple non-linear paths using different non-linear functions, different arrangements of the same non-linear functions with respect to the weight matrices, or both (Diao, section 3.1 and fig. 1—each domain adapter comprises a path, and each includes a non-linear layer. ReLU is given as an example of a non-linear function; other non-linear functions are an obvious variation. The Mixture-of-Adapters Gate, for example, described in section 3.3, includes a Sigmoid function).
Regarding Claim 4, Diao/Li teaches wherein the non-linear modifier includes at least one linear path (Diao, section 3 and fig. 1—although each adapter includes a nonlinearity layer, they also include linear down projection and up projection layers. Li describes linear matrix multiplication in section 2.2 and fig. 1. An adapter path with only linear layers is an obvious variation).
Regarding Claim 5, Diao/Li teaches wherein the multiple non-linear paths include the same weight matrices (Li, section 2.2 and fig. 2—the domain adapters use several weight matrices, some of which are shared {i.e. the same weight matrices}).
Regarding Claim 6, Diao/Li teaches wherein the multiple non-linear paths share at least some weights (Li, section 2.2 and fig. 2—the domain adapters use several weight matrices, some of which are shared {i.e. shared weights}).
Regarding Claim 7, Diao/Li teaches wherein weights in the weight matrices of the multiple non-linear paths come from a common pool of parameters, such that to tune the non-linear modifier, the AI trainer updates the common pool of parameters (Li, section 2.4).
Regarding Claim 8, Diao/Li teaches wherein the non-linear modifier comprises:
a path splitter configured to direct the received data to each of the paths (Diao, section 3.1 and fig. 1—the Add & Norm layer is a path splitter that splits data to each path through the multiple domain adapters); and
a path combiner configured to combine outputs of each of the paths to submit a combined output back to the base model (Diao, section 3.3 and fig. 1—the Mixture-of-Adapters Gate is a path combiner).
Regarding Claim 9, Diao/Li teaches wherein the path combiner combines the outputs using an operation including one or a combination of: an identity, a duplication, a permutation, a polynomial basis expansion, a Fourier basis expansion, an addition, a multiplication, a division, a subtraction, a modulo-addition, a modulo-product, a Kronecker product, a Kronecker sum, a Hadamard product, a concatenation, a log-sum-exp, an affine transform, a convolution, randomization, a normalization, a nonlinear activation operation, and variants thereof (Diao, section 3.3—the path combiner uses a weighted sum, which comprises addition and multiplication, among other calculations).
Regarding Claim 10, Diao/Li teaches wherein the operation of the path combiner includes a parameter learned during the tuning of the AI trainer (Diao, section 3.3—the path combiner is trained, thus including learned parameters).
Regarding Claim 11, Diao/Li teaches wherein the AI architecture of the non-linear path includes a bottleneck configuration of multiple layers (Diao, section 3.1 and fig. 1; also Li, section 2.3.1).
Regarding Claims 13 and 19, Diao/Li teaches wherein the AI trainer further comprises a weight constructor comprising a pool of parameters and a set of hyperparameters forming rules of propagation of the parameters from the pool of parameters into the weight matrices of the multiple paths of the non-linear modifier, and wherein the weight constructor is configured to: update the pool of parameters and the set of hyperparameters for given weights of the base model; and propagate the parameters from the pool of parameters to different weight matrices of different paths according to the trained hyperparameters (Li, section 2.4—the shared weight matrices comprise a pool of parameters, and weight matrices are updated dynamically, propagating parameters to the different weight matrices).
Regarding Claim 14, Diao/Li teaches wherein the AI trainer updates weights of the adapter for frozen weights of the base model (Diao, section 3.1 and fig. 1—parameters {weights} of the adapters are updated during training while parameters {weights} of the base model are frozen).
Regarding Claim 15, Diao/Li teaches wherein weight matrices of the adapter have lower dimensions than weight matrices of the base model (Diao, sections 3.1 and 3.2—the adapters use parameter efficient tuning methods to keep the parameter size low, indicating lower dimensions in the weight matrices of the adapters compared to the base model. Li, section 2.3.1 also describes the adapters as small modules with smaller parameter dimensions than the base model).
Regarding Claim 16, Diao/Li teaches wherein weight matrices of the adapter are coming from a pool of parameters updated by the AI trainer during the tuning (Li, section 2.4), and wherein a number of parameters in the pool of parameters is more than 1000 times less than a number of parameters of the base model (Diao, sections 3.1 and 3.2 and Li, section 2.3.1—the number of parameters of the adapter is small compared to that of the base model. A number that is more than 1000 times less is an obvious variation).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. This art includes:
Song, Chenyang, et al. (“Conpet: Continual parameter-efficient tuning for large language models,” arXiv preprint arXiv:2309.14763 (2023)) teaches tuning of adapters to a large language model with a selector that chooses the top-k adapters to process an input and concatenates results from the adapters
Wang et al. (U.S. 2025/0103876) teaches an ensemble of LoRA adapters to a large language model
Maschmeyer et al. (U.S. 2025/0094025) teaches a large language model system with a LoRA repository that applies multiple LoRA adapters based on user selections
Bang et al. (U.S. 2025/0131262) teaches a personalized model generated from a base model and a pool of adapters that are combined by a weighted sum or tensor
Any inquiry concerning this communication or earlier communications from the examiner should be directed to HAL W SCHNEE whose telephone number is (571) 270-1918. The examiner can normally be reached M-F 7:30 a.m. - 6:00 p.m.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michael Huntley can be reached at 303-297-4307. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/HAL SCHNEE/Primary Examiner, Art Unit 2129