Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
This Office Action is sent in response to Application’s Communication received on 05/01/2026 for application number 18/592250. The Office hereby acknowledges receipt of the following and placed of record in file: Specification, Drawing, Abstract, Oath/Declaration, and Claims.
Claims (1-8), (9-16) and (17-20) are presented for examination.
Information Disclosure Statement
The information disclosure statements (IDS) submitted on 02/29/2024 and 02/29/2024 were filed prior to current Office Action. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1 and 9 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. The claims require the additional weight deltas to be: “based on the input data guided by the initial weights and the one or more previous weight deltas”. The term “guided by” does not identify an objective mathematical relationship, it is unclear whether “guided by” requires use of the disclosed forgetting-regularization loss; mere consideration of the previous weights; a penalty based on the previous deltas; avoidance of overlapping feature-space locations; or any calculation using the previous parameters. The specification provides examples, but the claims do not state which relationship is required. It is also uncertain how “guided by” differs from the later requirement that updated weights be “based on” those parameters. Therefore, claims 1 and 9 are rejected under 35 U.S.C. 112(b) as Indefinite because the phrase “guided by the initial weights and the one or more previous weight deltas” fails to define the required relationship weight deltas with reasonable clarity. The claims do not specify what operation or degree of influence constitutes being “guided by” those parameters.
Claim Rejections - 35 USC§ 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefore, subject to the conditions and requirements of this title.
Claims (17-18) are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception (i.e., a law of nature, a natural phenomenon, or an abstract idea) without significantly more.
Step 1: Claims (17-18) are drawn to a method each of which is within the four statutory categories (e.g., a process, a machine).
Step 2A - Prong One: In prong one of step 2A, the claims are analyzed to evaluate whether they recite a judicial exception.
Claim 17. A method comprising:
obtaining, using at least one processing device of an electronic device, input data associated with a user request for a trained machine learning model;
identifying, using the at least one processing device, one or more customized tokens associated with the input data, the one or more customized tokens associated with one or more of multiple previous concepts learned by the trained machine learning model;
identifying, using the at least one processing device, key, value, and query features based on the input data and the one or more customized tokens;
performing, using the at least one processing device, key-value projection using the key features, the value features, and weights of the trained machine learning model to generate projected features; and
generating, using the at least one processing device, a response to the user request based on the query features and the projected features;
wherein the weights of the trained machine learning model are modified by sequentially teaching the trained machine learning model one or more new concepts over time.
Claim 18. The method of Claim 17, wherein:
the user request comprises a request to generate an image containing one or more specified contents, the one or more specified contents associated with one or more of the multiple previous concepts learned by the trained machine learning model;
the one or more customized tokens are associated with the one or more of the multiple previous concepts learned by the trained machine learning model; and
the method further comprises generating a new customized token for each of the one or more new concepts.
Claim 17 expressly recites “performing … key-value projection using the key features, the value features, and weights of the trained machine learning model to generate projected features.” This limitation describes a mathematical calculation: projection of feature representations using model weights. Paragraph [0054] expressly characterizes this operation as a mathematical projection; paragraphs [0052] and [0067] explain the attention features and weighted transformations. Therefore, falls within mathematical concepts
Step 2A Prong 2:
The additional elements are considered both individually and in their ordered interaction with the projection. The device supplies execution: the request and customized tokens select the information to be processed; the feature-identification stage supplies representations; and the response stage uses the projected features with the query features. The final clause requires weights produced through sequential teaching. The complete sequence is therefore addressed, including the relationship between previously learned concepts, their tokens and the model weights.
The specification describes technical benefits: reducing catastrophic forgetting, limiting additional parameters and avoiding long-term retention of training data. Those benefits deserve substantive consideration. Claim 17, however, does not specify how earlier knowledge is protected when weights change or how the projection or feature-processing architecture is improved. It requires the use of customized representations in a model that has learned concepts in sequence, without specifying an update mechanism that accounts for earlier concept updates. [P, ¶¶0027–0031, 0054–0059].
Likewise, the customized tokens constrain the model’s input representation, but the claim does not require the disclosed random initialization, replacement of concept names, or protection of earlier embeddings. The specification links reduced token interference to those further techniques. Identifying a token associated with a prior concept, at the breadth claimed, does not itself require that interference-reduction mechanism. [P, ¶¶0060–0065, 0078].
This is a claim-scope finding, not a requirement to recite every embodiment or numerical performance threshold. The combination remains instructions to use the projection within a broadly specified learned-model response process. Obtaining request data is preparatory; the response is the broadly specified result of the processing. The electronic device is not claimed through an architecture that provides a separate technical improvement. These elements do not meaningfully integrate the calculation into an improved technical process.
Step 2B:
The additional elements are reconsidered as an ordered combination. No further limitation changes their identified roles: input acquisition, concept-representation selection, model execution and generation of the requested informational result. The unrestricted sequential-training clause does not supply an additional implementation of knowledge preservation. Accordingly, the proposed rejection rests on instructions to apply the calculation in that model-processing environment, together with preparatory and result-producing activity; the combination does not supply an inventive concept beyond those roles. [A1, §2106.05(f)–(h)]
The specification describes the processor as selectable from general processor types and describes ordinary memory and interfaces. This supports treating the recited hardware, as claimed, as an implementation tool.
Step 1 and Step 2A, Prong One
Claim 18 incorporates claim 17 and remains a statutory process. It inherits the mathematical key-value projection identified above. Its added limitations specify an image-generation request for content associated with previously learned concepts, reiterate the token-to-concept association, and require generating a new customized token for each new concept. They do not independently recite a named mathematical calculation.
Step 2A, Prong Two:
The image-request limitation gives the inherited response process a concrete purpose. It is nevertheless stated at the level of desired image content; it does not require a particular synthesis, denoising, rendering or image-correction technique. The rejection therefore does not turn on the mere fact that an image is digital. It turns on the absence of a claimed technical method by which mathematical processing improves image generation.
Generating a new customized token for each concept is an operative limitation and is evaluated with the inherited token/feature/projection sequence. The new token provides a concept-specific handle for later requests. However, claim 18 does not specify how that token’s representation is initialized or learned to reduce interference, or how older representations and concept knowledge are maintained while new concepts are taught. The specification’s more particular token-learning and continual-adaptation techniques remain outside this claim. [P, ¶¶0060–0065, 0078]
Viewed as a whole, claim 18 narrows the requested output and extends the concept identifiers used by claim 17’s model-processing pipeline. On the proposed construction, it still claims use of that pipeline to obtain a desired result rather than the disclosed technical mechanism for improving continual customization.
Step 2B:
The additional image-request and token-generation limitations are reconsidered with all inherited limitations. At their claimed breadth, they specify the application of the projection and the representation of the concepts to be processed; they do not add further technical implementation beyond those functional roles. Merely restricting the desired response to image content does not supply significantly more.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 3, 9 and 11 are rejected under AIA 35 U.S.C. 103(a) as being unpatentable over Hyder et al. NPL publication 2022: Incremental Task Learning with Incremental Rank Updates (hereinafter Hyder) in view of Chen et al. US Patent Application Publication US 20220383126 A1 (hereinafter Chen).
Regarding claim 1, Hyder teaches obtaining, using at least one processing device of an electronic device, input data associated with a new concept to be learned by a trained machine learning model (Abstract, page. 1, ¶ 2, §3, pp. 5–7, Eqs. (2)– (4), wherein Hyder describes using new tasks to train a trained machine learning) identifying, using the at least one processing device, initial weights of the trained machine learning model and one or more previous weight deltas associated with the trained machine learning model; identifying, using the at least one processing device, one or more additional weight deltas based on the input data and guided by the initial weights and the one or more previous weight deltas (§3, pp. 5–7, Eqs. (2)– (4) wherein Hyder uses additional weight and factors to train the trained machine learning model based on the existing task and new added tasks).
PNG
media_image1.png
320
684
media_image1.png
Greyscale
Hyder teaches identifying, using the at least one processing device, one or more additional weight deltas based on the input data and guided by the initial weights and the one or more previous weight deltas (§3, pp. 5–7, Eqs. (2)– (4) wherein Hyder describes weight for every layer learned as low-rank matrix for every task wherein the task 1 represents the initial weight and previous weight) wherein the one or more additional weight deltas are integrated into the trained machine learning model by identifying updated weights for the trained machine learning model based on the initial weights, the one or more previous weight deltas, and the one or more additional weight deltas (Abstract, (§3, pp. 5–7, Eqs. (2)– (4) wherein Hyder describes adding weights for each layer as a linear combination of several rank-I matrices. The network is updated based on new task weight and learn a rank-I (or low-rank) matrix and add that to the weights of every layer. Wherein the additional selector vector that assigns different weights to the low-rank matrices learned for the previous tasks and provides better accuracy and forgetting).
Hyder does not teach and integrating, using the at least one processing device, the one or more additional weight deltas into the trained machine learning model.
However in analogous art of sequential customization of text-to-image diffusion models, Chen teaches integrating, using the at least one processing device, the one or more additional weight deltas into the trained machine learning model ([0015-0018], wherein Chen incorporates LoRA that allows he training of each of multiple dense layers in the neural network indirectly by injecting and optimizing their rank decomposition matrices A and B, while keeping the original matrices of pretrained weights unchanged. Wherein a low-rank factorization matrices are injected into other neural network models to adapt them to specific tasks and domains).
It would have been obvious to a person in the ordinary skill in the art before the effective filing date of the claimed invention to combine Hyder with Chen by incorporating the method of integrating, using the at least one processing device, the one or more additional weight deltas into the trained machine learning model of Chen into the method of using the at least one processing device, one or more additional weight deltas based on the input data and guided by the initial weights and the one or more previous weight deltas of Hyder for the purpose of incorporating neural network-based model base model weight matrices for each of multiple neural network layers. First low-rank factorization matrices are added to corresponding base model weight matrices to form a first domain model. The low-rank factorization matrices are treated as trainable parameters. The first domain model is trained with first domain specific training data without modifying base model weight matrices. (Chen: [0004]).
Regarding claim 3, Hyder as modified by Chen teaches wherein: the one or more previous weight deltas are associated with one or more concepts previously learned by the trained machine learning model; the one or more additional weight deltas are associated with the new concept; and identifying the one or more additional weight deltas comprises performing sequential, self-regulating low-rank adaptation based on the one or more previous weight deltas (Abstract, [0015],[0022-0027], [0075], Claims 1–5 and 19 text, wherein Chen obtains neural network-based model base model weight matrices for each of multiple neural network layers. First low-rank factorization matrices are added to corresponding base model weight matrices to form a first domain model. The low-rank factorization matrices are treated as trainable parameters. The first domain model is trained with first domain specific training data without modifying base model weight matrices), (Abstract, page. 1, ¶ 2, §3, pp. 5–7, Eqs. (2)– (4) wherein Hyder focuses on task-incremental continual learning in which data for every task are provided in a sequential manner to train/update the network. Wherein Algorithms 1–2: retain earlier low-rank factors, add factors for a new task, and learn selectors that weight old and new contributions. Old factors participate in the model used to optimize the new task.).
Claim 9 is similar in scope to claim 1 therefore the claims are rejected under similar rationale.
Claim 11 is similar in scope to claim 3 therefore the claims are rejected under similar rationale.
Claims 2 and 10 are rejected under AIA 35 U.S.C. 103(a) as being unpatentable over Hyder et al. NPL publication 2022: Incremental Task Learning with Incremental Rank Updates (hereinafter Hyder) in view of Chen et al. US Patent Application Publication US 20220383126 A1 (hereinafter Chen) and further in view of Kumari et al. NPL publication 2022, Multi-Concept Customization of Text-to-Image Diffusion (hereinafter Kumari).
Regarding claim 2, Hyder and Chen do not teach wherein: the trained machine learning model is configured to perform a key-value projection as part of at least one cross-attention function; the at least one cross-attention function is configured to receive input from a text encoder of the trained machine learning model; and the text encoder is configured to provide at least one embedding of at least one customized token associated at least one concept learned by the trained machine learning model.
However in analogous art of sequential customization of text-to-image diffusion models, Kumari teaches wherein: the trained machine learning model is configured to perform a key-value projection as part of at least one cross-attention function; the at least one cross-attention function is configured to receive input from a text encoder of the trained machine learning model; and the text encoder is configured to provide at least one embedding of at least one customized token associated at least one concept learned by the trained machine learning model (Page. 3-4, wherein Kumari identifies a small subset of model weights, namely the key and value mapping from text to latent features in the cross-attention layers. Wherein Kumari optimizes key and value projection matrices in the diffusion model cross-attention layers along with the modifier token).
It would have been obvious to a person in the ordinary skill in the art before the effective filing date of the claimed invention to combine Kumari with Hyder and Chen by incorporating the method of wherein: the trained machine learning model is configured to perform a key-value projection as part of at least one cross-attention function; the at least one cross-attention function is configured to receive input from a text encoder of the trained machine learning model; and the text encoder is configured to provide at least one embedding of at least one customized token associated at least one concept learned by the trained machine learning model of Kumari into the method of using the at least one processing device, one or more additional weight deltas based on the input data and guided by the initial weights and the one or more previous weight deltas of Hyder and Chen for the purpose of train for multiple concepts or combine multiple fine-tuned models into one via closed-form constrained optimization (Kumari: [0004]).
Claim 10 is similar in scope to claim 2 therefore the claims are rejected under similar rationale.
Claims 4-5 and 12-13 are rejected under AIA 35 U.S.C. 103(a) as being unpatentable over Hyder et al. NPL publication 2022: Incremental Task Learning with Incremental Rank Updates (hereinafter Hyder) in view of Chen et al. US Patent Application Publication US 20220383126 A1 (hereinafter Chen) and further in view of Desjardins et al. US Patent Application Publication US 20190236482 A1 (hereinafter Desjardins).
Regarding claim 4, Hyder and Chen do not teach:
PNG
media_image2.png
330
712
media_image2.png
Greyscale
However in analogous art of sequential customization of text-to-image diffusion models, Desjardins teaches general forgetting-regularization wherein new-task training constrained by earlier parameter importance; an objective combines new-task performance with a weighted penalty for changing earlier-task parameters (Claims 1, 4, 6, 8–10 and 18 text, Abstract, [0035], [0042] wherein new-task training constrained by earlier parameter importance; an objective combines new-task performance with a weighted penalty for changing earlier-task parameters).
It would have been obvious to a person in the ordinary skill in the art before the effective filing date of the claimed invention to combine Desjardins with Hyder and Chen by incorporating the method of claim 4 of Desjardins into the method of using the at least one processing device, one or more additional weight deltas based on the input data and guided by the initial weights and the one or more previous weight deltas of Hyder and Chen for the purpose of allowing the machine learning model to learn the second task without forgetting the first task (e.g., by retaining memory about the first task), the system trains the machine learning model on the new training data by adjusting the first values of the parameters to optimize an objective function that depends in part on a penalty term that is based on the determined measures of importance of the parameters to the machine learning model with respect to the first task (Desjardins: [0055]).
Regarding claim 5, Hyder as modified by Chen and Desjardins teaches:
PNG
media_image3.png
246
710
media_image3.png
Greyscale
Chen teaches first domain language model is trained at operation 230 with first domain specific training data without modifying base model weight matrices. Training may include the use of a loss function using standard backpropagation, calculating a gradient for every parameter and updating weights by subtracting the gradients ([0031], [0042]). (Claims 1, 4, 6, 8–10 and 18 text, Abstract, [0035], [0042] wherein new-task training constrained by earlier parameter importance; an objective combines new-task performance with a weighted penalty for changing earlier-task parameters).
Claim 12 is similar in scope to claim 4 therefore the claims are rejected under similar rationale.
Claim 13 is similar in scope to claim 5 therefore the claims are rejected under similar rationale.
Claims 6 and 14 are rejected under AIA 35 U.S.C. 103(a) as being unpatentable over Hyder et al. NPL publication 2022: Incremental Task Learning with Incremental Rank Updates (hereinafter Hyder) in view of Chen et al. US Patent Application Publication US 20220383126 A1 (hereinafter Chen) and further in view of Xue et al. NPL Publication 2022, Meta-attention for ViT-backed Continual Learning (hereinafter Xue).
Regarding claim 6, Hyder and Chen do not teach identifying, using the at least one processing device, a hard-attention mask based on a learnable mask tensor containing learnable mask parameters that are parameterized using a categorical distribution; wherein the updated weights for the trained machine learning model are identified based on the initial weights, the one or more previous weight deltas, and the hard-attention mask applied to the one or more additional weight deltas.
However in analogous art of sequential customization of text-to-image diffusion models, Xue teaches identifying, using the at least one processing device, a hard-attention mask based on a learnable mask tensor containing learnable mask parameters that are parameterized using a categorical distribution; wherein the updated weights for the trained machine learning model are identified based on the initial weights, the one or more previous weight deltas, and the hard-attention mask applied to the one or more additional weight deltas (§§3.2.1–3.2.2, pp. 3–5, Eqs. (6)–(8), wherein Xue dynamically assigns an attention mask with continuous values to modify the standard information interaction between all image tokens to learn adaptive communication patterns when sequentially studying new tasks).
It would have been obvious to a person in the ordinary skill in the art before the effective filing date of the claimed invention to combine Xue with Hyder and Chen by incorporating the method of identifying, using the at least one processing device, a hard-attention mask based on a learnable mask tensor containing learnable mask parameters that are parameterized using a categorical distribution; wherein the updated weights for the trained machine learning model are identified based on the initial weights, the one or more previous weight deltas, and the hard-attention mask applied to the one or more additional weight deltas of Xue into the method of using the at least one processing device, one or more additional weight deltas based on the input data and guided by the initial weights and the one or more previous weight deltas of Hyder and Chen for the purpose of assigning special-designed attention masks to the self-attention block to create adapted token interaction patterns in a task-specific manner by dynamically activating and isolating corresponding image tokens when incrementally learning new tasks. (Xue: §3.2.1, p. 4).
Regarding claim 7, Hyder as modified by Chen and Xue teach wherein identifying the hard-attention mask comprises applying multi-layer perceptron (MLP) parameterization to the learnable mask tensor to modify the learnable mask parameters prior to being parameterized using the categorical distribution (page. 8, paragraph 1, page. 13, paragraph 1, wherein Hyder incorporates multilayer perceptron (MLP) and factorization and rank selection for hyperparameter tuning, and wherein low-rank increments approach can be generalized to other type of networks and layers, wherein convolutional kernels have four dimensional weight tensors as opposed to the two-dimensional weight matrices of fully connected layers).
Claim 14 is similar in scope to claim 6 therefore the claims are rejected under similar rationale.
Claims 8 and 16 are rejected under AIA 35 U.S.C. 103(a) as being unpatentable over Hyder et al. NPL publication 2022: Incremental Task Learning with Incremental Rank Updates (hereinafter Hyder) in view of Chen et al. US Patent Application Publication US 20220383126 A1 (hereinafter Chen) and further in view of Xue et al. NPL Publication 2022, Meta-attention for ViT-backed Continual Learning (hereinafter Xue) and further in view Guo et al. NPL Publication 2020, Parameter-Efficient Transfer Learning With Dief Pruning (hereinafter Guo).
Regarding claim 8, Hyder, Chen and Xue do teach wherein identifying the hard-attention mask comprises performing sparsity regularization to reduce a number of non-zero values contained in the hard-attention mask.
However in analogous art of sequential customization of text-to-image diffusion models, Guo teaches wherein identifying the hard-attention mask comprises performing sparsity regularization to reduce a number of non-zero values contained in the hard-attention mask (Abstract, §1, pp. 1-2, §§3–3.3, pp. 2–4: Guo teaches a learned difference vector added to a fixed pretrained model; elementwise mask on trainable differences; sparsity regularization; subsequent magnitude pruning and fixed-mask tuning. Guo describes Diff pruning that becomes parameter-efficient as the number of tasks increases, as it requires storing only the nonzero positions and weights of the diff vector for each task, while the cost of storing the shared pretrained model remains constant).
It would have been obvious to a person in the ordinary skill in the art before the effective filing date of the claimed invention to combine Guo with Hyder, Chen and Xue by incorporating the method of wherein identifying the hard-attention mask comprises performing sparsity regularization to reduce a number of non-zero values contained in the hard-attention mask of Guo into the method of using the at least one processing device, one or more additional weight deltas based on the input data and guided by the initial weights and the one or more previous weight deltas of Hyder, Chen and Xue for the purpose of incorporating parameter-efficiency with pretrained models that is to learn sparse models for each task where a subset of the final model parameters are exactly zero. (Guo: §1, p. 2).
Claim 16 is similar in scope to claim 8 therefore the claims are rejected under similar rationale.
Claims 17-20 are rejected under AIA 35 U.S.C. 103(a) as being unpatentable over Hyder et al. NPL publication 2022: Incremental Task Learning with Incremental Rank Updates (hereinafter Hyder) in view of Chen et al. US Patent Application Publication US 20220383126 A1 (hereinafter Chen) and further in view of Kumari et al. NPL publication 2022, Multi-Concept Customization of Text-to-Image Diffusion (hereinafter Kumari).
Regarding claim 17, Hyder teaches obtaining, using at least one processing device of an electronic device, input data associated with a user request for a trained machine learning model (Abstract, page. 1, ¶ 2, §3, pp. 5–7, Eqs. (2)– (4), wherein Hyder describes using new tasks to train a trained machine learning).
Hyder does not teach identifying, using the at least one processing device, one or more customized tokens associated with the input data, the one or more customized tokens associated with one or more of multiple previous concepts learned by the trained machine learning model; identifying, using the at least one processing device, key, value, and query features based on the input data and the one or more customized tokens; performing, using the at least one processing device, key-value projection using the key features, the value features, and weights of the trained machine learning model to generate projected features; and generating, using the at least one processing device, a response to the user request based on the query features and the projected features; wherein the weights of the trained machine learning model are modified by sequentially teaching the trained machine learning model one or more new concepts over time.
However in analogous art of sequential customization of text-to-image diffusion models, Kumari teaches identifying, using the at least one processing device, one or more customized tokens associated with the input data, the one or more customized tokens associated with one or more of multiple previous concepts learned by the trained machine learning model; identifying, using the at least one processing device, key, value, and query features based on the input data and the one or more customized tokens; performing, using the at least one processing device, key-value projection using the key features, the value features, and weights of the trained machine learning model to generate projected features; and generating, using the at least one processing device, a response to the user request based on the query features and the projected features; wherein the weights of the trained machine learning model are modified by sequentially teaching the trained machine learning model one or more new concepts over time (§3.1, pp. 2–4, Figs. 2 and 4: concept-specific modifier tokens and updating key/value projections in diffusion cross-attention. §3.2, pp. 4–5: joint training or merging separately customized models. §4.2, p. 7: an explicit sequential-training baseline on two concepts, with forgetting of the first.).
It would have been obvious to a person in the ordinary skill in the art before the effective filing date of the claimed invention to combine Kumari with Hyder and Chen by incorporating the method teaches identifying, using the at least one processing device, one or more customized tokens associated with the input data, the one or more customized tokens associated with one or more of multiple previous concepts learned by the trained machine learning model; identifying, using the at least one processing device, key, value, and query features based on the input data and the one or more customized tokens; performing, using the at least one processing device, key-value projection using the key features, the value features, and weights of the trained machine learning model to generate projected features; and generating, using the at least one processing device, a response to the user request based on the query features and the projected features; wherein the weights of the trained machine learning model are modified by sequentially teaching the trained machine learning model one or more new concepts over time of Kumari into the method of using the at least one processing device, one or more additional weight deltas based on the input data and guided by the initial weights and the one or more previous weight deltas of Hyder and Chen for the purpose of train for multiple concepts or combine multiple fine-tuned models into one via closed-form constrained optimization (Kumari: [0004]).
Regarding claim 18, Hyder as modified by Chen and Kumari teach wherein: the user request comprises a request to generate an image containing one or more specified contents, the one or more specified contents associated with one or more of the multiple previous concepts learned by the trained machine learning model; the one or more customized tokens are associated with the one or more of the multiple previous concepts learned by the trained machine learning model; and the method further comprises generating a new customized token for each of the one or more new concepts (§3.1, pp. 2–4, Figs. 2 and 4 wherein Kumari describes individual concept and train them jointly. To denote the target concepts, Kumari uses different modifier tokens initialized with different rarely-occurring tokens and optimize them along with cross-attention key and value matrices for each layer).
Regarding claim 19, Hyder as modified by Chen and Kumari teach for each of the one or more new concepts: obtaining, using the at least one processing device, additional data associated with the new concept for the trained machine learning model; generating, using the at least one processing device, one or more additional weight deltas based on the additional data; and identifying, using the at least one processing device, updated weights for the trained machine learning model based on initial weights of the trained machine learning model, one or more previous weight deltas associated with at least one of the multiple previous concepts, and the one or more additional weight deltas (§3.1, pp. 2–4, Figs. 2 and 4 wherein Kumari describes identifying previous weights and previous trained visual concepts and provides method for training the machine learning based on the new concepts), (Abstract, [0015],[0022-0027], [0075], Claims 1–5 and 19 text, wherein Chen obtains neural network-based model base model weight matrices for each of multiple neural network layers. First low-rank factorization matrices are added to corresponding base model weight matrices to form a first domain model. The low-rank factorization matrices are treated as trainable parameters. The first domain model is trained with first domain specific training data without modifying base model weight matrices), (Abstract, page. 1, ¶ 2, §3, pp. 5–7, Eqs. (2)– (4) wherein Hyder focuses on task-incremental continual learning in which data for every task are provided in a sequential manner to train/update the network. Wherein Algorithms 1–2: retain earlier low-rank factors, add factors for a new task, and learn selectors that weight old and new contributions. Old factors participate in the model used to optimize the new task.).
Regarding claim 20, Hyder as modified by Chen and Kumari teach deleting the additional data from the electronic device after the one or more additional weight deltas are generated ([0060], [0063] wherein Chen removes the first low-rank factorization matrices and adding to the base model weight matrices, corresponding second low-rank factorization matrices treated as trainable parameters that are trained with second domain specific training data without modifying base model weight matrices).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to HASSAN MRABI whose telephone number is (571)272-8875. The examiner can normally be reached on Monday-Friday, 7:30am-5pm. Alt, Friday, EST.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Viker Lamardo can be reached on 571-270-5871. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/HASSAN MRABI/Examiner, Art Unit 2144