DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 1-20 are pending and have been examined. Claims 1, 11, and 20 are independent.
This Application was published as U.S. 20250166599.
Apparent priority: October 20, 2023.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 02/06/2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 (Statutory Category). Claim 1 is directed to a method (a process), claim 11 is directed to a computing system (a machine), and claim 20 is directed to non-transitory computer-readable storage media (an article of manufacture). Each independent claim is therefore directed to a statutory category of invention.
Step 2A, Prong One (Recitation of a Judicial Exception). Independent claims 1, 11, and 20 recite, in pertinent part, processing a natural language instruction to generate an instruction representation based on a meaning of the natural language instruction, translating the instruction representation into data indicating an intent of the natural language instruction, providing the natural language instruction and that data onward, and generating a response based on the natural language instruction and that data. Under their broadest reasonable interpretation, these limitations recite mental processes, that is, concepts performed in the human mind, including observation, evaluation, judgment, and opinion. A person who reads a written instruction, forms an understanding of what the instruction means, determines from that understanding what the instruction is asking for, passes the instruction together with that determination to a second person, and composes a response to the instruction in light of both, performs each recited step mentally or with the aid of pen and paper. Reciting that the steps are carried out by a first language model, a translation module, and a second language model does not remove the limitations from the mental process grouping, because a claim that recites a judicial exception performed by a generically recited computer component still recites the judicial exception. The claims therefore recite an abstract idea.
Step 2A, Prong Two (Integration into a Practical Application). The judicial exception is not integrated into a practical application. Beyond the mental processes, claim 1 recites a first language model, a translation module comprising an interface between the first language model and a second language model, and a second language model trained with domain specific knowledge. Claim 11 additionally recites processing circuitry in communication with storage media, the processing circuitry configured to execute the machine learning system. Claim 20 additionally recites non-transitory computer-readable storage media having instructions encoded thereon configured to cause processing circuitry to perform the recited operations. Each of these additional elements is recited at a high level of generality. The claims do not recite the architecture, the training, or the internal operation of either language model or of the translation module beyond the functional results those components achieve, and claims 8 and 18 recite the machine learning model of the translation module as a transformer network, a Recurrent Neural Network (RNN), or a state space model, each a conventional model type named at the level of the class. These additional elements amount to mere instructions to apply the mental processes using generic computer components, and to generally linking the use of the judicial exception to a particular technological environment. The claims do not recite an improvement to the functioning of a computer or to any other technology or technical field; the asserted advance lies in interpreting an instruction and composing a response to it, not in any improvement to the way the recited components operate. Accordingly, the additional elements do not integrate the abstract idea into a practical application, and the claims are directed to the abstract idea.
Step 2B (Inventive Concept). The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration, the first language model, the second language model, the translation module, the processing circuitry, the storage media, and the non-transitory computer-readable storage media are recited at a high level of generality and perform the well-understood, routine, and conventional functions of receiving data, processing data with a machine learning model, passing data between components, and outputting data. That these components are conventional is evidenced by Feng and Moriarty of record, discussed below: Feng discloses a general-purpose language model cooperating with a smaller specialized language model through an interface, and Moriarty discloses the recited processing circuitry, storage media, and non-transitory computer-readable storage media. Considered individually and as an ordered combination, the additional elements amount to no more than mere instructions to apply the abstract idea using generic components and do not provide an inventive concept. The claims are not patent eligible.
The dependent claims have been considered and do not cure the deficiencies of the independent claims. Each dependent claim is addressed separately below.
Claim 2 recites that generating a response comprises generating a programming code based on the natural language instruction. Composing program code that carries out a stated instruction is a step a person can perform mentally or with pen and paper, and narrowing the form the response takes further describes the mental process itself; the claim adds no additional element beyond the generically recited language models addressed above and is neither integrated into a practical application nor significantly more.
Claim 3 recites that the programming code implements a user desired function and that the natural language instruction specifies an input format and an expected behavior of the user desired function. These limitations describe the content of the instruction that is read and of the response that is composed, which is the information on which the mental process operates; the claim supplies no additional element and is neither integrated into a practical application nor significantly more.
Claim 4 recites generating, by the first language model, a hidden representation indicative of the meaning of the natural language instruction. Forming a representation of what an instruction means is the mental step of comprehension restated at the level of its result, and the first language model that produces it remains a generically recited tool; the claim supplies no additional element beyond those addressed above.
Claim 5 recites concatenating and/or adding, by the translation module, the hidden representation with one or more prefix embeddings to form the data. Concatenating or adding numeric vectors is a mathematical calculation, which is itself a judicial exception, and performing that calculation with the generically recited translation module is neither integrated into a practical application nor significantly more.
Claim 6 recites that the translation module comprises a machine learning model and that the method further comprises training that model to translate the instruction representation into the data indicating the intent using a dataset containing pairs of the natural language instructions and corresponding responses. Training a model on paired examples is a mathematical optimization performed on a generically recited model, and the dataset is the data on which that mathematics operates; the claim supplies no additional element.
Claim 7 recites freezing the first language model and the second language model while training the translation module, or freezing the first language model while training the translation module and training the second language model. Holding the parameters of a model fixed while a training computation runs further specifies the mathematical training procedure of claim 6 and supplies no additional element.
Claim 8 recites that the machine learning model comprises a transformer network, a Recurrent Neural Network (RNN), or a state space model. Naming a conventional model type recites a generic machine learning component employed as a tool to perform the abstract idea; naming the component neither integrates the exception into a practical application nor provides an inventive concept.
Claim 9 recites tokenizing the natural language instruction into a sequence of instruction tokens and generating the hidden representation based on that sequence. Dividing text into units and forming an understanding from the resulting sequence are steps a person can perform mentally or with pen and paper, and they further describe the data on which the generically recited model operates; the claim supplies no additional element.
Claim 10 recites that the first language model is a large language model (LLM) and that the second language model is a smaller domain-specific model. Specifying the relative size and the domain of the generically recited models describes those models at the same high level of generality addressed above and supplies no additional element.
Claim 12 recites that the machine learning system configured to generate the response is further configured to generate a programming code based on the natural language instruction. Composing program code that carries out a stated instruction is a step a person can perform mentally or with pen and paper, and the machine learning system that performs it is recited generically; the claim supplies no additional element.
Claim 13 recites that the programming code implements a user desired function and that the natural language instruction specifies an input format and an expected behavior of the user desired function. These limitations describe the content of the instruction and of the response and supply no additional element; the claim is neither integrated into a practical application nor significantly more.
Claim 14 recites that the machine learning system is further configured to generate, by the first language model, a hidden representation indicative of the meaning of the natural language instruction. Forming a representation of what an instruction means is the mental step of comprehension restated at the level of its result, performed by the generically recited first language model; the claim supplies no additional element.
Claim 15 recites concatenating and/or adding, by the translation module, the hidden representation with one or more prefix embeddings to form the data. Concatenating or adding numeric vectors is a mathematical calculation, which is itself a judicial exception, and performing that calculation with the generically recited translation module is neither integrated into a practical application nor significantly more.
Claim 16 recites that the translation module comprises a machine learning model and that the machine learning system is further configured to train that model to translate the instruction representation into the data indicating the intent using a dataset containing pairs of the natural language instructions and corresponding responses. Training a model on paired examples is a mathematical optimization performed on a generically recited model, and the dataset is the data on which that mathematics operates; the claim supplies no additional element.
Claim 17 recites freezing the first language model and the second language model while training the translation module, or freezing the first language model while training the translation module and training the second language model. Holding the parameters of a model fixed while a training computation runs further specifies the mathematical training procedure of claim 16 and supplies no additional element.
Claim 18 recites that the machine learning model comprises a transformer network, a Recurrent Neural Network (RNN), or a state space model. Naming a conventional model type recites a generic machine learning component employed as a tool to perform the abstract idea and provides no inventive concept.
Claim 19 recites that the first language model is a large language model (LLM) and that the second language model is a smaller domain-specific model. Specifying the relative size and the domain of the generically recited models describes those models at the same high level of generality addressed above and supplies no additional element.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 10, 11, 19, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Feng et al., "COOK: Empowering General-Purpose Language Models with Modular and Collaborative Knowledge," arXiv:2305.09955v1, May 17, 2023, https://arxiv.org/abs/2305.09955v1, hereinafter Feng, in view of Moriarty et al. (U.S. 20250005282).
Regarding claim 1, Feng discloses:
1. A method for generating responses by a Machine Learning (ML) system, the method comprising: (Feng discloses a machine learning framework that generates responses: "we propose CooK, a novel framework to empower general-purpose LLMs with modular and Collaborative Knowledge through the integration of smaller, but specialized language models" (Feng, Sec. 1, third para.), the framework operating "until the LLM answers 'No' and generates a knowledge-informed response" (Feng, Sec. 2.3, Top-Down Approach, final sentence).)
processing, by a first language model, a natural language instruction to generate an instruction representation based on a meaning of the natural language instruction; (In Feng's top-down approach the general-purpose LLM processes the query first and generates a textual representation of what the instruction requires: "We first ask the LLM a yes/no question to determine whether external knowledge is needed for the given query q" and, in automatic selection, "We further prompt the LLM with 'What kind of information do you need?' and select one specialized LM based on its response rq" (Feng, Sec. 2.3, Top-Down Approach, second para. and Automatic Selection para.). The general-purpose LLM is the first language model, the query q is the natural language instruction, and the response rq, generated from the meaning of the query, is the instruction representation.)
translating, by a translation module comprising an interface between the first language model and a second language model, the instruction representation into data indicating an intent of the natural language instruction, wherein the second language model is trained with domain specific knowledge; (Feng discloses a translation module comprising an interface between the two models, the relevance filter and selection mechanism through which the LLM's response activates a specialized LM: "We adopt a separate encoder-based LM enc(.) that maps a token sequence to a feature vector and cosine similarity sim(.,.) to measure relevance." (Feng, Sec. 2.2, Relevance Filter para., fifth sentence.) That interface translates the instruction representation rq into a determination of the kind of information the instruction calls for: "We further prompt the LLM with 'What kind of information do you need?' and select one specialized LM based on its response rq. Concretely, we identify which LM description {s1, ..., sn} is most relevant to rq with the relevance filter (Sec. 2.2) and activate the corresponding LM" (Feng, Sec. 2.3, Automatic Selection para.). In the worked example, rq states what the query is asking about: "The state Tom Brady is from." (Feng, Table 9.) The relevance filter and selection mechanism of Feng are the translation module of the claim, and their output identifying the kind of information the instruction calls for is the data indicating the intent of the natural language instruction. Feng also discloses that the second model is trained with domain specific knowledge: "we propose to curate specialized LMs that are much smaller than black-box LLMs, trained on diversified knowledge corpora from a wide range of domains and sources" (Feng, Sec. 2.1, second para.).)
providing, by the translation module, the natural language instruction and the data indicating the intent of the natural language instruction to the second language model; and (Feng discloses providing the instruction to the second model: "these specialized LMs are selectively activated and used with prompted generation. Formally, given the query q, specialized LM c defines a mapping c(q): q → dq where q is used as prompt to generate a continuation as the knowledge document dq" (Feng, Sec. 2.1, second para.). The relevance filter and selection mechanism, the translation module as set forth above, perform this providing by activating the selected specialized LM with the query as its prompt.)
generating, by the second language model, a response based on the natural language instruction and the data indicating the intent of the natural language instruction. (The knowledge document dq generated by the specialized LM from the query is the response (Feng, Sec. 2.1, second para.).)
Feng does not teach the aspects of the claim in which the data indicating the intent is provided to and used by the second language model.
Moriarty discloses:
A method for generating responses by a Machine Learning (ML) system, the method comprising: (Moriarty discloses a text analysis system with machine learning models that receives a natural language input and returns a result: "an input text for performing a text analysis task may be received" (Moriarty, para. [0041]; see also Fig. 6, step 610), the system performing "text analysis tasks, such as summarization, comparison, question answering, or adding introductory or conclusory sections" and returning "a result 156 which can be passed back as text analysis 158" (Moriarty, para. [0017]).)
processing, by a first language model, a natural language instruction to generate an instruction representation based on a meaning of the natural language instruction; (Moriarty discloses processing the input to generate a representation of it: "the input text may be parsed, tokenized, transformed into a feature vector or other representation which may be input to a classification system (e.g., using machine-learning models) or a similarity search" (Moriarty, para. [0043]; see also Fig. 6, step 620). The feature vector or other representation generated from the input text is an instruction representation based on its content.)
translating, by a translation module comprising an interface between the first language model and a second language model, the instruction representation into data indicating an intent of the natural language instruction, wherein the second language model is trained with domain specific knowledge; (Moriarty discloses determining, from the input, the domain it calls for and deriving the domain-indicating data: the representation of the input text is input to the classification system or similarity search "to identify the domain-specialty" (Moriarty, para. [0043]), and the domain entities are "extracted from the input text using an machine learning model trained to recognize entities of a domain in a given text" (Moriarty, para. [0043]; Abstract; Fig. 6, step 620), performed by domain entity extraction 144 (Moriarty, para. [0017]; Fig. 1). Moriarty selects among pluralities of domain-specific models based on the determined domain: "Corresponding machine learning models for entity recognition and pre-trained large language models fine-tuned to the domain may be identified (e.g., legal entity extraction models and fine-tuned large language models may be identified and used)" (Moriarty, para. [0042]; see also claim 8). Moriarty further discloses that the second model is trained with domain specific knowledge, being "a pre-trained large language model fine-tuned to the domain" (Moriarty, para. [0044]; Fig. 1, element 142).)
providing, by the translation module, the natural language instruction and the data indicating the intent of the natural language instruction to the second language model; and (Moriarty provides both components to the second model: "the one or more domain entities may be inserted as part of generating instructions to perform the text analysis task using a pre-trained large language model fine-tuned to the domain" (Moriarty, para. [0044]; see also para. [0038], insertion of the recognized entities into the instruction prompt), and "the pre-trained large language model fine-tuned to the domain may be caused to perform the text analysis task on the input text using the generated instructions that include the domain entities" (Moriarty, para. [0045]; see also para. [0035] and Fig. 1, extracted entities in task analysis instruction 154).)
generating, by the second language model, a response based on the natural language instruction and the data indicating the intent of the natural language instruction. (The result of the text analysis task, generated by the fine-tuned large language model from the input text and the entity-bearing instructions, is the response (Moriarty, paras. [0045]-[0046]; para. [0017], result 156; Fig. 6, steps 640 and 650).)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the framework of Feng such that the determination of the kind of information the instruction calls for, made at the interface between the general-purpose language model and the specialized language model, is provided together with the query to the selected specialized language model and used in generating its response, as Moriarty inserts the domain-indicating entities into the instructions supplied to the domain-fine-tuned model and causes that model to perform the task on the input text using those instructions. One of ordinary skill in the art would have been motivated to make this modification in order to condition the response-generating model on the determined domain of the instruction as Moriarty expressly teaches that guiding the performance of the task with the terms present in the text serves to "reduce hallucination and improve summary completeness" (Moriarty, para. [0014]).
Regarding claim 10, the combination of Feng and Moriarty teaches:
10. The method of claim 1, wherein the first language model is a large language model (LLM), and wherein the second language model is a smaller domain-specific model. (Feng discloses this size and role relation expressly: the specialized LMs are "much smaller than black-box LLMs" (Feng, Sec. 2.1, second para.), "with 100x fewer parameters than the general-purpose LLM" (Feng, Sec. 1, final para.), and "one specialized LM with 1.3B parameters successfully updates the parametric knowledge of the 175B Codex" (Feng, Sec. 4, MidtermQA para.).)
Claim 11 is a system claim with limitations corresponding to the limitations of Claim 1 and is rejected under similar rationale. Additionally:
Regarding claim 11, Feng discloses:
11. … processing circuitry in communication with storage media, the processing circuitry configured to execute a machine learning system comprising a first language model, a second language model and a translation module, the machine learning system configured to: (Feng's framework executes on processing circuitry in communication with storage media: "We used a GPU cluster with 16 NVIDIA A40 GPUs, 1988G memory, and 104 CPU cores for the experiments." (Feng, Appendix E, Computation Resources Details para.). The first language model, the second language model, and the translation module of the executed framework are the general-purpose LLM, the specialized LM, and the relevance filter and selection mechanism as set forth for claim 1.)
Claim 19 is a system claim with limitations corresponding to the limitations of Claim 10 and is rejected under a similar rationale.
Claim 20 is a computer program product claim with limitations corresponding to the limitations of method Claim 1 and is rejected under similar rationale. Additionally:
Feng does not teach the recited non-transitory computer-readable storage media.
Moriarty discloses non-transitory computer-readable storage media storing instructions executable by processing circuitry: the methods are implemented on computer systems that include "one or more processors executing program instructions stored on one or more computer-readable storage media coupled to the processors" (Moriarty, para. [0051]), the media being non-transitory: "a non-transitory, computer-readable storage medium may include storage media or memory media such as magnetic or optical media" (Moriarty, para. [0056]; see also claim 14).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to encode the instructions that cause the processing circuitry to perform the operations of the combined framework of Feng and Moriarty on the non-transitory computer-readable storage media of Moriarty. One of ordinary skill in the art would have been motivated to make this modification in order that the recited operations be carried out by the one or more processors of a computing device, as Moriarty expressly teaches (Moriarty, paras. [0051] and [0056]).
Claims 2, 3, 12, and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Feng in view of Moriarty as applied to claims 1 and 11 above, and further in view of Chen et al., "Evaluating Large Language Models Trained on Code," arXiv:2107.03374v1, Jul. 7, 2021, https://arxiv.org/abs/2107.03374v1, hereinafter Chen.
Regarding claim 2, Feng in view of Moriarty teaches the method of claim 1 as set forth above. The combination of Feng and Moriarty does not disclose the following, which Chen teaches:
2. The method of claim 1, wherein generating a response comprises: generating a programming code based on the natural language instruction. (Chen discloses a specialized language model that generates program code from a natural language instruction: "we focus on the task of generating standalone Python functions from docstrings" (Chen, Sec. 1, third para.), synthesizing "programs from docstrings" (Chen, Abstract).)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to employ the specialized code-generation model of Chen as the domain-specific second language model of the combination, Chen expressly describing Codex as "a specialized GPT model" (Chen, Sec. 1, second para.). One of ordinary skill in the art would have been motivated to make this modification in order to serve programming-domain instructions with generated programs whose functional correctness is verified by unit tests (Chen, Abstract; Sec. 1, third para.), and Feng evidences the compatibility of Codex within its framework by adopting it directly: "We use Codex (CODE-DAVINCI-002) [Chen et al., 2021] as the default, general-purpose, black-box LLM" (Feng, Sec. 3, first para.).
Regarding claim 3, Feng in view of Moriarty and further in view of Chen teaches the method of claim 2 as set forth above. The combination of Feng and Moriarty does not disclose the following, which Chen teaches:
3. The method of claim 2, wherein the programming code implements a user desired function, and wherein the natural language instruction specifies an input format and an expected behavior of the user desired function. (Chen discloses instructions that specify both the input format and the expected behavior of the desired function: "Each problem includes a function signature, docstring, body, and several unit tests" (Chen, Sec. 2.2, first para.), and the model "needs to be able to follow instructions to implement the functionality specified in the docstring" (Chen, Sec. 4.2, sixth para.). The function signature specifies the input format, and the docstring specifies the expected behavior of the user desired function.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to further modify the combination such that the natural language instruction supplied to the code-generating second language model specifies the input format and the expected behavior of the user desired function, in the form of the function signature and the docstring of Chen. One of ordinary skill in the art would have been motivated to make this modification in order to state the function to be implemented in a form against which the generated program can be evaluated for functional correctness, Chen expressly teaching that each problem "includes a function signature, docstring, body, and several unit tests" (Chen, Sec. 2.2, first para.) and that the model "needs to be able to follow instructions to implement the functionality specified in the docstring" (Chen, Sec. 4.2, sixth para.).
Claim 12 is a system claim with limitations corresponding to the limitations of Claim 2 and is rejected under a similar rationale.
Claim 13 is a system claim with limitations corresponding to the limitations of Claim 3 and is rejected under a similar rationale.
Claims 4-9 and 14-18 are rejected under 35 U.S.C. 103 as being unpatentable over Feng in view of Moriarty as applied to claims 1 and 11 above, and further in view of Ivison et al., "HINT: Hypernetwork Instruction Tuning for Efficient Zero-Shot Generalisation," arXiv:2212.10315v1, Dec. 20, 2022, https://arxiv.org/abs/2212.10315v1, hereinafter Ivison.
Regarding claim 4, Feng in view of Moriarty teaches the method of claim 1 as set forth above. The combination of Feng and Moriarty does not disclose the following, which Ivison teaches:
4. The method of claim 1, wherein generating the instruction representation further comprises: generating, by the first language model, a hidden representation indicative of the meaning of the natural language instruction. (Ivison discloses an encoder that is a pretrained language model and that generates continuous contextual representations of the instruction: "Our hypernetwork consists of two core elements: an encoder to transform instruction and few-shot text into continuous (contextual) representations, and a generator to then convert these embeddings into the parameters-efficient modules described above." (Ivison, Sec. 3.2, first para.) "To encode our text, we simply use a pretrained language model encoder (e.g., T5)." (Ivison, Sec. 3.2, Encoder para.) The pretrained language model encoder of Ivison is a language model that generates, from the natural language instruction, continuous contextual representations, which are the hidden representation indicative of the meaning of the natural language instruction.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to further modify the combination such that the first language model generates, in processing the natural language instruction, a hidden representation of the instruction in the manner of Ivison's pretrained language model encoder, which the translation module then translates. One of ordinary skill in the art would have been motivated to make this modification in order to process the instruction once into a reusable representation instead of carrying the instruction in every model input, as Ivison expressly teaches that its models "convert task instructions and examples using a pretrained text encoder into parameter-efficient modules inserted into an underlying model, eliminating the need to include instructions in the model input" (Ivison, Abstract).
Regarding claim 5, Feng in view of Moriarty and further in view of Ivison teaches the method of claim 4 as set forth above. The combination of Feng and Moriarty does not disclose the following, which Ivison teaches:
5. The method of claim 4, wherein translating, by the translation module, the instruction representation into data indicating an intent of the natural language instruction comprises: concatenating and/or adding, by the translation module, the hidden representation with one or more prefix embeddings to form the data. (Ivison discloses generating prefixes and concatenating them within the attention computation: "Following Li and Liang (2021), we concatenate prefixes to the key and values of the self- and cross-attentions in every layer." (Ivison, Sec. 3.1, Prefixes para.) Ivison further concatenates the encoded representation of the instruction with the model input: "we take the hypernetwork-encoded sequence and concatenate it with the rest of the input to use in the decoder cross-attention" (Ivison, Sec. 3.3, Fusion In Decoder para.). The generated prefixes are the one or more prefix embeddings of the claim, the hypernetwork-encoded sequence is the hidden representation of claim 4, and their concatenation forms the data.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to further modify the combination such that the translation module forms the data indicating the intent by concatenating the hidden representation of the natural language instruction with one or more prefix embeddings, in the manner of Ivison. One of ordinary skill in the art would have been motivated to make this modification because the concatenation adds a "small amount of extra computation, but yields substantial performance improvements" (Ivison, Sec. 3.3, Fusion In Decoder para.).
Regarding claim 6, Feng in view of Moriarty and further in view of Ivison teaches the method of claim 4 as set forth above, and further teaches:
6. The method of claim 4, wherein the translation module comprises a machine learning model, the method further comprising: training the machine learning model of the translation module to translate the instruction representation into the data indicating the intent of the natural language instruction using a dataset containing pairs of the natural language instructions and corresponding responses. (The translation module of the combination comprises a machine learning model: "Domain entity recognition 110, which may be a locally hosted (e.g., on a same system as text analysis system 140) or remotely hosted machine learning model that is trained to recognize entities in given text for a domain" (Moriarty, para. [0015]). Moriarty trains its machine learning models, including the discretely trained entity detection model, on labeled paired data: "a model training coordinator 235 may be used for training the machine learning models with labeled training data, such as annotated transcripts" (Moriarty, para. [0032]), and "the medical entity detection model and the role identification model are discretely trained for the specific entity detection/role identification" (Moriarty, para. [0028]). The training data pairs each natural language text with its corresponding ground truth output: training data set 102 contains input text 104a paired with analysis task ground truth 104b (Moriarty, para. [0015]; Fig. 1), and the labeled training data comprises "previously provided summaries" (Moriarty, para. [0032]). The paired natural language texts and corresponding ground truth outputs are the dataset containing pairs of the natural language instructions and corresponding responses, and the model so trained produces the domain entities that are the data indicating the intent, as set forth for claim 1.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to train the machine learning model of the translation module of the combination using a dataset containing pairs of the natural language instructions and corresponding responses, as Moriarty trains the machine learning models of its text analysis system with labeled training data pairing texts with their corresponding ground truth outputs (Moriarty, paras. [0015] and [0032]). One of ordinary skill in the art would have been motivated to make this modification in order that the domain data the translation module derives reflect the analysis tasks the system performs, as Moriarty expressly teaches that guiding the task with the terms present in the text serves to "reduce hallucination and improve summary completeness" (Moriarty, para. [0014]).
Regarding claim 7, Feng in view of Moriarty and further in view of Ivison teaches the method of claim 6 as set forth above, and further teaches:
7. The method of claim 6, further comprising: at least one of 1) freezing the first language model and the second language model while training the translation module and 2) freezing the first language model while training the translation module and training the second language model. (The claim recites alternatives, of which the second is taught by the combination. The first language model of Feng is a black-box model that is never trained: the general-purpose LLMs "are prohibitively expensive to train or adapt, Cook specifically focuses on augmenting black-box LLMs" (Feng, Sec. 1, third para.). The second language model of Feng is trained: each specialized LM starts from an existing LM checkpoint and is "further trained on a specific knowledge corpora Di with the causal language modeling objective" (Feng, Sec. 2.1, second para.). The machine learning model of the translation module is trained as set forth for claim 6 (Moriarty, paras. [0015], [0028], and [0032]).)
No further modification of the combination is required; the freezing of the first language model and the training of the translation module and the second language model follow from the combination as set forth for claims 1, 4, and 6.
Regarding claim 8, Feng in view of Moriarty and further in view of Ivison teaches the method of claim 6 as set forth above. The combination of Feng and Moriarty does not disclose the following, which Ivison teaches:
8. The method of claim 6, wherein the machine learning model comprises a transformer network, a Recurrent Neural Network (RNN) or a state space model. (The claim recites species in the alternative, of which the transformer species is taught. Ivison's hypernetwork is built from transformer components: "To encode our text, we simply use a pretrained language model encoder (e.g., T5)." (Ivison, Sec. 3.2, Encoder para.), and its generator uses "a single multi-head cross-attention layer" (Ivison, Sec. 6.2, Full Decoder vs Multi-head Attention para.); Ivison likewise describes the models of its framework as transformers: "The underlying model is simply a transformer model (Vaswani et al., 2017)" (Ivison, Sec. 3, second para.). The T5 encoder and multi-head cross-attention generator of the hypernetwork are a transformer network.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to further modify the combination such that the machine learning model of the translation module is the transformer network of Ivison. One of ordinary skill in the art would have been motivated to make this modification because the multi-head cross-attention generator "performs better at a much cheaper cost" (Ivison, Sec. 6.2, Full Decoder vs Multi-head Attention para.).
Regarding claim 9, Feng in view of Moriarty and further in view of Ivison teaches the method of claim 4 as set forth above, and further teaches:
9. The method of claim 4, wherein processing the natural language instruction comprises tokenizing the natural language instruction into a sequence of instruction tokens, and wherein generating the hidden representation comprises generating the hidden representation based on the sequence of instruction tokens. (Ivison discloses tokenizing the instruction and generating the hidden representation from the resulting token sequence: the pretrained language model encoder transforms the instruction into the continuous contextual representations (Ivison, Sec. 3.2, first para. and Encoder para., as quoted for claim 4), the encoder inputs being sequences of tokens: "Median sequence length, given in number of T5 tokens" (Ivison, Appendix A, Table 5 caption).)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to further modify the combination such that the natural language instruction is tokenized into the sequence of instruction tokens from which the hidden representation is generated. One of ordinary skill in the art would have been motivated to make this modification in order to place the instruction in the token sequence form that the pretrained language model encoder processes, as Ivison's encoder operates on sequences of T5 tokens (Ivison, Appendix A, Table 5 caption).
Claim 14 is a system claim with limitations corresponding to the limitations of Claim 4 and is rejected under a similar rationale.
Claim 15 is a system claim with limitations corresponding to the limitations of Claim 5 and is rejected under a similar rationale.
Claim 16 is a system claim with limitations corresponding to the limitations of Claim 6 and is rejected under a similar rationale.
Claim 17 is a system claim with limitations corresponding to the limitations of Claim 7 and is rejected under a similar rationale.
Claim 18 is a system claim with limitations corresponding to the limitations of Claim 8 and is rejected under a similar rationale.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
a. J. Li et al., "BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models," arXiv:2301.12597: a querying transformer trained as a bridge between a frozen image encoder and a frozen large language model, whose projected query embeddings are prepended to the input text embeddings as soft prompts for the frozen language model.
b. X. L. Li and Liang, "Prefix-Tuning: Optimizing Continuous Prompts for Generation," arXiv:2101.00190v1: continuous task-specific prefix vectors prepended as virtual tokens to steer a frozen language model without fine-tuning.
c. Song et al., "MPNet: Masked and Permuted Pre-training for Language Understanding," arXiv:2004.09297: the encoder-based language model adopted by Feng as the encoder in the relevance filter.
d. Irving et al. (US 12,450,464 B2): a language model system employing auxiliary scoring and rule models in processing natural language inputs.
e. Dong et al. (US 12,591,746 B2): a hierarchy of trained virtual token generators generating, based on user inputs, fixed length virtual token embeddings that prompt a large language model.
f. Ye and Ren, "Learning to Generate Task-Specific Adapters from Task Description," ACL 2021, pp. 646-653: generating adapter parameters for an underlying model from natural language task descriptions.
g. Phang et al., "HyperTuning: Toward Adapting Large Language Models without Back-propagation," arXiv:2211.12485: a hypernetwork trained to generate parameter-efficient model adaptations for an underlying language model from task instructions.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to YUVAL H. LEVENTAL whose telephone number is (571) 270-3130. The examiner can normally be reached Monday-Friday, 8:00 AM - 5:00 PM.
If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, PIERRE-LOUIS DESIR, can be reached at (571) 272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/YUVAL HAIM LEVENTAL/Examiner, Art Unit 2659
/FARIBA SIRJANI/Primary Examiner, Art Unit 2659