Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
2. This office action is in response to the original filing of 04/15/2024. Claims 1-60 are pending and have been considered below.
Claim Rejections - 35 USC § 101
3. 35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-60 are rejected under 35 U.S.C. 101 because the claimed invention is directed to abstract ideas without significantly more.
Claim 1:
Step 1: The claim is directed to a method, falling under one of the four statutory categories of invention.
Step 2A Prong 1: The claim recites following abstract ideas:
The limitations “encoding, with an encoder of the encoder decoder architecture, features of the multi modal inputs to form the multi modal prompt, the multi modal prompt comprising embedded features of mixed modalities from the at least two different input modality types”; and “providing the prompt to a decoder of the encoder decoder architecture to cause the decoder to output the response based on the multi modal prompt, the decoder configured to output the response without prior training on at least one of the multi modal inputs received from the user” 2106.04(a)(2)(I)(C) “Mathematical Calculations A claim that recites a mathematical calculation, when the claim is given its broadest reasonable interpretation in light of the specification, will be considered as falling within the "mathematical concepts" grouping. A mathematical calculation is a mathematical operation (such as multiplication) or an act of calculating using mathematical methods to determine a variable or number, e.g., performing an arithmetic operation such as exponentiation. There is no particular word or set of words that indicates a claim recites a mathematical calculation. That is, a claim does not have to recite the word "calculating" in order to be considered a mathematical calculation. For example, a step of "determining" a variable or number using mathematical methods or "performing" a mathematical operation may also be considered mathematical calculations when the broadest reasonable interpretation of the claim in light of the specification encompasses a mathematical calculation.
2A – Prong 2: This judicial exception is not integrated into a practical application. In particular, claim 1 recites the additional elements:
“receiving multi modal inputs from a user, the multi modal inputs comprising at least two different input modality types” amount to insignificant extra solution activity like mere data gathering, MPEP 2106.05(g)).
“readable medium” and “computer” merely uses a computer as a tool to perform an abstract idea, MPEP 2106.05(f)). These computer components are recited at a high-level of generality (i.e., as a generic processor performing a generic computer function of state transition probability calculation) such that it amounts no more than mere instructions to apply the exception using a generic computer component.
2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
“receiving multi modal inputs from a user, the multi modal inputs comprising at least two different input modality types” amount to insignificant extra solution activity like mere data gathering, MPEP 2106.05(g)).
“readable medium” and “computer” merely uses a computer as a tool to perform an abstract idea, MPEP 2106.05(f)). These computer components are recited at a high-level of generality (i.e., as a generic processor performing a generic computer function of state transition probability calculation) such that it amounts no more than mere instructions to apply the exception using a generic computer component.
Claim 31:
Step 1: The claim is directed to a method, falling under one of the four statutory categories of invention.
Step 2A Prong 1: The claim recites following abstract ideas:
The limitations “encoding, with an encoder of the encoder decoder architecture, features of the multi modal inputs to form the multi modal prompt, the multi modal prompt comprising embedded features of mixed modalities from the at least two different input modality types”; and “providing the prompt to a decoder of the encoder decoder architecture to cause the decoder to output the response based on the multi modal prompt, the decoder configured to output the response without prior training on at least one of the multi modal inputs received from the user” 2106.04(a)(2)(I)(C) “Mathematical Calculations A claim that recites a mathematical calculation, when the claim is given its broadest reasonable interpretation in light of the specification, will be considered as falling within the "mathematical concepts" grouping. A mathematical calculation is a mathematical operation (such as multiplication) or an act of calculating using mathematical methods to determine a variable or number, e.g., performing an arithmetic operation such as exponentiation. There is no particular word or set of words that indicates a claim recites a mathematical calculation. That is, a claim does not have to recite the word "calculating" in order to be considered a mathematical calculation. For example, a step of "determining" a variable or number using mathematical methods or "performing" a mathematical operation may also be considered mathematical calculations when the broadest reasonable interpretation of the claim in light of the specification encompasses a mathematical calculation.
2A – Prong 2: This judicial exception is not integrated into a practical application. In particular, claim 1 recites the additional elements:
“receiving multi modal inputs from a user, the multi modal inputs comprising at least two different input modality types” amount to insignificant extra solution activity like mere data gathering, MPEP 2106.05(g)).
2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
“receiving multi modal inputs from a user, the multi modal inputs comprising at least two different input modality types” amount to insignificant extra solution activity like mere data gathering, MPEP 2106.05(g)).
Claim 2 recites “wherein the multi modal inputs having the at least two different input modality types comprise two or more of text, image, video, audio, signal, byte sequence, code, and electromagnetic inputs” amount to insignificant extra solution activity like mere data gathering, MPEP 2106.05(g)).
Claim 3 recites “wherein the electromagnetic inputs comprise radiofrequency (RF) waves, microwaves, light waves, and/or infrared radiation” amount to insignificant extra solution activity like mere data gathering, MPEP 2106.05(g)).
Claim 4 recites “wherein the at least two different input modality types comprises at least three different input modality types amount to insignificant extra solution activity like mere data gathering, MPEP 2106.05(g))..
Claim 5 recites “wherein the operations further comprise receiving context information from the user, encoding the context information “amount to insignificant extra solution activity like mere data gathering, MPEP 2106.05(g), and causing the decoder to output the response based on the multi modal prompt and encoded context information amount to insignificant extra solution activity like mere data gathering, MPEP 2106.05(g).
Claim 6 recites “wherein the encoder need not be retrained to encode different multimodal inputs from the user, and instead is configured to be reused; and wherein the encoder is configured to encode both the features of the multi modal inputs to form the multi modal prompt and the context information to feed the decoder directly, without any added layers for combining features of different modes” amounts to no more than mere instructions to apply the exception using a generic computer component. .
Claim 7 recites “wherein the trained parameterized model comprises a large language model” amounts to no more than mere instructions to apply the exception using a generic computer component.
Claim 8 recites “wherein the trained parameterized model comprises a transformer” amounts to no more than mere instructions to apply the exception using a generic computer component.
Claim 9 recites “wherein the trained parameterized model further comprises a parietal space amount to insignificant extra solution activity like mere data gathering, MPEP 2106.05(g)..
Claim 10 recites “wherein the parameterized model comprises one or more neural networks” amounts to no more than mere instructions to apply the exception using a generic computer component.
Claim 11 recites “wherein the encoder comprises a first neural network amounts to no more than mere instructions to apply the exception using a generic computer component.
Claim12 recites “wherein the decoder comprises a second neural network” amounts to no more than mere instructions to apply the exception using a generic computer component.
Claim 13 recites “wherein the trained parameterized model and/or the encoder decoder architecture comprises one or more adapters” amounts to no more than mere instructions to apply the exception using a generic computer component.
Claim 14 recites “wherein the multi modal prompt comprises a single prompt, no matter how many different input modality types are included in the multi modal inputs received from the user” amount to insignificant extra solution activity like mere data gathering, MPEP 2106.05(g).
Claim 15 recites “wherein only key features of each of the multi modal inputs are encoded to form the multi modal prompt such that the multi modal prompt is relatively low dimensional compared to a dimensionality of any of the multi modal inputs, the key features being more predictive than other features of correct outputs during training of the parameterized model” amount to insignificant extra solution activity like mere data gathering, MPEP 2106.05(g).
Claim recites “wherein training of the parameterized model is supervised or unsupervised”. amount to insignificant extra solution activity like mere data gathering, MPEP 2106.05(g).
Claim 17 recites “wherein the training configures the parameterized model to learn a generic associativity of multi modal prompts, and once trained, to be deployed to output the zero-shot learning response to the multi modal prompt, without finetuning on new data types” amount to insignificant extra solution activity like mere data gathering, MPEP 2106.05(g).
Claim 18 recites “wherein the parameterized model is configured to solve a task involving new multi modal inputs by finding a closest match to the multi modal prompt in an embedding space, and then assigning the multi modal prompt to a most relevant class based on a similarity of the multi modal prompt to the most relevant class; wherein the decoder comprises a transformer decoder; and wherein, given a new input modality feature, the transformer decoder is finetuned for a task that uses the new input modality of the feature, such that the parameterized model adapts how to best project input features into an internal embedding space of the parameterized model” amount to insignificant extra solution activity like mere data gathering, MPEP 2106.05(g)..
Claim 19 recites “wherein the decoder comprises a multi-attention head configured to receive the multi modal prompt and guide generation of the output response” amount to insignificant extra solution activity like mere data gathering, MPEP 2106.05(g)..
Claim 20 recites “wherein the multi modal inputs having the at least two different input modality types comprise a first input comprising text, and a second input comprising an image, a video, audio input, a signal, a byte sequence, code, or an electromagnetic input” amount to insignificant extra solution activity like mere data gathering, MPEP 2106.05(g)..
Claim 21 recites “wherein the multi modal inputs having the at least two different input modality types comprise a first input comprising an image, a video, audio input, a signal, a byte sequence, code, or an electromagnetic input, and a second input comprising a different one of the image, video, audio input, signal, byte sequence, code, or electromagnetic input” amount to insignificant extra solution activity like mere data gathering, MPEP 2106.05(g).
Claim 22 recites “wherein encoding the features of the multi modal inputs to form the multi modal prompt and outputting the zero-shot learning response to the multi modal prompt decouples a training dataset from application of the parameterized model such that the parameterized model is trained to have generic associativity capabilities instead of outputting responses based a particular training dataset” amount to insignificant extra solution activity like mere data gathering, MPEP 2106.05(g).
Claim 23 recites “wherein at least a portion of the response output by the trained parameterized model is provided as feedback to the trained parameterized model” amounts to no more than mere instructions to apply the exception using a generic computer component.
Claim 24 recites “wherein the portion of the response output by the trained parameterized model provided as feedback is used as input for subsequent responses by the trained parameterized model” amount to insignificant extra solution activity like mere data gathering, MPEP 2106.05(g).
Claim 25 recites “wherein the feedback is configured to iteratively refine the input to the trained parameterized model, while the trained parameterized model itself remains the same” amount to insignificant extra solution activity like mere data gathering, MPEP 2106.05(g).
Claim 26 recites “wherein the feedback comprises code and/or output of executed code.
Claim 27 recites “wherein the feedback is used as input that is separate from, and in addition to, the multi modal inputs from the user” amount to insignificant extra solution activity like mere data gathering, MPEP 2106.05(g).
Claim 28 recites “wherein the trained parameterized model is configured to store embedded features of mixed modalities from prior prompts in a feature database to create a library of features, to be used in combination with later prompts and/or context information to output responses” amount to insignificant extra solution activity like mere data gathering, MPEP 2106.05(g).
Claim 29 recites “wherein using stored features to output responses to later prompts comprises performing a hierarchical feature search of the feature database and/or an external database to efficiently identify features related to a user query that can be provided as input to the trained parameterized model” amount to insignificant extra solution activity like mere data gathering, MPEP 2106.05(g).
Claim 30 recites “wherein the parameterized model is configured to solve a task involving new multi modal inputs by finding a closest match to the multi modal prompt in an embedding space, based on a result of the hierarchical feature search and/or the context information, and then assigning the multi modal prompt to a most relevant class based on a similarity of the multi modal prompt to the most relevant class” amount to insignificant extra solution activity like mere data gathering, MPEP 2106.05(g).
Claims 31-60 are similar in scope as claims 1-30, respectively; therefore, they are rejected under the same rationale.
Claim Rejections - 35 USC § 103
4. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
5. Claim(s) 1-2, 4-6, 8-22, 28, 31-32, 34-36, 38-52, 58 is/are rejected under 35 U.S.C. 103 as being unpatentable over Dirik et al. (A Dive into Vision-Language Models, 2023) in view of ZHANG et al. (CN 115393849 A).
Claim 1. Dirik discloses a non-transitory computer readable medium having instructions thereon, the instructions when executed by a computer, causing the computer to output a zero-shot learning response to a multi modal prompt (multimodal learning involving image, video, text, audio, body gestures, facial expressions, and physiological signals… zero-shot image classification in which an image and prompts are provided to obtain a predicted result)(Introduction) using a trained parameterized model, the trained parameterized model comprising encoder decoder architecture (Transformer-based image and text encoders and describes a VisionEncoderDecoderModel having a pretrained Transformer vision model as encoder and a pretrained language model as decoder) (Learning Strategies), the instructions causing the computer to perform operations comprising:
receiving multi modal inputs from a user, the multi modal inputs comprising at least two different input modality types (image and natural-language modalities and identifies image, video, text, audio and physiological inputs as multimodal inputs) (Introduction); and
providing the prompt to a decoder of the encoder decoder architecture to cause the decoder to output the response based on the multi modal prompt, the decoder configured to output the response without prior training on at least one of the multi modal inputs received from the user (zero-shot image classification and zero-shot video classification, including use of pretrained models for zero-shot downstream tasks…) (abstract, Sections: Introduction, Contrastive Learning Supporting Vision-Language Models in Transformers).
Dirik does not explicitly disclose encoding, with an encoder of the encoder decoder architecture, features of the multi modal inputs to form the multi modal prompt, the multi modal prompt comprising embedded features of mixed modalities from the at least two different input modality types.
However, ZHANG discloses disclose encoding, with an encoder of the encoder decoder architecture, features of the multi modal inputs to form the multi modal prompt, the multi modal prompt comprising embedded features of mixed modalities from the at least two different input modality types (obtains a visual prompt vector from image features, encodes the visual prompt and business text, and generates corresponding visual and textual encoding vectors…. then combines the visual prompt vector and entity prompt vector to produce a multimodal prompt vector and performs autoregressive decoding) (pp. 3, 6, 10, 12) Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Dirik further in view of ZHANG to incorporate the above cited feature. One would have been motivated to do so to a predictable multimodal, prompt-driven, zero-shot encoder/decoder system.
Claim 2. Dirik and ZHANG disclose the medium of claim 1, Dirik further discloses wherein the multi modal inputs having the at least two different input modality types comprise two or more of text, image, video, audio, signal, byte sequence, code, and electromagnetic inputs (identify image, video, text, audio, body gestures, facial expressions and physiological signals as modalities used in multimodal learning… video/text datasets and models supporting video-text zero-shot classification) (Introduction).
Claim 4. Dirik and ZHANG disclose the medium of claim 1, Dirik further discloses wherein the at least two different input modality types comprises at least three different input modality types (numerous multimodal combinations, including image, video, text, audio, body gestures, facial expressions and physiological signals..) (Introduction).
Claim 5. Dirik and ZHANG disclose the medium of claim 1, ZHANG further discloses wherein the operations further comprise receiving context information from the user, encoding the context information, and causing the decoder to output the response based on the multi modal prompt and encoded context information (the complementary encoder/decoder architecture in which encoded multimodal information is provided to the decoder) (p. 3). One would have been motivated to do so to a predictable multimodal, prompt-driven, zero-shot encoder/decoder system.
Claim 6. Dirik and ZHANG disclose the medium of claim 5, Dirik further discloses wherein the encoder need not be retrained to encode different multimodal inputs from the user, and instead is configured to be reused; and wherein the encoder is configured to encode both the features of the multi modal inputs to form the multi modal prompt and the context information to feed the decoder directly, without any added layers for combining features of different modes (the use of pretrained image and language models and expressly notes that such models may be mixed and matched, including pretrained vision models used as encoders and pretrained language models as decoders) (Supporting Vision-Language Models in Transformers).
Claim 8. Dirik and ZHANG disclose the medium of claim 1, Dirik further discloses wherein the trained parameterized model comprises a transformer (states that contemporary image/text encoders predominantly employ Transformer architectures) (Supporting Vision-Language Models in Transformers).
Claim 9. Dirik and ZHANG disclose the medium of claim 8, Dirik further discloses wherein the trained parameterized model further comprises a parietal space (mapping image and text inputs to the same feature space, such that distances between embeddings indicate correspondence… Aligning images and texts to a joint feature space in a contrastive manner) (p. 4 Contrastive Learning).
Claim 10. Dirik and ZHANG disclose the medium of claim 1, ZHANG further discloses wherein the parameterized model comprises one or more neural networks (a visual-language pretrained model, language-model encoding, and language-model decoding.. based on the visual characteristic of the sample service image extracted by the visual-language pre-training model, and mapping the visual characteristic of the sample service image to the input space of the pre-training language model based on the initial multilayer sensing network, obtaining the sample visual prompting vector) (p. 3). One would have been motivated to do so to a predictable multimodal, prompt-driven, zero-shot encoder/decoder system.
Claim 11. Dirik and ZHANG disclose the medium of claim 1, ZHANG further discloses wherein the encoder comprises a first neural network (a visual-language pretrained model, language-model encoding, and language-model decoding) (p. 13). One would have been motivated to do so to a predictable multimodal, prompt-driven, zero-shot encoder/decoder system.
Claim 12. Dirik and ZHANG disclose the medium of claim 1, ZHANG further discloses wherein the decoder comprises a second neural network (a visual-language pretrained model, language-model encoding, and language-model decoding) (p. 13). One would have been motivated to do so to a predictable multimodal, prompt-driven, zero-shot encoder/decoder system.
Claim 13. Dirik and ZHANG disclose the medium of claim 1, Dirik further discloses wherein the trained parameterized model and/or the encoder decoder architecture comprises one or more adapters (the multimodal vision-language environment into which such parameter-efficient adaptation would be applied) (p. 3 Learning Strategies).
Claim 14. Dirik and ZHANG disclose the medium of claim 1, ZHANG further discloses wherein the multi modal prompt comprises a single prompt, no matter how many different input modality types are included in the multi modal inputs received from the user (..combine visual and entity prompt information into a multimodal prompt vector) (p. 3). One would have been motivated to do so to a predictable multimodal, prompt-driven, zero-shot encoder/decoder system.
Claim 15. Dirik and ZHANG disclose the medium of claim 1, Dirik further discloses wherein only key features of each of the multi modal inputs are encoded to form the multi modal prompt such that the multi modal prompt is relatively low dimensional compared to a dimensionality of any of the multi modal inputs, the key features being more predictive than other features of correct outputs during training of the parameterized model (extraction of image and text features and their projection into a common feature space… Aligning images and texts to a joint feature space in acontrastive manner)(pp. 3-5).
Claim 16. Dirik and ZHANG disclose the medium of claim 1, Dirik further discloses wherein training of the parameterized model is supervised or unsupervised (multiple pretraining strategies, including contrastive learning, PrefixLM, multimodal fusion, masked-language modeling and image-text matching.)( pp. 5,7 Contrastive Learning).
Claim 17. Dirik and ZHANG disclose the medium of claim 16, Dirik further discloses wherein the training configures the parameterized model to learn a generic associativity of multi modal prompts, and once trained, to be deployed to output the zero-shot learning response to the multi modal prompt, without finetuning on new data types (identify zero-shot generalization as a major capability of multimodal vision-language models and describes zero-shot image and video classification) (pp. 1-3, 15).
Claim 18. Dirik and ZHANG disclose the medium of claim 1, Dirik further discloses wherein the parameterized model is configured to solve a task involving new multi modal inputs by finding a closest match to the multi modal prompt in an embedding space, and then assigning the multi modal prompt to a most relevant class based on a similarity of the multi modal prompt to the most relevant class; wherein the decoder comprises a transformer decoder; and wherein, given a new input modality feature, the transformer decoder is finetuned for a task that uses the new input modality of the feature, such that the parameterized model adapts how to best project input features into an internal embedding space of the parameterized model (mapping image and text inputs into the same feature space and determining correspondence based upon the distance between embeddings)(pp. 4-6, 15).
Claim 19. Dirik and ZHANG disclose the medium of claim 1, Dirik further discloses wherein the decoder comprises a multi-attention head configured to receive the multi modal prompt and guide generation of the output response (multimodal fusion using cross-attention and identifies Transformer architectures for multimodal systems) (pp. 1, 2, 4, 6-7).
Claim 20. Dirik and ZHANG disclose the medium of claim 1, Dirik further discloses wherein the multi modal inputs having the at least two different input modality types comprise a first input comprising text, and a second input comprising an image, a video, audio input, a signal, a byte sequence, code, or an electromagnetic input (combinations involving text with image, video and audio, and discusses video-text and image-text models)(pp. 1-4).
Claim 21. Dirik and ZHANG disclose the medium of claim 1, Dirik further discloses wherein the multi modal inputs having the at least two different input modality types comprise a first input comprising an image, a video, audio input, a signal, a byte sequence, code, or an electromagnetic input, and a second input comprising a different one of the image, video, audio input, signal, byte sequence, code, or electromagnetic input (multimodal systems using image, video, audio and physiological signals… further identifies video-text datasets and multimodal models, including X-CLIP for video and text)(pp. 14-15, 19).
Claim 22. Dirik and ZHANG disclose the medium of claim 1, Dirik further discloses wherein encoding the features of the multi modal inputs to form the multi modal prompt and outputting the zero-shot learning response to the multi modal prompt decouples a training dataset from application of the parameterized model such that the parameterized model is trained to have generic associativity capabilities instead of outputting responses based a particular training dataset (pretrained multimodal models being transferred to various downstream tasks, including zero-shot image classification, image segmentation, object detection and visual question answering) (pp. 7, 9).
Claim 28. Dirik and ZHANG disclose the medium of claim 1, Dirik further discloses wherein the trained parameterized model is configured to store embedded features of mixed modalities from prior prompts in a feature database to create a library of features, to be used in combination with later prompts and/or context information to output responses (multimodal embeddings and datasets containing image/text and video/text pairs) (pp. 5, 12).
Claims 31, 32-36, 38-52 and 58 represent the method of claims 1-2, 4-6, 8-22, 28, respectively and are rejected along the same rationale.
5. Claim(s) 3 and 33 is/are rejected under 35 U.S.C. 103 as being unpatentable over Dirik et al. (A Dive into Vision-Language Models, 2023) in view of ZHANG et al. (CN 115393849 A) and further in view of Green et al. (US 9,499,185).
Claim 3. Dirik and ZHANG disclose the medium of claim 2, but fail to explicitly disclose wherein the electromagnetic inputs comprise radiofrequency (RF) waves, microwaves, light waves, and/or infrared radiation.
However, Green discloses wherein the electromagnetic inputs comprise radiofrequency (RF) waves, microwaves, light waves, and/or infrared radiation (Col. 3, lines 44-50). Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Dirik further in view of Green to incorporate the above cited feature. One would have been motivated to do so to provide more cost-effective.
Claims 33 represents the method of claim 3 and are rejected along the same rationale.
6. Claim(s) 7, 23-27, 29-30, 37, 53-57, and 59-60 is/are rejected under 35 U.S.C. 103 as being unpatentable over Dirik et al. (A Dive into Vision-Language Models, 2023) in view of ZHANG et al. (CN 115393849 A) and further in view of YOUNG (KR 102506404 B1).
Claim 7. Dirik and ZHANG disclose the medium of claim 1, but fail to explicitly disclose wherein the trained parameterized model comprises a large language model.
However, YOUNG discloses a trained language model having multiple Transformer blocks, including self-attention, cross-attention and feed-forward neural-network layers (p. 3, abstract). Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Dirik further in view of YOUNG to incorporate the above cited feature. One would have been motivated to do so to improve accuracy and speed in language model.
Claim 23. Dirik and ZHANG disclose the medium of claim 1, but fail to explicitly disclose wherein at least a portion of the response output by the trained parameterized model is provided as feedback to the trained parameterized model.
However, YOUNG discloses a prompt-generation pipeline in which generated/improved prompt information is supplied to a pretrained language model to produce decision-making text (abstract, pp. 6-8). Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Dirik further in view of YOUNG to incorporate the above cited feature. One would have been motivated to do so to improve accuracy and speed in language model.
Claim 24. Dirik ZHANG and YOUNG disclose the medium of claim 23, YOUNG further discloses wherein the portion of the response output by the trained parameterized model provided as feedback is used as input for subsequent responses by the trained parameterized model (Decision-making simulation device using a trained language model)(pp. 2-3). One would have been motivated to do so to improve accuracy and speed in language model.
Claim 25. Dirik ZHANG and YOUNG disclose the medium of claim 24, Dirik further discloses wherein the feedback is configured to iteratively refine the input to the trained parameterized model, while the trained parameterized model itself remains the same (p. 9).
Claim 26. Dirik ZHANG and YOUNG disclose the medium of claim 23, YOUNG further discloses wherein the feedback comprises code and/or output of executed code (expressly notes that GPT-3 can perform tasks including simple web coding (p. 7)…Some programs may be configured such that motion types are displayed as objects on the screen. For example,the operation form of some programs may be displayed in the form of an 'execution window' as an object on the screen. For example, the execution window may include a document editing window outputted as a document editing program is executed and a web browser window outputted as a web browser application is executed) (p.2). One would have been motivated to do so to improve accuracy and speed in language model.
Claim 27. Dirik ZHANG and YOUNG disclose the medium of claim 23, YOUNG further discloses wherein the feedback is used as input that is separate from, and in addition to, the multi modal inputs from the user (..the text encoder 20 according to an embodiment of the present invention refers toa downsampling artificial neural network module that uses raw prompt information in the form of text as input data and prompt vectors as output data. . Specifically, the text encoder 20 according to an embodiment of the present invention is a tokenization module (raw prompt information in the form of text information as input data and token information including at least one token (eg, [CLS] . It can include a module that takes input data and prompt vectors, which are tensors, as output data) (p.4). One would have been motivated to do so to improve accuracy and speed in language model.
Claim 29. Dirik ZHANG and YOUNG disclose the medium of claim 28, YOUNG further discloses wherein using stored features to output responses to later prompts comprises performing a hierarchical feature search of the feature database and/or an external database to efficiently identify features related to a user query that can be provided as input to the trained parameterized model (..the exemplary indexing module 60 according to an embodiment of the present invention. As shown in FIG. 6, the Example indexing module 60 according to an embodiment of the present invention includes an Example database in which text information of web pages and app pages of Wikipedia, blogs, GitHub, Instagram, Facebook, etc. Connected and pre-stored external text data, environment image vector, prompt vector, and task information are used as input data, and Example storage location information and prompt-task relevance score for the storage location of the Example text in the Example database are obtained. It includes an artificial neural network module (Example indexing artificial neural network module) as output data, generates text data whose prompt-task relevance score is higher than a specific value as Example text information, and transforms the generated Example text information into transformer-based Refers to a module that generates an Example text vector by encoding with an encoder)(p. 6). One would have been motivated to do so to improve accuracy and speed in language model.
Claim 30. Dirik ZHANG and YOUNG disclose the medium of claim 29, YOUNG further discloses wherein the parameterized model is configured to solve a task involving new multi modal inputs by finding a closest match to the multi modal prompt in an embedding space, based on a result of the hierarchical feature search and/or the context information, and then assigning the multi modal prompt to a most relevant class based on a similarity of the multi modal prompt to the most relevant class (a multimodal classification architecture receiving a cross-attention representation and outputting a task class and confidence score)(abstract, p. 7). One would have been motivated to do so to improve accuracy and speed in language model.
Claims 37, 53-57 and 59-60 represent the method of claims 7, 23-27 and 29-30, respectively and are rejected along the same rationale.
Conclusion
8. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure (See PTO-892).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Phenuel S. Salomon whose telephone number is (571) 270-1699. The examiner can normally be reached on Mon-Fri 7:00 A.M. to 4:00 P.M. (Alternate Friday Off) EST.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Usmaan Saeed can be reached on (571) 272-4046. The fax phone number for the organization where this application or proceeding is assigned is 571-273-3800.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PHENUEL S SALOMON/Primary Examiner, Art Unit 2146