DETAILED ACTION
This action is in response to the submission filed 29 December 2023 for application 18/400,477. Currently claims 1-20 are pending and have been examined.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Acknowledgment is made of applicants’ claim for domestic priority based on the provisional application 63/500,551 filed on 05 May 2023.
Information Disclosure Statement
Information disclosure statements (IDS) were submitted on 27 March 2024, 01 November 2024, 29 July 2025, and 09 October 2025. The submissions are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statements are being considered by the examiner.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1 - 20 are rejected under 35 U.S.C. 101 because the claimed invention is directed towards abstract ideas without significantly more.
Regarding claims 1-7:
According to the first step (Step 1) of the 101 analysis, claims 1-7 are directed to a method of generating a multi-modal task output for a text instruction relating to a plurality of inputs of different modalities (process) and falls within one of the four statutory categories (i.e., process, machine, manufacture, or composition of matter).
Regarding claim 1:
In step (Step 2A, prong 1) of the analysis, the limitations of:
encoding, the first input of the first modality into a first encoded representation conditioned on the text instruction;
encoding, the second input of the second modality into a second encoded representation conditioned on the text instruction;
and generating, the multi-modal task output based on an input combining the first encoded representation, the second encoded representation, and the text instruction.
Under the broadest reasonable interpretation, the above limitations are process steps that cover mental processes including an observation, evaluation, judgment or opinion that could be performed in the mind or with the aid of pencil and paper but for the recitation of a generic computer component because for example, one can look at the first and second encoded representations (which could be descriptions of an image and video) of the different inputs and the text instruction (which could be a question) along with it and evaluate it together to generate an output (which could be an answer to the question) . If a claim, under its broadest reasonable interpretation, covers a mental process but for the recitation of generic computer components, then it falls within the “Mental Process” grouping of abstract ideas.
In the next step (Step 2A, prong 2) of the analysis, the limitations:
encoding, by a first multimodal encoder adapted for the first modality;
encoding, by a second multimodal encoder adapted for the second modality;
by a neural network based language model
are considered to be additional elements and it does not integrate the abstract idea into a practical application because the additional elements are recited so generically (no details whatsoever are provided other than that it is a method of encoding, by a first multimodal encoder adapted for the first modality; encoding, by a second multimodal encoder adapted for the second modality; by a neural network based language model) that it represents no more than mere instructions to apply the judicial exception on a computer. As discussed in MPEP 2106.05(f), mere instructions to implement an abstract idea on a computer as a tool to perform an abstract idea is not indicative of integration into a practical application.
In the same step (Step 2A, prong 2) of the analysis, the limitation:
receiving, via a data interface, a first input of a first modality, a second input of a second modality, and the text instruction relating to the first and the second inputs;
is considered to be an additional element and as recited represents insignificant extra-solution activity because it is mere data gathering. See MPEP 2106.05(g), discussing limitations that the Federal Circuit has considered to be insignificant extra-solution activity.
Accordingly, at Step 2A, prong two, after considering all claim elements individually and as an ordered combination, it is determined that the claims do not integrate the judicial exception into a practical application.
In the last step (Step 2B) of the analysis, the additional elements do not amount to significantly more than the judicial exceptions. As explained with respect to Step 2A Prong Two, the method of encoding, by a first multimodal encoder adapted for the first modality; encoding, by a second multimodal encoder adapted for the second modality; by a neural network based language model, is at best the equivalent of merely adding the words “apply it” to the judicial exception. See MPEP 2106.05(f). Even when considered in combination, mere instructions to apply an exception cannot provide an inventive concept and does not amount to significantly more than the judicial exception.
In the same step (Step 2B) of the analysis, as discussed above the additional element of receiving, via a data interface, a first input of a first modality, a second input of a second modality, and the text instruction relating to the first and the second inputs, which is recited at a high level of generality and amounts to extra-solution activity of receiving data i.e. pre-solution activity of gathering data for use in the claimed process. The courts have found limitations directed to obtaining information electronically, recited at a high level of generality, to be well-understood, routine, and conventional (see MPEP 2106.05(d)(II), “receiving or transmitting data over a network”, "electronic record keeping," and "storing and retrieving information in memory"). These limitations therefore remain insignificant extra-solution activity even upon reconsideration, and do not amount to significantly more.
Accordingly, at Step 2B, after considering all claim elements individually and as an ordered combination, it is determined that the claims do not amount to significantly more than the judicial exception. The claim is not patent eligible.
Regarding claim 2:
In the next step (Step 2A, prong 2) of the analysis, the limitations:
wherein the first modality is one of: image, video, audio, or 3D.
However, it does not integrate the abstract idea into a practical application because the additional element is generally linking the use of a judicial exception to a particular technological environment such as image, video, audio, or 3D. As discussed in MPEP 2106.05(h), generally linking the use of a judicial exception to a particular technological environment or field of use is not indicative of integration into a practical application.
Accordingly, at Step 2A, prong two, after considering all claim elements individually and as an ordered combination, it is determined that the claims do not integrate the judicial exception into a practical application.
In the last step (Step 2B) of the analysis, the additional element does not amount to significantly more than the abstract idea because the additional element does not meaningfully limit the judicial exception when considered both individually and as a combination. As discussed in MPEP 2106.05(h), employing generic computer functions to execute an abstract idea, even when limiting the use of the idea to one particular environment, it does not add significantly more, and thus fails to add an inventive concept to the claim. Even when considered in combination, mere instructions to apply an exception cannot provide an inventive concept and does not amount to significantly more than the judicial exception.
Accordingly, at Step 2B, after considering all claim elements individually and as an ordered combination, it is determined that the claims do not amount to significantly more than the judicial exception. The claim is not patent eligible.
Regarding claim 3:
In the next step (Step 2A, prong 2) of the analysis, the limitation:
wherein the second modality is a different modality than the first modality, and wherein the second modality is one of: image, video, audio, or 3D.
is considered to be an additional element and it does not integrate the abstract idea into a practical application because the additional element is generally linking the use of a judicial exception to a particular technological environment such as image, video, audio, or 3D. As discussed in MPEP 2106.05(h), generally linking the use of a judicial exception to a particular technological environment or field of use is not indicative of integration into a practical application.
Accordingly, at Step 2A, prong two, after considering all claim elements individually and as an ordered combination, it is determined that the claims do not integrate the judicial exception into a practical application.
In the last step (Step 2B) of the analysis, the additional element does not amount to significantly more than the abstract idea because the additional element does not meaningfully limit the judicial exception when considered both individually and as a combination. As discussed in MPEP 2106.05(h), employing generic computer functions to execute an abstract idea, even when limiting the use of the idea to one particular environment, it does not add significantly more, and thus fails to add an inventive concept to the claim. Even when considered in combination, mere instructions to apply an exception cannot provide an inventive concept and does not amount to significantly more than the judicial exception.
Accordingly, at Step 2B, after considering all claim elements individually and as an ordered combination, it is determined that the claims do not amount to significantly more than the judicial exception. The claim is not patent eligible.
Regarding claim 4:
In step (Step 2A, prong 1) of the analysis, the limitations of:
encoding, adapted for the one or more additional modalities, the one or more additional inputs of one or more additional modalities into additional encoded representations conditioned on the text instruction;
wherein the generating the multi-modal task output is further based on the additional encoded representations.
Under the broadest reasonable interpretation, the above limitations are process steps that cover mental processes including an observation, evaluation, judgment or opinion that could be performed in the mind or with the aid of pencil and paper but for the recitation of a generic computer component because for example, one can look at the additional encoded representations and evaluate it to generate the multi-modal task output. If a claim, under its broadest reasonable interpretation, covers a mental process but for the recitation of generic computer components, then it falls within the “Mental Process” grouping of abstract ideas.
In the next step (Step 2A, prong 2) of the analysis, the limitations:
by respective multimodal encoders adapted for the one or more additional modalities;
are considered to be additional elements and it does not integrate the abstract idea into a practical application because the additional elements are recited so generically (no details whatsoever are provided other than that it is a method of encoding, by respective multimodal encoders adapted for the one or more additional modalities) that it represents no more than mere instructions to apply the judicial exception on a computer. As discussed in MPEP 2106.05(f), mere instructions to implement an abstract idea on a computer as a tool to perform an abstract idea is not indicative of integration into a practical application.
In the same step (Step 2A, prong 2) of the analysis, the limitation:
receiving, via the data interface, one or more additional inputs of one or more additional modalities;
is considered to be an additional element and as recited represents insignificant extra-solution activity because it is mere data gathering. See MPEP 2106.05(g), discussing limitations that the Federal Circuit has considered to be insignificant extra-solution activity.
Accordingly, at Step 2A, prong two, after considering all claim elements individually and as an ordered combination, it is determined that the claims do not integrate the judicial exception into a practical application.
In the last step (Step 2B) of the analysis, the additional elements do not amount to significantly more than the judicial exceptions. As explained with respect to Step 2A Prong Two, the method of encoding, by respective multimodal encoders adapted for the one or more additional modalities, is at best the equivalent of merely adding the words “apply it” to the judicial exception. See MPEP 2106.05(f). Even when considered in combination, mere instructions to apply an exception cannot provide an inventive concept and does not amount to significantly more than the judicial exception.
In the same step (Step 2B) of the analysis, as discussed above the additional element of receiving, via the data interface, one or more additional inputs of one or more additional modalities, which is recited at a high level of generality and amounts to extra-solution activity of receiving data i.e. pre-solution activity of gathering data for use in the claimed process. The courts have found limitations directed to obtaining information electronically, recited at a high level of generality, to be well-understood, routine, and conventional (see MPEP 2106.05(d)(II), “receiving or transmitting data over a network”, "electronic record keeping," and "storing and retrieving information in memory"). These limitations therefore remain insignificant extra-solution activity even upon reconsideration, and do not amount to significantly more.
Accordingly, at Step 2B, after considering all claim elements individually and as an ordered combination, it is determined that the claims do not amount to significantly more than the judicial exception. The claim is not patent eligible.
Regarding claim 5:
In step (Step 2A, prong 1) of the analysis, the limitation of:
wherein the generating the multi-modal task output is further based on a first prefix indicating the first modality and a second prefix indicating the second modality.
Under the broadest reasonable interpretation, the above limitations are process steps that cover mental processes including an observation, evaluation, judgment or opinion that could be performed in the mind or with the aid of pencil and paper but for the recitation of a generic computer component because for example, one can look at the additional encoded representations and evaluate it to generate the multi-modal task output. If a claim, under its broadest reasonable interpretation, covers a mental process but for the recitation of generic computer components, then it falls within the “Mental Process” grouping of abstract ideas.
In the next step (Step 2A, prong 2) of the analysis, it does not integrate into a practical application because it does not add any additional elements that integrate the abstract idea into practical application.
Accordingly, at Step 2A, prong two, after considering all claim elements individually and as an ordered combination, it is determined that the claims do not integrate the judicial exception into a practical application.
In the last step (Step 2B) of the analysis, it does not add any additional elements that amount to significantly more than the abstract idea and thus fails to add an inventive concept.
Accordingly, at Step 2B, after considering all claim elements individually and as an ordered combination, it is determined that the claims do not amount to significantly more than the judicial exception. The claim is not patent eligible.
Regarding claim 6:
In the next step (Step 2A, prong 2) of the analysis, the limitations:
encoding, by a modality-specific encoder, the first input into a first vector representation, wherein the encoding the first input of the first modality into the first encoded representation includes generating, by the first multimodal encoder, the first encoded representation based on cross-attending the first vector representation to the text instruction.
is considered to be an additional element and it does not integrate the abstract idea into a practical application because the additional element is recited so generically (no details whatsoever are provided other than encoding, by a modality-specific encoder, the first input into a first vector representation, wherein the encoding the first input of the first modality into the first encoded representation includes generating, by the first multimodal encoder, the first encoded representation based on cross-attending the first vector representation to the text instruction) that it represents no more than mere instructions to apply the judicial exception on a computer. As discussed in MPEP 2106.05(f), mere instructions to implement an abstract idea on a computer as a tool to perform an abstract idea is not indicative of integration into a practical application.
Accordingly, at Step 2A, prong two, after considering all claim elements individually and as an ordered combination, it is determined that the claims do not integrate the judicial exception into a practical application.
In the last step (Step 2B) of the analysis, the additional element does not amount to significantly more than the judicial exceptions. As explained with respect to Step 2A Prong Two, the method wherein encoding, by a modality-specific encoder, the first input into a first vector representation, wherein the encoding the first input of the first modality into the first encoded representation includes generating, by the first multimodal encoder, the first encoded representation based on cross-attending the first vector representation to the text instruction, is at best the equivalent of merely adding the words “apply it” to the judicial exception. See MPEP 2106.05(f). Even when considered in combination, mere instructions to apply an exception cannot provide an inventive concept and does not amount to significantly more than the judicial exception.
Accordingly, at Step 2B, after considering all claim elements individually and as an ordered combination, it is determined that the claims do not amount to significantly more than the judicial exception. The claim is not patent eligible.
Regarding claim 7:
In the next step (Step 2A, prong 2) of the analysis, the limitations:
wherein the encoding the first input of the first modality into the first encoded representation further includes cross-attending a plurality of vector queries to the text instruction.
is considered to be an additional element and it does not integrate the abstract idea into a practical application because the additional element is recited so generically (no details whatsoever are provided other than wherein the encoding the first input of the first modality into the first encoded representation further includes cross-attending a plurality of vector queries to the text instruction) that it represents no more than mere instructions to apply the judicial exception on a computer. As discussed in MPEP 2106.05(f), mere instructions to implement an abstract idea on a computer as a tool to perform an abstract idea is not indicative of integration into a practical application.
Accordingly, at Step 2A, prong two, after considering all claim elements individually and as an ordered combination, it is determined that the claims do not integrate the judicial exception into a practical application.
In the last step (Step 2B) of the analysis, the additional element does not amount to significantly more than the judicial exceptions. As explained with respect to Step 2A Prong Two, the method wherein the encoding the first input of the first modality into the first encoded representation further includes cross-attending a plurality of vector queries to the text instruction, is at best the equivalent of merely adding the words “apply it” to the judicial exception. See MPEP 2106.05(f). Even when considered in combination, mere instructions to apply an exception cannot provide an inventive concept and does not amount to significantly more than the judicial exception.
Accordingly, at Step 2B, after considering all claim elements individually and as an ordered combination, it is determined that the claims do not amount to significantly more than the judicial exception. The claim is not patent eligible.
Regarding claims 8-14:
According to the first step (Step 1) of the 101 analysis, claims 8-14 are directed to a system for generating a multi-modal task output for a text instruction relating to a plurality of inputs of different modalities, the system comprising: a memory that stores a neural network based language model and a plurality of processor executable instructions (manufacture) and falls within one of the four statutory categories (i.e., process, machine, manufacture, or composition of matter).
Regarding claim 8:
In step (Step 2A, prong 2) of the analysis, the limitation of:
system for generating a multi-modal task output for a text instruction relating to a plurality of inputs of different modalities, the system comprising: a memory that stores a neural network based language model and a plurality of processor executable instructions:
is considered to be an additional element and it does not integrate the abstract idea into a practical application because the additional element is recited so generically (no details whatsoever are provided other than that it is a system for generating a multi-modal task output for a text instruction relating to a plurality of inputs of different modalities, the system comprising: a memory that stores a neural network based language model and a plurality of processor executable instructions) that it represents no more than mere instructions to apply the judicial exception on a computer. As discussed in MPEP 2106.05(f), mere instructions to implement an abstract idea on a computer as a tool to perform an abstract idea is not indicative of integration into a practical application.
The rest of the limitations of claim 8 are substantially similar to claim 1 and therefore is rejected on similar grounds as claim 1 as explained above.
Regarding claim 9:
Claim 9 is substantially similar to claim 2 and therefore is rejected on similar grounds as claim 2.
Regarding claim 10:
Claim 10 is substantially similar to claim 3 and therefore is rejected on similar grounds as claim 3.
Regarding claim 11:
Claim 11 is substantially similar to claim 4 and therefore is rejected on similar grounds as claim 4.
Regarding claim 12:
Claim 12 is substantially similar to claim 5 and therefore is rejected on similar grounds as claim 5.
Regarding claim 13:
Claim 13 is substantially similar to claim 6 and therefore is rejected on similar grounds as claim 6.
Regarding claim 14:
Claim 14 is substantially similar to claim 7 and therefore is rejected on similar grounds as claim 7.
Regarding claims 15-20:
According to the first step (Step 1) of the 101 analysis, claims 15-20 are directed to a non-transitory machine-readable medium comprising a plurality of machine-executable instructions which, when executed by one or more processors, are adapted to cause the one or more processors to perform operations (manufacture) and falls within one of the four statutory categories (i.e., process, machine, manufacture, or composition of matter).
Regarding claim 15:
In step (Step 2A, prong 2) of the analysis, the limitation of:
A non-transitory machine-readable medium comprising a plurality of machine-executable instructions which, when executed by one or more processors, are adapted to cause the one or more processors to perform operations:
is considered to be an additional element and it does not integrate the abstract idea into a practical application because the additional element is recited so generically (no details whatsoever are provided other than that it is a non-transitory machine-readable medium comprising a plurality of machine-executable instructions which, when executed by one or more processors, are adapted to cause the one or more processors to perform operations) that it represents no more than mere instructions to apply the judicial exception on a computer. As discussed in MPEP 2106.05(f), mere instructions to implement an abstract idea on a computer as a tool to perform an abstract idea is not indicative of integration into a practical application.
The rest of the limitations of claim 15 are substantially similar to claim 1 and therefore is rejected on similar grounds as claim 1 as explained above.
Regarding claim 16:
Claim 16 is substantially similar to claims 2 and 3 and therefore is rejected on similar grounds as claim 2 and 3.
Regarding claim 17:
Claim 17 is substantially similar to claim 4 and therefore is rejected on similar grounds as claim 4.
Regarding claim 18:
Claim 18 is substantially similar to claim 5 and therefore is rejected on similar grounds as claim 5.
Regarding claim 19:
Claim 19 is substantially similar to claim 6 and therefore is rejected on similar grounds as claim 6.
Regarding claim 20:
Claim 20 is substantially similar to claim 7 and therefore is rejected on similar grounds as claim 7.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1, 2, 3, 5, 8-10, 12, 15, 16, and 18 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Zhu et al (Knowledge Transfer with Visual Prompt in Multi-modal Dialogue Understanding and Generation, 2022).
Regarding claim 1:
Zhu teaches: A method of generating a multi-modal task output for a text instruction relating to a plurality of inputs of different modalities, the method comprising: receiving, via a data interface, a first input of a first modality, a second input of a second modality, and the text instruction relating to the first and the second inputs ([Abstract] In this work, we propose a knowledge transfer method with visual prompt (VPTG) fusing multi-modal data, which is a flexible module that can utilize the text-only seq2seq model to handle VD tasks. The VPTG conducts text-image co-learning and multi-modal information fusion with visual prompts and visual knowledge distillation. Moreover, we also realize visual knowledge transfer through distillation between two different models’ text representations, so that the seq2seq model can actively learn visual semantic representations. [Page 8, Column 2, Paragraph 1] Figure 1: Description of the Multi-modal Dialogue Understanding and Generation (MDUG) task. From step1 to step 3, the video is about a priest, and the subtitles are snippets of wedding vows. For the response generation of step 4, supposing that only dialogue text context was taken, the previous dialog text: “OK, then” is inadequate for generating the expected output: “you may kiss the bride.” [Page 9, Column 1, Paragraph 1] In this work, we mainly focus on video visual dialogue such as the Multi-modal Dialogue Understanding and Generation (MDUG) dataset (Wang et al., 2022b). Compared to image captioning and image visual dialogue, it requires modeling long distance image sequences, which is more challenging and practical. The MDUG task proposes a multi-modal dialogue task in the video field. It needs the system to generate a response of the current frame based on multi-modal video scene and historical dialogue information, where historical video clips frame and text captions are mapped one-to-one. The video clips and visual images have much abundant and useful information about the plot development. It is easy to pick up on their movements and expressions from visual information. For example, in the last frame of Figure 1. On the one hand, from the body movements of people such as they gradually face each other and a smile on the man’s face, we can observe that the man is going to kiss his bride, so models can infer the “kiss” action in generated response. Note: First modality corresponds to video, first input of a first modality corresponds to the first frame of video in Figure 1. Second input of a second modality corresponds to image);
encoding, by a first multimodal encoder adapted for the first modality, the first input of the first modality into a first encoded representation conditioned on the text instruction ([Page 8, Column 2, Paragraph 1] Figure 1: Description of the Multi-modal Dialogue Understanding and Generation (MDUG) task. From step1 to step 3, the video is about a priest, and the subtitles are snippets of wedding vows. For the response generation of step 4, supposing that only dialogue text context was taken, the previous dialog text: “OK, then” is inadequate for generating the expected output: “you may kiss the bride.”. [Page 9, Column 2, Paragraph 2] In addition, to improve the visual modeling ability of language models, we conduct visual knowledge transfer by transferring visual representations to visual prompt and using it to prompt the seq2seq model modeling multi-modal data. Specifically, the “answer text” feature is also provided to the encoder output “[CLS]” vector of the seq2seq model for distillation. [Page 13, Column 1, Paragraph 3] In training LKL, it performs gradient decoupling (stop-gradient operator) for V CLIP text (x) and Encoder1. Note: Vclip is the video clip and corresponds to the 1st modality. Encoder 1 corresponds to a first multimodal encoder adapted for the first modality);
encoding, by a second multimodal encoder adapted for the second modality, the second input of the second modality into a second encoded representation conditioned on the text instruction ([Page 8, Abstract] The VPTG conducts text-image co-learning and multi-modal information fusion with visual prompts and visual knowledge distillation. Specifically, we construct visual prompts from visual representations and then induce sequence-to-sequence (seq2seq) models to fuse visual information and textual contexts by visual-text patterns. [Page 13, Column 1, Paragraph 3] where X is the training set of all image-text pairs. w0 ∈ Rd×k is a trainable weights vector. The text predictor encoder (Encoder2) is trained simultaneously by the response generation task. Note: Image corresponds to the 2nd modality. Encoder 2 corresponds to a second multimodal encoder adapted for the second modality);
and generating, by a neural network based language model, the multi-modal task output based on an input combining the first encoded representation, the second encoded representation, and the text instruction ([Page 10, Column 1, Section 2.1, Paragraph 2] In the VD task, given some frame or a video clip, a dialog history context, the agent has to ground in image and text, infer context from history, and generate text response accurately. It requires multi-dimensional modeling based on visual information to generate accurate descriptions, which has been used to help visually impaired people better understand the visual content of the environment. The MDUG dataset is a VD dataset that aims to generate an interactive response based on the image captions context history and video clips image content. The traditional multi-modal fusion method first uses the visual model to extract the image features and then uses the neural network such as LSTM. [Page 13, Column 1, Paragraph 1] assume that the last hidden state output among two encoders and text can be defined as p1(t | p) and p2(t | z). There are two transformer encoders in the VPTG, where we call the visual predictor encoder as Encoder1, the text predictor encoder as Encoder2).
Regarding claim 2:
Zhu teaches: The method of claim 1, wherein the first modality is one of: image, video, audio, or 3D ([Page 8, Column 2, Paragraph 1] Figure 1: Description of the Multi-modal Dialogue Understanding and Generation (MDUG) task. From step1 to step 3, the video is about a priest. Note: First modality corresponds to video in Figure 1).
Regarding claim 3:
Zhu teaches: The method of claim 2, wherein the second modality is a different modality than the first modality, and wherein the second modality is one of: image, video, audio, or 3D ([Abstract] The VPTG conducts text-image co-learning and multi-modal information fusion with visual prompts and visual knowledge distillation. [Page 9, Column 1, Paragraph 1] The video clips and visual images have much abundant and useful information about the plot development. [Page 10, Column 1, Paragraph 3] We present a useful method, which can be used in almost all seq2seq models. And it conducts visual prompts and visual knowledge transfer to jointly learn images and text, and effectively generate a response. We explore the task with multi-modal information representation, co-learning, and fusion. Note: Image corresponds to second modality and as shown above video corresponds to first modality. This shows that second modality is a different modality than the first modality).
Regarding claim 5:
Zhu teaches: The method of claim 1, wherein the generating the multi-modal task output is further based on a first prefix indicating the first modality and a second prefix indicating the second modality ([Page 9, Column 1, Paragraph 1] Compared to image captioning and image visual dialogue, it requires modeling long distance image sequences, which is more challenging and practical. The MDUG task proposes a multi-modal dialogue task in the video field. It needs the system to generate a response of the current frame based on multi-modal video scene and historical dialogue information, where historical video clips frame and text captions are mapped one-to-one. The video clips and visual images have much abundant and useful information about the plot development. It is easy to pick up on their movements and expressions from visual information. For example, in the last frame of Figure 1. On the one hand, from the body movements of people such as they gradually face each other and a smile on the man’s face, we can observe that the man is going to kiss his bride, so models can infer the “kiss” action in generated response. On the other hand, from the wedding vows context, it’s easy to infer their roles as bride and groom. Therefore, this example demonstrates the importance of combining images and texts for the MDUG task. Note: Image captioning corresponds to first prefix and text captions corresponds to a second prefix).
Regarding claim 8:
Zhu teaches: A system for generating a multi-modal task output for a text instruction relating to a plurality of inputs of different modalities, the system comprising: a memory that stores a neural network based language model and a plurality of processor executable instructions ([Abstract] In this work, we propose a knowledge transfer method with visual prompt (VPTG) fusing multi-modal data, which is a flexible module that can utilize the text-only seq2seq model to handle VD tasks. [Page 10, Column 1, Last Paragraph] The MDUG dataset is a VD dataset that aims to generate an interactive response based on the image captions context history and video clips image content. The traditional multi-modal fusion method first uses the visual model to extract the image features and then uses the neural network such as LSTM. Note: Neural network shows that there is a computer with memory and processor being used).
Claim 8 recites substantially same limitations as claim 1 except for the above limitation and is therefore rejected for same rationale as claim 1.
Regarding claim 9:
Claim 9 is substantially similar to claim 2 and therefore is rejected on similar grounds as claim 2.
Regarding claim 10:
Claim 10 is substantially similar to claim 3 and therefore is rejected on similar grounds as claim 3.
Regarding claim 12:
Claim 12 is substantially similar to claim 5 and therefore is rejected on similar grounds as claim 5.
Regarding claim 15:
Zhu teaches: A non-transitory machine-readable medium comprising a plurality of machine-executable instructions which, when executed by one or more processors, are adapted to cause the one or more processors to perform operations comprising: ([Abstract] In this work, we propose a knowledge transfer method with visual prompt (VPTG) fusing multi-modal data, which is a flexible module that can utilize the text-only seq2seq model to handle VD tasks. [Page 10, Column 1, Last Paragraph] The MDUG dataset is a VD dataset that aims to generate an interactive response based on the image captions context history and video clips image content. The traditional multi-modal fusion method first uses the visual model to extract the image features and then uses the neural network such as LSTM. Note: Neural network shows that there is a computer with a non-transitory machine-readable medium comprising a plurality of machine-executable instructions which, when executed by one or more processors, are adapted to cause the one or more processors to perform operations).
Claim 15 recites substantially same limitations as claim 1 except for the above limitation and is therefore rejected for same rationale as claim 1.
Regarding claim 16:
Claim 16 is substantially similar to claim 2 and therefore is rejected on similar grounds as claim 2.
Regarding claim 18:
Claim 18 is substantially similar to claim 5 and therefore is rejected on similar grounds as claim 5.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 4, 6, 7, 11, 13, 14, 17, 19, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Zhu et al (Knowledge Transfer with Visual Prompt in Multi-modal Dialogue Understanding and Generation, 2022) in view of Jaegle et al (Perceiver: General Perception with Iterative Attention, 2021).
Regarding claim 4:
Zhu teaches: The method of claim 1 (as shown above).
Zhu further teaches: conditioned on the text instruction ([Abstract] Moreover, we also realize visual knowledge transfer through distillation between two different models’ text representations, so that the seq2seq model can actively learn visual semantic representations).
However, Zhu does not explicitly disclose: further comprising: receiving, via the data interface, one or more additional inputs of one or more additional modalities; and encoding, by respective multimodal encoders adapted for the one or more additional modalities, the one or more additional inputs of one or more additional modalities into additional encoded representations, wherein the generating the multi-modal task output is further based on the additional encoded representations.
Jaegle teaches, in an analogous system: further comprising: receiving, via the data interface, one or more additional inputs of one or more additional modalities ([Abstract] Biological systems perceive the world by simultaneously processing high-dimensional inputs from modalities as diverse as vision, audition, touch, proprioception, etc. In this paper we introduce the Perceiver – a model that builds upon Transformers and hence makes few architectural assumptions about the relationship between its inputs, but that also scales to hundreds of thousands of inputs, like ConvNets. We show that this architecture is competitive with or outperforms strong, specialized models on classification tasks across various modalities: images, point clouds, audio, video, and video+audio. The Perceiver obtains performance comparable to ResNet-50 and ViT on ImageNet without 2D convolutions by directly attending to 50,000 pixels. It is also competitive in all modalities in AudioSet. [Page 4, Column 2, Paragraph 1] The second block shows performance when the inputs are RGB values concatenated with 2D Fourier features (FF) – the same that the Perceiver receives);
and encoding, by respective multimodal encoders adapted for the one or more additional modalities, the one or more additional inputs of one or more additional modalities into additional encoded representations, wherein the generating the multi-modal task output is further based on the additional encoded representations ([Page 8, Column 2, Paragraph 3] Audio + video. In this experiment we feed the Perceiver both the 12,544 space-time patches and either 480 raw audio vectors or 4,800 spectrogram values. Since modalities are fused at input, audio and video inputs need to have the same number of channels. We achieve this by concatenating a learned, modality-specific encoding to each input. As video has more channels, we use an embedding of size 4 for video inputs and make the audio encoding as large as necessary for the input channels between the two input arrays. This encoding doubles as a modality-specific position encoding (as discussed in Sec. 3.2), and we found it worked better than simply passing the audio encoding through a linear layer to match the video).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Zhu to incorporate the teachings of Jaegle to use receiving, via the data interface, one or more additional inputs of one or more additional modalities and encoding, by respective multimodal encoders adapted for the one or more additional modalities, the one or more additional inputs of one or more additional modalities into additional encoded representations, wherein the generating the multi-modal task output is further based on the additional encoded representations. One would have been motivated to do this modification because doing so would give the benefit of scaling to hundreds of thousands of inputs as taught by Jaegle [Abstract].
Regarding claim 6:
Zhu teaches: The method of claim 1 (as shown above).
Zhu further teaches: further comprising: encoding, by a modality-specific encoder, the first input into a first vector representation, wherein the encoding the first input of the first modality into the first encoded representation includes generating, by the first multimodal encoder, the first encoded representation based on the first vector representation to the text instruction ([Page 8, Column 2, Paragraph 1] Figure 1: Description of the Multi-modal Dialogue Understanding and Generation (MDUG) task. From step1 to step 3, the video is about a priest, and the subtitles are snippets of wedding vows. For the response generation of step 4, supposing that only dialogue text context was taken, the previous dialog text: “OK, then” is inadequate for generating the expected output: “you may kiss the bride.” Note: First modality corresponds to video, first input of a first modality corresponds to the first frame of video in Figure 1. [Page 13, Column 1, Paragraphs 1-2] There are two transformer encoders in the VPTG, where we call the visual predictor encoder as Encoder1, the text predictor encoder as Encoder2. p1(t | p) ∝ V CLIP text, p2(t | z) ∝ V seq2seq CLS (5) where t is input dialogue text, p is the input frame image; z is the visual prompt according to p; V CLIP text ∈ Rk is the representation of image in the visual predictor. The p1 represent the Encoder1, and the p2 represent the Encoder2. We close the gap between V seq2seq CLS and V CLIP text by minimizing the KL-divergence).
However, Zhu does not explicitly disclose: cross-attending.
Jaegle teaches, in an analogous system: cross-attending ([Page 4, Column 2, Paragraph 3] cross-attend. [Page 6, Column 2, Paragraph 1] We trained models for 120 epochs with an initial learning rate of 0.004, decaying it by a factor of 10 at [84, 102, 114] epochs. The best-performing Perceiver we identified on ImageNet attends to the input image 8 times, each time processing the full 50,176-pixel input array using a cross-attend module and a latent Transformer with 6 blocks and one cross-attend module with a single head per block. We found that sharing the initial cross-attention with subsequent cross-attends led to instability in training, so we share all cross-attends after the first. The dense subblock of each Transformer block doesn’t use a bottleneck. We used a latent array with 512 indices and 1024 channels, and position encodings generated with 64 bands and a maximum resolution of 224 pixels. On ImageNet, we found that models of this size overfit without weight sharing, so we use a model that shares weights for all but the first cross-attend and latent Transformer modules. The resulting model has ∼ 45 million parameters, making it comparable in size to convolutional models used on ImageNet.).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Zhu of generating an encoded representation to incorporate the teachings of Jaegle of using cross-attending. One would have been motivated to do this modification because doing so would give the benefit of better performance as taught by Jaegle [Page 4, Column 2, Paragraph 2].
Regarding claim 7:
The system of Zhu and Jaegle teaches: The method of claim 6 (as shown above).
Zhu further teaches: wherein the encoding the first input of the first modality into the first encoded representation further includes a plurality of vector queries to the text instruction ([Page 8, Column 2, Paragraph 1] Figure 1: Description of the Multi-modal Dialogue Understanding and Generation (MDUG) task. From step1 to step 3, the video is about a priest, and the subtitles are snippets of wedding vows. For the response generation of step 4, supposing that only dialogue text context was taken, the previous dialog text: “OK, then” is inadequate for generating the expected output: “you may kiss the bride.” Note: First modality corresponds to video and steps 1 through step 4 in Figure 1 corresponds to a plurality of vector queries to the text instruction).
However, Zhu does not explicitly disclose: cross-attending.
Jaegle further teaches, in an analogous system: cross-attending ([Page 4, Column 2, Paragraph 3] cross-attend).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Zhu to incorporate the teachings of Jaegle to use cross-attending. One would have been motivated to do this modification because doing so would give the benefit of better performance as taught by Jaegle [Page 4, Column 2, Paragraph 2].
Regarding claim 11:
Claim 11 is substantially similar to claim 4 and therefore is rejected on similar grounds as claim 4.
Regarding claim 13:
Claim 13 is substantially similar to claim 6 and therefore is rejected on similar grounds as claim 6.
Regarding claim 14:
Claim 14 is substantially similar to claim 7 and therefore is rejected on similar grounds as claim 7.
Regarding claim 17:
Claim 17 is substantially similar to claim 4 and therefore is rejected on similar grounds as claim 4.
Regarding claim 19:
Claim 19 is substantially similar to claim 6 and therefore is rejected on similar grounds as claim 6.
Regarding claim 20:
Claim 20 is substantially similar to claim 7 and therefore is rejected on similar grounds as claim 7.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Kwon et al (MASKED VISION AND LANGUAGE MODELING FOR MULTI-MODAL REPRESENTATION LEARNING, 2023) discloses how to use masked signal modeling in vision and language (V+L) representation learning. Instead of developing masked language modeling (MLM) and masked image modeling (MIM) independently, we propose to build joint masked vision and language modeling, where the masked signal of one modality is reconstructed with the help from another modality. This is motivated by the nature of image-text paired data that both of the image and the text convey almost the same information but in different formats. The masked signal reconstruction of one modality conditioned on another modality can also implicitly learn cross-modal alignment between language tokens and image patches. Our experiments on various V+L tasks show that the proposed method, along with common V+L alignment losses, achieves state-of-the-art performance in the regime of millions of pre-training data. Also, we outperform the other competitors by a significant margin in limited data scenarios.
Singh et al (US 20220230061 A1) discloses Method for Generating a Multimodal Question-answering Model for Providing Multi-modal Answer to Query, Involves Generating Response to Text-based Query Comprising Portion of Text Passage or Image Among Multiple Images According to Respective Relevance Scores.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to CHAITANYA RAMESH JAYAKUMAR whose telephone number is (571)272-3369. The examiner can normally be reached Mon-Fri 9am-1pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Omar Fernandez Rivas can be reached at (571)272-2589. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/C.R.J./ Examiner, Art Unit 2128
/OMAR F FERNANDEZ RIVAS/Supervisory Patent Examiner, Art Unit 2128