Prosecution Insights
Last updated: October 01, 2026
Application No. 18/639,519

Multimodal Learning from Structured and Unstructured Data

Non-Final OA §101§103
Filed
Apr 18, 2024
Priority
May 17, 2023 — provisional 63/467,120
Examiner
JUNG, DONG YOON
Art Unit
Tech Center
Assignee
Google LLC
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
16 currently pending
Career history
7
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§101 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority The present application has a provisional application No. 63/467,120 filed on May 17, 2023. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Regarding Claim 1 Step 1 – whether the claim falls within any statutory category. See MPEP 2016.03 Claim 1 is a method claim thus it falls into one of the four categories of statutory subject matter. Step 2A Prong 1 – whether the claim recites a judicial exception. See MPEP 2106.04, subsection II. Regarding independent claim 1, following limitations recite a judicial exception: “receiving, by one or more processors, multimodal training data comprising un-masked training examples and masked training examples” [Mental Process] – Receiving training data that comprises un-masked and masked, which are simply data that has been partially omitted or not, which involves observations, evaluations, judgments, and opinions that is capable of being performed in the human mind with the assistance of paper and pen “performing, by the one or more processors, one or more pretraining iterations” [Mental Process] – performing one or more pretraining iterations simply means it will continue the procedures certain number of counts, which involves observations, evaluations, judgments, and opinions that is capable of being performed in the human mind with the assistance of paper and pen “generating modality-specific encoded representations of the un-masked training examples and masked training examples” [Mental Process] – generating modality-specific encoded representations of training examples simply transforming the data according to its type, which involves observations, evaluations, judgments, and opinions that is capable of being performed in the human mind with the assistance of paper and pen “determining a plurality of modality-specific masking losses measuring the similarity between the modality-specific encoded representations of the un-masked training examples and the masked training examples” [Mental Process] – determining losses measuring similarity between two data sets requires to compare the transformed data, which involves observations, evaluations, judgments, and opinions that is capable of being performed in the human mind with the assistance of paper and pen “generating a first fused encoded representation of the un-masked training example encoded representations and a second fused encoded representation of the masked training example encoded representations” [Mental Process] – generating fused encoded representation of un-masked and masked training examples require to combine all types of examples as one, which involves observations, evaluations, judgments, and opinions that is capable of being performed in the human mind with the assistance of paper and pen “determining a multimodal masking loss measuring the similarity between the first fused encoded representation and the second fused encoded representation” [Mental Process] – determining multimodal masking loss measuring the similarity simply require to use the fused encoded representations of un-masked and masked training examples to compare the similarity between them, which involves observations, evaluations, judgments, and opinions that is capable of being performed in the human mind with the assistance of paper and pen “updating, by the one or more processors, one or more weights of the multimodal model in accordance with both the plurality of modality-specific masking losses and the multimodal masking loss” [Mental Process] – updating weights in accordance with two types of losses requires to compare them by check how similar they are then update the related numbers, which involves observations, evaluations, judgments, and opinions that is capable of being performed in the human mind with the assistance of paper and pen Step 2A Prong 2 – whether the claim recites additional elements that integrate the exception into a practical application of the exception? Regarding Claim 1, the claim recites additional elements of “by one or more processors” Processors are recited at a high level of generality and is merely adding words “apply it” to the judicial exception. (see MPEP 2106.05(d)) “multimodal model” Multimodal model is recited at a high level of generality and is merely adding words “apply it” to the judicial exception. (see MPEP 2106.05(f)) [Even when viewed in combination, the additional elements do no more than automate the mental processes that a person could perform, using computer components as a tool, thus the claim as a whole does not integrate into a practical application.] Step 2B – whether the claim as a whole amount to significantly more than the judicial exception? I.e. Are there any additional elements (features/limitations/step) recited in the claim beyond the abstract idea? The claim does not provide an inventive concept (significantly more than the abstract idea). The claim is ineligible. As explained above, the additional element [1] is considered merely computer components that are just to store and execute code-based instructions which are considered a mere instruction to apply an exception and amount to storing and receiving information in memory, which is well-understood, routine, conventional activity (See MPEP 2106.05(d), subsection II). This limitation remains a mere instruction to apply an exception. The additional element [2] is considered a mere instruction to apply an exception to the generic computer components or machine-learning components that simply run mathematical calculations and mental processes (see MPEP 2106.05(f)). This limitation remains a mere instruction to apply an exception even upon reconsideration. These limitations remain a mere instruction to apply an exception even upon reconsideration. Even when considered in combination, the additional element represents a mere instruction to apply an exception, which cannot provide an inventive concept. Regarding Claim 2 Step 1 – whether the claim falls within any statutory category. See MPEP 2016.03 Claim 2 is a dependent claim of 1, thus it falls within the same category of statutory subject matter. Step 2A Prong 1 – whether the claim recites a judicial exception. See MPEP 2106.04, subsection II. As Claim 2 does not have any abstract idea by itself, thus uses all the limitations of Claim 1. Step 2A Prong 2 – whether the claim recites additional elements that integrate the exception into a practical application of the exception? The claim 2 does not recite any additional elements other than abstract ideas, so it does not integrate into a practical application. Thus, this claim is directed to the abstract idea. Regarding Claim 3 Step 1 – whether the claim falls within any statutory category. See MPEP 2016.03 Claim 3 is a dependent claim of 1, thus it falls within the same category of statutory subject matter. Step 2A Prong 1 – whether the claim recites a judicial exception. See MPEP 2106.04, subsection II. As Claim 3 does not have any abstract idea by itself, thus uses all the limitations of Claim 1. Step 2A Prong 2 – whether the claim recites additional elements that integrate the exception into a practical application of the exception? The claim 3 does not recite any additional elements other than abstract ideas, so it does not integrate into a practical application. Thus, this claim is directed to the abstract idea. Regarding Claim 4 Step 1 – whether the claim falls within any statutory category. See MPEP 2016.03 Claim 4 is a dependent claim of 3, thus it falls within the same category of statutory subject matter. Step 2A Prong 1 – whether the claim recites a judicial exception. See MPEP 2106.04, subsection II. As Claim 4 does not have any abstract idea by itself, thus uses all the limitations of Claim 3. Step 2A Prong 2 – whether the claim recites additional elements that integrate the exception into a practical application of the exception? The claim 4 does not recite any additional elements other than abstract ideas, so it does not integrate into a practical application. Thus, this claim is directed to the abstract idea. Regarding Claim 5 Step 1 – whether the claim falls within any statutory category. See MPEP 2016.03 Claim 5 is a dependent claim of 4, thus it falls within the same category of statutory subject matter. Step 2A Prong 1 – whether the claim recites a judicial exception. See MPEP 2106.04, subsection II. Regarding dependent claim 5, following limitations recite a judicial exception: “wherein generating the modality-specific encoded representations comprises generating, by the one or more processors, a structured data encoded representation using a structured data encoder of the multimodal model” [Mental Process] – generating modality-specific encoded representations comprising a structured data simply transforming the structured data, which involves observations, evaluations, judgments, and opinions that is capable of being performed in the human mind with the assistance of paper and pen Step 2A Prong 2 – whether the claim recites additional elements that integrate the exception into a practical application of the exception? Regarding Claim 1, the claim recites additional elements of “by one or more processors” Processors are recited at a high level of generality and is merely adding words “apply it” to the judicial exception. (see MPEP 2106.05(d)) “multimodal model” Multimodal model is recited at a high level of generality and is merely adding words “apply it” to the judicial exception. (see MPEP 2106.05(f)) [Even when viewed in combination, the additional elements do no more than automate the mental processes that a person could perform, using computer components as a tool, thus the claim as a whole does not integrate into a practical application.] Step 2B – whether the claim as a whole amount to significantly more than the judicial exception? I.e. Are there any additional elements (features/limitations/step) recited in the claim beyond the abstract idea? The claim does not provide an inventive concept (significantly more than the abstract idea). The claim is ineligible. As explained above, the additional element [1] is considered merely computer components that are just to store and execute code-based instructions which are considered a mere instruction to apply an exception and amount to storing and receiving information in memory, which is well-understood, routine, conventional activity (See MPEP 2106.05(d), subsection II). This limitation remains a mere instruction to apply an exception. The additional element [2] is considered a mere instruction to apply an exception to the generic computer components or machine-learning components that simply run mathematical calculations and mental processes (see MPEP 2106.05(f)). This limitation remains a mere instruction to apply an exception even upon reconsideration. These limitations remain a mere instruction to apply an exception even upon reconsideration. Even when considered in combination, the additional element represents a mere instruction to apply an exception, which cannot provide an inventive concept. Regarding Claim 6 Step 1 – whether the claim falls within any statutory category. See MPEP 2016.03 Claim 6 is a dependent claim of 3, thus it falls within the same category of statutory subject matter. Step 2A Prong 1 – whether the claim recites a judicial exception. See MPEP 2106.04, subsection II. Regarding dependent claim 6, following limitations recite a judicial exception: “receiving, by the one or more processors, labeled training data comprising labeled training examples corresponding to a machine learning task” [Mental Process] – Receiving labeled training data is simply getting the labeled or truth data to be compared, which involves observations, evaluations, judgments, and opinions that is capable of being performed in the human mind with the assistance of paper and pen “performing, by the one or more processors, one or more fine-tuning iterations” [Mental Process] – performing one or more fine-tuning iterations simply means it will continue the procedures certain number of counts, which involves observations, evaluations, judgments, and opinions that is capable of being performed in the human mind with the assistance of paper and pen “processing the labeled training data through the multimodal model, determining a task-specific loss measuring performance of the multimodal model in performing the machine learning task, and updating weights of the multimodal model in accordance with the task-specific loss” [Mental Process] – processing the labeled training data, determining a task-specific loss and updating weights constitute data manipulation, mathematical evaluation, and numerical calculations, which involves observations, evaluations, judgments, and opinions that is capable of being performed in the human mind with the assistance of paper and pen Step 2A Prong 2 – whether the claim recites additional elements that integrate the exception into a practical application of the exception? Regarding Claim 1, the claim recites additional elements of “by one or more processors” Processors are recited at a high level of generality and is merely adding words “apply it” to the judicial exception. (see MPEP 2106.05(d)) “multimodal model” Multimodal model is recited at a high level of generality and is merely adding words “apply it” to the judicial exception. (see MPEP 2106.05(f)) [Even when viewed in combination, the additional elements do no more than automate the mental processes that a person could perform, using computer components as a tool, thus the claim as a whole does not integrate into a practical application.] Step 2B – whether the claim as a whole amount to significantly more than the judicial exception? I.e. Are there any additional elements (features/limitations/step) recited in the claim beyond the abstract idea? The claim does not provide an inventive concept (significantly more than the abstract idea). The claim is ineligible. As explained above, the additional element [1] is considered merely computer components that are just to store and execute code-based instructions which are considered a mere instruction to apply an exception and amount to storing and receiving information in memory, which is well-understood, routine, conventional activity (See MPEP 2106.05(d), subsection II). This limitation remains a mere instruction to apply an exception. The additional element [2] is considered a mere instruction to apply an exception to the generic computer components or machine-learning components that simply run mathematical calculations and mental processes (see MPEP 2106.05(f)). This limitation remains a mere instruction to apply an exception even upon reconsideration. These limitations remain a mere instruction to apply an exception even upon reconsideration. Even when considered in combination, the additional element represents a mere instruction to apply an exception, which cannot provide an inventive concept. Regarding Claim 7 Step 1 – whether the claim falls within any statutory category. See MPEP 2016.03 Claim 7 is a dependent claim of 6, thus it falls within the same category of statutory subject matter. Step 2A Prong 1 – whether the claim recites a judicial exception. See MPEP 2106.04, subsection II. Regarding dependent claim 7, following limitations recite a judicial exception: “after performing the one or more pretraining iterations and the one or more fine-tuning iterations, processing, by the one or more processors, instances of multimodal data input through the multimodal model to generate output in accordance with the machine learning task, wherein: one or more instances comprise data of each modality of the plurality of modalities present in the multimodal training data, and each instance other than the one or more instances comprises different respective combinations of at least partially missing data of at least one modality of the plurality of modalities present in the multimodal training data” [Mental Process] – processing instances labeled training data, which the data comprises data of each type and their combinations, and generating output in accordance with a task constitute mental analysis, which involves observations, evaluations, judgments, and opinions that is capable of being performed in the human mind with the assistance of paper and pen Step 2A Prong 2 – whether the claim recites additional elements that integrate the exception into a practical application of the exception? Regarding Claim 1, the claim recites additional elements of “by one or more processors” Processors are recited at a high level of generality and is merely adding words “apply it” to the judicial exception. (see MPEP 2106.05(d)) “multimodal model” Multimodal model is recited at a high level of generality and is merely adding words “apply it” to the judicial exception. (see MPEP 2106.05(f)) [Even when viewed in combination, the additional elements do no more than automate the mental processes that a person could perform, using computer components as a tool, thus the claim as a whole does not integrate into a practical application.] Step 2B – whether the claim as a whole amount to significantly more than the judicial exception? I.e. Are there any additional elements (features/limitations/step) recited in the claim beyond the abstract idea? The claim does not provide an inventive concept (significantly more than the abstract idea). The claim is ineligible. As explained above, the additional element [1] is considered merely computer components that are just to store and execute code-based instructions which are considered a mere instruction to apply an exception and amount to storing and receiving information in memory, which is well-understood, routine, conventional activity (See MPEP 2106.05(d), subsection II). This limitation remains a mere instruction to apply an exception. The additional element [2] is considered a mere instruction to apply an exception to the generic computer components or machine-learning components that simply run mathematical calculations and mental processes (see MPEP 2106.05(f)). This limitation remains a mere instruction to apply an exception even upon reconsideration. These limitations remain a mere instruction to apply an exception even upon reconsideration. Even when considered in combination, the additional element represents a mere instruction to apply an exception, which cannot provide an inventive concept. Regarding Claim 8 Step 1 – whether the claim falls within any statutory category. See MPEP 2016.03 Claim 8 is a dependent claim of 1, thus it falls within the same category of statutory subject matter. Step 2A Prong 1 – whether the claim recites a judicial exception. See MPEP 2106.04, subsection II. Regarding dependent claim 8, following limitations recite a judicial exception: “generating the first fused encoded representation and the second fused encoded representation comprises processing, by the one or more processors, the modality-specific encoded representations through a plurality of cross-attention layers of an attention-based transformer” [Mental Process] – generating the first and second fused encoded representations of the specific type of inputs simply require going through data transformation, which involves observations, evaluations, judgments, and opinions that is capable of being performed in the human mind with the assistance of paper and pen Step 2A Prong 2 – whether the claim recites additional elements that integrate the exception into a practical application of the exception? Regarding Claim 1, the claim recites additional elements of “by one or more processors” Processors are recited at a high level of generality and is merely adding words “apply it” to the judicial exception. (see MPEP 2106.05(d)) “a plurality of cross-attention layers of an attention-based transformer” The layers are recited at a high level of generality and is merely adding words “apply it” to the judicial exception. (see MPEP 2106.05(f)) [Even when viewed in combination, the additional elements do no more than automate the mental processes that a person could perform, using computer components as a tool, thus the claim as a whole does not integrate into a practical application.] Step 2B – whether the claim as a whole amount to significantly more than the judicial exception? I.e. Are there any additional elements (features/limitations/step) recited in the claim beyond the abstract idea? The claim does not provide an inventive concept (significantly more than the abstract idea). The claim is ineligible. As explained above, the additional element [1] is considered merely computer components that are just to store and execute code-based instructions which are considered a mere instruction to apply an exception and amount to storing and receiving information in memory, which is well-understood, routine, conventional activity (See MPEP 2106.05(d), subsection II). This limitation remains a mere instruction to apply an exception. The additional element [2] is considered a mere instruction to apply an exception to the generic computer components or machine-learning components that simply run mathematical calculations and mental processes (see MPEP 2106.05(f)). This limitation remains a mere instruction to apply an exception even upon reconsideration. These limitations remain a mere instruction to apply an exception even upon reconsideration. Even when considered in combination, the additional element represents a mere instruction to apply an exception, which cannot provide an inventive concept. Regarding Claim 9 Step 1 – whether the claim falls within any statutory category. See MPEP 2016.03 Claim 9 is a dependent claim of 1, thus it falls within the same category of statutory subject matter. Step 2A Prong 1 – whether the claim recites a judicial exception. See MPEP 2106.04, subsection II. Regarding dependent claim 9, following limitations recite a judicial exception: “a negative cosine similarity between a projection of the first fused encoded representation and the second fused representation, and a negative cosine similarity between a projection of the second fused representation and the first fused encoded representation” [Mathematical Calculations] – a negative cosine similarity constitutes an abstract idea of Mathematical Calculation which requires to calculate the similarity between two inputs. Step 2A Prong 2 – whether the claim recites additional elements that integrate the exception into a practical application of the exception? The claim 9 does not recite any additional elements other than abstract ideas, so it does not integrate into a practical application. Thus, this claim is directed to the abstract idea. Regarding Claim 10 Step 1 – whether the claim falls within any statutory category. See MPEP 2016.03 Claim 10 is a dependent claim of 9, thus it falls within the same category of statutory subject matter. Step 2A Prong 1 – whether the claim recites a judicial exception. See MPEP 2106.04, subsection II. Regarding dependent claim 10, following limitations recite a judicial exception: “calculating a total loss as the sum of the plurality of modality-specific masking losses and the multimodal masking loss, each of the masking losses weighted by a respective hyperparameter value” [Mathematical Calculations] – calculating a total loss as a weighted sum recites a mathematical formula and basic arithmetic operation, which recites an abstract idea “updating the one or more weights of the multimodal model using backpropagation with gradient descent and the total loss” [Mental Process] - updating weights in accordance with two types of losses requires to compare them by check how similar they are then update the related numbers, which involves observations, evaluations, judgments, and opinions that is capable of being performed in the human mind with the assistance of paper and pen Step 2A Prong 2 – whether the claim recites additional elements that integrate the exception into a practical application of the exception? Regarding Claim 10, the claim recites additional elements of “using backpropagation with gradient descent and the total loss” The layers are recited at a high level of generality and is merely adding words “apply it” to the judicial exception. (see MPEP 2106.05(f)) [Even when viewed in combination, the additional elements do no more than automate the mental processes that a person could perform, using computer components as a tool, thus the claim as a whole does not integrate into a practical application.] Step 2B – whether the claim as a whole amount to significantly more than the judicial exception? I.e. Are there any additional elements (features/limitations/step) recited in the claim beyond the abstract idea? The claim does not provide an inventive concept (significantly more than the abstract idea). The claim is ineligible. As explained above, the additional element [1] is considered a mere instruction to apply an exception to the generic computer components or machine-learning components that simply run mathematical calculations and mental processes (see MPEP 2106.05(f)). This limitation remains a mere instruction to apply an exception even upon reconsideration. This limitation remains a mere instruction to apply an exception even upon reconsideration. Even when considered in combination, the additional element represents a mere instruction to apply an exception, which cannot provide an inventive concept. Regarding Claim 11 Step 1 – whether the claim falls within any statutory category. See MPEP 2016.03 Claim 11 is a system claim thus it falls into one of the four categories of statutory subject matter. Step 2A Prong 1 – whether the claim recites a judicial exception. See MPEP 2106.04, subsection II. Regarding independent claim 11, following limitations recite a judicial exception: “receive multimodal input data” [Mental Process] – Receiving multimodal input data can be simply done with human minds or through receiving physical data, which involves observations, evaluations, judgments, and opinions that is capable of being performed in the human mind with the assistance of paper and pen “process the multimodal input data through a multimodal model, pretrained in accordance with one or more total losses comprising: a plurality of modality-specific masking losses from generating modality-specific encoded representations of masked and un-masked training examples, and one or more multimodal masking losses from generating fused encoded representations of the modality-specific encoded representations” [Mental Process] – processing the input data in accordance with total losses comprising both modality-specific losses and multimodal masking losses constitute mental analysis, which involves observations, evaluations, judgments, and opinions that is capable of being performed in the human mind with the assistance of paper and pen “generate, in response to receiving the multimodal input data, model output from the multimodal model” [Mental Process] – generating output based on the inputs simply constitute mental analysis of in-out data generation, which involves observations, evaluations, judgments, and opinions that is capable of being performed in the human mind with the assistance of paper and pen Step 2A Prong 2 – whether the claim recites additional elements that integrate the exception into a practical application of the exception? Regarding Claim 11, the claim recites additional elements of “one or more processors” Processors are recited at a high level of generality and is merely adding words “apply it” to the judicial exception. (see MPEP 2106.05(d)) “multimodal model” Multimodal model is recited at a high level of generality and is merely adding words “apply it” to the judicial exception. (see MPEP 2106.05(f)) [Even when viewed in combination, the additional elements do no more than automate the mental processes that a person could perform, using computer components as a tool, thus the claim as a whole does not integrate into a practical application.] Step 2B – whether the claim as a whole amount to significantly more than the judicial exception? I.e. Are there any additional elements (features/limitations/step) recited in the claim beyond the abstract idea? The claim does not provide an inventive concept (significantly more than the abstract idea). The claim is ineligible. As explained above, the additional element [1] is considered merely computer components that are just to store and execute code-based instructions which are considered a mere instruction to apply an exception and amount to storing and receiving information in memory, which is well-understood, routine, conventional activity (See MPEP 2106.05(d), subsection II). This limitation remains a mere instruction to apply an exception. The additional element [2] is considered a mere instruction to apply an exception to the generic computer components or machine-learning components that simply run mathematical calculations and mental processes (see MPEP 2106.05(f)). This limitation remains a mere instruction to apply an exception even upon reconsideration. These limitations remain a mere instruction to apply an exception even upon reconsideration. Even when considered in combination, the additional element represents a mere instruction to apply an exception, which cannot provide an inventive concept. Regarding Claim 12 Step 1 – whether the claim falls within any statutory category. See MPEP 2016.03 Claim 12 is a system claim thus it falls into one of the four categories of statutory subject matter. Step 2A Prong 1 – whether the claim recites a judicial exception. See MPEP 2106.04, subsection II. Regarding independent claim 12, following limitations recite a judicial exception: “a modality-specific masking loss is a measurement of the similarity between modality-specific encoded representations of the un-masked training examples and modality-specific encoded representations of the masked training examples” [Mathematical Concept/Calculation] – loss being a measurement of the similarity between two data representations requires to be computed using mathematical formulas and calculations “the one or more multimodal masking losses are measurements of similarities between fused encoded representations generated from modality-specific encoded representations for the masked training examples and fused encoded representations generated from modality-specific encoded representations for the un-masked training examples” [Mathematical Concept/Calculation] – loss being a measurement of the similarity between two data representations requires to be computed using mathematical formulas and calculations Step 2A Prong 2 – whether the claim recites additional elements that integrate the exception into a practical application of the exception? The claim 12 does not recite any additional elements other than abstract ideas, so it does not integrate into a practical application. Thus, this claim is directed to the abstract idea. Regarding Claims 13-17 Claims 13-17 have similar limitations of Claims 1, 10, 3, 7, 6, respectively. For the reasons described above with respect to Claims 1, 10, 3, 7, 6, these judicial exceptions are not meaningfully integrated into a practical application, or significantly more than the abstract ideas. The claims do not provide anything more than the abstract ideas of mental processes and mathematical calculations that are practically capable of being performed with the assistance of pen and paper. Therefore, Claims 13-17 also recite abstract ideas that do not integrate into a practical application or amount to significantly more than judicial exception, and thus are rejected under U.S.C. 101. Regarding Claims 18-20 Claims 18-20 have similar limitations of Claims 11, 12, 13, respectively. For the reasons described above with respect to Claims 11, 12, 13, these judicial exceptions are not meaningfully integrated into a practical application, or significantly more than the abstract ideas. The claims do not provide anything more than the abstract ideas of mental processes and mathematical calculations that are practically capable of being performed with the assistance of pen and paper. Therefore, Claims 18-20 also recite abstract ideas that do not integrate into a practical application or amount to significantly more than judicial exception, and thus are rejected under U.S.C. 101. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1, 2 8, 9 are rejected under 35 U.S.C. 103 as being unpatentable over Liu et al. (Liu), Non-Patent Literature listed in IDS filed on 04/18/2024, “OPT: Omni-Perception Pre-Trainer for Cross-Modal Understanding and Generation”, Published in 07/06/2021, arXiv, 10 Pages in view of Chen et al. (Chen), Non-Patent Literature listed in IDS filed on 04/18/2024, “Exploring Simple Siamese Representation Learning”, Published in 09/20/2020, arXiv, 10 Pages. As to independent Claim 1, Liu teaches a method of pretraining a multimodal model, the method comprising: receiving, by one or more processors, multimodal training data comprising un-masked training examples and masked training examples ( Liu, Pg6, Section 5.1, Lines1-4, "We mainly use Open Images [18] dataset with localized narratives and synchronized speech that are provided by [29] as our pre-training dataset. Only text-image-audio triplets are used for pre-training" Pg2, Left Column, Paragraph1, Lines7-9, "Token-level modeling predicts the semantics of masked tokens given the unmasked inputs" Pg5, Section 4.2, Paragraph2, Lines3-8, "Modality-level masking is in parallel with token-level masking mechanism. It masks out one or two modalities from the input. specifically, each modality is independently masked out with a probability of 0.3, and the case when all modalities are masked is skipped" Pg6, Section5.2, Lines13-15, "Models are trained on 4 Tesla V100 GPUs with a total batch size of 10,240 for 100,000 iterations, and early stop is performed", wherein Liu teaches that the models are trained on GPUs (hereinafter any processors will be referring to this) which inherently comprises a system to train the model that consists processors. Liu also explicitly teaches receiving multimodal training data consisting of image, text and audio modalities. Liu also teaches generating and receiving both masked and unmasked training examples through two masking schemes of token-level masking, which masks a portion of tokens within an input sequence while maintaining the remaining tokens as un-masked, and modality-level masking, which masks entire modalities independently with a probability of 0.3 that the case where all modalities are masked is skipped that the model receives inputs containing a combination of masked and remaining un-masked modalities, rendering it functionally equivalent to the claimed invention); and performing, by the one or more processors, one or more pretraining iterations comprising ( Liu, Pg6, Section 5.1, Lines1-4, "We mainly use Open Images [18] dataset with localized narratives and synchronized speech that are provided by [29] as our pre-training dataset. Only text-image-audio triplets are used for pre-training"): generating modality-specific encoded representations of the un-masked training examples and masked training examples ( Liu, Pg3, Abstract, Lines4-6, "OPT is constructed in an encoder-decoder framework, including three single-modal encoders to generate token-based embeddings for each modality..." Pg3, Section3.1, Paragraph1, "Text Encoder. Following BERT [7], we first tokenize all words by WordPieces [15] to obtain the token sequence... The final embedding for each token is obtained via summing up its token embedding and position embedding..." Pg4, Section4.1, Paragraph2, Lines4-6, "Following BERT, we randomly mask 15% words with the special token in [MASK] the sentence" Pg3. Section3.1, Paragraph2, "Vision Encoder. We use Faster R-CNN [34] pre-trained on Visual Genome dataset [16] to extract the visual representations (pooled ROI features) for each image region...The final visual embedding for each region is obtained by summing up the two FC outputs and then passing through an LN layer" Pg3. Section3.1, Paragraph3, "Audio Encoder. We use pre-trained wav2vec 2.0 [3] to obtain the audio tokens and extract the features for each token. The final audio embedding is obtained by passing the audio features through an LN layer” Pg3, Section3.2, Lines1-9, "After processing the input text, image and audio using single-modal encoders, we obtain the initial text embedding T, image embedding V and audio embedding A. To model the cross-modal interations among textual words, visual regions and audio tokens, we introduce a Transformer based cross-modal encoder. Specifically, we first combine T, V and A to get a long sequence of token features, and then feed the sequence of features into the cross-modal encoder to learn contextualized representations M, M = CrossEncoder([T;V;A])", wherein Liu explicitly discloses single-modal encoders (the corresponding modality-specific encoders) for input modalities. Also, Liu discloses both token-level masking and modality-level masking, where the single-modal encoders generate encoded representations for both un-masked inputs (complete sequences T,V,A or un-masked tokens T_\m, V_\m, A_\m) and masked inputs (inputs containing [MASK] tokens/masked features or modality level masked inputs), rendering it functionally equivalent to the claimed invention), determining a plurality of modality-specific masking losses measuring the similarity between the modality-specific encoded representations of the un-masked training examples and the masked training examples ( Liu, Pg4, Section4.1, Pagagraph2, "Masked Language Modeling (MLM)… The goal is to predict these masked words based on the observation of their surrounding words T_\m, all image regions V and all audio tokens A, by minimizing the negative log-likelihood: PNG media_image1.png 22 307 media_image1.png Greyscale " Pg4, Section4.1, Paragraph3, Lines1-4 and Equation3, "Masked Vision Modeling (MVM)... Similar to MLM, we also propose Masked Vision Modeling (MVM) to predict the correct image regions given contextual regions and other input modalities", PNG media_image2.png 26 291 media_image2.png Greyscale Pg4, Section4.1 Paragraph6, Lines1-6 and Equation6, "Masked Audio Modeling (MAM). For MAM, we mask audio features with a probability of 15%. Then the model is trained to reconstruct masked audio A_m, given the remaining audio tokens A_\m and all information from other modalities (i.e. text and image). Here, we propose two objectives for MAM, which share the same objective base", PNG media_image3.png 26 287 media_image3.png Greyscale Pg5, Equation8, PNG media_image4.png 124 324 media_image4.png Greyscale Pg4, Right Column, Paragraph1, Lines1-7, “The first objective is Masked Visual Feature Regression (MVFR), which regresses the cross-modal encoder output of each masked region to its input ROI visual features Vm m. We use an additional FC layer to transform the out V put of cross-modal encoder to the same dimensional space as the input visual feature. Then we apply L2 regression between the two”, wherein Liu explicitly discloses computing distinct losses for each modality as shown in the three equations using masked and unmasked representations of (T_\m, V_\m, A_\m) and (T_m, V_m, A_m) where at least audio and vision modalities use idea of similarity to compute the loss, rendering it functionally equivalent to the claimed invention.) generating a fused encoded representation ( Liu, Pg3, Section3.2, Lines1-9, "After processing the input text, image and audio using single-modal encoders, we obtain the initial text embedding T, image embedding V and audio embedding A. To model the cross-modal interations among textual words, visual regions and audio tokens, we introduce a Transformer based cross-modal encoder. Specifically, we first combine T, V and A to get a long sequence of token features, and then feed the sequence of features into the cross-modal encoder to learn contextualized representations M, M = CrossEncoder([T;V;A])" wherein Liu explicitly teaches about generating fused representations M, rendering it functionally equivalent to the claimed invention), updating, by the one or more processors, one or more weights of the multimodal model ( Liu, Pg4, Section4, "We propose three levels of pre-training tasks: (1) token-level modeling, including masked language modeling (MLM), masked vision modeling (MVM), and masked audio modeling (MAM); (2) modality-level modeling, including denoising text reconstruction (DTR) and denoising image reconstruction (DIR); and (3) sample-level modeling" Pg6, Section5.2, Lines13-17, "Models are trained on 4 Tesla V100 GPUs with a total batch size of 10,240 for 100,000 iterations, and early stop is performed. We adopt the Adam optimizer with an initial learning rate of 5e-5", wherein Liu explicitly updates these network weights end-to-end over 100,000 iterations using the backpropagation of the Adam optimizer, rendering it functionally equivalent to the claimed invention.) However, Liu does not explicitly teach (highlighted sections) generating a first fused encoded representation of the un-masked training example encoded representations and a second fused encoded representation of the masked training example encoded representations determining a multimodal masking loss measuring the similarity between the first fused encoded representation and the second fused encoded representation updating, by the one or more processors, one or more weights of the multimodal model in accordance with both the plurality of modality-specific masking losses and the multimodal masking loss As mentioned above, Liu teaches about generating fused encoding representations but is silent about “a first fused encoded representation of the un-masked training example encoded representations and a second fused encoded representation of the masked training example encoded representations.” From the same field of endeavor, Chen teaches this ( Chen, Pg2, Section3, Lines1-9, PNG media_image5.png 278 464 media_image5.png Greyscale , Pg3, Left Column, Line1 and Equation2, PNG media_image6.png 86 418 media_image6.png Greyscale Wherein Chen discloses using two augmented inputs x1 and x2 into the encoder f to generate z1,z2 (the corresponding first and second fused encoding representations) and the output vectors p1,p2. If this pipeline is combined with Liu’s fused encoding representations, it would create a first and a second fused encoded representations of masked and unmasked training examples, rendering it functionally equivalent to the claimed invention.) Chen further teaches determining a multimodal masking loss measuring the similarity between the first fused encoded representation and the second fused encoded representation ( Chen, Pg2, Section3, Lines7-9 and Equation1, PNG media_image7.png 138 462 media_image7.png Greyscale Pg3, Left Column, Line1 and Equation2, PNG media_image6.png 86 418 media_image6.png Greyscale wherein Chen discloses f(x1) and f(x2) correspond to the first and second fused representations generated by the encoder f as mentioned above, while p1 and p2 represent their projected vectors transformed by the prediction head h. Consequently, the term D(p1, z2) calculates the negative cosine similarity between the projection of the first representation,p1, and the second representation,z2, whereas D(p2,z1) calculates the negative cosine similarity between the projection of the second representation,p2, and the first representation,z1. Combining these terms in the symmetrized formulation in Equation9 to compute the multimodal masking losses of the first and second fused representations mentioned above renders it functionally equivalent to the claimed invention.) As mentioned above, Liu teaches about updating the weights of the multimodal model, but is silent about updating the weights in accordance with both the plurality of modality-specific masking losses and the multimodal masking loss. Chen teaches this ( Pg3, Left Column, Line1 and Equation2, PNG media_image6.png 86 418 media_image6.png Greyscale , Wherein Chen, as mentioned above, computes the multimodal masking loss which can be directly incorporated into Liu’s backpropagation of Adam optimizer that updates the weights, which relies on using the loss function, rendering it functionally equivalent to the claimed invention when combined.) Both Liu and Chen are analogous to the claimed invention as they are from the same field of endeavor of self-supervised representation learning utilizing Siamese neural network architectures to minimize similarity loss between generated representations. Therefore, it would have been obvious to one of ordinary skills in the art, before the effective filing date, to combine the multimodal pre-training scheme of Liu with dual-view similarity loss computation of Chen. The motivation is as recited by Chen (Chen, Pg1, Abstract, Lines3-12, " These models maximize the similarity between two augmentations of one image, subject to certain conditions for avoiding collapsing solutions. In this paper, we report surprising empirical results that simple Siamese networks can learn meaningful representations even using none of the following: (i) negative sample pairs, (ii) large batches, (iii) momentum encoders. Our experiments show that collapsing solutions do exist for the loss and structure, but a stop-gradient operation plays an essential role in preventing collapsing" Pg2, Section3, Lines1-7, “Our architecture (Figure 1) takes as input two randomly augmented views x1 and x2 from an image x. The two views are processed by an encoder network f consisting of a backbone (e.g., ResNet [19]) and a projection MLP head [8]. The encoder f shares weights between the two views. A prediction MLP head [15], denoted as h, transforms the output of one view and matches it to the other view”) such that incorporating a similarity loss between different (masked and unmasked) representations enables the multimodal model to directly maximize cross-modal feature alignment and improve representation learning efficiency without requiring complex negative sampling or momentum encoders, while utilizing a stop-gradient operation prevents output collapsing during backpropagation, yielding predictable training stability. As to dependent Claim 2, The combination of Liu and Chen teaches, as mentioned above, all the limitations of Claim 1. It teaches the overall architecture of receiving multimodal masked and unmasked training data to generate modality specific data representation of the data according to its type, which then are fused together to generate fused representations of masked and unmasked data to train the multimodal model by backpropagation. Liu further teaches the method of claim 1, wherein the masked training examples comprise data that is least partially masked or removed from the un-masked training examples ( Liu, Pg3, Figure1, Lines4-6, "We introduce two masking mechanisms: (1) token-level masking, in order for token-level modeling; and (2) modality-level masking, in order for modality-level modeling and enabling arbitrary number of input modalities" Pg4, Section4.1, Paragraph2, Lines4-6, "Following BERT, we randomly mask 15% words with the special token in [MASK] the sentence" Pg4, Right Column, Lines1-2, "We sample image regions and then mask their visual features with a probability of 15%" Pg4, Lines Right Column, Masked Audio Modeling(MAM) Section, Lines1-2, "For MAM, we mask audio features with a probability of 15%" Pg5, Section 4.2, Paragraph2, Lines3-8, "Modality-level masking is in parallel with token-level masking mechanism. It masks out one or two modalities from the input. specifically, each modality is independently masked out with a probability of 0.3, and the case when all modalities are masked is skipped", wherein Liu explicitly discloses the pre-training data preparation pipeline which implements a dual masking scheme comprising token-level masking and modality-level masking. Specifically, Liu’s token-level masking (masking 15% of text words, visual features, or audio tokens) teaches data that is ‘partially masked’ and Liu’s modality-level masking (randomly drops entire input streams from one or two modalities) teaches data that is ‘removed from the un-masked training examples’, rendering it functionally equivalent to the claimed invention.) As to dependent Claim 8, The combination of Liu and Chen teaches, as mentioned above, all the limitations of Claim 1. It teaches the overall architecture of receiving multimodal masked and unmasked training data to generate modality specific data representation of the data according to its type, which then are fused together to generate fused representations of masked and unmasked data to train the multimodal model by backpropagation. Liu further teaches the method of claim 1, wherein generating the first fused encoded representation and the second fused encoded representation comprises processing, by the one or more processors, the modality-specific encoded representations through a plurality of cross-attention layers of an attention-based transformer ( Liu, Pg2, Left Column, Lines2-5, "We then learn joint contextualized representations for these three modalities through a Transformer based cross-modal encoder, which provides interaction among modalities at different representation depths" Pg3, Section3.2, Lines3-9, "To model the cross-modal interations among textual words, visual regions and audio tokens, we introduce a Transformer based cross-modal encoder. Specifically, we first combine T, V and A to get a long sequence of token features, and then feed the sequence of features into the cross-modal encoder to learn contextualized representation M" Pg6, Section5.2, Lines7-10, "For the cross-modal encoder, we use the BERT-base model [7] with 12 layers of Transformer blocks. Each block has 12 attention heads and the hidden size is 768", wherein Liu teaches generating fused (joint) encoded representations by processing modality-specific encoded representations through an attention-based transformer by disclosing that initial single-modal embeddings (T,V,A) are combined into a sequence and fed into a Transformer-based cross-modal encoder to compute contextualized cross-modal representations M. Although Liu discloses performing self-attention over the combined sequence [T; V; A], this provides attention/interaction across different modalities rendering it functionally equivalent to processing through a plurality of cross-attention layers. Furthermore, it would have been obvious to one of ordinary skills in the art, before the effective filing date, to modify Liu’s encoder to utilize a plurality of cross-attention layers as they are a simple substitution of know techniques within the art of multimodal Transformers, since both self-attention over concatenated inputs and dedicated cross-attention layers are well-known architectural alternatives used to achieve cross-modal interactions. Executing such a substitution represents a predictable design choice that achieves the same expected function of extracting fused multimodal representations, without producing any unexpected or non-obvious results.) As to dependent Claim 9, The combination of Liu and Chen teaches, as mentioned above, all the limitations of Claim 1. It teaches the overall architecture of receiving multimodal masked and unmasked training data to generate modality specific data representation of the data according to its type, which then are fused together to generate fused representations of masked and unmasked data to train the multimodal model by backpropagation. However, Liu does not explicitly teach the method of claim 1, wherein the multimodal masking loss comprises: a negative cosine similarity between a projection of the first fused encoded representation and the second fused representation, and a negative cosine similarity between a projection of the second fused representation and the first fused encoded representation. From the same field of endeavor, Chen teaches them ( Chen, Pg2, Section3, Lines7-9 and Equation1, PNG media_image7.png 138 462 media_image7.png Greyscale Pg3, Left Column, Line1 and Equation2, PNG media_image6.png 86 418 media_image6.png Greyscale , wherein Chen discloses f(x1) and f(x2) correspond to the first and second fused representations generated by the encoder f, while p1 and p2 represent their projected vectors transformed by the prediction head h. Consequently, the term D(p1, z2) calculates the negative cosine similarity between the projection of the first representation,p1, and the second representation,z2, whereas D(p2,z1) calculates the negative cosine similarity between the projection of the second representation,p2, and the first representation,z1. Combining these terms in the symmetrized formulation in Equation2 to compute the multimodal masking losses of the first and second fused representations mentioned above renders it functionally equivalent to the claimed invention.) Both Liu and Chen are analogous to the claimed invention as they are from the same field of endeavor of self-supervised representation learning utilizing Siamese neural network architectures to minimize similarity loss between generated representations. Therefore, it would have been obvious to one of ordinary skills in the art, before the effective filing date, to combine the multimodal pre-training scheme of Liu with dual-view similarity loss computation of Chen. The motivation is as recited by Chen (Chen, Pg1, Abstract, Lines3-12, " These models maximize the similarity between two augmentations of one image, subject to certain conditions for avoiding collapsing solutions. In this paper, we report surprising empirical results that simple Siamese networks can learn meaningful representations even using none of the following: (i) negative sample pairs, (ii) large batches, (iii) momentum encoders. Our experiments show that collapsing solutions do exist for the loss and structure, but a stop-gradient operation plays an essential role in preventing collapsing" Pg2, Section3, Lines1-7, “Our architecture (Figure 1) takes as input two randomly augmented views x1 and x2 from an image x. The two views are processed by an encoder network f consisting of a backbone (e.g., ResNet [19]) and a projection MLP head [8]. The encoder f shares weights between the two views. A prediction MLP head [15], denoted as h, transforms the output of one view and matches it to the other view”) such that incorporating a similarity loss between different (masked and unmasked) representations enables the multimodal model to directly maximize cross-modal feature alignment and improve representation learning efficiency without requiring complex negative sampling or momentum encoders, while utilizing a stop-gradient operation prevents output collapsing during backpropagation, yielding predictable training stability. Claims 3-7 are rejected under 35 U.S.C. 103 as being unpatentable over Liu and Chen as mentioned in Claim 1 in further view of Arik et al. (Arik), Non-Patent Literature listed in IDS filed on 04/18/2024, “TabNet: Attentive Interpretable Tabular Learning”, Published in May 2021, AAAI, 9 Pages. As to dependent Claim 3, The combination of Liu and Chen teaches, as mentioned above, all the limitations of Claim 1. It teaches the overall architecture of receiving multimodal masked and unmasked training data to generate modality specific data representation of the data according to its type, which then are fused together to generate fused representations of masked and unmasked data to train the multimodal model by backpropagation. Liu, as mentioned above, uses multimodal training data that comprises types of text, visual, and audio, but Liu does not teach The method of claim 1, wherein the multimodal training data comprises data of a plurality of modalities, comprising structured data. From the same field of endeavor, Arik teaches this limitation ( Arik, Pg6679, Right Column, Lines1-5, "In addition, unlike tree learning, DNNs enable gradient descent based end-to-end learning for tabular data which can have a multitude of benefits: (i) efficiently encoding multiple data types like images along with tabular data" Pg6679, Introduction, Lines1-3, "Deep neural networks (DNNs) have shown notable success with images (He et al. 2015), text (Lai et al. 2015) and audio (Amodei et al. 2015)" Pg6679, Introduction, Lines5-9, "One data type that has yet to see such success with a canonical architecture is tabular data. Despite being the most common data type in real-world AI (as it is comprised of any categorical and numerical features)", wherein Arik explicitly identify tabular data, composed of categorical numerical features, as structured data, while separately enumerating unstructured modalities including text, image, and audio, rendering it functionally equivalent to the claimed invention of using training data comprises structured data and unstructured data such as text, image or audio.) Liu, Chen and Arik are analogous to the claimed invention as they are from the same field of endeavor of self-supervised multimodal representation learning and attention-based deep neural network architectures for processing heterogeneous data. Therefore, it would have been obvious to one of ordinary skills in the art, before the effective filing date, to combine the multimodal pre-training scheme of Liu and dual-view similarity loss computation of Chen with the tabular data encoding architecture of Arik. The motivation is as recited by Arik (Arik, Pg6679, Right Column, Lines1-10, "In addition, unlike tree learning, DNNs enable gradient descent-based end-to-end learning for tabular data which can have a multitude of benefits: (i) efficiently encoding multiple data types like images along with tabular data; (ii) alleviating the need for feature engineering, which is currently a key aspect in tree-based tabular data learning methods; (iii) learning from streaming data and perhaps most importantly (iv) end-to-end models allow representation learning which enables many valuable application scenarios") such that the combination of the three would integrate Arik’s tabular data encoding into Liu’s multimodal framework while applying Chen’s dual-view similarity loss to align masked and unmasked representations, thereby achieving a unified, end-to-end deep learning model capable of jointly learning representations across both structured and unstructured modalities without relying on disjoint tree-based models or manual feature engineering. As to dependent Claim 4, The combination of Liu, Chen and Arik teaches, as mentioned above, all the limitations of Claim 3. It teaches that the multimodal training data comprises structured data or as well as the unstructured data like text, image, audio. However, both Liu and Chen do not teach structured data comprises tabular feature data, time-series data, or both the tabular feature data and the time-series data. From the same field of endeavor, Arik teaches this limitation ( Arik, Pg6679, Introduction, Lines5-9, "One data type that has yet to see such success with a canonical architecture is tabular data. Despite being the most common data type in real-world AI (as it is comprised of any categorical and numerical features)" Pg6681, Left Column, Paragraph3, Lines2-3, "We use the raw numerical features and consider mapping of categorical features with trainable embeddings" Pg6684, Right Column, Paragraph1, Lines1-4, "Rossmann Store Sales 6: The task is forecasting the store sales from static and time-varying features. We observe that TabNet outperforms commonly-used methods. The time features (e.g. day) obtain high importance", wherein Arik discloses TabNet which is a canonical deep learning architecture designed specifically for processing tabular data composed of categorical and numerical features. Furthermore, Arik discloses processing structured inputs containing temporal and sequential components, as demonstrated in its forecasting application using static and time-varying features (e.g., time features such as day), rendering it functionally equivalent to the claimed invention.) Liu, Chen and Arik are analogous to the claimed invention as they are from the same field of endeavor of self-supervised multimodal representation learning and attention-based deep neural network architectures for processing heterogeneous data. Therefore, it would have been obvious to one of ordinary skills in the art, before the effective filing date, to combine the multimodal pre-training scheme of Liu and dual-view similarity loss computation of Chen with the tabular data encoding architecture of Arik. The motivation is as recited by Arik (Arik, Pg6679, Right Column, Lines1-10, "n addition, unlike tree learning, DNNs enable gradient descent-based end-to-end learning for tabular data which can have a multitude of benefits: (i) efficiently encoding multiple data types like images along with tabular data; (ii) alleviating the need for feature engineering, which is currently a key aspect in tree-based tabular data learning methods; (iii) learning from streaming data and perhaps most importantly (iv) end-to-end models allow representation learning which enables many valuable application scenarios") such that the combination of the three would integrate Arik’s tabular data encoding into Liu’s multimodal framework while applying Chen’s dual-view similarity loss to align masked and unmasked representations, thereby achieving a unified, end-to-end deep learning model capable of jointly learning representations across both structured and unstructured modalities without relying on disjoint tree-based models or manual feature engineering. As to dependent Claim 5, The combination of Liu, Chen and Arik teaches, as mentioned above, all the limitations of Claim 4. It teaches that the structured multimodal training data that comprises tabular feature data, time-series data. However, Liu and Chen do not teach generating the modality-specific encoded representations comprises generating, by the one or more processors, a structured data encoded representation using a structured data encoder of the multimodal model. From the same field of endeavor, Arik further teaches this limitation ( Arik, Pg6681, Left Column, Paragraph3, Lines1-3, "Fig. 4 shows the TabNet architecture for encoding tabular data. We use the raw numerical features and consider mapping of categorical features with trainable embeddings" Pg6681, Left Column, Paragraph3, Lines7-8, "TabNet’s encoding is based on sequential multi-step processing with N_steps decision steps" Pg6682, Figure4, Line1, "Figure 4: (a) TabNet encoder, composed of a feature transformer, an attentive transformer and feature masking" Pg6680, Figure2, Lines2-3, "Unsupervised representation learning by masked self-supervised learning results in an improved encoder model for the supervised learning task", wherein Arik discloses a dedicated neural network encoder architecture, specifically designated as the "TabNet encoder", that ingests raw numerical and embedded categorical tabular features to produce output encoded representations via sequential multi-step decision processing. Because Arik provides a structured data encoder designed to process tabular structured inputs into feature-rich latent encoded representations with machine learning frameworks, rendering it functionally equivalent to the claimed invention.) Liu, Chen and Arik are analogous to the claimed invention as they are from the same field of endeavor of self-supervised multimodal representation learning and attention-based deep neural network architectures for processing heterogeneous data. Therefore, it would have been obvious to one of ordinary skills in the art, before the effective filing date, to combine the multimodal pre-training scheme of Liu and dual-view similarity loss computation of Chen with the tabular data encoding architecture of Arik. The motivation is as recited by Arik (Arik, Pg6679, Right Column, Lines1-10, "n addition, unlike tree learning, DNNs enable gradient descent-based end-to-end learning for tabular data which can have a multitude of benefits: (i) efficiently encoding multiple data types like images along with tabular data; (ii) alleviating the need for feature engineering, which is currently a key aspect in tree-based tabular data learning methods; (iii) learning from streaming data and perhaps most importantly (iv) end-to-end models allow representation learning which enables many valuable application scenarios") such that the combination of the three would integrate Arik’s tabular data encoding into Liu’s multimodal framework while applying Chen’s dual-view similarity loss to align masked and unmasked representations, thereby achieving a unified, end-to-end deep learning model capable of jointly learning representations across both structured and unstructured modalities without relying on disjoint tree-based models or manual feature engineering. As to dependent Claim 6, The combination of Liu, Chen and Arik teaches, as mentioned above, all the limitations of Claim 3. It teaches that the multimodal training data comprises structured data or as well as the unstructured data like text, image, audio. Liu further teaches the method of claim 3, further comprising: receiving, by the one or more processors, labeled training data comprising labeled training examples corresponding to a machine learning task ( Liu, Pg6, Table1, PNG media_image8.png 186 355 media_image8.png Greyscale Pg5, Section4.3 Lines16-17, "where gt(T,V,A) is the one-hot vector of ground-truth label" Pg2, Paragraph2, Lines1-4, "We extensively validate OPT on a number of downstream tasks, including cross-modal retrieval, multi-modal classification, visual question answering, cross-modal text generation" wherein Liu explicitly teaches receiving labeled datasets such as COCO Captions and VQA datasets which consist of paired multimodal training examples along with ground-truth one-hot vectors(gt) to train and evaluate the model across various downstream machine learning tasks, rendering it functionally equivalent to the claimed invention); and performing, by the one or more processors, one or more fine-tuning iterations of ( Liu, Pg1, Right Column, Paragraph1, Lines4-8, "OPT is pre-trained on large amounts of language-vision-audio triplets with a multi-task pretext learning scheme, and can effectively adapt to downstream understanding and generation tasks given single-, two-, or three-modal inputs" Pg6, Section5.2, Lines13-15, "Models are trained on 4 Tesla V100 GPUs with a total batch size of 10,240 for 100,000 iterations, and early stop is performed" Pg6, Right Column, Multi-Modal Classification Section, Lines2-5, "We add a linear layer after the average pooling output of cross-modal encoder for classification. We freeze our pre-trained model and only linear layer is learned" wherein Liu describes how the pretrained model adapts to downstream tasks through iterative training loops over up to 100,000 iterations, either updating model parameters or optimizing task-specific classification head layers, rendering it functionally equivalent to the claimed invention): processing the labeled training data through the multimodal model ( Pg3, Section3.2, Lines6-9, "Specifically, we first combine T, V and A to get a long sequence of token features, and then feed the sequence of features into the cross-modal encoder to learn contextualized representations M" Pg4, Lines4-6, "When performing on down-stream, the two decoders is used to generate results, e.g., image captioning, text-to-image generation" Pg6, Right Column, Multi-Modal Classification Section, Lines2-4, "We add a linear layer after the average pooling output of cross-modal encoder for classification", wherein Liu processes labeled multimodal inputs by combining token features across modalities (T,V,A) and passing them through the Transformer-based cross-modal encoder and task decoders/classifiers during downstream task execution, rendering it functionally equivalent to the claimed invention), determining a task-specific loss measuring performance of the multimodal model in performing the machine learning task ( Pg4, Right Column, Paragraph2, Lines9-10, "The final objective minimizes the cross-entropy (CE) loss: PNG media_image9.png 41 346 media_image9.png Greyscale " Pg5, Right Column, Section4.3, Lines14-15, "The loss function is the binary cross entropy (BCE) loss: PNG media_image10.png 41 314 media_image10.png Greyscale " Pg5, Right Column, Paragraph1, Lines3-4, "The loss function is, PNG media_image11.png 30 313 media_image11.png Greyscale " wherein Liu teaches computing distinct, task-specific objective functions such as Cross-Entropy loss for classification tasks, Binary Cross-Entropy for sample-level matching, and negative log-likelihood for sequence generation, by comparing predicted model outputs against ground-truth target labels(gt), rendering it functionally equivalent to the claimed invention), and updating weights of the multimodal model in accordance with the task-specific loss ( Liu, Pg6, Section5.2, Lines16-17, "We adopt the Adam optimizer with an initial learning rate of 5e-5" Pg6, Right Column, Multi-Modal Classification Section, Lines4-5, "We freeze our pre-trained model and only linear layer is learned" Pg4, Section4.1, Paragraph2, Line10, "where _theta_ is the trainable parameters", wherein Liu describes optimizing trainable parameters, theta, using backpropagation with the Adam optimizer, wherein weights of the model (or added task-specific linear layers) are updated to minimize the task loss, rendering it functionally equivalent to the claimed invention.) As to dependent Claim 7, The combination of Liu, Chen and Arik teaches, as mentioned above, all the limitations of Claim 6. It teaches about using the labeled training data that corresponds to the machine learning tasks to perform fine-tunings by processing the data, measures loss of the data, and updating the weights of the model. Liu further teaches the method of claim 6, further comprising: after performing the one or more pretraining iterations and the one or more fine-tuning iterations, processing, by the one or more processors, instances of multimodal data input through the multimodal model to generate output in accordance with the machine learning task ( Liu, Pg1, Right Column, Paragraph1, Lines4-8, "OPT is pre-trained on large amounts of language-vision-audio triplets with a multi-task pretext learning scheme, and can effectively adapt to downstream understanding and generation tasks given single-, two-, or three-modal inputs" Pg4, Lines4-6, "When performing on down-stream, the two decoders is used to generate results, e.g., image captioning, text-to-image generation" Pg2, Paragraph2, Lines1-4, "We extensively validate OPT on a number of downstream tasks, including cross-modal retrieval, multi-modal classification, visual question answering, cross-modal text generation", wherein Liu describes how the model, after undergoing multimodal pretraining and task adaption, processes input instances through its cross-modal encoder and decoders to produce outputs across various downstream tasks such as classification, cross-modal retrieval, text generation, and text-to-image synthesis, rendering it functionally equivalent to the claimed invention), wherein: one or more instances comprise data of each modality of the plurality of modalities present in the multimodal training data ( Liu, Pg1, Right Column, Paragraph1, Lines4-8, "OPT is pre-trained on large amounts of language-vision-audio triplets with a multi-task pretext learning scheme, and can effectively adapt to downstream understanding and generation tasks given single-, two-, or three-modal inputs" Pg6, Right Column, Paragraph1, Lines11-15, "When with multimodal features, the performance can be further improved. In particular, adding text feature brings the largest improvement. On the basis of image+text feature, we further add audio feature, and find the performance still increases" Pg6, Table2, PNG media_image12.png 254 370 media_image12.png Greyscale Pg8, Figure3, Line1, "Some results of text-to-image generation and text generation", "Both: In this image I can see the candle...", PNG media_image13.png 542 1136 media_image13.png Greyscale , wherein Liu explicitly teaches evaluating downstream test instances where all three pretraining modalities, text, vision, and audio, are concurrently provided as input to the model during inference, as demonstrated in its full three-modality classification experiments in Table2 and combined text-generation tasks) each instance other than the one or more instances comprises different respective combinations of at least partially missing data of at least one modality of the plurality of modalities present in the multimodal training data ( Liu, Pg2, Left Column, Third Bullet, “OPT can effectively adapt to and perform competitively on a series of cross-modal understanding and generation downstream tasks with parial or all modalities as inputs” Pg6, Table2, PNG media_image12.png 254 370 media_image12.png Greyscale Pg8, Figure3, PNG media_image13.png 542 1136 media_image13.png Greyscale Pg6, Right Column, Lines7-12, “When only using image features, our OPT outperforms ResNet-50 and ResNet-101 by a large margin. When only using text or audio feature, our model also obtains promising results, which indicates that the model has learnt the associations between different modalities” wherein Liu describes processing inference instances where on or more modalities present during pre-training are missing, resulting in various combinations of partial modality inputs by testing downstream performance across multiple partial-modality input configurations, including single-modality inputs (text-only, image-only, audio-only) and two-modality combinations (image+text, audio+text, audio+image), demonstrating the model yields accurate task outputs even when specific pre-training modalities are entirely omitted, rendering it functionally equivalent to the claimed invention.) Claims 10-14, 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over Liu and Chen as mentioned in Claim 1 in further view of Chae et al. (Chae), Non-Patent Literature, “End-to-end Multimodal Emotion and Gender Recognition with Dynamic Joint Loss Weights”, Published in September 2018, arXiv, 5 Pages. As to dependent Claim 10, The combination of Liu and Chen teaches, as mentioned above, all the limitations of Claim 9. It teaches that the multimodal masking loss uses negative cosine similarity between the two fused representations in both ways. Liu further teaches the method of claim 9, wherein updating the one or more weights comprises: calculating a total loss as the sum of the plurality of modality-specific masking losses and the multimodal masking loss, each of the masking losses weighted by a respective hyperparameter value ( Liu, Pg4, Section4.1, Pagagraph2, "Masked Language Modeling (MLM)… The goal is to predict these masked words based on the observation of their surrounding words T_\m, all image regions V and all audio tokens A, by minimizing the negative log-likelihood: PNG media_image1.png 22 307 media_image1.png Greyscale " Pg4, Section4.1, Paragraph3, Lines1-4 and Equation3, "Masked Vision Modeling (MVM)... Similar to MLM, we also propose Masked Vision Modeling (MVM) to predict the correct image regions given contextual regions and other input modalities", PNG media_image2.png 26 291 media_image2.png Greyscale Pg4, Section4.1 Paragraph6, Lines1-6 and Equation6, "Masked Audio Modeling (MAM). For MAM, we mask audio features with a probability of 15%. Then the model is trained to reconstruct masked audio A_m, given the remaining audio tokens A_\m and all information from other modalities (i.e. text and image). Here, we propose two objectives for MAM, which share the same objective base", PNG media_image3.png 26 287 media_image3.png Greyscale , wherein Liu, as mentioned in Claim1, explicitly discloses computing distinct losses for each modality as shown in the three equations using masked and unmasked representations of (T_\m, V_\m, A_\m) and (T_m, V_m, A_m), rendering it functionally equivalent to the claimed invention); and updating the one or more weights of the multimodal model using backpropagation with gradient descent and the loss ( Liu, Pg6, Section5.2, Lines16-17, "We adopt the Adam optimizer with an initial learning rate of 5e-5", wherein Liu teaches optimizing these combined multi-task objectives, mentioned above, using the Adam optimizer, which is a standard gradient descent algorithm that relies on backpropagation to calculate loss gradients from the total objective and update the network weights, rendering it functionally equivalent to the claimed invention.) However, Liu does not explicitly teach (Highlighted sections) calculating a total loss as the sum of the plurality of modality-specific masking losses and the multimodal masking loss, each of the masking losses weighted by a respective hyperparameter value updating the one or more weights of the multimodal model using backpropagation with gradient descent and the total loss. From the same field of endeavor, Chen teaches the multimodal masking loss ( Chen, Pg2, Section3, Lines7-9 and Equation1, PNG media_image7.png 138 462 media_image7.png Greyscale Pg3, Left Column, Line1 and Equation2, PNG media_image6.png 86 418 media_image6.png Greyscale , Pg4, Section4.2, Lines5-7, “Its gradient has the same direction as the gradient of D(z1, z2), with the magnitude scaled by ½” wherein Chen, as mentioned in Claim9, teaches about computing multimodal masking loss L in Equation2 using the negative cosine, rendering it functionally equivalent to the claimed invention.) Both Liu and Chen are analogous to the claimed invention as they are from the same field of self-supervised multimodal representation learning and deep neural network optimization. Therefore, it would have been obvious to one of ordinary skills in the art, before the effective filing date, to combine the multi-task pretraining framework of Liu comprising a plurality of modality-specific masking losses with the hyper-parameter weighted loss aggregation scheme of Chen wherein individual loss terms are weighted by respective scaling coefficients. The motivation is taught by Chen ( Chen, Pg2, Left Column, Lines3-6, “Siamese networks can naturally introduce inductive biases for modeling invariance, as by definition “invariance” means that two observations of the same concept should produce the same outputs”, Pg4, Section4.2, Lines5-7, “Its gradient has the same direction as the gradient of D(z1, z2), with the magnitude scaled by ½” ) such that in multi-task multimodal representation learning, individual loss functions operate on fundamentally different numerical scales and produce gradient updates of varying magnitudes. Integrating hyper-parameter weighting factors for each loss component provides a routine mechanism to rebalance the relative gradient contributions of the modality-specific masking losses against the multimodal masking loss during backpropagation, which would prevent any single dominant loss component from overwhelming the backpropagation process, thereby stabilizing gradient descent updates, help numerical convergence, and prevent representation collapse across all combined modalities. However, both Liu and Chen do not explicitly teach calculating a total loss as the sum of the plurality of modality-specific masking losses and the multimodal masking loss, each of the masking losses weighted by a respective hyperparameter value; From the same field of endeavor, Chae teaches this ( Chae, Pg3, Right Column, Lines1-3, “Then, the joint loss can be calculated as follows: LJoint = we x Le + wg x Lg (3) where the values for we and wg vary between 0 and 1 such that we + wg = 1”, wherein Chae explicitly discloses joint loss, Ljoint (the corresponding total loss) that is the sum of two different losses, Le and Lg, for multimodal model with the weights (the corresponding respective hyperparameter value), rendering it functionally equivalent to the claimed invention as Le and Lg are replaced with Liu and Chen’s modality-specific loss and multimodal loss.) Liu, Chen, and Chae are analogous to the claimed invention as they are from the same field of endeavor of multi-task deep neural network learning and multi-modal representation modeling. Therefore, it would have been obvious to one of ordinary skills in the art, before the effective filing date, to combine the multi-modal encoder-decoder pre-training framework of Liu, the predictor and stop-gradient similarity loss optimization of Chen, with the dynamic joint loss weighting mechanism updated via backpropagation of Chae. The motivation is taught by Chae (Chae, Pg1, Right Column, Paragraph2, Lines3-11, “In multi-task learning, the performance of a model is highly dependent on the joint loss weights of each task, so selecting proper weights is an important issue for multi-task learning approaches. One method to determine the joint loss weights consists in using prior knowledge of models [12]. However, these static weights can cause adverse effect because each of the loss values continuously changes while training the models”) such that the combination would dynamically adjust task weights during backpropagation to adaptively balance tasks as their loss values continuously evolve, thereby preventing performance degradation caused by static/empirical weight selection and optimizing the overall joint loss across all integrated tasks. As to independent Claim 11, Liu teaches a system, comprising: One or more processors configured to: ( Liu, Pg6, Section5.2, Lines13-15, "Models are trained on 4 Tesla V100 GPUs with a total batch size of 10,240 for 100,000 iterations, and early stop is performed" ) receive multimodal input data; process the multimodal input data through a multimodal model, pretrained ( Liu, Pg1, Right Column, Paragraph1, Lines4-8, "OPT is pre-trained on large amounts of language-vision-audio triplets with a multi-task pretext learning scheme, and can effectively adapt to downstream understanding and generation tasks given single-, two-, or three-modal inputs" Pg2, Left Column, Third Bullet, "OPT can effectively adapt to and perform competitively on a series of cross-modal understanding and generation downstream tasks with parial or all modalities as inputs" Pg6, Right Column, Paragraph1, Lines6-15, "When only using image features, our OPT outperforms ResNet-50 and ResNet-101 by a large margin. When only using text or audio feature, our model also obtains promising results, which indicates that the model has learnt the associations between different modalities. When with multimodal features, the performance can be further improved. In particular, adding text feature brings the largest improvement. On the basis of image+text feature, we further add audio feature, and find the performance still increases" Pg6, Table2, PNG media_image12.png 254 370 media_image12.png Greyscale , wherein Liu explicitly discloses receiving and processing inference instances with missing pre-trained modalities. By evaluating downstream performance across single-modality and two-modality configurations, Liu demonstrates that the underlying architecture processes a partial-modality inputs to produce accurate task outputs despite omitting specific pre-trained modalities, rendering it functionally equivalent to the claimed invention); a plurality of modality-specific masking losses from generating modality-specific encoded representations of masked and un-masked training examples ( Liu, Pg3, Abstract, Lines4-6, "OPT is constructed in an encoder-decoder framework, including three single-modal encoders to generate token-based embeddings for each modality..." Pg3, Section3.1, Paragraph1, "Text Encoder. Following BERT [7], we first tokenize all words by WordPieces [15] to obtain the token sequence... The final embedding for each token is obtained via summing up its token embedding and position embedding..." Pg4, Section4.1, Paragraph2, Lines4-6, "Following BERT, we randomly mask 15% words with the special token in [MASK] the sentence" Pg3. Section3.1, Paragraph2, "Vision Encoder. We use Faster R-CNN [34] pre-trained on Visual Genome dataset [16] to extract the visual representations (pooled ROI features) for each image region...The final visual embedding for each region is obtained by summing up the two FC outputs and then passing through an LN layer" Pg3. Section3.1, Paragraph3, "Audio Encoder. We use pre-trained wav2vec 2.0 [3] to obtain the audio tokens and extract the features for each token. The final audio embedding is obtained by passing the audio features through an LN layer” Pg3, Section3.2, Lines1-9, "After processing the input text, image and audio using single-modal encoders, we obtain the initial text embedding T, image embedding V and audio embedding A. To model the cross-modal interations among textual words, visual regions and audio tokens, we introduce a Transformer based cross-modal encoder. Specifically, we first combine T, V and A to get a long sequence of token features, and then feed the sequence of features into the cross-modal encoder to learn contextualized representations M, M = CrossEncoder([T;V;A])", Liu, Pg4, Section4.1, Pagagraph2, "Masked Language Modeling (MLM)… The goal is to predict these masked words based on the observation of their surrounding words T_\m, all image regions V and all audio tokens A, by minimizing the negative log-likelihood: PNG media_image1.png 22 307 media_image1.png Greyscale " Pg4, Section4.1, Paragraph3, Lines1-4 and Equation3, "Masked Vision Modeling (MVM)... Similar to MLM, we also propose Masked Vision Modeling (MVM) to predict the correct image regions given contextual regions and other input modalities", PNG media_image2.png 26 291 media_image2.png Greyscale Pg4, Section4.1 Paragraph6, Lines1-6 and Equation6, "Masked Audio Modeling (MAM). For MAM, we mask audio features with a probability of 15%. Then the model is trained to reconstruct masked audio A_m, given the remaining audio tokens A_\m and all information from other modalities (i.e. text and image). Here, we propose two objectives for MAM, which share the same objective base", PNG media_image3.png 26 287 media_image3.png Greyscale , wherein Liu explicitly discloses single-modal encoders (the corresponding modality-specific encoders) for input modalities. Also, Liu discloses both token-level masking and modality-level masking, where the single-modal encoders generate encoded representations for both un-masked inputs (complete sequences T,V,A or un-masked tokens T_\m, V_\m, A_\m) and masked inputs (inputs containing [MASK] tokens/masked features or modality level masked inputs). Liu also explicitly discloses computing distinct losses for each modality as shown in the three equations using masked and unmasked representations of (T_\m, V_\m, A_\m) and (T_m, V_m, A_m), rendering it functionally equivalent to the claimed invention.), and generating a fused encoded representation of the modality-specific encoded representations ( Liu, Pg3, Section3.2, Lines1-9, "After processing the input text, image and audio using single-modal encoders, we obtain the initial text embedding T, image embedding V and audio embedding A. To model the cross-modal interations among textual words, visual regions and audio tokens, we introduce a Transformer based cross-modal encoder. Specifically, we first combine T, V and A to get a long sequence of token features, and then feed the sequence of features into the cross-modal encoder to learn contextualized representations M, M = CrossEncoder([T;V;A])" wherein Liu explicitly teaches about generating fused representations M, rendering it functionally equivalent to the claimed invention); and generate, in response to receiving the multimodal input data, model output from the multimodal model ( Liu, Pg2, Left Column, Lines5-8, "Finally, two cross-modal decoders take the outputs of the cross-modal encoder to generate text and image respectively in autoregressive manners" Pg2, Paragraph2, Lines1-4, "We extensively validate OPT on a number of downstream tasks, including cross-modal retrieval, multi-modal classification, visual question answering, cross-modal text generation" wherein Liu teaches generating downstream model outputs (such as classification labels, retrieved items, or generated text/images) via decoders or task heads upon processing the received multimodal input data, rendering it functionally equivalent to the claimed invention.) However, Liu does not explicitly teach (Highlighted parts) one or more multimodal masking losses from generating fused encoded representations of the modality-specific encoded representations Liu teaches about generating a fused encoded representation, but not multiple of them and using these representations to generate one or more multimodal masking losses. From the same field of endeavor, Chen teaches this ( Chen, Pg2, Section3, Lines1-9, PNG media_image5.png 278 464 media_image5.png Greyscale Pg3, Left Column, Line1 and Equation2, PNG media_image6.png 86 418 media_image6.png Greyscale , wherein Chen discloses using two augmented inputs x1 and x2 into the encoder f to generate z1,z2 (the corresponding first and second fused encoding representations) and the output vectors p1,p2. If this pipeline is combined with Liu’s fused encoding representations, it would create multiple, first and second, fused encoded representations of the modality-specific encoded representations. As f(x1) and f(x2) correspond to the first and second fused representations generated by the encoder f as mentioned above, while p1 and p2 represent their projected vectors transformed by the prediction head h. Consequently, the term D(p1, z2) calculates the negative cosine similarity between the projection of the first representation,p1, and the second representation,z2, whereas D(p2,z1) calculates the negative cosine similarity between the projection of the second representation,p2, and the first representation,z1. Combining these terms in the symmetrized formulation in Equation2 to compute the multimodal masking losses of the first and second fused representations mentioned above renders it functionally equivalent to the claimed invention.) Both Liu and Chen are analogous to the claimed invention as they are from the same field of self-supervised multimodal representation learning and deep neural network optimization. Therefore, it would have been obvious to one of ordinary skills in the art, before the effective filing date, to combine the multi-task pretraining framework of Liu comprising a plurality of modality-specific masking losses with the hyper-parameter weighted loss aggregation scheme of Chen wherein individual loss terms are weighted by respective scaling coefficients. The motivation is taught by Chen ( Chen, Pg2, Left Column, Lines3-6, “Siamese networks can naturally introduce inductive biases for modeling invariance, as by definition “invariance” means that two observations of the same concept should produce the same outputs”, Pg4, Section4.2, Lines5-7, “Its gradient has the same direction as the gradient of D(z1, z2), with the magnitude scaled by ½” ) such that in multi-task multimodal representation learning, individual loss functions operate on fundamentally different numerical scales and produce gradient updates of varying magnitudes. Integrating hyper-parameter weighting factors for each loss component provides a routine mechanism to rebalance the relative gradient contributions of the modality-specific masking losses against the multimodal masking loss during backpropagation, which would prevent any single dominant loss component from overwhelming the backpropagation process, thereby stabilizing gradient descent updates, help numerical convergence, and prevent representation collapse across all combined modalities. However, both Liu and Chen do not explicitly teach process the multimodal input data through a multimodal model, pretrained in accordance with one or more total losses As mentioned above, Liu and Chen teach about the modality-specific losses and the multimodal loss but is silent about combining these losses to create one or more total losses. From the same field of endeavor, Chae teaches this ( Chae, Pg3, Right Column, Lines1-3, “Then, the joint loss can be calculated as follows: LJoint = we x Le + wg x Lg (3) where the values for we and wg vary between 0 and 1 such that we + wg = 1”, wherein Chae explicitly discloses joint loss, Ljoint (the corresponding total loss) that is the sum of two different losses, Le and Lg, for multimodal model, rendering it functionally equivalent to the claimed invention as Le and Lg are replaced with Liu and Chen’s modality-specific loss and multimodal loss.) Liu, Chen, and Chae are analogous to the claimed invention as they are from the same field of endeavor of multi-task deep neural network learning and multi-modal representation modeling. Therefore, it would have been obvious to one of ordinary skills in the art, before the effective filing date, to combine the multi-modal encoder-decoder pre-training framework of Liu, the predictor and stop-gradient similarity loss optimization of Chen, with the dynamic joint loss weighting mechanism updated via backpropagation of Chae. The motivation is taught by Chae (Chae, Pg1, Right Column, Paragraph2, Lines3-11, “In multi-task learning, the performance of a model is highly dependent on the joint loss weights of each task, so selecting proper weights is an important issue for multi-task learning approaches. One method to determine the joint loss weights consists in using prior knowledge of models [12]. However, these static weights can cause adverse effect because each of the loss values continuously changes while training the models”) such that the combination would dynamically adjust task weights during backpropagation to adaptively balance tasks as their loss values continuously evolve, thereby preventing performance degradation caused by static/empirical weight selection and optimizing the overall joint loss across all integrated tasks. As to dependent Claim 12, The combination of Liu, Chen and Chae teaches, as mentioned above, all the limitations of Claim 11. It teaches the overall system of receiving multimodal inputs to process them with the pre-trained model, by computing modality-specific losses and multimodal loss which are combined with a respective hyperparameter to produce a total loss, to produce the multimodal outputs. Liu further teaches a modality-specific masking loss is a measurement of the similarity between modality-specific encoded representations of the un-masked training examples and modality-specific encoded representations of the masked training examples( Liu, Pg3, Abstract, Lines4-6, "OPT is constructed in an encoder-decoder framework, including three single-modal encoders to generate token-based embeddings for each modality..." Pg3, Section3.1, Paragraph1, "Text Encoder. Following BERT [7], we first tokenize all words by WordPieces [15] to obtain the token sequence... The final embedding for each token is obtained via summing up its token embedding and position embedding..." Pg4, Section4.1, Paragraph2, Lines4-6, "Following BERT, we randomly mask 15% words with the special token in [MASK] the sentence" Pg3. Section3.1, Paragraph2, "Vision Encoder. We use Faster R-CNN [34] pre-trained on Visual Genome dataset [16] to extract the visual representations (pooled ROI features) for each image region...The final visual embedding for each region is obtained by summing up the two FC outputs and then passing through an LN layer" Pg3. Section3.1, Paragraph3, "Audio Encoder. We use pre-trained wav2vec 2.0 [3] to obtain the audio tokens and extract the features for each token. The final audio embedding is obtained by passing the audio features through an LN layer” Pg3, Section3.2, Lines1-9, "After processing the input text, image and audio using single-modal encoders, we obtain the initial text embedding T, image embedding V and audio embedding A. To model the cross-modal interations among textual words, visual regions and audio tokens, we introduce a Transformer based cross-modal encoder. Specifically, we first combine T, V and A to get a long sequence of token features, and then feed the sequence of features into the cross-modal encoder to learn contextualized representations M, M = CrossEncoder([T;V;A])", Pg4, Section4.1, Pagagraph2, "Masked Language Modeling (MLM)… The goal is to predict these masked words based on the observation of their surrounding words T_\m, all image regions V and all audio tokens A, by minimizing the negative log-likelihood: PNG media_image1.png 22 307 media_image1.png Greyscale " Pg4, Section4.1, Paragraph3, Lines1-4 and Equation3, "Masked Vision Modeling (MVM)... Similar to MLM, we also propose Masked Vision Modeling (MVM) to predict the correct image regions given contextual regions and other input modalities", PNG media_image2.png 26 291 media_image2.png Greyscale Pg4, Section4.1 Paragraph6, Lines1-6 and Equation6, "Masked Audio Modeling (MAM). For MAM, we mask audio features with a probability of 15%. Then the model is trained to reconstruct masked audio A_m, given the remaining audio tokens A_\m and all information from other modalities (i.e. text and image). Here, we propose two objectives for MAM, which share the same objective base", PNG media_image3.png 26 287 media_image3.png Greyscale , Pg5, Equation8, PNG media_image4.png 124 324 media_image4.png Greyscale Pg4, Right Column, Paragraph1, Lines1-7, “The first objective is Masked Visual Feature Regression (MVFR), which regresses the cross-modal encoder output of each masked region to its input ROI visual features Vm m. We use an additional FC layer to transform the out V put of cross-modal encoder to the same dimensional space as the input visual feature. Then we apply L2 regression between the two” wherein Liu explicitly discloses single-modal encoders (the corresponding modality-specific encoders) for input modalities. Also, Liu discloses both token-level masking and modality-level masking, where the single-modal encoders generate encoded representations for both un-masked inputs (complete sequences T,V,A or un-masked tokens T_\m, V_\m, A_\m) and masked inputs (inputs containing [MASK] tokens/masked features or modality level masked inputs). Liu also explicitly discloses computing distinct losses for each modality as shown in the three equations using masked and unmasked representations of (T_\m, V_\m, A_\m) and (T_m, V_m, A_m) where at least audio and vision modalities use idea of similarity to compute the loss, rendering it functionally equivalent to the claimed invention), and fused encoded representations generated from modality-specific encoded representations ( Liu, Pg3, Section3.2, Lines1-9, "After processing the input text, image and audio using single-modal encoders, we obtain the initial text embedding T, image embedding V and audio embedding A. To model the cross-modal interations among textual words, visual regions and audio tokens, we introduce a Transformer based cross-modal encoder. Specifically, we first combine T, V and A to get a long sequence of token features, and then feed the sequence of features into the cross-modal encoder to learn contextualized representations M, M = CrossEncoder([T;V;A])", wherein Liu explicitly teaches about generating fused representations M, rendering it functionally equivalent to the claimed invention.) However, Liu does not explicitly teach the one or more multimodal masking losses are measurements of similarities between fused encoded representations generated from modality-specific encoded representations for the masked training examples and fused encoded representations generated from modality-specific encoded representations for the un-masked training examples. As mentioned above, Liu teaches about generating a fused encoded representation, M, from the modality-specific encoded representations. Liu does not teach that the output from M should be comprising both masked and unmasked fused encoded representations and using them to compute the multimodal masking loss. From the same field of endeavor, Chen teaches this ( Chen, Pg2, Section3, Lines1-9, PNG media_image5.png 278 464 media_image5.png Greyscale Pg3, Left Column, Line1 and Equation2, PNG media_image6.png 86 418 media_image6.png Greyscale , wherein Chen discloses using two augmented inputs x1 and x2 into the encoder f to generate z1,z2 (the corresponding first and second fused encoding representations) and the output vectors p1,p2. If this pipeline is combined with Liu’s fused encoding representations, it would create multiple, first and second, fused encoded representations of the modality-specific encoded representations. As f(x1) and f(x2) correspond to the first and second fused representations generated by the encoder f as mentioned above, while p1 and p2 represent their projected vectors transformed by the prediction head h. Consequently, the term D(p1, z2) calculates the negative cosine similarity between the projection of the first representation,p1, and the second representation,z2, whereas D(p2,z1) calculates the negative cosine similarity between the projection of the second representation,p2, and the first representation,z1. Combining these terms in the symmetrized formulation in Equation2 to compute the multimodal masking losses of the first and second fused representations mentioned above renders it functionally equivalent to the claimed invention.) As to dependent Claim 13, The combination of Liu, Chen and Chae teaches, as mentioned above, all the limitations of Claim 12. It teaches about the modality-specific masking loss and the multimodal masking loss. Liu further teaches the system of claim 12, wherein the one or more processors are further configured to: receive training data comprising the masked and un-masked training examples ( Liu, Pg6, Section 5.1, Lines1-4, "We mainly use Open Images [18] dataset with localized narratives and synchronized speech that are provided by [29] as our pre-training dataset. Only text-image-audio triplets are used for pre-training" Pg2, Left Column, Paragraph1, Lines7-9, "Token-level modeling predicts the semantics of masked tokens given the unmasked inputs" Pg5, Section 4.2, Paragraph2, Lines3-8, "Modality-level masking is in parallel with token-level masking mechanism. It masks out one or two modalities from the input. specifically, each modality is independently masked out with a probability of 0.3, and the case when all modalities are masked is skipped" Pg6, Section5.2, Lines13-15, "Models are trained on 4 Tesla V100 GPUs with a total batch size of 10,240 for 100,000 iterations, and early stop is performed", wherein Liu teaches that the models are trained on GPUs (hereinafter any processors will be referring to this) which inherently comprises a system to train the model that consists processors. Liu also explicitly teaches receiving multimodal training data consisting of image, text and audio modalities. Liu also teaches generating and receiving both masked and unmasked training examples through two masking schemes of token-level masking, which masks a portion of tokens within an input sequence while maintaining the remaining tokens as un-masked, and modality-level masking, which masks entire modalities independently with a probability of 0.3 that the case where all modalities are masked is skipped that the model receives inputs containing a combination of masked and remaining un-masked modalities, rendering it functionally equivalent to the claimed invention); and perform one or more pretraining iterations, comprising ( Liu, Pg6, Section 5.1, Lines1-4, "We mainly use Open Images [18] dataset with localized narratives and synchronized speech that are provided by [29] as our pre-training dataset. Only text-image-audio triplets are used for pre-training"): generating modality-specific encoded representations of the un-masked training examples and masked training examples ( Liu, Pg3, Abstract, Lines4-6, "OPT is constructed in an encoder-decoder framework, including three single-modal encoders to generate token-based embeddings for each modality..." Pg3, Section3.1, Paragraph1, "Text Encoder. Following BERT [7], we first tokenize all words by WordPieces [15] to obtain the token sequence... The final embedding for each token is obtained via summing up its token embedding and position embedding..." Pg4, Section4.1, Paragraph2, Lines4-6, "Following BERT, we randomly mask 15% words with the special token in [MASK] the sentence" Pg3. Section3.1, Paragraph2, "Vision Encoder. We use Faster R-CNN [34] pre-trained on Visual Genome dataset [16] to extract the visual representations (pooled ROI features) for each image region...The final visual embedding for each region is obtained by summing up the two FC outputs and then passing through an LN layer" Pg3. Section3.1, Paragraph3, "Audio Encoder. We use pre-trained wav2vec 2.0 [3] to obtain the audio tokens and extract the features for each token. The final audio embedding is obtained by passing the audio features through an LN layer” Pg3, Section3.2, Lines1-9, "After processing the input text, image and audio using single-modal encoders, we obtain the initial text embedding T, image embedding V and audio embedding A. To model the cross-modal interations among textual words, visual regions and audio tokens, we introduce a Transformer based cross-modal encoder. Specifically, we first combine T, V and A to get a long sequence of token features, and then feed the sequence of features into the cross-modal encoder to learn contextualized representations M, M = CrossEncoder([T;V;A])", wherein Liu explicitly discloses single-modal encoders (the corresponding modality-specific encoders) for input modalities. Also, Liu discloses both token-level masking and modality-level masking, where the single-modal encoders generate encoded representations for both un-masked inputs (complete sequences T,V,A or un-masked tokens T_\m, V_\m, A_\m) and masked inputs (inputs containing [MASK] tokens/masked features or modality level masked inputs), rendering it functionally equivalent to the claimed invention), determining the plurality of modality-specific masking losses from the modality-specific encoded representations ( Liu, Pg4, Section4.1, Pagagraph2, "Masked Language Modeling (MLM)… The goal is to predict these masked words based on the observation of their surrounding words T_\m, all image regions V and all audio tokens A, by minimizing the negative log-likelihood: PNG media_image1.png 22 307 media_image1.png Greyscale " Pg4, Section4.1, Paragraph3, Lines1-4 and Equation3, "Masked Vision Modeling (MVM)... Similar to MLM, we also propose Masked Vision Modeling (MVM) to predict the correct image regions given contextual regions and other input modalities", PNG media_image2.png 26 291 media_image2.png Greyscale Pg4, Section4.1 Paragraph6, Lines1-6 and Equation6, "Masked Audio Modeling (MAM). For MAM, we mask audio features with a probability of 15%. Then the model is trained to reconstruct masked audio A_m, given the remaining audio tokens A_\m and all information from other modalities (i.e. text and image). Here, we propose two objectives for MAM, which share the same objective base", PNG media_image3.png 26 287 media_image3.png Greyscale Pg5, Equation8, PNG media_image4.png 124 324 media_image4.png Greyscale Pg4, Right Column, Paragraph1, Lines1-7, “The first objective is Masked Visual Feature Regression (MVFR), which regresses the cross-modal encoder output of each masked region to its input ROI visual features Vm m. We use an additional FC layer to transform the out V put of cross-modal encoder to the same dimensional space as the input visual feature. Then we apply L2 regression between the two”, wherein Liu explicitly discloses computing distinct losses for each modality as shown in the three equations using masked and unmasked representations of (T_\m, V_\m, A_\m) and (T_m, V_m, A_m) where at least audio and vision modalities use idea of similarity to compute the loss, rendering it functionally equivalent to the claimed invention.) generating a fused encoded representation ( Liu, Pg3, Section3.2, Lines1-9, "After processing the input text, image and audio using single-modal encoders, we obtain the initial text embedding T, image embedding V and audio embedding A. To model the cross-modal interations among textual words, visual regions and audio tokens, we introduce a Transformer based cross-modal encoder. Specifically, we first combine T, V and A to get a long sequence of token features, and then feed the sequence of features into the cross-modal encoder to learn contextualized representations M, M = CrossEncoder([T;V;A])" wherein Liu explicitly teaches about generating fused representations M, rendering it functionally equivalent to the claimed invention), updating, by the one or more processors, one or more weights of the multimodal model ( Liu, Pg4, Section4, "We propose three levels of pre-training tasks: (1) token-level modeling, including masked language modeling (MLM), masked vision modeling (MVM), and masked audio modeling (MAM); (2) modality-level modeling, including denoising text reconstruction (DTR) and denoising image reconstruction (DIR); and (3) sample-level modeling" Pg6, Section5.2, Lines13-17, "Models are trained on 4 Tesla V100 GPUs with a total batch size of 10,240 for 100,000 iterations, and early stop is performed. We adopt the Adam optimizer with an initial learning rate of 5e-5", wherein Liu explicitly updates these network weights end-to-end over 100,000 iterations using the backpropagation of the Adam optimizer, rendering it functionally equivalent to the claimed invention.) However, Liu does not explicitly teach (highlighted sections) generating a first fused encoded representation of the un-masked training example encoded representations and a second fused encoded representation of the masked training example encoded representations determining the multimodal masking loss from the first and the second fused encoded representations updating, by the one or more processors, one or more weights of the multimodal model in accordance with both the plurality of modality-specific masking losses and the multimodal masking loss As mentioned above, Liu teaches about generating fused encoding representations but is silent about “a first fused encoded representation of the un-masked training example encoded representations and a second fused encoded representation of the masked training example encoded representations.” From the same field of endeavor, Chen teaches this ( Chen, Pg2, Section3, Lines1-9, PNG media_image5.png 278 464 media_image5.png Greyscale , Pg3, Left Column, Line1 and Equation2, PNG media_image6.png 86 418 media_image6.png Greyscale Wherein Chen discloses using two augmented inputs x1 and x2 into the encoder f to generate z1,z2 (the corresponding first and second fused encoding representations) and the output vectors p1,p2. If this pipeline is combined with Liu’s fused encoding representations, it would create a first and a second fused encoded representations of masked and unmasked training examples, rendering it functionally equivalent to the claimed invention.) Chen further teaches determining a multimodal masking loss measuring the similarity between the first fused encoded representation and the second fused encoded representation ( Chen, Pg2, Section3, Lines7-9 and Equation1, PNG media_image7.png 138 462 media_image7.png Greyscale Pg3, Left Column, Line1 and Equation2, PNG media_image6.png 86 418 media_image6.png Greyscale wherein Chen discloses f(x1) and f(x2) correspond to the first and second fused representations generated by the encoder f as mentioned above, while p1 and p2 represent their projected vectors transformed by the prediction head h. Consequently, the term D(p1, z2) calculates the negative cosine similarity between the projection of the first representation,p1, and the second representation,z2, whereas D(p2,z1) calculates the negative cosine similarity between the projection of the second representation,p2, and the first representation,z1. Combining these terms in the symmetrized formulation in Equation9 to compute the multimodal masking losses of the first and second fused representations mentioned above renders it functionally equivalent to the claimed invention.) As mentioned above, Liu teaches about updating the weights of the multimodal model, but is silent about updating the weights in accordance with both the plurality of modality-specific masking losses and the multimodal masking loss. Chen teaches this ( Pg3, Left Column, Line1 and Equation2, PNG media_image6.png 86 418 media_image6.png Greyscale , Wherein Chen, as mentioned above, computes the multimodal masking loss which can be directly incorporated into Liu’s backpropagation of Adam optimizer that updates the weights, which relies on using the loss function, rendering it functionally equivalent to the claimed invention when combined.) Both Liu and Chen are analogous to the claimed invention as they are from the same field of endeavor of self-supervised representation learning utilizing Siamese neural network architectures to minimize similarity loss between generated representations. Therefore, it would have been obvious to one of ordinary skills in the art, before the effective filing date, to combine the multimodal pre-training scheme of Liu with dual-view similarity loss computation of Chen. The motivation is as recited by Chen (Chen, Pg1, Abstract, Lines3-12, " These models maximize the similarity between two augmentations of one image, subject to certain conditions for avoiding collapsing solutions. In this paper, we report surprising empirical results that simple Siamese networks can learn meaningful representations even using none of the following: (i) negative sample pairs, (ii) large batches, (iii) momentum encoders. Our experiments show that collapsing solutions do exist for the loss and structure, but a stop-gradient operation plays an essential role in preventing collapsing" Pg2, Section3, Lines1-7, “Our architecture (Figure 1) takes as input two randomly augmented views x1 and x2 from an image x. The two views are processed by an encoder network f consisting of a backbone (e.g., ResNet [19]) and a projection MLP head [8]. The encoder f shares weights between the two views. A prediction MLP head [15], denoted as h, transforms the output of one view and matches it to the other view”) such that incorporating a similarity loss between different (masked and unmasked) representations enables the multimodal model to directly maximize cross-modal feature alignment and improve representation learning efficiency without requiring complex negative sampling or momentum encoders, while utilizing a stop-gradient operation prevents output collapsing during backpropagation, yielding predictable training stability. As to dependent Claim 14, it is a system claim that contains similar limitations of Claim 10 and thus rejected under the same rationale. As to independent Claim 18, it is a non-transitory computer-readable medium claim that contains similar limitations of Claim 11 and thus rejected under the same rationale. As to dependent Claim 19, it is a non-transitory computer-readable medium claim that contains similar limitations of Claim 12 and thus rejected under the same rationale. As to dependent Claim 20, it is a non-transitory computer-readable medium claim that contains similar limitations of Claim 13 and thus rejected under the same rationale. Claims 15-17 are rejected under 35 U.S.C. 103 as being unpatentable over Liu, Chen and Chae as mentioned in Claim 14 in further view of Arik et al. (Arik), Non-Patent Literature listed in IDS filed on 04/18/2024, “TabNet: Attentive Interpretable Tabular Learning”, Published in May 2021, AAAI, 9 Pages. As to dependent Claim 15, The combination of Liu, Chen and Chae teaches, as mentioned above, all the limitations of Claim 13. It teaches the overall architecture of receiving multimodal masked and unmasked training data to generate modality specific data representation of the data according to its type, which then are fused together to generate fused representations of masked and unmasked data to train the multimodal model by backpropagation. Liu, as mentioned above, uses multimodal training data that comprises types of text, visual, and audio, but Liu does not teach The method of claim 1, wherein the multimodal training data comprises data of a plurality of modalities, comprising structured data. From the same field of endeavor, Arik teaches this limitation ( Arik, Pg6679, Right Column, Lines1-5, "In addition, unlike tree learning, DNNs enable gradient descent based end-to-end learning for tabular data which can have a multitude of benefits: (i) efficiently encoding multiple data types like images along with tabular data" Pg6679, Introduction, Lines1-3, "Deep neural networks (DNNs) have shown notable success with images (He et al. 2015), text (Lai et al. 2015) and audio (Amodei et al. 2015)" Pg6679, Introduction, Lines5-9, "One data type that has yet to see such success with a canonical architecture is tabular data. Despite being the most common data type in real-world AI (as it is comprised of any categorical and numerical features)", wherein Arik explicitly identify tabular data, composed of categorical numerical features, as structured data, while separately enumerating unstructured modalities including text, image, and audio, rendering it functionally equivalent to the claimed invention of using training data comprises structured data and unstructured data such as text, image or audio.) Liu, Chen, Chae and Arik are analogous to the claimed invention as they are from the same field of endeavor of self-supervised multimodal representation learning and attention-based deep neural network architectures for processing heterogeneous data. Therefore, it would have been obvious to one of ordinary skills in the art, before the effective filing date, to combine the multimodal pre-training scheme of Liu and dual-view similarity loss computation of Chen, the dynamic joint loss weighting mechanism updated via backpropagation of Chae with the tabular data encoding architecture of Arik. The motivation is as recited by Arik (Arik, Pg6679, Right Column, Lines1-10, "In addition, unlike tree learning, DNNs enable gradient descent-based end-to-end learning for tabular data which can have a multitude of benefits: (i) efficiently encoding multiple data types like images along with tabular data; (ii) alleviating the need for feature engineering, which is currently a key aspect in tree-based tabular data learning methods; (iii) learning from streaming data and perhaps most importantly (iv) end-to-end models allow representation learning which enables many valuable application scenarios") such that the combination of them would integrate Arik’s tabular data encoding into Liu’s multimodal framework while applying Chen’s dual-view similarity loss to align masked and unmasked representations and Chae’s dynamic adjustment of task weights during backpropagation to adaptively balance tasks as their loss values continuously evolve, thereby achieving a unified, end-to-end deep learning model capable of jointly learning representations across both structured and unstructured modalities without relying on disjoint tree-based models or manual feature engineering. As to dependent Claim 16, The combination of Liu, Chen and Chae teaches, as mentioned above, all the limitations of Claim 15. It teaches about the multimodal training data comprises structured data and text data, image data, video data, or audio data. Liu further teaches the system of claim 15, wherein the multimodal input data comprises: one or more instances comprise data of each modality of the plurality of modalities present in the multimodal training data ( Liu, Pg1, Right Column, Paragraph1, Lines4-8, "OPT is pre-trained on large amounts of language-vision-audio triplets with a multi-task pretext learning scheme, and can effectively adapt to downstream understanding and generation tasks given single-, two-, or three-modal inputs" Pg6, Right Column, Paragraph1, Lines11-15, "When with multimodal features, the performance can be further improved. In particular, adding text feature brings the largest improvement. On the basis of image+text feature, we further add audio feature, and find the performance still increases" Pg6, Table2, PNG media_image12.png 254 370 media_image12.png Greyscale Pg8, Figure3, Line1, "Some results of text-to-image generation and text generation", "Both: In this image I can see the candle...", PNG media_image13.png 542 1136 media_image13.png Greyscale , wherein Liu explicitly teaches evaluating downstream test instances where all three pretraining modalities, text, vision, and audio, are concurrently provided as input to the model during inference, as demonstrated in its full three-modality classification experiments in Table2 and combined text-generation tasks), and each instance other than the one or more instances comprises different respective combinations of at least partially missing data of at least one modality of the plurality of modalities present in the multimodal training data ( Liu, Pg2, Left Column, Third Bullet, “OPT can effectively adapt to and perform competitively on a series of cross-modal understanding and generation downstream tasks with parial or all modalities as inputs” Pg6, Table2, PNG media_image12.png 254 370 media_image12.png Greyscale Pg8, Figure3, PNG media_image13.png 542 1136 media_image13.png Greyscale Pg6, Right Column, Lines7-12, “When only using image features, our OPT outperforms ResNet-50 and ResNet-101 by a large margin. When only using text or audio feature, our model also obtains promising results, which indicates that the model has learnt the associations between different modalities” wherein Liu describes processing inference instances where on or more modalities present during pre-training are missing, resulting in various combinations of partial modality inputs by testing downstream performance across multiple partial-modality input configurations, including single-modality inputs (text-only, image-only, audio-only) and two-modality combinations (image+text, audio+text, audio+image), demonstrating the model yields accurate task outputs even when specific pre-training modalities are entirely omitted, rendering it functionally equivalent to the claimed invention.) As to dependent Claim 17, The combination of Liu, Chen and Chae teaches, as mentioned above, all the limitations of Claim 11. It teaches the overall system of receiving multimodal inputs to process them with the pre-trained model, by computing modality-specific losses and multimodal loss which are combined with a respective hyperparameter to produce a total loss, to produce the multimodal outputs. Liu further teaches the system of claim 11, wherein in generating the model output, the one or more processors are configured to: receive labeled training data comprising labeled training examples corresponding to a machine learning task ( Liu, Pg6, Table1, PNG media_image8.png 186 355 media_image8.png Greyscale Pg5, Section4.3 Lines16-17, "where gt(T,V,A) is the one-hot vector of ground-truth label" Pg2, Paragraph2, Lines1-4, "We extensively validate OPT on a number of downstream tasks, including cross-modal retrieval, multi-modal classification, visual question answering, cross-modal text generation" wherein Liu explicitly teaches receiving labeled datasets such as COCO Captions and VQA datasets which consist of paired multimodal training examples along with ground-truth one-hot vectors(gt) to train and evaluate the model across various downstream machine learning tasks, rendering it functionally equivalent to the claimed invention); and perform one or more fine-tuning iterations of: ( Liu, Pg1, Right Column, Paragraph1, Lines4-8, "OPT is pre-trained on large amounts of language-vision-audio triplets with a multi-task pretext learning scheme, and can effectively adapt to downstream understanding and generation tasks given single-, two-, or three-modal inputs" Pg6, Section5.2, Lines13-15, "Models are trained on 4 Tesla V100 GPUs with a total batch size of 10,240 for 100,000 iterations, and early stop is performed" Pg6, Right Column, Multi-Modal Classification Section, Lines2-5, "We add a linear layer after the average pooling output of cross-modal encoder for classification. We freeze our pre-trained model and only linear layer is learned" wherein Liu describes how the pretrained model adapts to downstream tasks through iterative training loops over up to 100,000 iterations, either updating model parameters or optimizing task-specific classification head layers, rendering it functionally equivalent to the claimed invention): processing the labeled training data through the multimodal model ( Pg3, Section3.2, Lines6-9, "Specifically, we first combine T, V and A to get a long sequence of token features, and then feed the sequence of features into the cross-modal encoder to learn contextualized representations M" Pg4, Lines4-6, "When performing on down-stream, the two decoders is used to generate results, e.g., image captioning, text-to-image generation" Pg6, Right Column, Multi-Modal Classification Section, Lines2-4, "We add a linear layer after the average pooling output of cross-modal encoder for classification", wherein Liu processes labeled multimodal inputs by combining token features across modalities (T,V,A) and passing them through the Transformer-based cross-modal encoder and task decoders/classifiers during downstream task execution, rendering it functionally equivalent to the claimed invention), determining a task-specific loss measuring performance of the multimodal model in performing the machine learning task ( Pg4, Right Column, Paragraph2, Lines9-10, "The final objective minimizes the cross-entropy (CE) loss: PNG media_image9.png 41 346 media_image9.png Greyscale " Pg5, Right Column, Section4.3, Lines14-15, "The loss function is the binary cross entropy (BCE) loss: PNG media_image10.png 41 314 media_image10.png Greyscale " Pg5, Right Column, Paragraph1, Lines3-4, "The loss function is, PNG media_image11.png 30 313 media_image11.png Greyscale " wherein Liu teaches computing distinct, task-specific objective functions such as Cross-Entropy loss for classification tasks, Binary Cross-Entropy for sample-level matching, and negative log-likelihood for sequence generation, by comparing predicted model outputs against ground-truth target labels(gt), rendering it functionally equivalent to the claimed invention), and updating one or more weights of the multimodal model in accordance with the task-specific loss ( Liu, Pg6, Section5.2, Lines16-17, "We adopt the Adam optimizer with an initial learning rate of 5e-5" Pg6, Right Column, Multi-Modal Classification Section, Lines4-5, "We freeze our pre-trained model and only linear layer is learned" Pg4, Section4.1, Paragraph2, Line10, "where _theta_ is the trainable parameters", wherein Liu describes optimizing trainable parameters, theta, using backpropagation with the Adam optimizer, wherein weights of the model (or added task-specific linear layers) are updated to minimize the task loss, rendering it functionally equivalent to the claimed invention.) Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Jonathan et al. US Patent Publication, US-20230039210, Published in February 2023 Xi et al. Chinese Patent Application, CN-112784902-A, Published in May 2021 Kollada et al. WO Patent Application, WO-2022256193-A2, Published in December 2022 Bin et al. Chinese Patent Application, CN-114840734-A, Published in August 2022 HOEHNE et al. WO Patent Application, WO-2022184516-A1, Published in September 2022 Any inquiry concerning this communication or earlier communications from the examiner should be directed to DONG YOON JUNG whose telephone number is (571)270-0198. The examiner can normally be reached 8am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Cesar Paula can be reached at (571) 272-4128. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /DONG YOON JUNG/ Examiner, Art Unit 2145 /CESAR B PAULA/ Supervisory Patent Examiner, Art Unit 2145
Read full office action

Prosecution Timeline

Apr 18, 2024
Application Filed
Sep 10, 2026
Non-Final Rejection mailed — §101, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month