Prosecution Insights
Last updated: October 02, 2026
Application No. 18/626,959

SYSTEMS AND METHODS FOR GESTURE GENERATION FROM TEXT AND NON-SPEECH

Non-Final OA §101§103§112
Filed
Apr 04, 2024
Priority
Apr 06, 2023 — provisional 63/457,561
Examiner
BALAKRISHNAN, VIJAY MURALI
Art Unit
Tech Center
Assignee
Datum Point Labs Inc.
OA Round
1 (Non-Final)
41%
Grant Probability
Moderate
1-2
OA Rounds
1y 5m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 41% of resolved cases
41%
Career Allowance Rate
11 granted / 27 resolved
-19.3% vs TC avg
Strong +73% interview lift
Without
With
+73.3%
Interview Lift
resolved cases with interview
Typical timeline
3y 11m
Avg Prosecution
15 currently pending
Career history
44
Total Applications
across all art units

Statute-Specific Performance

§101
27.5%
-12.5% vs TC avg
§103
36.6%
-3.4% vs TC avg
§102
12.6%
-27.4% vs TC avg
§112
23.3%
-16.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 27 resolved cases

Office Action

§101 §103 §112
DETAILED ACTION This nonfinal action is in response to application 18/626,959 filed on 04/04/2024 with priority to provisional application 63/457,561 filed on 04/06/2023. Claims 1-20 are pending in the application. Claims 1, 8, and 14 are independent claims. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement (IDS) filed 07/12/2024 has been fully considered. Specification The specification is objected to because the title of the invention, “Systems and Methods for Gesture Generation from Text and Non-Speech”, is not clearly indicative of the invention to which the claims are directed. The recited gesture generation from “non-speech” appears to be contradictory to claimed subject matter, which expressly recites utilization of text and speech input for training a gesture generation model (e.g., claims 2, 4-6, 15, 17-19) and generating gestures (e.g., claims 9, 12). The specification also does not account for this discrepancy, as it only briefly recites “gesture generation from text and non-speech” in [¶ 0003] and [¶ 0014] without providing any further clarification on the term “non-speech”, or providing description of an embodiment that expressly utilizes “non-speech”. Either clarification regarding the usage of “non-speech” in the title, or correction of the title and corresponding description ([¶ 0003] and [¶ 0014]) to remove usage of the term “non-speech”, is requested. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 8-13 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Regarding claim 8, it recites the limitation “generating, via the generator, a first pose”. There is insufficient antecedent basis for the term “the generator” in the claims. Consequently, the intended scope of the invention is unclear. For purposes of examination, the limitation is interpreted as “generating, via a generator, a first pose”. Regarding claim 9, it recites the limitation “generating, via a generator, reconstructed speech based on speech features”. It is unclear if the recited “generator” is in reference to the “generator” previously recited in parent claim 8, or an entirely separate generator. Consequently, the intended scope of the invention is unclear For purposes of examination, the limitation is interpreted as “generating, via the generator, reconstructed speech based on speech features”. Regarding claims 10-13, they inherit the deficiencies of their parent claims. Consequently, they are also rejected under 35 U.S.C. 112(b) as being indefinite for depending on an indefinite parent claim. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The analysis of the claims will follow the 2019 Revised Patent Subject Matter Eligibility Guidance, 84 Fed. Reg. 50 (“2019 PEG”). Independent Claims (Claim 1, Claim 8, Claim 14): Step 1: Claim 1 is drawn to a method, claim 8 is drawn to a method, and claim 14 is drawn to a system/apparatus. Therefore, each of these claims falls under one of the four categories of statutory subject matter (process/method, machine/apparatus, manufacture/product, or composition of matter). Step 2A Prong 1: Claims 1, 8, and 14 each recite a judicially recognized exception of an abstract idea. Claim 1 recites, inter alia: masking a subset of the multimodal input; generating a multimodal embedding based on the masked multimodal input; generating multimodal features based on the multimodal embedding; generating multimodal output based on the multimodal features; – These limitations recite a series of generic data observation and processing steps that result in a translation of data into a numerical representation, and thereby recite a series of algorithmic steps that could be reasonably performed by a human using pen and paper, and/or recite a process of mathematical calculation. computing a loss based on the multimodal input and the multimodal output – This limitation recites using mathematical methods to calculate a difference between determined values, and thereby recites mathematical calculation. Claim 8 recites, inter alia: generating an embedding based on the multimodal input; generating multimodal features based on the embedding; generating first motion features based on the multimodal features; generating second motion features based on the first motion features and the multimodal features; and generating a first pose based on the first motion features and second pose based on the second motion features – These limitations recite a series of generic data observation and processing steps that result in a translation of data into a numerical representation, and thereby recite a series of algorithmic steps that could be reasonably performed by a human using pen and paper, and/or recite a process of mathematical calculation. Generating “poses” from determined “motion features” may further be e.g., reasonably ascribed to a workflow of visualizing and drafting sketches based on determined features, which is a process that could be reasonably performed by a human on pen and paper. Claim 14 recites substantially similar abstract idea limitations to those found in claim 1, and thereby recites the same judicial exception. Step 2A Prong 2: The following additional elements recited in claims 1, 8, and 14 do not integrate the recited judicial exceptions into a practical application. Claim 1 additionally recites: A method for training a co-speech gesture generation model, the method comprising: [generating] via an embedder; [generating] via an encoder; [generating] via a generator – These limitations do no more than invoke generic machine learning (ML) model components as tools to perform steps of an existing abstract procedure, and thereby amount to mere instructions to “apply” an exception. Generally linking the recited abstract procedure to implementation on a generic ML model does not provide integration into a practical application. receiving, via a data interface, a multimodal input – This limitation amounts to a step of merely gathering data to enable further analysis, and therefore recites insignificant extra-solution activity. wherein the encoder includes one or more attention layers connecting different modalities – This limitation amounts to an insignificant step of model implementation, as it does no more than ascribe a generic attention mechanism to a model component (encoder), wherein said component is merely being invoked as a tool to perform a step of the recited abstract procedure. updating parameters of the encoder based on the loss – Generic recitation of parameter updation does no more than generally link the recited abstract procedure to implementation on an ML model, and thereby does not provide integration into a practical application. Claim 8 additionally recites: A method for co-speech gesture generation, the method comprising: [generating] via an embedder; [generating] via an encoder; [generating] via a decoder; [generating] via the decoder; and [generating] via the generator – These limitations do no more than invoke generic machine learning (ML) model components as tools to perform steps of an existing abstract procedure, and thereby amount to mere instructions to “apply” an exception. Generally linking the recited abstract procedure to implementation on a generic ML model does not provide integration into a practical application. receiving, via a data interface, multimodal input – This limitation amounts to a generic step of mere data gathering to enable further analysis, and therefore recites insignificant extra-solution activity. Claim 14 recites substantially similar additional elements to those found in claim 1, and further recites: A system [for training], the system comprising: a memory that stores a plurality of processor-executable instructions; one or more processors that read and execute the plurality of processor-executable instructions from the memory to perform operations – These limitations amount to mere instructions to implement an abstract idea on a computer or computer components. Step 2B: The additional elements recited in claims 1, 8, and 14, viewed individually or as an ordered combination, do not provide an inventive concept or otherwise amount to significantly more than the recited abstract ideas themselves. Claim 1 additionally recites: A method for training a co-speech gesture generation model, the method comprising: [generating] via an embedder; [generating] via an encoder; [generating] via a generator – Mere instructions to “apply” an exception on generic machine learning (ML) model components do not provide an inventive concept or significantly more to the recited abstract idea. receiving, via a data interface, a multimodal input – Receiving data via network components is well-understood, routine, and conventional activity (see MPEP § 2106.05(d); “Receiving or transmitting data over a network”) and therefore does not provide an inventive concept or significantly more to the recited abstract idea. wherein the encoder includes one or more attention layers connecting different modalities – Utilization of attention mechanisms in transformer models to process input features is well-understood, routine, and conventional activity (see Cristina, “The Transformer Attention Mechanism”, [pages 1-2]) and therefore does not provide an inventive concept or significantly more to the recited abstract idea. updating parameters of the encoder based on the loss – Generally linking the recited abstract procedure to implementation on an ML model via mere recitation of parameter updation does not provide an inventive concept or significantly more to the recited abstract idea. Claim 8 additionally recites: A method for co-speech gesture generation, the method comprising: [generating] via an embedder; [generating] via an encoder; [generating] via a decoder; [generating] via the decoder; and [generating] via the generator – Mere instructions to “apply” an exception on generic machine learning (ML) model components do not provide an inventive concept or significantly more to the recited abstract idea. receiving, via a data interface, multimodal input – Receiving data via network components is well-understood, routine, and conventional activity (see MPEP § 2106.05(d); “Receiving or transmitting data over a network”) and therefore does not provide an inventive concept or significantly more to the recited abstract idea. Claim 14 recites substantially similar additional elements to those found in claim 1, and further recites: A system [for training], the system comprising: a memory that stores a plurality of processor-executable instructions; one or more processors that read and execute the plurality of processor-executable instructions from the memory to perform operations – Mere instructions to implement an abstract idea on a computer or computer components do not provide an inventive concept or significantly more to the recited abstract idea. Even when considered as an ordered combination, the additional elements recited in the claims ultimately do no more than place the claims in the context of generic implementation of an abstract procedure on machine learning model components, rather than being directed towards improvement of the operation of a machine learning model itself. As such, claims 1, 8, and 14 are not patent eligible. Dependent Claims (Claims 2-7, Claims 9-13, Claims 15-20): Dependent claims 2-7, 9-13, and 15-20 narrow the scope of independent claims 1, 8, and 14, and likewise narrow the recited judicial exceptions. They recite abstract idea limitations that are similar to those recited within the independent claims (i.e., mental processes and/or mathematical concepts), and thereby merely expand on the already recited exceptions. The dependent claims also do not recite any further additional elements that successfully integrate the recited judicial exceptions into a practical application or provide significantly more than the recited abstract ideas themselves. Consequently, claims 2-7, 9-13, and 15-20 are also rejected under 35 U.S.C. 101. Step 1: Claims 2-7 are drawn to a method, claims 9-13 are drawn to a method, and claims 15-20 are drawn to a system/apparatus. Therefore, each of these claims falls under one of the four categories of statutory subject matter (process/method, machine/apparatus, manufacture/product, or composition of matter). Step 2A Prong 1: Claims 2-7, 9-13, and 15-20 each recite a judicially recognized exception of an abstract idea. Claim 2 recites, inter alia: generating a second multimodal embedding based on the second multimodal input, wherein the second multimodal embedding includes a speech embedding, text embedding, and a pose embedding; generating a second multimodal output based on the second multimodal embedding; – These limitations recite a series of generic data observation and processing steps that result in a translation of data into a numerical representation, and thereby recite a series of algorithmic steps that could be reasonably performed by a human using pen and paper, and/or recite a process of mathematical calculation. computing a first embedding loss based on the text embedding and the speech embedding; computing a second embedding loss based on the text embedding and the pose embedding; – These limitations recite using mathematical methods to calculate a difference between determined values, and thereby recite mathematical calculation. Claim 3 recites, inter alia: computing a third embedding loss based on the second multimodal input and the second multimodal output – This limitation recites using mathematical methods to calculate a difference between determined values, and thereby recite mathematical calculation. Claim 4 recites the same judicial exception as claim 1. Claim 5 recites, inter alia: generat[ing] a text embedding in the multimodal embedding based on the text input; generat[ing] a speech embedding in the multimodal embedding based on the speech input; generat[ing] a pose embedding in the multimodal embedding based on the pose input – These limitations recite a series of generic data observation and processing steps that result in a translation of data into a numerical representation, and thereby recite a series of algorithmic steps that could be reasonably performed by a human using pen and paper, and/or recite a process of mathematical calculation. Claim 6 recites, inter alia: wherein masking the multimodal input comprises at least one of: removing all text input based on a first probability; removing a first subset of the text input based on a second probability; replacing a second subset of the text input with random words based on a third probability; removing all speech input based on a fourth probability; removing a first subset of the speech input based on a fifth probability; replacing a second subset of the speech input with random speech based on a sixth probability; removing all pose input based on a seventh probability; removing a first subset of the pose input based on an eighth probability; or replacing a second subset of the pose input with random poses based on a ninth probability – This limitation recites a series of data observation and alteration steps resulting in replacement/removal of data, and thereby recites a process of evaluation that a human could reasonably perform in the mind or using pen and paper. Claim 7 recites, inter alia: wherein the generating a multimodal embedding includes first processing an intermediate representation based on the multimodal input – These limitation recites a generic data processing step that results in a translation of data into a numerical representation, and thereby recite an algorithmic step that could be reasonably performed by a human using pen and paper, and/or recites a process of mathematical calculation. Claim 9 recites, inter alia: generating reconstructed speech based on speech features and reconstructed text based on text features – These limitations recite a series of generic data observation and processing steps that result in a translation of data into a numerical representation, and thereby recite a series of algorithmic steps that could be reasonably performed by a human using pen and paper, and/or recite a process of mathematical calculation. computing a first loss based on reconstructed speech, speech input, reconstructed text, and text input; computing a second loss based on the first pose, the second pose, and the pose input – These limitations recite using mathematical methods to calculate a difference between determined values, and thereby recite mathematical calculation. Claim 10 recites the same judicial exception as claim 8. Claim 11 recites the same judicial exception as claim 9. Claim 12 recites the same judicial exception as claim 8. Claim 13 recites the same judicial exception as claim 8. Claims 14-20 recite substantially similar abstract idea limitations to those found in claims 1-7, and thereby recite the same judicial exceptions. Step 2A Prong 2: Claims 6-7 and 19-20 do not recite any further additional elements besides those recited in the independent claims, and the following additional elements recited in claims 2-5, 9-13, and 15-18 do not integrate the recited judicial exceptions into a practical application. Claim 2 additionally recites: receiving, via the data interface, a second multimodal input – This limitation amounts to a step of merely gathering data to enable further analysis, and therefore recites insignificant extra-solution activity. updating parameters of the embedder and the generator based on the first embedding loss and second embedding loss – Generic recitation of parameter updation does no more than generally link the recited abstract procedure to implementation on an ML model, and thereby does not provide integration into a practical application. Claim 3 additionally recites: wherein updating parameters of the embedder and generator is further based on the third embedding loss – Generic recitation of parameter updation does no more than generally link the recited abstract procedure to implementation on an ML model, and thereby does not provide integration into a practical application. Claim 4 additionally recites: wherein the multimodal input includes speech input, text input, and pose input – This limitation merely specifies input data as being within the realm of text/speech data, and thereby does no more than generally link the recited judicial exception to the field of use of natural language processing. Claim 5 additionally recites: wherein the embedder includes a text embedder, speech embedder, and pose embedder, wherein the text embedder [generates a text embedding]; the speech embedder [generates a speech embedding], and the pose embedder [generates a pose embedding] – These limitations do no more than invoke generic machine learning (ML) model components as tools to perform steps of an existing abstract procedure, and thereby amount to mere instructions to “apply” an exception. Generally linking the recited abstract procedure to implementation on a generic ML model does not provide integration into a practical application. Claim 9 additionally recites: wherein the multimodal input includes text input, speech input, and pose input – This limitation merely specifies input data as being within the realm of text/speech data, and thereby does no more than generally link the recited judicial exception to the field of use of natural language processing. updating parameters of the embedder, encoder, generator, and decoder based on the first loss and the second loss – Generic recitation of parameter updation does no more than generally link the recited abstract procedure to implementation on an ML model, and thereby does not provide integration into a practical application. Claim 10 additionally recites: wherein the encoder includes one or more attention layers connecting different modalities – This limitation amounts to an insignificant step of model implementation, as it does no more than ascribe a generic attention mechanism to a model component (encoder), wherein said component is merely being invoked as a tool to perform a step of the recited abstract procedure. Claim 11 additionally recites: wherein the decoder includes one or more attention layers, wherein a query associated with one or more attention layers is based on the first pose and a key and value associated with one or more attention layers are based on the multimodal features – This limitation amounts to an insignificant step of model implementation, as it does no more than ascribe a transformer architecture-type attention mechanism (query, key, value) to a model component (decoder), wherein said component is merely being invoked as a tool to perform a step of the recited abstract procedure. Claim 12 additionally recites: wherein multimodal input includes text input and speech input – This limitation merely specifies input data as being within the realm of text/speech data, and thereby does no more than generally link the recited judicial exception to the field of use of natural language processing. Claim 13 additionally recites: wherein the embedder is a neural network comprising at least one fully-connected layer – This limitation does no more than generally invoke a neural network architecture as a tool to perform steps of the recited abstract procedure, and thereby amounts to mere instructions to “apply” an exception. Claims 15-18 recite substantially similar additional elements to those found in claims 2-5, and thereby also do not integrate the recited judicial exceptions into a practical application. Step 2B: The additional elements recited in claims 2-5, 9-13, and 15-18, viewed individually or as an ordered combination, do not provide an inventive concept or otherwise amount to significantly more than the recited abstract ideas themselves. Claim 2 additionally recites: receiving, via the data interface, a second multimodal input – Receiving data via network components is well-understood, routine, and conventional activity (see MPEP § 2106.05(d); “Storing and retrieving information in memory”, “Receiving or transmitting data over a network”) and thereby does not provide an inventive concept or significantly more to the recited abstract idea. updating parameters of the embedder and the generator based on the first embedding loss and second embedding loss – Generic recitation of parameter updation does no more than generally link the recited abstract procedure to implementation on an ML model, and thereby does not provide integration into a practical application. Claim 3 additionally recites: wherein updating parameters of the embedder and generator is further based on the third embedding loss – Generally linking the recited abstract procedure to implementation on an ML model via mere recitation of parameter updation does not provide an inventive concept or significantly more to the recited abstract idea. Claim 4 additionally recites: wherein the multimodal input includes speech input, text input, and pose input – Generally linking the recited judicial exception to the field of use of natural language processing does not provide an inventive concept or significantly more to the recited abstract idea. Claim 5 additionally recites: wherein the embedder includes a text embedder, speech embedder, and pose embedder, wherein the text embedder [generates a text embedding]; the speech embedder [generates a speech embedding], and the pose embedder [generates a pose embedding] – Mere instructions to “apply” an exception on generic machine learning (ML) model components do not provide an inventive concept or significantly more to the recited abstract idea. Claim 9 additionally recites: wherein the multimodal input includes text input, speech input, and pose input – Generally linking the recited judicial exception to the field of use of natural language processing does not provide an inventive concept or significantly more to the recited abstract idea. updating parameters of the embedder, encoder, generator, and decoder based on the first loss and the second loss – Generally linking the recited abstract procedure to implementation on an ML model via mere recitation of parameter updation does not provide an inventive concept or significantly more to the recited abstract idea. Claim 10 additionally recites: wherein the encoder includes one or more attention layers connecting different modalities – Utilization of attention mechanisms in transformer models to process input features is well-understood, routine, and conventional activity (see Cristina, “The Transformer Attention Mechanism”, [pages 1-2]) and therefore does not provide an inventive concept or significantly more to the recited abstract idea. Claim 11 additionally recites: wherein the decoder includes one or more attention layers, wherein a query associated with one or more attention layers is based on the first pose and a key and value associated with one or more attention layers are based on the multimodal features – Utilization of attention mechanisms in transformer models to process input features is well-understood, routine, and conventional activity (see Cristina, “The Transformer Attention Mechanism”, [pages 1-2]) and therefore does not provide an inventive concept or significantly more to the recited abstract idea. Claim 12 additionally recites: wherein multimodal input includes text input and speech input – Generally linking the recited judicial exception to the field of use of natural language processing does not provide an inventive concept or significantly more to the recited abstract idea. Claim 13 additionally recites: wherein the embedder is a neural network comprising at least one fully-connected layer – Mere instructions to “apply” an exception on a neural network architecture do not provide an inventive concept or significantly more to the recited abstract idea. Claims 15-18 recite substantially similar additional elements to those found in claims 2-5, and thereby also do not integrate the recited judicial exceptions into a practical application. Even when considered as an ordered combination, the additional elements recited in the claims ultimately do no more than place the claims in the context of being generally linked to natural language processing and mere implementation of an abstract procedure on machine learning model components and/or transformer model components, rather than being directed towards improvement of the operation of a machine learning model itself. As such, claims 2-5, 9-13, and 15-18 also are not patent eligible. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 4-7, 14, and 17-20 are rejected under 35 U.S.C. 103 as being unpatentable over Yoon et al. (“Speech Gesture Generation from the Trimodal Context of Text, Audio, and Speaker Identity”, available arXiv 09/04/2020), hereinafter Yoon, in view of Shi et al. (“Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Predction”, available arXiv 3 Mar 2022), hereinafter Shi. The examiner notes that Yoon was cited in the IDS filed 07/12/2024. Regarding claim 1, Yoon teaches A method for training a co-speech gesture generation model (“In this paper, we present an automatic gesture generation model that uses the multimodal context of speech text, audio, and speaker identity to reliably generate gestures. By incorporating a multimodal context and an adversarial training scheme, the proposed model outputs gestures that are humanlike and that match with speech content and rhythm” [Yoon Abstract]), the method comprising: receiving, via a data interface, a multimodal input; (“The gesture generation model is trained on the TED gesture dataset (Yoon et al. 2019), which is a large-scale, English-language dataset for data-driven gesture generation research. The dataset includes speech from various speakers, so it is suitable for learning individual gesture styles. We added 471 additional TED videos to the data of (Yoon et al. 2019), for a total of 1,766 videos. Extracted human poses from TED videos, speech audio, and transcribed English speech text are available…The dataset was divided into training, validation, and test sets. Thee division was done at the video level. We used the training set for training the model, the validation set for tuning the systems, and the test set for qualitative results and human evaluation. The final number of 34-frame sequences in each data partition were 199,384; 26,795; and 25,930.” [Yoon pages 5-6 TED Gesture Dataset]) generating, via an encoder, multimodal features based on multimodal input (“We propose a neural network architecture consisting of three encoders for input speech modalities and a decoder for gesture generation. Figure 2 shows the overall architecture. Three modalities—text, audio, and speaker identity (ID)—are encoded with different encoder networks and transferred to the gesture generator” [Yoon page 4 Overall Architecture]; see Fig. 2 including Speaker ID, Speech Audio, and Speech Text converted to Feature Vectors – “The architecture of the proposed gesture generation model. The generator generates a sequence of human poses from a sequence of context feature vectors that contain the encoded features of speech text, speech audio, and speaker identity (ID). The features of text, audio, and speaker ID are depicted as red, blue, and green arrows, respectively” [Yoon page 4]) generating, via a generator, multimodal output based on the multimodal features; (“A gesture is represented as a sequence of human poses, and the generator, which is a recurrent neural network, generates poses frame-by-frame from an input sequence of feature vectors containing encoded speech context” [Yoon page 4 Overall Architecture]; see Fig. 2 including Generator – “The generator generates a sequence of human poses from a sequence of context feature vectors that contain the encoded features of speech text, speech audio, and speaker identity (ID)” [Yoon page 4]) computing a loss based on the multimodal input and the multimodal output; and updating parameters of the encoder based on the loss (“The model is trained using the losses below. We use LG to train the encoders and gesture generator and LD to train the discriminator PNG media_image1.png 165 720 media_image1.png Greyscale where where t is the length of the gesture sequence, di represents the ith pose, represented as directional vectors, in a training sample. When training the encoder and gesture generator, we minimized the difference between human poses d in the training examples and the corresponding generated poses dˆ using the Huber loss (Huber 1964). This loss LHuber G can be interpreted as a once-diffeerentiable combination of the L1 and L2 losses, and is therefore sometimes called the smooth L1 loss.” [Yoon page 6 Training Loss Function]) However, Yoon does not expressly teach masking a subset of the multimodal input and generating, via an embedder prior to an encoder, a multimodal embedding based on the masked multimodal input, and wherein the encoder includes one or more attention layers connecting different modalities. In the same field of endeavor, Shi teaches a means of multimodal speech processing (“Video recordings of speech contain correlated audio and visual information, providing a strong signal for speech representation learning from the speaker’s lip movements and the produced sound. We introduce Audio-Visual Hidden Unit BERT (AV-HuBERT), a self-supervised representation learning framework for audio-visual speech, which masks multi-stream video input and predicts automatically discovered and iteratively refined multimodal hidden units” [Shi Abstract]) that mask[s] a subset of the multimodal input and generat[es], via an embedder prior to an encoder, a multimodal embedding based on the masked multimodal input (“Our research builds on Audio HuBERT (Hsu et al., 2021a) which is a self-supervised learning framework for speech and audio. It alternates between two steps: feature clustering and masked prediction… Using (A1:T ; za 1:T ) pairs, the second step learns new feature representations by minimizing a masked prediction loss, similar to masked language modeling in BERT (Devlin et al., 2019)” [Shi page 3 Preliminary: AudioHuBERT]; “To perform the masked prediction task, the model first encodes PNG media_image2.png 27 41 media_image2.png Greyscale using a ResNet into an intermediate visual feature sequence PNG media_image3.png 27 43 media_image3.png Greyscale , which is then corrupted into PNG media_image4.png 31 47 media_image4.png Greyscale via a binary mask M. Specifically, PNG media_image5.png 37 117 media_image5.png Greyscale is replaced with a learned masked embedding. We adopt the same strategy in HuBERT to generate span masks. The masked visual features PNG media_image6.png 34 39 media_image6.png Greyscale are encoded into a sequence of contextualized features e1:T via a transformer encoder followed by a linear projection layer. The loss is computed over the masked regions and optionally over unmasked ones (when α >= 0):” [Shi page 3 Single-Modal & Cross-Modal Visual HuBERT]; “Our primary model in this work is Audio-Visual HuBERT (AV-HuBERT), shown in figure 1, which is trained iteratively by alternating between feature clustering and masked prediction in a similar way to the Visual HuBERT” [Shi page 4 Audio-Visual HuBERT]), and wherein the encoder includes one or more attention layers connecting different modalities (“The AV-HuBERT model consumes both acoustic and image frames for the masked prediction training, which enables better modeling and distillation of the correlations between the two modalities. Specifically, image sequences and acoustic features pass through their light-weight modality-specific encoders to produce intermediate features, which are then fused and fed into a shared backbone transformer encoder to predict masked cluster assignments” [Shi page 4 Audio-visual input]; “We consider two model configurations: BASE with 12 transformer blocks and LARGE with 24 transformer blocks. For BASE and LARGE, the embedding dimension/feed-forward dimension/attention heads in each transformer block are 768/3072/12 and 1024/4096/16 respectively” [Shi page 6 Setup]) It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated masking a subset of the multimodal input and generating, via an embedder prior to an encoder, a multimodal embedding based on the masked multimodal input, and and wherein the encoder includes one or more attention layers connecting different modalities as taught by Shi into Yoon because they are both directed towards multimodal speech processing. Incorporating the masked prediction task taught by Shi would enable the gesture generation model of Yoon to better learn multimodal structure (“In contrast to XDC, AV-HuBERT is trained with a BERT-like masked prediction loss, which forces the model to learn the structure within the multimodal input and was shown in Hsu et al. (2021c) to be more resilient to bad cluster assignments compared to unmasked cluster prediction” [Shi page 2 Related Work]) and would prevent over-reliance on any single input modality during sequence generation (“To prevent the model’s over-reliance on the audio stream in our joint model, we only use a linear layer to encode acoustic input to force the audio encoder to learn simple features. Additionally, before fusing audio and visual inputs into the backbone transformer encoder, dropout is applied to mask the full features of one modality; we refer to it as modality dropout… Note that modality drop out is applied at the sequence level instead of at the frame-level, which effectively tasks AV-HuBERT to perform masked prediction with visual-only, audio-only, or audio-visual input. Modality dropout prevents the model from ignoring video input and encourages the model to produce the prediction regardless of what modalities are used as input.” [Shi page 4 Modality dropout]). Regarding claim 4, the combination of Yoon and Shi teaches the limitations of parent claim 1, and Yoon further teaches wherein the multimodal input includes speech input, text input, and pose input (see Fig. 2 – Speech Text, Speech Audio, Speaker ID , and Seed Pose are input to generator [Yoon page 4]) Regarding claim 5, the combination of Yoon and Shi teaches the limitations of parent claim 4, and Shi further teaches wherein the embedder includes a text embedder, speech embedder, and pose embedder, wherein the text embedder generates a text embedding in the multimodal embedding based on the text input, the speech embedder generates a speech embedding in the multimodal embedding based on the speech input, and the pose embedder generates a pose embedding in the multimodal embedding based on the pose input (“To this end, we train an audio encoder based on the aligned audio frame sequence A1:T in parallel to the visual encoder. The iterative training alternates between the two encoders. In each iteration, an audio encoder Ea is utilized to generate target cluster assignments za 1:T . The visual encoder Ev is trained subsequently with (I1:T ; za 1:T ). The za 1:T is also used to train the next iteration of the audio encoder Ea for refinement” [Shi page 3 Cross-modal Visual HuBERT]; “The AV-HuBERT model consumes both acoustic and image frames for the masked prediction training, which enables better modeling and distillation of the correlations between the two modalities. Specifically, image sequences and acoustic features pass through their light-weight modality-specific encoders to produce intermediate features, which are then fused and fed into a shared backbone transformer encoder to predict masked cluster assignments” [Shi page 4 Audio-visual input]; Shi expressly teaches utilization of specific encoders (i.e., embedders) for each modality prior to multimodal fusion). Regarding claim 6, the combination of Yoon and Shi teaches the limitations of parent claim 4, and Shi further teaches wherein masking the multimodal input comprises at least one of: removing all text input based on a first probability; (“Additionally, before fusing audio and visual inputs into the backbone transformer encoder, dropout is applied to mask the full features of one modality; we refer to it as modality dropout. With a probability pm, both modalities are used as input. When only one modality is used, the audio stream is selected with a probability of pa” [Shi page 4 Modality dropout]). removing a first subset of the text input based on a second probability; (“The audio and visual segments are masked independently using two different masking probabilities ma and mv. We hypothesize that the difficulty of the masked prediction task differs for each modality: inferring the masked targets given the audio stream is more straightforward than using the lip movement stream” [Shi page 5 Masking by substitution]) replacing a second subset of the text input with random words based on a third probability; (“We propose a novel masking strategy for AV-HuBERT that masks segments in the visual stream by substituting them with random segments from the same video” [Shi page 5 Masking by substitution]). The correspond alternative limitations for speech and pose modalities would likewise be taught by Shi via substantially similar reasons to those detailed above (when considered in combination with the text, speech, and pose input taught by Yoon). Regarding claim 7, the combination of Yoon and Shi teaches the limitations of parent claim 1, and Shi further teaches wherein the generating a multimodal embedding includes first processing an intermediate representation based on the multimodal input (“The AV-HuBERT model consumes both acoustic and image frames for the masked prediction training, which enables better modeling and distillation of the correlations between the two modalities. Specifically, image sequences and acoustic features pass through their light-weight modality-specific encoders to produce intermediate features, which are then fused and fed into a shared backbone transformer encoder to predict masked cluster assignments” [Shi page 4 Audio-visual input]) Regarding claims 14 and 17-20, they are system/apparatus claims that correspond to the method of claims 1 and 4-7, which are already taught by the combination of Yoon and Shi as detailed above. Yoon further teaches A system for training a co-speech gesture generation model, the system comprising: a memory that stores a plurality of processor-executable instructions; one or more processors that read and execute the plurality of processor-executable instructions from the memory to perform the claimed operations (“The model was trained for 100 epochs. An Adam optimizer with β1 = 0:5 and β2 = 0:999 was used, and the learning rate was 0.0005….The trained encoders and generator are used at the synthesis stage. As the model is lightweight enough, the synthesis can be done in real time. A single synthesis generating 30 poses takes 10 ms on a GPU (NVIDIA RTX 2080 Ti) and 80 ms on a CPU (Intel i7-5930K)” [Yoon page 6 Training Loss Function]). Consequently, claims 14 and 17-20 are rejected for the same reasons as claims 1 and 4-7. Claims 2-3 and 15-16 are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Yoon and Shi, as applied to claim 1 above, further in view of Shvetsova et al. (“Everything at Once – Multi-modal Fusion Transformer for Video Retrieval”, available conference 2022), hereinafter Shvetsova. Regarding claim 2, the combination of Yoon and Shi further teaches the limitations of parent claim 1, and Yoon further teaches receiving, via the data interface, a second multimodal input; (“The gesture generation model is trained on the TED gesture dataset (Yoon et al. 2019), which is a large-scale, English-language dataset for data-driven gesture generation research. The dataset includes speech from various speakers, so it is suitable for learning individual gesture styles. We added 471 additional TED videos to the data of (Yoon et al. 2019), for a total of 1,766 videos. Extracted human poses from TED videos, speech audio, and transcribed English speech text are available…The dataset was divided into training, validation, and test sets. Thee division was done at the video level. We used the training set for training the model, the validation set for tuning the systems, and the test set for qualitative results and human evaluation. The final number of 34-frame sequences in each data partition were 199,384; 26,795; and 25,930.” [Yoon pages 5-6 TED Gesture Dataset]). Shi further teaches generating, via the embedder, a second multimodal embedding based on the second multimodal input (“To perform the masked prediction task, the model first encodes PNG media_image2.png 27 41 media_image2.png Greyscale using a ResNet into an intermediate visual feature sequence PNG media_image3.png 27 43 media_image3.png Greyscale , which is then corrupted into PNG media_image4.png 31 47 media_image4.png Greyscale via a binary mask M. Specifically, PNG media_image5.png 37 117 media_image5.png Greyscale is replaced with a learned masked embedding” [Shi page 3 Single-Modal & Cross-Modal Visual HuBERT], and Yoon further teaches wherein the second multimodal embedding includes a speech embedding, text embedding, and a pose embedding; (see Fig. 2 – Speech Text, Speech Audio, Speaker ID , and Seed Pose are input to generator [Yoon page 4]) and generating, via the generator, a second multimodal output based on the second multimodal embedding; (“A gesture is represented as a sequence of human poses, and the generator, which is a recurrent neural network, generates poses frame-by-frame from an input sequence of feature vectors containing encoded speech context” [Yoon page 4 Overall Architecture]; see Fig. 2 including Generator – “The generator generates a sequence of human poses from a sequence of context feature vectors that contain the encoded features of speech text, speech audio, and speaker identity (ID)” [Yoon page 4]) However, the combination of Yoon and Shi does not expressly teach wherein the embedder and the generator are trained by: computing a first embedding loss based on the text embedding and the speech embedding; computing a second embedding loss based on the text embedding and the pose embedding; and updating parameters of the embedder and the generator based on the first embedding loss and second embedding loss. In the same field of endeavor, Shvetsova teaches a means of multimodal feature representation and cross-modal learning (“In this work, we present a multi-modal, modality agnostic fusion transformer that learns to exchange information between multiple modalities, such as video, audio, and text, and integrate them into a fused representation in a joined multi-modal embedding space” [Shvetsova Abstract]) wherein the embedder and the generator are trained (“To train the model, we propose a combinatorial loss function which considers contrastive loss between all possible and available input combinations. For example, in the case of vision, text, and audio, the loss is based on each modality embedding alone as well as based on pairwise vision-text, audio-text, and text-audio combinations as shown in Figure 1. The resulting model is thus able to fuse any number of input modalities at test time” [Shvetsova page 2 Introduction]; “We train the system with a combinatorial input. Namely, we apply it to joint sets of input tokens from all possible combinations of modalities… In this way, we can obtain a fused representation from multiple modalities: the combination (t, v) will result in a fused representation of text and video modalities denoted as tv, resp. for va - video and audio, and ta - text and audio” [Shevtsova page 4 Multi-modal Fusion Transformer]) by: computing a first embedding loss based on the text embedding and the speech embedding; (“Unlike other methods [1, 2, 12, 44] that learn how to bring modalities together by training with three pairwise single-modality contrastive losses, Lt v between (t, v), Lt a between (t, a), and Lv a between (v, a), we force tokens to exchange information between modalities while enabling additional contrastive losses: Lt va between (t, va), Lv ta between (v, ta), and La tv between (a, tv), and introduce our combinatorial loss:… PNG media_image7.png 93 544 media_image7.png Greyscale ” [Shvetsova page 5 Combinatorial Loss]) computing a second embedding loss based on the text embedding and the pose embedding; (see Lt_v in combinatorial loss function as detailed in [Shvetsova page 5 Combinatorial Loss] above) and updating parameters of the embedder and the generator based on the first embedding loss and second embedding loss. (“By combining both aspects, the processing of all possible modality combinations and the training of the system with the proposed combinatorial loss, we obtain a multi-modal fusion transformer that learns how to attend tokens from one modality to the tokens from all other modalities” [Shvetsova page 5 Combinatorial Loss]) It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated wherein the embedder and the generator are trained by: computing a first embedding loss based on the text embedding and the speech embedding; computing a second embedding loss based on the text embedding and the pose embedding; and updating parameters of the embedder and the generator based on the first embedding loss and second embedding loss as taught by Shvetsova into Yoon and Shi because they are likewise directed towards means of multimodal feature representation. Given that Shi already expressly teaches the implementation of a fused multimodal embedding (“Specifically, image sequences and acoustic features pass through their light-weight modality-specific encoders to produce intermediate features, which are then fused and fed into a shared backbone transformer encoder to predict masked cluster assignments” [Shi page 4 Audio-visual input]), a person of ordinary skill in the art would recognize the value of incorporating the combinatorial loss function taught by Shvetsova to further optimize cross-modal attention and feature alignment prior to deep fusion (“By combining both aspects, the processing of all possible modality combinations and the training of the system with the proposed combinatorial loss, we obtain a multi-modal fusion transformer that learns how to attend tokens from one modality to the tokens from all other modalities” [Shvetsova page 5 Combinatorial Loss]) Regarding claim 3, the combination of Yoon, Shi, and Shvetsova teaches the limitations of parent claim 2, and Yoon further teaches computing a third embedding loss based on the second multimodal input and the second multimodal output, and wherein updating parameters of the embedder and generator is further based on the third embedding loss (“The model is trained using the losses below. We use LG to train the encoders and gesture generator and LD to train the discriminator PNG media_image1.png 165 720 media_image1.png Greyscale where where t is the length of the gesture sequence, di represents the ith pose, represented as directional vectors, in a training sample. When training the encoder and gesture generator, we minimized the difference between human poses d in the training examples and the corresponding generated poses dˆ using the Huber loss (Huber 1964). This loss LHuber G can be interpreted as a once-differentiable combination of the L1 and L2 losses, and is therefore sometimes called the smooth L1 loss.” [Yoon page 6 Training Loss Function]). Regarding claims 15-16, they are system/apparatus claims that correspond to the method of claims 2-3 which are already taught by the combination of Yoon, Shi, and Shvetsova as detailed above. Consequently, claims 15-16 are rejected for the same reasons as claims 2-3. Claims 8, 10, and 12-13 are rejected under 35 U.S.C. 103 as being unpatentable over Yoon et al. (“Speech Gesture Generation from the Trimodal Context of Text, Audio, and Speaker Identity”, available arXiv 09/04/2020, cited in IDS filed 07/12/2024), hereinafter Yoon, in view of Shi et al. (“Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Predction”, available arXiv 3 Mar 2022), hereinafter Shi, and Athanasiou et al. (“TEACH: Temporal Action Composition for 3D Humans”, available conference 2022), hereinafter Athanasiou. Regarding claim 8, Yoon teaches A method for co-speech gesture generation (“In this paper, we present an automatic gesture generation model that uses the multimodal context of speech text, audio, and speaker identity to reliably generate gestures. By incorporating a multimodal context and an adversarial training scheme, the proposed model outputs gestures that are humanlike and that match with speech content and rhythm” [Yoon Abstract]), the method comprising: receiving, via a data interface, multimodal input; (“The gesture generation model is trained on the TED gesture dataset (Yoon et al. 2019), which is a large-scale, English-language dataset for data-driven gesture generation research. The dataset includes speech from various speakers, so it is suitable for learning individual gesture styles. We added 471 additional TED videos to the data of (Yoon et al. 2019), for a total of 1,766 videos. Extracted human poses from TED videos, speech audio, and transcribed English speech text are available…The dataset was divided into training, validation, and test sets. Thee division was done at the video level. We used the training set for training the model, the validation set for tuning the systems, and the test set for qualitative results and human evaluation. The final number of 34-frame sequences in each data partition were 199,384; 26,795; and 25,930.” [Yoon pages 5-6 TED Gesture Dataset]) and generating, via the generator, a first pose based on first multimodal input and a second pose based on second multimodal input (“A gesture is represented as a sequence of human poses, and the generator, which is a recurrent neural network, generates poses frame-by-frame from an input sequence of feature vectors containing encoded speech context” [Yoon page 4 Overall Architecture]; see Fig. 2 including Generator – “The generator generates a sequence of human poses from a sequence of context feature vectors that contain the encoded features of speech text, speech audio, and speaker identity (ID)” [Yoon page 4]) However, Yoon does not expressly teach generating, via an embedder prior to an encoder, an embedding based on the multimodal output and generating, via an encoder, multimodal features based on the embedding. In the same field of endeavor, Shi teaches a means of multimodal speech processing (“Video recordings of speech contain correlated audio and visual information, providing a strong signal for speech representation learning from the speaker’s lip movements and the produced sound. We introduce Audio-Visual Hidden Unit BERT (AV-HuBERT), a self-supervised representation learning framework for audio-visual speech, which masks multi-stream video input and predicts automatically discovered and iteratively refined multimodal hidden units” [Shi Abstract]) that generat[es], via an embedder prior to an encoder, a multimodal embedding based on the multimodal input ( “To perform the masked prediction task, the model first encodes PNG media_image2.png 27 41 media_image2.png Greyscale using a ResNet into an intermediate visual feature sequence PNG media_image3.png 27 43 media_image3.png Greyscale , which is then corrupted into PNG media_image4.png 31 47 media_image4.png Greyscale via a binary mask M. Specifically, PNG media_image5.png 37 117 media_image5.png Greyscale is replaced with a learned masked embedding. We adopt the same strategy in HuBERT to generate span masks. The masked visual features PNG media_image6.png 34 39 media_image6.png Greyscale are encoded into a sequence of contextualized features e1:T via a transformer encoder followed by a linear projection layer. The loss is computed over the masked regions and optionally over unmasked ones (when α >= 0):” [Shi page 3 Single-Modal & Cross-Modal Visual HuBERT]; “Our primary model in this work is Audio-Visual HuBERT (AV-HuBERT), shown in figure 1, which is trained iteratively by alternating between feature clustering and masked prediction in a similar way to the Visual HuBERT” [Shi page 4 Audio-Visual HuBERT]), and generating, via an encoder, multimodal features based on the embedding. (“The AV-HuBERT model consumes both acoustic and image frames for the masked prediction training, which enables better modeling and distillation of the correlations between the two modalities. Specifically, image sequences and acoustic features pass through their light-weight modality-specific encoders to produce intermediate features, which are then fused and fed into a shared backbone transformer encoder to predict masked cluster assignments” [Shi page 4 Audio-visual input]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated generating, via an embedder prior to an encoder, an embedding based on the multimodal output and generating, via an encoder, multimodal features based on the embedding as taught by Shi into Yoon because they are both directed towards multimodal speech processing. Incorporating the masked prediction task taught by Shi would enable the gesture generation model of Yoon to better learn multimodal structure (“In contrast to XDC, AV-HuBERT is trained with a BERT-like masked prediction loss, which forces the model to learn the structure within the multimodal input and was shown in Hsu et al. (2021c) to be more resilient to bad cluster assignments compared to unmasked cluster prediction” [Shi page 2 Related Work]) and would prevent over-reliance on any single input modality during sequence generation (“To prevent the model’s over-reliance on the audio stream in our joint model, we only use a linear layer to encode acoustic input to force the audio encoder to learn simple features. Additionally, before fusing audio and visual inputs into the backbone transformer encoder, dropout is applied to mask the full features of one modality; we refer to it as modality dropout… Note that modality drop out is applied at the sequence level instead of at the frame-level, which effectively tasks AV-HuBERT to perform masked prediction with visual-only, audio-only, or audio-visual input. Modality dropout prevents the model from ignoring video input and encourages the model to produce the prediction regardless of what modalities are used as input.” [Shi page 4 Modality dropout]). However, the combination of Yoon and Shi does not expressly teach generating, via a decoder, first motion features based on the multimodal features; and generating, via the decoder, second motion features based on the first motion features and the multimodal features. In the same field of endeavor, Athanasiou teaches a means of gesture generation and temporal motion decoding (“Given a series of natural language descriptions, our task is to generate 3D human motions that correspond semantically to the text, and follow the temporal order of the instructions. In particular, our goal is to enable the synthesis of a series of actions, which we refer to as temporal action composition… Our approach, called TEACH for “TEmporal Action Compositions for Human motions”, produces realistic human motions for a wide variety of actions and temporal compositions from language descriptions” [Athanasiou Abstract]) that generat[es], via a decoder, first motion features based on the encoder-outputted features; (“Motion decoder. We use the same decoder architecture as in TEMOS [31], which generates a sequence of poses from a single embedding. This Transformer-based motion decoder takes the current latent vector zi and Fi positional encodings (in the form of sinusoidal functions) as input, and generates the sequence of human motions” [Athanasiou page 4 Motion decoder]; see Motion Decoder Mdec in Figure 2 which outputs H^i 1…Fi motion features) and generat[es], via the decoder, second motion features based on the first motion features and the encoder-outputted features (“Moreover, the last P frames of the previous generated motion, _ Hi−1 Fi−1−P:Fi−1 , are encoded into motion features IFi−1−P:Fi−1 (Past Encoder). Then, we combine the features from the previous action, Ii−1 Fi−1−j, j ∈ N, and the current text features along with learnable tokens (μtoken, Σtoken and SEP), and pass them as inputs to the Past-conditioned Text-Encoder, which generates the distribution parameters μi and Σi. μi and Σi are treated as parameters of a Gaussian distribution, from which we sample and decode the final motion” [Athanasiou page 4 Architecture]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated generating, via a decoder, first motion features based on the multimodal features; and generating, via the decoder, second motion features based on the first motion features and the multimodal features as taught by Athanasiou into Yoon and Shi because both Yoon and Athanasiou are directed towards gesture generation and temporal motion decoding. Incorporating the past-conditioned motion decoder taught by Athanasiou into the gesture generation model would improve continuity between generated pose frames by providing necessary temporal context to ensure natural transitions between consecutive outputs (“One of the key challenges in synthesizing long action sequences given a stream of textual prompts is how to ensure continuity within the transitions between actions. Independently generating one motion per action would not guarantee temporal smoothness. In our framework, we find that encoding the next action conditioned on the last few frames of the previous action is a simple and effective solution” [Athanasiou page 2 Introduction]). Regarding claim 10, the combination of Yoon, Shi, and Athanasiou teaches the limitations of parent claim 8, and Shi further teaches wherein the encoder includes one or more attention layers connecting different modalities (The AV-HuBERT model consumes both acoustic and image frames for the masked prediction training, which enables better modeling and distillation of the correlations between the two modalities. Specifically, image sequences and acoustic features pass through their light-weight modality-specific encoders to produce intermediate features, which are then fused and fed into a shared backbone transformer encoder to predict masked cluster assignments” [Shi page 4 Audio-visual input]; “We consider two model configurations: BASE with 12 transformer blocks and LARGE with 24 transformer blocks. For BASE and LARGE, the embedding dimension/feed-forward dimension/attention heads in each transformer block are 768/3072/12 and 1024/4096/16 respectively” [Shi page 6 Setup]). Regarding claim 12, the combination of Yoon, Shi, and Athanasiou teaches the limitations of parent claim 8, and Yoon further teaches wherein multimodal input includes text input and speech input (see Fig. 2 – Speech Text, Speech Audio, Speaker ID , and Seed Pose are input to generator [Yoon page 4]). Regarding claim 13, the combination of Yoon, Shi, and Athanasiou teaches the limitations of parent claim 8, and Shi further teaches wherein the embedder is a neural network comprising at least one fully-connected layer (“The masked visual features ~fv1:T are encoded into a sequence of contextualized features e1:T via a transformer encoder followed by a linear projection layer” [Shi page 3 Single-modal Visual HuBERT]). Claims 9 and 11 are rejected under 35 U.S.C. 103 as being unpatentable over Yoon, Shi, and Athanasiou, further in view of Lu et al. (“UNIFIED-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks”, available arXiv 10/04/2022), hereinafter Lu. Regarding claim 9, the combination of Yoon, Shi, and Athanasiou teaches the limitations of parent claim 8, and Yoon further teaches wherein the multimodal input includes text input, speech input, and pose input (see Fig. 2 – Speech Text, Speech Audio, Speaker ID , and Seed Pose are input to generator [Yoon page 4]). Athanasiou further teaches computing a reconstruction loss based on the first pose, the second pose, and the pose input; (“Reconstruction loss. From the two forward passes, we generate the motions _ H1 1:F1 and _ H2 1:F2 . We enforce them to be close to the corresponding ground truth motions H1 1:F1 and H2 1:F2 via the following loss terms: PNG media_image8.png 60 529 media_image8.png Greyscale ” [Athanasiou page 4 Reconstruction loss]) and updating parameters of the overall model based on the reconstruction loss (“We do one backward pass, which optimizes the reconstruction loss and the KL loss on the two segments jointly.” [Athanasiou page 4 Data handling]). However, the combination does not expressly teach generating, via a generator, reconstructed speech based on speech features and reconstructed text based on text features; computing a first loss based on reconstructed speech, speech input, reconstructed text, and text input; and updating parameters of the embedder, encoder, generator, and decoder based on the first loss and the second loss. In the same field of endeavor, Lu teaches a means of multimodal sequence-to-sequence learning (see Figure 1 – “UNIFIED-IO is a single sequence-to-sequence model that performs a variety of tasks in computer vision and NLP using a unified architecture without a need for either task or modality-specific branches. This broad unification is achieved by homogenizing every task’s input and output into a sequence of discrete vocabulary tokens. UNIFIED-IO supports modalities as diverse as images, masks, keypoints, boxes, and text, and tasks as varied as depth estimation, inpainting, semantic segmentation, captioning, and reading comprehension” [Lu page 2]) that generat[es], via a generator, reconstructed multimodal (e.g., text, image) data based on multimodal (e.g., text, image) features (“UNIFIED-IO is trained in two stages – A pre-training stage that uses unsupervised losses from text, image, and paired image-text data, and a massive multi-task stage where the model is jointly trained on a large variety of tasks…Pre-training. To learn good representations from large-scale webly supervised image and text data, we consider two pre-training tasks: text span denoising and masked image denoising. The text span denoising task follows Raffel et al. (2020) – randomly corrupt 15% of the tokens and replace the consecutive corrupted tokens with a unique mask token. The masked image denoising task follows Bao et al. (2022) and He et al. (2022) – randomly masked 75% of the image patches, and the goal is to recover the whole image. When another modality is present, i.e. image or text, the model can use information from that modality to complete the tasks” [Lu page 6 Training]) comput[es] a first loss based on reconstructed speech, speech input, reconstructed text, and text input; ([Lu page 6 Training] as detailed above; Lu discloses a pretraining stage that calculates reconstruction losses particular to each specific modality (e.g., text, image – could be configured to speech)) and updat[es] parameters of the overall model based on the first loss ([Lu page 6 Training] as detailed above). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated generating, via a generator, reconstructed speech based on speech features and reconstructed text based on text features; computing a first loss based on reconstructed speech, speech input, reconstructed text, and text input; and updating parameters of the embedder, encoder, generator, and decoder based on the first loss and the second loss as taught by Lu into Yoon, Shi, and Athanasiou because both Yoon and Lu are directed towards multimodal sequence-to-sequence learning. Incorporating the pre-training objective of Lu (adapted to the particular input modalities of the given combination) would expand on the BERT-inspired masked pre-training objectives of Shi and further preserve model generalizability (“Vision and language pre-training has become standard practice for multi-modal models, including unified and non-unified models requiring task-specific heads to train from scratch during finetuning. Many initial pre-training strategies were inspired by BERT (Devlin et al., 2019) and included masked-language-modeling, image-text-matching, or mask-region-modeling objectives, often supplemented with objectives using the predictions of a strong object detector model (e.g, VILBERT (Lu et al., 2019), LXMERT (Tan & Bansal, 2019), VisualBERT (Li et al., 2019))…. The generalized masked-data-modeling pretraining objective used in UNIFIED-IO is similar to ones used in several recent works (Wang et al., 2022c; Peng et al., 2022; Singh et al., 2022)” [Lu page 11 Related Work]). Regarding claim 11, the combination of Yoon, Shi, Athanasiou, and Liu teaches the limitations of parent claim 9, and Athanasiou further teaches wherein the decoder includes one or more attention layers, (“This Transformer-based motion decoder takes the current latent vector zi and Fi positional encodings (in the form of sinusoidal functions) as input, and generates the sequence of human motions.” [Athanasiou page 4 Motion decoder]; A transformer-based decoder inherently consists of attention layers) wherein a query associated with one or more attention layers is based on the first pose (“Moreover, the last P frames of the previous generated motion, PNG media_image9.png 44 183 media_image9.png Greyscale are encoded into motion features PNG media_image10.png 39 165 media_image10.png Greyscale (Past Encoder)” [Athanasiou page 3 Architecture]) and a key and value associated with one or more attention layers are based on the multimodal features (“Then, we combine the features from the previous action, Ii−1 Fi−1−j, j ∈ N, and the current text features along with learnable tokens (μtoken, Σtoken and SEP), and pass them as inputs to the Past-conditioned Text-Encoder” [Athanasiou page 3 Architecture]). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Cristina (“The Transformer Attention Mechanism”, available online 6 Jan 2023) discloses a review of transformer models as commonly implemented in the art. Any inquiry concerning this communication or earlier communications from the examiner should be directed to VIJAY M BALAKRISHNAN whose telephone number is (571) 272-0455. The examiner can normally be reached 10am-5pm EST Mon-Thurs. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, JENNIFER WELCH can be reached on (571) 272-7212. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /V.M.B./ Examiner, Art Unit 2143 /JENNIFER N WELCH/Supervisory Patent Examiner, Art Unit 2143
Read full office action

Prosecution Timeline

Apr 04, 2024
Application Filed
Aug 03, 2026
Non-Final Rejection mailed — §101, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743623
INFORMATION PROCESSING DEVICE AND MACHINE LEARNING METHOD THAT OPTIMIZE A DECODING PROCESS USING BACK-PROPAGATION
4y 0m to grant Granted Sep 22, 2026
Patent 12731026
METHOD AND SYSTEM FOR PROGRAM SAMPLING USING NEURAL NETWORK
3y 10m to grant Granted Sep 08, 2026
Patent 12711407
REASONING METHOD BASED ON STRUCTURAL ATTENTION MECHANISM FOR KNOWLEDGE-BASED QUESTION ANSWERING AND COMPUTING APPARATUS FOR PERFORMING THE SAME
3y 8m to grant Granted Aug 18, 2026
Patent 12645933
Method and System for Training a Neural Network for Generating Universal Adversarial Perturbations
4y 7m to grant Granted Jun 02, 2026
Patent 12619871
INTERPRETABLE NEURAL NETWORK ARCHITECTURE USING CONTINUED FRACTIONS
3y 11m to grant Granted May 05, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
41%
Grant Probability
99%
With Interview (+73.3%)
3y 11m (~1y 5m remaining)
Median Time to Grant
Low
PTA Risk
Based on 27 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month