DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 18 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor, or for pre-AIA the applicant regards as the invention.
Dependent claim 18 depends upon independent claim 16. The claim 16 recites “receiving a runtime action sentence; converting the runtime action sentence into a set of runtime motion tokens; iteratively unmasking runtime motion tokens of the set of runtime motion tokens using a trained denoise transformer; and transforming the unmasked runtime motion tokens into a runtime skeletal representation of the human motion using a trained vector quantized variational autoencoder (VQ-VAE) model”. However, the claim 18 recites the same above limitations “receiving a runtime action sentence; converting the runtime action sentence into a set of runtime motion tokens; iteratively unmasking runtime motion tokens of the set of runtime motion tokens using the trained denoise transformer; and transforming the unmasked runtime motion tokens into a runtime skeletal representation of the human motion using the trained VQ-VAE model”. The issue is persons of ordinary skill in the art reading the specification is not able to understand what Applicant regards as the invention. Therefore, the claim is rejected under U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-4 are rejected under 35 U.S.C. 103 as being unpatentable over LEE et al (U.S. Patent Application Publication 2025/0371782 A1) in view of Qin et al (U.S. Patent Application Publication 2026/0044993 A1).
Regarding claim 1, LEE discloses a computer-implemented method for multi-motion generation, comprising:
training a variational autoencoder (VAE) model (Paragraph [0005], Variational Auto-encoders (VAEs) ...; FIG. 1; paragraph [0058], in an example system 100, the motion decoder 122 is a convolutional decoder, and the motion encoder 126 is a convolutional encoder of the autoencoder 120, which is embodied in a vector quantization variable autoencoder (VQVAE); paragraphs [0061]-[0062], FIG. 2 shows an example training method 200 for the system 100 ... An example training method 200 first trains at 202 the autoencoder, e.g., the VQVAE 120 ...) based on a skeletal representation of human motion (Paragraph [0032], human motion is typically represented as a temporal sequence of 3D points, e.g., human meshes or skeletons ...), one or more motion tokens associated with the skeletal representation of the human motion (Paragraph [0045], in an initial training phase, the motion encoder and motion decoder can be trained to encode short human motion into learned discrete (specific) tokens), and a dynamic transition probability based on a distance between the one or more motion tokens (Paragraphs [0099]-[0105], for each motion, experiments ranked the Euclidean distance to 32 text descriptions of 1 positive and 31 negatives. The Top-1, Top-2, and Top-3 accuracy were determined ... an inference time as shown in FIG. 6 and described by example above. To train these transition latent vectors, an example method randomly substituted part of the quantized latent vectors Z into the transition latent vectors while training the VQVAE).
However, LEE does not specifically disclose training a denoise transformer by performing self-attention based on the one or more motion tokens and cross-attention based on the one or more motion tokens and an action sentence.
In additional, Qin discloses (Abstract, embodiments described herein provide a generation model comprising a video-specific variational auto-encoder (VAE) for effective compression of video pixel information with reduced spatial and temporal dimensions and a video diffusion transformer (vDiT) to generate latent representations of frames ...) training a denoise transformer (Paragraph [0033], FIG. 2A is a simplified diagram illustrating an example training framework 200 of the t2v generation model described in FIG. 1 ...; paragraph [0039], details of the training and inference process of the denoising diffusion model of vDiT 110 may be provided below in relation to FIG. 2B; paragraphs [0049]-[0052], the Transformer-like denoising network (denoising model 222)) by performing self-attention based on the one or more motion tokens (Paragraphs [0057]-[0059], FIG. 4 is a simplified diagram illustrating a data pipeline 400 for generating video-text training data for training the text-to-video generation model described in
FIG. 1, according to some embodiments. First, a long-video clipping module 402 splits long videos into manageable clips ... an aesthetic scoring module 406 analyzes aesthetics and motion dynamics across frames to eliminate static video clips and inconsistent frames; paragraph [0040], vDiT 110 may comprise a stack of spatial-temporal transformer blocks as illustrated in FIG. 2A. Each transformer module comprises one or more modulation layers to scale and/or shift a representation vector, a spatial self-attention layer 215 to capture spatial information and a temporal self-attention layer 220 to capture temporal characteristics from the encoded video, and a feed forward layer to generate an output from the Transformer block ...; paragraph [0041], In FIG. 2C, the temporal self-attention layer 220 may adopt Rotary Positional Embedding (RoPE) to encode temporal information, e.g., to compute attentions between an input matrix having a size of (B, H, W) capturing visual information of each frame in the batch ... In FIG. 2D, the spatial self-attention layer 215 may adopt sinusoidal encoding to encode spatial information, e.g., to compute attentions between an input matrix having a size of (B, T) capturing the batch size and the total number of frames in the input video, input matrix (H, W) capturing spatially distributed visual content on each frame and the input vector having a size of C representing the number of image channels of each frame) and cross-attention based on the one or more motion tokens and an action sentence (Paragraph [0115], the video diffusion model comprises a spatial attention layer (e.g., 215 in FIG. 2A), a temporal attention layer (e.g., 220 in FIG. 2A) and a text-video cross-attention layer. he spatial attention layer outputs attention weights capturing spatial information of an input vector relating to the training video. The temporal attention layer outputs attention weights capturing temporal characteristics of an input vector relating to the training video. The text-video cross-attention layer output attention weights capturing relationships between embeddings of the training text and spatial and/or temporal portions of the training video).
Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method for training a model for generating a representation of long-term motion from a text input taught by LEE incorporate the teachings of Qin, applying the training framework for a denoising diffusion model taught by Qin to add a denoise transformer into the artificial intelligence system and provide the method for training the denoise transformer. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify LEE according to the relied-upon teachings of Qin to obtain the invention as specified in claim.
Regarding claim 2, the combination of LEE in view of Qin discloses everything claimed as applied above (see claim 1), and LEE further disclose wherein the VAE model is a vector quantized variational autoencoder (VQ-VAE) model (FIG. 1; paragraph [0058], in an example system 100, the motion decoder 122 is a convolutional decoder, and the motion encoder 126 is a convolutional encoder of the autoencoder 120, which is embodied in a vector quantization variable autoencoder (VQVAE)).
Regarding claim 3, the combination of LEE in view of Qin discloses everything claimed as applied above (see claim 1), and LEE further disclose wherein the VAE model (FIG. 1; paragraph [0058], in an example system 100, the motion decoder 122 is a convolutional decoder, and the motion encoder 126 is a convolutional encoder of the autoencoder 120, which is embodied in a vector quantization variable autoencoder (VQVAE)) converts the skeletal representation of the human motion into one or more motion tokens (Paragraph [0032], human motion is typically represented as a temporal sequence of 3D points, e.g., human meshes or skeletons ...; paragraphs [0054]-[0056], the text encoder 102 is configured to receive a text input ... The text input includes a plurality of text groups, chunks, or phrases, such as but not limited to sentences. A plurality of phrases may be embodied in a paragraph or other text group. Each phrase can describe an action, such that the plurality of phrases describes at least two actions. The plurality of phrases have an arbitrary length, that is, they can have any suitable length, and such a length can be independent of the length of a corresponding motion ... The text encoder 102 is configured to predict a latent representation comprising a continuous stream of latent vectors conditioned on the text input and the duration. Each latent vector represents a fixed length of motion ...; paragraph [0066], the text encoder 102 predicts at 306 a latent vector representation conditioned on the text input and the duration(s) that is embodied in a continuous stream of latent vectors. Each latent vector represents a fixed length of motion, and comprises a vector in a discrete latent space) and reconstructs the human motion from the one or more motion tokens (FIG. 3; paragraph [0067], the stream of latent vectors is then passed through the motion decoder 122, e.g., in a trained 1D convolutional VQVAE decoder. The motion decoder 122 decodes the (e.g., quantized and concatenated) latent representation at 310 to continuously reconstruct the desired long-term motion representation; paragraph [0104], transition latent vector: To avoid the need for long-term data, example methods can rely on transition tokens, which are designed to chain sequences together at inference. Such transition tokens can be obtained by masking at training time).
Regarding claim 4, the combination of LEE in view of Qin discloses everything claimed as applied above (see claim 1).
However, LEE does not specifically disclose wherein the training the denoise transformer by performing the self-attention is based on relative positional encoding (RPE).
In additional, Qin discloses wherein the training the denoise transformer (Paragraph [0033], FIG. 2A is a simplified diagram illustrating an example training framework 200 of the t2v generation model described in FIG. 1 ...; paragraph [0039], details of the training and inference process of the denoising diffusion model of vDiT 110 may be provided below in relation to FIG. 2B; paragraphs [0049]-[0052], the Transformer-like denoising network (denoising model 222)) by performing the self-attention is based on relative positional encoding (RPE) (Paragraph [0026], the Video-diffusion Transformer (vDiT) may incorporate transformer blocks with both temporal and spatial self-attention layers to encode spatial and temporal position information from latent representations of videos).
Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method for training a model for generating a representation of long-term motion from a text input taught by LEE incorporate the teachings of Qin, applying the training framework for a denoising diffusion model taught by Qin to add a denoise transformer into the artificial intelligence system for performing the self-attention is based on relative positional encoding. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify LEE according to the relied-upon teachings of Qin to obtain the invention as specified in claim.
Claims 5-10 are rejected under 35 U.S.C. 103 as being unpatentable over LEE et al (U.S. Patent Application Publication 2025/0371782 A1) in view of Qin et al (U.S. Patent Application Publication 2026/0044993 A1) in view of Wang et al (U.S. Patent Application Publication 2025/0181974 A1).
Regarding claim 5, the combination of LEE in view of Qin discloses everything claimed as applied above (see claim 1), and LEE further disclose comprising:
receiving a runtime action sentence (FIG. 7; paragraph [0073], the computing device 700 may receive the input 740, such as an input text, from a user via the user interface; FIG. 3; paragraph [0065], the text encoder receives at 302 a text input comprising a plurality of phrases, where each phrase describes an action);
converting the runtime action sentence into a set of runtime motion tokens (FIG. 1; paragraphs [0054]-[0056], the text encoder 102 is configured to receive a text input ... The text input includes a plurality of text groups, chunks, or phrases, such as but not limited to sentences. A plurality of phrases may be embodied in a paragraph or other text group. Each phrase can describe an action, such that the plurality of phrases describes at least two actions. The plurality of phrases have an arbitrary length, that is, they can have any suitable length, and such a length can be independent of the length of a corresponding motion ... The text encoder 102 is configured to predict a latent representation comprising a continuous stream of latent vectors conditioned on the text input and the duration. Each latent vector represents a fixed length of motion ...; paragraph [0066], the text encoder 102 predicts at 306 a latent vector representation conditioned on the text input and the duration(s) that is embodied in a continuous stream of latent vectors. Each latent vector represents a fixed length of motion, and comprises a vector in a discrete latent space); and
transforming the unmasked runtime motion tokens (FIG. 3; paragraph [0067], the stream of latent vectors is then passed through the motion decoder 122, e.g., in a trained 1D convolutional VQVAE decoder. The motion decoder 122 decodes the (e.g., quantized and concatenated) latent representation at 310 to continuously reconstruct the desired long-term motion representation; paragraph [0104], transition latent vector: To avoid the need for long-term data, example methods can rely on transition tokens, which are designed to chain sequences together at inference. Such transition tokens can be obtained by masking at training time. Thus, the unmasked runtime motion tokens will be transformed) into a runtime skeletal representation of the human motion (Paragraph [0032], human motion is typically represented as a temporal sequence of 3D points, e.g., human meshes or skeletons ...) using the trained VAE model (Paragraph [0005], Variational Auto-encoders (VAEs) ...; FIG. 1; paragraph [0058], in an example system 100, the motion decoder 122 is a convolutional decoder, and the motion encoder 126 is a convolutional encoder of the autoencoder 120, which is embodied in a vector quantization variable autoencoder (VQVAE)).
However, LEE does not specifically disclose iteratively unmasking runtime motion tokens of the set of runtime motion tokens using the trained denoise transformer.
In additional, Wang discloses (Abstract, a system for controlling a character in a virtual environment includes at least one processor and at least one nontransitory computer-readable medium comprising instructions that are executable by the at least one processor. The instructions include receiving a text input indicating one or more motions of the character, generating a plurality of text-based tokens and a plurality of motion tokens, appending a plurality of masked tokens to the plurality of motion tokens, performing a masked transformer routine to generate a plurality of predicted motion tokens based on the plurality of masked tokens, the plurality of motion tokens, and the plurality of text-based tokens, decoding the plurality of predicted motion tokens and the plurality of motion tokens into a motion sequence of the character, where the motion sequence includes the one or more motions and a predicted motion of the character, and controlling the character based on the motion sequence) iteratively unmasking runtime motion tokens of the set of runtime motion tokens using the trained denoise transformer (Paragraphs [0033]-[0034], the MMM system includes a motion tokenizer (MT) module that transforms human motion into a sequence of discrete tokens (e.g., discrete units, elements, and/or character strings that represent segments of the input data) in a latent space and a conditional masked motion transformer (CMMT) module that predicts randomly masked motion tokens, which are conditioned on the discrete tokens generated by the MT module ... the MMM system is trained based on a two-stage approach. In the first stage, the MT module is trained to convert and quantize raw motion data into a sequence of discrete motion tokens in the latent space based on a motion codebook. In the second stage, motion tokens generated by the MT module are randomly masked, and the parameters of the CMMT module are iteratively and selectively modified until the CMMT can accurately predict the masked tokens based on the unmasked tokens and an input text. See claim 1, the trained denoise transformer).
Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method for training a model for generating a representation of long-term motion from a text input taught by LEE incorporate the teachings of Wang, applying the system for controlling a character in a virtual environment based on a text input taught by Wang to provide the trained transformer to iterative the unmasking motion tokens until the vector quantization variable autoencoder (VQVAE) transform the unmasked runtime motion tokens into a runtime skeletal representation of the human motion. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify LEE according to the relied-upon teachings of Wang to obtain the invention as specified in claim.
Regarding claim 6, the combination of LEE in view of Qin discloses everything claimed as applied above (see claim 1), and LEE further disclose comprising:
receiving a runtime action sentence including a first action and a second action (FIG. 7; paragraph [0073], the computing device 700 may receive the input 740, such as an input text, from a user via the user interface; FIG. 3; paragraph [0065], the text encoder receives at 302 a text input comprising a plurality of phrases, where each phrase describes an action);
converting the runtime action sentence into a first set of runtime motion tokens corresponding to the first action and a second set of runtime motion tokens corresponding to the second action (FIG. 1; paragraphs [0054]-[0056], the text encoder 102 is configured to receive a text input ... The text input includes a plurality of text groups, chunks, or phrases, such as but not limited to sentences. A plurality of phrases may be embodied in a paragraph or other text group. Each phrase can describe an action, such that the plurality of phrases describes at least two actions. The plurality of phrases have an arbitrary length, that is, they can have any suitable length, and such a length can be independent of the length of a corresponding motion ... The text encoder 102 is configured to predict a latent representation comprising a continuous stream of latent vectors conditioned on the text input and the duration. Each latent vector represents a fixed length of motion ...; paragraph [0066], the text encoder 102 predicts at 306 a latent vector representation conditioned on the text input and the duration(s) that is embodied in a continuous stream of latent vectors. Each latent vector represents a fixed length of motion, and comprises a vector in a discrete latent space);
transforming the first set of unmasked runtime motion tokens and the second set of unmasked runtime motion tokens (FIG. 3; paragraph [0067], the stream of latent vectors is then passed through the motion decoder 122, e.g., in a trained 1D convolutional VQVAE decoder. The motion decoder 122 decodes the (e.g., quantized and concatenated) latent representation at 310 to continuously reconstruct the desired long-term motion representation; paragraph [0104], transition latent vector: To avoid the need for long-term data, example methods can rely on transition tokens, which are designed to chain sequences together at inference. Such transition tokens can be obtained by masking at training time. Thus, the unmasked runtime motion tokens will be transformed) into a runtime skeletal representation of the human motion (Paragraph [0032], human motion is typically represented as a temporal sequence of 3D points, e.g., human meshes or skeletons ...) using the trained VAE model (Paragraph [0005], Variational Auto-encoders (VAEs) ...; FIG. 1; paragraph [0058], in an example system 100, the motion decoder 122 is a convolutional decoder, and the motion encoder 126 is a convolutional encoder of the autoencoder 120, which is embodied in a vector quantization variable autoencoder (VQVAE)).
However, LEE does not specifically disclose iteratively unmasking runtime motion tokens of the first set of runtime motion tokens using the trained denoise transformer to generate a first set of unmasked runtime motion tokens;
iteratively unmasking runtime motion tokens of the second set of runtime motion tokens using the trained denoise transformer to generate a second set of unmasked runtime motion tokens.
In additional, Wang discloses (Abstract, a system for controlling a character in a virtual environment includes at least one processor and at least one nontransitory computer-readable medium comprising instructions that are executable by the at least one processor. The instructions include receiving a text input indicating one or more motions of the character, generating a plurality of text-based tokens and a plurality of motion tokens, appending a plurality of masked tokens to the plurality of motion tokens, performing a masked transformer routine to generate a plurality of predicted motion tokens based on the plurality of masked tokens, the plurality of motion tokens, and the plurality of text-based tokens, decoding the plurality of predicted motion tokens and the plurality of motion tokens into a motion sequence of the character, where the motion sequence includes the one or more motions and a predicted motion of the character, and controlling the character based on the motion sequence) iteratively unmasking runtime motion tokens of the first set of runtime motion tokens using the trained denoise transformer to generate a first set of unmasked runtime motion tokens (Paragraphs [0033]-[0034], the MMM system includes a motion tokenizer (MT) module that transforms human motion into a sequence of discrete tokens (e.g., discrete units, elements, and/or character strings that represent segments of the input data) in a latent space and a conditional masked motion transformer (CMMT) module that predicts randomly masked motion tokens, which are conditioned on the discrete tokens generated by the MT module ... the MMM system is trained based on a two-stage approach. In the first stage, the MT module is trained to convert and quantize raw motion data into a sequence of discrete motion tokens in the latent space based on a motion codebook. In the second stage, motion tokens generated by the MT module are randomly masked, and the parameters of the CMMT module are iteratively and selectively modified until the CMMT can accurately predict the masked tokens based on the unmasked tokens and an input text);
iteratively unmasking runtime motion tokens of the second set of runtime motion tokens using the trained denoise transformer to generate a second set of unmasked runtime motion tokens (Paragraphs [0033]-[0034], the MMM system includes a motion tokenizer (MT) module that transforms human motion into a sequence of discrete tokens (e.g., discrete units, elements, and/or character strings that represent segments of the input data) in a latent space and a conditional masked motion transformer (CMMT) module that predicts randomly masked motion tokens, which are conditioned on the discrete tokens generated by the MT module ... the MMM system is trained based on a two-stage approach. In the first stage, the MT module is trained to convert and quantize raw motion data into a sequence of discrete motion tokens in the latent space based on a motion codebook. In the second stage, motion tokens generated by the MT module are randomly masked, and the parameters of the CMMT module are iteratively and selectively modified until the CMMT can accurately predict the masked tokens based on the unmasked tokens and an input text).
Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method for training a model for generating a representation of long-term motion from a text input taught by LEE incorporate the teachings of Wang, applying the system for controlling a character in a virtual environment based on a text input taught by Wang to provide the trained transformer to iterative the unmasking motion tokens until the vector quantization variable autoencoder (VQVAE) transform the unmasked runtime motion tokens into a runtime skeletal representation of the human motion. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify LEE according to the relied-upon teachings of Wang to obtain the invention as specified in claim.
Regarding claim 7, the combination of LEE in view of Qin in view of Wang discloses everything claimed as applied above (see claim 6).
However, LEE does not specifically disclose comprising performing independent sampling wherein the denoise transformer iteratively unmasks the runtime motion tokens for the first set of runtime motion tokens independent of the second set of runtime motion tokens.
In additional, Wang discloses comprising performing independent sampling (Paragraph [0042], FIGS. 1A and 2A-2B, the MT module 22 may be pretrained or trained to accurately generate a plurality of motion tokens based on training text inputs indicating one or more sample motions of the character) wherein the denoise transformer iteratively unmasks the runtime motion tokens (Paragraph [0034], motion tokens generated by the MT module are randomly masked, and the parameters of the CMMT module are iteratively and selectively modified until the CMMT can accurately predict the masked tokens based on the unmasked tokens and an input text) for the first set of runtime motion tokens (Paragraph [0046], the TBT module 24 is configured to generate a plurality of text-based tokens based on the text input) independent of the second set of runtime motion tokens (Paragraph [0047], the plurality of motion tokens generated by the MT module 22).
Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method for training a model for generating a representation of long-term motion from a text input taught by LEE incorporate the teachings of Wang, applying the system for controlling a character in a virtual environment based on a text input taught by Wang to provide the trained transformer to iterative the unmasking motion tokens until the vector quantization variable autoencoder (VQVAE) transform the unmasked runtime motion tokens into a runtime skeletal representation of the human motion. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify LEE according to the relied-upon teachings of Wang to obtain the invention as specified in claim.
Regarding claim 8, the combination of LEE in view of Qin in view of Wang discloses everything claimed as applied above (see claim 6), and LEE further disclose comprising performing joint sampling (Paragraph [0006], in the presence of text inputs, human motion generation can also be cast into a machine-translation problem. A joint cross-modal latent space can also be used; paragraph [0145], ... in the trained model, the trained text encoder receives a text input comprising a plurality of phrases, where each phrase describes an action, and processes the received text input and a received duration to predict a latent sequence, the trained quantization module quantizes the latent sequence from the predicted latent sequence, and the trained motion decoder decodes the quantized latent sequence to output parameters for generating the representation of long-term motion; and wherein the representation of long-term motion is provided for one or more of (a) display on at least one display, and (b) control of at least one autonomous device. In combination with any of the above features in this paragraph, the autoencoder may comprise a vector quantization variable autoencoder (VQVAE). ... In combination with any of the above features in this paragraph, the entity may be a human. In combination with any of the above features in this paragraph, the entity may be a robot. In combination with any of the above features in this paragraph, training an autoencoder may use a gradient estimator. In combination with any of the above features in this paragraph, training an autoencoder may use a reconstruction loss; wherein the reconstruction loss may comprise an L1-loss, a reconstructed joint loss, and/or a velocity ...).
In additional, In additional, Wang discloses wherein the denoise transformer (Paragraph [0042], FIGS. 1A and 2A-2B, the MT module 22 may be pretrained or trained to accurately generate a plurality of motion tokens based on training text inputs indicating one or more sample motions of the character) iteratively unmasks the runtime motion tokens (Paragraph [0034], motion tokens generated by the MT module are randomly masked, and the parameters of the CMMT module are iteratively and selectively modified until the CMMT can accurately predict the masked tokens based on the unmasked tokens and an input text) for the first set of runtime motion tokens (Paragraph [0046], the TBT module 24 is configured to generate a plurality of text-based tokens based on the text input) and the second set of runtime motion tokens concurrently (Paragraph [0047], the plurality of motion tokens generated by the MT module 22).
Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method for training a model for generating a representation of long-term motion from a text input taught by LEE incorporate the teachings of Wang, applying the system for controlling a character in a virtual environment based on a text input taught by Wang to provide the trained transformer to iterative the unmasking motion tokens until the vector quantization variable autoencoder (VQVAE) transform the unmasked runtime motion tokens into a runtime skeletal representation of the human motion. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify LEE according to the relied-upon teachings of Wang to obtain the invention as specified in claim.
Regarding claim 9, the combination of LEE in view of Qin in view of Wang discloses everything claimed as applied above (see claim 6).
However, LEE does not specifically disclose wherein training the denoise transformer includes performing normalization on the one or more motion tokens.
In additional, Qin discloses wherein training the denoise transformer (Paragraph [0033], FIG. 2A is a simplified diagram illustrating an example training framework 200 of the t2v generation model described in FIG. 1 ...; paragraph [0039], details of the training and inference process of the denoising diffusion model of vDiT 110 may be provided below in relation to FIG. 2B; paragraphs [0049]-[0052], the Transformer-like denoising network (denoising model 222)) includes performing normalization on the one or more motion tokens (Paragraph [0038], the resulting latent representation 205 (after compressed both spatially and temporally) may be used to train the vDiT 110. The vDiT 110 may comprise a latent diffusion model which may be trained with denoising loss and uses Diffusion Transformer (DiT) as the diffusion backbone. For example, during training, a noise 208 may be iteratively added to the latent representation 205 to form a noised latent representation Zt 209, which is in turn input to the vDiT 110. The vDiT 110 is trained to estimate and/or remove the added noise 208 from the noised latent representation 209 to reconstruct a latent representation 213. Such denoising step is repeated iteratively so that over a number of iterations (e.g., 50 iterations), the reconstructed latent representation 213 may be considered as a denoised version of the noised latent representation Zt 209, which is supposedly close to the original latent representation 205. In this way, vDiT 110 may be trained by comparing the reconstructed latent representation 213 and the original latent representation 205).
Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method for training a model for generating a representation of long-term motion from a text input taught by LEE incorporate the teachings of Qin, applying the training framework for a denoising diffusion model taught by Qin to add a denoise transformer into the artificial intelligence system for performing the self-attention is based on relative positional encoding. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify LEE according to the relied-upon teachings of Qin to obtain the invention as specified in claim.
Regarding claim 10, the combination of LEE in view of Qin in view of Wang discloses everything claimed as applied above (see claim 6).
However, LEE does not specifically disclose wherein the training the denoise transformer includes performing the cross-attention based an action token derived from the action sentence.
In additional, Qin discloses wherein the training the denoise transformer includes performing the cross-attention based an action token derived from the action sentence (Paragraph [0115], the video diffusion model comprises a spatial attention layer (e.g., 215 in FIG. 2A), a temporal attention layer (e.g., 220 in FIG. 2A) and a text-video cross-attention layer. he spatial attention layer outputs attention weights capturing spatial information of an input vector relating to the training video. The temporal attention layer outputs attention weights capturing temporal characteristics of an input vector relating to the training video. The text-video cross-attention layer output attention weights capturing relationships between embeddings of the training text and spatial and/or temporal portions of the training video).
Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method for training a model for generating a representation of long-term motion from a text input taught by LEE incorporate the teachings of Qin, applying the training framework for a denoising diffusion model taught by Qin to add a denoise transformer into the artificial intelligence system and provide the method for training the denoise transformer. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify LEE according to the relied-upon teachings of Qin to obtain the invention as specified in claim.
Claims 11-20 are rejected under 35 U.S.C. 103 as being unpatentable over LEE et al (U.S. Patent Application Publication 2025/0371782 A1) in view of Wang et al (U.S. Patent Application Publication 2025/0181974 A1) in view of Qin et al (U.S. Patent Application Publication 2026/0044993 A1).
Regarding claim 11, LEE discloses a system for multi-motion generation, comprising:
a memory storing one or more instructions (Paragraphs [0132]-[0133], example systems, methods, and embodiments may be implemented within a system 1100 or a portion thereof such as illustrated in FIG. 11 ... Memory 1110 may also be provided in whole or in part by external storage in communication with the processor 1108; paragraph [0143], a storage medium (or a data carrier, or a computer-readable medium) comprises, stored thereon, the computer program or the computer-executable instructions ...);
a processor executing one or more of the instructions stored on the memory to perform (Paragraph [0133], memory 1110 may also be provided in whole or in part by external storage in communication with the processor 1108 ...; paragraph [0143], a storage medium (or a data carrier, or a computer-readable medium) comprises, stored thereon, the computer program or the computer-executable instructions for performing one of the methods described herein when it is performed by a processor):
receiving a runtime action sentence (Paragraph [0073], the computing device 700 may receive the input 740, such as an input text, from a user via the user interface; FIG. 3; paragraph [0065], the text encoder receives at 302 a text input comprising a plurality of phrases, where each phrase describes an action);
converting the runtime action sentence into a set of runtime motion tokens (FIG. 1; paragraphs [0054]-[0056], the text encoder 102 is configured to receive a text input ... The text input includes a plurality of text groups, chunks, or phrases, such as but not limited to sentences. A plurality of phrases may be embodied in a paragraph or other text group. Each phrase can describe an action, such that the plurality of phrases describes at least two actions. The plurality of phrases have an arbitrary length, that is, they can have any suitable length, and such a length can be independent of the length of a corresponding motion ... The text encoder 102 is configured to predict a latent representation comprising a continuous stream of latent vectors conditioned on the text input and the duration. Each latent vector represents a fixed length of motion ...; paragraph [0066], the text encoder 102 predicts at 306 a latent vector representation conditioned on the text input and the duration(s) that is embodied in a continuous stream of latent vectors. Each latent vector represents a fixed length of motion, and comprises a vector in a discrete latent space); and
transforming the unmasked runtime motion tokens (FIG. 3; paragraph [0067], the stream of latent vectors is then passed through the motion decoder 122, e.g., in a trained 1D convolutional VQVAE decoder. The motion decoder 122 decodes the (e.g., quantized and concatenated) latent representation at 310 to continuously reconstruct the desired long-term motion representation; paragraph [0104], transition latent vector: To avoid the need for long-term data, example methods can rely on transition tokens, which are designed to chain sequences together at inference. Such transition tokens can be obtained by masking at training time. Thus, the unmasked runtime motion tokens will be transformed) into a runtime skeletal representation of the human motion (Paragraph [0032], human motion is typically represented as a temporal sequence of 3D points, e.g., human meshes or skeletons ...) using a trained variational autoencoder (VAE) model (Paragraph [0005], Variational Auto-encoders (VAEs) ...; FIG. 1; paragraph [0058], in an example system 100, the motion decoder 122 is a convolutional decoder, and the motion encoder 126 is a convolutional encoder of the autoencoder 120, which is embodied in a vector quantization variable autoencoder (VQVAE)),
wherein the trained VAE model is trained (Paragraph [0005], Variational Auto-encoders (VAEs) ...; FIG. 1; paragraph [0058], in an example system 100, the motion decoder 122 is a convolutional decoder, and the motion encoder 126 is a convolutional encoder of the autoencoder 120, which is embodied in a vector quantization variable autoencoder (VQVAE); paragraphs [0061]-[0062], FIG. 2 shows an example training method 200 for the system 100 ... An example training method 200 first trains at 202 the autoencoder, e.g., the VQVAE 120 ...) based on a skeletal representation of human motion (Paragraph [0032], human motion is typically represented as a temporal sequence of 3D points, e.g., human meshes or skeletons ...), one or more motion tokens associated with the skeletal representation of the human motion (Paragraph [0045], in an initial training phase, the motion encoder and motion decoder can be trained to encode short human motion into learned discrete (specific) tokens), and a dynamic transition probability based on a distance between the one or more motion tokens (Paragraphs [0099]-[0105], for each motion, experiments ranked the Euclidean distance to 32 text descriptions of 1 positive and 31 negatives. The Top-1, Top-2, and Top-3 accuracy were determined ... an inference time as shown in FIG. 6 and described by example above. To train these transition latent vectors, an example method randomly substituted part of the quantized latent vectors Z into the transition latent vectors while training the VQVAE).
However, LEE does not specifically disclose iteratively unmasking runtime motion tokens of the set of runtime motion tokens using a trained denoise transformer; and
wherein the trained denoise transformer is trained by performing self-attention based on the one or more motion tokens and cross-attention based on the one or more motion tokens and an action sentence.
In additional, Wang discloses (Abstract, a system for controlling a character in a virtual environment includes at least one processor and at least one nontransitory computer-readable medium comprising instructions that are executable by the at least one processor. The instructions include receiving a text input indicating one or more motions of the character, generating a plurality of text-based tokens and a plurality of motion tokens, appending a plurality of masked tokens to the plurality of motion tokens, performing a masked transformer routine to generate a plurality of predicted motion tokens based on the plurality of masked tokens, the plurality of motion tokens, and the plurality of text-based tokens, decoding the plurality of predicted motion tokens and the plurality of motion tokens into a motion sequence of the character, where the motion sequence includes the one or more motions and a predicted motion of the character, and controlling the character based on the motion sequence) iteratively unmasking runtime motion tokens of the set of runtime motion tokens using a trained Paragraphs [0033]-[0034], the MMM system includes a motion tokenizer (MT) module that transforms human motion into a sequence of discrete tokens (e.g., discrete units, elements, and/or character strings that represent segments of the input data) in a latent space and a conditional masked motion transformer (CMMT) module that predicts randomly masked motion tokens, which are conditioned on the discrete tokens generated by the MT module ... the MMM system is trained based on a two-stage approach. In the first stage, the MT module is trained to convert and quantize raw motion data into a sequence of discrete motion tokens in the latent space based on a motion codebook. In the second stage, motion tokens generated by the MT module are randomly masked, and the parameters of the CMMT module are iteratively and selectively modified until the CMMT can accurately predict the masked tokens based on the unmasked tokens and an input text).
Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method for training a model for generating a representation of long-term motion from a text input taught by LEE incorporate the teachings of Wang, applying the system for controlling a character in a virtual environment based on a text input taught by Wang to provide the trained transformer to iterative the unmasking motion tokens until the vector quantization variable autoencoder (VQVAE) transform the unmasked runtime motion tokens into a runtime skeletal representation of the human motion. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify LEE according to the relied-upon teachings of Wang to obtain the invention as specified in claim.
However, the combination of LEE in view of Wang does not specifically disclose a trained denoise transformer; and
wherein the trained denoise transformer is trained by performing self-attention based on the one or more motion tokens and cross-attention based on the one or more motion tokens and an action sentence.
In additional, Qin discloses (Abstract, embodiments described herein provide a generation model comprising a video-specific variational auto-encoder (VAE) for effective compression of video pixel information with reduced spatial and temporal dimensions and a video diffusion transformer (vDiT) to generate latent representations of frames ...) a trained denoise transformer (Paragraph [0033], FIG. 2A is a simplified diagram illustrating an example training framework 200 of the t2v generation model described in FIG. 1 ...; paragraph [0039], details of the training and inference process of the denoising diffusion model of vDiT 110 may be provided below in relation to FIG. 2B; paragraphs [0049]-[0052], the Transformer-like denoising network (denoising model 222)); and
wherein the trained denoise transformer is trained (Paragraph [0033], FIG. 2A is a simplified diagram illustrating an example training framework 200 of the t2v generation model described in FIG. 1 ...; paragraph [0039], details of the training and inference process of the denoising diffusion model of vDiT 110 may be provided below in relation to FIG. 2B; paragraphs [0049]-[0052], the Transformer-like denoising network (denoising model 222)) by performing self-attention based on the one or more motion tokens (Paragraphs [0057]-[0059], FIG. 4 is a simplified diagram illustrating a data pipeline 400 for generating video-text training data for training the text-to-video generation model described in FIG. 1, according to some embodiments. First, a long-video clipping module 402 splits long videos into manageable clips ... an aesthetic scoring module 406 analyzes aesthetics and motion dynamics across frames to eliminate static video clips and inconsistent frames; paragraph [0040], vDiT 110 may comprise a stack of spatial-temporal transformer blocks as illustrated in FIG. 2A. Each transformer module comprises one or more modulation layers to scale and/or shift a representation vector, a spatial self-attention layer 215 to capture spatial information and a temporal self-attention layer 220 to capture temporal characteristics from the encoded video, and a feed forward layer to generate an output from the Transformer block ...; paragraph [0041], In FIG. 2C, the temporal self-attention layer 220 may adopt Rotary Positional Embedding (RoPE) to encode temporal information, e.g., to compute attentions between an input matrix having a size of (B, H, W) capturing visual information of each frame in the batch ... In FIG. 2D, the spatial self-attention layer 215 may adopt sinusoidal encoding to encode spatial information, e.g., to compute attentions between an input matrix having a size of (B, T) capturing the batch size and the total number of frames in the input video, input matrix (H, W) capturing spatially distributed visual content on each frame and the input vector having a size of C representing the number of image channels of each frame) and cross-attention based on the one or more motion tokens and an action sentence (Paragraph [0115], the video diffusion model comprises a spatial attention layer (e.g., 215 in FIG. 2A), a temporal attention layer (e.g., 220 in FIG. 2A ) and a text-video cross-attention layer. he spatial attention layer outputs attention weights capturing spatial information of an input vector relating to the training video. The temporal attention layer outputs attention weights capturing temporal characteristics of an input vector relating to the training video. The text-video cross-attention layer output attention weights capturing relationships between embeddings of the training text and spatial and/or temporal portions of the training video).
Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method for training a model for generating a representation of long-term motion from a text input taught by LEE in view of Wang incorporate the teachings of Qin, applying the training framework for a denoising diffusion model taught by Qin to add a denoise transformer into the artificial intelligence system and provide the method for training the denoise transformer. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify LEE in view of Wang according to the relied-upon teachings of Qin to obtain the invention as specified in claim.
Regarding claim 12, the combination of LEE in view of Wang in view of Qin discloses everything claimed as applied above (see claim 11), and LEE further disclose wherein the VAE model is a vector quantized variational autoencoder (VQ-VAE) model (FIG. 1; paragraph [0058], in an example system 100, the motion decoder 122 is a convolutional decoder, and the motion encoder 126 is a convolutional encoder of the autoencoder 120, which is embodied in a vector quantization variable autoencoder (VQVAE)).
Regarding claim 13, the combination of LEE in view of Wang in view of Qin discloses everything claimed as applied above (see claim 11), and LEE further disclose wherein the VAE model (FIG. 1; paragraph [0058], in an example system 100, the motion decoder 122 is a convolutional decoder, and the motion encoder 126 is a convolutional encoder of the autoencoder 120, which is embodied in a vector quantization variable autoencoder (VQVAE)) converts the skeletal representation of the human motion into one or more motion tokens (Paragraph [0032], human motion is typically represented as a temporal sequence of 3D points, e.g., human meshes or skeletons ...; paragraphs [0054]-[0056], the text encoder 102 is configured to receive a text input ... The text input includes a plurality of text groups, chunks, or phrases, such as but not limited to sentences. A plurality of phrases may be embodied in a paragraph or other text group. Each phrase can describe an action, such that the plurality of phrases describes at least two actions. The plurality of phrases have an arbitrary length, that is, they can have any suitable length, and such a length can be independent of the length of a corresponding motion ... The text encoder 102 is configured to predict a latent representation comprising a continuous stream of latent vectors conditioned on the text input and the duration. Each latent vector represents a fixed length of motion ...; paragraph [0066], the text encoder 102 predicts at 306 a latent vector representation conditioned on the text input and the duration(s) that is embodied in a continuous stream of latent vectors. Each latent vector represents a fixed length of motion, and comprises a vector in a discrete latent space) and reconstructs the human motion from the one or more motion tokens (FIG. 3; paragraph [0067], the stream of latent vectors is then passed through the motion decoder 122, e.g., in a trained 1D convolutional VQVAE decoder. The motion decoder 122 decodes the (e.g., quantized and concatenated) latent representation at 310 to continuously reconstruct the desired long-term motion representation; paragraph [0104], transition latent vector: To avoid the need for long-term data, example methods can rely on transition tokens, which are designed to chain sequences together at inference. Such transition tokens can be obtained by masking at training time).
Regarding claim 14, the combination of LEE in view of Wang in view of Qin discloses everything claimed as applied above (see claim 11).
However, LEE does not specifically disclose wherein the training the denoise transformer by performing the self-attention is based on relative positional encoding (RPE).
In additional, Qin discloses wherein the training the denoise transformer (Paragraph [0033], FIG. 2A is a simplified diagram illustrating an example training framework 200 of the t2v generation model described in FIG. 1 ...; paragraph [0039], details of the training and inference process of the denoising diffusion model of vDiT 110 may be provided below in relation to FIG. 2B; paragraphs [0049]-[0052], the Transformer-like denoising network (denoising model 222)) by performing the self-attention is based on relative positional encoding (RPE) (Paragraph [0026], the Video-diffusion Transformer (vDiT) may incorporate transformer blocks with both temporal and spatial self-attention layers to encode spatial and temporal position information from latent representations of videos).
Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method for training a model for generating a representation of long-term motion from a text input taught by LEE incorporate the teachings of Qin, applying the training framework for a denoising diffusion model taught by Qin to add a denoise transformer into the artificial intelligence system for performing the self-attention is based on relative positional encoding. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify LEE according to the relied-upon teachings of Qin to obtain the invention as specified in claim.
Regarding claim 15, the combination of LEE in view of Wang in view of Qin discloses everything claimed as applied above (see claim 11). and LEE further disclose wherein the processor performs:
independent sampling (Paragraph [0006], in the presence of text inputs, human motion generation can also be cast into a machine-translation problem. A joint cross-modal latent space can also be used; paragraph [0145], ... in the trained model, the trained text encoder receives a text input comprising a plurality of phrases, where each phrase describes an action, and processes the received text input and a received duration to predict a latent sequence, the trained quantization module quantizes the latent sequence from the predicted latent sequence, and the trained motion decoder decodes the quantized latent sequence to output parameters for generating the representation of long-term motion; and wherein the representation of long-term motion is provided for one or more of (a) display on at least one display, and (b) control of at least one autonomous device. In combination with any of the above features in this paragraph, the autoencoder may comprise a vector quantization variable autoencoder (VQVAE). ... In combination with any of the above features in this paragraph, the entity may be a human. In combination with any of the above features in this paragraph, the entity may be a robot. In combination with any of the above features in this paragraph, training an autoencoder may use a gradient estimator. In combination with any of the above features in this paragraph, training an autoencoder may use a reconstruction loss; wherein the reconstruction loss may comprise an L1-loss, a reconstructed joint loss, and/or a velocity ...);
joint sampling (Paragraph [0006], in the presence of text inputs, human motion generation can also be cast into a machine-translation problem. A joint cross-modal latent space can also be used) wherein the trained denoise transformer iteratively unmasks the runtime motion tokens for the first set of runtime motion tokens and the second set of runtime motion tokens concurrently (Paragraph [0145], ... in the trained model, the trained text encoder receives a text input comprising a plurality of phrases, where each phrase describes an action, and processes the received text input and a received duration to predict a latent sequence, the trained quantization module quantizes the latent sequence from the predicted latent sequence, and the trained motion decoder decodes the quantized latent sequence to output parameters for generating the representation of long-term motion; and wherein the representation of long-term motion is provided for one or more of (a) display on at least one display, and (b) control of at least one autonomous device. In combination with any of the above features in this paragraph, the autoencoder may comprise a vector quantization variable autoencoder (VQVAE). ... In combination with any of the above features in this paragraph, the entity may be a human. In combination with any of the above features in this paragraph, the entity may be a robot. In combination with any of the above features in this paragraph, training an autoencoder may use a gradient estimator. In combination with any of the above features in this paragraph, training an autoencoder may use a reconstruction loss; wherein the reconstruction loss may comprise an L1-loss, a reconstructed joint loss, and/or a velocity ...); and
transforming the first set of unmasked runtime motion tokens and the second set of unmasked runtime motion tokens (FIG. 3; paragraph [0067], the stream of latent vectors is then passed through the motion decoder 122, e.g., in a trained 1D convolutional VQVAE decoder. The motion decoder 122 decodes the (e.g., quantized and concatenated) latent representation at 310 to continuously reconstruct the desired long-term motion representation; paragraph [0104], transition latent vector: To avoid the need for long-term data, example methods can rely on transition tokens, which are designed to chain sequences together at inference. Such transition tokens can be obtained by masking at training time. Thus, the unmasked runtime motion tokens will be transformed) into a runtime skeletal representation of the human motion (Paragraph [0032], human motion is typically represented as a temporal sequence of 3D points, e.g., human meshes or skeletons ...) using the trained VAE model (Paragraph [0005], Variational Auto-encoders (VAEs) ...; FIG. 1; paragraph [0058], in an example system 100, the motion decoder 122 is a convolutional decoder, and the motion encoder 126 is a convolutional encoder of the autoencoder 120, which is embodied in a vector quantization variable autoencoder (VQVAE)) based on the independent sampling (Paragraph [0065], at training time, durations can be extracted from the data, while at inference, durations can be either treated as an input or sampled from a prior (e.g., an average duration obtained from training data of action described by a phrase)) and the joint sampling (Paragraph [0006], in the presence of text inputs, human motion generation can also be cast into a machine-translation problem. A joint cross-modal latent space can also be used).
In additional, Wang discloses wherein the trained denoise transformer (Paragraph [0042], FIGS. 1A and 2A-2B, the MT module 22 may be pretrained or trained to accurately generate a plurality of motion tokens based on training text inputs indicating one or more sample motions of the character) iteratively unmasks runtime motion tokens Paragraph [0034], motion tokens generated by the MT module are randomly masked, and the parameters of the CMMT module are iteratively and selectively modified until the CMMT can accurately predict the masked tokens based on the unmasked tokens and an input text) for a first set of runtime motion tokens associated with a first action (Paragraph [0046], the TBT module 24 is configured to generate a plurality of text-based tokens based on the text input) independent of a second set of runtime motion tokens associated with a second action (Paragraph [0047], the plurality of motion tokens generated by the MT module 22).
Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method for training a model for generating a representation of long-term motion from a text input taught by LEE incorporate the teachings of Wang, applying the system for controlling a character in a virtual environment based on a text input taught by Wang to provide the trained transformer to iterative the unmasking motion tokens until the vector quantization variable autoencoder (VQVAE) transform the unmasked runtime motion tokens into a runtime skeletal representation of the human motion. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify LEE according to the relied-upon teachings of Wang to obtain the invention as specified in claim.
Regarding claim 16, LEE discloses a system for multi-motion generation, comprising:
a memory storing one or more instructions (Paragraphs [0132]-[0133], example systems, methods, and embodiments may be implemented within a system 1100 or a portion thereof such as illustrated in FIG. 11 ... Memory 1110 may also be provided in whole or in part by external storage in communication with the processor 1108; paragraph [0143], a storage medium (or a data carrier, or a computer-readable medium) comprises, stored thereon, the computer program or the computer-executable instructions ...);
a processor executing one or more of the instructions stored on the memory to perform (Paragraph [0133], memory 1110 may also be provided in whole or in part by external storage in communication with the processor 1108 ...; paragraph [0143], a storage medium (or a data carrier, or a computer-readable medium) comprises, stored thereon, the computer program or the computer-executable instructions for performing one of the methods described herein when it is performed by a processor):
receiving a runtime action sentence (Paragraph [0073], the computing device 700 may receive the input 740, such as an input text, from a user via the user interface; FIG. 3; paragraph [0065], the text encoder receives at 302 a text input comprising a plurality of phrases, where each phrase describes an action);
converting the runtime action sentence into a set of runtime motion tokens FIG. 1; paragraphs [0054]-[0056], the text encoder 102 is configured to receive a text input ... The text input includes a plurality of text groups, chunks, or phrases, such as but not limited to sentences. A plurality of phrases may be embodied in a paragraph or other text group. Each phrase can describe an action, such that the plurality of phrases describes at least two actions. The plurality of phrases have an arbitrary length, that is, they can have any suitable length, and such a length can be independent of the length of a corresponding motion ... The text encoder 102 is configured to predict a latent representation comprising a continuous stream of latent vectors conditioned on the text input and the duration. Each latent vector represents a fixed length of motion ...; paragraph [0066], the text encoder 102 predicts at 306 a latent vector representation conditioned on the text input and the duration(s) that is embodied in a continuous stream of latent vectors. Each latent vector represents a fixed length of motion, and comprises a vector in a discrete latent space); and
transforming the unmasked runtime motion tokens (FIG. 3; paragraph [0067], the stream of latent vectors is then passed through the motion decoder 122, e.g., in a trained 1D convolutional VQVAE decoder. The motion decoder 122 decodes the (e.g., quantized and concatenated) latent representation at 310 to continuously reconstruct the desired long-term motion representation; paragraph [0104], transition latent vector: To avoid the need for long-term data, example methods can rely on transition tokens, which are designed to chain sequences together at inference. Such transition tokens can be obtained by masking at training time. Thus, the unmasked runtime motion tokens will be transformed) into a runtime skeletal representation of the human motion (Paragraph [0032], human motion is typically represented as a temporal sequence of 3D points, e.g., human meshes or skeletons ...) using a trained vector quantized variational autoencoder (VQ-VAE) model (Paragraph [0005], Variational Auto-encoders (VAEs) ...; FIG. 1; paragraph [0058], in an example system 100, the motion decoder 122 is a convolutional decoder, and the motion encoder 126 is a convolutional encoder of the autoencoder 120, which is embodied in a vector quantization variable autoencoder (VQVAE)),
wherein the trained VQ-VAE model is trained (Paragraph [0005], Variational Auto-encoders (VAEs) ...; FIG. 1; paragraph [0058], in an example system 100, the motion decoder 122 is a convolutional decoder, and the motion encoder 126 is a convolutional encoder of the autoencoder 120, which is embodied in a vector quantization variable autoencoder (VQVAE); paragraphs [0061]-[0062], FIG. 2 shows an example training method 200 for the system 100 ... An example training method 200 first trains at 202 the autoencoder, e.g., the VQVAE 120 ...) based on a skeletal representation of human motion (Paragraph [0032], human motion is typically represented as a temporal sequence of 3D points, e.g., human meshes or skeletons ...), one or more motion tokens associated with the skeletal representation of the human motion (Paragraph [0045], in an initial training phase, the motion encoder and motion decoder can be trained to encode short human motion into learned discrete (specific) tokens), and a dynamic transition probability based on a distance between the one or more motion tokens (Paragraphs [0099]-[0105], for each motion, experiments ranked the Euclidean distance to 32 text descriptions of 1 positive and 31 negatives. The Top-1, Top-2, and Top-3 accuracy were determined ... an inference time as shown in FIG. 6 and described by example above. To train these transition latent vectors, an example method randomly substituted part of the quantized latent vectors Z into the transition latent vectors while training the VQVAE).
However, LEE does not specifically disclose iteratively unmasking runtime motion tokens of the set of runtime motion tokens using a trained denoise transformer; and
wherein the trained denoise transformer is trained by performing self-attention based on the one or more motion tokens and cross-attention based on the one or more motion tokens and an action sentence.
In additional, Wang discloses (Abstract, a system for controlling a character in a virtual environment includes at least one processor and at least one nontransitory computer-readable medium comprising instructions that are executable by the at least one processor. The instructions include receiving a text input indicating one or more motions of the character, generating a plurality of text-based tokens and a plurality of motion tokens, appending a plurality of masked tokens to the plurality of motion tokens, performing a masked transformer routine to generate a plurality of predicted motion tokens based on the plurality of masked tokens, the plurality of motion tokens, and the plurality of text-based tokens, decoding the plurality of predicted motion tokens and the plurality of motion tokens into a motion sequence of the character, where the motion sequence includes the one or more motions and a predicted motion of the character, and controlling the character based on the motion sequence) iteratively unmasking runtime motion tokens of the set of runtime motion tokens using a trained Paragraphs [0033]-[0034], the MMM system includes a motion tokenizer (MT) module that transforms human motion into a sequence of discrete tokens (e.g., discrete units, elements, and/or character strings that represent segments of the input data) in a latent space and a conditional masked motion transformer (CMMT) module that predicts randomly masked motion tokens, which are conditioned on the discrete tokens generated by the MT module ... the MMM system is trained based on a two-stage approach. In the first stage, the MT module is trained to convert and quantize raw motion data into a sequence of discrete motion tokens in the latent space based on a motion codebook. In the second stage, motion tokens generated by the MT module are randomly masked, and the parameters of the CMMT module are iteratively and selectively modified until the CMMT can accurately predict the masked tokens based on the unmasked tokens and an input text).
Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method for training a model for generating a representation of long-term motion from a text input taught by LEE incorporate the teachings of Wang, applying the system for controlling a character in a virtual environment based on a text input taught by Wang to provide the trained transformer to iterative the unmasking motion tokens until the vector quantization variable autoencoder (VQVAE) transform the unmasked runtime motion tokens into a runtime skeletal representation of the human motion. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify LEE according to the relied-upon teachings of Wang to obtain the invention as specified in claim.
However, the combination of LEE in view of Wang does not specifically disclose a trained denoise transformer; and
wherein the trained denoise transformer is trained by performing self-attention based on the one or more motion tokens and cross-attention based on the one or more motion tokens and an action sentence.
In additional, Qin discloses (Abstract, embodiments described herein provide a generation model comprising a video-specific variational auto-encoder (VAE) for effective compression of video pixel information with reduced spatial and temporal dimensions and a video diffusion transformer (vDiT) to generate latent representations of frames ...) a trained denoise transformer (Paragraph [0033], FIG. 2A is a simplified diagram illustrating an example training framework 200 of the t2v generation model described in FIG. 1 ...; paragraph [0039], details of the training and inference process of the denoising diffusion model of vDiT 110 may be provided below in relation to FIG. 2B; paragraphs [0049]-[0052], the Transformer-like denoising network (denoising model 222)); and
wherein the trained denoise transformer is trained (Paragraph [0033], FIG. 2A is a simplified diagram illustrating an example training framework 200 of the t2v generation model described in FIG. 1 ...; paragraph [0039], details of the training and inference process of the denoising diffusion model of vDiT 110 may be provided below in relation to FIG. 2B; paragraphs [0049]-[0052], the Transformer-like denoising network (denoising model 222)) by performing self-attention based on the one or more motion tokens (Paragraphs [0057]-[0059], FIG. 4 is a simplified diagram illustrating a data pipeline 400 for generating video-text training data for training the text-to-video generation model described in FIG. 1, according to some embodiments. First, a long-video clipping module 402 splits long videos into manageable clips ... an aesthetic scoring module 406 analyzes aesthetics and motion dynamics across frames to eliminate static video clips and inconsistent frames; paragraph [0040], vDiT 110 may comprise a stack of spatial-temporal transformer blocks as illustrated in FIG. 2A. Each transformer module comprises one or more modulation layers to scale and/or shift a representation vector, a spatial self-attention layer 215 to capture spatial information and a temporal self-attention layer 220 to capture temporal characteristics from the encoded video, and a feed forward layer to generate an output from the Transformer block ...; paragraph [0041], In FIG. 2C, the temporal self-attention layer 220 may adopt Rotary Positional Embedding (RoPE) to encode temporal information, e.g., to compute attentions between an input matrix having a size of (B, H, W) capturing visual information of each frame in the batch ... In FIG. 2D, the spatial self-attention layer 215 may adopt sinusoidal encoding to encode spatial information, e.g., to compute attentions between an input matrix having a size of (B, T) capturing the batch size and the total number of frames in the input video, input matrix (H, W) capturing spatially distributed visual content on each frame and the input vector having a size of C representing the number of image channels of each frame) and cross-attention based on the one or more motion tokens and an action sentence (Paragraph [0115], the video diffusion model comprises a spatial attention layer (e.g., 215 in FIG. 2A), a temporal attention layer (e.g., 220 in FIG. 2A) and a text-video cross-attention layer. he spatial attention layer outputs attention weights capturing spatial information of an input vector relating to the training video. The temporal attention layer outputs attention weights capturing temporal characteristics of an input vector relating to the training video. The text-video cross-attention layer output attention weights capturing relationships between embeddings of the training text and spatial and/or temporal portions of the training video).
Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method for training a model for generating a representation of long-term motion from a text input taught by LEE in view of Wang incorporate the teachings of Qin, applying the training framework for a denoising diffusion model taught by Qin to add a denoise transformer into the artificial intelligence system and provide the method for training the denoise transformer. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify LEE in view of Wang according to the relied-upon teachings of Qin to obtain the invention as specified in claim.
Regarding claim 17, the combination of LEE in view of Wang in view of Qin discloses everything claimed as applied above (see claim 16).
However, LEE does not specifically disclose wherein the training the denoise transformer by performing the self-attention is based on relative positional encoding (RPE).
In additional, Qin discloses wherein the training the denoise transformer (Paragraph [0033], FIG. 2A is a simplified diagram illustrating an example training framework 200 of the t2v generation model described in FIG. 1 ...; paragraph [0039], details of the training and inference process of the denoising diffusion model of vDiT 110 may be provided below in relation to FIG. 2B; paragraphs [0049]-[0052], the Transformer-like denoising network (denoising model 222)) by performing the self-attention is based on relative positional encoding (RPE) (Paragraph [0026], the Video-diffusion Transformer (vDiT) may incorporate transformer blocks with both temporal and spatial self-attention layers to encode spatial and temporal position information from latent representations of videos).
Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method for training a model for generating a representation of long-term motion from a text input taught by LEE incorporate the teachings of Qin, applying the training framework for a denoising diffusion model taught by Qin to add a denoise transformer into the artificial intelligence system for performing the self-attention is based on relative positional encoding. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify LEE according to the relied-upon teachings of Qin to obtain the invention as specified in claim.
Regarding claim 18, the combination of LEE in view of Wang in view of Qin discloses everything claimed as applied above (see claim 16).
The claim recites the same limitations of claim 16. Therefore, the combination of LEE in view of Wang in view of Qin discloses comprising:
receiving a runtime action sentence (see claim 16);
converting the runtime action sentence into a set of runtime motion tokens (see claim 16);
iteratively unmasking runtime motion tokens of the set of runtime motion tokens using the trained denoise transformer (see claim 16).; and
transforming the unmasked runtime motion tokens into a runtime skeletal representation of the human motion using the trained VQ-VAE model (see claim 16).
Regarding claim 19, the combination of LEE in view of Wang in view of Qin discloses everything claimed as applied above (see claim 16), and LEE further disclose
comprising:
receiving a runtime action sentence including a first action and a second action (FIG. 7; paragraph [0073], the computing device 700 may receive the input 740, such as an input text, from a user via the user interface; FIG. 3; paragraph [0065], the text encoder receives at 302 a text input comprising a plurality of phrases, where each phrase describes an action);
converting the runtime action sentence into a first set of runtime motion tokens corresponding to the first action and a second set of runtime motion tokens corresponding to the second action (FIG. 1; paragraphs [0054]-[0056], the text encoder 102 is configured to receive a text input ... The text input includes a plurality of text groups, chunks, or phrases, such as but not limited to sentences. A plurality of phrases may be embodied in a paragraph or other text group. Each phrase can describe an action, such that the plurality of phrases describes at least two actions. The plurality of phrases have an arbitrary length, that is, they can have any suitable length, and such a length can be independent of the length of a corresponding motion ... The text encoder 102 is configured to predict a latent representation comprising a continuous stream of latent vectors conditioned on the text input and the duration. Each latent vector represents a fixed length of motion ...; paragraph [0066], the text encoder 102 predicts at 306 a latent vector representation conditioned on the text input and the duration(s) that is embodied in a continuous stream of latent vectors. Each latent vector represents a fixed length of motion, and comprises a vector in a discrete latent space);
transforming the first set of unmasked runtime motion tokens and the second set of unmasked runtime motion tokens (FIG. 3; paragraph [0067], the stream of latent vectors is then passed through the motion decoder 122, e.g., in a trained 1D convolutional VQVAE decoder. The motion decoder 122 decodes the (e.g., quantized and concatenated) latent representation at 310 to continuously reconstruct the desired long-term motion representation; paragraph [0104], transition latent vector: To avoid the need for long-term data, example methods can rely on transition tokens, which are designed to chain sequences together at inference. Such transition tokens can be obtained by masking at training time. Thus, the unmasked runtime motion tokens will be transformed) into a runtime skeletal representation of the human motion (Paragraph [0032], human motion is typically represented as a temporal sequence of 3D points, e.g., human meshes or skeletons ...) using the trained VAE model (Paragraph [0005], Variational Auto-encoders (VAEs) ...; FIG. 1; paragraph [0058], in an example system 100, the motion decoder 122 is a convolutional decoder, and the motion encoder 126 is a convolutional encoder of the autoencoder 120, which is embodied in a vector quantization variable autoencoder (VQVAE)).
However, LEE does not specifically disclose iteratively unmasking runtime motion tokens of the first set of runtime motion tokens using the trained denoise transformer to generate a first set of unmasked runtime motion tokens;
iteratively unmasking runtime motion tokens of the second set of runtime motion tokens using the trained denoise transformer to generate a second set of unmasked runtime motion tokens.
In additional, Wang discloses iteratively unmasking runtime motion tokens of the first set of runtime motion tokens using the trained denoise transformer to generate a first set of unmasked runtime motion tokens (Paragraphs [0033]-[0034], the MMM system includes a motion tokenizer (MT) module that transforms human motion into a sequence of discrete tokens (e.g., discrete units, elements, and/or character strings that represent segments of the input data) in a latent space and a conditional masked motion transformer (CMMT) module that predicts randomly masked motion tokens, which are conditioned on the discrete tokens generated by the MT module ... the MMM system is trained based on a two-stage approach. In the first stage, the MT module is trained to convert and quantize raw motion data into a sequence of discrete motion tokens in the latent space based on a motion codebook. In the second stage, motion tokens generated by the MT module are randomly masked, and the parameters of the CMMT module are iteratively and selectively modified until the CMMT can accurately predict the masked tokens based on the unmasked tokens and an input text);
iteratively unmasking runtime motion tokens of the second set of runtime motion tokens using the trained denoise transformer to generate a second set of unmasked runtime motion tokens (Paragraphs [0033]-[0034], the MMM system includes a motion tokenizer (MT) module that transforms human motion into a sequence of discrete tokens (e.g., discrete units, elements, and/or character strings that represent segments of the input data) in a latent space and a conditional masked motion transformer (CMMT) module that predicts randomly masked motion tokens, which are conditioned on the discrete tokens generated by the MT module ... the MMM system is trained based on a two-stage approach. In the first stage, the MT module is trained to convert and quantize raw motion data into a sequence of discrete motion tokens in the latent space based on a motion codebook. In the second stage, motion tokens generated by the MT module are randomly masked, and the parameters of the CMMT module are iteratively and selectively modified until the CMMT can accurately predict the masked tokens based on the unmasked tokens and an input text).
Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method for training a model for generating a representation of long-term motion from a text input taught by LEE incorporate the teachings of Wang, applying the system for controlling a character in a virtual environment based on a text input taught by Wang to provide the trained transformer to iterative the unmasking motion tokens until the vector quantization variable autoencoder (VQVAE) transform the unmasked runtime motion tokens into a runtime skeletal representation of the human motion. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify LEE according to the relied-upon teachings of Wang to obtain the invention as specified in claim.
Regarding claim 20, the combination of LEE in view of Wang in view of Qin discloses everything claimed as applied above (see claim 19), and LEE further disclose comprising:
performing independent sampling (Paragraph [0006], in the presence of text inputs, human motion generation can also be cast into a machine-translation problem. A joint cross-modal latent space can also be used; paragraph [0145], ... in the trained model, the trained text encoder receives a text input comprising a plurality of phrases, where each phrase describes an action, and processes the received text input and a received duration to predict a latent sequence, the trained quantization module quantizes the latent sequence from the predicted latent sequence, and the trained motion decoder decodes the quantized latent sequence to output parameters for generating the representation of long-term motion; and wherein the representation of long-term motion is provided for one or more of (a) display on at least one display, and (b) control of at least one autonomous device. In combination with any of the above features in this paragraph, the autoencoder may comprise a vector quantization variable autoencoder (VQVAE). ... In combination with any of the above features in this paragraph, the entity may be a human. In combination with any of the above features in this paragraph, the entity may be a robot. In combination with any of the above features in this paragraph, training an autoencoder may use a gradient estimator. In combination with any of the above features in this paragraph, training an autoencoder may use a reconstruction loss; wherein the reconstruction loss may comprise an L1-loss, a reconstructed joint loss, and/or a velocity ...); and
performing joint sampling (Paragraph [0006], in the presence of text inputs, human motion generation can also be cast into a machine-translation problem. A joint cross-modal latent space can also be used) wherein the denoise transformer iteratively unmasks the runtime motion tokens for the first set of runtime motion tokens and the second set of runtime motion tokens concurrently (Paragraph [0145], ... in the trained model, the trained text encoder receives a text input comprising a plurality of phrases, where each phrase describes an action, and processes the received text input and a received duration to predict a latent sequence, the trained quantization module quantizes the latent sequence from the predicted latent sequence, and the trained motion decoder decodes the quantized latent sequence to output parameters for generating the representation of long-term motion; and wherein the representation of long-term motion is provided for one or more of (a) display on at least one display, and (b) control of at least one autonomous device. In combination with any of the above features in this paragraph, the autoencoder may comprise a vector quantization variable autoencoder (VQVAE). ... In combination with any of the above features in this paragraph, the entity may be a human. In combination with any of the above features in this paragraph, the entity may be a robot. In combination with any of the above features in this paragraph, training an autoencoder may use a gradient estimator. In combination with any of the above features in this paragraph, training an autoencoder may use a reconstruction loss; wherein the reconstruction loss may comprise an L1-loss, a reconstructed joint loss, and/or a velocity ...).
In additional, Wang discloses wherein the denoise transformer (Paragraph [0042], FIGS. 1A and 2A-2B, the MT module 22 may be pretrained or trained to accurately generate a plurality of motion tokens based on training text inputs indicating one or more sample motions of the character) iteratively unmasks the runtime motion tokens (Paragraph [0034], motion tokens generated by the MT module are randomly masked, and the parameters of the CMMT module are iteratively and selectively modified until the CMMT can accurately predict the masked tokens based on the unmasked tokens and an input text) for the first set of runtime motion tokens (Paragraph [0046], the TBT module 24 is configured to generate a plurality of text-based tokens based on the text input) independent of the second set of runtime motion tokens (Paragraph [0047], the plurality of motion tokens generated by the MT module 22).
Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method for training a model for generating a representation of long-term motion from a text input taught by LEE incorporate the teachings of Wang, applying the system for controlling a character in a virtual environment based on a text input taught by Wang to provide the trained transformer to iterative the unmasking motion tokens until the vector quantization variable autoencoder (VQVAE) transform the unmasked runtime motion tokens into a runtime skeletal representation of the human motion. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify LEE according to the relied-upon teachings of Wang to obtain the invention as specified in claim.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Xilin Guo whose telephone number is (571)272-5786. The examiner can normally be reached Monday - Friday 9:00 AM-5:30 PM EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Hajnik can be reached at 571-272-7642. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/XILIN GUO/Primary Examiner, Art Unit 2616