Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Status
Applicant’s amendments filed on 08/06/2026 have been received and considered. Claims 1-20 are pending. Claims 1, 7-9, 11, and 17-19 have been amended. No new claims have been added or cancelled.
Response to Arguments
Applicant’s arguments, see pgs. 7-12, filed 08/06/2026, with respect to the rejections under 35 U.S.C. 103 have been fully considered and are persuasive.
More specifically, the Applicant argues that regarding claim 1, “Modifying Villanueva to automatically extract the expression style directly from the input audio as taught by Zhang would fundamentally destroy Villanueva's stated purpose of decoupling the speech-driven and speech-agnostic components to allow manual control of facial expressions.” [Remarks, pg. 9, par. 3, lines 1-3]. The Examiner agrees.
The Applicant also argues that regarding claim 9, “Kim does not process an input speech audio signal.” [Remarks, pg. 10, full par. 1, line 1] and “In even further contrast, Kim does not use transformer blocks to generate the style vector.” [Remarks, pg. 10, full par. 4, line 1]. The Examiner agrees.
Furthermore, the Applicant argues that the combination of Huang and Zhang fails to teach the amended limitation “comprising a transformer encoder configured to extract a latent code from the input speech audio signal” [Remarks, pg. 9, par. 1]. The Examiner agrees.
Therefore, the rejections under 35 U.S.C. 103 of claims 1, 9, and 11 have been withdrawn. Dependent claims 2-10 and 12-20 have also been withdrawn.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-2, 6, 9, 11-12, 16, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Huang et al. (U.S. 2025/0061634 A1, hereinafter Huang) in view of Zhang et al. (“Speech Emotion Recognition with Complementary Acoustic Representations”, hereinafter Zhang).
Regarding claim 1, Huang teaches a method for speech-driven three dimensional (3D) facial animation, the method comprising: ([0023] “Approaches in accordance with various embodiments can generate animation that is representative of one or more characters uttering speech represented by audio data. This can include, for example, high resolution, full three-dimensional (3D) facial animation”)
receiving, by a processor, an input speech audio signal; ([0023] “One or more deep neural networks, such as a frame-based convolutional neural network (CNN), or recurrent neural network (RNN), and/or a transformer-based model can take as input raw audio”, where “In at least one embodiment, computer system 900 may include, without limitation, processor 902 that may include, without limitation, one or more execution units 908 to perform machine learning model training and/or inferencing according to techniques described herein.” [0131])
inputting, by the processor, the input speech audio signal and the speaker style vector into a mesh generation model ([0039] “A deep neural network 206, as described in more detail later herein, can output a vector that encodes position or motion data for various points on a mesh for one or more facial components, and can feed this output vector (or another output, such as a global transformation matrix) to a renderer 216 that can apply these values to one or more meshes for this character in order to guide the animation.”, where “One or more deep neural networks… can take as input raw audio, extract features from the raw audio, and receive one or more component vectors with which a character is to be animated to utter speech contained in an audio segment extracted from the input raw audio.” [0023], and “one or more component vectors, such as a style vector or an emotion vector, that indicates one or more emotions” [0023])
and generating vertex position information for a 3D facial animation based on the input speech audio signal and the speaker style vector; ([0038] “During inferencing, the network may receive the audio input 202 (e.g., only audio data in some embodiments) as input, and may infer a set of vertex positions 214 for various facial components (e.g., head, face, eyeballs, jaw, tongue), which can then be fed to a renderer 216 (e.g., a rendering engine of an animation or video synthesis system) in order to generate a frame of animation 218, which may be one of a series of frames that provide the animation upon presentation or playback. As discussed in more detail elsewhere herein, emotion or style vector data may also be provided as input to the deep neural network 206”)
and outputting, by the processor, the vertex position information ([0080] “Finally, the vertex positions of the mesh can be exported for each frame in the shot.”)
Huang fails to teach generating, by the processor, a speaker style vector from the input speech audio signal based on a speaker style embedding model comprising a transformer encoder configured to extract a latent code from the input speech audio signal. However, this is known in the art as taught by Zhang.
Zhang teaches generating, by the processor, a speaker style vector from the input speech audio signal based on a speaker style embedding model comprising a transformer encoder configured to extract a latent code from the input speech audio signal (Fig. 1, where “The Overall network architecture of the proposed framework where a CNN encoder and a Transformer encoder learn acoustic representations in parallel… the input to the Transformer is MFCC. The latent acoustic embeddings created by the two complementary encoders are fused to predict the emotion of the speech sequence.” [Fig. 1])
Zhang is analogous to the claimed invention, as both relate to speech style recognition. Zhang further teaches that “Research efforts based on Transformers have achieved great success in ASR [8] and multimodal emotion recognition [9]” [pg. 847, col. 1, lines 2-3] and “Transformers, on the other hand, promote global information by exploiting long-range dependencies through self-attention, which have shown to be effective in NLP [3] and computer vision [4]” [pg. 846, col. 1, par. 2, lines 5-8]. Therefore, it would be obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teachings of Zhang to Huang, as transformer encoders are successful in emotion recognition in speech input by taking advantage of global acoustic features.
Regarding claim 2, the combination of Huang and Zhang teaches the method of claim 1, further comprising: displaying the 3D facial animation with animated movements based on the vertex position information (Huang; [0038] “During inferencing, the network may receive the audio input 202 (e.g., only audio data in some embodiments) as input, and may infer a set of vertex positions 214 for various facial components (e.g., head, face, eyeballs, jaw, tongue), which can then be fed to a renderer 216 (e.g., a rendering engine of an animation or video synthesis system) in order to generate a frame of animation 218, which may be one of a series of frames that provide the animation upon presentation or playback. As discussed in more detail elsewhere herein, emotion or style vector data may also be provided as input to the deep neural network 206”, and “In some embodiments, a user may also be able to provide, as a type of style input, adjustment to specific feature points or facial components in the display.” [0056]).
Regarding claim 6, the combination of Huang and Zhang teaches the method of claim 1, wherein the mesh generation model is trained based on a Mean Squared Error (MSE) loss function (Huang; [0084] “One error metric that can be used is the mean of squared differences between the desired output y and the output produced by the network 9.”)
Regarding claim 9, the combination of Huang and Zhang teaches the method of claim 1, wherein the generating the speaker style vector includes: inputting feature vectors based on the input speech audio signal to a transformer encoder; (Zhang; Fig. 1, where “the input to the Transformer is MFCC” [Fig. 1, line 4], and “The MFCC features extracted from the audio signal are positionally encoded and projected to obtain the query Q, the key K and the value V, respectively.” [pg. 847, col. 1, full par. 5, lines 4-6]).
and generating the speaker style vector based on an output of the transformer encoder (“The input to the CNN is logMel while the input to the Transformer is MFCC. The latent acoustic embeddings created by the two complementary encoders are fused to predict the emotion of the speech sequence.”).
Zhang is analogous to the claimed invention, as both relate to speech style recognition. Zhang further teaches that “Research efforts based on Transformers have achieved great success in ASR [8] and multimodal emotion recognition [9]” [pg. 847, col. 1, lines 2-3] and “Transformers, on the other hand, promote global information by exploiting long-range dependencies through self-attention, which have shown to be effective in NLP [3] and computer vision [4]” [pg. 846, col. 1, par. 2, lines 5-8]. Therefore, it would be obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teachings of Zhang to Huang, as transformer encoders are successful in emotion recognition in speech input by taking advantage of global acoustic features.
Regarding claim 11, claim 11 recites substantially similar limitations to claim 1, but in a device form. The combination of Huang and Zhang further teaches an artificial intelligence (AI) device, comprising: (Huang; [0037] “In this system 200, the component vector 204 is fed into an articulation network portion 210 of the deep neural network 206 at multiple levels, including at least a beginning and an end of the network to help condition the network.”)
a memory configured to store facial animation information; (Huang; [0064] “This animation can then be provided 512 for purposes such as presentation or storage, among other such options.”, where “In at least one embodiment memory device 1120 can operate as system memory for system 1100, to store data 1122 and instruction 1121 for use when one or more processor(s) 1102 executes an application or process.” [0151])
and a controller configured to: (Huang; [0151] “In at least one embodiment, memory controller 1116 also couples with an optional external graphics processor 1112, which may communicate with one or more graphics processor(s) 1108 in processor(s) 1102 to perform graphics and media operations.”)
Regarding claim 12, claim 12 recites substantially similar limitations to claim 2, therefore, will be rejected under the same rationale as claim 2.
Regarding claim 16, claim 16 recites substantially similar limitations to claim 6, therefore, will be rejected under the same rationale as claim 6.
Regarding claim 19, claim 19 recites substantially similar limitations to claim 9, therefore, will be rejected under the same rationale as claim 9.
Claims 3-4 and 13-14 are rejected under 35 U.S.C. 103 as being unpatentable over Huang in view of Zhang, and further in view of Khakhulin et al. (U.S. Patent No. 12,169,900, hereinafter Khakhulin).
Regarding claim 3, the combination of Huang and Zhang teaches the method of claim 1, but fails to teach wherein the mesh generation model is trained based on a two dimensional (2D) photometric loss function based on inverse rendering of predicted vertex position information and corresponding ground truth vertex position information. However, this is known in the art as taught by Khakhulin.
Khakhulin teaches wherein the mesh generation model is trained based on a two dimensional (2D) photometric loss function based on inverse rendering of predicted vertex position information and corresponding ground truth vertex position information ([col. 3, lines 9-13] “The rendering may include reconstructing the predicted image and the segmentation mask based by comparing the predicted image and the segmentation mask with a ground-truth image and a mask of the ground-truth image via a photometric loss.”, where “The providing the predicted mesh may include: rendering the initial mesh into an xyz-coordinate texture” [col. 2, lines 60-61], “A key feature of the method according to the disclosure is usage of solely 2D supervision on geometry” [col. 8, lines 14-15], and “In the method according to the disclosure, learned is geometry without any ground truth 3D supervision during training or pre-training (on top of the pretrained DECA estimator). For that utilized are… photometric losses” [col. 10, lines 29-34]).
Khakhulin is analogous to the claimed invention, as both relate to creating a 3D animation of a head mesh. Khakhulin further teaches “Photometric loss terms not only allow to obtain photorealistic renders but also aid in learning proper geometric reconstructions.” [col. 11, lines 26-28]. Therefore, it would be obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teachings of Khakhulin to the combination of Huang and Zhang in order to create photorealistic and proper geometric 3D meshes to be used.
Regarding claim 4, the combination of Huang, Zhang, and Khakhulin teaches the method of claim 3, wherein the 2D photometric loss function includes a mask parameter configured to remove background information to isolate a face ([col. 3, lines 9-13] “The rendering may include reconstructing the predicted image and the segmentation mask based by comparing the predicted image and the segmentation mask with a ground-truth image and a mask of the ground-truth image via a photometric loss.”, where “a method for 3D-reconstruction of human head for obtaining render of human image, using a single source image, wherein face shape extracted from the single source image, head pose, the facial expression extracted” [col. 5, lines 50-53], see FIG. 1)
Khakhulin is analogous to the claimed invention, as both relate to creating a 3D animation of a head mesh. Khakhulin further teaches “Photometric loss terms not only allow to obtain photorealistic renders but also aid in learning proper geometric reconstructions.” [col. 11, lines 26-28]. Therefore, it would be obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teachings of Khakhulin to the combination of Huang and Zhang in order to create photorealistic and proper geometric 3D meshes to be used.
Regarding claim 13, claim 13 recites substantially similar limitations to claim 3, therefore, will be rejected under the same rationale as claim 3.
Regarding claim 14, claim 14 recites substantially similar limitations to claim 4, therefore, will be rejected under the same rationale as claim 4.
Claims 5 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Huang in view of Zhang and Khakhulin, and further in view of Shu et al. (“Feature-metric Loss for Self-supervised Learning of Depth and Egomotion”, hereinafter Shu).
Regarding claim 5, the combination of Huang, Zhang, and Khakhulin teaches the method of claim 3, but fails to teach wherein the 2D photometric loss function is based on a pixel difference between two 2D images. However, this is known in the art as taught by Shu.
Shu teaches that it is known in the art, wherein the 2D photometric loss function is based on a pixel difference between two 2D images. ([pg. 4, par. 1, lines 10-11] “the per-pixel loss which measures the photometric difference, i.e, photometric loss”, and “photometric loss, which is defined as the photometric difference between a pixel warped from source view by estimated depth and pose and the pixel captured in the target view” [pg. 2, par. 1, lines 3-5]). Shu is also analogous to the art, as both relate to training machine learning models for 3D reconstruction. Therefore, it would be obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teachings of Shu to the combination of Huang, Zhang, and Khakhulin for the photometric loss to be based on pixel differences as it is known in the art of 3D reconstruction training for machine learning models.
Regarding claim 15, claim 15 recites substantially similar limitations to claim 5, therefore, rejected under the same rationale as claim 5.
Claims 7 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Huang in view of Zhang, and further in view of Villanueva Aylagas et al. (U.S. Patent No. 12,406,419, hereinafter Villanueva).
Regarding claim 7, the combination of Huang and Zhang teaches the method of claim 1, but fails to teach wherein the vertex position information includes a tensor of dimension NUMFRAMES × VERTEXCOUNT × 3, where NUMFRAMES is a number of frames based on a length of input speech audio signal, VERTEXCOUNT is a number of vertices based on a mesh for the 3D facial animation, and 3 corresponds to x, y and z coordinates. However, this is known in the art as taught by Villanueva.
Villanueva teaches wherein the vertex position information includes a tensor of dimension NUMFRAMES × VERTEXCOUNT × 3, where NUMFRAMES is a number of frames based on a length of input speech audio signal, VERTEXCOUNT is a number of vertices based on a mesh for the 3D facial animation, and 3 corresponds to x, y and z coordinates. (Villanueva; [col. 9, lines 56-59] “The mesh data at for an animation frame may be provided as a vector of mesh vertex coordinates. For example at may be a vector of size 3V representing three-dimensional coordinates for V vertices.”, where “The acoustic features for an audio frame 202 for a time step t” [col. 9, lines 49-52], and “the cVAE is being trained to reconstruct facial animation data a2 for an animation frame using a sequence of acoustic features from three audio frames” [col. 9, lines 62-65]. Note: audio frame with time step t is mapped to NUMFRAMES, where NUMFRAMES as taught by Villanueva is three audio frames from t = 0 to t = 2 (see FIG. 2, audio frame 202)).
Villanueva is analogous to the claimed invention, as both relate to generating 3D facial animations from speech audio and style vectors. Villanueva further teaches that “the described systems and methods can be used to generate animations that are more appropriate and realistic for speech audio.” [col. 5, lines 24-26]. Therefore, it would be obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teachings of Villanueva to the combination of Huang and Zhang in order to generate facial animations that are realistic and match the input speech audio.
Regarding claim 17, claim 17 recites substantially similar limitations to claim 7, therefore, will be rejected under the same rationale as claim 7.
Claims 8 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Huang in view of Zhang, and further in view of Tashiro et al. (U.S. 2021/0304478 A1, hereinafter Tashiro).
Regarding claim 8, the combination of Huang and Zhang teaches the method of claim 1, but fails to teach wherein the mesh generation model is trained based on augmented training data that includes 3D animation data generated based on 2D videos. However, this is known in the art as taught by Tashiro.
Tashiro teaches wherein the mesh generation model is trained based on augmented training data that includes 3D animation data generated based on 2D videos ([0016] “The mesh-tracking based dynamic 4D modeling for ML deformation training includes… Machine learning based deformation synthesis and animation using standard MoCAP animation workflow includes using single-view or multi-view 2D videos of MoCAP actors as input”).
Tashiro is analogous to the claimed invention, as both relates to animating multi-dimensional meshes of a human’s face. Tashiro further teaches “Unlike the prior art 3D scan technology, the deformation training implementation described herein is able to generate dynamic face and full-body modeling with implicit deformations by machine learning (ML), that is the synthesis of arbitrary novel action of face expression or body language with natural deformation.” [0012]. Therefore, it would be obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teachings of Tashiro to the combination of Huang and Zhang, as the training as taught by Tashiro would allow for natural and dynamic face expressions.
Regarding claim 18, claim 18 recites substantially similar limitations to claim 8, therefore, will be rejected under the same rationale as claim 8.
Claims 10 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Huang in viewof Zhang, and further in view of Fan et al. (“FaceFormer: Speech-Driven 3D Facial Animation with Transformers”, hereinafter Fan).
Regarding claim 10, the combination of Huang and Zhang teaches the method of claim 1, wherein the mesh generation model includes a vertex decoder (Huang; [0063] “The decoder may then generate an output for use by one or more networks to generate 416 motion vectors for various point associated with the one or more features of the output vector. For example, the neural network may generate a set of motion vectors (or vertices or deformation values, etc.) for one or more facial feature points of one or more facial feature points of the character.” [col. 8, lines 23-25]), but fails to teach the same decoder a transformer-based vertex decoder configured with causal self-attention and cross-modal attention. However, this is known in the art as taught by Fan.
Fan teaches a transformer-based vertex decoder configured with causal self-attention and cross-modal attention ([Figure 2, lines 1-2] “Overview of FaceFormer. An encoder-decoder model with Transformer architecture takes raw audio as input and autoregressively generates a sequence of animated 3D face meshes.” where “Also, we devise two biased attention mechanisms well suited to this specific task, including the biased cross modal multi-head (MH) attention and the biased causal MH self-attention with a periodic positional encoding strategy” [pg. 1, col. 1, Abstract, lines 12-15]).
Fan is analogous to the claimed invention, as both relates to 3D facial mesh animations. Fan further teaches that “the latter offers abilities to generalize to longer audio sequences” [pg. 1, col. 1, Abstract, lines 17-18], as “Prior works typically focus on learning phoneme-level features of short audio windows with limited context, occasionally resulting in inaccurate lip movements” [pg. 1, col. 1, Abstract, lines 3-6]. Therefore, it would be obvious for one of ordinary skill of the art before the effective filing date of the claimed invention to incorporate the teachings of Fan to the combination of Huang and Zhang in order to be able to animate 3D facial motions for longer audios with accurate lip movements.
Regarding claim 20, claim 20 recites substantially similar limitations to claim 10, therefore, will be rejected under the same rationale as claim 10.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALICIA HA whose telephone number is (571)272-3601. The examiner can normally be reached Mon-Thurs 9:30 AM - 6:30 PM, and Fri 9:30 AM - 1:30 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kee Tung can be reached at (571) 272-7794. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ALICIA HA/Examiner, Art Unit 2611
/KEE M TUNG/Supervisory Patent Examiner, Art Unit 2611