DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Claims 1, 10, and 18 are rejected on the ground of nonstatutory double patenting as being unpatentable over claim 10 of U.S. Patent No. 12,614,329 in view of Kim et al. (US 2025/0265756; hereinafter “Kim”) because the patented claim (including the limitations from the parent, patented claim) teaches each limitation, or a trivial variation, of the current claims. The Kim reference is used to show that an emotion is a type of speaking style.
The following table illustrates a sample mapping of the limitations of claim 1 of the current application when compared against the pertinent limitations of claim 10 (including the limitations of parent claim 1) of the patent. The remaining claims can be mapped in a similar manner. NOTE: in the technology area of computer graphics, different statutory categories (e.g. a system, a processor, a method, a computer-readable medium, etc.) that execute the same computational steps are considered trivial variations of each other.
Current Application
Patent (pertinent limitations)
1. One or more processors, comprising: one or more circuits to:
1. A method, comprising: [see note about statutory categories above]
[claim 1] identify an indication of a speaking style for an animation of a mesh; generate a configuration input for a machine-learning model based at least on the indication of the speaking style;
[claim 1] an emotion vector indicative of one or more emotions associated with the speech … generate an animation … [claim 10] at least one deformable mesh
[claim 1] and generate, using the machine-learning model and based at least on the configuration input and input audio data, a set of vertex deltas
[claim 1] using a neural network and based at least in part on audio data corresponding to speech and an emotion vector … feature position data … rigid transformations of defined rigid portions of the one or more feature points … [claim 10] the one or more feature points include one or more vertices
[claim 1] corresponding to the animation of the mesh, the animation synchronized at least in part with the input audio data
[claim 1] generate an animation of the character appearing to utter the speech [synchronizing the visualization with the audio is implicit]
Although the patented claims recite “emotion,” the Kim reference illustrates that an emotion is a type of speaking style: “the inference module may determine multiple emotion styles” (Kim, para. 68); “the target audio source may be generated by reflecting an emotion style … the emotion animation includes audio or facial expressions reflecting emotions” (Kim, para. 95). Before the effective filing date of the claimed invention, it would have been obvious to one having ordinary skill in the art to use an emotion as a type of speaking style in the patented claims in order to improve the resulting animation.
Claims 1, 2, 4-6, 9, 10, and 17-19 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1, 3, 6-8, 10, 11, 17, 18, and 20 of copending application 18/917,641. Although the claims at issue are not identical, they are not patentably distinct from each other because the copending claims teach each limitation, or a trivial variation, of the current claims. The Examiner acknowledges that the claims of application 18/917,641 appear to be focused on the training of the claimed machine-learning model, whereas the claims of the current application are focused on the use of the claimed machine-learning model. However, as shown below, the copending claims do recite each limitation (or a trivial variation) of the current claims, and therefore the claims are not patentably distinct. In most cases, training a particular machine-learning model is not patentably distinct from using that particular machine-learning model in the same way that it was trained to perform.
The following table illustrates a mapping of the conflicting claims:
Current Application
1
2
4-6
9-10
17-18
19
Copending Application
1
3
6-8
10-11
17-18
20
The following table illustrates a sample mapping of the limitations of claim 1 of the current application when compared against the limitations of claim 1 of the copending application. The remaining claims can be mapped in a similar manner.
Current Application
18/917,641 (pertinent limitations)
1. One or more processors, comprising: one or more circuits to:
1. One or more processors comprising: one or more circuits to:
identify an indication of a speaking style for an animation of a mesh; generate a configuration input for a machine-learning model based at least on the indication of the speaking style;
identify an animation for a mesh … an indication of a speaking style … the machine-learning model generates output vertex deltas for the mesh given an input speaking style
and generate, using the machine-learning model and based at least on the configuration input and input audio data, a set of vertex deltas
the machine-learning model generates output vertex deltas for the mesh given an input speaking style and input audio data.
corresponding to the animation of the mesh, the animation synchronized at least in part with the input audio data
an animation for a mesh corresponding to audio data [synchronizing the visualization with the audio is implicit]
Claim Objections
The claims are objected to because of the following informalities: Claim 15 recites “the machine-learning layer” in line 1, which appears to contain a typographical error. Appropriate correction is required.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1-4, 7-10, 13-16, and 18-20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Li et al. (US 2022/0392131; hereinafter “Li”).
Regarding claim 1, Li discloses One or more processors, comprising: one or more circuits (“a processor executing instructions,” para. 58) to: identify an indication of a speaking style (“an input designating a desired style to use,” para. 22) for an animation of a mesh (“the head to animate may be input as a 3D mesh,” para. 25); generate a configuration input for a machine-learning model based at least on the indication of the speaking style (“a style embedding corresponding to the desired style,” para. 22); and generate, using the machine-learning model and based at least on the configuration input and input audio data (“predict 3D facial landmarks corresponding to the specified style,” para. 22; “predicted 3D facial landmarks may be generated from input speech,” para. 23), a set of vertex deltas corresponding to the animation of the mesh (“facial landmarks may be selected from (or based on) points in the 3D mesh,” para. 25; “predict corresponding facial landmarks that reflect displacement of the initial structure,” para. 42; “the neural network can be trained to predict displacement from the initial structure,” para. 4), the animation synchronized at least in part with the input audio data (“the present audio-driven animation techniques can synchronize multiple components of a talking head,” para. 20).
Regarding claim 2, Li discloses generate a style vector for the configuration input using the indication of the speaking style (“a style embedding corresponding to the desired style,” para. 22).
Regarding claim 3, Li discloses receive the indication of the speaking style in response to an interaction with a graphical element of a graphical user interface (“user input selecting the particular speaking style,” published claim 2).
Regarding claim 4, Li discloses generate a transformed mesh corresponding to at least one frame of the animation by applying the set of vertex deltas to the mesh (“the network can learn to predict deformation or displacement of landmarks … the network can learn to predict 3D facial landmarks that can drive animations,” para. 21; “facial landmarks may be selected from (or based on) points in the 3D mesh,” para. 25).
Regarding claim 7, Li discloses generate a plurality of sets of vertex deltas (“facial landmarks may be selected from (or based on) points in the 3D mesh,” para. 25; “the neural network can be trained to predict displacement from the initial structure,” para. 4) for a plurality of frames of the animation using the machine-learning model (“a plurality of animation frames are generated,” para. 59) and based at least on the configuration input (“an input designating a desired style to use,” para. 22) and respective windows of the input audio data (“an audio feature vector is extracted from an audio chunk from a sliding window of the audio signal,” para. 59).
Regarding claim 8, Li discloses generate the set of vertex deltas by decoding an output of the machine-learning model (“the landmark decoder to predict 3D facial landmarks corresponding to the specified style,” para. 22).
Regarding claim 9, Li discloses wherein the one or more processors are comprised in at least one of: … a system for generating synthetic data … (“generating an animation of a talking head,” abstract).
Regarding claim 10, Li discloses A system, comprising: one or more processors (“a processor executing instructions,” para. 58) to: receive, in response to input to a graphical user interface, an indication of a speaking style (“an input designating a desired style to use,” para. 22; “user input selecting the particular speaking style,” published claim 2) for animating a facial mesh (“the head to animate may be input as a 3D mesh,” para. 25); provide the indication of the speaking style and audio data as input to a machine-learning model (“predict 3D facial landmarks corresponding to the specified style,” para. 22; “predicted 3D facial landmarks may be generated from input speech,” para. 23) to generate a set of vertex deltas for the facial mesh (“facial landmarks may be selected from (or based on) points in the 3D mesh,” para. 25; “predict corresponding facial landmarks that reflect displacement of the initial structure,” para. 42; “the neural network can be trained to predict displacement from the initial structure,” para. 4); and generate at least one frame of an animation using the set of vertex deltas and the facial mesh (“a plurality of animation frames are generated,” para. 59).
Regarding claim 13, Li discloses provide the audio data as input according to a sliding window (“an audio feature vector is extracted from an audio chunk from a sliding window of the audio signal,” para. 59); and generate the animation of the facial mesh to synchronize with the audio data (“the present audio-driven animation techniques can synchronize multiple components of a talking head,” para. 20).
Regarding claim 14, Li discloses present the animation of the facial mesh via the graphical user interface (“a user operating application 107 on client device 105 … a user might upload an audio clip and an image to be animated, and server 120 might return an animated video,” para. 36).
Regarding claim 15, Li discloses wherein the machine-learning layer comprises a set of multilayer perceptron layers and a set of decoder layers (“Single-style landmark decoder 340 may comprise a multilayer perceptron (MLP) network (e.g., three layers, in a simple example),” para. 48), and wherein the one or more processors are to: provide the indication of the speaking style as input to the set of multilayer perceptron layers (“multi-style style encoder 440 may have a similar structure as speech content encoder 430, but followed by a network (e.g., one-layer MLP network) to transform the output into a style embedding,” para. 51); and provide the audio data as input to the set of decoder layers (“single-style landmark decoder 340 may decode the speech content embedding,” para. 48).
Regarding claim 16, Li discloses generate the set of vertex deltas by decoding an output of the machine-learning model (“the landmark decoder to predict 3D facial landmarks corresponding to the specified style,” para. 22).
Regarding claims 18-20, they are rejected using the same citations and rationales described in the rejections of claims 1-3, respectively.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 5, 6, 11, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Li in view of Cudeiro et al. (“Capture, Learning, and Synthesis of 3D Speaking Styles”; hereinafter “Cudeiro”).
Regarding claim 5, Li does not disclose wherein the mesh is a blended mesh, and wherein the one or more circuits are to: generate the blended mesh based at least on a first mesh corresponding to a first identity and a second mesh corresponding to a second identity.
In the same art of audio-driven facial animation, Cudeiro teaches wherein the mesh (“face template mesh,” pg. 2, col. 2, para. 3) is a blended mesh, and wherein the one or more circuits are to: generate the blended mesh based at least on a first mesh corresponding to a first identity and a second mesh corresponding to a second identity (“receives as input a subject-specific template T and the raw audio signal … The speech features and the final convolutional layer are conditioned on the subject labels to learn subject-specific styles when trained across multiple subjects,” pg. 4, sec. 4, paras. 2-4; “Conditioning on different subjects during inference results in different speaking styles … We generate new intermediate speaking styles by convex combinations of conditions,” pg. 7, sec. 7.2, para. 4).
Before the effective filing date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Cudeiro to Li. The motivation would have been “enables the model both to generalize to new subjects not seen during training and to synthesize different speaker styles” (Cudeiro, pg. 2, col. 2, para. 1).
Regarding claim 6, the combination of Li and Cudeiro renders obvious wherein the indication of the speaking style comprises a first weight value for the first identity and a second weight value of the second identity, and wherein the one or more circuits are to: generate the blended mesh further based at least on the first weight value and the second weight value (“VOCA’s compatibility with FLAME allows alteration of the identity-dependent facial shape by adding weighted shape blendshapes from FLAME,” Cudeiro, pg. 4, col. 2, para. 5; “Conditioning on different subjects … We generate new intermediate speaking styles by convex combinations of conditions,” pg. 7, sec. 7.2, paras. 3-4; a “convex combination” by definition is a linear/weighted combination).
Regarding claim 11, Li does not disclose generate the facial mesh based at least on a blend of at least two facial meshes according to the indication of the speaking style.
In the same art of audio-driven facial animation, Cudeiro teaches generate the facial mesh based at least on a blend of at least two facial meshes according to the indication of the speaking style (“receives as input a subject-specific template T and the raw audio signal … The speech features and the final convolutional layer are conditioned on the subject labels to learn subject-specific styles when trained across multiple subjects,” pg. 4, sec. 4, paras. 2-4; “Conditioning on different subjects during inference results in different speaking styles … We generate new intermediate speaking styles by convex combinations of conditions,” pg. 7, sec. 7.2, para. 4).
Before the effective filing date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Cudeiro to Li. The motivation would have been “enables the model both to generalize to new subjects not seen during training and to synthesize different speaker styles” (Cudeiro, pg. 2, col. 2, para. 1).
Regarding claim 17, the combination of Li and Cudeiro renders obvious wherein the system is comprised in at least one of: … a system for generating synthetic data … (“generating an animation of a talking head,” Li, abstract).
Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Li in view of Moser (US 2024/0257431).
Regarding claim 12, Li does not disclose receive the indication of the speaking style in response to a slider input at the graphical user interface.
In the same art of facial animation, Moser teaches receive the indication of the [identity] in response to a slider input at the graphical user interface (“an output image depicting a face that is a blend of characteristics of a plurality of input entities,” abstract; “the amount of blending from each identity may be related to the proximity of the corresponding slider 402 to the corresponding identity icon 406,” para. 216).
Before the effective filing date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the identity blending slider of Moser to the speaking style selection of Li. Note that in Li, the selection of the speaking style corresponds to the selection of a particular identity: “To make style embedding different for each speaking style (e.g., each speaker)” (Li, para. 51). Therefore, applying the identity slider of Moser to Li would render obvious using a slider for selecting an identity/speaking style. The motivation would have been “There is a desire in the field of computer-generated (CG) animation and/or manipulation of facial images … to morph images of the face to some form of blend between two or more identities” (Moser, para. 3).
Pertinent Prior Art
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
Fu et al. (“Mimic: Speaking Style Disentanglement for Speech-Driven 3D Facial Animation”) and Thambiraja et al. (“Imitator: Personalized Speech-driven 3D Facial Animation”) teach inputting a speaking style and an audio signal into a machine-learning model to generate an output animated mesh using vertex displacements.
Additionally, Villanueva Aylagas et al. (US 12,406,419), Villanueva Aylagas et al. (US 12,592,018), Sinha et al. (US 2023/0351662), del Val Santos et al. (US 2023/0123486), and Zhou et al. (US 2021/0233299) each teach inputting a speaking style and an audio signal into a machine-learning model to generate an output animated mesh using vertex displacements.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Ryan McCulley whose telephone number is (571)270-3754. The examiner can normally be reached Monday through Friday, 8:00am - 4:30pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kee Tung can be reached at (571) 272-7794. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/RYAN MCCULLEY/Primary Examiner, Art Unit 2611