Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Status
In response to the election filed 5/27/2026, claims 1-7 are withdrawn, claims 13-19 are cancelled, and new claims 20-24 are added. Claims 1-12 and 20-24 are pending, New Claims 20-24 substantially correspond to elected claims 8-12, thus, claims 8-12 and 20-24 are under examination.
Election/Restrictions
Claims 1-7 are withdrawn from further consideration pursuant to 37 CFR 1.142(b), as being drawn to a nonelected invention, there being no allowable generic or linking claim. Applicant timely traversed the restriction (election) requirement in the reply filed on 5/27/2026.
Applicant's election with traverse of Group II (claims 8-12) in the reply filed on 5/27/2026 is acknowledged. The traversal is on the ground(s) that the invention does not pose serious search/examination burden. Applicant’s reply filed 5/27/2026, pp. 7-9. This is not found persuasive.
In response to the argument that Invention Group I and Invention group II are highly integrated, it is respectfully disagreed. The applicant is providing statements regarding the state of the art without any supporting evidence, and argument does not constitute evidence It is customary that models are trained and then release for use, where how the model has been trained does not dictate how the model will be used, and be, the model is used without any knowledge of how the model was trained, and indeed independently of how the model was trained, with several commonly used models which are not open source being distributed and in use without any knowledge to the end user of how the model was trained, the data which was used, and at times, even the architecture and the size of the model, for example the AI models used by Photoshop, the Nano Banana model by Google, etc. Indeed, it is customary for users to employ within a same workflow exchangeable models, workflow and model being separate.
In response to the argument that undue search/examination burden doesn’t exists (Applicant’s reply pp. 8), it is respectfully disagreed. There would be a serious search and/or examination burden when (a) the invention has acquired a separate status in the art in view of their different classification. In this case, Invention Group I is classified in G06T2207/20081 and GO6N3/08-G06N3/0985 whereas Invention Group II is classified in G06T13/40 and G06N3/0475. Thus, the serious search and/or examination burden exists.
In response to the argument that Applicant bears a financial burden for restriction (Applicant’s reply pp. 8-9), According to MPEP, there are two criteria for a proper requirement for restriction between patentably distinct inventions: (A) The inventions must be independent (see MPEP § 802.01, § 806.06, § 808.01) or distinct as claimed (see MPEP § 806.05 - § 806.05(j)); and (B) There would be a serious search and/or examination burden on the examiner if restriction is not required (see MPEP § 803.02, § 808, and § 808.02). MPEP 803(I). Here, the supra responses already show that both criteria (A) and (b) are met. Thus, the restriction requirement is proper.
The requirement is still deemed proper and is therefore made FINAL.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 02/05/2025, 03/17/2025, AND 05/05/2026 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 8-9, 11-12, 20-21, and 23-24 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Cao et al. ("Real-time 3D neural facial animation from binocular video." ACM Transactions on Graphics (TOG) 40.4 (2021): 1-17.; IDS REF) (hereinafter Cao).
Regarding claim 8, Cao discloses A video rendering method, applied to a computer device, the method comprising: (page 87:1, “We present a method for performing real-time facial animation of a 3D avatar from binocular video.”; Data capture setup in Fig. 2)
obtaining a to-be-rendered video comprising a target object; (page 87:1, “We present a method for performing real-time facial animation of a 3D avatar from binocular video."; (a) Data Capture in Fig. 1; 4 FACE ESTIMATION "this section describes how to produce realistic avatar animation by fitting the model to a binocular video…")
PNG
media_image1.png
515
1012
media_image1.png
Greyscale
mapping a facial action of the target object in the to-be-rendered video based on a three- dimensional face model, to obtain an intermediate video comprising a three-dimensional face; (Fig. 4; “The face mesh M and relighted texture T…”;
PNG
media_image2.png
560
569
media_image2.png
Greyscale
4.1 Illumination Invariant Face Model, “Given an input image pair, {𝐼 𝑣 }, our goal is to estimate the full state of the face, comprising the rigid head pose and facial expression parameters. In this work, we use the Deep Appearance Model (DAM) proposed by [Lombardi et al. 2018] as the parametric face model.”)
obtaining a target rendering model associated with the target object; (Fig. 4; Fig. 10; 4.1 Illumination Invariant Face Model, “Given an input image pair, {𝐼 𝑣 }, our goal is to estimate the full state of the face, comprising the rigid head pose and facial expression parameters. In this work, we use the Deep Appearance Model (DAM) proposed by [Lombardi et al. 2018] as the parametric face model”)
and rendering the intermediate video based on the target rendering model, to obtain a target video. (4.1 Illumination Invariant Face Model, "The face mesh 𝑀 and relighted texture 𝑇ˆ 𝑣 are then rendered into the original image space to get a relighted avatar.")
Regarding claim 9, Cao discloses The method according to claim 8, wherein the obtaining the to-be- rendered video comprising a target object comprises: obtaining a virtual object generation model established based on the target object; (Fig.4; 4.1 Illumination Invariant Face Model, ” The head pose, viewpoint vector and face mesh are taken as input to the lighting model G𝜙 to generate gain map 𝐺𝑣 and bias map 𝐵𝑣 , which are used to relight the texture.”; 3.2 Algorithm overview, “First, given a PS-DAM learned from a multi-camera system [Lombardi et al. 2018], an off-line analysis-by-synthesis method is applied on each frame to estimate accurate rigid and non-rigid face parameters in novel lighting environments (ğ4.2)”)
and generating the to-be-rendered video based on the virtual object generation model. (Fig.4; 4.1 Illumination Invariant Face Model, ” The head pose, viewpoint vector and face mesh are taken as input to the lighting model G𝜙 to generate gain map 𝐺𝑣 and bias map 𝐵𝑣 , which are used to relight the texture.”; 3.2 Algorithm overview, “First, given a PS-DAM learned from a multi-camera system [Lombardi et al. 2018], an off-line analysis-by-synthesis method is applied on each frame to estimate accurate rigid and non-rigid face parameters in novel lighting environments (ğ4.2). The pairs of images and face parameters are used as training data to train an encoder for real-time inference (ğ 5.2). Second, in the real-time facial animation stage, we take a user’s new input video and combine a coarse mesh tracking algorithm (ğ5.1) with the encoder obtained from the off-line training step to achieve pixel-precise facial animation in real-time. The environments and lighting of the testing scenario may be different
from those in the training data. To accurately recover the facial motions of testing images with new environments and lighting, we apply a few-shot learning strategy to adapt the encoder to the test images (ğ 5.3).”)
Regarding claim 11, Cao discloses The method according to claim 8, wherein the rendering the intermediate video based on the target rendering model, to obtain the target video comprises:
rendering each frame of the intermediate video based on the target rendering model, to obtain rendered pictures with a same quantity as frames of the intermediate video; and
combining the rendered pictures based on a correspondence between each of the rendered pictures and each frame of the intermediate video, to obtain the target video. (Fig. 1; Fig. 2; Fig. 4;
PNG
media_image3.png
317
719
media_image3.png
Greyscale
3.2 Algorithm overview, “Fig. 2 (b)(c) shows the overview of our algorithm that has two stages. First, given a PS-DAM learned from a multi-camera system [Lombardi et al. 2018], an off-line analysis-by-synthesis method is applied on each frame to estimate accurate rigid and non-rigid face parameters in novel lighting environments (ğ4.2). The pairs of images and face parameters are used as training data to train an encoder for real-time inference (ğ 5.2). Second, in the real-time facial animation stage, we take a user’s new input video and combine a coarse mesh tracking algorithm (ğ5.1) with the encoder obtained from the off-line training step to achieve pixel-precise facial animation in real-time. The environments and lighting of the testing scenario may be different
from those in the training data.”; Examiner’s note: each frame corresponds to a same quantity.)
Regarding claim 12, Cao discloses The method according to claim 8, wherein the mapping the facial action of the target object in the to-be-rendered video based on the three-dimensional face model, to obtain the intermediate video comprising the three-dimensional face comprises:
cropping each frame of the to-be-rendered video, and reserving a facial region of the target object in each frame of the to-be-rendered video, to obtain a face video; and (Fig. 2; Fig.1; Fig.4; page 87:1 “First, we solve for our illumination model’s parameters by applying analysis-by-synthesis on a short video recording.“; page 87:4 “At each frame we can get a pair of images {𝐼𝑣}𝑣∈[0,1]. Each capture lasted about 8 minutes, during which the participant would follow along with a video showing a variety of facial expressions and sentences to read aloud, similar to the procedure in [Lombardi et al. 2018].”)
PNG
media_image4.png
421
857
media_image4.png
Greyscale
mapping the facial action of the target object in the face video based on the three- dimensional face model, to obtain the intermediate video. (4.1 Illumination Invariant Face Model, “Given an input image pair, {𝐼 𝑣 }, our goal is to estimate the full state of the face, comprising the rigid head pose and facial expression parameters. In this work, we use the Deep Appearance Model (DAM) proposed by [Lombardi et al . 2018] as the parametric face model.DAM generates mesh and view-dependent texture as a function of a facial expression code z ∈ R256...To fully describe the appearance effects brought by illumination differences, these gain and bias maps depend on the specific illu-mination condition, rigid head pose, facial expression and viewing direction. Thus, we parameterize these maps using a neural net-work, that takes rigid head pose [r, t], the face mesh 𝑀 and camera viewpoint vector v𝑣 as input:")
Regarding claims 20, 21, 23, and 24, the claims are computer device claims of claims 8, 9, 11, and 12 respectively except a processor, a memory, computer-executable instruction and the computer device (Cao, 6.2, AR/VR Application, Windows machine with 64 cores CPU and 2 Nvidia GTX 2080 GPUs). The claims are similar in scope to claims 8, 9, 11, and 12 respectively and they are rejected under same rationale as claims 8, 9, 11, and 12 respectively.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 10 and 22 are rejected under 35 U.S.C. 103 as being unpatentable over Cao et al. ("Real-time 3D neural facial animation from binocular video." ACM Transactions on Graphics (TOG) 40.4 (2021): 1-17.; IDS REF) (hereinafter Cao) in view of Tang et al. ("Humanoid audio–visual avatar with emotive text-to-speech synthesis." IEEE Transactions on multimedia 10.6 (2008): 969-981.; IDS REF) (hereinafter Tang).
Regarding claim 9, Cao does not disclose The method according to claim 9, wherein the generating the to-be- rendered video based on the virtual object generation model comprises:
obtaining text for generating the to-be-rendered video;
converting the text into a speech of the target object, wherein content of the speech corresponds to content of the text;
obtaining at least one group of lip synchronization parameters based on the speech;
inputting the at least one group of lip synchronization parameters into the virtual object generation model, wherein the virtual object generation model drives a face of a virtual object corresponding to the target object to perform a corresponding action based on the at least one group of lip synchronization parameters, to obtain a virtual video associated with the at least one group of lip synchronization parameters; and
rendering the virtual video to obtain the to-be-rendered video.
Tang teaches The method according to claim 9, wherein the generating the to-be- rendered video based on the virtual object generation model comprises:
obtaining text for generating the to-be-rendered video;
converting the text into a speech of the target object, wherein content of the speech corresponds to content of the text; (Fig. 1; IV. SYSTEM FRAMEWORK AND APPROACHES, "In this framework, an input textual message is first converted into emotive synthetic speech by emotive speech synthesis.)
PNG
media_image5.png
320
626
media_image5.png
Greyscale
obtaining at least one group of lip synchronization parameters based on the speech; (IV. SYSTEM FRAMEWORK AND APPROACHES, “At the same time, a phoneme sequence with exact timing is generated.”
inputting the at least one group of lip synchronization parameters into the virtual object generation model, wherein the virtual object generation model drives a face of a virtual object corresponding to the target object to perform a corresponding action based on the at least one group of lip synchronization parameters, to obtain a virtual video associated with the at least one group of lip synchronization parameters; and
rendering the virtual video to obtain the to-be-rendered video. (IV. SYSTEM FRAMEWORK AND APPROACHES, “The phoneme sequence is then mapped into a viseme sequence that defines the speech gestures. The emotional state decides the facial expressions, which will be combined with the speech gestures synchronously and naturally. The key frame technique [7] is utilized to animate the viseme sequence. Our primary research is focused on emotive speech synthesis, emotional facial expression animation, and the co-articulation of speech gestures and facial expressions.”; A. 3-D Audio–Visual Avatar Modeling, “Our previous work in our lab on 3-D face modeling (iFace) and animation [7], [6] lays a partial foundation of this research. The iFace system provides a research platform for face modeling and animation. It takes as input the Cyberware™ scanner data of a person’s face and fits the data with a generic head model. The output is a customized geometric 3-D face model ready for animation. The generic head model, as shown in Fig. 2, consists of nearly all the components of the head (i.e., face, eyes, ears, teeth, tongue, etc.). The surfaces of these components are approximated by triangular meshes. There are a total of 2240 vertices and 2946 triangles to ensure the closeness of the approximation. In order to customize the generic head model for a particular person, we need to obtain the range map (Fig. 3, left) and texture map (Fig. 3, right) of that person by scanning his or her face using a Cyberware™ scanner. On the face component of the generic head model, 35 feature points are explicitly defined. If we were to unfold the face component onto a 2-D plane, those feature points would triangulate the whole face into multiple local patches. By manually selecting 35 corresponding feature points on the texture map (and thus on the range map), as shown in Fig. 3, right, we can compute the 2-D positions of all the vertices of the face meshes in the range and texture maps. As the range map contains 3-D face geometry of the person, we can then deform the face component of the generic head model by displacing the vertices of the triangular meshes by certain amounts as determined by the corresponding range information collected at the corresponding positions. The remaining components of the generic head model (hair, eyebrows, etc.) are automatically adjusted by shifting, rotating and scaling. Manual adjustments and fine tunes are needed where the scanner data are missing or noisy. Fig. 4, left, shows a personalized head model. In addition, texture can be mapped onto the customized head model to achieve photo-realistic appearance, as shown in Fig. 4, right. Based on the above 3-D face modeling tool, a text-driven audio–visual avatar was built. The Microsoft text-to-speech synthesis engine was used to convert any text input into speech and a phoneme sequence with detailed timing information. Each phoneme in the sequence is converted into a corresponding viseme by table lookup, which comprises a key frame at some proper time instant for animation. The facial shapes of the frames between two successive key frames are generated by interpolation.”)
As both Cao and Tang are from the same field of endeavor, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Cao to include wherein the generating the to-be- rendered video based on the virtual object generation model comprises: obtaining text for generating the to-be-rendered video; converting the text into a speech of the target object, wherein content of the speech corresponds to content of the text; obtaining at least one group of lip synchronization parameters based on the speech; inputting the at least one group of lip synchronization parameters into the virtual object generation model, wherein the virtual object generation model drives a face of a virtual object corresponding to the target object to perform a corresponding action based on the at least one group of lip synchronization parameters, to obtain a virtual video associated with the at least one group of lip synchronization parameters; and rendering the virtual video to obtain the to-be-rendered video, in the context of facial animation rendering, according to the teaching of Tang, in order to combine speech gesture (i.e., lip movements due to speech production) and facial expressions naturally and realistically (I. Introduction of Tang).
Regarding claim 22, claim 22 is the computer device claim of claim 10 except computer-executable instruction and a computer device (Cao, 6.2, AR/VR Application, Windows machine with 64 cores CPU and 2 Nvidia GTX 2080 GPUs) and is accordingly rejected under same rationale.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Hyorim Park whose telephone number is (571)272-3859. The examiner can normally be reached Monday - Friday.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alicia Harrington can be reached at (571) 272-2330. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Hyorim Park/Examiner, Art Unit 2615
/ALICIA M HARRINGTON/Supervisory Patent Examiner, Art Unit 2615