Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant correctly identified that the claim rejections under 35 U.S.C. 112(b) pertain to claims 2, 8, 14, and 18, not 4, 8, 14, and 18.
Applicant's arguments filed 6/15/2026 concerning 35 U.S.C. 112(b) have been fully considered but they are not persuasive.
Applicant states the following: “In view of these portions and the specification and drawings as a whole, one of ordinary skill in the art would readily apprehend the scope of the rejected claims, as well as the object-to terminology recited thereby. Applicant therefore respectfully submits that such claims are not indefinite (Remarks, 6/15/2026, pg. 7-8).” It is respectfully submitted that the phrase “substantially disentangled” does not indicate what degree of disentanglement is required by the claim, and different reasonable people of ordinary skill in the art would differ on said degree. As such, the rejection under 35 U.S.C. 112(b) is maintained.
Applicant's arguments filed 6/15/2026 concerning 35 U.S.C. 103 have been fully considered but they are not persuasive.
Applicant states the following: “Villanueva fails to describe generation of "joint trajectory data for a character model" in any manner, let alone in the specific manner recited by claim 1. The passage on which the Office relies discloses that its facial animation rig "may be skeletal-based, wherein the controls define the configuration (e.g. rotation) of various joints of a skeleton, and the mesh is generated/deformed based on the configuration of the joints". (Id. at 13/2-6.) However, this facial animation rig is a downstream model that consumes mesh data to deform a mesh; it is not an output generated by the cVAE, and the "configuration (e.g. rotation) of various joints" at a given pose is not a trajectory of joint movement (Remarks, 6/15/2026, pg. 9).” It is respectfully submitted that Villanueva teaches a facial animation rig (element 304) controlled by a machine learning model (element 303). Applicant cites the following portion of Villanueva: “The facial animation rig may be skeletal-based, wherein the controls define the configuration (e.g. rotation) of various joints of a skeleton, and the mesh is generated/deformed based on the configuration of the joints.” This demonstrates that the joints of a skeleton (considered joint data) and the mesh deformation based on the movement of the joints (considered face vertex displacement data) are the direct result of the machine learning output, which satisfies the limitations of the claim as written.
Applicant states the following: “Yu's reverse-diffusion process operates on a noisy latent representation that "may be repeatedly and progressively fed into [the] denoising model ... to gradually remove noise from the latent representation vector based on the conditioning input." (Yu at [0027].) Yu's conditioning input may include multimodal data ("a text prompt, a conditioning image, an attention map, or other conditioning inputs"), but these conditioning inputs only affect or guide Yu's reverse-diffusion process - they are not themselves processed via reverse diffusion. (See Yu at [0027]-[0028] and [0040].) Yu fails to describe processing a combined representation of multimodal input data via reverse diffusion to generate an intermediate representation; rather, the thing processed via reverse diffusion in Yu is the noisy latent representation of image data - and only image data (Remarks, 6/15/2026, pg. 9-10).” It is respectfully submitted that the conditioning input and the noisy latent representation are both inputted into the denoising model, after which reverse diffusion uses the two inputs in order to create an image, creating a combined representation. This input is taken and processed by the training framework as shown in Figure 1 of Yu, providing the desired image as an output. As such, Yu satisfies the limitation as claimed.
In response to applicant’s argument that there is no teaching, suggestion, or motivation to combine the references, the examiner recognizes that obviousness may be established by combining or modifying the teachings of the prior art to produce the claimed invention where there is some teaching, suggestion, or motivation to do so found either in the references themselves or in the knowledge generally available to one of ordinary skill in the art (Remarks, 6/15/2026, pg. 10). See In re Fine, 837 F.2d 1071, 5 USPQ2d 1596 (Fed. Cir. 1988), In re Jones, 958 F.2d 347, 21 USPQ2d 1941 (Fed. Cir. 1992), and KSR International Co. v. Teleflex, Inc., 550 U.S. 398, 82 USPQ2d 1385 (2007). In this case, both references pertain to neural networks. As such, substituting characteristics and practices between the two would be obvious to one of ordinary skill in that art.
Applicant states the following: “For example, with respect to dependent claim 2 (which recites "receiving a fused multimodal representation of the multimodal input data and of a corresponding plurality of substantially disentangled latent representations of input modalities of the multimodal input data"), the Office asserts that Yu's "text prompt (202) and time (304) are considered disentangled." (Office Action at 4, citing Yu at FIG. 3 and [0029]-[0020], [0036].) However, text prompt 202 is merely one of the raw multimodal inputs described in Yu - a natural-language input the user provides to describe the desired image. (Yu at [0030].) It is not a "latent representation" of anything, but rather the raw conditioning text. The same applies to "time 304," which Yu describes as "an input to both the fixed diffusion model 212 and the trainable diffusion model 214 which indicates to the model which iteration of denoising is currently being performed." (Yu at the FIG. 3 discussion.) Time 304 is therefore merely a diffusion iteration index, not a latent representation of an input modality. These are not the features of dependent claim 2 (Remarks, 6/15/2026, pg. 11).” It is respectfully submitted that, as claimed, latent representation is considered to be a broad term which encompasses the time and text prompt input. Paragraph 0046 of Applicant’s Specifications state the following: “As used herein, a latent representation refers to a high-dimensional encoding of input data, which captures the features and underlying structure of the data. In embodiments of techniques described herein, this encoding is used to facilitate various processing tasks such as disentanglement, translation, and reconstruction by the MMC 215. The latent representation serves as a compact and informative summary of the input data, enabling efficient manipulation and analysis of complex multimodal inputs.” As shown in Figure 3, random noise, time, and text prompts serve as input data and are encoded. As such, Yu satisfies the limitation as claimed.
Applicant states the following: “The Office asserts that Villanueva teaches "interaction with an avatar in a virtual digital environment, the avatar having a set of body features and a set of facial features" and for "generating one or more animation sequences for both the set of body features and the set of facial features". (Id.) However, Villanueva fails to describe generating an animation sequence for a set of body features in any manner, let alone in the specific manner recited by claim 13. Villanueva's generative models output facial mesh data and, separately, tongue mesh data, and Villanueva fails to describe any generation of an animation sequence "for both the set of body features and the set of facial features", as recited (Remarks, 6/15/2026, pg. 12).” It is respectfully submitted that the tongue is part of the body and its features satisfy the limitation of body features. As such, Villanueva satisfies the limitation as claimed.
For the reasons stated, the rejection of all the aforementioned claims under 35 U.S.C. 103 are maintained. As such, the rejections of the dependent claims are also maintained.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 2, 8, 14, and 18 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
The term “substantially disentangled latent” in claim 2 is a relative term which renders the claim indefinite. The term “substantially disentangled latent” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. It is unclear what level of disentanglement rises to the level of substantial, and the claim does not have any clear metes and bounds.
Concerning claim 8, 14, and 18, see the rejection of claim 2.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1-3, 5-9, and 11-12 is/are rejected under 35 U.S.C. 103 as being unpatentable over US Publication 2024/0386623 A1 to Yu et al. (hereinafter Yu) in view of US Patent 12406419 B1 to Villanueva Aylagas et al. (hereinafter Villanueva Aylagas).
Concerning claim 1,
Yu discloses receiving a combined representation of multimodal input data based on a plurality of input modalities (0030-0031; 0036; Figure 3);
processing the combined representation via reverse diffusion to generate an intermediate representation (0027); and
iteratively processing the intermediate representation via a U-Net structure (0031).
Yu does not disclose generating, using machine learning, face vertex displacement data and joint trajectory data for a character model.
Villanueva Aylagas teaches generating, using machine learning, face vertex displacement data and joint trajectory data for a character model (Col. 12, ln 64-Col. 13, ln 27).
It would have been obvious for one with ordinary skill in the art before the effective filing date of the claimed invention to incorporate the machine learning animation control from Villanueva Aylagas with the machine learning image generation mechanics from Yu as both concern machine learning generation. The generation and control of character model animations shown in Villanueva Aylagas would make the generation systems of Yu more robust.
Concerning claim 2,
Yu discloses receiving the combined representation of multimodal input data comprises receiving a fused multimodal representation of the multimodal input data and of a corresponding plurality of substantially disentangled latent representations of input modalities of the multimodal input data (0029-0030; 0036; Figure 3; Wherein the encoding of the task instruction (206) to the visual condition (204) is considered a fused multimodal representation and the text prompt (202) and time (304) are considered disentangled. See the 112b rejection above).
Concerning claim 3,
Yu discloses iteratively processing the intermediate representation via a U-Net structure comprises refining the intermediate representation via a control network coupled to the U-Net structure, the control network comprising a decoder network having a plurality of zero-convolution layers (0040-0043; Figure 3; Figure 5).
Concerning claim 5,
Yu discloses processing the combined representation comprises applying a time encoder to the combined representation to incorporate temporal information into the intermediate representation (0040, Figure 3).
Concerning claim 6,
Yu does not disclose generating an animated representation of the character model based at least in part on the face vertex displacement data and the joint trajectory data; and
providing the animated representation to an environment-specific adapter to animate the character model within a virtual digital environment corresponding to the environment-specific adapter.
Villanueva Aylagas teaches generating an animated representation of the character model based at least in part on the face vertex displacement data and the joint trajectory data (Col. 12, ln 64-Col. 13, ln 27); and
providing the animated representation to an environment-specific adapter to animate the character model within a virtual digital environment corresponding to the environment-specific adapter (Col. 3, ln 66-Col. 4, ln 4; Col. 5, ln 7-26; Col. 7, ln 1-18).
Concerning claim 7, see the rejection of claim 1.
Concerning claim 8, see the rejection of claim 2.
Concerning claim 9, see the rejection of claim 3.
Concerning claim 11, see the rejection of claim 5.
Concerning claim 12, see the rejection of claim 6.
Claim(s) 4, 10, 13, and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over US Publication 2024/0386623 A1 to Yu et al. in view of US Patent 12406419 B1 to Villanueva Aylagas et al. and further in view of “Future of NPCs? OpenAI + UnrealEngine5 + Text to speech” by Sushidad (hereinafter Sushidad).
Concerning claim 4,
Yu discloses multimodal input data.
Yu does not disclose receiving generated speech data from a large language model (LLM), the generated speech data being output by the LLM based on input data.
Sushidad teaches receiving generated speech data from a large language model (LLM), the generated speech data being output by the LLM based on input data (0:00-1:00; Description).
It would have been obvious for one with ordinary skill in the art before the effective filing date of the claimed invention to incorporate the LLM integration into an NPC avatar as taught by Sushidad with the machine learning image generation mechanics from Yu as both concern machine learning generation. The LLM integration into an NPC avatar would make the generation systems of Yu more multifaceted.
Concerning claim 13,
Yu discloses receiving multimodal input data comprising a plurality of input modalities (0030-0031; 0036; Figure 3)
providing the multimodal input data as input to one or more neural networks (0030-0031; 0036; Figure 3); and
Yu does not disclose interaction with a non-player character (NPC) in a virtual digital environment, the NPC having a set of body features and a set of facial features
based on output of the one or more neural networks in response to the input data, generating one or more animation sequences for both the set of body features and the set of facial features.
Villanueva Aylagas teaches interaction with an avatar in a virtual digital environment, the avatar having a set of body features and a set of facial features (Col. 12, ln 64-Col. 13, ln 27);
based on output of the one or more neural networks in response to the input data, generating one or more animation sequences for both the set of body features and the set of facial features (Col. 12, ln 64-Col. 13, ln 27).
It would have been obvious for one with ordinary skill in the art before the effective filing date of the claimed invention to incorporate the machine learning animation control from Villanueva Aylagas with the machine learning image generation mechanics from Yu as both concern machine learning generation. The generation and control of character model animations shown in Villanueva Aylagas would make the generation systems of Yu more robust.
Sushidad teaches a non-player character (NPC) (0:00-1:00; Description).
It would be an obvious to try to make the avatar disclosed in Villanueva Aylagas an NPC as taught by Sushidad, as there are only a few known configurations of what an avatar in a game can be (player character, non-player character, etc.). In other words, there are a finite number of identified, predictable solutions, a person of ordinary skill has good reason to pursue the known options within his or her technical grasp. The fact that a combination was obvious to try shows it was obvious under 35 U.S.C. 103.” KSR Int’l Co. v. Teleflex Inc., 127 S.Ct. 1727, 1742, 82 USPQ2d 1385, 1396 (2007).
Concerning claim 10, see the rejection of claim 4.
Concerning claim 17, see the rejection of claim 13.
Claim(s) 14-16 and 18-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over US Publication 2024/0386623 A1 to Yu et al. in view of US Patent 12406419 B1 to Villanueva Aylagas et al. further in view of “Future of NPCs? OpenAI + UnrealEngine5 + Text to speech” by Sushidad and further in view of US Publication 2024/0013464 A1 to Ravichandran et al. (hereinafter Ravichandran).
Concerning claim 14,
Yu discloses generating, via the one or more neural networks, a combined representation of the multimodal input data and the substantially disentangled latent representations (0029-0030; 0036; Figure 3; See the 112b rejection above).
Yu does not disclose disentangling, via the one or more neural networks, a set of encoded latent representations of the plurality of input modalities to generate a substantially disentangled latent representation corresponding to each input modality of the plurality of input modalities;
generating, via the one or more neural networks, speech data for the NPC based on providing the combined representation to a large-language model (LLM).
Ravichandran teaches disentangling, via the one or more neural networks, a set of encoded latent representations of the plurality of input modalities to generate a substantially disentangled latent representation corresponding to each input modality of the plurality of input modalities (0075; Figure 6; See the 112b rejection above).
It would have been obvious for one with ordinary skill in the art before the effective filing date of the claimed invention to incorporate the multimodal disentanglement in the context of machine learning and avatars from Ravichandran with the machine learning image generation mechanics from Yu as both concern machine learning generation. The machine learning disentanglement process of Ravichandran would make the generation systems of Yu more robust.
Sushidad teaches generating, via the one or more neural networks, speech data for the NPC based on providing the combined representation to a large-language model (LLM) (0:00-1:00; Description).
Concerning claim 15,
Yu discloses one or more neural networks and using reverse diffusion and combined representation (0027; 0030-0031; 0036; Figure 3).
Yu does not disclose generating face vertex displacement data and joint trajectory data for the NPC based at least in part on the generated speech data and on the combined representation.
Villanueva Aylagas teaches generating face vertex displacement data and joint trajectory data for the NPC (Col. 12, ln 64-Col. 13, ln 27).
Concerning claim 16,
generating an animated representation of the avatar based at least in part on the face vertex displacement data, the joint trajectory data, and the generated speech data (Col. 3, ln 66-Col. 4, ln 4; Col. 5, ln 7-26; Col. 7, ln 1-18); and
providing the animated representation to one or more environment-specific adapters to animate the avatar within the virtual digital environment (Col. 3, ln 66-Col. 4, ln 4; Col. 5, ln 7-26; Col. 7, ln 1-18).
Concerning claim 18, see the rejection of claim 14.
Concerning claim 19, see the rejection of claim 15.
Concerning claim 20, see the rejection of claim 16.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ISHAYU SINGH whose telephone number is (571)272-3179. The examiner can normally be reached Flex.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Dmitry Suhol can be reached at (571) 272-4430. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/I.S./Examiner, Art Unit 3715
/DMITRY SUHOL/Supervisory Patent Examiner, Art Unit 3715