Prosecution Insights
Last updated: August 17, 2026
Application No. 18/487,577

PERSONALIZED MACHINE-LEARNED MODEL ENSEMBLES FOR RENDERING OF PHOTOREALISTIC FACIAL REPRESENTATIONS

Non-Final OA §103
Filed
Oct 16, 2023
Examiner
STATZ, BENJAMIN TOM
Art Unit
2611
Tech Center
2600 — Communications
Assignee
Charter Communications Operating LLC
OA Round
3 (Non-Final)
33%
Grant Probability
At Risk
3-4
OA Rounds
0m
Est. Remaining
58%
With Interview

Examiner Intelligence

Grants only 33% of cases
33%
Career Allowance Rate
2 granted / 6 resolved
-28.7% vs TC avg
Strong +25% interview lift
Without
With
+25.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 8m
Avg Prosecution
21 currently pending
Career history
39
Total Applications
across all art units

Statute-Specific Performance

§101
3.6%
-36.4% vs TC avg
§103
65.7%
+25.7% vs TC avg
§102
8.6%
-31.4% vs TC avg
§112
10.7%
-29.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 6 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Applicant’s arguments with respect to claim(s) 1, 11, and 17 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-3, 8, 10, and 17-19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Soares et al. (US 11830182 B1, hereinafter "Soares") in view of Oz et al. (US 20230247180 A1, hereinafter "Oz"), Garrido et al. (US 20180089880 A1; hereinafter "Garrido"), and Chen et al. (US 20230222721 A1; hereinafter "Chen"). Regarding claim 1, Soares discloses: A method, comprising: obtaining, by a computing system comprising one or more processor devices, information descriptive of video data that depicts a face of the particular user (para. 18 (col. 4 lines 37-44) “According to one or more embodiments, training module 122 may train a model, such as a neural network, based on image data from a single subject or multiple subjects. In one or more embodiment, network device may capture image data of a person or people presenting one or more facial expressions. In one or more embodiments, the image data may be in the form of still images, or video images, such as a series of frames.”); processing, by the computing system, the video data with a plurality of machine-learned models (fig. 2 shows a process including training 2 distinct models (elements 220 and 240) as part of generating the final dynamic texture model, para. 22 (col. 5 lines 30-33) “Referring to FIG. 2, a flow diagram is illustrated in which mesh and texture autoencoders are trained from a given sequence”, para. 23-28 (cols. 5-6) explains the process, fig. 6 shows the use of the machine-learned models to render an avatar, para. 43-44 (cols. 9-10) explains the process) of a user-specific model ensemble (models are specific to a user, para. 25 (col. 6 lines 9-11) “In one or more embodiments, the specific user's expression model may be stored for use during runtime”) for photorealistic facial representation (para. 9 (col. 1 lines 56-57) states that the goal is “to generate a photorealistic avatar”) to obtain a corresponding plurality of model outputs (fig. 6 shows the use of the machine-learned models to render an avatar, para. 43-44 (cols. 9-10) explains the process, para. 45 (col. 10 lines 19-37) states that each rendered avatar is unique to a particular user), wherein the plurality of machine-learned models comprises one or more of: a machine-learned mesh representation model trained to generate a Three- Dimensional (3D) polygonal mesh representation of the face of the particular user (fig. 2, expression model 220 based on mesh representation 210, para. 23 (col. 5 lines 38-41) “the mesh and texture autoencoders may be trained from a series of images of one or more users in which the users are providing a particular expression”); a machine-learned texture representation model trained to generate a plurality of textures representative of the face of the particular user (fig. 2, texture model 240, para. 23 (col. 5 lines 38-41) “the mesh and texture autoencoders may be trained from a series of images of one or more users in which the users are providing a particular expression”); or one or more subsurface anatomical representation models trained to generate one or more respective sub-surface model outputs, each comprising a representation of a different sub-surface anatomy of the face of the particular user (para. 33 (col. 7 lines 44-58) “the training module 122 trains a neural network to map facial expression to the blood flow texture, such as a texture autoencoder. As described above, in one or more embodiments, the facial expression may be represented by latent variables descriptive of the 3D geometry of the face presenting the facial expression. In one or more embodiments, at 335, training module 122 generates a texture map from the calculated offset for the one or more facial expressions. The blood flow texture may be a 2D blood flow map that indicates a coloration offset from the albedo texture for the subject. The trained texture autoencoder and corresponding texture decoder may receive as input image data and/or an indication of the expression (e.g., the latent vector described above), and output the 2D blood flow texture map”); and optimizing, by the computing system, at least one machine-learned model of the plurality of machine-learned models based on a loss function that evaluates at least one model output of the plurality of model outputs (Soares para. 12 (col. 2 lines 25-29) “The aim of an autoencoder is to learn a representation for a set of data in an optimized form. A trained autoencoder will have an encoder portion, a decoder portion, and latent variables, which represent the optimized representation of the data”, para. 23 (col. 5 lines 38-41) “the mesh and texture autoencoders may be trained from a series of images of one or more users in which the users are providing a particular expression”; one of ordinary skill in the art would recognize that autoencoder networks are typically optimized using a loss function); based on the information descriptive of the video data, generating, by the computing system, a rendering of a 3D photorealistic representation of the face of the particular user (col. 4 lines 1-12 “Returning to client device 175, avatar module 186 renders an avatar, for example, depicting a user of client device 175 or a user of a device communicating with client device 175. In one or more embodiments, the avatar module 186 obtains a blood flow texture map from network device 100 and renders the avatar based on the blood flow texture map. The avatar may be rendered according to additional data, such as head pose, lighting condition, and a view vector. According to one or more embodiments, the head pose, lighting condition, and view vector may be determined based on data obtained from camera 176, depth sensor 178, and/or other sensors that are part of client device 175.”). Soares does not explicitly teach: identifying, by the computing system, a difference between the information descriptive of the video data and prior video data depicting the face of the particular user, wherein the difference comprises a change of the face of the particular user; and performing the previously listed steps responsive to identifying the change. Oz teaches: identifying, by the computing system, a difference between the information descriptive of the video data and prior video data depicting the face of the particular user, wherein the difference comprises a change of the face of the particular user ([0192] “During the conference call, at the Encoder, an additional mechanism is added. This mechanism, a Change Detector (CD), analyzes the expressions made by the participant user and finds when the participant makes a new expression—one that was not taken into account when the models were created. This can easily be performed as expressions—as mentioned above—can be modeled using a 100-component vector. Therefore, the CD can identify situations when the participant's expressions are composed of vectors that were not used for training. The CD may decide that a new vector detected may require re-modelling if the newly detected vector is distant from all previously used vectors by some metric.”); and updating a facial model responsive to identifying the change ([0192] “The CDs would store these unmodelled conditions and at some time after the conference call ends, would either create new models locally and then update a central server with the new models, or would send the relevant pictures as captured by a camera to the central server which would be able to improve the existing models faster”). Soares and Oz are both analogous to the claimed invention because they are in the same field of 3D facial model generation. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the invention of Soares with the teachings of Oz to only perform the machine-learning model update steps when necessitated by a change in a user’s face. The motivation would have been to increase efficiency by eliminating unnecessary computation. The combination of Soares in view of Oz does not explicitly teach: obtaining, by a computing system comprising one or more processor devices from a computing device associated with a particular user during a teleconference session between the computing device and one or more second computing devices, information descriptive of video data that depicts a face of the particular user. Garrido teaches: obtaining, by a computing system comprising one or more processor devices from a computing device associated with a particular user during a teleconference session between the computing device and one or more second computing devices, information descriptive of video data that depicts a face of the particular user ([0034] “The block 220 displays the real-time video analysis stage by the source device in order to generate avatar data. The video analysis may consist of at least two operations: first, to identify the defining characteristics of the sending user 211 in order to create a user model. Second, to track motions and changes in those characteristics. Tracking information is used to mimic the expressions, movements, and gestures of the sending user 211 by the animated avatar 231 at block 230.” [0036] “In an alternative embodiment, the sending user's identification operation may be performed in parallel with the tracking operation. In other words, while the sending user's identification operation is being performed, the source device communicates the tracking data to the receiving device.”; the remainder of the paragraph provides additional detail; [0039] “Whether the sending user's model is developed in advance of a communication session or in parallel, in order to create a model the video analysis unit may identify facial landmarks of the sending user, identify their physical characteristics (e.g. size, shape, and position), and measure the relative distance and angle between those different components.”). Garrido is analogous to the claimed invention because it is in the same field of 3D facial model generation. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the invention of Soares in view of Oz with the teachings of Garrido to generate a facial model during a communication session, as the facial data is still being analyzed, instead of beforehand. The motivation would have been for user convenience, allowing them to immediately enter a user conference rather than wait until their model has been generated. The combination of Soares in view of Oz and Garrido does not explicitly teach: transmitting, by the computing system during the teleconference session, the rendering of the 3D photorealistic representation of the face of the particular user to one or more second computing devices. Chen teaches: transmitting, by the computing system during the teleconference session, the rendering of the 3D photorealistic representation of the face of the particular user to one or more second computing devices ([0063] “The modified video stream depicting the video conference participant in an avatar form may be transmitted to other video conference participants for display on their local device”). Chen is analogous to the claimed invention because it is in the same field of 3D facial model generation. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the invention of Soares in view of Oz and Garrido with the teachings of Chen to transmit the rendered images to another device. The motivation would have been to offload computation to another device such as a central server. Regarding claim 2, the combination of Soares in view of Oz, Garrido, and Chen teaches: The method of claim 1, wherein the method further comprises: generating, by the computing system, at least one optimized model output with the at least one machine-learned model of the user-specific model ensemble (Soares para. 12 (col. 2 lines 25-29) “The aim of an autoencoder is to learn a representation for a set of data in an optimized form. A trained autoencoder will have an encoder portion, a decoder portion, and latent variables, which represent the optimized representation of the data.”, para. 43-44 (cols. 9-10) describes generating mesh and texture outputs using the optimized latent variables). Regarding claim 3, the combination of Soares in view of Oz, Garrido, and Chen teaches: updating, by the computing system, a user-specific model output repository (Oz [0192] “The CDs would store these unmodelled conditions and at some time after the conference call ends, would either create new models locally and then update a central server with the new models, or would send the relevant pictures as captured by a camera to the central server which would be able to improve the existing models faster”; [0190] suggests that users’ models are stored for significant periods of time across multiple video conference sessions) for photorealistic facial representation (Soares para. 9 (col. 1 lines 56-57) states that the goal is “to generate a photorealistic avatar”) based on the at least one optimized model output (Soares para. 12 (col. 2 lines 25-29) “The aim of an autoencoder is to learn a representation for a set of data in an optimized form. A trained autoencoder will have an encoder portion, a decoder portion, and latent variables, which represent the optimized representation of the data.”, para. 43-44 (cols. 9-10) describes generating mesh and texture outputs using the optimized latent variables), wherein the user-specific model output repository stores an optimized instance of each of the plurality of model outputs (Soares col. 4 lines 65-67 “The result of the training may be a model that provides the blood flow texture maps. The model or models may be stored in model store 145.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the invention of Soares in view of Oz, Garrido, and Chen with the additional teachings of Oz to update a model repository with newer model versions so that the models do not need to be re-updated or generated from scratch, saving time and computation. Regarding claim 8, the combination of Soares in view of Oz, Garrido, and Chen teaches: The method of claim 1, wherein processing the information descriptive of the video data with the plurality of machine-learned models of the user-specific model ensemble for photorealistic facial representation to obtain the corresponding plurality of model outputs comprises: processing, by the computing system, the information descriptive of the video data with a blood-flow mapping model of the one or more subsurface anatomical representation models to obtain a sub-surface model output indicative of a mapping of a blood flow anatomy of the face of the particular user (Soares para. 33 (col. 7 lines 44-58) “the training module 122 trains a neural network to map facial expression to the blood flow texture, such as a texture autoencoder. As described above, in one or more embodiments, the facial expression may be represented by latent variables descriptive of the 3D geometry of the face presenting the facial expression. In one or more embodiments, at 335, training module 122 generates a texture map from the calculated offset for the one or more facial expressions. The blood flow texture may be a 2D blood flow map that indicates a coloration offset from the albedo texture for the subject. The trained texture autoencoder and corresponding texture decoder may receive as input image data and/or an indication of the expression (e.g., the latent vector described above), and output the 2D blood flow texture map”). Regarding claim 10, the combination of Soares in view of Oz, Garrido, and Chen teaches the method of claim 1, wherein the method further comprises: generating, by the computing system, model update information descriptive of optimizations made to the at least one machine-learned model; and transmitting, by the computing system, the model update information to a computing device associated with the particular user (Chen fig. 4, [0051] “In step 420, an electronic version or copy the trained machine learning network may be distributed to multiple client devices. For example, the trained machine learning network may be transmitted to and locally stored on client devices. The machine learning network may be updated and further trained from time to time and the machine learning network may be distributed to a client device 150, 151, and stored locally.”). Chen is analogous to the claimed invention because it is in the same field of 3D avatar generation. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the machine-learned 3D avatar model generation of Soares in view of Oz, Garrido, and Chen with the additional teachings of Chen to add the ability to update machine learning models stored on a user’s personal device. The motivation would have been to reduce the invention’s need for centralized computing hardware, saving on costs and upkeep. Regarding claim 17, it is rejected with the same references, rationale, and motivation to combine as claim 1 because its limitations substantially correspond to the limitations of claim 1, with the additional limitation of: A non-transitory computer-readable storage medium that includes executable instructions (Soares col. 13 lines 18-30 “Storage 865 may include one more non-transitory computer-readable storage mediums including, for example, magnetic disks (fixed, floppy, and removable) and tape, optical media such as CD-ROMs and digital video disks (DVDs), and semiconductor memory devices such as Electrically Programmable Read-Only Memory (EPROM), and Electrically Erasable Programmable Read-Only Memory (EEPROM). Memory 860 and storage 865 may be used to tangibly retain computer program instructions or code organized into one or more modules and written in any desired computer programming language. When executed by, for example, processor 805 such computer program code may implement one or more of the methods described herein.”). Regarding claims 18 and 19, they are rejected with the same references, rationale, and motivation to combine as claims 2 and 3 respectively because their limitations substantially correspond to the limitations of claims 2 and 3 respectively. Claim(s) 4, 6, 7, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Soares (US 11830182 B1) in view of Garrido (US 20180089880 A1), Oz (US 20230247180 A1), and Chen (US 20230222721 A1) as applied to claims 1 and 17 above, and further in view of Shang et al. (US 20210027511 A1; hereinafter "Shang"). Regarding claim 4, the combination of Soares in view of Oz, Garrido, and Chen teaches: The method of claim 1, wherein the video data depicts the face of the particular user performing an expression that is unique to the particular user (col. 4 lines 39-42 “In one or more embodiment, network device may capture image data of a person or people presenting one or more facial expressions.”); and wherein generating the rendering of the 3D photorealistic representation comprises: using, by the computing system, the plurality of model outputs to render the 3D photorealistic representation of the face of the particular user performing the expression unique to the particular user (fig. 2; col. 5 lines 38-41 “According to one or more embodiments, the mesh and texture autoencoders may be trained from a series of images of one or more users in which the users are providing a particular expression.”; col. 5 lines 54-59 “The flowchart continues at 210, where the training module 122 converts the expression images to meshes. Each set of expression images may be converted into an expressive 3D mesh representation using photogrammetry or similar geometry reconstruction method and used to train an expression mesh autoencoder neural network (see block 215)”; col. 6 lines 41-45 “Returning to block 215, after the expression autoencoder is trained, the flowchart also continues at block 245, where a latent network is trained to translate mesh latents to texture latents. According to one or more embodiments, an expression of a user drives the texture of the user.”). The combination of Soares in view of Oz, Garrido, and Chen does not explicitly teach that the expression is a microexpression. Shang teaches detecting microexpressions in input images in the process of generating an animated 2D avatar ([0062] “Process 200 trains (210) a model to generate emotion embeddings. In certain embodiments, emotion embeddings can be used to provide a measure of emotion as an input to an inference engine, rather than determining emotion from landmarks. This can allow the emotional response to be more robust because pixel data can be used to gauge emotion, allowing for the capture of micro-expressions that may not be readily detectable in landmarks or other animation parameters.”). Shang is analogous to the claimed invention because it is in the same field of 3D facial model generation and animation. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the invention of Soares in view of Oz, Garrido, and Chen with the teachings of Shang to add the level of precision and detail required to capture users’ microexpressions and emulate them via the generated avatar. Expressing these small-scale facial movements would add to the believability of the avatar’s movement, furthering the goal of photorealism. Regarding claim 6, the combination of Soares in view of Oz, Garrido, and Chen and further in view of Shang teaches: The method of claim 4, wherein the information descriptive of the video data comprises a plurality of key frames from the video data (Soares para. 23 (col. 5 lines 44-50) “… the training module 122 captures or otherwise obtains expression images. In one or more embodiments, the expression images may be captured as a series of frames, such as a video, or may be captured from still images or the like. The expression images may be acquired from numerous individuals, or a single individual”). Regarding claim 7, the combination of Soares in view of Oz, Garrido, and Chen teaches: The method of claim 4, wherein the information descriptive of the video data comprises a motion capture information derived from the video data (Chen [0054] “At step 450, the machine learning network determines facial expression values such as one or more action unit values with an associated action intensity value. In some embodiments, only an action unit value is determined. For example, an image of a user may depict that the user's eyes are closed, and the user's head is slightly turned to the left. The trained machine learning network may output two pairs of action unit values and corresponding intensity values of 43, 1 and 51, 0.5. Action unit value 43 would indicate that the eyes are closed, and the intensity values 1 would maximum action (i.e., eyes closed all the way). Action unit value 51 would indicate head turned to the left, and the intensity value 0.5 would indicate pronounced action (i.e., head turned half-way to the left).”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the machine-learned 3D avatar model generation of Soares in view of Oz, Garrido, and Chen with the additional teachings of Chen to incorporate Chen’s quantitative motion-capture system for expressions. The motivation would have been to reduce the amount of data necessary to represent an expression animation, benefiting users with poor network connections. Regarding claim 20, it is rejected with the same references, rationale, and motivation to combine as claim 4 because its limitations substantially correspond to the limitations of claim 4. Claim(s) 9 is/are rejected under 35 U.S.C. 103 as being unpatentable over Soares (US 11830182 B1) in view of Garrido (US 20180089880 A1), Oz (US 20230247180 A1), and Chen (US 20230222721 A1) as applied to claim 1 above, and further in view of Raman et al. (Mesh-Tension Driven Expression-Based Wrinkles for Synthetic Faces, hereinafter "Raman"). Regarding claim 9, the combination of Soares in view of Oz, Garrido, and Chen teaches: The method of claim 1, but does not explicitly teach: wherein processing the information descriptive of the video data with the plurality of machine-learned models of the user-specific model ensemble for photorealistic facial representation to obtain the corresponding plurality of model outputs comprises: processing, by the computing system, the information descriptive of the video data with a skin tension mapping model of the one or more subsurface anatomical representation models to obtain a sub-surface model output indicative of a mapping of a skin tension anatomy of the face of the particular user. Raman teaches: wherein processing the information descriptive of the video data with the plurality of machine-learned models of the user-specific model ensemble for photorealistic facial representation to obtain the corresponding plurality of model outputs comprises: processing, by the computing system, the information descriptive of the video data with a skin tension mapping model of the one or more subsurface anatomical representation models to obtain a sub-surface model output indicative of a mapping of a skin tension anatomy of the face of the particular user (pg. 2 col. 1 “Our central idea is to capture complex wrinkling effects for an identity from high-resolution scans of their posed expressions. We store all these possible wrinkles into albedo and displacement textures we refer to as wrinkle maps. At synthesis, for any arbitrary expression beyond those represented in the source scans, we blend between the neutral and wrinkle textures using a notion of the tension in the face mesh to obtain dynamic wrinkling effects.”). Raman and the combination of Soares in view of Oz, Garrido, and Chen are analogous to the claimed invention because they are in the same field of 3D facial model generation using machine learning. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the avatar generation method of Soares in view of Oz, Garrido, and Chen with the skin tension mapping of Raman. The motivation would have been to further improve the photorealism of the generated avatars by modeling detailed skin wrinkles. Claim(s) 11-14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Soares et al. (US 11830182 B1, hereinafter "Soares") in view of Chen et al. (US 20230222721 A1; hereinafter "Chen"), Oz et al. (US 20230247180 A1, hereinafter "Oz"), Garrido et al. (US 20180089880 A1; hereinafter "Garrido"), and Shang et al. (US 20210027511 A1; hereinafter "Shang"). Regarding claim 11, Soares discloses: A computing system, comprising: a memory; and one or more processor devices coupled to the memory (fig. 7, para. 47 (col. 11 lines 48-55) “Storage 755 may store media (e.g., audio, image and video files), computer program instructions or software, preference information, device profile information, pre-generated models, frameworks, and any other suitable data. When executed by processor module 730 and/or graphics hardware 735 such computer program code may implement one or more of the methods described herein (e.g., see FIGS. 1-6)”) configured to: use a plurality of optimized model outputs (para. 12 (col. 2 lines 25-29) “The aim of an autoencoder is to learn a representation for a set of data in an optimized form. A trained autoencoder will have an encoder portion, a decoder portion, and latent variables, which represent the optimized representation of the data”, para. 23 (col. 5 lines 38-41) “the mesh and texture autoencoders may be trained from a series of images of one or more users in which the users are providing a particular expression”, para. 24 (col. 5 lines 59-61) “As part of the training process of the expression mesh autoencoder, mesh latents may be obtained”, para. 27 (col. 6 lines 30-32) “In one or more embodiments, texture latents may be obtained based on the training of the texture autoencoder”) to generate a Three- Dimensional (3D) photorealistic representation (para. 9 (col. 1 lines 56-57) states that the goal is “to generate a photorealistic avatar”) of the face of the particular user (para. 20 (col. 5 lines 1-3) “avatar module 186 renders an avatar, for example, depicting a user of client device 175 or a user of a device communicating with client device 175”), wherein the plurality of optimized model outputs are obtained from a corresponding plurality of machine-learned models (para. 12 (col. 2 lines 25-29) “The aim of an autoencoder is to learn a representation for a set of data in an optimized form. A trained autoencoder will have an encoder portion, a decoder portion, and latent variables, which represent the optimized representation of the data”, para. 23 (col. 5 lines 38-41) “the mesh and texture autoencoders may be trained from a series of images of one or more users in which the users are providing a particular expression”) of a user-specific model ensemble (models are unique to a user, para. 25 (col. 6 lines 9-11) “In one or more embodiments, the specific user's expression model may be stored for use during runtime”) for photorealistic facial representation, and wherein the plurality of machine-learned models comprises one or more of: a machine-learned mesh representation model trained to generate a 3D polygonal mesh representation of the face of the particular user (fig. 2, expression model 220 based on mesh representation 210, para. 22 (col. 5 lines 31-32) “mesh and texture autoencoders are trained from a given sequence”); a machine-learned texture representation model trained to generate a plurality of textures representative of the face of the particular user (fig. 2, texture model 240, para. 22 (col. 5 lines 31-32) “mesh and texture autoencoders are trained from a given sequence”); or one or more subsurface anatomical representation models trained to generate one or more respective sub-surface model outputs, each comprising a representation of a different sub-surface anatomy of the face of the particular user (para. 33 (col. 7 lines 44-58) “the training module 122 trains a neural network to map facial expression to the blood flow texture, such as a texture autoencoder. As described above, in one or more embodiments, the facial expression may be represented by latent variables descriptive of the 3D geometry of the face presenting the facial expression. In one or more embodiments, at 335, training module 122 generates a texture map from the calculated offset for the one or more facial expressions. The blood flow texture may be a 2D blood flow map that indicates a coloration offset from the albedo texture for the subject. The trained texture autoencoder and corresponding texture decoder may receive as input image data and/or an indication of the expression (e.g., the latent vector described above), and output the 2D blood flow texture map”); generate a rendering of the 3D photorealistic representation of the face of the particular user (para. 35 (col. 8 lines 13-14) “Referring to FIG. 5, a flow chart is depicted in which an avatar is rendered utilizing a blood texture map.”; also see fig. 2) performing the expression unique to the particular user (fig. 2; col. 5 lines 38-41 “According to one or more embodiments, the mesh and texture autoencoders may be trained from a series of images of one or more users in which the users are providing a particular expression.”; col. 5 lines 54-59 “The flowchart continues at 210, where the training module 122 converts the expression images to meshes. Each set of expression images may be converted into an expressive 3D mesh representation using photogrammetry or similar geometry reconstruction method and used to train an expression mesh autoencoder neural network (see block 215)”; col. 6 lines 41-45 “Returning to block 215, after the expression autoencoder is trained, the flowchart also continues at block 245, where a latent network is trained to translate mesh latents to texture latents. According to one or more embodiments, an expression of a user drives the texture of the user.”). Soares does not explicitly teach that the computing system is configured to: obtain, from a computing device associated with a particular user, motion capture information indicative of a face of the particular user performing a microexpression unique to the particular user; or generate a rendering of the 3D photorealistic representation of the face of the particular user based on the motion capture information, transmit, during the teleconference session, the rendering of the 3D photorealistic representation of the face of the particular user to the one or more second computing devices of a teleconference session that includes the computing device and the one or more second computing devices. Chen teaches: obtain, from a computing device associated with a particular user ([0053] “In step 440, each frame from the video (or the identified group of pixels) is input into the local version of the machine learning network stored on the client device.”), motion capture information indicative of a face of the particular user performing an expression ([0054] “At step 450, the machine learning network determines facial expression values such as one or more action unit values with an associated action intensity value. In some embodiments, only an action unit value is determined. For example, an image of a user may depict that the user's eyes are closed, and the user's head is slightly turned to the left. The trained machine learning network may output two pairs of action unit values and corresponding intensity values of 43, 1 and 51, 0.5. Action unit value 43 would indicate that the eyes are closed, and the intensity values 1 would maximum action (i.e., eyes closed all the way). Action unit value 51 would indicate head turned to the left, and the intensity value 0.5 would indicate pronounced action (i.e., head turned half-way to the left).”), generate a rendering of the 3D photorealistic representation of the face of the particular user based on the motion capture information ([0061] “At step 540, the system 100 determines facial expression values of each frame of the video stream and applies the facial expression values to the avatar model.”), transmit, during the teleconference session, the rendering of the representation of the face of the particular user to the one or more second computing devices of a teleconference session that includes the computing device and the one or more second computing devices ([0063] “The modified video stream depicting the video conference participant in an avatar form may be transmitted to other video conference participants for display on their local device”). Soares and Chen are both analogous to the claimed invention because they are in the same field of 3D avatar generation using machine learning. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the machine-learned 3D avatar model generation of Soares with two aspects of the invention of Chen: the quantitative motion-capture system for expressions (which would have reduced the amount of data necessary to represent an expression animation, benefiting users with poor network connections), and the ability to directly transmit the modified video feed to other teleconference participants’ devices (which would have provided the option to render the modified video feed containing the avatar on a centralized computing platform, benefiting users with devices that have poor processing power). The combination of Soares in view of Chen does not explicitly teach: identify a difference between the motion capture information and prior motion capture information indicative of the face of the particular user, wherein the difference comprises a change in the face of the particular user; and performing the previously listed steps responsive to the change. Oz teaches: identify a difference between the motion capture information and prior motion capture information indicative of the face of the particular user, wherein the difference comprises a change in the face of the particular user ([0192] “During the conference call, at the Encoder, an additional mechanism is added. This mechanism, a Change Detector (CD), analyzes the expressions made by the participant user and finds when the participant makes a new expression—one that was not taken into account when the models were created. This can easily be performed as expressions—as mentioned above—can be modeled using a 100-component vector. Therefore, the CD can identify situations when the participant's expressions are composed of vectors that were not used for training. The CD may decide that a new vector detected may require re-modelling if the newly detected vector is distant from all previously used vectors by some metric.”); and updating a facial model responsive to identifying the change ([0192] “The CDs would store these unmodelled conditions and at some time after the conference call ends, would either create new models locally and then update a central server with the new models, or would send the relevant pictures as captured by a camera to the central server which would be able to improve the existing models faster”); and updating a facial model responsive to identifying the change ([0192] “The CDs would store these unmodelled conditions and at some time after the conference call ends, would either create new models locally and then update a central server with the new models, or would send the relevant pictures as captured by a camera to the central server which would be able to improve the existing models faster”). Oz is analogous to the claimed invention because it is the same field of 3D facial model generation. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the invention of Soares in view of Chen with the teachings of Oz to only perform the machine-learning model update steps when necessitated by a change in a user’s face. The motivation would have been to increase efficiency by eliminating unnecessary computation. The combination of Soares in view of Chen and Oz does not explicitly teach: obtaining motion capture information during a teleconference session between the computing device and one or more second computing devices. Garrido teaches obtaining information for a facial model during a teleconference session between the computing device and one or more second computing devices ([0034] “The block 220 displays the real-time video analysis stage by the source device in order to generate avatar data. The video analysis may consist of at least two operations: first, to identify the defining characteristics of the sending user 211 in order to create a user model. Second, to track motions and changes in those characteristics. Tracking information is used to mimic the expressions, movements, and gestures of the sending user 211 by the animated avatar 231 at block 230.” [0036] “In an alternative embodiment, the sending user's identification operation may be performed in parallel with the tracking operation. In other words, while the sending user's identification operation is being performed, the source device communicates the tracking data to the receiving device.”; the remainder of the paragraph provides additional detail; [0039] “Whether the sending user's model is developed in advance of a communication session or in parallel, in order to create a model the video analysis unit may identify facial landmarks of the sending user, identify their physical characteristics (e.g. size, shape, and position), and measure the relative distance and angle between those different components.”). Garrido is analogous to the claimed invention because it is in the same field of 3D facial model generation. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the invention of Soares in view of Chen and Oz with the teachings of Garrido to generate a facial model during a communication session, as the facial data is still being analyzed, instead of beforehand. The motivation would have been for user convenience, allowing them to immediately enter a user conference rather than wait until their model has been generated. The combination of Soares in view of Chen, Oz, and Garrido does not explicitly teach that the expression is a microexpression. Shang teaches detecting microexpressions in input images in the process of generating an animated 2D avatar ([0062] “Process 200 trains (210) a model to generate emotion embeddings. In certain embodiments, emotion embeddings can be used to provide a measure of emotion as an input to an inference engine, rather than determining emotion from landmarks. This can allow the emotional response to be more robust because pixel data can be used to gauge emotion, allowing for the capture of micro-expressions that may not be readily detectable in landmarks or other animation parameters.”). Shang is analogous to the claimed invention because it is in the same field of 3D facial model generation and animation. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the invention of Soares in view of Chen, Oz, and Garrido with the teachings of Shang to add the level of precision and detail required to capture users’ microexpressions and emulate them via the generated avatar. Expressing these small-scale facial movements would add to the believability of the avatar’s movement, furthering the goal of photorealism. Regarding claim 12, the combination of Soares in view of Chen, Oz, Garrido, and Shang teaches: The computing system of claim 11, wherein using the plurality of optimized model outputs to generate the 3D photorealistic representation of the face of the particular user comprises: obtaining a model output comprising the 3D polygonal mesh representation of the face of the particular user (Soares fig. 6, para. 44 (col. 10 lines 16-19) “…a 3D mesh 625, which may be output from the expression neural network model 605”); and applying a model output comprising the plurality of textures representative of the face of the particular user to the 3D polygonal mesh representation of the face of the particular user (Soares fig. 6, para. 43 (col. 9 lines 46-48) “FIG. 6 depicts and example flow diagram of generating an avatar utilizing a blood flow map from a trained neural network, such as a texture”, para. 44 (col. 10 lines 10-19) “As described above, the blood flow neural network model may map the expression (e.g., the latent vector representing the expression) to the 2D blood flow map 620. The flow diagram continues at 630, where the avatar module generates the avatar. According to one or more embodiments, the avatar module renders the avatar by applying the 2D blood flow map 620 to a 3D mesh 625, which may be output from the expression neural network model 605”). Regarding claim 13, the combination of Soares in view of Chen, Oz, Garrido, and Shang teaches: The computing system of claim 12, wherein using the plurality of optimized model outputs to generate the 3D photorealistic representation of the face of the particular user further comprises: applying one or more sub-surface model outputs to the 3D polygonal mesh representation of the face of the particular user, wherein each of the one or more sub-surface model outputs represents a different sub-surface anatomy of the face of the particular user (Soares fig. 6, para. 43 (col. 9 lines 46-48) “FIG. 6 depicts and example flow diagram of generating an avatar utilizing a blood flow map from a trained neural network, such as a texture”, para. 44 (col. 10 lines 10-19) “As described above, the blood flow neural network model may map the expression (e.g., the latent vector representing the expression) to the 2D blood flow map 620. The flow diagram continues at 630, where the avatar module generates the avatar. According to one or more embodiments, the avatar module renders the avatar by applying the 2D blood flow map 620 to a 3D mesh 625, which may be output from the expression neural network model 605”). Regarding claim 14, the combination of Soares in view of Chen, Oz, Garrido, and Shang teaches: The computing system of claim 13, wherein generating the rendering of the 3D photorealistic representation of the face of the particular user comprises: obtaining a microexpression animation for the microexpression unique to the particular user (Chen [0055] “At step 460, the system 100 applies the determined action unit value and corresponding intensity value pairs to an avatar model. Blendshapes of the avatar model are then selected based on the determined action unit values.”, Shang teaches the use of microexpressions unique to the particular user as described for claim 11); and animating the 3D polygonal mesh representation of the face of the particular user based on the microexpression animation (Chen [0055] “A 3D animation of the avatar model is then rendered using the selected blendshapes. The selected blend shapes morph or adjust the mesh geometry of the avatar model.”, Shang teaches the use of microexpressions unique to the particular user as described for claim 11). Shang is analogous to the claimed invention because it is in the same field of 3D facial model generation and animation. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the 3D facial expression animation of Soares in view of Chen, Oz, Garrido, and Shang with the additional teachings of Shang to add the level of precision and detail required to capture users’ microexpressions and emulate them via the generated avatar. Expressing these small-scale facial movements would add to the believability of the avatar’s movement, furthering the goal of photorealism. Claim(s) 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Soares (US 11830182 B1) in view of Chen (US 20230222721 A1), Oz (US 20230247180 A1), Garrido (US 20180089880 A1), and Shang (US 20210027511 A1) as applied to claim 14 above, and further in view of Orvalho et al. (US 20210328954 A1, hereinafter "Orvalho"). Regarding claim 15, the combination of Soares in view of Chen, Oz, Garrido, and Shang teaches: the computing system of claim 14, as well as the microexpression unique to the particular user (Soares and Shang, see claim 11), but does not explicitly teach wherein obtaining the microexpression animation for the microexpression unique to the particular user comprises: processing the motion capture information with an animator model of the plurality of machine-learned models of the user-specific model ensemble for photorealistic facial representation to obtain a model output comprising the microexpression animation. Orvalho teaches wherein obtaining the microexpression animation for the microexpression unique to the particular user comprises: processing the motion capture information (fig. 8, transform matrix 802, [0089] “In this example, the transform matrix 802 generates a matrix based on decomposition of the captured data stream (710 in FIG. 7) into symbols and one or more matrices”) with an animator model of the plurality of machine-learned models of the user-specific model ensemble for photorealistic facial representation to obtain a model output comprising the microexpression animation ([0090] “In various embodiments, based on the matrix from the transform matrix 802, a selection is made of one or more applicable base expressions 804 from a set of base expressions and a determination is made of a combination of two or more of those base expressions that most closely mimics the expression of the first user 734 per the data stream 710. Machine learning may be used to fine tune the determination. In various embodiments, expression symbols and weights are determined.”, [0101] “Referring to FIG. 8, the time 814 aspect is identified and may be represented by a plurality of bytes (eight bytes in the example) in order to provide expression stream and corresponding time information for the expression stream for customizing the 3D animatable model.”). Orvalho is analogous to the claimed invention because it pertains to the same issue of capturing and processing a user’s expression based on video data. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the invention of Soares in view of Chen, Oz, Garrido, and Shang with the invention of Orvalho to add a machine learning-based method of translating a user’s motion capture data into a facial animation. The motivation would have been to expand upon and improve Chen’s method of animating an avatar based on motion capture data. Claim(s) 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Soares (US 11830182 B1) in view of Chen (US 20230222721 A1), Oz (US 20230247180 A1), Garrido (US 20180089880 A1), and Shang (US 20210027511 A1) as applied to claim 14 above, and further in view of Beijing Deepscience Technology Co., Ltd. (CN 110992455 B, hereinafter "Beijing Deepscience"). Regarding claim 16, the combination of Soares in view of Chen, Oz, Garrido, and Shang teaches the computing system of claim 14, wherein obtaining the microexpression animation for the microexpression unique to the particular user comprises: retrieving the microexpression from a user-specific model output repository that stores an optimized instance of each of the plurality of model outputs. Soares teaches storing an optimized instance of each of the plurality of model outputs (para. 12 (col. 2 lines 25-29) “The aim of an autoencoder is to learn a representation for a set of data in an optimized form. A trained autoencoder will have an encoder portion, a decoder portion, and latent variables, which represent the optimized representation of the data.”, para. 43-44 (col. 9-10) describes generating mesh and texture outputs using the optimized latent variables) in a model output repository (para. 19 (col. 4 lines 65-67) “The result of the training may be a model that provides the blood flow texture maps. The model or models may be stored in model store 145”). Chen teaches a user-specific avatar repository (fig. 1A, elements 130 “Avatar Model Repository” and 134 “Avatar Model Customization Repository”, [0022] “The avatar model repository may store and/or maintain avatar models for selection and use with the video communication platform 140… The avatar model customization repository 134 may include customizations, style, coloring, clothing, facial feature sizing and other customizations made be a user to a particular avatar”, [0028] “The changes made to the particular avatar are stored or saved in the avatar model customization repository 134”), and retrieving an avatar from the avatar repository ([0033] describes rendering a video using an avatar from the repository, which necessitates retrieval). Shang teaches detecting microexpression information as previously discussed in the rejection of claim 11). The combination of Soares in view of Chen, Oz, Garrido, and Shang does not explicitly teach a repository for animation. Beijing Deepscience teaches a facial animation repository ([n0023] “The facial animation data-driven method provided by the present invention uses real actor performances to create a large facial animation database, which is built based on FACS (Facial Action Coding System), reducing the workload required to create various facial expressions while maintaining the naturalness of human facial expressions and allowing real-time manipulation of data before visualization.”) Soares, Chen, Shang, and Beijing Deepscience are all analogous to the claimed invention because they each pertain to the same issue of capturing user’s facial movements based on video data. Furthermore, the FACS system used by the facial animation repository of Beijing Deepscience is the same system used by Chen to quantify users’ expressions from motion capture data. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the machine learning model output repository of Soares, the user-specific avatar repository and retrieval system of Chen, the microexpressions of Shang, and the facial animation database of Beijing Deepscience to create a repository for machine-learned, user-specific facial animations capable of the precision and detail required to capture microexpressions and generate a corresponding avatar animation. The motivation would have been to create a system to reuse user’s captured and recorded animations whenever possible, reducing the overall amount of computational processing required by the invention. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to BENJAMIN STATZ whose telephone number is (571)272-6654. The examiner can normally be reached Mon-Fri 8am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Tammy Goddard can be reached at (571)272-7773. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /BENJAMIN TOM STATZ/Examiner, Art Unit 2611
Read full office action

Prosecution Timeline

Show 2 earlier events
Oct 27, 2025
Examiner Interview Summary
Oct 27, 2025
Applicant Interview (Telephonic)
Oct 30, 2025
Response Filed
Jan 16, 2026
Final Rejection mailed — §103
Mar 16, 2026
Response after Non-Final Action
Apr 16, 2026
Request for Continued Examination
Apr 19, 2026
Response after Non-Final Action
Aug 07, 2026
Non-Final Rejection mailed — §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
33%
Grant Probability
58%
With Interview (+25.0%)
2y 8m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 6 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month