Prosecution Insights
Last updated: October 02, 2026
Application No. 18/062,239

MULTI-MODAL THREE-DIMENSIONAL FACE MODELING AND TRACKING FOR GENERATING EXPRESSIVE AVATARS

Final Rejection §103
Filed
Dec 06, 2022
Priority
Oct 13, 2022 — RO A-2022-00630
Examiner
RICHER, AARON M
Art Unit
2617
Tech Center
2600 — Communications
Assignee
Microsoft Technology Licensing, LLC
OA Round
2 (Final)
52%
Grant Probability
Moderate
3-4
OA Rounds
0m
Est. Remaining
73%
With Interview

Examiner Intelligence

Grants 52% of resolved cases
52%
Career Allowance Rate
252 granted / 481 resolved
-9.6% vs TC avg
Strong +21% interview lift
Without
With
+20.7%
Interview Lift
resolved cases with interview
Typical timeline
3y 9m
Avg Prosecution
25 currently pending
Career history
506
Total Applications
across all art units

Statute-Specific Performance

§101
9.8%
-30.2% vs TC avg
§103
54.9%
+14.9% vs TC avg
§102
12.7%
-27.3% vs TC avg
§112
19.9%
-20.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 481 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Applicant's arguments filed 21 May 2026 have been fully considered but they are not persuasive. Applicant’s arguments with respect to the prior art have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 1, 2, 11, and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Parra Pozo (U.S. Publication 2022/0413433) in view of Fasogbon (WO 2021/173489). As to claim 1, Parra Pozo discloses a computer system for generating an expressive avatar using multi-modal three-dimensional face modeling and tracking (p. 2, sections 0028-0029; p. 3, section 0037; p. 6, sections 0054-0056; a hologram of a user, reading on an avatar, is generated based on modes such as audio, color images, and depth images to create a 3D mesh model; the hologram includes facial expressions and the user is tracked over time), the computer system comprising: a processor coupled to a storage system that stores instructions (p. 5, section 0049), which, upon execution by the processor, cause the processor to: receive initialization data describing an initial state of a facial model (p. 9, section 0091-p. 10, sections 0094; an initial model, generated from a pre-scan, describing a face and body is generated and received at a program block executing on the processor); receive a plurality of multi-modal data signals; perform a fitting process using the received initialization data and the received plurality of multi-modal data signals (p. 10, section 0095-p. 11, section 0100; based on the initial model and received updates for depth imagery, color imagery, and audio, a process refines the mesh model to more closely resemble/fit the user); and determine a set of parameters based on the fitting process, wherein the determined set of parameters describes an updated state of the facial model (p. 10-11, section 0097; p. 15, sections 0141-0143; as part of the user fitting process, a number of geometry parameters representing the face, such as shape, texture, etc. are determined to represent the updated face mesh model). Parra Pozo discloses a fitting process comprising calculating a combination of loss functions (p. 3, section 0036; p. 13, section 0120; p. 15, sections 0141-0143). Parra Pozo does not disclose, but Fasogbon does disclose the process calculating a combination of weighted loss functions wherein each of the weighted loss functions corresponds to a different one of the plurality of multi-modal data signals (p. 16, section 0049; weighted loss functions corresponding to different modes, for example, geometry, and visual color, are combined to train a system). The motivation for this is to tune the network to both smooth and maintain high level detail, and more accurately produce human faces where missing parts exist. It would have been obvious to one skilled in the art before the effective filing date of the claimed invention to modify Parra Pozo to calculate a combination of weighted loss functions wherein each of the weighted loss functions corresponds to a different one of the plurality of multi-modal data signals in order to tune the network to both smooth and maintain high level detail, and more accurately produce human faces where missing parts exist as taught by Fasogbon. As to claim 2, Parra Pozo discloses wherein performing the fitting process comprises iteratively performing: simulating a measurement using the initialization data; comparing the simulated measurement with an actual measurement derived from the plurality of multi-modal data signals; and updating the initialization data based on the comparison of the simulated measurement and the actual measurement (p. 15, sections 0139-0141; based on the initial image, measurements of 3D geometry are predicted/simulated; the predicted/simulated guess is compared with ground truth for the 3D vertices based on actual position and color measurements from depth and color signals; the model is then trained by updating the prediction/simulation). As to claim 11, see the rejection to claim 1. As to claim 12, see the rejection to claim 2. Claims 3 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Parra Pozo in view of Fasogbon and further in view of Weber (U.S. Publication 2024/0078726). As to claim 3, Parra Pozo discloses wherein the set of parameters is determined based on the updated initialization data of an iteration of the fitting process where the comparison of the simulated measurement and the actual measurement is evaluated based on loss (p. 15, sections 0139-0141; based on the initial image, measurements of 3D geometry are predicted/simulated; the predicted/simulated guess is compared with ground truth for the 3D vertices based on actual position and color measurements from depth and color signals; the model is then trained by updating the prediction/simulation based on loss). Parra Pozo does not explicitly disclose, but Weber does disclose that the loss calculation is whether a loss threshold is satisfied (p. 1, section 0007; p. 4, section 0043-p. 5, section 0046; geometry loss, which characterizes how accurately positions of facial geometry have been measured in an iteration compared to ground truth/actual geometry features is evaluated against a loss threshold to determine whether to stop estimation passes or not). The motivation for this is to generate a model that learns to perform face swapping in a geometrically consistent manner (p. 5, section 0046). It would have been obvious to one skilled in the art before the effective filing date of the claimed invention to modify Parra Pozo and Fasogbon to use satisfaction of a loss threshold as a condition in order to generate a model that learns to perform face swapping in a geometrically consistent manner as taught by Weber. As to claim 13, see the rejection to claim 3. Claims 4, 14, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Parra Pozo in view of Fasogbon and further in view of Huschyn (U.S. Publication 2018/0316860) As to claim 4, Parra Pozo does not disclose, but Huschyn discloses wherein the plurality of multi-modal data signals comprises a first data signal received from an eye camera, a second data signal received from an antenna, and a third data signal received from a microphone (p. 3, section 0032-p. 4, section 0033; p. 5-6, section 0050; input data to determine a face model comes from an eye tracking camera and a microphone; other sensor inputs are also mentioned and the inputs are transferred, in some cases wirelessly, to a server that processes them, necessitating a component acting as an antenna to receive the signals). The motivation for this is to locate who is speaking and transfer that information to a server than can analyze it. It would have been obvious to one skilled in the art before the effective filing date of the claimed invention to modify Parra Pozo and Fasogbon to use a plurality of multi-modal data signals comprising a first data signal received from an eye camera, a second data signal received from an antenna, and a third data signal received from a microphone in order to locate who is speaking and transfer that information to a server than can analyze it as taught by Huschyn. As to claim 14, see the rejection to claim 4. As to claim 20, Parra Pozo discloses a head-mounted display (figs. 2a, 2b; p. 1, sections 0004-0005) for generating an expressive avatar using multi-modal three-dimensional face modeling and tracking (p. 2, sections 0028-0029; p. 3, section 0037; p. 6, sections 0054-0056; a hologram of a user, reading on an avatar, is generated based on modes such as audio, color images, and depth images to create a 3D mesh model; the hologram includes facial expressions and the user is tracked over time), the wearable device comprising: and a processor coupled to a storage system that stores instructions (p. 5, section 0049), which, upon execution by the processor, cause the processor to: receive initialization data describing an initial state of a facial model (p. 9, section 0091-p. 10, sections 0094; an initial model, generated from a pre-scan, describing a face and body is generated and received at a program block executing on the processor); perform a fitting process using the received initialization data and the received plurality of multi-modal data signals (p. 10, section 0095-p. 11, section 0100; based on the initial model and received updates for depth imagery, color imagery, and audio, a process refines the mesh model to more closely resemble/fit the user) by iteratively performing: simulating a measurement using the initialization data; comparing the simulated measurement with an actual measurement derived from the plurality of multi-modal data signals; and updating the initialization data based on the comparison of the simulated measurement and the actual measurement (p. 15, sections 0139-0141; based on the initial image, measurements of 3D geometry are predicted/simulated; the predicted/simulated guess is compared with ground truth for the 3D vertices based on actual position and color measurements from depth and color signals; the model is then trained by updating the prediction/simulation); and determine a set of parameters based on the fitting process, wherein the determined set of parameters describes an updated state of the facial model (p. 10-11, section 0097; p. 15, sections 0141-0143; as part of the user fitting process, a number of geometry parameters representing the face, such as shape, texture, etc. are determined to represent the updated face mesh model). Parra Pozo discloses a fitting process comprising calculating a combination of loss functions (p. 3, section 0036; p. 13, section 0120; p. 15, sections 0141-0143). Parra Pozo does not disclose, but Fasogbon does disclose the process calculating a combination of weighted loss functions wherein each of the weighted loss functions corresponds to a different one of the plurality of multi-modal data signals (p. 16, section 0049; weighted loss functions corresponding to different modes, for example, geometry, and visual color, are combined to train a system). The motivation for this is to tune the network to both smooth and maintain high level detail, and more accurately produce human faces where missing parts exist. It would have been obvious to one skilled in the art before the effective filing date of the claimed invention to modify Parra Pozo to calculate a combination of weighted loss functions wherein each of the weighted loss functions corresponds to a different one of the plurality of multi-modal data signals in order to tune the network to both smooth and maintain high level detail, and more accurately produce human faces where missing parts exist as taught by Fasogbon. Parra Pozo does not disclose, but Huschyn discloses a set of antennas; a set of eye cameras; and a microphone, and receiving a plurality of multi-modal data signals comprising a first data signal from the set of antennas, a second data signal from the set of eye cameras, and a third data signal from the microphone (p. 3, section 0032-p. 4, section 0033; p. 5-6, section 0050; input data to determine a face model comes from an eye tracking camera and a microphone; other sensor inputs are also mentioned and the inputs are transferred, in some cases wirelessly, to a server that processes them, necessitating a component acting as an antenna to receive the signals). Motivation for the combination is given in the rejection to claim 4. Claims 6, 7, 16, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Parra Pozo in view of Fasogbon and further in view of Savvides (U.S. Publication 2024/0320964). As to claim 6, Parra Pozo does not disclose, but Savvides discloses wherein the initialization data comprises a set of initial parameters including an identity parameter describing an initial expression, an expression parameter describing an initial expression, and a pose parameter describing an initial pose of the facial model (p. 1, section 0005; p. 1, section 0012; p. 2 sections 0015-0016; initial parameters for a face with an initial expression including identity, expression, and pose are received and some of the parameters including expression and pose can be modified). The motivation for this is that the changeable initial parameters allow for generation of images with variations in characteristics to enable training a network. It would have been obvious to one skilled in the art before the effective filing date of the claimed invention to modify Parra Pozo and Fasogbon to have the initialization data comprise a set of initial parameters including an identity parameter describing an initial expression, an expression parameter describing an initial expression, and a pose parameter describing an initial pose of the facial model in order to allow for generation of images with variations in characteristics to enable training a network as taught by Savvides. As to claim 7, Parra Pozo does not disclose, but Savvides discloses wherein the determined set of parameters includes the identity parameter of the set of initial parameters (p. 1, section 0005; p. 1, section 0012; p. 2 sections 0015-0016; the identity parameter is preserved through other parameter changes). Motivation for the combination is given in the rejection to claim 6. As to claim 16, see the rejection to claim 6. As to claim 17, see the rejection to claim 7. Claims 8, 9, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Parra Pozo in view of Fasogbon and further in view of Ivanov (U.S. Publication 2022/0091571) and Washington (WO 2020/097505). As to claim 8, Parra Pozo does not disclose, but Ivanov discloses, wherein the plurality of multi-modal data signals comprises a data signal received from a set of antennas (p. 3, section 0059-p. 4, section 0062; capacitance and other electrical characteristics are measured, at least in part by receiving signals by connecting an antenna analyzer to an antenna system), and wherein performing the fitting process includes simulating a capacitance value (p. 1, section 0004; p. 8, section 0105; p. 11, section 0137; a machine learning system predicts/simulates a capacitance value as part of fitting a model to a material). The motivation for this is to more accurately characterize physical properties (p. 1, sections 0005-0007). It would have been obvious to one skilled in the art before the effective filing date of the claimed invention to modify Parra Pozo and Fasogbon to use a data signal received from a set of antennas and have performing the fitting process include simulating a capacitance value in order to more accurately characterize physical properties as taught by Ivanov. Further, Parra Pozo does not disclose, but Washington discloses, that the capacitance value model uses a parallel plate capacitor model (p. 3, section 0011; p. 4-5, section 0016; a neural network is used to calibrate capacitance for a pressure sensor that uses parallel plates). The motivation for this is that such parallel plate models can allow repeated normal force to deform the thin film without failure (see abstract). It would have been obvious to one skilled in the art before the effective filing date of the claimed invention to modify Parra Pozo, Fasogbon, and Ivanov to have the capacitance value model use a parallel plate capacitor model in order to model/calibrate a system that allows repeated normal force to deform the thin film without failure as taught by Washington. As to claim 9, Parra Pozo does not disclose, but Ivanov discloses wherein the storage system stores further instructions, which, upon execution by the processor, cause the processor to: perform a calibration process to map simulated capacitance values to actual capacitance values (p. 11, section 0137; p. 13, sections 0166-0167; models, including capacitance prediction models, are tested and predicted/simulated values are compared/mapped to actual values to determine an error). Motivation for the combination is given in the rejection to claim 8. As to claim 18, see the rejection to claim 8. Conclusion Claims 5, 10, 15, and 19 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. As to claims 5 and 15, Parra Pozo teaches the general concept of fitting based on loss, as discussed in the rejection to claim 3. Parra Pozo does not disclose specifically solving ψ* an argmin function with weighted eyecam, RF, audio, and loss parameters specifically as well as a regularization parameter for enforcing prior constraints, and other art combinable with Parra Pozo does not appear to teach solving such an equation along with the other limitations of claims 5 and 15 and the claims on which they depend. As to claims 10 and 19, Ivanov discloses simulating capacitance value, which would include calculating capacitances for various components. Ivanov does not disclose where this simulation specifically includes partitioning an antenna within the set of antennas into a plurality of antenna triangles, determining a plurality of antenna-face triangle pairs by: for each antenna triangle, determining a face triangle that is closest to the antenna triangle based on a distance metric, and calculating a capacitance for each of the plurality of antenna-face triangle pairs. Further, other art combinable with Parra Pozo and Ivanov does not appear to teach such simulation steps along with the other limitations of claims 10 and 19 and the claims on which they depend. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to AARON M RICHER whose telephone number is (571)272-7790. The examiner can normally be reached 9AM-5PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, King Poon can be reached at (571)272-7440. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /AARON M RICHER/Primary Examiner, Art Unit 2617
Read full office action

Prosecution Timeline

Dec 06, 2022
Application Filed
Feb 25, 2026
Non-Final Rejection mailed — §103
May 21, 2026
Response Filed
Aug 19, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749146
GENERATING IMAGE BLENDING WEIGHTS
4y 9m to grant Granted Sep 29, 2026
Patent 12743852
Prediction of Mechanical Properties of Sedimentary Rocks based on a Grain to Grain Parametric Cohesive Contact Model
2y 5m to grant Granted Sep 22, 2026
Patent 12718460
GENERATION OF CURATED TRAINING DATA FOR DIFFUSION MODELS
3y 9m to grant Granted Aug 25, 2026
Patent 12705817
High Accuracy Texture Filtering in Computer Graphics
7y 4m to grant Granted Aug 11, 2026
Patent 12705695
METHOD TO SELECT RESOLUTION VALUES
3y 10m to grant Granted Aug 11, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
52%
Grant Probability
73%
With Interview (+20.7%)
3y 9m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 481 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month