Prosecution Insights
Last updated: August 17, 2026
Application No. 18/910,636

SYNTHESIZING FACIAL EXPRESSIONS AND SPEECH BASED ON FACIAL MICROMOVEMENTS

Non-Final OA §103
Filed
Oct 09, 2024
Priority
Jul 20, 2022 — provisional 63/390,653 +6 more
Examiner
TUCKER, WESLEY J
Art Unit
Tech Center
Assignee
Apple Inc.
OA Round
1 (Non-Final)
84%
Grant Probability
Favorable
1-2
OA Rounds
1y 2m
Est. Remaining
90%
With Interview

Examiner Intelligence

Grants 84% — above average
84%
Career Allowance Rate
610 granted / 729 resolved
+23.7% vs TC avg
Moderate +6% lift
Without
With
+5.9%
Interview Lift
resolved cases with interview
Typical timeline
3y 0m
Avg Prosecution
18 currently pending
Career history
741
Total Applications
across all art units

Statute-Specific Performance

§101
13.8%
-26.2% vs TC avg
§103
37.3%
-2.7% vs TC avg
§102
37.3%
-2.7% vs TC avg
§112
8.4%
-31.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 729 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 181-184, 187-190, 192-196 and 199-200 are rejected under 35 U.S.C. 103 as being unpatentable over the combination of USPNs 2021/0174934 to Kilmer et al, 2021/0035585 to Gupta, and 2023/0021339 to Bosnak et al. With regard to claim 181, Kilmer discloses a non-transitory computer readable medium containing instructions that when executed by at least one processor cause the at least one processor to perform operations for generating synthesized representations of facial expressions (paragraph [0042], a computer program installed on a computer performs gathering facial images, processing images and outputting image analysis results), the operations comprising: controlling at least one coherent light source in a manner enabling illumination of a portion of a face (paragraphs [0055]-[0057] and Fig. 2, light is projected onto and imaged face); receiving output signals from a light detector, wherein the output signals correspond to reflections of coherent light from the portion of the face (paragraphs [0055]-[0057] and [0062]-[0066], The face is imaged using the light and the speckle pattern projected onto the face); applying speckle analysis on the output signals to determine speckle analysis-based facial skin micromovements (paragraphs [0055]-[0057] and [0062]-[0066], The face is imaged using the light and the speckle pattern projected onto the face); using the determined speckle analysis-based facial skin micromovements to identify at least one change in a facial expression during the time period (paragraphs [0045]-[0050], Facial Action Coding System or FACS is used to determine facial expressions using coded muscle movements in combination with the speckle analysis of tracking movement of facial muscles. The FACS system tracks muscles and determines facial expressions according to the detected facial muscle action units); and Kilmer does not disclose using the determined speckle analysis-based facial skin micromovements to identify at least one word prevocalized or vocalized during a time period. Gupta teaches a similar facial element tracking system to that of Kilmer and further teaches using facial features to perform lip reading i.e. words being vocalized or mouthed (paragraphs [0035], [0058], [0069] and [0076]-[0078]). Therefore it would have been obvious to one of ordinary skill in the art before time of filing to use the lip-reading by tracking facial features of Gupta, in combination with the facial feature tracking of Kilmer in order to determine what a person is saying in addition to facial expression tracking in order to gather more information about the person being monitored. Kilmer and Gupta do not explicitly disclose during the time period, outputting data for causing a virtual representation of the face to mimic the at least one change in the facial expression in conjunction with an audio presentation of the at least one word. Bosnak discloses a system for monitoring and recognizing both a person’s facial expressions and their speech in order to determine emotional metrics (paragraphs [0008]-[0010]) and further teaches outputting an avatar or virtual representation that mimics facial expressions and voice/speech audio (paragraphs [0057]-[0058] and [0107]). Therefore it would have been obvious to one of ordinary skill in the art before time of filing to output a virtual representation or avatar of an imaged person as taught by Bosnak in combination with the facial expression recognition taught by Kilmer and the voice/lip-reading recognition taught by Gupta in order to create an accurate representation of an observed person. With regard to claim 182, Kilmer discloses the non-transitory computer readable medium of claim 18‎1, wherein controlling the at least one coherent light source in a manner enabling illumination of the portion of the face includes projecting a light pattern on the portion of the face (paragraphs [0055]-[0057] and [0062]-[0066], The face is imaged using the light and the speckle pattern projected onto the face). With regard to claim 183, Kilmer discloses the non-transitory computer readable medium of claim 182, wherein the light pattern includes a plurality of spots (paragraphs [0055]-[0057] and [0062]-[0066], The face is imaged using the light and the speckle pattern projected onto the face including spots). With regard to claim 184, Kilmer discloses the non-transitory computer readable medium of claim 182, wherein the portion of the face includes cheek skin (Fig. 2, and (paragraphs [0055]-[0057] and [0062]-[0066], The face is imaged using the light and the speckle pattern projected onto the face including the cheeks of the face). With regard to claim 187, Kilmer discloses the non-transitory computer readable medium of claim 181, wherein the output signals from the light detector emanate from a non-wearable device (Fig. 2, The camera and speckle detection/tracking sensor does not appear to be worn). With regard to claim 188, Kilmer discloses the non-transitory computer readable medium of claim 181, wherein the determined speckle analysis-based facial skin micromovements are associated with recruitment of at least one of: a zygomaticus muscle, an orbicularis oris muscle, a genioglossus muscle, a risorius muscle, or a levator labii superioris alaeque nasi muscle (Figs. 1 and 5, and paragraphs [0045]-[0047] and [0066], Kilmer discloses that many different facial muscles are tracked and the FACS system used assigns values to different facial muscles for detecting facial expressions). With regard to claim 189, the combination of Kilmer, Gupta and Bosnak discloses the non-transitory computer readable medium of claim 181, wherein the at least one change in the facial expression during the period of time includes speech-related facial expressions and non-speech-related facial expressions. Kilmer discloses recognizing specific facial expressions by observing specific facial muscles. Gupta teaches the recognition of speech-related facial expressions through the process of lip-reading facial/mouth element movement. With regard to claim 190, the combination of Kilmer, Gupta and Bosnak discloses the non-transitory computer readable medium of claim 189, and Bosnak discloses wherein the virtual representation of the face is associated with an avatar of an individual from whom the output signals are derived, and wherein mimicking the at least one change in the facial expression includes causing visual changes to the avatar that reflect at least one of the speech-related facial expressions and the non-speech-related facial expressions (paragraphs [0008]-[0010], [0057]-[0058] and [0107], Bosnak discloses a system for monitoring and recognizing both a person’s facial expressions and their speech in order to determine emotional metrics and further teaches outputting an avatar or virtual representation that mimics facial expressions and voice/speech audio. Therefore it would have been obvious to one of ordinary skill in the art before time of filing to output a virtual representation or avatar of an imaged person as taught by Bosnak in combination with the facial expression recognition taught by Kilmer and the voice/lip-reading recognition taught by Gupta in order to create an accurate representation of an observed person. With regard to claim 192, Gupta teaches wherein the audio presentation of the at least one word is based on a recording of an individual (paragraph [0076], Gupta teaches speech-to-text recognition using recorded audio data). With regard to claim 193, Bosnak teaches wherein the audio presentation of the at least one word is based on a synthesized voice (paragraphs [0008]-[0010], [0057]-[0058] and [0107], Bosnak discloses a system for monitoring and recognizing both a person’s facial expressions and their speech in order to determine emotional metrics and further teaches outputting an avatar or virtual representation that mimics facial expressions and voice/speech audio). With regard to claim 194, Bosnak discloses wherein the synthesized voice corresponds with a voice of an individual from whom the output signals are derived (paragraphs [0008]-[0010], [0057]-[0058] and [0107], Bosnak discloses a system for monitoring and recognizing both a person’s facial expressions and their speech in order to determine emotional metrics and further teaches outputting an avatar or virtual representation that mimics facial expressions and voice/speech audio). With regard to claim 195, Bosnak discloses wherein the synthesized voice corresponds with template voice selected by an individual from whom the output signals are derived (paragraphs [0059], [0104]-[0105] and [0109]-[0110], the synthetic voice is created based on the recorded voices and modeled for tone, cadence, speech rate, etc.). With regard to claim 196, Bosnak discloses wherein the operations further include determining an emotional state of an individual from whom the output signals are derived based at least in part on the facial skin micromovements and augmenting the virtual representation of the face to reflect the determined emotional state (paragraphs [0008] and [0016], FACS system is used to monitor facial action micromovements and the monitored movements are used to determine the facial expression and to generate the synthetic avatar to match the emotional state associated with the recognized facial expressions). With regard to claim 199, the discussion of claim 181 applies. The method is disclosed in the operation of the computer program system of claim 181. With regard to claim 200, the discussion of claim 181 applies. The references o Kilmer, Gupta and Bosnak all disclose systems that use processors and cameras to perform the steps discussed in claim 181. Claim 186 is rejected under 35 U.S.C. 103 as being unpatentable over the combination of USPNs 2021/0174934 to Kilmer et al, 2021/0035585 to Gupta, and 2023/0021339 to Bosnak et al. and further in view of 2007/0047768 to Gordon et al. With regard to claim 186, Kilmer, Gupta and Bosnak disclose the non-transitory computer readable medium of claim 181, but do not explicitly disclose wherein the output signals from the light detector emanate from a wearable device. Gordon discloses a system similar to Kilmer that tracks facial movement data and uses speckle patterns on the facial skin (paragraph [0030]). Gordon further teaches that tracking device can be worn by the subject (paragraph [0024] and Figs. 1A and 1B). Therefore it would be obvious to one of ordinary skill in the art before time of filing to use a wearable device as taught by Gordon in combination with the speckle projection and imaging of Kilmer in order to image the user’s face consistently while t user may move their head in order to gather consistently imaged facial motion. Claim 191 is rejected under 35 U.S.C. 103 as being unpatentable over the combination of USPNs 2021/0174934 to Kilmer et al, 2021/0035585 to Gupta, and 2023/0021339 to Bosnak et al. and further in view of 2011/0252344 to Van Os. With regard to claim 191, Bosnak discloses an avatar but does not explicitly disclose wherein the visual changes to the avatar involve changing a color of at least a portion of the avatar. The ability to control the appearance of an avatar such as in a video game or messaging application is very well known in the art. Van Os discloses ana avatar in which the color of the avatar can be changed (paragraph [0042]). Therefore it would have been obvious to one of ordinary skill in the art before time of filing to use a color changeable avatar as taught by Van Os in combination with the avatar generation of Bosnak in order to allow a user to customize the appearance of the avatar. Claims 197-198 are rejected under 35 U.S.C. 103 as being unpatentable over the combination of USPNs 2021/0174934 to Kilmer et al, 2021/0035585 to Gupta, and 2023/0021339 to Bosnak et al. and further in view of 2008/0301577 to Kotlyar With regard to claim 197, Bosnak discloses an avatar for conveying facial expressions but does not explicitly disclose wherein the operations further include receiving a selection of a desired emotional state, and augmenting the virtual representation of the face to reflect the selected emotional state. The act of manipulating an avatar is well known in the art. Kotlyar discloses a dating application that uses avatars to communicate facial expressions and emotions and teaches that the user has control of how the avatar’s represented desired emotional state is conveyed (paragraphs [0004], [0009], [0017], [0038] and [0057]). Therefore it would have been obvious to one of ordinary skill in the art before time of filing to use allow a user to manipulate the facial expressions to reflect a desired emotional state of the avatar as taught by Kotlyar in combination with the avatar representation of Bosnak. With regard to claim 198, the discussion of claim 197 applies. It would have been obvious to one of ordinary skill in the art before time of filing to omit undesired facial expressions as well in order to convey only the desired facial expressions to convey the desired emotion. Allowable Subject Matter Claim 185 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Contact Information Any inquiry concerning this communication or earlier communications from the examiner should be directed to WESLEY J TUCKER whose telephone number is (571)272-7427. The examiner can normally be reached 9AM-5PM Monday-Friday. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, JOHN VILLECCO can be reached at 571-272-7319. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /WESLEY J TUCKER/Primary Examiner, Art Unit 2661
Read full office action

Prosecution Timeline

Oct 09, 2024
Application Filed
Jul 31, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12694651
DATASET-AWARE AND INVARIANT LEARNING FOR FACE RECOGNITION
3y 6m to grant Granted Jul 28, 2026
Patent 12688607
SYSTEM AND METHOD FOR MODEL-FREE, ONE-SHOT OBJECT POSE ESTIMATION VIA COORDINATE REGRESSION
2y 2m to grant Granted Jul 21, 2026
Patent 12679060
PRESS MACHINE AND METHOD OF MONITORING IMAGE OF PRESS MACHINE
3y 0m to grant Granted Jul 14, 2026
Patent 12682681
AGE VERIFICATION
3y 0m to grant Granted Jul 14, 2026
Patent 12682616
TRAINING SYSTEM FOR COMPUTER VISION MODEL
2y 8m to grant Granted Jul 14, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
84%
Grant Probability
90%
With Interview (+5.9%)
3y 0m (~1y 2m remaining)
Median Time to Grant
Low
PTA Risk
Based on 729 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month