Prosecution Insights
Last updated: October 04, 2026
Application No. 18/149,476

REDUCING BIAS IN VISUAL SPEECH RECOGNITION

Non-Final OA §103
Filed
Jan 03, 2023
Examiner
MAC, GARY
Art Unit
2127
Tech Center
2100 — Computer Architecture & Software
Assignee
Technology Innovation Institute - Sole Proprietorship LLC
OA Round
3 (Non-Final)
41%
Grant Probability
Moderate
3-4
OA Rounds
7m
Est. Remaining
79%
With Interview

Examiner Intelligence

Grants 41% of resolved cases
41%
Career Allowance Rate
9 granted / 22 resolved
-14.1% vs TC avg
Strong +38% interview lift
Without
With
+38.3%
Interview Lift
resolved cases with interview
Typical timeline
4y 4m
Avg Prosecution
19 currently pending
Career history
53
Total Applications
across all art units

Statute-Specific Performance

§101
36.9%
-3.1% vs TC avg
§103
43.3%
+3.3% vs TC avg
§102
7.1%
-32.9% vs TC avg
§112
11.4%
-28.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 22 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 06/24/2026 has been entered. Response to Arguments Applicant’s argument filed 06/24/2026 have been fully considered but they are not persuasive regarding 103 rejections. Applicant’s Argument: On page 8-11 of Applicant’s response to rejections under 35 U.S.C. 103, applicant states that the cited references do not teach the claimed invention. Kaliouby generates static facial images and does not teach generating synthetic talking videos with audio or text for speech content. Kumar uses a GAN model to recognize speech from digital video by generating viseme sequence and Kumar does not teach generating synthetic training samples to reduce bias. Yi discloses a deep neural network model for generating talking face videos with personalized head pose and fails to teach generating videos depicting subjects speaking with lip movements characteristics of any demographic group. The Examiner fails to provide proper motivation to combine the cited references. Examiner’s Response: Applicant’s argument is not persuasive. In response to applicant's arguments against the references individually, one cannot show nonobviousness by attacking references individually where the rejections are based on combinations of references. See In re Keller, 642 F.2d 413, 208 USPQ 871 (CCPA 1981); In re Merck & Co., 800 F.2d 1091, 231 USPQ 375 (Fed. Cir. 1986). Claim 1 recites the training of a GAN model to generate synthetic samples for under-represented groups and the synthetic samples are used to train a visual speech recognition model. Kaliouby (par. 33) explicitly discloses the use of a GAN model to generate additional samples of underrepresented demographic groups. Kaliouby is cited to teach the generation of new synthetic training samples of underrepresented groups to mitigate bias in neural network training. Although Kaliouby may not explicitly disclose the GAN generates a video with audio sample, Kaliouby (par. 26 and 69) discloses that training data may include images and voice data collected for multiple users from devices such as video cameras and smartphones. Under the broadest reasonable interpretation, a video can be decomposed into individual static frames depicting an individual. Kaliouby (par. 42) further discloses analyzing the facial features to detect facial expressions, moods, emotions, and cognitive states. Kumar (abstract) discloses the use of a GAN model to recognize speech from a digital video and teaches a visual speech recognition system. Kaliouby discloses the training data can consist of video and audio data and the implementation of a GAN model to be trained on under-represented groups. Kaliouby also discloses the analysis of the facial features of a person. It would be obvious to one of the ordinary skills in the art to combine Kumar and Kaliouby because Kumar improves the GAN model by analyzing the video and audio training data to extract viseme sequences by analyzing the facial features of a person from an under-represented group. Yi (abstract, Figure 4) teaches a model that generates a talking face video by taking audio signal of a source person and a video of a target person as input. Yi discloses that the experiment works well for people of different races and ages. Kaliouby discloses that training data needs to be generated for under-represented demographic groups. Kumar discloses speech recognition from digital videos of a person talking. Kaliouby in combination of Kumar teaches the analysis of digital videos of underrepresented demographic groups for visual speech recognition. It would be obvious to one of the ordinary skills in the art to combine Yi with Kumar and Kaliouby because the extracted speech of a underrepresented demographic group can be used as input to combine with a video of a target person to generate new synthetic training data. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-4, 8-13, 15-16, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Kaliouby (US20220101146A1) in view of Kumar (US20230252993A1) and Yi, “Audio-driven Talking Face Video Generation with Learning-based Personalized Head Pose”. Kaliouby and Kumar references are cited in PTO-892 dated 12/16/2025. Yi reference is cited in PTO-892 dated 04/16/2026. Regarding claim 1, Kaliouby teaches: “A method for generating synthetic samples for training a ” (abstract, [0008], Training of a neural network using deep learning is disclosed using training and synthetic data. The trained neural network can be used in a variety of analysis tasks such as facial recognition, emotion analysis, and/or mental state analysis.) “obtaining a set of training data configured to train the ” ([0026], Facial images are obtained for training the model and may include images captured by a video camera.) “deriving, for each sample included in the set of training data, a number of groups in which the sample corresponds, wherein each group is associated with a corresponding group type” ([0028-0029], The facial images are obtained for training the model and the facial images are partitioned into multiple subgroups. The subgroups can be based on age bands, facial hair, or ethnicity.) “identifying, for each group, whether the group is an under-represented group based on a determination of whether the group is associated with fewer samples than an amount of samples associated with other groups of a same group type” ([0042-0043, 0050, Figure 6], Bias occurs when too few facial images represent one or more demographic groups are present in the training dataset. The detect bias block is used to detect bias in facial images. From Figure 6, the age group 0-17 and 65+ are under-represented because there is not a lot of video data showing these groups.) “generating, by a sample generation model, a plurality of synthetic samples for the identified under-represented groups, wherein the sample generation model is a generative adversarial network (GAN) model trained using data associated with the identified under-represented groups, wherein each synthetic sample comprises a wherein the input data comprises at least one of: a face image and audio associated with an identified under-represented group, a mouth area image and audio associated with an identified under-represented group, a face image and a natural language text associated with an identified under-represented group, a mouth area image and a natural language text associated with an identified under-represented group, a face image and a text-audio pair associated with an identified under-represented group, or a mouth area image and a text-audio pair associated with an identified under-represented group, and wherein the plurality of synthetic samples to be generated for each under-represented group is determined based on a difference between samples associated with the under-represented group and the samples that are associated with the other groups of the same group type” ([0023, 0033, 0043, 0050, 0069], Bias mitigation can include adding additional images to the training dataset by using a GAN to generate synthetic images of underrepresented demographic groups. The GAN model can generate additional images by using real images from a specific demographic as input. The facial image and audio data may be obtained as training data. Data can be collected from multiple people. The claim limitation only requires the teaching of any of the listed input data and the reference teaches input data consisting of a face image and audio. Bias occurs when too few facial images represent one or more demographic groups are present in the training dataset. The generated data can include specific demographics such as race and ethnicity and specific facial characteristics such as hairstyle and facial markings.) “training the ” ([0035-0037], The training dataset is augmented using additional synthetic images. The training dataset may contain real and synthetic images. The augmented training dataset is used to further train the model.) Kaliouby does not explicitly disclose an implementation of “a visual speech recognition (VSR) model”, “the set of training data including a series of samples, with each sample providing video of a subject and corresponding text specifying speech content”, “wherein each synthetic sample comprises a video depicting a subject speaking with lip movements characteristic of an identified under-represented group and corresponding audio or text specifying speech content associated with the video”, and “wherein the VSR is configured to derive speech content from a video input”. However, Kumar discloses in the same field of endeavor: “A method for generating synthetic samples for training a visual speech recognition (VSR) model, the method comprising” (abstract, A GAN is used to recognize speech from a digital video.) “obtaining a set of training data configured to train the VSR model, the set of training data including a series of samples, with each sample providing video of a subject and corresponding text specifying speech content” ([0043, 0088], A digital video is input into the visual speech recognition system. The training dataset may also include utterances of varying lengths and transcribed utterances (text specifying speech content).) “training the VSR model using the set of training data and the plurality of synthetic samples generated by the sample generation model, wherein the VSR is configured to derive speech content from a video input” ([0047, 0079], A visual speech recognition system trains a GAN to generate viseme sequence predictions from digital videos. The generated viseme sequence further trains the discriminator.) It would be obvious to one of the ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teaching of “a visual speech recognition (VSR) model”, “the set of training data including a series of samples, with each sample providing video of a subject and corresponding text specifying speech content”, and “wherein the VSR is configured to derive speech content from a video input” from Kumar into the teaching of Kaliouby. Doing so can improve the performance of a model by implementing a GAN model to provide additional samples for training a system in recognizing speech from a digital video (Kumar, abstract). Kaliouby in view of Kumar does not explicitly disclose an implementation of “wherein each synthetic sample comprises a video depicting a subject speaking with lip movements characteristic of an identified under-represented group and corresponding audio or text specifying speech content associated with the video”. However, Yi discloses in the same field of endeavor: “generating, by a sample generation model, a plurality of synthetic samples for the identified identified and corresponding audio or text specifying speech content associated with the video, ...” ([abstract; pg. 2, col. 1, par. 1-4; pg. 3, col. 1, par. 1-4; pg. 7, Figure 4], The study describes a method of generating high-quality talking face videos using a memory-augmented GAN model. The model can transfer an audio signal of arbitrary source person into a high-quality talking face video of arbitrary target person. The experiment shows that the proposed method works well for people of different ages and races. The talking face generation considers lip motion and face expression.) It would be obvious to one of the ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teaching of “wherein each synthetic sample comprises a video depicting a subject speaking with lip movements characteristic of an identified under-represented group and corresponding audio or text specifying speech content associated with the video” from Yi into the teaching of Kaliouby in view of Kumar. Doing so can improve a generative model in producing additional synthetic data of high-quality talking face videos by inputting an audio signal from a source person and a video of a target person (Yi, abstract). Regarding claim 8: Claim 8 recites an article of manufacture that performs the same process as described in Claim 1. Therefore claim 8 is rejected under the same reasons mention for claim 1. The additional elements of claim 8 is addressed below by Kaliouby: “A non-transitory computer-readable storage medium containing program instructions for a method being executed by an application, the application comprising code for one or more components that are called by the application during runtime, wherein execution of the program instructions by one or more processors of a computer system causes the one or more processors to perform steps comprising” ([0034, 0078], The system can include a computer program product embodied in a non-transitory computer readable medium for machine learning, the computer program product comprising code which causes one or more processors to perform operations disclosed.) Regarding claim 15: Claim 15 recites a process that performs the same method as described in Claim 1. Therefore claim 15 is rejected under the same reasons mention for claim 1. The additional elements of claim 15 is addressed below by Kaliouby: “... the set of training data including a series of samples, with each sample providing video of a subject and corresponding audio and/or text specifying speech content” ([0028, 0069], The facial image and audio data may be obtained as training data. Video and audio data can be collected on multiple people.) “specifying a number of group types and, for each group type, a series of groups” ([0029], The training data can be partitioned into multiple subgroups. The group types may include age, gender, and ethnicity. For age, multiple age bands can also be defined for the group type.) “generating, by a generative adversarial network (GAN) model trained using data associated with the identified under-represented groups, a plurality of synthetic samples for the identified under-represented groups, ..., and wherein the plurality of synthetic samples comprise features associated with one or more identified under-represented groups” ([0033], The GAN may be used to generate additional images of underrepresented demographic groups. The synthetic images may include facial images of more young people showing different facial features.) Regarding claims 2 and 11, Kaliouby teaches: “wherein each group is associated with a corresponding group type, wherein each group type relates to any of: a predicted age, gender, with/without beard or moustache, accent, ethnicity and/or other attributes of a subject depicted in each sample” ([0029], The facial images are obtained for training the model and the facial images are partitioned into multiple subgroups. The subgroups can be based on age bands, facial hair, or ethnicity.) Regarding claims 3 and 12, Kaliouby teaches: “processing each sample with any of an age estimation model, a gender classification model, and a cultural background prediction model to predict a series of attributes of the sample, wherein the series of attributes are used in deriving each group in which the sample corresponds” ([0074-0075], The two or more classifier models can be used to determine demographic data, which may include age, race, gender, or religion.) Regarding claims 4, 13, and 18, Kaliouby teaches: “performing a histogram analysis to determine each under-represented group as being associated with fewer samples than the amount of samples associated with other groups of the same group type” ([0050, Figure 6], From Figure 6, the age group 0-17 and 65+ are under-represented because there is not a lot of video data showing these groups. Plot 620 shows a histogram of the test sample distribution.) Regarding claim 9, Kaliouby teaches: “wherein the plurality of synthetic samples comprise features associated with one or more identified under-represented groups” ([0033], The GAN may be used to generate additional images of underrepresented demographic groups. The synthetic images may include facial images of more young people showing different facial features.) Regarding claims 10 and 16, Kaliouby teaches: “wherein the plurality of synthetic samples to be generated for each under-represented group is determined based on a difference between samples associated with the under-represented group and the samples that are associated with the other groups of the same group type” ([0043, 0050], Bias mitigation can include adding additional images to the training dataset by using a GAN to generate synthetic images of underrepresented demographic groups. Bias occurs when too few facial images represent one or more demographic groups are present in the training dataset.) Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to GARY MAC whose telephone number is (703)756-1517. The examiner can normally be reached Monday - Friday 8:00 AM - 5:00 PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Abdullah Kawsar can be reached at (571) 270-3169. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /GARY MAC/Examiner, Art Unit 2127 /BRENT JOHNSTON HOOVER/Primary Examiner, Art Unit 2127
Read full office action

Prosecution Timeline

Jan 03, 2023
Application Filed
Dec 16, 2025
Non-Final Rejection mailed — §103
Mar 16, 2026
Response Filed
Apr 16, 2026
Final Rejection mailed — §103
Jun 24, 2026
Request for Continued Examination
Jun 26, 2026
Response after Non-Final Action
Sep 04, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12724983
NATURAL LANGUAGE GENERATION USING KNOWLEDGE GRAPH INCORPORATING TEXTUAL SUMMARIES
1y 5m to grant Granted Sep 01, 2026
Patent 12711442
METHODS AND SYSTEMS FOR SERVER FAILURE PREDICTION USING SERVER LOGS
5y 3m to grant Granted Aug 18, 2026
Patent 12699910
EXPLAINABILITY FOR ARTIFICIAL INTELLIGENCE-BASED DECISIONS
3y 11m to grant Granted Aug 04, 2026
Patent 12688426
METHOD AND DEVICE FOR COMPRESSING NEURAL NETWORK
4y 8m to grant Granted Jul 21, 2026
Patent 12626130
METHOD AND DEVICE FOR COMPRESSING NEURAL NETWORK
4y 5m to grant Granted May 12, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
41%
Grant Probability
79%
With Interview (+38.3%)
4y 4m (~7m remaining)
Median Time to Grant
High
PTA Risk
Based on 22 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month