Prosecution Insights
Last updated: August 17, 2026
Application No. 18/986,673

ADAPTIVE VISUAL SPEECH RECOGNITION

Non-Final OA §102§103
Filed
Dec 18, 2024
Priority
Jun 18, 2021 — GR 20210100402 +2 more
Examiner
COLUCCI, MICHAEL C
Art Unit
Tech Center
Assignee
DeepMind Technologies Limited
OA Round
1 (Non-Final)
76%
Grant Probability
Favorable
1-2
OA Rounds
1y 5m
Est. Remaining
91%
With Interview

Examiner Intelligence

Grants 76% — above average
76%
Career Allowance Rate
765 granted / 1009 resolved
+15.8% vs TC avg
Strong +15% interview lift
Without
With
+15.2%
Interview Lift
resolved cases with interview
Typical timeline
3y 1m
Avg Prosecution
34 currently pending
Career history
1050
Total Applications
across all art units

Statute-Specific Performance

§101
14.1%
-25.9% vs TC avg
§103
61.2%
+21.2% vs TC avg
§102
8.7%
-31.3% vs TC avg
§112
4.8%
-35.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1009 resolved cases

Office Action

§102 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION Double Patenting The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory obviousness-type double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); and In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on a nonstatutory double patenting ground provided the conflicting application or patent either is shown to be commonly owned with this application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. Effective January 1, 1994, a registered attorney or agent of record may sign a terminal disclaimer. A terminal disclaimer signed by the assignee must fully comply with 37 CFR 3.73(b). Claims 2, 13, and 21 (with 5-7 and 16-18) with dependent claims thereof, are rejected on the ground of nonstatutory obviousness-type double patenting as being unpatentable over claims 1, 13, and 14 and any dependent claims thereof of U.S. Patent No. 12211488. Although the conflicting claims are not identical, they are not patentably distinct from each other because said claims of the instant application includes all of the features of said claims of U.S. Patent No. 12211488. It would have been obvious to one of ordinary skill in the art to omit the step of using ground truth per se and other semantic differences otherwise amounting to a broader representation than the related case, In re Karlson 136 USPQ 184 (1963): "Omission of an element and its function is an obvious expedient if the remaining elements perform the same functions as before" Present invention U.S. Patent No. 12211488 2. (New) A method performed by one or more computers, the method comprising:receiving a video that includes a plurality of video frames that depict a first speaker;obtaining a first learned embedding characterizing the first speaker; andprocessing a first input comprising (i) the video and (ii) the first embedding using a neural network having a plurality of parameters, wherein the neural network is configured to process the video and the first embedding in accordance with trained values of the parameters to generate a speech recognition output that defines text corresponding to speech being spoken by the first speaker in the video. 3. (New) The method of claim 2, wherein the neural network is configured to:generate, from the first embedding, an additional input channel; andcombine the additional channel with one or more of the frames in the video prior to processing the frames in the video to generate the speech recognition output. 4. (New) The method of claim 2, wherein the neural network comprises a plurality of hidden layers, and wherein the neural network is configured to, for at least one of the hidden layers:generate, from the first embedding, an additional hidden channel; andcombine the hidden channel and an output of the hidden layer prior to providing the output for processing by another hidden layer of the visual speech recognition neural network. 5. (New) The method of claim 2, wherein the first learned embedding of the first speaker has been learned on a set of adaptation data. 6. (New) The method of claim 5, further comprising:obtaining pre-trained values for the model parameters that have been determined by training the neural network on training data comprising training examples corresponding to a plurality of speakers that are different from the first speaker, wherein determining the first embedding comprises determining the first embedding using the pre-trained values and the set of adaptation data. 7. (New) The method of claim 6, wherein determining the first embedding comprises:initializing the first embedding; andupdating the first embedding by repeatedly performing operations comprising:processing each of one or more adaptation videos in the adaptation data and the first embedding using the neural network in accordance with current values of the parameters to generate a respective speech recognition output for each of the one or more adaptation videos; and updating the first embedding to minimize the loss function. 8. (New) The method of claim 7, wherein updating the first embedding to minimize the loss function comprises:backpropagating gradients of the loss function through the neural network to determine a gradient of the loss function with respect to the first embedding; andupdating the first embedding using the gradient of the loss function with respect to the first embedding. 9. (New) The method of claim 7, wherein the current values are equal to the pre-trained values and to the trained values and wherein the model parameters are fixed while determining the first embedding. 10. (New) The method of claim 7, wherein the operations further comprise:updating the current values of the parameters of the neural network based on gradients of the loss function with respect to the parameters of the neural network, and wherein the trained values are equal to the current values after determining the first embedding vector. 11. (New) The method of claim 2, further comprising:applying a decoder to the speech recognition output for the video to generate a sequence of one or more words being spoken by the first speaker in the video. 12. (NEw) The method of claim 2, wherein the speech recognition output comprises, for each of the video frames, a respective probability distribution over a vocabulary of text elements. 13. (New) A method performed by one or more computers, the method comprising:receiving a video that includes a plurality of video frames that depict a first speaker;obtaining a first learned embedding characterizing the first speaker; andprocessing a first input comprising (i) the video and (ii) the first embedding using a neural network having a plurality of parameters, wherein the neural network is configured to process the video and the first embedding in accordance with trained values of the parameters to generate a speech recognition output that defines text corresponding to speech being spoken by the first speaker in the video. 14. (New) The system of claim 13, wherein the neural network is configured to:generate, from the first embedding, an additional input channel; andcombine the additional channel with one or more of the frames in the video prior to processing the frames in the video to generate the speech recognition output. 15. (New) The system of claim 13, wherein the neural network comprises a plurality of hidden layers, and wherein the neural network is configured to, for at least one of the hidden layers:generate, from the first embedding, an additional hidden channel; andcombine the hidden channel and an output of the hidden layer prior to providing the output for processing by another hidden layer of the visual speech recognition neural network. 16. (New) The system of claim 13, wherein the first learned embedding of the first speaker has been learned on a set of adaptation data. 17. (New) The system of claim 16, the operations further comprising:obtaining pre-trained values for the model parameters that have been determined by training the neural network on training data comprising training examples corresponding to a plurality of speakers that are different from the first speaker, wherein determining the first embedding comprises determining the first embedding using the pre-trained values and the set of adaptation data. 18. (New) The system of claim 17, wherein determining the first embedding comprises:initializing the first embedding; andupdating the first embedding by repeatedly performing operations comprising:processing each of one or more adaptation videos in the adaptation data and the first embedding using the neural network in accordance with current values of the parameters to generate a respective speech recognition output for each of the one or more adaptation videos; and updating the first embedding to minimize the loss function. 19. (New) The system of claim 18, wherein updating the first embedding to minimize the loss function comprises:backpropagating gradients of the loss function through the neural network to determine a gradient of the loss function with respect to the first embedding; andupdating the first embedding using the gradient of the loss function with respect to the first embedding. 20. (New) The system of claim 19, wherein the current values are equal to the pre-trained values and to the trained values and wherein the model parameters are fixed while determining the first embedding. 21. (New) One or more non-transitory computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:receiving a video that includes a plurality of video frames that depict a first speaker;obtaining a first learned embedding characterizing the first speaker; andprocessing a first input comprising (i) the video and (ii) the first embedding using a neural network having a plurality of parameters, wherein the neural network is configured to process the video and the first embedding in accordance with trained values of the parameters to generate a speech recognition output that defines text corresponding to speech being spoken by the first speaker in the video. 1. (Currently Amended) A method performed by one or more computers, the method comprising:receiving a video that includes a plurality of video frames that depict a first speaker;obtaining a first embedding characterizing the first speaker, comprising:obtaining adaptation data for the first speaker, the adaptation data comprising one or more adaptation videos of the first speaker and a respective ground truth transcription for each of the one or more adaptation videos, anddetermining the first embedding for the first speaker using the adaptation data by minimizing a loss function that measures, for each of the one or more adaptation videos, a respective error between the ground truth transcription of the adaptation video and a respective speech recognition output for the adaptation video; and processing a first input comprising (i) the video and (ii) the first embedding using a visual speech recognition neural network having a plurality of parameters, wherein the visual speech recognition neural network is configured to process the video and the first embedding in accordance with trained values of the parameters to generate a speech recognition output that defines a sequence of one or more words being spoken by the first speaker in the video. 2. (Original) The method of claim 1, wherein the visual speech recognition neural network is configured to:generate, from the first embedding, an additional input channel; andcombine the additional channel with one or more of the frames in the video prior to processing the frames in the video to generate the speech recognition output. 3. (Previously Presented) The method of claim 1, wherein the visual speech recognition neural network comprises a plurality of hidden layers, and wherein the neural network is configured to, for at least one of the hidden layers:generate, from the first embedding, an additional hidden channel; andcombine the hidden channel and an output of the hidden layer prior to providing the output for processing by another hidden layer of the visual speech recognition neural network. 4. (Cancelled) 5. (Currently Amended) The method of claim 4claim 1,further comprising:obtaining pre-trained values for the model parameters that have been determined by training the visual speech recognition neural network on training data comprising training examples corresponding to a plurality of speakers that are different from the first speaker, wherein determining the first embedding comprises determining the first embedding using the pre-trained values and the adaptation data. 6. (Currently Amended) The method of claim 5, wherein determining the first embedding comprises:initializing the first embedding; andupdating the first embedding by repeatedly performing operations comprising:processing each of one or more of the adaptation videos segments in the adaptation data and the first embedding using the visual speech recognition neural network in accordance with current values of the parameters to generate a respective speech recognition output for each of the one or more adaptation videos segments; andupdating the first embedding to minimize [[a]]the loss function that measures, for each of the one or more video segments, a respective error between the ground truth transcription of the video segment and the respective speech recognition output for the video segment. 7. (Currently Amended) The method of claim 6, wherein updating the first embedding to minimize [[a]]the loss function that measures, for each of the one or more video segments, a respective error between the ground truth transcription of the video segment and the respective speech recognition output for the video segment comprises:backpropagating gradients of the loss function through the visual speech recognition neural network to determine a gradient of the loss function with respect to the first embedding; and updating the first embedding using the gradient of the loss function with respect to the first embedding. 8. (Previously Presented) The method of claim 6, wherein the current values are equal to the pre-trained values and to the trained values and wherein the model parameters are fixed while determining the first embedding. 9. (Previously Presented) The method of claim 6, wherein the operations further comprise:updating the current values of the parameters of the visual speech recognition neural network based on gradients of the loss function with respect to the parameters of the visual speech recognition neural network, and wherein the trained values are equal to the current values after determining the first embedding vector. 10. (Previously Presented) The method of claim 1, further comprising:applying a decoder to the speech recognition output for the video to generate the sequence of one or more words being spoken by the first speaker in the video. 11. (Previously Presented) The method of claim 1, wherein the speech recognition output comprises, for each of the video frames, a respective probability distribution over a vocabulary of text elements. 12. (Cancelled) 13. (Currently Amended) One or more computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:receiving a video that includes a plurality of video frames that depict a first speaker;obtaining a first embedding characterizing the first speaker, comprising:obtaining adaptation data for the first speaker, the adaptation data comprising one or more adaptation videos of the first speaker and a respective ground truth transcription for each of the one or more adaptation videos, anddetermining the first embedding for the first speaker using the adaptation data by minimizing a loss function that measures, for each of the one or more adaptation videos, a respective error between the ground truth transcription of the adaptation video and a respective speech recognition output for the adaptation video; and processing a first input comprising (i) the video and (ii) the first embedding using a visual speech recognition neural network having a plurality of parameters, wherein the visual speech recognition neural network is configured to process the video and the first embedding in accordance with trained values of the parameters to generate a speech recognition output that defines a sequence of one or more words being spoken by the first speaker in the video. 14. (Currently Amended) A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to perform operations comprising:receiving a video that includes a plurality of video frames that depict a first speaker;obtaining a first embedding characterizing the first speaker, comprising:obtaining adaptation data for the first speaker, the adaptation data comprising one or more adaptation videos of the first speaker and a respective ground truth transcription for each of the one or more adaptation videos, anddetermining the first embedding for the first speaker using the adaptation data by minimizing a loss function that measures, for each of the one or more adaptation videos, a respective error between the ground truth transcription of the adaptation video and a respective speech recognition output for the adaptation video; and processing a first input comprising (i) the video and (ii) the first embedding using a visual speech recognition neural network having a plurality of parameters, wherein the visual speech recognition neural network is configured to process the video and the first embedding in accordance with trained values of the parameters to generate a speech recognition output that defines a sequence of one or more words being spoken by the first speaker in the video. 15. (New) The system of claim 14, wherein the visual speech recognition neural network is configured to:generate, from the first embedding, an additional input channel; andcombine the additional channel with one or more of the frames in the video prior to processing the frames in the video to generate the speech recognition output. 16. (New) The system of claim 14, wherein the visual speech recognition neural network comprises a plurality of hidden layers, and wherein the neural network is configured to, for at least one of the hidden layers:generate, from the first embedding, an additional hidden channel; andcombine the hidden channel and an output of the hidden layer prior to providing the output for processing by another hidden layer of the visual speech recognition neural network. 17. (New) The system of claim 14, the operations further comprising:obtaining pre-trained values for the model parameters that have been determined by training the visual speech recognition neural network on training data comprising training examples corresponding to a plurality of speakers that are different from the first speaker, wherein determining the first embedding comprises determining the first embedding using the pre-trained values and the adaptation data. 18. (New) The system of claim 17, wherein determining the first embedding comprises:initializing the first embedding; andupdating the first embedding by repeatedly performing updating operations comprising:processing each of one or more of the adaptation videos in the adaptation data and the first embedding using the visual speech recognition neural network in accordance with current values of the parameters to generate a respective speech recognition output for each of the one or more adaptation videos; andupdating the first embedding to minimize the loss function. 19. (New) The system of claim 18, wherein updating the first embedding to minimize the loss function comprises:backpropagating gradients of the loss function through the visual speech recognition neural network to determine a gradient of the loss function with respect to the first embedding; and updating the first embedding using the gradient of the loss function with respect to the first embedding. 20. (New) The system of claim 18, wherein the current values are equal to the pre-trained values and to the trained values and wherein the model parameters are fixed while determining the first embedding. 21. (New) The system of claim 18, wherein the updating operations further comprise:updating the current values of the parameters of the visual speech recognition neural network based on gradients of the loss function with respect to the parameters of the visual speech recognition neural network, and wherein the trained values are equal to the current values after determining the first embedding vector. 22. (New) The system of claim 14, the operations further comprising:applying a decoder to the speech recognition output for the video to generate the sequence of one or more words being spoken by the first speaker in the video. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claims 2-5, 11-16, and 21 rejected under 35 U.S.C. 102(a)(2) as being anticipated by US 20210065712 A1 HOLM; Steffen (hereinafter HOLM). Re claim 2, HOLM teaches 2. (New) A method performed by one or more computers, the method comprising: (fig. 2 and abstract) receiving a video that includes a plurality of video frames that depict a first speaker; (from a video, image frames received for video with a speaker for ASR 0052 with fig. 2 and 0043 0049 with fig. 1b) obtaining a first learned embedding characterizing the first speaker; and (embedded vector for speaker features 0073…sourced from a video, image frames received for video with a speaker for ASR 0052 with fig. 2 and 0043 0049 with fig. 1b) processing a first input comprising (i) the video and (ii) the first embedding using a neural network having a plurality of parameters, wherein the neural network is configured to process the video and the first embedding in accordance with trained values of the parameters to generate a speech recognition output that defines text corresponding to speech being spoken by the first speaker in the video. (a neural network learns from training data using video and speaker data extracted in which words are discovered 0045 and 0007 0011 0055-0058 0073, embedded vector for speaker features 0073…sourced from a video, image frames received for video with a speaker for ASR 0052 with fig. 2 and 0043 0049 with fig. 1b) Re claim 13, this claim has been rejected for teaching a broader, or narrower claim based on general inclusion of hardware alone (e.g. processor, memory, instructions), representation of claim 1 omitting/including hardware for instance, otherwise amounting to a virtually identical scope For instance, see fig. 2 of HOLM which contains the hardware/components Re claim 21, this claim has been rejected for teaching a broader, or narrower claim based on general inclusion of hardware alone (e.g. processor, memory, instructions), representation of claim 1 omitting/including hardware for instance, otherwise amounting to a virtually identical scope For instance, see fig. 2 of HOLM which contains the hardware/components Re claims 3 and 14, HOLM teaches 3. (New) The method of claim 2, wherein the neural network is configured to: generate, from the first embedding, an additional input channel; and (embedding vector for instance, channels such as 0047-0049, and a neural network learns from training data using video and speaker data extracted in which words are discovered 0045 and 0007 0011 0055-0058 0073, embedded vector for speaker features 0073…sourced from a video, image frames received for video with a speaker for ASR 0052 with fig. 2 and 0043 0049 with fig. 1b) combine the additional channel with one or more of the frames in the video prior to processing the frames in the video to generate the speech recognition output. (channels such as 0047-0049, a neural network learns from training data using video and speaker data extracted in which words are discovered 0045 and 0007 0011 0055-0058 0073, embedded vector for speaker features 0073…sourced from a video, image frames received for video with a speaker for ASR 0052 with fig. 2 and 0043 0049 with fig. 1b) Re claims 4 and 15, HOLM teaches 4. (New) The method of claim 2, wherein the neural network comprises a plurality of hidden layers, and wherein the neural network is configured to, for at least one of the hidden layers: generate, from the first embedding, an additional hidden channel; and (a neural network e.g. DNN with layers, learns from training data using video and speaker data extracted in which words are discovered 0045 and 0007 0011 0055-0058 0073, embedded vector for speaker features 0073…sourced from a video, image frames received for video with a speaker for ASR 0052 with fig. 2 and 0043 0049 with fig. 1b) combine the hidden channel and an output of the hidden layer prior to providing the output for processing by another hidden layer of the visual speech recognition neural network. (channels such as 0047-0049, a neural network e.g. DNN with layers, learns from training data using video and speaker data extracted in which words are discovered 0045 and 0007 0011 0055-0058 0073, embedded vector for speaker features 0073…sourced from a video, image frames received for video with a speaker for ASR 0052 with fig. 2 and 0043 0049 with fig. 1b) Re claims 5 and 16, HOLM teaches 5. (New) The method of claim 2, wherein the first learned embedding of the first speaker has been learned on a set of adaptation data. (adaptation, learning, updating, training, etc. e.g. 0060, and a neural network learns from training data using video and speaker data extracted in which words are discovered 0045 and 0007 0011 0055-0058 0073, embedded vector for speaker features 0073…sourced from a video, image frames received for video with a speaker for ASR 0052 with fig. 2 and 0043 0049 with fig. 1b) Re claim 11, HOLM teaches 11. (New) The method of claim 2, further comprising: applying a decoder to the speech recognition output for the video to generate a sequence of one or more words being spoken by the first speaker in the video. (words extracted, via a neural network that learns from training data using video and speaker data extracted in which words are discovered 0045 and 0007 0011 0055-0058 0073, embedded vector for speaker features 0073…sourced from a video, image frames received for video with a speaker for ASR 0052 with fig. 2 and 0043 0049 with fig. 1b) Re claim 12, HOLM teaches 12. (New) The method of claim 2, wherein the speech recognition output comprises, for each of the video frames, a respective probability distribution over a vocabulary of text elements. (acoustic probability e.g. 0059, words extracted, via a neural network that learns from training data using video and speaker data extracted in which words are discovered 0045 and 0007 0011 0055-0058 0073, embedded vector for speaker features 0073…sourced from a video, image frames received for video with a speaker for ASR 0052 with fig. 2 and 0043 0049 with fig. 1b) Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 6 and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over US 20210065712 A1 HOLM; Steffen (hereinafter HOLM) in view of US 20220044687 A1 Perret; Valentin Alain Jean et al. (hereinafter Perret). Re claims 6 and 17, while HOLM teaches a neural network that learns from training data using video and speaker data extracted in which words are discovered using embedded vector for speaker features sourced from a video, image frames received for video with a speaker for ASR, it fails to teach different speakers per se as follows: 6. (New) The method of claim 5, further comprising: obtaining pre-trained values for the model parameters that have been determined by training the neural network on training data comprising training examples corresponding to a plurality of speakers that are different from the first speaker, wherein determining the first embedding comprises determining the first embedding using the pre-trained values and the set of adaptation data. (Perret different speakers using neural networks 0041) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of HOLM to incorporate the above claim limitations as taught by Perret to allow for combining prior art elements according to known methods to yield predictable results such as utilizing different speaker profiles, for a fused deep neural network architecture that can accept a stream of sound as input (e.g., a raw audio waveform) and can output its latent representation in the form of an identity embedding (or vector), which can thus learn a representation directly from a raw waveform with no spectral transformation or language transcription needed, utilizing a speaker inventory that can store and maintain identity embeddings that were generated at any point in time during a conversation. Allowable Subject Matter Claims 7-10 and 18-20 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. After searching through patent and non-patent literature, there was no evidence that there exists a limitation in direct relation or an obvious variant to such limitations as a whole as precisely limited. When searching for a secondary prior art for the limitation as recited in the above claims, the most relevant topics pertained to material from the same Inventor and Assignee but did not teach or suggest the aforementioned complex limitations as a whole as precisely limited. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. US 20220310058 A1 ZHAO; Sheng et al Different target speaker voices Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHAEL C COLUCCI whose telephone number is (571)270-1847. The examiner can normally be reached on M-F 9 AM - 5 PM. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Flanders can be reached at (571)272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MICHAEL COLUCCI/Primary Examiner, Art Unit 2655 (571)-270-1847 Examiner FAX: (571)-270-2847 Michael.Colucci@uspto.gov
Read full office action

Prosecution Timeline

Dec 18, 2024
Application Filed
Jul 20, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12706085
SYSTEMS AND METHODS FOR CONTINUAL LEARNING FOR END TO-END AUTOMATIC SPEECH RECOGNITION
2y 5m to grant Granted Aug 11, 2026
Patent 12688851
VOICE BASED ACTIVATION DETECTION
2y 4m to grant Granted Jul 21, 2026
Patent 12682896
SYSTEM AND METHOD FOR THE GENERATION OF WORKLISTS FROM INTERACTION RECORDINGS
2y 7m to grant Granted Jul 14, 2026
Patent 12664976
QUERY REPLAY FOR PERSONALIZED RESPONSES IN AN LLM POWERED ASSISTANT
2y 6m to grant Granted Jun 23, 2026
Patent 12664977
FLY PARAMETER COMPRESSION AND DECOMPRESSION TO FACILITATE FORWARD AND/OR BACK PROPAGATION AT CLIENTS DURING FEDERATED LEARNING
2y 1m to grant Granted Jun 23, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
76%
Grant Probability
91%
With Interview (+15.2%)
3y 1m (~1y 5m remaining)
Median Time to Grant
Low
PTA Risk
Based on 1009 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month