Prosecution Insights
Last updated: August 17, 2026
Application No. 18/815,537

TWO-PASS END TO END SPEECH RECOGNITION

Final Rejection §102§103
Filed
Aug 26, 2024
Priority
Dec 04, 2019 — provisional 62/943,703 +2 more
Examiner
COLUCCI, MICHAEL C
Art Unit
2655
Tech Center
2600 — Communications
Assignee
Google LLC
OA Round
2 (Final)
76%
Grant Probability
Favorable
3-4
OA Rounds
1y 2m
Est. Remaining
91%
With Interview

Examiner Intelligence

Grants 76% — above average
76%
Career Allowance Rate
765 granted / 1009 resolved
+13.8% vs TC avg
Strong +15% interview lift
Without
With
+15.2%
Interview Lift
resolved cases with interview
Typical timeline
3y 1m
Avg Prosecution
34 currently pending
Career history
1050
Total Applications
across all art units

Statute-Specific Performance

§101
14.1%
-25.9% vs TC avg
§103
61.2%
+21.2% vs TC avg
§102
8.7%
-31.3% vs TC avg
§112
4.8%
-35.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1009 resolved cases

Office Action

§102 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION The terminal disclaimer filed on 06/06/2026 disclaiming the terminal portion of any patent granted on this application which would extend beyond the expiration date of U.S. Patent No. 12073824 has been reviewed and is accepted. The terminal disclaimer has been recorded. Response to Arguments Applicant's arguments with respect to claims 1, 10, and 19 have been considered but are moot in view of the new ground(s) of rejection. Applicant’s arguments are directed to the amended subject matter; new prior art citations and explanation using existing art of Phillips is provided below. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1, 10, and 19 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by US 20110054899 A1 Phillips; Michael S. et al. (hereinafter Phillips). Re claim 1, Phillips teaches 1. A method implemented by one or more processors, the method comprising: (fig. 1 professor on hardware) training a two-pass automatic speech recognition (ASR) model to generate a text representation of a spoken utterance, wherein training the ASR model comprises: (ASR operations using speech recognition models for training using past utterances, user history, context, and corrections thereof for instance 0071 fig. 1 and fig. 2… e.g. the first pass and second pass involving a user command with corrections or two separate interactions, both are used to train the model 0105 and fig. 7b-7c) training a first-pass portion of the ASR model, wherein training the first-pass portion of the ASR model comprises updating one or more portions of a shared encoder portion of the ASR model (the first pass is a literal first interaction with an ASR model by a user such as a first command 0105 and fig. 7b-7c, an ASR model containing encoded data or an encoder per se as in 0177 and 0185, ASR operations using speech recognition models for training using past utterances, user history, context, and corrections thereof for instance 0071 fig. 1 and fig. 2) AND/OR updating one or more portions of a recurrent neural network transformer (RNN-T) decoder portion of the first-pass portion of the ASR model based on processing a plurality of training instances; and training a second-pass portion of the ASR model, wherein training the second-pass portion of the ASR model comprises updating one or more portions of an additional encoder portion of the ASR model (the second pass can be a literal second interaction with an ASR model by a user such as the correction to a first input of a first command 0105 and fig. 7b-7c, an ASR model containing encoded data or an encoder per se as in 0177 and 0185, ASR operations using speech recognition models for training using past utterances, user history, context, and corrections thereof for instance 0071 fig. 1 and fig. 2) AND/OR updating one or more portions of a listen attend spell (LAS) decoder based on processing the plurality of training instances; subsequent to training the ASR model, processing audio data capturing a spoken utterance using the ASR model to generate a text representation of the spoken utterance, (ASR inherently and explicitly in Phillips, converts speech to text, particularly after iterations of model training, e.g. the first pass and second pass involving a user command 0105 and fig. 7b-7c, an ASR model containing encoded data or an encoder per se as in 0177 and 0185, ASR operations using speech recognition models for training using past utterances, user history, context, and corrections thereof for instance 0071 fig. 1 and fig. 2) , wherein the audio data capturing the spoken utterance comprises a sequence of segments and wherein processing the audio data capturing the spoken utterance using the ASR model comprises: (each word is a segment in a sequence utterance, as in 0104 with fig. 7b, the user can select the suggested options specifically by replacing or selecting words by speaking as in 0064-0066 using ASR models, such actions would be considered a second pass per se, the alternative interpretation is any pass that is in the future at a separate time) using both the first-pass portion of the ASR model and the second pass portion of the ASR model in processing each of the segments; (utilizing the concept that each word is a segment in a sequence utterance, as in 0104 with fig. 7b, the user can select the suggested options specifically by replacing or selecting words by speaking as in 0064-0066 using ASR models, such actions would be considered a second pass per se, the alternative interpretation is any pass that is in the future at a separate time) causing a client device to perform one or more actions based on the text representation of the spoken utterance. (in fig. 7c the action is opening and populating the app fields just by uttering a command… ASR inherently and explicitly in Phillips, converts speech to text, particularly after iterations of model training, e.g. the first pass and second pass involving a user command 0105 and fig. 7b-7c, an ASR model containing encoded data or an encoder per se as in 0177 and 0185, ASR operations using speech recognition models for training using past utterances, user history, context, and corrections thereof for instance 0071 fig. 1 and fig. 2) Re claim 10, this claim has been rejected for teaching a broader, or narrower claim based on general inclusion of hardware alone (e.g. processor, memory, instructions), representation of claim 1 omitting/including hardware for instance, otherwise amounting to a virtually identical scope For instance, see fig. 1 and 0194 which contains the necessary memory and processors. Re claim 19, this claim has been rejected for teaching a broader, or narrower claim based on general inclusion of hardware alone (e.g. processor, memory, instructions), representation of claim 1 omitting/including hardware for instance, otherwise amounting to a virtually identical scope For instance, see fig. 1 and 0194 which contains the necessary memory and processors. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 2 and 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over US 20110054899 A1 Phillips; Michael S. et al. (hereinafter Phillips) in view of US 20190318261 A1 Deng; Yue et al. (hereinafter Deng). Re claims 2 and 11, Phillips teaches 2. The method of claim 1, wherein training the first-pass portion of the ASR model further comprises: (in fig. 7c the action is opening and populating the app fields just by uttering a command… ASR inherently and explicitly in Phillips, converts speech to text, particularly after iterations of model training, e.g. the first pass and second pass involving a user command 0105 and fig. 7b-7c) for each of the plurality of training instances and until one or more conditions are satisfied: (confidence levels met 0083… in fig. 7c the action is opening and populating the app fields just by uttering a command… ASR inherently and explicitly in Phillips, converts speech to text, particularly after iterations of model training, e.g. the first pass and second pass involving a user command 0105 and fig. 7b-7c) processing an instance of training audio data portion of the training instance using the shared encoder to generate shared encoder training output, wherein the training audio data captures a spoken training utterance; (an ASR model containing encoded data or an encoder per se as in 0177 and 0185, ASR operations using speech recognition models for training using past utterances, user history, context, and corrections thereof for instance 0071 fig. 1 and fig. 2) However, while Phillips teaches encoder or encoding based ASR model learning, it fails to teach RNN concepts: processing the shared encoder training output using the RNN-T decoder portion to generate predicted RNN-T training output; (Deng a sequence learning on real time data RNN is analogous to RNN-T under BRI, training by using sequenced learning on real tie data through loss in comparing prediction with ground truth and updating encoders per se 0055-0057 and 0065) determining a loss based on comparing the predicted RNN-T training output and a ground truth text representation of the training utterance; (Deng a sequence learning on real time data RNN is analogous to RNN-T under BRI, training by using sequenced learning on real tie data through loss in comparing prediction with ground truth and updating encoders per se 0055-0057 and 0065) updating the one or more portions of the shared encoder (Deng a sequence learning on real time data RNN is analogous to RNN-T under BRI, training by using sequenced learning on real tie data through loss in comparing prediction with ground truth and updating encoders per se 0055-0057 and 0065) and/or updating the one or more portions of the RNN-T decoder based on the loss. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Phillips to incorporate the above claim limitations as taught by Deng to allow for simple substitution of one known element for another to obtain predictable results such as the trainable model with encoding and ASR of Phillips with an RNN equivalent for speech to allow for superior mapping of variable-length audio to text by leveraging sequential memory, improving both recognition accuracy and adaptability to new data, by having the encoder learn complex temporal dynamics and context from raw audio e.g. streamed for using the ground truth (transcriptions) to update these encoders, often in an end-to-end (E2E) fashion functionally analogous to RNN with sequenced data. Claims 6-8 and 15-17 is/are rejected under 35 U.S.C. 103 as being unpatentable over US 20110054899 A1 Phillips; Michael S. et al. (hereinafter Phillips) in view of US 20190318261 A1 Deng; Yue et al. (hereinafter Deng) and further in view of US 20180166066 A1 Dimitriadis; Dimitrios B. et al. (hereinafter Dimitriadis). Re claims 6 and 15, while the combination teaches encoder or encoding based ASR model learning, it fails to teach speaking start/stop and RNN concepts: 6. The method of claim 2, wherein the RNN-T training output includes an end of query token indicating a human speaker has finished speaking using the first-pass portion of the ASR model. (Dimitriadis user start and stop areas determined by applying RNN to sequenced frames of real time audio data 0041 with 0063-0064 fig. 4-6 and fig. 9, evidenced further in 0050-0054 for training or having the RNN learn per se, a label analogous to a token with nth number of passes possible, to prevent erroneous speech) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Phillips in view of Deng to incorporate the above claim limitations as taught by Dimitriadis to allow for simple substitution of one known element for another to obtain predictable results such as the trainable model with encoding and ASR of Phillips with an RNN and start/stop labeling for nth iterations of speech to train a model analogous to the model of Phllips with a weighting to enforce good learning avoid or prevent garbage or wrong recognitions, to allow for improved speaker diarization and accuracy of who is talking and when they start/stop, with learning to prevent erroneous speech being analogous to loss per se which RNNs utilize to handle arbitrary input sequences. Re claims 7 and 16, while the combination teaches encoder or encoding based ASR model learning, it fails to teach speaking start/stop and RNN concepts: 7. The method of claim 6, wherein determining the human speaker has finished speaking the utterance comprises determining the human speaker has finished speaking the utterance in response to identifying the end of query token in the RNN-T training output. (Dimitriadis user start and stop areas determined by applying RNN to sequenced frames of real time audio data 0041 with 0063-0064 fig. 4-6 and fig. 9, evidenced further in 0050-0054 for training or having the RNN learn per se, a label analogous to a token with nth number of passes possible, to prevent erroneous speech) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Phillips in view of Deng to incorporate the above claim limitations as taught by Dimitriadis to allow for simple substitution of one known element for another to obtain predictable results such as the trainable model with encoding and ASR of Phillips with an RNN and start/stop labeling for nth iterations of speech to train a model analogous to the model of Phllips with a weighting to enforce good learning avoid or prevent garbage or wrong recognitions, to allow for improved speaker diarization and accuracy of who is talking and when they start/stop, with learning to prevent erroneous speech being analogous to loss per se which RNNs utilize to handle arbitrary input sequences. Re claims 8 and 17, while the combination teaches encoder or encoding based ASR model learning, it fails to teach speaking start/stop and RNN concepts: 8. The method of claim 7, wherein training the ASR model comprises penalizing the RNN-T decoder portion for generating the end of query token too early or too late. (Dimitriadis RNNs use loss to predict outcomes and to prevent erroneous speech 0073, user start and stop areas determined by applying RNN to sequenced frames of real time audio data 0041 with 0063-0064 fig. 4-6 and fig. 9, evidenced further in 0050-0054 for training or having the RNN learn per se, a label analogous to a token with nth number of passes possible) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Phillips in view of Deng to incorporate the above claim limitations as taught by Dimitriadis to allow for simple substitution of one known element for another to obtain predictable results such as the trainable model with encoding and ASR of Phillips with an RNN and start/stop labeling, in which an RNN inherently uses loss i.e. penalization under BRI, for nth iterations of speech to train a model analogous to the model of Phllips with a weighting to enforce good learning avoid or prevent garbage or wrong recognitions, to allow for improved speaker diarization and accuracy of who is talking and when they start/stop, with learning to prevent erroneous speech being analogous to loss per se which RNNs utilize to handle arbitrary input sequences. Allowable Subject Matter Claims 3-5 and 12-14 Claims 9 and 18 Are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. After searching through patent and non-patent literature, there was no evidence that there exists a limitation in direct relation or an obvious variant to such limitations as a whole as precisely limited. When searching for a secondary prior art for the limitation as recited in the above claims, the most relevant topics pertained to material from the same Inventor and Assignee but did not teach or suggest the aforementioned complex limitations as a whole as precisely limited. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. US 20180247643 A1 BATTENBERG; Eric et al. LAS concepts Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHAEL COLUCCI whose telephone number is (571)270-1847. The examiner can normally be reached on M-F 9 AM - 7 PM. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Flanders can be reached at (571)272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MICHAEL COLUCCI/Primary Examiner, Art Unit 2655 (571)-270-1847 Examiner FAX: (571)-270-2847 Michael.Colucci@uspto.gov
Read full office action

Prosecution Timeline

Aug 26, 2024
Application Filed
Mar 09, 2026
Non-Final Rejection mailed — §102, §103
Jun 09, 2026
Response Filed
Jul 16, 2026
Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12706085
SYSTEMS AND METHODS FOR CONTINUAL LEARNING FOR END TO-END AUTOMATIC SPEECH RECOGNITION
2y 5m to grant Granted Aug 11, 2026
Patent 12688851
VOICE BASED ACTIVATION DETECTION
2y 4m to grant Granted Jul 21, 2026
Patent 12682896
SYSTEM AND METHOD FOR THE GENERATION OF WORKLISTS FROM INTERACTION RECORDINGS
2y 7m to grant Granted Jul 14, 2026
Patent 12664976
QUERY REPLAY FOR PERSONALIZED RESPONSES IN AN LLM POWERED ASSISTANT
2y 6m to grant Granted Jun 23, 2026
Patent 12664977
FLY PARAMETER COMPRESSION AND DECOMPRESSION TO FACILITATE FORWARD AND/OR BACK PROPAGATION AT CLIENTS DURING FEDERATED LEARNING
2y 1m to grant Granted Jun 23, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
76%
Grant Probability
91%
With Interview (+15.2%)
3y 1m (~1y 2m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 1009 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month