Prosecution Insights
Last updated: August 17, 2026
Application No. 18/191,711

On-Device Multilingual Speech Recognition

Final Rejection §102§103
Filed
Mar 28, 2023
Examiner
SHARMA, NEERAJ
Art Unit
2659
Tech Center
2600 — Communications
Assignee
Google LLC
OA Round
2 (Final)
85%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
97%
With Interview

Examiner Intelligence

Grants 85% — above average
85%
Career Allowance Rate
395 granted / 466 resolved
+22.8% vs TC avg
Moderate +12% lift
Without
With
+11.9%
Interview Lift
resolved cases with interview
Typical timeline
2y 8m
Avg Prosecution
29 currently pending
Career history
487
Total Applications
across all art units

Statute-Specific Performance

§101
17.4%
-22.6% vs TC avg
§103
46.6%
+6.6% vs TC avg
§102
28.3%
-11.7% vs TC avg
§112
6.1%
-33.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 466 resolved cases

Office Action

§102 §103
DETAILED ACTION Introduction 1. This office action is in response to Applicant's submission filed on 03/28/2023. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claims 1-20 are currently pending and examined below. Drawings 2. The drawings filed on 03/28/2023 have been accepted and considered by the Examiner. Information Disclosure Statement 3. The Information Statement (IDS) filed on 08/01/2025 has been accepted/considered and is in compliance with the provisions of 37 CFR 1.97. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) The claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention. 4. Claims 1-2, 8, 10-12, 18 and 20 are rejected under 35 U.S.C. 102 (a) (1) as being anticipated by Ramanna (U.S. Patent Application Publication # 2023/0016962 A1). With regards to claim 1, Ramanna teaches a computer-implemented method executed on data processing hardware of a client device that causes the data processing hardware to perform operations comprising while a first language pack for use in recognizing speech in a first language is loaded onto the client device, receiving a sequence of input audio frames generated from input audio data characterizing an utterance (Para 13, teaches a multilingual Natural Language Understanding or NLU model platform. Using the disclosed platform and related techniques, NLU models to support different languages can be created and managed. A primary NLU model is created in a first human language such as English. Para 14, teaches receiving user input utterances); processing, by a language identification (ID) predictor model, each corresponding input audio frame in the sequence of input audio frames to determine a language ID event associated with the corresponding input audio frame that indicates a predicted language for the corresponding input audio frame (Para 14, further teaches that the primary NLU model is configured to understand user input utterances. Para 26, illustrates an example whereby the primary NLU model for understanding statements in English can be integrated with a virtual agent application. This allows the virtual agent application to interpret the user context of user input statements provided in English. When a user inputs the statement “I need to change my password,” the primary NLU model predicts that the intent is a “Reset Password” intent); obtaining a sequence of speech recognition events for the sequence of input audio frames, each speech recognition event comprising a respective speech recognition result in the first language determined by the first language pack for a corresponding one of the input audio frames (Para 26, illustrates an example whereby the primary NLU model for understanding statements in English can be integrated with a virtual agent application. This allows the virtual agent application to interpret the user context of user input statements provided in English. When a user inputs the statement “I need to change my password,” the primary NLU model predicts that the intent is a “Reset Password” intent); based on determining that the language ID events are indicative of the utterance including a language switch from the first language to a second language loading, from memory hardware of the client device, a second language pack onto the client device for use in recognizing speech in the second language (Para 13, teaches that secondary NLU models that are consistent with the first NLU model are then created to support additional languages. Para 35, teaches that by maintaining a consistent application to the NLU model interface, the software application can be configured to easily enable/disable support for different languages by switching NLU models); and rewinding the input audio data buffered by an audio buffer to a time of the corresponding input audio frame associated with the language ID event that first indicated the second language as the predicted language (Para 55, teaches an example wherein configured intent has the name “Hardware request” and includes 3 utterances and 3 associated entities. The last displayed intent, configured intent, has the name “Account access” and also includes 3 utterances and 3 associated entities. Each intent also displays timestamps corresponding to a creation date and last updated time); emitting, using the respective speech recognition results determined by the first language pack for only the corresponding input audio frames associated language ID events that indicate the first language as the predicted language, a first transcription for a first portion of the utterance (Para 52, teaches that primary language field displays the label “Primary language” and the value “English—en” to indicate that the primary language and the language of the primary NLU model is English. Secondary language field displays the label “* Translate into” and the value “French—fr” to indicate that the language of this secondary NLU model is French. Translation selection element allows the NLU model designer to select from multiple options for translation services including a software option, a manual option, and a third-party option. If the currently selected option is the software option, then it automatically translates the language content using machine intelligence); and processing, using the second language pack loaded onto the client device, the rewound buffered audio data to generate a second transcription for a second portion of the utterance (Para 52, teaches that primary language field displays the label “Primary language” and the value “English—en” to indicate that the primary language and the language of the primary NLU model is English. Secondary language field displays the label “* Translate into” and the value “French—fr” to indicate that the language of this secondary NLU model is French. Translation selection element allows the NLU model designer to select from multiple options for translation services including a software option, a manual option, and a third-party option. If the currently selected option is the software option, then it automatically translates the language content using machine intelligence). With regards to claim 2, Ramanna teaches the computer-implemented method of claim 1, wherein: the first transcription comprises one or more words in the first language and the second transcription comprises one or more words in the second language (Para 52, teaches that primary language field displays the label “Primary language” and the value “English—en” to indicate that the primary language and the language of the primary NLU model is English. Secondary language field displays the label “* Translate into” and the value “French—fr” to indicate that the language of this secondary NLU model is French. Translation selection element allows the NLU model designer to select from multiple options for translation services including a software option, a manual option, and a third-party option. If the currently selected option is the software option, then it automatically translates the language content using machine intelligence). With regards to claim 8, Ramanna teaches the computer-implemented method of claim 1, wherein each speech recognition event comprising the respective speech recognition result in the first language further comprises an indication that the respective speech recognition result comprises a partial result or a final result (Para 28, teaches a final result, wherein as part of the workflow process, language content for the new intent is automatically translated. For example, a new “Show Account Status” intent is added that responds to the user input statement “What is my account status?” The new intent is added to each secondary NLU model along with the appropriate translation of the user statement “What is my account status?” into each model's configured language. Para 46, teaches a partial result wherein probabilistic prediction and scoring is used to rank user intent). With regards to claim 10, Ramanna teaches the computer-implemented method of claim 1, wherein the first language pack and the second language pack each comprise at least one of: an automated speech recognition (ASR) model; parameters/configurations of the ASR model; an external language model; neural network types; an acoustic encoder; components of a speech recognition decoder; or the language ID predictor model (Para 13, teaches a language ID predictor model and an external language model in its disclosure of a multilingual NLU model. Para 15, teaches use of neural network types in its disclosure of a secondary NLU machine learning model and corresponding machine learning weights). With regards to claims 11-12, 18 and 20, these are system claims for the corresponding method claims 1-2, 8 and 10. These two sets of claims are related as method and apparatus of using the same, with each claimed system element's function corresponding to the claimed method step. Accordingly, claims 11-12, 18 and 20 are similarly rejected under the same rationale as applied above with respect to method claims 1-2, 8 and 10. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 5. Claims 3 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Ramanna in view of Lavilla (U.S. Patent Application Publication # 2020/0160845 A1). With regards to claim 3, although Ramanna teaches probabilistic prediction and scoring (Para 46), it pertains to user intent determination. Ramanna may not explicitly detail the limitation wherein the language ID event determined by the language ID predictor model that indicates the predicted language for each corresponding input audio frame further comprises a probability score indicating a likelihood that the corresponding input audio frame includes the predicted language. This is taught by Lavilla (Para 23, teaches a deep neural network or DNN classifier is used to perform one or more semantic classifications on one or more portions of the audio stream. An output of the DNN classifier indicates a mathematical likelihood of the presence of a particular semantic class in the audio stream. DNN output include probabilistic or statistical predictive data values for each target class, where a target class corresponds to an enrolled language. Para 29, further teaches sample may refer to a temporal portion of digital data extracted from an audio stream, window may refer to a time interval over which features are extracted or scores are computed for a sample and segment may refer to a portion of the audio stream that contains one or more content classes). Ramanna and Lavilla can be considered as analogous art as they belong to a similar field of endeavor in language processing. It would thus have been obvious to one having ordinary skill in the art to advantageously combine the teachings of Lavilla (Use of probabilistic scoring for language prediction) with those of Ramanna (Use of probabilistic scoring for user intent prediction in a multilingual processing environment) so as to reduce delays between receipt of an audio sample and the output of a label indicating the classification of its speech content (Lavilla, para 16). With regards to claim 13, this is a system claim for the corresponding method claim 3. These two claims are related as method and apparatus of using the same, with each claimed system element's function corresponding to the claimed method step. Accordingly, claim 13 is similarly rejected under the same rationale as applied above with respect to method claim 3. Allowable Subject Matter 6. Claims 4-7, 9, 14-17 and 19 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The prior art of record, alone or in combination, does not currently suggest or teach the invention as outlined in these claims. More detailed reasons for allowance will be outlined as and when the Application proceeds to allowability. Conclusion 7. The following prior art, made of record but not relied upon, is considered pertinent to applicant's disclosure: Hannun (U.S. Patent Application Publication # 2023/0352011 A1), Wu (U.S. Patent Application Publication # 2008/0040099 A1). These references are also included in the PTO-892 form attached with this office action. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. If you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). In case you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. Any inquiry concerning this communication or earlier communications from the examiner should be directed to NEERAJ SHARMA whose contact information is given below. The examiner can normally be reached on Monday to Friday 8 am to 5 pm. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre Louis-Desir can be reached on 571-272-7799 (Direct Phone). The fax number for the organization where this application or proceeding is assigned is 571-273-8300. /NEERAJ SHARMA/ Primary Examiner, Art Unit 2659 571-270-5487 (Direct Phone) 571-270-6487 (Direct Fax) neeraj.sharma@uspto.gov (Direct Email)
Read full office action

Prosecution Timeline

Mar 28, 2023
Application Filed
May 27, 2025
Response after Non-Final Action
May 08, 2026
Non-Final Rejection mailed — §102, §103
Jul 30, 2026
Applicant Interview (Telephonic)
Jul 31, 2026
Response Filed
Aug 12, 2026
Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12706083
DYNAMIC TRANSLATION RELAY SYSTEM
2y 2m to grant Granted Aug 11, 2026
Patent 12688853
SYSTEM FOR ENABLING PROCESSING RICH MEDIA DATA
2y 6m to grant Granted Jul 21, 2026
Patent 12676144
DEEPFAKE DETECTION
2y 4m to grant Granted Jul 07, 2026
Patent 12676153
SPEECH RECOGNITION METHOD AND APPARATUS, ELECTRONIC DEVICE, AND COMPUTER-READABLE STORAGE MEDIUM
2y 2m to grant Granted Jul 07, 2026
Patent 12664695
TEXT AND IMAGE GENERATION FOR CREATION OF IMAGERY FROM AUDIBLE INPUT
3y 2m to grant Granted Jun 23, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
85%
Grant Probability
97%
With Interview (+11.9%)
2y 8m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 466 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month