Prosecution Insights
Last updated: August 18, 2026
Application No. 18/965,193

Language Agnostic Multilingual End-To-End Streaming On-Device ASR System

Non-Final OA §103
Filed
Dec 02, 2024
Priority
Oct 06, 2021 — provisional 63/262,161 +1 more
Examiner
ABEBE, DANIEL DEMELASH
Art Unit
Tech Center
Assignee
Google LLC
OA Round
1 (Non-Final)
90%
Grant Probability
Favorable
1-2
OA Rounds
9m
Est. Remaining
97%
With Interview

Examiner Intelligence

Grants 90% — above average
90%
Career Allowance Rate
927 granted / 1034 resolved
+29.7% vs TC avg
Moderate +7% lift
Without
With
+7.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 5m
Avg Prosecution
16 currently pending
Career history
1047
Total Applications
across all art units

Statute-Specific Performance

§101
12.7%
-27.3% vs TC avg
§103
31.3%
-8.7% vs TC avg
§102
26.7%
-13.3% vs TC avg
§112
9.6%
-30.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1034 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Double Patenting The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b). The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13. The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer. Claims 1-20 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1-20 of U.S. Patent No. 12183322. Although the claims at issue are not identical, they are not patentably distinct from each other because the claims in the present application define an invention that is merely an obvious variation of the invention claimed in the patent for the following reasons. Comparing the claims, such as claim 1 of the present application and claim 11 of the patent, it is clear that all the elements of the claim 1 are found in claim 11. The difference is claim 11 comprises additional elements including, a prediction network configured to receive, as input, a sequence of non-blank symbols output by a final softmax layer; and generate, at each of the plurality of output steps, a hidden representation, that are not included in claim 1, therefore represents a species of the generic invention of the application claims. Since it has been held that the generic invention is anticipated by the species, claim 1 of the present application is anticipated by claim 11 of the patent. Claims 2-10 are anticipated by claims 12-20 of the patent. Claims 11-20 are anticipated by claims 1-10 of the patent. 18965193 12183322 1. A computer-implemented method when executed by data processing hardware causes the data processing hardware to perform operations comprising: receiving, as input to an automated speech recognition (ASR) model, a sequence of acoustic frames characterizing one or more utterances; generating, by an encoder of the ASR model, at each of a plurality of output steps, a higher order feature representation for a corresponding acoustic frame in the sequence of acoustic frames, wherein the encoder comprises a stack of multi-headed attention layers; generating, by a joint network of the ASR model, at each of the plurality of output steps, a probability distribution over possible speech recognition hypotheses based on the higher order feature representation generated by the encoder at each of the plurality of output steps; and generating, by an endpointer model, at each of the plurality of output steps, a classification for the higher order feature representation generated for each acoustic frame in the sequence of acoustic frames as either speech, initial silence, intermediate silence, or final silence. 11. A computer-implemented method when executed by data processing hardware causes the data processing hardware to perform operations comprising: receiving, as input to a multilingual automated speech recognition (ASR) model, a sequence of acoustic frames characterizing one or more utterances; generating, by an encoder of the multilingual ASR model, at each of a plurality of output steps, a higher order feature representation for a corresponding acoustic frame in the sequence of acoustic frames, wherein the encoder comprises a stack of multi-headed attention layers; generating, by a prediction network of the multilingual ASR model, at each of the plurality of output steps, a hidden representation based on a sequence of non-blank symbols output by a final softmax layer; generating, by a first joint network of the multilingual ASR model, at each of the plurality of output steps, a probability distribution over possible speech recognition hypotheses based on the hidden representation generated by the prediction network at each of the plurality of output steps and the higher order feature representation generated by the encoder at each of the plurality of output steps; predicting, by a second joint network of the multilingual ASR model, an end of utterance (EOU) token at an end of each utterance based on the hidden representation generated by the prediction network at each of the plurality of output steps and the higher order feature representation generated by the encoder at each of the plurality of output steps; and classifying, by a multilingual endpointer model, each acoustic frame in the sequence of acoustic frames as either speech, initial silence, intermediate silence, or final silence. Examiner’s Note Examiner has cited particular columns and line numbers or figures in the references as applied to the claims below for the convenience of the applicant. Although the specified citations are representative of the teachings in the art and are applied to the specific limitations within the individual claim, other passages and figures may apply as well. It is respectfully requested from the applicant, in preparing the responses, to fully consider the references in entirety as potentially teaching all or part of the claimed invention, as well as the context of the passage as taught by the prior art or disclosed by the examiner. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1, 4-5, 7-11, 14-15 and 17-20 are rejected under 35 U.S.C. 103 as being unpatentable over Ramakrishnan/Rama et al. (US 2004/0015352) and in view of Kannan et al. (US 2020/0380215). As to claim 1, Rama teaches a computer-implemented method when executed by data processing hardware causes the data processing hardware to perform operations comprising: receiving 101, as input to an automated speech recognition (ASR) model, a sequence of acoustic frames characterizing one or more utterances; generating/extracting 110, a higher order feature representation 102 for a corresponding acoustic frame in the sequence of acoustic frames (Pars.26-30); Generating (120-142), by a joint network of the ASR model, at each of the plurality of output steps, a probability distribution over possible speech recognition hypotheses based on the higher order feature representation generated by the encoder at each of the plurality of output steps (Pars.39, 50-53); and Generating 150, by an endpointer model, at each of the plurality of output steps, a classification for the higher order feature representation generated for each acoustic frame in the sequence of acoustic frames as either speech, initial silence, intermediate silence, or final silence (Abstract; Pars.101-105). PNG media_image1.png 292 688 media_image1.png Greyscale It is noted that Rama doesn’t explicitly teach where generating/extracting the features comprise an encoder of the ASR model, comprising a stack of multi-headed attention layers. However, Kannan teaches an automated speech recognition (ASR) system for performing speech recognition using sequence-to-sequence models, where the system receives audio data for an utterance and provides features indicative of acoustic characteristics of the utterance as input to an encoder neural network comprising stack of attention layers (Figs.1-4; Pars.5-15). Combining the analogous teachings would be obvious to one of ordinary skill in the art before the time of applicant’s invention for the purpose of generating efficient features that accurately represent the acoustics characteristics of the utterance. As to claim 4, Kanan teaches, wherein the ASR model comprises a multilingual end-to-end (E2E) speech recognition model (Pars.3-5, 52). As to claim 5, Kannan teaches wherein each multilingual training utterance is concatenated with a corresponding language id/vector (Fig.1, #118; Pars.5, 13, 15, 52-55). As to claim 7, Kannan teaches wherein the sequence of acoustic frames characterizes a first utterance spoken in a first language followed by a second utterance spoken in a second language different than the first language (Figs.2A-2B). As to claim 8, Kannan teaches wherein the stack of multi-headed attention layers comprises a stack of conformer layers (Figs.2-3). As to claim 9, Kannan teaches wherein: the higher order feature representation generated at each of the plurality of output steps comprises a first higher order feature representation; and the encoder comprises: a first sub-encoder configured to generate, at each of the plurality of output steps, the first higher order feature representation for the corresponding acoustic frame in the sequence of acoustic frames; and a second sub-encoder configured to receive, as input, the first higher feature representation generated at each of the plurality of output steps and generate, at each of the plurality of output steps, a second higher order feature representation for the corresponding acoustic frame in the sequence of acoustic frames (Figs.1-3). As to claim 10, Kannan wherein the operations further comprise: generating, by a prediction network of the ASR model, at each of the plurality of output steps, a hidden representation based on a sequence of non-blank symbols output by a final softmax layer 240, wherein generating the probability distribution over possible speech recognition hypotheses comprises generating, by the joint network of the ASR model, at each of the plurality of output steps, the probability distribution over possible speech recognition hypotheses based on the hidden representation generated by the prediction network 220 at each of the plurality of output steps and the second higher order feature representation generated by the second sub-encoder at each of the plurality of output steps (Fig.1). Regarding claims 11, 14-15 and 17-20, the corresponding system comprising the steps similar to the claims addressed above, is analogous therefore rejected as being unpatentable over Rama and in view of Kannan for the foregoing reasons. Allowable Subject Matter Claims 2-3, 6, 12-13 and 16 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The following is a statement of reasons for the indication of allowable subject matter: claims 2 and 12 allowable because the prior arts do not teach wherein the operations further comprise triggering, by a microphone closer, a microphone closing event in response to the endpointer model classifying a higher order feature representation for a corresponding acoustic frame as final silence. claims 6 and 16 allowable because the prior arts do not teach multilingual training utterances concatenated with a corresponding domain ID representing a voice search domain comprise end of utterance (EOU) training tokens; and multilingual training utterances concatenated with a corresponding domain ID representing a domain other than the voice search domain do not include any EOU training tokens. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Prabhavalkar et al (US 2020/0027444), (Figs.1-4) PNG media_image2.png 514 708 media_image2.png Greyscale Any inquiry concerning this communication or earlier communications from the examiner should be directed to DANIEL DEMELASH ABEBE whose telephone number is (571)272-7615. The examiner can normally be reached monday-friday 7-4. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at 571-272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /DANIEL ABEBE/ Primary Examiner, Art Unit 2657
Read full office action

Prosecution Timeline

Dec 02, 2024
Application Filed
Jul 31, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12706107
METHOD AND APPARATUS FOR ENCODING/DECODING AUDIO SIGNAL
2y 3m to grant Granted Aug 11, 2026
Patent 12694874
MULTILINGUAL COMMAND LINE INTERFACE BOTS ENABLED WITH A REAL-TIME LANGUAGE ANALYSIS SERVICE
2y 7m to grant Granted Jul 28, 2026
Patent 12694881
METHODS, ENCODER AND DECODER FOR HANDLING ENVELOPE REPRESENTATION COEFFICIENTS
2y 3m to grant Granted Jul 28, 2026
Patent 12681689
ELECTRONIC DEVICE AND CONTROLLING METHOD OF ELECTRONIC DEVICE
2y 4m to grant Granted Jul 14, 2026
Patent 12646514
Power-Sensitive Control of Virtual Agents
2y 10m to grant Granted Jun 02, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
90%
Grant Probability
97%
With Interview (+7.3%)
2y 5m (~9m remaining)
Median Time to Grant
Low
PTA Risk
Based on 1034 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month