Prosecution Insights
Last updated: August 17, 2026
Application No. 18/493,365

AUTOMATED PREDICTION OF PRONUNCIATION OF TEXT ENTITIES BASED ON CO-EMITTED SPEECH RECOGNITION PREDICTIONS

Non-Final OA §101
Filed
Oct 24, 2023
Examiner
SHAH, PARAS D
Art Unit
2655
Tech Center
2600 — Communications
Assignee
Google LLC
OA Round
3 (Non-Final)
73%
Grant Probability
Favorable
3-4
OA Rounds
11m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 73% — above average
73%
Career Allowance Rate
480 granted / 654 resolved
+11.4% vs TC avg
Strong +31% interview lift
Without
With
+31.2%
Interview Lift
resolved cases with interview
Typical timeline
3y 9m
Avg Prosecution
20 currently pending
Career history
677
Total Applications
across all art units

Statute-Specific Performance

§101
18.3%
-21.7% vs TC avg
§103
47.4%
+7.4% vs TC avg
§102
13.4%
-26.6% vs TC avg
§112
10.7%
-29.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 654 resolved cases

Office Action

§101
DETAILED ACTION This communication is in response to the Amendments and Arguments filed on 06/16/2026. Claims 1-3, 5-10, 12-16, and 18-20 are pending and have been examined. Any previous objection/rejection not mentioned in this office action has been withdrawn by the Examiner. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 04/16/2026 has been entered. Compact Prosecution The Examiner recommends the following amendments to overcome the 35 USC 101 abstract rejection below. More specifically, including: How the last limitation of the claim is used to update the “pronunciation model” iteratively based on the co-emitted text entities and update of the phoneme space” (see [0039] of as filed spec) Include the connection to the usage of this updated pronunciation model by the ASR in a subsequent iteration to show the improvement of the ASR model as a result of this process. Response to Amendments and Arguments With respect to the Applicant’s amendments of the independent claims in view of the 35 USC 103 rejections, the Examiner agrees with the arguments and therefore the prior art rejections have been withdrawn. However, upon further consideration, the 35 USC 101 abstract rejections have been reinstated. Although, the prior examiner’s reasoning has been considered, the claims are directed to an abstract idea for the reasons mentioned below. The Examiner also provided specific amendments below that would overcome these rejections and tie the claims to a practical application as noted above. On 10/13/2025, the Applicant presented arguments with respect to the 35 USC 101 rejections. The Examiner will address those in this OA. The Applicant asserts on page 7: The Examiner's rejection fails to appreciate the specific, non-abstract limitations that shift the "character of the claim" away from a mere abstract idea and toward a tangible technological solution that improves an Automated Speech Recognition (ASR) system, as described in the specification. The claimed invention recites several limitations that are inherently technical and cannot be reasonably performed in the human mind, nor are they mere "insignificant extra-solution activities" that amount to generic computer functions. Namely, the claims now require: "generating, via processing circuitry, an encoding of allowable pronunciations of the text sample within a phoneme space." The Examiner disagrees with this assertion. The Examiner as noted above in the 35 USC 101 section clearly has shown how the claims are directed towards an abstract idea. The Applicant claims that the claims provide an improvement to ASR as described in the Specification. Although this may be true, the claims do not reflect this. The claims simply have one instance of the usage of ASR and that is simply to generate text and co-emitted text of the audio sample. The Examiner has also noted that this ASR per the spec is a general purpose ASR. Therefore, the updated encoding of allowable pronunciations are not used again by the ASR not is the ASR updated in any way and then used to show this improvement. Applicant assertions on page 8: A human mind, even one of a speech pathologist, cannot practically generate a multidimensional "encoding of allowable pronunciations" and locate them precisely "within a phoneme space." The specification explains that this "phoneme space" is a sophisticated technical construct: "The phoneme space can be modeled in two dimensions or can have more than two dimensions" and its "position of pronunciations... can correspond to acoustic features of the pronunciation" ([0020]). This is a technical step involving the computation and representation of complex, multi-dimensional acoustic and phonetic data, which is only possible with specialized processing circuitry. The output of this step-a defined, encoded data structure within a computable space-is a machine-readable artifact that is essential for the ASR system's function, not a result of human thought. The claims also require: "for each corresponding co-emitted text sample... determining, using a corresponding phoneme model associated with corresponding text sample, one or more corresponding pronunciations of the corresponding co-emitted text sample." . This step describes a machine-driven data processing task that involves accessing and utilizing a technical component-a "corresponding phoneme model"-for each of the one or more co-emitted text samples. The sheer complexity and volume of data involved in a real-time ASR system, especially one that generates and processes multiple co-emitted text entities ([0031]), is far beyond the capacity of the human mind to practically manage or compute simultaneously. The human mind might recall a phonetic approximation, but it cannot determine pronunciations for multiple simultaneous hypotheses by accessing and running multiple associated, technical phoneme models. The Examiner disagrees with these assertions. As noted above a human can reasonably construct a phoneme space in 2D form. It is unclear based on the Applicant’s arguments why 2D is not something a human can plot on a cartesian coordinate. The claims only note that the text (pronunciation) is encoded onto phoneme space. Thus, the usage of multidimensional data of acoustic and phonetic is misleading as the audio sample is converted to text form and is no longer in acoustic and phonetic for. Further, the encoded data structure is not defined per the claims and generally represents the numeric representation of the text which is a vectorized version. This is not “machine readable artifact” essential for the ASR systems function. This Applicant appears to be limiting the interpretation of “encoding” as the claims do not tie this to the ASR system’s function as being asserted. The Applicant next asserts that the “phoneme model” is a machine driven processing task. The Examiner disagrees as there is nothing recited regarding this phoneme model and under BRI can be reasonably interpreted as converting text to phonetic representations using a mapping rule or pre-existing rules. A human can easily convert text to phonetic reprentations or pronunciations can be created. This does not require a machine. With respect to the last paragraph of the response, the complexity of the of the stated claimed processes where a human “cannot determine pronunciations for multiple simultaneous hypotheses by accessing and running multiple associated, technical phoneme models”, the examiner respectfully disagrees. The Examiner notes that the fact that a human cant process it as quick as a machine can is not a proper justification that the claims are patent eligible. The examiner recommends the Applicant to review the case law of Versata Dev. Group v. SAP Am., Inc., 793 F.3d 1306, 1335, 115 USPQ2d 1681, 1702. The Applicant asserts on page 9-10: The claims now also require: "encoding the corresponding pronunciations of the one or more co-emitted text samples as a convex hull in the phoneme space." This is a highly specific geometric and computational operation that is entirely machine- implemented. A human mind is not "equipped" (2025 Guidance Memo, Sec. II.A) to perform the process of "encoding... as a convex hull" in a multi-dimensional phoneme space ([0020], [0038]). This limitation is not a simple comparison but a particular, technical mechanism for "computing an area of intersection" ([0038]), which involves algorithmic determination of an extreme subset of points to form a boundary in a high-dimensional space. The Examiner's rejection of related dependent claim 4 as a "mental process" is therefore in error, as the complex geometry and computation of a convex hull is a technical task, not a human one. Lastly, the independent claims are further amended to specify: "updating... the encoding... by removing any pronunciations of the text sample that fall outside of the convex hull based on pronunciations of the one or more co-emitted text samples." This is a precise, automated data modification step that relies on the previously generated technical features: the multi-dimensional phoneme space encoding and the convex hull geometric boundary. The human mind is capable of an abstract "evaluation" or "comparison," but it cannot practically execute a complex data removal operation that involves checking the coordinates of potentially hundreds or thousands of encoded pronunciations against the mathematical boundaries of a multi-dimensional convex hull ([0038], [0039]). This action is a specific technical refinement of the ASR system's data (the pronunciation model) that improves its future performance ([0039]), clearly setting the claim apart from a simple mental step. The independent claims, when considered as a whole, describe a particular technical solution for refining an ASR pronunciation model. The use of "phoneme space," "phoneme model," and the highly specific "encoding... as a convex hull" to perform data-driven "updating" are all limitations that are far beyond the practical capability of the human mind. Therefore, the claim is not merely an automation of a mental process but is directed to a specific, technical method for improving a computer's functionality in the technical field of Automated Speech Recognition, particularly for refining pronunciation data models, a process that is not practically performable in the human mind. The "character of the claim" is directed to a non- abstract, machine-implemented solution, and the rejection under 35 U.S.C. §101 must be withdrawn. The Examiner respectfully disagrees for the reasons mentioned below with respect to the newly added limitation. The Applicant has not shown or explained how a “convex hull” cannot be performed in the human mind once the data has been plotted. As the examiner noted above, the convex hull is similar to a rubber band test to determine the majority of the data the polygon that is created represents. The Applicant notes that the action requires “technical refinement of the ASR systems data”. The Examiner notes that this aspect is not clear as to how this is a technical refinement of ASR systems data since the claim scope is not reflective of this aspect. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The independent claims 1, 8, and 15 relate to a method and system and system and CRM thus relating to a statutory category. The claims further recite per claim 1 “generating, via processing circuitry, an encoding of allowable pronunciations of the text sample within a phoneme space; receiving an audio sample including the text sample, the audio sample comprising speech spoken by a user; processing, using automatic speech recognition (ASR), via the processing circuity, the audio sample to generate predicted text samples corresponding to the audio sample, the predicted text samples including the text sample and one or more co-emitted text samples; outputting, via the processing circuitry, the text sample; for each corresponding co-emitted text sample of the one or more co-emitted text samples, determining, using a corresponding phoneme model associated with corresponding text sample, one or more corresponding pronunciations of the corresponding co-emitted text sample; encoding the corresponding pronunciations of the one or more co-emitted text samples as a convex hull in the phoneme space; and updating, via the processing circuitry, the encoding of allowable pronunciations of the text sample by refining allowable pronunciations of the text sample towards a centroid of the convex hull and removing any pronunciations of the text sample that fall outside of the convex hull.” The limitation of claims 1 of “generating…”, “receiving…”, “processing...”, “outputting…”. “determining…”, “encoding…” and “updating…” as drafted covers mental activities. More specifically, a human can generate a XY coordinate mapping of allowable pronunciations mapped on a cartesian coordinate system of a finite number by first converting the pronunciation into a numeric representation. Then, receiving further audio from another person, Generating, by the human a text representation and further other related text representations. Then, converting the text representations into phonetic representations. Then converting these phonetic representations into numeric representation to determine where each fits in the cartesian coordinate. Further, removing outliers and retaining those pronunciations that are close to the convex hull. The Examiner notes that a convex hull is in essence similar to the rubber band test, where the rubber band forms a polygon on the outermost points which forms the convex hull. Thus, based on the plots a human can determine a polygon visually and then create this convex hull to remove that. This judicial exception is not integrated into a practical application. In particular, claims 1 and 8 recites additional limitations of “processing circuitry” and “ASR”, claim 14 recites “processor”, “computer readable storage media”,. Each of these elements are being used as a tool on which the stated steps operate on. Paragraph [0066] describe the processing circuitry and processor and CRM can be implemented using general purpose computers and notes any “medium” can be used. Para [0014] notes that the ASR model can be in exemplary a end-to-end speech recognition mode which is a general purpose model. Accordingly, since there are no additional elements then the abstract idea cannot be integrated into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claims are not patent eligible. With respect to claim 2, 9, and 16, the claim relates to “wherein the encoding of allowable pronunciations is generated based on a measure of pronunciation certainty of the text sample.” This reads on a human converting the pronunciations based on a confidence of the pronunciation associated with the text sample. No additional limitations are present. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception other than those already mention in the independent claims. With respect to claim 3, 10 and 17, the claims relate to “wherein the text sample is outputted based on syntactic context of the audio sample.” This relates to a human converting audio to text based on syntax such as part of speech. No additional limitations are present. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. With respect to claim 5, 12, and 18, the claim relates to “wherein the updating the encoding of allowable pronunciations of the text sample further includes updating a predicted accuracy of allowable pronunciations of the text sample based on the pronunciations of the one or more co- emitted text samples.” This relates to the human updating the encoded set of pronunciations known based on the variations determined from the other user’s audio input. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. With respect to claim 6, 13, and 19, the claim relates to “wherein the updating the encoding of allowable pronunciations of the text sample further includes generating allowable pronunciations using a grapheme-to phoneme model.” This relates to a human compiling and updating a mapping between text to phonetic representations based on the updated pronunciations. No additional elements are present. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. With respect to claims 7 and 20, the claim relates to “wherein the pronunciations of the one or more co-emitted text samples are inputs to the grapheme-to phoneme model.” This relates a human using the variations of text and their pronunciations to provide to a mapping to remember the similar variations. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. These claims further do not remedy the judicial exception being integrated into a practical application and further fail to include additional elements that are sufficient to amount to significantly more than the judicial exception. Allowable Subject Matter Claim 1-3, 5-10, 12-16 and 18-20 would be allowable if rewritten or amended to overcome the rejection(s) under 35 U.S.C. 101, set forth in this Office action. The following is a statement of reasons for the indication of allowable subject matter: The closest prior art of record, Beutnagel teaches a method for predicting pronunciation of a text sample, comprising: (Beutnagel teaches predicting plausible pronunciations. Beutnagel at 4:52 - 5:15.) generating, via processing circuitry, an encoding of allowable pronunciations of the text sample within a phoneme space; (Beutnagel teaches generating candidate pronunciations for a word (i.e., text entity) wherein the candidates are generated in order to ensure the correct pronunciation is selected (i.e., the correct pronunciation is "allowable"). Beutnagel at 5:5 - 6:19. Further, Beutnagel teaches generating a plurality of candidate pronunciations phonetically (i.e., encoding of allowable pronunciations). Beutnagel at 5:16 - 6:41. Further, Beutnagel teaches a computer coupled to memory executing software such as a speech recognition engine. Beutnagel at 2:35 – 2:56. Further still, Beutnagel teaches generating predicted pronunciations for a text sample based upon phonemes (i.e., within a phoneme space.) Beutnagel at Fig. 2 and 4:12 - 7:50.) receiving an audio sample including the text sample, the audio sample comprising speech spoken by a user; (Beutnagel teaches receiving an audio sample containing an audio sample corresponding to a text sample and generating predicted pronunciations for the text. Beutnagel at Fig. 2 and 4:12 - 5:56.) processing, via the processing circuitry, the audio sample to generate predicted text samples corresponding to an audio sample, the predicted text samples including the text sample and one or more co-emitted text samples; (Beutnagel teaches selecting the N-best results of the system to ensure the correct pronunciation is selected. Beutnagel at 6:20 - 6:64. As such, the N-best results concept, in conjunction with the generation of multiple pronunciation candidates, demonstrates that Beutnagel’s candidate pronunciations are associated with an audio sample (e.g., the user specifying the word Peabody verbally in 4:12 - 4:26.) Further, Beutnagel teaches processing audio samples using speech recognition in part of a process for generating predicted pronunciations of text samples and audio samples. Beutnagel at Fig. 2 and 4:12 – 7:50.) outputting, via the processing circuitry, the text sample; (Beutnagel teaches returning the results (i.e., outputting) which are the N-best ranked answers. Beutnagel at 6:20 - 6:41.) Beutnagel, however, does not teach updating, via the processing circuitry, the encoding of allowable pronunciations of the text sample based on pronunciations of the one or more co-emitted text samples. Fanty teaches updating, via the processing circuitry, the encoding of allowable pronunciations of the text sample …. (Fanty teaches updating a pronunciation dictionary of allowable pronunciations (i.e., an encoding of allowable pronunciations) by comparing the pronunciations of one pattern with a series of replacement or alternative phonemes (i.e., co-emitted text entities). Fanty at 6:63 - 7:65.) Beutnagel-Fanty, however, do not alone teach for each corresponding co-emitted text sample of the one or more co-emitted text samples, determining, using a corresponding phoneme model associated with corresponding text sample, one or more corresponding pronunciations of the corresponding co-emitted text sample; Aleksic teaches for each corresponding co-emitted text sample of the one or more co-emitted text samples, determining, using a corresponding phoneme model associated with corresponding text sample, one or more corresponding pronunciations of the corresponding co-emitted text sample; (Aleksic teaches, for each candidate transcription generating grammars (i.e., allowable pronunciations) and calculating confidence scores for the grammars for each of the candidate transcriptions. Aleksic at 9:5 - 9:26 and 11:56 - 12:16. Further, Aleksic teaches using a sequence-to-sequence neural network with specific layers that correspond to each sample utterance (i.e., a corresponding phoneme model associated with the corresponding text samples.) Aleksic at 3:61 - 4:25.) Fox teaches encoding the dialect of the utterances using a convex hull; (Fox teaches encoding multiple vowels used in dialects and generations into a convex hull (i.e., encoding corresponding pronunciations of a text sample into a phoneme space) Fox at section II. subsection D. Vowel space area computations.) and … removing any dialects that fall outside of the convex hull. (Fox teaches including boundaries for the convex hull (i.e., excluding results outside the boundaries). Fox at section II. subsection D. Vowel space area computations.) However, none of the cited references either alone or in combination thereof teaches the combination of limitations as recited in the independent claims. More specifically, the limitations of “updating, via the processing circuitry, the encoding of allowable pronunciations of the text sample by refining allowable pronunciations of the text sample towards a centroid of the convex hull and removing any pronunciations of the text sample that fall outside of the convex hull” in combination with the other elements of the claim. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Sahajpal et al. (“Transcription of Text by Incremental Support Vector Machine”) is cited to disclose use of “convex hull” for learning different phonemic data (see sect. 4). Khan et al, (“Diversity by Phonetics and its Application in Neural Machine Translation”) is cited disclose determination of a convex hull of pinyin if embeddings (see sect 5). Machery (“Discriminative Training and Acoustic Modeling for Automatic Speech Recognition”) is cited to disclose use of a convex hull for MER training on lattices (See sect 8.3). Any inquiry concerning this communication or earlier communications from the examiner should be directed to PARAS D SHAH whose telephone number is (571)270-1650. The examiner can normally be reached Monday-Thursday 7:30AM-2:30PM, 5PM-7PM (EST), Friday 8AM-noon (EST). Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, PARAS D SHAH can be reached at 571-270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Paras D Shah/ Supervisory Patent Examiner, Art Unit 2653 07/10/2026
Read full office action

Prosecution Timeline

Oct 24, 2023
Application Filed
Aug 08, 2025
Non-Final Rejection mailed — §101
Oct 13, 2025
Response Filed
Feb 09, 2026
Final Rejection mailed — §101
Apr 16, 2026
Request for Continued Examination
Apr 20, 2026
Response after Non-Final Action
Aug 05, 2026
Non-Final Rejection mailed — §101 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12670920
Joint Acoustic Echo Cancellation (AEC) and Personalized Noise Suppression (PNS)
3y 4m to grant Granted Jun 30, 2026
Patent 12670924
SPATIAL REGION BASED AUDIO SEPARATION
2y 7m to grant Granted Jun 30, 2026
Patent 12633295
THREE-DIMENSIONAL AUDIO SIGNAL CODING METHOD AND APPARATUS, AND ENCODER
2y 6m to grant Granted May 19, 2026
Patent 12586591
SOUND SIGNAL DECODING METHOD, SOUND SIGNAL DECODER, PROGRAM, AND RECORDING MEDIUM
3y 3m to grant Granted Mar 24, 2026
Patent 12579367
TWO-TOWER NEURAL NETWORK FOR CONTENT-AUDIENCE RELATIONSHIP PREDICTION
2y 4m to grant Granted Mar 17, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
73%
Grant Probability
99%
With Interview (+31.2%)
3y 9m (~11m remaining)
Median Time to Grant
High
PTA Risk
Based on 654 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month