Prosecution Insights
Last updated: October 04, 2026
Application No. 18/439,630

USING TEXT-INJECTION TO RECOGNIZE SPEECH WITHOUT TRANSCRIPTION

Non-Final OA §101§102§103
Filed
Feb 12, 2024
Priority
Mar 01, 2023 — provisional 63/487,821
Examiner
OGUNBIYI, OLUWADAMILOL M
Art Unit
2653
Tech Center
2600 — Communications
Assignee
Google LLC
OA Round
3 (Non-Final)
77%
Grant Probability
Favorable
3-4
OA Rounds
3m
Est. Remaining
96%
With Interview

Examiner Intelligence

Grants 77% — above average
77%
Career Allowance Rate
243 granted / 315 resolved
+15.1% vs TC avg
Strong +19% interview lift
Without
With
+19.4%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
28 currently pending
Career history
342
Total Applications
across all art units

Statute-Specific Performance

§101
20.8%
-19.2% vs TC avg
§103
49.9%
+9.9% vs TC avg
§102
11.2%
-28.8% vs TC avg
§112
13.3%
-26.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 315 resolved cases

Office Action

§101 §102 §103
DETAILED ACTION Claims 1 – 5, 7 – 15, and 17 – 20 are pending. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant’s submission filed on 08 May 2026 has been entered. Response to Amendment With regard to the Final Office Action from 27 March 2026, the Applicant has filed a response on 08 May 2026. Claims 6 and 16 have been cancelled. Claims 1 – 5, 10 – 15 and 20 were previously rejected under 35 U.S.C. 101 for being directed to a judicial exception without significantly more. The Applicant has amended the independent claims to incorporate subject matter from claims 6 and 16 which were not rejected under 35 U.S.C. 101. The incorporated subject matter introduces a limitation about ‘the speech recognition model comprises an audio encoder and a decoder, the audio encoder comprising a stack of self-attention layers each including multi-headed self-attention mechanism’ (as provided in claim 1). The Examiner notes that the audio encoder comprising a stack of self-attention layers with each including a multi-headed self-attention mechanism is not by itself a mental process, nor can it be described as any of the other indicated abstract ideas. As it currently stands, this limitation introduces an additional element presented as the audio encoder that comprises the stack of self-attention layers that include a multi-headed self-attention mechanism, and this integrates the mental process into a practical application, this being the training of the speech recognition model in order to learn to recognise speech in the target domain and phrases within the one or more classes of sensitive information. Due to the inclusion of this limitation in the independent claims, the Examiner hereby withdraws the 35 U.S.C. 101 rejection. Response to Arguments With regard to the 35 U.S.C. 103 rejection given to the claims, the Applicant’s arguments with respect to claims have been considered but are considered moot due to the new ground of rejection necessitated by the amendment to the independent claims. These claims will be addressed by their current presentation in the following section. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 2, 5, 7, 8, 9, 10, 11, 12, 15, 17, 18, 19 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Ganong, III et al. (US 2023/0395063 A1: hereafter — Ganong) in view of Bachtiger et al. (US 11,120,199 B1: hereafter — Bachtiger) further in view of Chen, Zhehuai, et al.1 (“Maestro: Matched speech text representations through modality matching.” arXiv preprint arXiv:2204.03409 (2022): hereafter — Chen). For claim 1, Ganong discloses a computer-implemented method when executed on data processing hardware causes the data processing hardware to perform operations (Ganong: [0020] — a processor (as a data processing hardware)) comprising: receiving training data comprising: transcribed speech utterances spoken in a general domain, each transcribed speech utterance paired with a corresponding transcription (Ganong: [0065] — having an input speech signal and a transcription of the input speech signal as its pair (this being an original domain whereby sensitive information is left in the speech and transcript)); modified speech utterances in a target domain, the modified speech utterances comprising utterances spoken in the target domain that have been modified to obfuscate one or more classes of sensitive information recited in the utterances, each modified speech utterance paired with a corresponding transcription that redacts the sensitive information obfuscated from the modified speech utterance (Ganong: [0070] — training involving the presence of obscured speech signal (the target domain being such a domain which has certain sensitive data removed); [0068] — an obscured input speech signal with the sensitive information removed); and unspoken textual utterances corresponding to the transcriptions of the modified speech utterances in the target domain, [[wherein the unspoken textual utterances are not paired with any corresponding audio representation]] and comprise fake [[random]] data inserted into redacted portions of the transcriptions of the modified speech utterances where the sensitive information recited in the modified speech utterance has been redacted (Ganong: [0055]–[0056] — replacing the sensitive portion of the transcription with fake information different from what was previously there; [0071] — training the speech processing model based on manually generated transcription of obscured speech signal (this transcription of the obscured speech signal belonging to unspoken utterances since the obscured speech was never uttered)); generating, using an alignment model, a corresponding alignment output for each unspoken textual utterance of the received training data (Ganong: [0071] — ‘aligning 428 the manually generated transcription of the obscured speech signal with the transcription’), [[wherein the alignment output comprises a textual representation having a number of frames that maps the unspoken textual utterance to speech frames, and wherein the alignment model generates the alignment output without converting the unspoken textual utterance into synthetic speech]]; and training a speech recognition model on the transcribed speech utterances, the modified speech utterances, [[and the alignment outputs generated for the unspoken textual utterances]] to teach the speech recognition model to learn to recognize speech in the target domain and phrases within the one or more classes of sensitive information (Ganong: [0070] — the training including the obscured speech signals (as the modified speech utterances); [0071] — training the speech processing model making use of the manually generated transcription, alignment with the transcription of the input speech (as the alignment outputs generated for the unspoken textual utterances, and also indicating that the transcription of the original speech is included for the purpose of training); [0056], FIGs. 6 & 7 — a transcription process that is able to identify the sensitive information in the speech), [[wherein the speech recognition model comprises an audio encoder and a decoder, the audio encoder comprising a stack of self-attention layers each including a multi-headed self-attention mechanism]]. The reference of Ganong provides teaching for inserting fake data into redacted portions of the transcriptions of the modified speech utterances, but differs from the claimed invention in that the claimed invention further provides teaching for inserting fake random information. This insertion of random information is however not new to the art as the reference of Bachtiger is now introduced to teach this as: unspoken textual utterances corresponding to the transcriptions of the modified speech utterances in the target domain, wherein the unspoken textual utterances are not paired with any corresponding audio representation and comprise fake random data inserted into redacted portions of the transcriptions of the modified speech utterances where the sensitive information recited in the modified speech utterance has been redacted (Bachtiger: Col 4 lines 44–48 — a redaction module which is able to substitute redacted portions of an utterance with random series of symbols such as ***-**-*** or #@$-$#-@#$#@ (which are the unspoken textual utterances that are not paired with any corresponding audio representation)). Hence, before the effective filing date of the claimed invention, one of ordinary skill in the art would have found it obvious to combine the known teaching of Bachtiger which replaces sensitive information with random data, with the teaching of inserting fake data into such redacted portions that contain sensitive information as taught by the reference of Ganong, to thereby come up with the claimed invention. The combination of both prior art elements would have provided the predictable result of being able to quickly generate and present information for replacing sensitive information without following a particular order, resulting in preventing information leaks that can be revealed through context. See KSR Int’l Co. v. Teleflex Inc., 550 U.S. 398, 415-421, 82 USPQ2d 1385, 1395-97 (2007). The combination of Ganong in view of Bachtiger provides teaching for generating alignment information and the training of a speech recognition model based on transcribed speech utterances and modified speech utterances. This combination differs from the claimed invention in that the claimed invention further provides teaching for generating alignment output information that comprises a textual representation mapping unspoken texts to speech frames, the training of a speech recognition model on the generated alignment outputs, wherein the speech recognition model includes an audio encoder. This however isn’t new to the art as the reference of Chen is now introduced to teach this as: generating, using an alignment model, a corresponding alignment output for each unspoken textual utterance of the received training data, wherein the alignment output comprises a textual representation having a number of frames that maps the unspoken textual utterance to speech frames, and wherein the alignment model generates the alignment output without converting the unspoken textual utterance into synthetic speech (Chen: Page 2 Col 2 Section 3.2.2 — ‘When learning from unspoken text, the speech-text alignment information is unavailable. We substitute it with durations predicted from a duration prediction model in a fashion similar to speech synthesis [22]. This model is trained on any available paired data to predict the duration of each token. The predicted duration on unspoken text is subsequently used to resample the initially learned text embeddings’ (teaching of mapping the unspoken text to predicted durations as the claimed speech frames, such that there is predicted duration information or frame information, that corresponds to the text tokens or unspoken textual utterances)); and training a speech recognition model on the transcribed speech utterances, the modified speech utterances, and the alignment outputs generated for the unspoken textual utterances (Chen: Page 2 Col 1 Sections 3.1, 3.2 — pre-training and training processes that include unspoken text and modality matching; Page 2 Col 2 Section 3.2.2 — when learning from unspoken text, the predicted duration on unspoken text is used (indicating a training process that uses unspoken text and their predicted duration, indicating an alignment output of the unspoken textual utterances) [[to teach the speech recognition model to learn to recognize speech in the target domain and phrases within the one or more classes of sensitive information]], wherein the speech recognition model comprises an audio encoder and a decoder, the audio encoder comprising a stack of self-attention layers each including a multi-headed self-attention mechanism (Chen: Page 3 Col 1 Section 4.2 — an encoder-decoder architecture, the encoder comprising layers of 8-headed self-attention blocks). Hence, before the effective filing date of the claimed invention, one of ordinary skill in the art would have found it obvious to improve upon the teaching of the combination of Ganong in view of Bachtiger which generates alignment information and trains a speech recognition model based on transcribed speech utterances and modified speech utterances, by applying the known technique of Chen which trains a speech recognition model based on unspoken text aligned with duration information, the training further comprising an encoder-decoder architecture with an encoder comprising a stack of self-attention layers with a multi-headed self-attention mechanism, to thereby come up with the claimed invention. The combination of both prior art elements would have resulted in the predictable result of the presence of self-attention layers which capture relationships between different and separate parts of a speech signal, leading to the generation of a model that can connect acoustic patterns with corresponding unspoken textual units over duration/frame information. See KSR Int’l Co. v. Teleflex Inc., 550 U.S. 398, 415-421, 82 USPQ2d 1385, 1395-97 (2007). For claim 2, claim 1 is incorporated and the combination of Ganong in view of Bachtiger further in view of Chen discloses the method of claim 1, wherein the one or more classes of sensitive information comprises at least one of personably identifiable information, protected health information, or dates (Ganong: [0055] — the sensitive information includes personally identifiable information). For claim 5, claim 1 is incorporated and the combination of Ganong in view of Bachtiger further in view of Chen discloses the method, wherein the transcribed speech utterances in the general domain comprise a greater number of hours of speech than the modified speech utterances (Ganong: [0098] — ‘In some implementations, transcription generation process 10 may selectively remove sensitive content from the training of a speech processing system. For example, some or all of sensitive content may be identified as noted above but instead of providing the sensitive content for transcription, transcription generation process 10 may dispose of the one or more sensitive content signals (e.g., sensitive content signals 1112). In this manner, a speech processing system (e.g., speech processing system 514) may be trained using only non-sensitive content from the input speech signal. (this indicating that the amount of the transcribed speech utterances containing the sensitive material would inherently contain more data than the utterances without the sensitive material, and thereby leading to longer hours of data)). For claim 7, claim 1 is incorporated and the combination of Ganong in view of Bachtiger further in view of Chen discloses the method, wherein the training data further comprises un-transcribed speech utterances spoken in the general domain, each un-transcribed speech utterance not paired with any corresponding transcription (Chen: Page 3 Col 2 Section 4.3 — training data that includes untranscribed speech from SpeechStew (SpeechStew being a database that contains multiple public speech datasets to thereby indicate general domain information, and, untranscribed speech also indicates speech that isn’t paired with a corresponding transcription)). For claim 8, claim 1 is incorporated and the combination of Ganong in view of Bachtiger further in view of Chen discloses the method, wherein training the speech recognition model comprises: for each un-transcribed speech utterance: generating a corresponding encoded representation of the un-transcribed utterance (Chen: Page 3 Col 1 Section 3.3 — using contrastive loss on speech encoder outputs to pretrain the model on untranscribed speech data (indicating the generation of encoded untranscribed utterance representations) (See also from Figure 1 that the speech encoder receives untranscribed speech/utterance)); and training the audio encoder on a contrastive loss applied on the corresponding encoded representation of the un-transcribed speech utterance (Chen: Page 3 Col 1 Section 3.3 — using contrastive loss on speech encoder outputs to pretrain the model on untranscribed speech data; Page 2 Col 2 Section 3.2.1 — training a speech encoder on contrastive loss); for each alignment output: generating a corresponding encoded representation of the alignment output (Chen: Page 3 Col 1 Section 3.3 — using contrastive loss on speech encoder outputs to pretrain the model (See also from Figure 1 that the speech encoder receives Paired Speech and text, the paired text here being an indication of the alignment output, thereby resulting in producing an encoded representation of the alignment output)); and training the audio encoder on a contrastive loss applied on the corresponding encoded representation of the alignment output (Chen: Page 3 Col 1 Section 3.3 — using contrastive loss on speech encoder outputs to pretrain the model; Page 2 Col 2 Section 3.2.1 — training a speech encoder on contrastive loss); and for each transcribed speech utterance: generating a corresponding encoded representation of the transcribed speech utterance (Chen: Page 3 Col 1 Section 3.3 — using contrastive loss on speech encoder outputs to pretrain the model (See also from Figure 1 that the speech encoder receives Paired Speech and text, the paired speech here being an indication of the transcribed speech utterance)); and training the audio encoder on a contrastive loss applied on the corresponding encoded representation of the transcribed speech utterance (Chen: Page 3 Col 1 Section 3.3 — using contrastive loss on speech encoder outputs to pretrain the model; Page 2 Col 2 Section 3.2.1 — training a speech encoder on contrastive loss). For claim 9, claim 1 is incorporated and the combination of Ganong in view of Bachtiger further in view of Chen discloses the method, wherein the decoder comprises one of a Connection Temporal Classification (CTC) decoder, a Listen Attend Spell (LAS) decoder, or Recurrent Neural Network-Transducer (RNN-T) decoder (Chen: Page 3 Col 1 – 2 Section 4.2 — RNN-T decoder, CTC-decoder, LAS-decoder). For claim 10, claim 1 is incorporated and the combination of Ganong in view of Bachtiger further in view of Chen discloses the method wherein generating the corresponding alignment output for each unspoken textual utterance of the received training (Chen: Page 2 Col 2 Section 3.2.2 — generating alignment data between predicted text and speech encoder output) data comprises: extracting an initial textual representation from the unspoken textual utterance (Chen: Page 2 Col 2 Section 3.2.2 — replicating initially learned text embeddings to match durations of speech embeddings (the text embeddings being the textual representation from unspoken textual utterances); Page 3 Col 1 Section 4.2 — a text embedding extractor (extracting text embeddings)); predicting a text chunk duration for each text chunk in the unspoken textual utterance (Chen: Page 2 Col 2 Section 3.2.2 — obtaining predicted durations for each text token); and upsampling the initial textual representation using the predicted text chunk duration for each text chunk in the unspoken textual utterance (Chen: Page 3 Col 1 Section 4.2 — upsampling the original text embedding to the target length of the specified duration). As for claim 11, system claim 11 and method claim 1 are related as apparatus and the method of using same, with each claimed element’s function corresponding to the claimed method step. Ganong in [0099] provides that the disclosure may take an entirely hardware embodiment, with [0020] providing one or more processor s as well as memory architectures, all these being suitable to read upon the limitations of this claim. Accordingly, claim 11 is similarly rejected under the same rationale as applied above with respect to method claim 1. As for claim 12, system claim 12 and method claim 2 are related as apparatus and the method of using same, with each claimed element’s function corresponding to the claimed method step. Accordingly, claim 12 is similarly rejected under the same rationale as applied above with respect to method claim 2. As for claim 15, system claim 15 and method claim 5 are related as apparatus and the method of using same, with each claimed element’s function corresponding to the claimed method step. Accordingly, claim 15 is similarly rejected under the same rationale as applied above with respect to method claim 5. As for claim 17, system claim 17 and method claim 7 are related as apparatus and the method of using same, with each claimed element’s function corresponding to the claimed method step. Accordingly, claim 17 is similarly rejected under the same rationale as applied above with respect to method claim 7. As for claim 18, system claim 18 and method claim 8 are related as apparatus and the method of using same, with each claimed element’s function corresponding to the claimed method step. Accordingly, claim 18 is similarly rejected under the same rationale as applied above with respect to method claim 8. As for claim 19, system claim 19 and method claim 9 are related as apparatus and the method of using same, with each claimed element’s function corresponding to the claimed method step. Accordingly, claim 19 is similarly rejected under the same rationale as applied above with respect to method claim 9. As for claim 20, system claim 20 and method claim 10 are related as apparatus and the method of using same, with each claimed element’s function corresponding to the claimed method step. Accordingly, claim 20 is similarly rejected under the same rationale as applied above with respect to method claim 10. Claims 3, 4, 13 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Ganong (US 2023/0395063 A1) in view of Bachtiger (US 11,120,199 B1) further in view of Chen, et al. (“Maestro: Matched speech text representations through modality matching.” arXiv preprint arXiv:2204.03409) as applied to claim 1 and 11, in view of Gaeta et al. (US 10,002,639 B1: hereafter — Gaeta). For claim 3, claim 1 is incorporated and the combination of Ganong in view of Bachtiger further in view of Chen provides teaching for identifying classes of sensitive information in an utterance. This combination however fails to teach the further limitations of this, for which the reference of Gaeta is now introduced to teach as the method, wherein the redacted portions of the transcriptions of the modified speech utterances are tagged with a class identifier identifying the class of sensitive information that has been redacted (Gaeta: FIG. 1, Col 2 line 66 – Col 3 line 2 — a table showing text that was identified as confidential and an indication of the type of the confidential information identified). The combination of Ganong in view of Bachtiger further in view of Chen provides teaching for identifying classes of sensitive information in an utterance, but differs from the claimed invention in that the claimed invention further provides teaching for the transcriptions to be tagged with a class identifier identifying the class of sensitive information being redacted. The reference of Gaeta is however introduced to teach this, as presented above. Hence, before the effective filing date of the claimed invention, one or ordinary skill in the art would have found it obvious to combine the known teaching of Gaeta which teaches the clear identification/tagging of the type of sensitive/confidential information, with the teaching of simply identifying sensitive information in speech as taught by the combination of Ganong in view of Bachtiger further in view of Chen, to thereby come up with the claimed invention. The combination of both prior art elements would have provided the predictable result of clearly informing a user observing the redacted transcription of the class of sensitive information encountered in an utterance. See KSR Int’l Co. v. Teleflex Inc., 550 U.S. 398, 415-421, 82 USPQ2d 1385, 1395-97 (2007). For claim 4, claim 3 is incorporated and the combination of Ganong in view of Bachtiger further in view of Chen and further in view of Gaeta teaches the method, wherein the fake random data inserted into each redacted portion of the transcriptions of the modified speech utterances is associated with the class of sensitive information identified by the class identifier at the redacted portion (Ganong: [0056] — replacing patients’ names, date of birth, medical history/prescription dosage information (the fake data inserted into each redacted portion being the same class of sensitive information that’s being redacted); Bachtiger: Col 4 lines 44–48 — a redaction module which is able to substitute redacted portions of an utterance with random series of symbols such as ***-**-*** or #@$-$#-@#$#@ (which are the unspoken textual utterances that are not paired with any corresponding audio representation)). As for claim 13, system claim 13 and method claim 3 are related as apparatus and the method of using same, with each claimed element’s function corresponding to the claimed method step. Accordingly, claim 13 is similarly rejected under the same rationale as applied above with respect to method claim 3. As for claim 14, system claim 14 and method claim 4 are related as apparatus and the method of using same, with each claimed element’s function corresponding to the claimed method step. Accordingly, claim 14 is similarly rejected under the same rationale as applied above with respect to method claim 4. Conclusion The prior art made of record and not relied upon is considered pertinent to Applicant’s disclosure. DOGGETT et al. (US 2021/0104241 A1) provides teaching for an alignment engine that generates a relevance time window that is duration-aligned to a text model output [0057]. Any inquiry concerning this communication or earlier communications from the Examiner should be directed to OLUWADAMILOLA M. OGUNBIYI whose telephone number is (571)272-4708. The Examiner can normally be reached Monday - Thursday (8:00 AM - 5:30 PM Eastern Standard Time). Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, Applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the Examiner by telephone are unsuccessful, the Examiner’s Supervisor, PARAS D. SHAH can be reached at (571) 270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /OLUWADAMILOLA M OGUNBIYI/ Examiner, Art Unit 2653 1 This reference appears to share an Applicant — ‘Google, Inc.’, with the instant application’s ‘Google LLC’, as well as some authors with some of the joint inventors of the instant application. This reference however qualifies as prior art under 35 U.S.C. 102(a)(1), and is being applied here in such a fashion. The limitations which this reference is applied to, also do not appear to be presented in the provisional application 63/487,861.
Read full office action

Prosecution Timeline

Feb 12, 2024
Application Filed
Sep 23, 2025
Non-Final Rejection mailed — §101, §102, §103
Dec 22, 2025
Response Filed
Mar 27, 2026
Final Rejection mailed — §101, §102, §103
May 08, 2026
Request for Continued Examination
May 09, 2026
Response after Non-Final Action
Sep 10, 2026
Non-Final Rejection mailed — §101, §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12725620
AUDIO ENCODER AND DECODER USING A FREQUENCY DOMAIN PROCESSOR , A TIME DOMAIN PROCESSOR, AND A CROSS PROCESSING FOR CONTINUOUS INITIALIZATION
3y 0m to grant Granted Sep 01, 2026
Patent 12718584
SENSOR FUSION FOR COLLISION DETECTION
3y 0m to grant Granted Aug 25, 2026
Patent 12694891
METHOD, DEVICE AND COMPUTER PROGRAM FOR EMOTION RECOGNITION FROM A REAL-TIME AUDIO SIGNAL
3y 7m to grant Granted Jul 28, 2026
Patent 12640154
Stylizing Text-to-Speech (TTS) Voice Response for Assistant Systems
3y 5m to grant Granted May 26, 2026
Patent 12608427
Drill Back To Original Audio Clip In Virtual Assistant Initiated Lists And Reminders
1y 11m to grant Granted Apr 21, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
77%
Grant Probability
96%
With Interview (+19.4%)
2y 11m (~3m remaining)
Median Time to Grant
High
PTA Risk
Based on 315 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month