Prosecution Insights
Last updated: October 01, 2026
Application No. 19/096,156

DYNAMICALLY DETERMINING WHETHER TO PERFORM CANDIDATE AUTOMATED ASSISTANT ACTION DETERMINED FROM SPOKEN UTTERANCE

Non-Final OA §DP
Filed
Mar 31, 2025
Priority
Aug 08, 2022 — provisional 63/396,108 +1 more
Examiner
CHAWAN, VIJAY B
Art Unit
Tech Center
Assignee
Google LLC
OA Round
1 (Non-Final)
88%
Grant Probability
Favorable
1-2
OA Rounds
1y 0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 88% — above average
88%
Career Allowance Rate
797 granted / 907 resolved
+27.9% vs TC avg
Moderate +12% lift
Without
With
+11.8%
Interview Lift
resolved cases with interview
Typical timeline
2y 6m
Avg Prosecution
8 currently pending
Career history
914
Total Applications
across all art units

Statute-Specific Performance

§101
21.7%
-18.3% vs TC avg
§103
14.4%
-25.6% vs TC avg
§102
34.4%
-5.6% vs TC avg
§112
8.7%
-31.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 907 resolved cases

Office Action

§DP
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Double Patenting The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b). The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13. The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer. Claims 1-20 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1-20 of U.S. Patent No. 12,266,358. Although the claims at issue are not identical, they are not patentably distinct from each other because claims 1-20 of the instant application are directed to device, and are similar in scope and content of the patented method claims 1-20 of the patent issued to the same Applicant. It is clear that all the elements of the application claims 1-20 are to be found in patented claims 1-20 (as the application claims 1-20 fully encompasses patented claims 1-20). The difference between the application claims and the patent claims lies in the fact that the patent claim includes many more elements and is thus much more specific. Thus the invention of claims 1-20 of the patent is in effect a “species” of the “generic” invention of the application claims 1-20. It has been held that the generic invention is “anticipated” by the “species”. See In re Goodman, 29 USPQ2d 2010 (Fed. Cir. 1993). Since application claims 1-20 is anticipated by claims 1-20 of the patent, it is not patentably distinct from of the patented claims. Application No: 19/096,156 Patent No: 12,266,358 1. A client device, comprising: one or more microphones; memory storing instructions; and one or more processors operable to execute the instructions to: process, independent of any explicit invocation of an automated assistant, audio data to generate a candidate automated assistant action and a confidence measure for the candidate automated assistant action, wherein the audio data is detected via one or more of the microphones while the client device is in an environment and wherein the audio data captures a spoken utterance of a user; generate one or more environment features that each reflects a corresponding current value for a corresponding dynamic state of the environment, wherein in generating the one or more environment features one or more of the processors are to generate the one or more environment features based on processing data from the client device and/or from one or more additional client devices in the environment, and wherein the one or more environment features comprise one or more of: a temporal feature indicative of one or more current temporal conditions, a spoken utterance origin feature indicative of an origination location and/or origination direction of the spoken utterance, a quantity of people feature that is indicative of a quantity of people in the environment, a user activity feature that is indicative of one or more activities in which the user is currently engaged, or an environment location feature that is indicative of one or more semantic classifications of the environment; determine, based on processing both the confidence measure for the candidate automated assistant action and the one or more environment features, whether to cause automatic performance of the candidate automated assistant action responsive to the spoken utterance; and in response to determining to cause automatic performance of the candidate automated assistant action: cause automatic performance of the candidate automated assistant action responsive to the spoken utterance; in response to not determining to cause automatic performance of the candidate automated assistant action: suppress any automatic performance of the candidate automated assistant action responsive to the spoken utterance. 1. A method implemented by one or more processors, the method comprising: performing, independent of any explicit invocation of an automated assistant, automatic speech recognition (ASR) on audio data, to generate ASR text that predicts a spoken utterance of a user, wherein the audio data is detected via one or more microphones of a client device in an environment and captures the spoken utterance of the user; generating, based on processing the ASR text: a candidate automated assistant action that corresponds to the ASR text, and a confidence measure for the candidate automated assistant action; generating one or more environment features that each reflects a corresponding current value for a corresponding dynamic state of the environment, wherein generating the one or more environment features is based on processing data from the client device and/or from one or more additional client devices in the environment, and wherein the one or more environment features comprise one or more of: a temporal feature indicative of one or more current temporal conditions, a spoken utterance origin feature indicative of an origination location and/or origination direction of the spoken utterance, a quantity of people feature that is indicative of a quantity of people in the environment, a user activity feature that is indicative of one or more activities in which the user is currently engaged, or an environment location feature that is indicative of one or more semantic classifications of the environment; determining whether to cause automatic performance of the candidate automated assistant action responsive to the spoken utterance, wherein determining whether to cause automatic performance of the candidate automated assistant action is based on processing both: the confidence measure for the candidate automated assistant action, and the one or more environment features; and in response to determining to cause automatic performance of the candidate automated assistant action: causing automatic performance of the candidate automated assistant action responsive to the spoken utterance; in response to not determining to cause automatic performance of the candidate automated assistant action: suppressing any automatic performance of the candidate automated assistant action responsive to the spoken utterance. 2. The client device of claim 1, wherein one or more of the processors are further operable to execute the instructions to: select, based on the candidate automated assistant action and from a plurality of candidate semantic categories, a semantic category for the candidate automated assistant action, wherein the semantic category is a genus category that encompasses a plurality of disparate intents, including an intent of the candidate automated assistant action; wherein determining whether to cause automatic performance of the candidate automated assistant action is further based on processing the semantic category. 2. The method of claim 1, further comprising: selecting, based on the candidate automated assistant action and from a plurality of candidate semantic categories, a semantic category for the candidate automated assistant action, wherein the semantic category is a genus category that encompasses a plurality of disparate intents, including an intent of the candidate automated assistant action; wherein determining whether to cause automatic performance of the candidate automated assistant action is further based on processing the semantic category. 3. The client device of claim 2, wherein in determining whether to cause automatic performance of the candidate automated assistant action one or more of the processors are to: process the confidence measure, the one or more environment features, and the semantic category using a trained machine learning (ML) model to generate ML output; and determine, based on the ML output, whether to cause automatic performance of the candidate automated assistant action. 3. The method of claim 2, wherein determining whether to cause automatic performance of the candidate automated assistant action comprises: processing the confidence measure, the one or more environment features, and the semantic category using a trained machine learning (ML) model to generate ML output; and determining, based on the ML output, whether to cause automatic performance of the candidate automated assistant action. 4. The client device of claim 3, wherein the ML output is a probability and wherein in determining, based on the ML output, whether to cause automatic performance of the candidate automated assistant action one or more of the processors are to compare the probability to a threshold. 4. The method of claim 3, wherein the ML output is a probability and wherein determining, based on the ML output, whether to cause automatic performance of the candidate automated assistant action comprises comparing the probability to a threshold. 5. The client device of claim 2, wherein in determining whether to cause automatic performance of the candidate automated assistant action one or more of the processors are to: identify a rule based on the rule being indexed in association with the semantic category; and determine whether to cause automatic performance of the candidate automated assistant action based on applying the confidence measure and the one or more environment features to the rule. 5. The method of claim 2, wherein determining whether to cause automatic performance of the candidate automated assistant action comprises: identifying a rule based on the rule being indexed in association with the semantic category; and determining whether to cause automatic performance of the candidate automated assistant action based on applying the confidence measure and the one or more environment features to the rule. 6. The client device of claim 2, wherein the candidate automated assistant action comprises an intent and one or more slot values for one or more corresponding slots of the intent. 6. The method of claim 2, wherein the candidate automated assistant action is generated based on performing natural language processing on the ASR text, and comprises an intent and one or more slot values for one or more corresponding slots of the intent. 7. The client device of claim 1, wherein in determining whether to cause automatic performance of the candidate automated assistant action one or more of the processors are to: process the confidence measure and the one or more environment features using a trained machine learning (ML) model to generate ML output; and determine, based on the ML output, whether to cause automatic performance of the candidate automated assistant action. 7. The method of claim 1, wherein determining whether to cause automatic performance of the candidate automated assistant action comprises: processing the confidence measure and the one or more environment features using a trained machine learning (ML) model to generate ML output; and determining, based on the ML output, whether to cause automatic performance of the candidate automated assistant action. 8. The client device of claim 1, wherein the one or more environment features comprise the quantity of people feature that is indicative of a quantity of people in the environment. 8. The method of claim 1, wherein the one or more environment features comprise the quantity of people feature that is indicative of a quantity of people in the environment. 9. The client device of claim 8, wherein in generating the quantity of people feature one or more of the processors are to: process the audio data and/or additional audio data to determine a quantity of unique human voices captured in the audio data and/or the additional audio data, wherein the additional audio data is detected, via the one or more microphones of the client device, prior to detection of the audio data; and generate the quantity of people feature as a function of the quantity of unique human voices. 9. The method of claim 8, wherein generating the quantity of people feature comprises: processing the audio data and/or additional audio data to determine a quantity of unique human voices captured in the audio data and/or the additional audio data, wherein the additional audio data is detected, via the one or more microphones of the client device, prior to detection of the audio data; and generating the quantity of people feature as a function of the quantity of unique human voices. 10. The client device of claim 9, wherein in processing the audio data and/or additional audio data to determine the quantity of unique human voices captured in the audio data and/or the additional audio data one or more of the processors are to: determine that given audio data, that corresponds to a candidate human voice, indicates that the candidate human voice originated from a speaker component as opposed to from a human speaker present in the environment; and filter, from inclusion in the quantity of unique human voices, the candidate human voice in response to determining that the given audio data indicates that the candidate human voice originated from the speaker component. 10. The method of claim 9, wherein processing the audio data and/or additional audio data to determine the quantity of unique human voices captured in the audio data and/or the additional audio data comprises: determining that given audio data, that corresponds to a candidate human voice, indicates that the candidate human voice originated from a speaker component as opposed to from a human speaker present in the environment; and filtering, from inclusion in the quantity of unique human voices, the candidate human voice in response to determining that the given audio data indicates that the candidate human voice originated from the speaker component. 11. The client device of claim 10, wherein in determining that the given audio data indicates that the candidate human voice originated from the speaker component one or more of the processors are to: determine that frequencies, of the given audio data, are all within a given frequency range that indicates origination from a speaker component. 11. The method of claim 10, wherein determining that the given audio data indicates that the candidate human voice originated from the speaker component comprises: determining that frequencies, of the given audio data, are all within a given frequency range that indicates origination from a speaker component. 12. The client device of claim 9, wherein in processing the audio data and/or additional audio data to determine the quantity of unique human voices one or more of the processors are to: process the audio data and/or the additional audio data using a text-independent speaker identification model to generate corresponding speaker embeddings; cluster the corresponding speaker embeddings; and determine the quantity of unique human voices based on a quantity of clusters from the clustering. 12. The method of claim 9, wherein processing the audio data and/or additional audio data to determine the quantity of unique human voices comprises: processing the audio data and/or the additional audio data using a text-independent speaker identification model to generate corresponding speaker embeddings; clustering the corresponding speaker embeddings; and determining the quantity of unique human voices based on a quantity of clusters from the clustering. 13. The client device of claim 12, wherein in determining the quantity of unique human voices based on the quantity of clusters from the clustering one or more of the processors are to: filter, from inclusion in the quantity of unique human voices, a given cluster of the clusters responsive to determining that given audio data, used to generate the corresponding speaker embeddings of the cluster, indicates origination from a speaker component as opposed to origination from a human speaker present in the environment. 13. The method of claim 12, wherein determining the quantity of unique human voices based on the quantity of clusters from the clustering comprises: filtering, from inclusion in the quantity of unique human voices, a given cluster of the clusters responsive to determining that given audio data, used to generate the corresponding speaker embeddings of the cluster, indicates origination from a speaker component as opposed to origination from a human speaker present in the environment. 14. The client device of claim 1, wherein the one or more environment features comprise two or more of: the temporal feature indicative of one or more current temporal conditions; the spoken utterance origin feature indicative of an origination location and/or origination direction of the spoken utterance; the quantity of people feature that is indicative of a quantity of people in the environment; the user activity feature that is indicative of one or more activities in which the user is currently engaged; or the environment location feature that is indicative of one or more semantic classifications of the environment. 14. The method of claim 1, wherein the one or more environment features comprise two or more of: the temporal feature indicative of one or more current temporal conditions; the spoken utterance origin feature indicative of an origination location and/or origination direction of the spoken utterance; the quantity of people feature that is indicative of a quantity of people in the environment; the user activity feature that is indicative of one or more activities in which the user is currently engaged; or the environment location feature that is indicative of one or more semantic classifications of the environment. 15. The client device of claim 1, wherein the client device is a battery powered mobile device and wherein the one or more environment features comprise a human-to-device feature that indicates a physical context between the client device and the user. 15. The method of claim 1, wherein the client device is a battery powered mobile device and wherein the one or more environment features comprise a human-to-device feature that indicates a physical context between the client device and the user. 16. The client device of claim 15, wherein the physical context between the client device and the user, indicated by the human-to-device feature is in a pocket of the user or being held and near a face of the user. 16. The method of claim 15, wherein the physical context between the client device and the user, indicated by the human-to-device feature is in a pocket of the user or being held and near a face of the user. 17. A client device, comprising: one or more microphones; memory storing instructions; and one or more processors operable to execute the instructions to: process, independent of any explicit invocation of an automated assistant, audio data to generate a candidate automated assistant action and a confidence measure for the candidate automated assistant action, wherein the audio data is detected via one or more of the microphones while the client device is in an environment and wherein the audio data captures a spoken utterance of a user; generate one or more environment features that each reflects a corresponding current value for a corresponding dynamic state of the environment, wherein in generating the one or more environment features one or more of the processors are to process data from the client device and/or from one or more additional client devices in the environment; determine whether to cause automatic performance of the first candidate automated assistant action or the second automated assistant action responsive to the spoken utterance, wherein in determining whether to cause automatic performance of the first candidate automated assistant action or the second automated assistant action responsive to the spoken utterance one or more of the processors are to: generate first output based on processing the one or more environment features along with one or more first features of the first automated assistant action, wherein the one or more first features comprise a first confidence measure for the first automated assistant action and/or a first semantic category for the first automated assistant action, generate second output based on processing the one or more environment features along with one or more second features of the second automated assistant action, wherein the one or more second features comprise a second confidence measure for the second automated assistant action and/or a second semantic category for the second automated assistant action, and selecting one of the first candidate action and the second candidate action based on comparing the first output to the second output; and cause, responsive to the spoken utterance, automatic performance of the selected one of the first candidate action and the second candidate action. 17. A method implemented by one or more processors, the method comprising: performing, independent of any explicit invocation of an automated assistant, automatic speech recognition (ASR) on audio data, to generate ASR text that predicts a spoken utterance of a user, wherein the audio data is detected via one or more microphones of a client device in an environment and captures the spoken utterance of the user; generating, based on processing the ASR text: a first candidate automated assistant action that corresponds to the ASR text; a second candidate automated assistant action that corresponds to the ASR text; generating one or more environment features that each reflects a corresponding current value for a corresponding dynamic state of the environment, wherein generating the one or more environment features is based on processing data from the client device and/or from one or more additional client devices in the environment; determining whether to cause automatic performance of the first candidate automated assistant action or the second automated assistant action responsive to the spoken utterance, wherein determining whether to cause automatic performance of the first candidate automated assistant action or the second automated assistant action responsive to the spoken utterance comprises: generating first output based on processing the one or more environment features along with one or more first features of the first automated assistant action, wherein the one or more first features comprise a first confidence measure for the first automated assistant action and/or a first semantic category for the first automated assistant action, generating second output based on processing the one or more environment features along with one or more second features of the second automated assistant action, wherein the one or more second features comprise a second confidence measure for the second automated assistant action and/or a second semantic category for the second automated assistant action, and selecting one of the first candidate action and the second candidate action based on comparing the first output to the second output; and causing, responsive to the spoken utterance, automatic performance of the selected one of the first candidate action and the second candidate action. 18. A client device, comprising: one or more microphones; memory storing instructions; and one or more processors operable to execute the instructions to: process, independent of any explicit invocation of an automated assistant, audio data to generate a candidate automated assistant action and a confidence measure for the candidate automated assistant action, wherein the audio data is detected via one or more of the microphones while the client device is in an environment and wherein the audio data captures a spoken utterance of a user; generate one or more environment features that each reflects a corresponding current value for a corresponding dynamic state of the environment, wherein generating the one or more environment features is based on processing data from the client device and/or from one or more additional client devices in the environment; determine whether to cause automatic performance of the candidate automated assistant action responsive to the spoken utterance, wherein in determining whether to cause automatic performance of the candidate automated assistant action one or more of the processors are to: process the confidence measure and the one or more environment features using a trained machine learning (ML) model to generate ML output; and determine, based on the ML output, whether to cause automatic performance of the candidate automated assistant action; and in response to determining to cause automatic performance of the candidate automated assistant action: cause automatic performance of the candidate automated assistant action responsive to the spoken utterance; in response to not determining to cause automatic performance of the candidate automated assistant action: suppress any automatic performance of the candidate automated assistant action responsive to the spoken utterance. 18. A method implemented by one or more processors, the method comprising: performing, independent of any explicit invocation of an automated assistant, automatic speech recognition (ASR) on audio data, to generate ASR text that predicts a spoken utterance of a user, wherein the audio data is detected via one or more microphones of a client device in an environment and captures the spoken utterance of the user; generating, based on processing the ASR text: a candidate automated assistant action that corresponds to the ASR text, and a confidence measure for the candidate automated assistant action; generating one or more environment features that each reflects a corresponding current value for a corresponding dynamic state of the environment, wherein generating the one or more environment features is based on processing data from the client device and/or from one or more additional client devices in the environment; determining whether to cause automatic performance of the candidate automated assistant action responsive to the spoken utterance, wherein determining whether to cause automatic performance of the candidate automated assistant action comprises: processing the confidence measure and the one or more environment features using a trained machine learning (ML) model to generate ML output; and determining, based on the ML output, whether to cause automatic performance of the candidate automated assistant action; and in response to determining to cause automatic performance of the candidate automated assistant action: causing automatic performance of the candidate automated assistant action responsive to the spoken utterance; in response to not determining to cause automatic performance of the candidate automated assistant action: suppressing any automatic performance of the candidate automated assistant action responsive to the spoken utterance. 19. The client device of claim 18, wherein the one or more environment features comprise a quantity of people feature that is indicative of a quantity of people in the environment. 19. The method of claim 18, wherein the one or more environment features comprise a quantity of people feature that is indicative of a quantity of people in the environment. 20. The client device of claim 19, wherein in generating the quantity of people feature one or more of the processors are to: process the audio data and/or additional audio data to determine a quantity of unique human voices captured in the audio data and/or the additional audio data, wherein the additional audio data is detected, via the one or more microphones of the client device, prior to detection of the audio data; and generate the quantity of people feature as a function of the quantity of unique human voices. 20. The method of claim 19, wherein generating the quantity of people feature comprises: processing the audio data and/or additional audio data to determine a quantity of unique human voices captured in the audio data and/or the additional audio data, wherein the additional audio data is detected, via the one or more microphones of the client device, prior to detection of the audio data; and generating the quantity of people feature as a function of the quantity of unique human voices. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Please see attached form PTO-892. The following is the closest available prior art applicable to Applicant invention Kracun et al., (US 2022/0148601 A1) teach techniques are described herein for multi-factor audio watermarking. A method includes: receiving audio data; processing the audio data to generate predicted output that indicates a probability of one or more hotwords being present in the audio data; determining that the predicted output satisfies a threshold that is indicative of the one or more hotwords being present in the audio data; in response to determining that the predicted output satisfies the threshold, processing the audio data using automatic speech recognition to generate a speech transcription feature; detecting a watermark that is embedded in the audio data; and in response to detecting the watermark: determining that the speech transcription feature corresponds to one of a plurality of stored speech transcription features; and in response to determining that the speech transcription feature corresponds to one of the plurality of stored speech transcription features, suppressing processing of a query included in the audio data. Garcia et al., (US 2019/0295544 A1) teach systems and processes for operating a virtual assistant to provide natural assistant interaction are provided. In accordance with one or more examples, a method includes, at an electronic device with one or more processors and memory: receiving a first audio stream including one or more utterances; determining whether the first audio stream includes a lexical trigger; generating one or more candidate text representations of the one or more utterances; determining whether at least one candidate text representation of the one or more candidate text representations is to be disregarded by the virtual assistant. If at least one candidate text representation is to be disregarded, one or more candidate intents are generated based on candidate text representations of the one or more candidate text representations other than the to be disregarded at least one candidate text representation. Prasad et al., (US 2018/0012593 A1) disclose features for detecting words in audio using contextual information in addition to automatic speech recognition results. A detection model can be generated and used to determine whether a particular word, such as a keyword or “wake word,” has been uttered. The detection model can operate on features derived from an audio signal, contextual information associated with generation of the audio signal, and the like. In some embodiments, the detection model can be customized for particular users or groups of users based usage patterns associated with the users. Any inquiry concerning this communication or earlier communications from the examiner should be directed to VIJAY B CHAWAN whose telephone number is (571)272-7601. The examiner can normally be reached 7-5 Monday thru Thursday. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richemond Dorvil can be reached at 571-272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /VIJAY B CHAWAN/Primary Examiner, Art Unit 2658
Read full office action

Prosecution Timeline

Mar 31, 2025
Application Filed
Sep 23, 2026
Non-Final Rejection mailed — §DP (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12738285
Transform Encoding/Decoding of Harmonic Audio Signals
2y 3m to grant Granted Sep 15, 2026
Patent 12724987
METHOD AND SYSTEM FOR PROCESSING DATA FOR DATA TRANSLATION
2y 4m to grant Granted Sep 01, 2026
Patent 12711308
SYSTEM AND METHOD FOR KNOWLEDGE-BASED AUDIO-TEXT MODELING VIA AUTOMATIC MULTIMODAL GRAPH CONSTRUCTION
2y 3m to grant Granted Aug 18, 2026
Patent 12711952
BACKGROUND AUDIO IDENTIFICATION FOR SPEECH DISAMBIGUATION
2y 3m to grant Granted Aug 18, 2026
Patent 12706087
METHODS AND SYSTEMS FOR VOICE CONTROL
3y 4m to grant Granted Aug 11, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
88%
Grant Probability
99%
With Interview (+11.8%)
2y 6m (~1y 0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 907 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month