Prosecution Insights
Last updated: October 02, 2026
Application No. 18/788,847

SYSTEMS AND METHODS FOR DETERMINING WHETHER TO TRIGGER A VOICE CAPABLE DEVICE BASED ON SPEAKING CADENCE

Final Rejection §103
Filed
Jul 30, 2024
Priority
Sep 24, 2018 — continuation of 10/861,444 +2 more
Examiner
SONIFRANK, RICHA MISHRA
Art Unit
2654
Tech Center
2600 — Communications
Assignee
Adeia Technologies Inc.
OA Round
2 (Final)
67%
Grant Probability
Favorable
3-4
OA Rounds
10m
Est. Remaining
92%
With Interview

Examiner Intelligence

Grants 67% — above average
67%
Career Allowance Rate
265 granted / 395 resolved
+5.1% vs TC avg
Strong +24% interview lift
Without
With
+24.4%
Interview Lift
resolved cases with interview
Typical timeline
3y 0m
Avg Prosecution
23 currently pending
Career history
418
Total Applications
across all art units

Statute-Specific Performance

§101
15.4%
-24.6% vs TC avg
§103
63.3%
+23.3% vs TC avg
§102
8.7%
-31.3% vs TC avg
§112
8.1%
-31.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 395 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority This application is a continuation of U.S. Patent Application No. 18/105,166, filed February 2, 2023, which is a continuation of U.S. Patent Application No. 17/089,357, filed November 4, 2020, now U.S. Patent No. 11,600,265, which is a continuation of U.S. Patent Application No. 16/139,453, filed September 24, 2018, now U.S. Patent No. 10,861,444 Response to Amendment Claims 1 and 11 are amended. Claims 1-20 are presented for examination. Response to Arguments Applicant arguments filed on 8/27/2026 have been considered. Following are the response: Rejections Under 35 U.S.C. § 101 In light of amendments and arguments rejection under § 101 is withdrawn. Rejections Under 35 U.S.C. § 103 Applicant’s arguments with respect to claims 1-20 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. And KSR, 550 U.S. at 418, 82 USPQ2d at 1396. Exemplary rationales that may support a conclusion of obviousness include: (A) Combining prior art elements according to known methods to yield predictable results; (B) Simple substitution of one known element for another to obtain predictable results; (C) Use of known technique to improve similar devices (methods, or products) in the same way; (D) Applying a known technique to a known device (method, or product) ready for improvement to yield predictable results; (E) "Obvious to try" – choosing from a finite number of identified, predictable solutions, with a reasonable expectation of success; (F) Known work in one field of endeavor may prompt variations of it for use in either the same field or a different one based on design incentives or other market forces if the variations are predictable to one of ordinary skill in the art; (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. See MPEP § 2143 for a discussion of the rationales listed above along with examples illustrating how the cited rationales may be used to support a finding of obviousness. See also MPEP § 2144 - § 2144.09 for additional guidance regarding support for obviousness determination. Claims 1-3, 5-8, 11-13 and 15-18 are rejected under 35 U.S.C. 103 as being unpatentable over Grima (US 20190295540) and further in view of Gruenstein ( US 20180233150) and further in view of Riediger (US 20150074507 ) Regarding claim 1, Grima teaches a method comprising: monitoring, using control circuitry of a voice capable device, for a plurality of voice inputs from a user to activate the voice capable device ( monitor ..for trigger phrase, Para 0006-007; trigger phrase activates the device), wherein the plurality of voice inputs are received using a microphone associated with the voice capable device ( microphone, Para 0007) ; determining, using the control circuitry, an position of a trigger word associated with a plurality of sentences respectively corresponding to the plurality of voice inputs ( Storing the voice trigger events as either validated or invalidated may provide a useful database of voice trigger events, from which the voice trigger validator is able to learn in order to further improve validation accuracy., Para 0018, 0067; voice trigger event can be trigger phrase, 0018); determining, using the control circuitry, a threshold maximum position of the trigger word based at least in part on the average position of the trigger word, wherein the threshold maximum position corresponds to the average number of other words ( threshold amount of time, Para 0066-0067- start or a specific pattern); receiving, using the microphone associated with the voice capable device, a voice input from the user to activate the voice capable device, wherein the voice input comprises the trigger word ( trigger word, Para 0018); determining, using the control circuitry, a position of the trigger word in the received voice input ( position of the trigger word in a voice input, Para 0056) ; and based at least in part on determining that the position of the trigger word in the received voice input is greater than the threshold maximum position of the trigger word, refraining from activating the voice capable device ( The counter 23-2 will time out if no start of speech is detected within a certain expected (predetermined or user-defined) period. If a trigger follows the start of speech, or vice versa, without the expected period then the trigger phrase validation step would then be activated and based on the counters 23-1 and 23-2. The trigger phrase validation block 24 may then indicate to a pass gate (driver) 25 that a trigger phrase has occurred, the pass gate 25 then in turn may allow the buffered trigger phrase to pass,, Para 0072, Fig 13; whether to command for certain actions, Para 0008-0010) Grima does not explicitly teach determining, using the control circuitry, an average time of a trigger word in relation to an average number of other words associated with a plurality of sentences respectively corresponding to the plurality of inputs and determining, using the control circuitry, a position of the trigger word in the received input, wherein the voice input comprises one or more sentences, and wherein the time of the trigger word is determined in relation to other words of the one or more sentences However, Gruenstein teaches determining, using the control circuitry, an average time of a trigger word in relation to an average number of other words associated with a plurality of sentences respectively corresponding to the plurality of inputs and determining, using the control circuitry, a position of the trigger word in the received input, wherein the voice input comprises one or more sentences, and wherein the time of the trigger word is determined in relation to other words of the one or more sentences Grima modified by Gruenstein does not teach determining, using the control circuitry, an average position of a trigger word in relation to an average number of other words associated with a plurality of sentences respectively corresponding to the plurality of inputs (when the client hotword detection module 106 determines that one or a portion of one of the first utterances satisfies the first threshold 108 of being a portion of the key phrase, the client hotword detection module 106 may determine whether a total length of the first utterances matches a length for the key phrases. For instance, the client hotword detection module 106 may determine that a time during which the one or more first utterances were spoken matches an average time for the key phrase to be spoken. The average time may be for a user of the client device 102 or for multiple different people, e.g., including the user of the client device 102, Para 0027) and determining, using the control circuitry, a position of the trigger word in the received input, wherein the position of the trigger word is determined in relation to other words of the one or more sentences(the client hotword detection module 106 may determine that a time during which the one or more first utterances were spoken matches an average time for the key phrase to be spoken. The average time may be for a user of the client device 102 or for multiple different people, e.g., including the user of the client device 102, Para 0027, 0024, 0030, 0042) It would have been obvious to POSITA having the teachings of Grima to further include the concept of Gruenstein before effective filing date so to determine whether the hotword was spoken ( Para 0027, Gruenstein) However, Riediger teach determining, using the control circuitry, an average position of a trigger word in relation to an average number of other words associated with a plurality of sentences respectively corresponding to the plurality of voice inputs and determining, using the control circuitry, a position of the trigger word in the received input, and wherein the position of the trigger word is determined in relation to other words of the one or more sentences ((e) calculating an averaged location for each annotated word, the averaged location defining an averaged placement of an annotated word within a target document; (f) identifying matches between the distance and averaged location of words of the target document and the distances and averaged location of at least one of the annotated words of the training documents; , Para 0007, 0108-0109) It would have been obvious to POSITA having the concept of Grima to further include the idea of determining a average position of a particular word in relation to word to align the particular word within a document and/or input (abstract, Riediger) Regarding claim 2, Grima as above in claim 1, teaches wherein the determining that the position of the trigger word in the received voice input is greater than the threshold maximum position of the trigger word comprises comparing the position of the trigger word in the received voice input to the threshold maximum position ( if the voice trigger is deemed to occur beyond the threshold amount of time from the start of speech, the voice trigger is ignored, Para 0055-0057) . Regarding claim 3, Grima modified by as above in claim 1, teaches wherein the determining that the position of the trigger word in the received voice input is greater than the threshold maximum position of the trigger word comprises: detecting a plurality of utterances comprising the trigger word from the plurality of voice inputs ( plural utterances, monitoring, Para 0066-0067, Grima; monitoring training, Para 0032, Gruenstein) ; storing, in a profile of the user, a subset of utterances from the plurality of utterances comprising the trigger word, wherein the subset of utterances comprises the trigger word in a position that is within the threshold maximum position of the trigger word ( learning based on the timing of the utterances, Para 0018, 0067, Grima; As shown in FIG. 4, information may be stored at block 420 for use in customizing wake word detection models. The information may include usage data, such as the audio provided by the client device 200, features extracted from the audio, ASR and/or NLU results generated using the audio or a subsequent audio stream, contextual information, and the like. The customization module 212 or some other module or component can use this information to retrain or update the wake word detection model, or to generate a new wake word detection model. In this way, the wake word detection model may be based on real-world data observed by the system that uses the model. As shown in FIG. 3, a customized wake word detection model may be provided to the user device 200 at [8]. The user device 200 can implement the customized wake word detection model at [9] by, e.g., replacing an existing detection model with the customized wake word detection model, Col 11, line 5-25- Fig 4 shows the utterances are stored to train the model which includes the utterance ( subset ) with wake-words; Interpretation – claim does not say the utterance which does not include the wakeword are not stored. ) ; and comparing the received voice input to each utterance of the stored subset of utterances ( valid or invalid based on learning, Para 0018, 0067, Grima; comparing if the position matches, Para 0027, 0032, Gruenstein) Regarding claim 5, Grima as above in claim 1, teaches wherein the determining that the position of the trigger word in the received voice input is greater than the threshold maximum position indicates that an inclusion of the trigger word in the received voice input is not intended to activate the voice capable device ( after the time the trigger is not activated ( invalidated) , Para 0018, Para 0069, 0073-0074) Regarding claim 6, Grima modified by Gruenstein as above in claim 1, teaches determining a threshold minimum position of the trigger word based at least in part on the average position of the trigger word; and based at least in part on determining that the position of the trigger word in the received voice input is greater than the threshold minimum position of the trigger word, activating the voice capable device ( The trigger phrase validation block 24 may then indicate to a pass gate (driver) 25 that a trigger phrase has occurred, the pass gate 25 then in turn may allow the buffered trigger phrase to pass, along with the associated audio data, to the speech recognition engine 26. The speech recognition engine 26 is operable to carry out further functions based on instructions spoken by a user, contained in the audio data, Para 0072; wherein based on timing the trigger phrase is validated, Para 0018, 0067 Fig 13-15; position or time of the hotword, Para 0027, Gruenstein) Regarding claim 7, as above in claim 6, teaches , wherein the activating the voice capable device comprises: determining whether a portion of the received voice input matches a word associated with a function of the voice capable device ( trigger, Para 0018, 0067-0069, Fig 13-16) ; and in response to determining that the portion of the voice input matches the word associated with the function of the voice capable device, performing the function of the voice capable device ( the trigger phrase validation block 24 may then indicate to a pass gate (driver) 25 that a trigger phrase has occurred, the pass gate 25 then in turn may allow the buffered trigger phrase to pass, along with the associated audio data, to the speech recognition engine 26. The speech recognition engine 26 is operable to carry out further functions based on instructions spoken by a user, contained in the audio data, Para 0072; wherein based on timing the trigger phrase is validated, Para 0018, 0067 Fig 13-15) Regarding claim 8, Grima as above in claim 1, teaches wherein the average position of the trigger word is the threshold maximum position of the trigger word ( position/time of the trigger word, Fig 12-16, Grima; position, Para 027 Gruenstein) Regarding claim 11, arguments analogous to claim 1, are applicable. In addition, Grima teaches memory and input/output circuitry ( Fig 1-2) Regarding claim 12, arguments analogous to claim 2, are applicable. Regarding claim 13, arguments analogous to claim 3, are applicable. Regarding claim 15, arguments analogous to claim 5, are applicable. Regarding claim 16, arguments analogous to claim 6, are applicable. Regarding claim 17, arguments analogous to claim 7, are applicable. Regarding claim 18, arguments analogous to claim 8, are applicable. Claims 4, 9, 14 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Grima ( US 20190295540) and further in view of Gruenstein ( US 20180233150) and further in view of Riediger (US 20150074507 ) and further in view of Chao ( US 20200135184) Regarding claim 4, Grima modified by Gruenstein , as above in claim 1, teaches further comprising: and determining that the user is attempting to activate the voice capable device ( trigger phrase validation, Para 0069, 0073) , receiving a value indicating the threshold maximum position in the received voice input where the trigger word appears ( based on the timing, Para 0073-0074, Grima; average time, Para 0027, Gruenstein; further Gruenstein teaches a server side system; ) Grima modified by Gruenstein and Riediger does not teach receiving the voice input from the user; transmitting a region and a language associated with the user to a database; and based at least in part on the transmitting, determine whether the trigger word was spoken However, Chao teaches receiving the voice input from the user; transmitting a region and a language associated with the user to a database; and based at least in part on the transmitting, determine whether the trigger word was spoken ( preemptively load user data before invocation, Para 0023; the user 214 can be operating an application 206 (i.e., APPLICATION_1) through a portable computing device 216, which provides basis for the assistant device 212 to select a particular language model. Alternatively, or additionally, the assistant device 212 can select a language model based on the user 214 being at a location 210. The table 220, or the user profile corresponding to the user 214, can provide a correspondence between a score for a language model and a context of the application and or the location. By identifying the context in which the user 214 is invoking the automated assistant, and comparing the contacts to the table 220, the assistant device 212 can determine the language model that has the highest score for the user 214, Para 0059, 0060, 0067; whether user is trying to invoke, Para 0019, 0053) It would have been obvious having the teachings of Grima modified by Gruenstein and Riediger to further include the concept of Chao to verify the user ( Para 0066, Chao) Regarding claim 14, arguments analogous to claim 4, are applicable. Regarding claim 9, Grima modified by Gruenstein, as above in claim 1, does not explicitly teaches detecting the trigger word based at least in part on a fingerprint associated with the trigger word, However, Chao teaches detecting the trigger word based at least in part on a fingerprint associated with the trigger word ( the score or probability can be based on a context in which the user is interacting with the automated assistant or the assistant device 212. For example, the user 214 can provide a spoken natural language input 218, such as “Assistant,” in order to invoke the automated assistant. The assistant device 212 can include an automated assistant interface that receives the spoken natural language input 218 for further processing at the assistant device 212. The assistant device 212 can employ a language model (e.g., an invocation phrase model) for determining a voice signature based on characteristics of the voice of the user 214. When the assistant device 212 has identified the voice signature of the user 214, the assistant device 212 can access a table 220 that identifies multiple different user profiles, corresponding to multiple different voice signatures, respectively, and a correspondence between the user profiles and different language models., Para 0059) It would have been obvious having the teachings of Grima modified Gruenstein and Riediger to further include the concept of Chao to verify the user ( Para 0066, Chao) Regarding claim 19, arguments analogous to claim 9, are applicable. Claims 10 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Grima ( US 20190295540) and further in view of Gruenstein ( US 20180233150) and further in view of Riediger (US 20150074507 ) and further in view of Mazzoleni( US 20190380660) Regarding claim 10, Grima modified by Gruenstein as above in claim 1, teach retrieving, from a profile of the user, demographic information corresponding to the user; storing, in the profile of the user, the template speaking cadence ( pattern) as the template speaking cadence of the user; comparing a cadence associated with the received voice input to the template speaking cadence (pattern) of the user; and based at least in part on the comparing, determining whether the received voice input intends to activate the voice capable device ( validate based on the speech pattern of the user speaking, and -based on the second speech pattern and the second trigger, the voice capable device is not activated, Para 0058-0059, Fig 8-9) Grima modified by Gruenstein identifying, based at least in part on the demographic information corresponding to the user, a template speaking cadence However, Mazzoleni teaches identifying, based at least in part on the demographic information corresponding to the user, a template speaking cadence ( The cadence threshold is from twenty to one hundred strides per minute and may be based on the demographic data 242, Para 0067, 0071, 0077) It would have been obvious having the teachings of Grima modified by Gruenstein and Riediger to further include the concept of Mazoleni since it well known that the demographic data and cadence and/or speech patterns are related and it would be obvious to use that so to get more information of the user speaking style based on their demographics and the result of Grima to differentiate the speaker would be predicable using this extra information. Regarding claim 20, arguments analogous to claim 10, are applicable. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. US 20160085565 – based on cadence and fingerprint THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Richa Sonifrank whose telephone number is (571)272-5357. The examiner can normally be reached M-T 7AM - 5:30PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Phan Hai can be reached at (571)272-6338. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Richa Sonifrank/Primary Examiner, Art Unit 2654
Read full office action

Prosecution Timeline

Jul 30, 2024
Application Filed
May 04, 2026
Non-Final Rejection mailed — §103
Aug 27, 2026
Response Filed
Sep 16, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748932
COMPRESSION OF WORD EMBEDDINGS FOR NATURAL LANGUAGE PROCESSING SYSTEMS
3y 10m to grant Granted Sep 29, 2026
Patent 12730977
SYSTEMS AND METHODS FOR AUTOMATED COMMUNICATION TRAINING
3y 4m to grant Granted Sep 08, 2026
Patent 12711329
CONTENT TRANSLATION USER INTERFACES
4y 2m to grant Granted Aug 18, 2026
Patent 12711323
QUESTION AND ANSWERING ON DOMAIN-SPECIFIC TABULAR DATASETS
2y 5m to grant Granted Aug 18, 2026
Patent 12705432
NATURAL LANGUAGE GOAL DETERMINATION
3y 1m to grant Granted Aug 11, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
67%
Grant Probability
92%
With Interview (+24.4%)
3y 0m (~10m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 395 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month