Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
The office action sent in response to Applicant’s communication received on 7/30/2024 for the application number 18788847. The office hereby acknowledges receipt of the following placed of record in the file: Specification, Abstract, Oath/Declaration and claims.
Priority
This application is a continuation of U.S. Patent Application No. 18/105,166, filed February 2, 2023, which is a continuation of U.S. Patent Application No. 17/089,357, filed November 4, 2020, now U.S. Patent No. 11,600,265, which is a continuation of U.S. Patent Application No. 16/139,453, filed September 24, 2018, now U.S. Patent No. 10,861,444
Status of the claims
Claims 1-20 are presented for examination.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 7/30/2024 filed before the mailing date of first office action. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 101.
Claim 11includes:
A system comprising: a memory; an input/output (I/O) circuitry; and a control circuitry configured to: (i)monitor for a plurality of voice inputs from a user to activate a voice capable device; (ii)determine an average position of a trigger word based at least in part on the plurality of voice inputs, wherein the average position of the trigger word is stored in the memory; (iii) a threshold maximum position of the trigger word based at least in part on the average position of the trigger word; wherein the I/O circuitry is configured to: (iv)receive a voice input from the user to activate the voice capable device, wherein the voice input comprises the trigger word; wherein the control circuitry is configured to: (v) determine a position of the trigger word in the received voice input; and wherein the I/O circuitry is configured to: (vi) based at least in part on determining that the position of the trigger word in the received voice input is greater than the threshold maximum position of the trigger word, refrain from activating the voice capable device.
Steps (iv) is a mere a data gathering activity
Steps (i), (ii) and (iii) can be performed by human as human can learn a behavior of the person speaking and write the average position of the word another human ask to command some device, for e.g. ask a person to turn on the device and based on writing it multiple times, a human can understand that pattern that if the person uses that word within a particular position they mean to call them otherwise not.
Steps (v) and (vi) can be perform by the human as based on the write down learned pattern of a person a human can determine whether the person mean to call a human or not
Step 1: This part of the eligibility analysis evaluates whether the claim falls within any statutory category. See MPEP 2106.03. The claim recites at least a system hence a machine. Thus, the claim is , recites a statutory categories of invention. (Step 1: YES).
Step 2A, Prong One: This part of the eligibility analysis evaluates whether the claim recites a judicial exception. As explained in MPEP 2106.04, subsection II, a claim “recites” a judicial exception when the judicial exception is “set forth” or “described” in the claim. As discussed above, the broadest reasonable interpretation of steps (i),(ii), (iii) , (v) and (vi) that those steps fall within the mental process groupings of abstract ideas because they cover concepts performed in the human mind, including observation, evaluation, judgment, and opinion. See MPEP 2106.04(a)(2), subsection III. As discussed a human can monitor a other person and learn their specific way of speaking, calculate an average of the position of th word using pen and paper to decide if the word meant for a specific command and act on it the next time a person speaks those words based on the specific position of the word. For e.g. Hey [name] vs this is their [name] or any patterns thereof. Hence, these steps can be performed by a human, using “observation, evaluation, judgment, [and] opinion,” because they involve making doing analysis on the given data which are mental tasks humans routinely do,' ” and thus can practically be performed in the human mind, In re Killian, 45 F.4th 1373, 1379 (Fed. Cir. 2022). Therefore, these limitations are considered together as a abstract idea for further analysis. (Step 2A, Prong One: YES).
Step 2A, Prong Two: This part of the eligibility analysis evaluates whether the claim as a whole integrates the recited judicial exception into a practical application of the exception or whether the claim is “directed to” the judicial exception. This evaluation is performed by (1) identifying whether there are any additional elements recited in the claim beyond the judicial exception, and (2) evaluating those additional elements individually and in combination to determine whether the claim as a whole integrates the exception into a practical application. See MPEP 2106.04(d). Claim requires a memory; an input/output (I/O) circuitry; and receive a voice input from the user to activate the voice capable device and a voice capable device which can be activated or deactivated. As discussed above step (iv) recites a data gathering step of receiving a voice signal which includes a trigger word. It is necessary to acquire the data in order to use the recited judicial exception to perform the judicial exception of acting on the voice signal. The “receiving” element does not impose any other meaningful limits on the claim. Therefore, the additional limitation is insignificant extra-solution activity. See MPEP 2106.05(g). The limitations of input/output circuitry, voice capable device provide nothing more than mere instructions to implement an abstract idea on a generic computer. See MPEP 2106.05(f). MPEP 2106.05(f) provides the following considerations for determining whether a claim simply recites a judicial exception with the words “apply it” (or an equivalent), such as mere instructions to implement an abstract idea on a computer: (1) whether the claim recites only the idea of a solution or outcome i.e., the claim fails to recite details of how a solution to a problem is accomplished; (2) whether the claim invokes computers or other machinery merely as a tool to perform an existing process; and (3) the particularity or generality of the application of the judicial exception. Even when viewed in combination, these additional elements do not integrate the recited judicial exception into a practical application (Step 2A, Prong Two: NO), and the claim is directed to the judicial exception. (Step 2A: YES).
Step 2B: This part of the eligibility analysis evaluates whether the claim as a whole amounts to significantly more than the recited exception, i.e., whether any additional element, or combination of additional elements, adds an inventive concept to the claim. See MPEP 2106.05. At Step 2A, Prong Two, the second additional element of using input/output circuitry was found to represent no more than mere instructions to apply the judicial exception on a computer using generic computer components. The analysis under Step 2A, Prong Two is carried through to Step 2B. Further, the first additional element in step (iv) was found to be insignificant extra-solution activity. However, a conclusion that an additional element is insignificant extra-solution activity in Step 2A should be re-evaluated in Step 2B. See MPEP 2106.05, subsection I.A. At Step 2B, the re-evaluation of the insignificant extra-solution activity consideration takes into account whether or not the extra-solution activity is well understood, routine, and conventional in the field. See MPEP 2106.05(g). Here, the step of receiving a voice signal data gathering that is recited at a high level of generality, and as discussed in the disclosure, is well-understood (background - Oftentimes, to perform a function associated with the voice activated device, a user must speak a trigger word that is used to prepare the voice activated device for performance of a function associated with the voice command.). Therefore, this limitation remains insignificant extra solution activity even upon reconsideration and does not amount to significantly more. Even when considered in combination, these additional elements represent mere instructions to apply an exception and insignificant extra-solution activity, and therefore do not provide an inventive concept (Step 2B: NO). The claim is not eligible.
Regarding claim 1, analysis analogous to claim 11, are applicable.
Regarding 2-8 and 12-18 human can learn a specific behavior pattern of the other person and act on it based on calculating the averages.
Regarding claim 10 and 20, human can determine based on the cadence where the person is from and their speaking style and act on a specific command based on those learned patterns.
Regarding claim 9 and 19- claim 9 and 19 recites an extra solutional activity of determining a fingerprint of the user. However, this is well routine a and conventional and same analysis in under step 2A, prong 2 and step 2b as above in claim 11, are applicable.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
And
KSR, 550 U.S. at 418, 82 USPQ2d at 1396. Exemplary rationales that may support a conclusion of obviousness include:
(A) Combining prior art elements according to known methods to yield predictable results;
(B) Simple substitution of one known element for another to obtain predictable results;
(C) Use of known technique to improve similar devices (methods, or products) in the same way;
(D) Applying a known technique to a known device (method, or product) ready for improvement to yield predictable results;
(E) "Obvious to try" – choosing from a finite number of identified, predictable solutions, with a reasonable expectation of success;
(F) Known work in one field of endeavor may prompt variations of it for use in either the same field or a different one based on design incentives or other market forces if the variations are predictable to one of ordinary skill in the art;
(G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention.
See MPEP § 2143 for a discussion of the rationales listed above along with examples illustrating how the cited rationales may be used to support a finding of obviousness. See also MPEP § 2144 - § 2144.09 for additional guidance regarding support for obviousness determination.
Claims 1-3, 5-8, 11-13 and 15-18 are rejected under 35 U.S.C. 103 as being unpatentable over Grima ( US 20190295540) and further in view of Gruenstein ( US 20180233150) and further in view of Prasad ( US 9697828)
Regarding claim 1, Grima teaches a method comprising: monitoring for a plurality of voice inputs from a user to activate a voice capable device ( monitor ..for trigger phrase, Para 0006-007; trigger phrase activates the device) ; determining position of a trigger word based at least in part on the plurality of voice inputs ( Storing the voice trigger events as either validated or invalidated may provide a useful database of voice trigger events, from which the voice trigger validator is able to learn in order to further improve validation accuracy., Para 0018, 0067; voice trigger event can be trigger phrase, 0018) ; determining a threshold maximum position of the trigger word based at least in part on the average position of the trigger word ( threshold amount of time, Para 0066-0067- start or a specific pattern) ; receiving a voice input from the user to activate the voice capable device, wherein the voice input comprises the trigger word ( trigger word, Para 0018) ; determining a position of the trigger word in the received voice input; and based at least in part on determining that the position of the trigger word in the received voice input is greater than the threshold maximum position of the trigger word, refraining from activating the voice capable device ( The counter 23-2 will time out if no start of speech is detected within a certain expected (predetermined or user-defined) period. If a trigger follows the start of speech, or vice versa, without the expected period then the trigger phrase validation step would then be activated and based on the counters 23-1 and 23-2. The trigger phrase validation block 24 may then indicate to a pass gate (driver) 25 that a trigger phrase has occurred, the pass gate 25 then in turn may allow the buffered trigger phrase to pass,, Para 0072, Fig 13; whether to command for certain actions, Para 0008-0010)
Although Grima mentions learning from the stored pattern which includes timing of the word ( positions of a word) , it does not explicitly teach it’s a an average position
However, Gruenstein teaches determining/learning based an average position of a word (the client hotword detection module 106 may determine that a time during which the one or more first utterances were spoken matches an average time for the key phrase to be spoken. The average time may be for a user of the client device 102 or for multiple different people, e.g., including the user of the client device 102, Para 0027)
It would have been obvious having the teachings of Grima to further include the concept of Gruenstein before effective filing date so to determine whether the hotword was spoken ( Para 0027, Gruenstein)
Grima does not explicitly teach the position based on time
In the same field of endeavor Prasad teaches time determines the location of the word ( position of the word is based on time stamp or pointer of a word, Col 9, line 25-30)
Grima has a base concept of activating a device based on the timing of the word in a sentence. Grima differed by the claimed invention based on the concept of position of the which is related to time. Prasad teaches this concept and it would have been obvious having the teachings of Grima to modify with the concept of Prasad before effective filing date since it would yield predicable result of determining when the key phase is spoken.
Regarding claim 2, Grima as above in claim 1, teaches wherein the determining that the position of the trigger word in the received voice input is greater than the threshold maximum position of the trigger word comprises comparing the position of the trigger word in the received voice input to the threshold maximum position ( if the voice trigger is deemed to occur beyond the threshold amount of time from the start of speech, the voice trigger is ignored, Para 0055-0057) .
Regarding claim 3, Grima modified by Prasad as above in claim 1, teaches wherein the determining that the position of the trigger word in the received voice input is greater than the threshold maximum position of the trigger word comprises: detecting a plurality of utterances comprising the trigger word from the plurality of voice inputs ( plural utterances, monitoring, Para 0066-0067, Grima; monitoring, Col 2, line 15-20, Prasad; monitoring training, Para 0032, Gruenstein) ; storing, in a profile of the user, a subset of utterances from the plurality of utterances comprising the trigger word, wherein the subset of utterances comprises the trigger word in a position that is within the threshold maximum position of the trigger word ( learning based on the timing of the utterances, Para 0018, 0067, Grima; As shown in FIG. 4, information may be stored at block 420 for use in customizing wake word detection models. The information may include usage data, such as the audio provided by the client device 200, features extracted from the audio, ASR and/or NLU results generated using the audio or a subsequent audio stream, contextual information, and the like. The customization module 212 or some other module or component can use this information to retrain or update the wake word detection model, or to generate a new wake word detection model. In this way, the wake word detection model may be based on real-world data observed by the system that uses the model. As shown in FIG. 3, a customized wake word detection model may be provided to the user device 200 at [8]. The user device 200 can implement the customized wake word detection model at [9] by, e.g., replacing an existing detection model with the customized wake word detection model, Col 11, line 5-25- Fig 4 shows the utterances are stored to train the model which includes the utterance ( subset ) with wake-words; Interpretation – claim does not say the utterance which does not include the wakeword are not stored. ) ; and comparing the received voice input to each utterance of the stored subset of utterances ( valid or invalid based on learning, Para 0018, 0067, Grima; comparing if the position matches, Para 0027, 0032, Gruenstein)
Regarding claim 5, Grima as above in claim 1, teaches wherein the determining that the position of the trigger word in the received voice input is greater than the threshold maximum position indicates that an inclusion of the trigger word in the received voice input is not intended to activate the voice capable device ( after the time the trigger is not activated ( invalidated) , Para 0018, Para 0069, 0073-0074)
Regarding claim 6, Grima modified by Gruenstein and Prasad as above in claim 1, teaches determining a threshold minimum position of the trigger word based at least in part on the average position of the trigger word; and based at least in part on determining that the position of the trigger word in the received voice input is greater than the threshold minimum position of the trigger word, activating the voice capable device ( The trigger phrase validation block 24 may then indicate to a pass gate (driver) 25 that a trigger phrase has occurred, the pass gate 25 then in turn may allow the buffered trigger phrase to pass, along with the associated audio data, to the speech recognition engine 26. The speech recognition engine 26 is operable to carry out further functions based on instructions spoken by a user, contained in the audio data, Para 0072; wherein based on timing the trigger phrase is validated, Para 0018, 0067 Fig 13-15; position or time of the hotword, Para 0027, Gruenstein)
Regarding claim 7, as above in claim 6, teaches , wherein the activating the voice capable device comprises: determining whether a portion of the received voice input matches a word associated with a function of the voice capable device ( trigger, Para 0018, 0067-0069, Fig 13-16) ; and in response to determining that the portion of the voice input matches the word associated with the function of the voice capable device, performing the function of the voice capable device ( the trigger phrase validation block 24 may then indicate to a pass gate (driver) 25 that a trigger phrase has occurred, the pass gate 25 then in turn may allow the buffered trigger phrase to pass, along with the associated audio data, to the speech recognition engine 26. The speech recognition engine 26 is operable to carry out further functions based on instructions spoken by a user, contained in the audio data, Para 0072; wherein based on timing the trigger phrase is validated, Para 0018, 0067 Fig 13-15)
Regarding claim 8, Grima as above in claim 1, teaches wherein the average position of the trigger word is the threshold maximum position of the trigger word ( position/time of the trigger word, Fig 12-16, Grima; position, Para 027 Gruenstein)
Regarding claim 11, arguments analogous to claim 1, are applicable. In addition, Grima teaches memory and input/output circuitry ( Fig 1-2)
Regarding claim 12, arguments analogous to claim 2, are applicable.
Regarding claim 13, arguments analogous to claim 3, are applicable.
Regarding claim 15, arguments analogous to claim 5, are applicable.
Regarding claim 16, arguments analogous to claim 6, are applicable.
Regarding claim 17, arguments analogous to claim 7, are applicable.
Regarding claim 18, arguments analogous to claim 8, are applicable.
Claims 4, 9, 14 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Grima ( US 20190295540) and further in view of Prasad ( US 9697828) and further in view of Gruenstein ( US 20180233150) and further in view of Chao ( US 20200135184)
Regarding claim 4, Grima modified by Prasad and Gruenstein , as above in claim 1, teaches further comprising: and determining that the user is attempting to activate the voice capable device ( trigger phrase validation, Para 0069, 0073) , receiving a value indicating the threshold maximum position in the received voice input where the trigger word appears ( based on the timing, Para 0073-0074, Grima; average time, Para 0027, Gruenstein; further Gruenstein teaches a server side system; )
Grima modified by Prasad and Gruenstein does not teach receiving the voice input from the user; transmitting a region and a language associated with the user to a database; and based at least in part on the transmitting, determine whether the trigger word was spoken
However, Chao teaches receiving the voice input from the user; transmitting a region and a language associated with the user to a database; and based at least in part on the transmitting, determine whether the trigger word was spoken ( preemptively load user data before invocation, Para 0023; the user 214 can be operating an application 206 (i.e., APPLICATION_1) through a portable computing device 216, which provides basis for the assistant device 212 to select a particular language model. Alternatively, or additionally, the assistant device 212 can select a language model based on the user 214 being at a location 210. The table 220, or the user profile corresponding to the user 214, can provide a correspondence between a score for a language model and a context of the application and or the location. By identifying the context in which the user 214 is invoking the automated assistant, and comparing the contacts to the table 220, the assistant device 212 can determine the language model that has the highest score for the user 214, Para 0059, 0060, 0067; whether user is trying to invoke, Para 0019, 0053)
It would have been obvious having the teachings of Grima modified by Prasad and Gruenstein to further include the concept of Chao to verify the user ( Para 0066, Chao)
Regarding claim 14, arguments analogous to claim 4, are applicable.
Regarding claim 9, Grima modified by Prasad and Gruenstein, as above in claim 1, does not explicitly teaches detecting the trigger word based at least in part on a fingerprint associated with the trigger word,
However, Chao teaches detecting the trigger word based at least in part on a fingerprint associated with the trigger word ( the score or probability can be based on a context in which the user is interacting with the automated assistant or the assistant device 212. For example, the user 214 can provide a spoken natural language input 218, such as “Assistant,” in order to invoke the automated assistant. The assistant device 212 can include an automated assistant interface that receives the spoken natural language input 218 for further processing at the assistant device 212. The assistant device 212 can employ a language model (e.g., an invocation phrase model) for determining a voice signature based on characteristics of the voice of the user 214. When the assistant device 212 has identified the voice signature of the user 214, the assistant device 212 can access a table 220 that identifies multiple different user profiles, corresponding to multiple different voice signatures, respectively, and a correspondence between the user profiles and different language models., Para 0059)
It would have been obvious having the teachings of Grima modified by Prasad and Gruenstein to further include the concept of Chao to verify the user ( Para 0066, Chao)
Regarding claim 19, arguments analogous to claim 9, are applicable.
Claims 10 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Grima ( US 20190295540) and further in view of Prasad ( US 9697828) and further in view of Gruenstein ( US 20180233150) and further in view of Mazzoleni( US 20190380660)
Regarding claim 10, Grima modified by Prasad and Gruenstein ,retrieving, from a profile of the user, demographic information corresponding to the user; storing, in the profile of the user, the template speaking cadence ( pattern) as the template speaking cadence of the user; comparing a cadence associated with the received voice input to the template speaking cadence (pattern) of the user; and based at least in part on the comparing, determining whether the received voice input intends to activate the voice capable device ( validate based on the speech pattern of the user speaking, and -based on the second speech pattern and the second trigger, the voice capable device is not activated, Para 0058-0059, Fig 8-9)
Grima modified by Prasad and Gruenstein identifying, based at least in part on the demographic information corresponding to the user, a template speaking cadence
However, Mazzoleni teaches identifying, based at least in part on the demographic information corresponding to the user, a template speaking cadence ( The cadence threshold is from twenty to one hundred strides per minute and may be based on the demographic data 242, Para 0067, 0071, 0077)
It would have been obvious having the teachings of Grima modified by Prasad and Gruenstein to further include the concept of Mazoleni since it well known that the demographic data and cadence and/or speech patterns are related and it would be obvious to use that so to get more information of the user speaking style based on their demographics and the result of Grima to differentiate the speaker would be predicable using this extra information.
Regarding claim 20, arguments analogous to claim 10, are applicable.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
US 20160085565 – based on cadence and fingerprint
US 20150074507 – average position of the word is determined
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Richa Sonifrank whose telephone number is (571)272-5357. The examiner can normally be reached M-T 7AM - 5:30PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Phan Hai can be reached at (571)272-6338. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Richa Sonifrank/Primary Examiner, Art Unit 2654