DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on July 14, 2026, has been entered.
Response to Arguments
Applicant's arguments, filed July 14, 2026, with respect to the rejections of claims 1 – 20 under 35 U.S.C. 103 have been fully considered but they are not persuasive.
On pages 7-8 of Applicant’s response, Applicant argues: “Kim discloses performing speaker identification on the audio input containing the trigger phrase to determine whether the speaker is a predetermined user. See Kim, Column 9, lines 15-20. Kim further discloses that, upon verification of the speaker's identity, the virtual assistant processes "audio input received subsequent to the audio input containing the trigger phrase." See Kim, Column 8, lines 23-38. However, Kim does not disclose separately determining that a second portion of the audio data-specifically, the portion containing the multiple terms of the utterance following the hotword-includes speech from the particular user who spoke the hotword. Kim verifies the speaker of the trigger phrase and then processes subsequent audio, but Kim does not verify that the subsequent command audio was also spoken by the verified user. In other words, Kim lacks any disclosure of processing the audio data to determine that the post-hotword command portion includes speech from the same particular user. Perotti does not cure this deficiency. Perotti discloses voiceprint authentication for individual keywords, where "[e]ach of the voiceprints may be associated with a different keyword." See Perotti, Column 8, lines 16-23; Column 9, lines 34-45. However, Perotti does not disclose or suggest determining that a second portion of audio data containing multiple terms following a hotword includes speech from the particular user who spoke the hotword. Perotti's authentication is tied to individual keywords, not to separately verifying that both a hotword portion and a subsequent command portion were spoken by the same user. Since "[a] reference is only good for what it clearly and definitely discloses," a person of ordinary skill in the art would not have been led by the cited references, individually or in combination, to arrive at the claimed invention, because the references fail to disclose all of the claimed limitations. In re Hughes, 145 U.S.P.Q. 467, 471 (C.C.P.A. 1965); In re Moreton, 129 U.S.P.Q. 227, 230 (C.C.P.A. 1961). For at least these reasons, Applicant respectfully submits that independent claims 1 and 11 and their dependent claims are patentable over the cited art. Applicant respectfully requests reconsideration of the pending claims and allowance.”.
However, Perotti (US Patent No. 10,360,916) recites, in column 12, lines 1-17, "In one or more embodiments, voiceprint confidence thresholds may be leveraged in a manner that facilitates user access of the commands 202 of the command library 219, while simultaneously increasing device security. For example, and still referring to FIG. 2, consider a situation in which the first voiceprint 224a includes a given voiceprint confidence threshold, and the second voiceprint 224b includes a different voiceprint confidence threshold. Further, the first keyword 222a may include a wakeup word, which must be matched prior to allowing user access to any other commands 202 (i.e., commands 202b-202n) of the command library 219. Accordingly, the first resource 226a may include a reference to all other keywords 222b-222n of the command library 219. In this way, a user utterance must first successfully match the first keyword 222a and the first voiceprint 224a in order for the user to access the commands 202b-202n.", disclosing matching a wakeup word to a voiceprint of a user, and requiring matching the wakeup word to a voiceprint of a user before allowing the user access to any other commands.
Perotti further recites, in column 11, lines 8-23, "In one or more embodiments, if a given resource 226 is associated with a keyword 222 that is associated with a voiceprint 224, then, in response to a successful keyword matching analysis of a user's utterance relative to the associated keyword 222, and a successful voiceprint comparison of the utterance relative to the associated voiceprint 224, the associated resource 226 may be accessed. In this way, a user may be provided access to the resource 226, or content to which the resource 226 refers. Accordingly, in such embodiments, if a keyword matching analysis and a voiceprint comparison analysis are both performed successfully for a command 202, then an authentication success event has occurred. However, in such embodiments, if either the keyword matching analysis or the voiceprint comparison analysis fails, then the authentication fails and resource access does not occur.", disclosing performing keyword matching analysis of a user's utterance and performing a voiceprint comparison of the utterance relative to a voiceprint associated with a keyword, where user authentication of a command requires matching a keyword for the command and a voiceprint for the keyword.
Perotti further recites, in column 9, lines 46-64, "As described herein, each voiceprint 224 includes the result of a prior analysis of a user speaking the phrase or words of the associated keyword 222. For example, a first voiceprint 224a may include the result of a prior analysis of a given user speaking a first keyword 222a; and a second voiceprint 224b may include the result of a prior analysis of the user speaking a second keyword 222b. In one or more embodiments, the analysis includes an analysis of one or more of a frequency, duration, and amplitude of the user's speech. In this way, each voiceprint 224 may comprise a model, function, or plot derived using such analysis. For example, using the exemplary listing of keywords, above, each of the voiceprints 224a-224g may include, respectively, a result of a prior analysis of a user speaking one of the keywords 222 selected from “play,” “pause,” “stop,” “next track,” “redial,” “call home,” “unlock my phone,” “answer,” “ignore,” “yes,” “no,” etc. Accordingly, each voiceprint 224 identifies elements of a human voice that may be used to uniquely identify the speaker.", disclosing that the voiceprints for authenticating keywords are determined from a particular user speaking the phrase or words of the associated keywords. In Perotti, voiceprints 224a-224g are the voiceprints of a particular user, determined from the particular user speaking the phrase or words of the associated keywords. In the embodiment where the first keyword 222a includes a wakeup word that must be matched prior to allowing user access to any other commands 202, and a user utterance must first successfully match the first keyword 222a and the first voiceprint 224a in order for the user to access the commands 202b-202n, the voiceprint of the particular user is used to determine that the wakeup word was spoken by the particular user. After allowing access to the commands 202b-202n based on determining that the wakeup word was spoken by the particular user, the command utterances are compared to voiceprints from the particular user to authorize the commands. This embodiment is shown in Figure 5 of Perotti:
PNG
media_image1.png
830
1076
media_image1.png
Greyscale
After the user speaks the wakeup phrase at 509, the voice of the user is authenticated at operation 512 using the user’s voiceprint associated with the wakeup phrase. When the user further speaks a command phrase at 517, the user is authenticated at operation 521 using the user’s voiceprint associated with the command phrase.
Perotti further recites, in column 5, lines 37-55, "The embodiments described herein provide for increased device security, while also improving user experience, battery life, and accuracy. For example, referring back to FIG. 1A, in response to a content of the utterance 105, the headset 104 may validate an identify of the user 102, and also present the user 102 with access to a resource on the headset 104 or a resource on the host device 106. In other words, the headset 104 may include multiple overloaded keywords that are not only linked to a functional resource, but also enable the biometric verification of the identity of the user 102. By overloading the keywords of the headset 104, the headset 102 may verify the identity of the user 102 with every command in a sequence of commands received from the user 102. Due to the continuous and recurring authentication of the user 102, the headset 104 can maintain a greater confidence in the identity of the user 102. In this way, the security of the headset 104 may be dramatically improved, and the headset 104 may become a token that can reliably confirm the identity of the user 102.", disclosing that the user is authentication with every command in a sequence of commands received from the user.
Perotti further recites, in column 17, lines 30-44, "As described above, the user 502 may continue to access commands of the headset 504 by speaking keywords. Further, each time the headset 504 detects a keyword within the speech of the user 502, the headset 504 may compare the relevant utterance to a previously stored voiceprint. In this way, the headset 504 may authenticate the user 502 in response to each command spoken by the user 502. Due to the continuous and recurring authentication of the user 502 by the headset 504, the headset 504 can maintain a greater confidence in the identity of the user 502. In this way, the security of the headset 504 may be dramatically improved relative to prior art devices, without negatively impacting the experience of the user 502, and the headset 504 may become a token that can reliably confirm the identity of the user 502 in communications to the host device 506.", disclosing that the user is authenticated the user in response to each command spoken by the user by comparing the relevant utterance to a previously stored voiceprint.
Therefore, the rejections of claims 1 – 20 under 35 U.S.C. 103 as being unpatentable over Kim et al. (US Patent No. 10,127,911), hereinafter Kim, in view of Perotti (US Patent No. 10,360,916), and the rejections of claims 6 – 8 and 16 – 18 under 35 U.S.C. 103 as being unpatentable over Kim in view of Perotti, and further in view of Foerster et al. (US Patent No. 9,418,656) are maintained.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1 – 20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 1 recites the limitation "the particular user" in lines 13, 22, and 25. There is insufficient antecedent basis for this limitation in the claim. This rejection can be overcome by changing "the particular user" in lines 13, 22, and 25 to "the particular person".
Claims 2 – 10 are also rejected as they depend from claim 1 and thus recite the limitations of claim 1, and do not resolve the indefinite language from claim 1.
Claim 11 recites the limitation "the particular user" in lines 16, 25, and 28. There is insufficient antecedent basis for this limitation in the claim. This rejection can be overcome by changing "the particular user" in lines 16, 25, and 28 to "the particular person".
Claims 12 – 20 are also rejected as they depend from claim 11 and thus recite the limitations of claim 11, and do not resolve the indefinite language from claim 11.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1 – 5, 9 – 15 and 19 – 20 are rejected under 35 U.S.C. 103 as being unpatentable over Kim et al. (US Patent No. 10,127,911), hereinafter Kim, in view of Perotti (US Patent No. 10,360,916).
Regarding claim 1, Kim discloses a computer-implemented method that when executed on data processing hardware causes the data processing hardware to perform operations (Column 4, line 61 - Column 5, line 3, "In some examples, a non-transitory computer-readable storage medium of memory 250 can be used to store instructions (e.g., for performing some or all of process 300, 400, 500, 600, or 700, described below) for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, processor-containing system, or other system that can fetch the instructions from the instruction execution system, apparatus, or device, and execute the instructions.") comprising:
receiving audio data characterizing an utterance comprising a hotword followed by multiple terms (Column 6, lines 28-48, "In some examples, process 300 can be performed by a system similar or identical to system 100 having a user device similar or identical to user device 102 configured to implement a virtual assistant capable of continuously (or intermittently over an extended period of time) monitoring an audio input for a receipt of a trigger phrase that initiates activation of the virtual assistant. For example, a user device implementing the virtual assistant can continuously or intermittently monitor sounds, speech, and the like detected by a microphone of the user device without performing an action, such as performing a task flow, generating an output response in an audible (e.g., speech) and/or visual form, or the like, in response to the monitored sounds and speech. However, in response to detecting the trigger phrase, the virtual assistant can perform a speaker identification process to ensure that the speaker of the trigger phrase is the intended operator of the virtual assistant. Upon verification of the identity of the speaker, the virtual assistant can be activated, causing the virtual assistant to process a subsequently received word or phrase and to respond accordingly."; Monitoring an audio input for a receipt of a trigger phrase and processing a subsequently received phrase reads on receiving audio data characterizing an utterance comprising a hotword followed by multiple terms.);
processing the audio data to: identify, using a hotword model, a presence of the hotword in the audio data (Column 6, lines 58-66, "At block 304, speech-to-text conversion can be performed on the audio input received at block 302 to determine whether the audio input includes user speech containing a predetermined trigger phrase. The trigger phrase can include any desired set of one or more predetermined words, such as “Hey Siri.” The trigger phrase can be used to activate the virtual assistant and signal to the virtual assistant that a user input, such as a request, command, or the like, will be subsequently provided."; Determine whether the audio input includes user speech containing a predetermined trigger phrase reads on identify a presence of the hotword in the audio data.);
and determine, using a speaker identification model, that a first portion of the audio data that includes the presence of the hotword spoken by a particular person (Column 6, lines 58-61, "At block 304, speech-to-text conversion can be performed on the audio input received at block 302 to determine whether the audio input includes user speech containing a predetermined trigger phrase.”; Column 9, lines 15-20, "At block 502, the user device can perform a speaker identification process on the audio input received at block 302 of process 300 to determine whether the speaker is a predetermined user (e.g., an authorized user of the device). Any desired speaker identification process can be used, such as an i-vector speaker identification process."; Performing a speaker identification process on the received audio input to determine whether the speaker is a predetermined user reads on determining that the utterance characterized by the audio data was spoken by a particular person.);
based on identifying the presence of the hotword in the audio data and determining that the first portion of the audio data that includes the presence of the hotword was spoken by the particular user: performing speech recognition on the audio data to generate a transcription of the utterance (Column 6, lines 58-61, "At block 304, speech-to-text conversion can be performed on the audio input received at block 302 to determine whether the audio input includes user speech containing a predetermined trigger phrase."; Column 8, lines 23-38, "At block 404, the user device can activate the virtual assistant by processing audio input received subsequent to the audio input containing the trigger phrase. For example, block 404 can include receiving the subsequent audio input, performing speech-to-text conversion on the subsequently received audio input to generate a textual representation of user speech contained in the subsequently received audio input, determining a user intent based on the textual representation, an acting on the determined user intent by performing one or more of the following: identifying a task flow with steps and parameters designed to accomplish the determined user intent; inputting specific requirements from the determined user intent into the task flow; executing the task flow by invoking programs, methods, services, APIs, or the like; and generating output responses to the user in an audible (e.g., speech) and/or visual form."; Column 12, lines 1-12, "At block 602, the user device can perform a speaker identification process on the audio input received at block 302 of process 300 in a manner similar or identical to that of block 502 of process 500. If it is determined that the speaker of the audio input is the predetermined user represented by the speaker profile, then process 600 can proceed to block 604 without adding the audio input to a speaker profile in a manner similar or identical to block 402 of process 400 or block 504 of process 500. At block 604, the virtual assistant can be activated and subsequently received audio input can be processed in a manner similar or identical to block 404 of process 400 or block 506 of process 500."; Performing speech-to-text conversion on the subsequently received audio input to generate a textual representation of user speech contained in the subsequently received audio input after determining that the audio input includes user speech containing a predetermined trigger phrase and determining that the speaker of the audio input is the predetermined user represented by the speaker profile reads on performing speech recognition on the audio data to generate a transcription of the utterance based on identifying the presence of the hotword in the audio data and determining that the first portion of the audio data that includes the presence of the hotword was spoken by the particular user.);
processing, using a command identifier, the transcription of the utterance to determine that the multiple terms of the utterance comprise a command directed toward a user computing device (Column 6, lines 63-66, "The trigger phrase can be used to activate the virtual assistant and signal to the virtual assistant that a user input, such as a request, command, or the like, will be subsequently provided."; Column 8, lines 23-38, "At block 404, the user device can activate the virtual assistant by processing audio input received subsequent to the audio input containing the trigger phrase. For example, block 404 can include receiving the subsequent audio input, performing speech-to-text conversion on the subsequently received audio input to generate a textual representation of user speech contained in the subsequently received audio input, determining a user intent based on the textual representation, an acting on the determined user intent by performing one or more of the following: identifying a task flow with steps and parameters designed to accomplish the determined user intent; inputting specific requirements from the determined user intent into the task flow; executing the task flow by invoking programs, methods, services, APIs, or the like; and generating output responses to the user in an audible (e.g., speech) and/or visual form."; Determining a user intent based on the textual representation, where the user intent is a command, reads on processing the transcription of the utterance to determine that the multiple terms of the utterance comprise a command directed toward a user computing device.);
and based on determining that the multiple terms of the utterance comprise the command [and that the second portion of the audio data that includes the multiple terms of the utterance following the hotword includes speech from the particular user], initiating performance of the command using the user computing device (Column 8, lines 23-38, "At block 404, the user device can activate the virtual assistant by processing audio input received subsequent to the audio input containing the trigger phrase. For example, block 404 can include receiving the subsequent audio input, performing speech-to-text conversion on the subsequently received audio input to generate a textual representation of user speech contained in the subsequently received audio input, determining a user intent based on the textual representation, an acting on the determined user intent by performing one or more of the following: identifying a task flow with steps and parameters designed to accomplish the determined user intent; inputting specific requirements from the determined user intent into the task flow; executing the task flow by invoking programs, methods, services, APIs, or the like; and generating output responses to the user in an audible (e.g., speech) and/or visual form."; Acting on the determined user intent by executing the task flow reads on initiating performance of the command using the user computing device based on determining that the multiple terms of the utterance comprise the command.).
Kim does not specifically disclose: based on identifying the presence of the hotword in the audio data and determining that the first portion of the audio data that includes the presence of the hotword was spoken by the particular user: processing the audio data to determine that a second portion of the audio data that includes the multiple terms of the utterance following the hotword includes speech from the particular user; and based on determining that the multiple terms of the utterance comprise the command and that the second portion of the audio data that includes the multiple terms of the utterance following the hotword includes speech from the particular user, initiating performance of the command using the user computing device.
Perotti teaches:
based on identifying the presence of the hotword in the audio data and determining that the first portion of the audio data that includes the presence of the hotword was spoken by the particular user: processing the audio data to determine that a second portion of the audio data that includes the multiple terms of the utterance following the hotword includes speech from the particular user (Column 11, lines 8-23, "In one or more embodiments, if a given resource 226 is associated with a keyword 222 that is associated with a voiceprint 224, then, in response to a successful keyword matching analysis of a user's utterance relative to the associated keyword 222, and a successful voiceprint comparison of the utterance relative to the associated voiceprint 224, the associated resource 226 may be accessed. In this way, a user may be provided access to the resource 226, or content to which the resource 226 refers. Accordingly, in such embodiments, if a keyword matching analysis and a voiceprint comparison analysis are both performed successfully for a command 202, then an authentication success event has occurred. However, in such embodiments, if either the keyword matching analysis or the voiceprint comparison analysis fails, then the authentication fails and resource access does not occur."; Column 12, lines 1-17, "In one or more embodiments, voiceprint confidence thresholds may be leveraged in a manner that facilitates user access of the commands 202 of the command library 219, while simultaneously increasing device security. For example, and still referring to FIG. 2, consider a situation in which the first voiceprint 224a includes a given voiceprint confidence threshold, and the second voiceprint 224b includes a different voiceprint confidence threshold. Further, the first keyword 222a may include a wakeup word, which must be matched prior to allowing user access to any other commands 202 (i.e., commands 202b-202n) of the command library 219. Accordingly, the first resource 226a may include a reference to all other keywords 222b-222n of the command library 219. In this way, a user utterance must first successfully match the first keyword 222a and the first voiceprint 224a in order for the user to access the commands 202b-202n."; Matching a wakeup word to a voiceprint of a user before allowing user access to any other commands reads on identifying the presence of the hotword in the audio data and determining that the first portion of the audio data that includes the presence of the hotword was spoken by the particular user, and performing a keyword matching analysis and a voiceprint comparison analysis for a command reads on processing the audio data to determine that a second portion of the audio data that includes the multiple terms of the utterance following the hotword includes speech from the particular user.);
and based on determining that the multiple terms of the utterance comprise the command and that the second portion of the audio data that includes the multiple terms of the utterance following the hotword includes speech from the particular user, initiating performance of the command using the user computing device (Column 12, lines 8-12, “Further, the first keyword 222a may include a wakeup word, which must be matched prior to allowing user access to any other commands 202 (i.e., commands 202b-202n) of the command library 219.”; Column 14, lines 32-39, "Furthermore, at step 308, while comparing the utterance, or portion thereof, with the voiceprint that is associated with the pre-determined keyword, a resource is identified. The resource is associated with the pre-determined keyword. Accordingly, the resource may be identified by virtue of being grouped with or linked to the pre-determined keyword. In one or more embodiments, the resource may include data, a function call, or a routine."; Column 14, lines 48-54, "Also, at step 310, in response to authenticating the user based on the comparison, the resource is accessed. In one or more embodiments, accessing the resource may include any execution, retrieval, or invocation operation that is suitable for the resource. For example, if the resource includes data that is stored on a headset or host device, the data may be retrieved."; Identifying a resource associated with a pre-determined keyword by comparing an utterance with a voiceprint, where the resource is associated with the pre-determined keyword and a wakeup word must be matched to a voiceprint prior to allowing user access to any other commands, reads on determining that the multiple terms of the utterance comprise the command and that the second portion of the audio data that includes the multiple terms of the utterance following the hotword includes speech from the particular user, and accessing the resource in response to authenticating the user based on the comparison reads on initiating performance of the command using the user computing device.).
Perotti is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Kim to incorporate the teachings of Perotti to match a wakeup word to a voiceprint of a user before allowing user access to any other commands, performing a keyword matching analysis and a voiceprint comparison analysis for a command, identify a resource associated with a pre-determined keyword by comparing an utterance with a voiceprint, where the resource is associated with the pre-determined keyword and a wakeup word must be matched to a voiceprint prior to allowing user access to any other commands, and access the resource in response to authenticating the user based on the comparison. Doing so would allow for confirming a user's identity in addition to causing the performance of the specific functionality that the user has requested (Perotti; Column 3, line 57 - Column 4, line 5).
Regarding claim 2, Kim in view of Perotti discloses the computer-implemented method as claimed in claim 1.
Kim further discloses:
wherein the audio data is captured by a microphone residing on the user computing device (Column 6, lines 50-53, "At block 302 of process 300, an audio input including user speech can be received at a user device. In some examples, a user device (e.g., user device 102) can receive the audio input including user speech via a microphone (e.g., microphone 230).").
Regarding claim 3, Kim in view of Perotti discloses the computer-implemented method as claimed in claim 1.
Kim further discloses:
wherein the speaker identification model is trained to recognize speech spoken by the particular person (Column 7, lines 21-30, "At block 305, the user device can generate a speaker profile, selectively perform speaker recognition using the speaker profile, and selectively activate the virtual assistant in response to positively identifying the speaker using speaker recognition. In some examples, the speaker profile can generally include one or more voice prints generated from an audio recording of a speaker's voice. The voice prints can be generated using any desired speech recognition technique, such as by generating i-vectors to represent speaker utterances."; Column 7, lines 37-48, "Specifically, at block 306, the user device can select one of multiple modes in which to operate. In some examples, the multiple modes can include a speaker profile building mode (represented by block 308) in which a speaker's voice can be modeled to generate a speaker profile, a speaker profile modifying mode (represented by block 310) in which a speaker profile can be used to verify the identity of a user and in which the speaker profile can be updated based on newly received user speech, and a static speaker profile mode in which an existing speaker profile can be used to verify the identity of a user and in which the speaker profile may not be changed based on newly received user speech."; Modeling a speaker's voice to generate a speaker profile to selectively perform speaker recognition reads on training the speaker identification model to recognize speech spoken by the particular person.).
Regarding claim 4, Kim in view of Perotti discloses the computer-implemented method as claimed in claim 3.
Kim further discloses:
wherein the speaker identification model is trained on previously collected speech data for the particular person (Column 7, lines 21-30, "At block 305, the user device can generate a speaker profile, selectively perform speaker recognition using the speaker profile, and selectively activate the virtual assistant in response to positively identifying the speaker using speaker recognition. In some examples, the speaker profile can generally include one or more voice prints generated from an audio recording of a speaker's voice. The voice prints can be generated using any desired speech recognition technique, such as by generating i-vectors to represent speaker utterances."; Generating voice prints from an audio recording of a speaker's voice reads on the speaker identification model being trained on previously collected speech data for the particular person.).
Regarding claim 5, Kim in view of Perotti discloses the computer-implemented method as claimed in claim 4.
Kim further discloses:
wherein the previously collected speech data characterizes utterances of various phrases the particular person is requested to repeat (Column 1, lines 33-38, "Some natural language processing systems can perform speaker identification to verify the identity of a user. These systems typically require the user to perform an enrollment process during which the user speaks a series of predetermined words or phrases to allow the natural language processing system to model the user's voice."; Performing an enrollment process during which the user speaks a series of predetermined words or phrases reads on the previously collected speech data characterizing utterances of various phrases the particular person is requested to repeat.).
Regarding claim 9, Kim in view of Perotti discloses the computer-implemented method as claimed in claim 1.
Kim further discloses:
wherein the data processing hardware resides on the user computing device (Column 3, lines 62-66, "Although the functionality of the virtual assistant is shown in FIG. 1 as including both a client-side portion and a server-side portion, in some examples, the functions of the assistant can be implemented as a standalone application installed on a user device."; Column 4, lines 7-10, "FIG. 2 is a block diagram of a user-device 102 according to various examples. As shown, user device 102 can include a memory interface 202, one or more processors 204, and a peripherals interface 206.").
Regarding claim 10, Kim in view of Perotti discloses the computer-implemented method as claimed in claim 1.
Kim further discloses:
wherein the user computing device comprises a smart phone, a laptop computer, a desktop computer, a smart speaker, or a smart watch (Column 3, lines 22-33, "As shown in FIG. 1, in some examples, a virtual assistant can be implemented according to a client-server model. The virtual assistant can include a client-side portion executed on a user device 102, and a server-side portion executed on a server system 110. User device 102 can include any electronic device, such as a mobile phone, tablet computer, portable media player, desktop computer, laptop computer, PDA, television, television set-top box, wearable electronic device, or the like, and can communicate with server system 110 through one or more networks 108, which can include the Internet, an intranet, or any other wired or wireless public or private network.").
Regarding claim 11, arguments analogous to claim 1 are applicable. In addition, Kim discloses a system comprising: data processing hardware; and memory hardware in communication with the data processing hardware and storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations (Column 4, line 61 - Column 5, line 3, "In some examples, a non-transitory computer-readable storage medium of memory 250 can be used to store instructions (e.g., for performing some or all of process 300, 400, 500, 600, or 700, described below) for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, processor-containing system, or other system that can fetch the instructions from the instruction execution system, apparatus, or device, and execute the instructions.") comprising the steps of claim 1.
Regarding claim 12, arguments analogous to claim 2 are applicable.
Regarding claim 13, arguments analogous to claim 3 are applicable.
Regarding claim 14, arguments analogous to claim 4 are applicable.
Regarding claim 15, arguments analogous to claim 5 are applicable.
Regarding claim 19, arguments analogous to claim 9 are applicable.
Regarding claim 20, arguments analogous to claim 10 are applicable.
Claims 6 – 8 and 16 – 18 are rejected under 35 U.S.C. 103 as being unpatentable over Kim in view of Perotti, and further in view of Foerster et al. (US Patent No. 9,418,656), hereinafter Foerster.
Regarding claim 6, Kim in view of Perotti discloses the computer-implemented method as claimed in claim 1, but does not specifically disclose: wherein processing the audio data to identify the presence of the hotword in the audio data comprises: processing, using the hotword model, the audio data to compute a hotword confidence score reflecting a likelihood that the audio data includes the hotword; and identifying, using the hotword model, the presence of the hotword based on the hotword confidence score.
Foerster teaches:
wherein processing the audio data to identify the presence of the hotword in the audio data comprises: processing, using the hotword model, the audio data to compute a hotword confidence score reflecting a likelihood that the audio data includes the hotword (Column 4, lines 25-31, "The audio subsystem of the computing device 115 provides the processed audio data 120 to a first stage of the hotworder. The first stage hotworder 125 may be a “coarse” hotworder. The first stage hotworder 125 performs a classification process that may be informed or trained using known utterances of the hotword, and computes a likelihood that the utterance 110 includes a hotword."; Column 5, lines 15-17, "Based on the classification process performed by the first stage hotworder 125, the first stage hotworder 125 computes a hotword confidence score.");
and identifying, using the hotword model, the presence of the hotword based on the hotword confidence score (Column 6, lines 51-60, "In some implementations, the speaker identification module 150 transmits a signal to the second stage hotworder 145 or to the first stage hotworder 125 indicating that the speaker identity confidence score satisfies a threshold and to cease storing or forwarding of the audio data 120. For example, the second stage hotworder 145 determines that the utterance 110 likely includes the hotword “OK computer,” and the second stage hotworder 145 transmits a signal to the first stage hotworder 125 instructing the first stage hotworder 125 to cease storing the audio data 125 into memory 140."; Indicating that the speaker identity confidence score satisfies a threshold reads on identifying the presence of the hotword based on the hotword confidence score.).
Foerster is considered to be analogous to the claimed invention because it is in the same field of speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Kim in view of Perotti to incorporate the teachings of Foerster to compute a hotword confidence score indicating the likelihood that an utterance includes a hotword and indicate that the speaker identity confidence score satisfies a threshold. Doing so would allow for discerning when an utterance is directed at the system as opposed to being directed at an individual present in the environment (Foerster; Column 1, lines 44-67).
Regarding claim 7, Kim in view of Perotti, and further in view of Foerster, discloses the computer-implemented method as claimed in claim 6.
Foerster further teaches:
wherein identifying the presence of the hotword based on the hotword confidence score comprises identifying the presence of the hotword based on determining that the hotword confidence score satisfies a hotword confidence score threshold (Column 6, lines 51-60, "In some implementations, the speaker identification module 150 transmits a signal to the second stage hotworder 145 or to the first stage hotworder 125 indicating that the speaker identity confidence score satisfies a threshold and to cease storing or forwarding of the audio data 120. For example, the second stage hotworder 145 determines that the utterance 110 likely includes the hotword “OK computer,” and the second stage hotworder 145 transmits a signal to the first stage hotworder 125 instructing the first stage hotworder 125 to cease storing the audio data 125 into memory 140."; Indicating that the speaker identity confidence score satisfies a threshold reads on identifying the presence of the hotword based on determining that the hotword confidence score satisfies a hotword confidence score threshold.).
Foerster is considered to be analogous to the claimed invention because it is in the same field of speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Kim in view of Perotti and further in view of Foerster to further incorporate the teachings of Foerster to indicate that the speaker identity confidence score satisfies a threshold. Doing so would allow for discerning when an utterance is directed at the system as opposed to being directed at an individual present in the environment (Foerster; Column 1, lines 44-67).
Regarding claim 8, Kim in view of Perotti, and further in view of Foerster, discloses the computer-implemented method as claimed in claim 6.
Foerster further teaches:
wherein the hotword confidence score is computed without the hotword model performing speech recognition on the audio data (Column 4, lines 25-31, "The audio subsystem of the computing device 115 provides the processed audio data 120 to a first stage of the hotworder. The first stage hotworder 125 may be a “coarse” hotworder. The first stage hotworder 125 performs a classification process that may be informed or trained using known utterances of the hotword, and computes a likelihood that the utterance 110 includes a hotword."; Computing the likelihood that an utterance includes a hotword by performing a classification process trained using known utterances of the hotword reads on computing the hotword confidence score without the hotword model performing speech recognition on the audio data.).
Foerster is considered to be analogous to the claimed invention because it is in the same field of speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Kim in view of Perotti and further in view of Foerster to further incorporate the teachings of Foerster to compute the likelihood that an utterance includes a hotword by performing a classification process trained using known utterances of the hotword. Doing so would allow for discerning when an utterance is directed at the system as opposed to being directed at an individual present in the environment (Foerster; Column 1, lines 44-67).
Regarding claim 16, arguments analogous to claim 6 are applicable.
Regarding claim 17, arguments analogous to claim 7 are applicable.
Regarding claim 18, arguments analogous to claim 8 are applicable.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to James Boggs whose telephone number is (571)272-2968. The examiner can normally be reached M-F 8:00 AM - 5:00 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at (571)272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JAMES BOGGS/Examiner, Art Unit 2657