DETAILED ACTION
Introduction
This office action is in response to Applicant’s Amended submission filed on 6/15/2026. Applicant has amended claims 1-2, 6, and 11-12. Claims 1-20 are pending and have been examined.
Response to Amendment and Arguments
35 U.S.C. 102/103 Rejections
Applicant’s amendments and arguments are considered but are either unpersuasive or moot in view of the new grounds of rejection that, if presented, were necessitated by the amendments to the Claims.
Applicant’s arguments are directed to material that is added by the most recent amendments to the Claims. Response, p. 7.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-2, 4, 6-12, 14, and 16-20 are rejected under 35 U.S.C. 103 as being unpatentable over Sharifi, in view of Oktem (US 20180308491).
Regarding Claim 1, Sharifi discloses: 1. A computer-implemented method that when executed on data processing hardware causes the data processing hardware to perform operations comprising: ([0003] One aspect of the disclosure provides a method for activating speaker-dependent warm words.)
receiving audio data corresponding to an utterance spoken by a particular user and captured in streaming audio by a user device; ([0003] The method includes receiving, at data processing hardware, audio data corresponding to an utterance spoken by a user and captured by an assistant-enabled device associated with the user. The utterance includes a command for a digital assistant to perform a long-standing operation.)
performing speaker identification on the audio data to identify an identity of the particular user that spoke the utterance; ([0006] the method also includes performing, by the data processing hardware, speaker identification on the audio data to identify the user that spoke the utterance. The speaker identification includes extracting, from the audio data corresponding to the utterance spoken by the user, a first speaker-discriminative vector representing characteristics of the utterance spoken by the user, and determining whether the extracted speaker-discriminative vector matches any enrolled speaker vectors stored on the assistant-enabled device. Each enrolled speaker vector is associated with a different respective enrolled user of the assistant-enabled device.)
Although Sharifi teaches/suggest in para 0029 that associating warm words with voice of particular user so that warm words are speaker dependent, accuracy in triggering respective action upon detecting the warm words is improved since only the particular user is permitted to speak the warm words, it is not clear that the model itself is actually conditioned or trained using speaker characteristics.
Oktem (previous mentioned in as art pertaining to the application) discloses: based on the identity of the particular user that spoke the utterance, obtaining a keyword detection model personalized for the particular user by conditioning a speaker- agnostic keyword detection model on speaker characteristic information associated with the particular user to adapt the keyword detection model to detect a presence of a keyword in audio for the particular user, the speaker-agnostic keyword detection model trained to detect the presence of the keyword in audio; ([0093] While the use of speaker identification features is generally described here, audio recordings may similarly be used. For example, the speech-enabled device 125 may store four audio recordings corresponding to the known user “Dad” saying “OK Computer” and then use the four audio recordings to generate a hotword detection model that can be later used to detect the known user “Dad” speaking the hotword. The hotword detection model may even be generated based on speaker identification features extracted from the four audio recordings. Accordingly, the description herein of the system 100 storing and using speaker identification features to detect a known user speaking a hotword may similarly apply to storing and using audio recordings to detect a known user speaking a hotword, and vice versa.)
and determining, using the keyword detection model personalized for the particular user, that the utterance comprises the keyword. ([0093] generate a hotword detection model that can be later used to detect the known user “Dad” speaking the hotword. The hotword detection model may even be generated based on speaker identification features extracted from the four audio recordings. Accordingly, the description herein of the system 100 storing and using speaker identification features to detect a known user speaking a hotword may similarly apply to storing and using audio recordings to detect a known user speaking a hotword, and vice versa.)
Sharifi and Oktem are considered analogous art. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Sharifi to combine the teaching of Oktem for the above mentioned feature, because the technique described enable personalization and security for the user (Oktem, [0005]).
Regarding Claim 2, Sharifi and Oktem disclose all of claim 1,
Oktem further discloses: wherein the keyword detection model is personalized for the particular user by: obtaining the speaker-agnostic keyword detection model, the speaker-agnostic keyword detection model trained to detect the presence of the keyword in audio; ([0093] While the use of speaker identification features is generally described here, audio recordings may similarly be used. For example, the speech-enabled device 125 may store four audio recordings corresponding to the known user “Dad” saying “OK Computer” and then use the four audio recordings to generate a hotword detection model that can be later used to detect the known user “Dad” speaking the hotword. The hotword detection model may even be generated based on speaker identification features extracted from the four audio recordings. Accordingly, the description herein of the system 100 storing and using speaker identification features to detect a known user speaking a hotword may similarly apply to storing and using audio recordings to detect a known user speaking a hotword, and vice versa.)
Sharifi further discloses: receiving one or more enrollment utterances spoken by the particular user; ([0040] In some examples, the enrolled speaker vector 154 for an enrolled user 200 includes a text-dependent enrolled speaker vector. For instance, the text-dependent enrolled speaker vector may be extracted from one or more audio samples of the respective enrolled user 200 speaking a predetermined term such as the hotword 110 (e.g., “Ok computer”) used for invoking the AED 104 to wake-up from a sleep state.)
extracting, from the one or more enrollment utterances spoken by the particular user, a speaker embedding characterizing voice characteristics of the particular user; ([0040] In some examples, the enrolled speaker vector 154 for an enrolled user 200 includes a text-dependent enrolled speaker vector. For instance, the text-dependent enrolled speaker vector may be extracted from one or more audio samples of the respective enrolled user 200 speaking a predetermined term such as the hotword 110 (e.g., “Ok computer”) used for invoking the AED 104 to wake-up from a sleep state. In other examples, the enrolled speaker vector 154 for an enrolled user 200 is text-independent obtained from one or more audio samples of the respective enrolled user 200 speaking phrases with different terms/words and of different lengths. In these examples, the text-independent enrolled speaker vector may be obtained over time from audio samples obtained from speech interactions the user 102 has with the AED 104 or other device linked to the same account.)
and conditioning the speaker-agnostic keyword detection model on the speaker embedding, the speaker embedding comprising the speaker characteristic information associated with the particular user. ([0040] In some examples, the enrolled speaker vector 154 for an enrolled user 200 includes a text-dependent enrolled speaker vector. For instance, the text-dependent enrolled speaker vector may be extracted from one or more audio samples of the respective enrolled user 200 speaking a predetermined term such as the hotword 110 (e.g., “Ok computer”) used for invoking the AED 104 to wake-up from a sleep state. In other examples, the enrolled speaker vector 154 for an enrolled user 200 is text-independent obtained from one or more audio samples of the respective enrolled user 200 speaking phrases with different terms/words and of different lengths. In these examples, the text-independent enrolled speaker vector may be obtained over time from audio samples obtained from speech interactions the user 102 has with the AED 104 or other device linked to the same account.)
Where the rationale for the combination would be similar the one already provided.
Regarding Claim 4, Sharifi and Oktem disclose all of claim 2,
Sharifi further discloses: wherein extracting the speaker embedding comprises extracting a text-independent speaker embedding or a text-dependent speaker embedding. ([0040] In some examples, the enrolled speaker vector 154 for an enrolled user 200 includes a text-dependent enrolled speaker vector. For instance, the text-dependent enrolled speaker vector may be extracted from one or more audio samples of the respective enrolled user 200 speaking a predetermined term such as the hotword 110 (e.g., “Ok computer”) used for invoking the AED 104 to wake-up from a sleep state. In other examples, the enrolled speaker vector 154 for an enrolled user 200 is text-independent obtained from one or more audio samples of the respective enrolled user 200 speaking phrases with different terms/words and of different lengths. In these examples, the text-independent enrolled speaker vector may be obtained over time from audio samples obtained from speech interactions the user 102 has with the AED 104 or other device linked to the same account.)
Regarding Claim 6, Sharifi and Oktem disclose all of claim 1,
Sharifi further discloses: wherein obtaining the keyword detection model personalized for the particular user comprises: using the identity of the particular user, retrieving the speaker characteristic information associated with the particular user from memory hardware in communication with the data processing hardware; ([0040] In some examples, the enrolled speaker vector 154 for an enrolled user 200 includes a text-dependent enrolled speaker vector. For instance, the text-dependent enrolled speaker vector may be extracted from one or more audio samples of the respective enrolled user 200 speaking a predetermined term such as the hotword 110 (e.g., “Ok computer”) used for invoking the AED 104 to wake-up from a sleep state. In other examples, the enrolled speaker vector 154 for an enrolled user 200 is text-independent obtained from one or more audio samples of the respective enrolled user 200 speaking phrases with different terms/words and of different lengths. In these examples, the text-independent enrolled speaker vector may be obtained over time from audio samples obtained from speech interactions the user 102 has with the AED 104 or other device linked to the same account.) [it is implied that the speaker vector 154 is stored to be access by the AED 104] as para 0031-0032 disclose that “The AED 104 includes data processing hardware 10 and memory hardware 12 storing instructions that when executed on the data processing hardware 10 cause the data processing hardware 10 to perform operations.” Also see para 0036, “the warm word selector 126 accesses a registry or table (e.g., stored on the memory hardware 12)”
Oktem further discloses: and conditioning the speaker-agnostic keyword detection model on the speaker characteristic information associated with the particular user to provide the keyword detection model personalized for the particular user, the speaker-agnostic keyword detection model trained to detect the presence of the keyword in audio. ([0093] While the use of speaker identification features is generally described here, audio recordings may similarly be used. For example, the speech-enabled device 125 may store four audio recordings corresponding to the known user “Dad” saying “OK Computer” and then use the four audio recordings to generate a hotword detection model that can be later used to detect the known user “Dad” speaking the hotword. The hotword detection model may even be generated based on speaker identification features extracted from the four audio recordings. Accordingly, the description herein of the system 100 storing and using speaker identification features to detect a known user speaking a hotword may similarly apply to storing and using audio recordings to detect a known user speaking a hotword, and vice versa.)
Where the rationale for the combination would be similar to the one already provided.
Regarding Claim 7, Sharifi and Oktem disclose all of claim 6,
Sharifi further discloses: wherein the speaker characteristic information associated with the particular user is stored on the memory hardware ([0040] In some examples, the enrolled speaker vector 154 for an enrolled user 200 includes a text-dependent enrolled speaker vector. For instance, the text-dependent enrolled speaker vector may be extracted from one or more audio samples of the respective enrolled user 200 speaking a predetermined term such as the hotword 110 (e.g., “Ok computer”) used for invoking the AED 104 to wake-up from a sleep state.) [it is implied that the speaker vector 154 is stored to be access by the AED 104] as para 0031-0032 disclose that “The AED 104 includes data processing hardware 10 and memory hardware 12 storing instructions that when executed on the data processing hardware 10 cause the data processing hardware 10 to perform operations.” Also see para 0036, “the warm word selector 126 accesses a registry or table (e.g., stored on the memory hardware 12)”
and comprises one or more enrollment utterances spoken by the particular user. ([0040] In some examples, the enrolled speaker vector 154 for an enrolled user 200 includes a text-dependent enrolled speaker vector. For instance, the text-dependent enrolled speaker vector may be extracted from one or more audio samples of the respective enrolled user 200 speaking a predetermined term such as the hotword 110 (e.g., “Ok computer”) used for invoking the AED 104 to wake-up from a sleep state.)
Regarding Claim 8, Sharifi and Oktem disclose all of claim 6,
Sharifi further discloses: wherein the speaker characteristic information associated with the particular is stored on the memory hardware ([0040] In some examples, the enrolled speaker vector 154 for an enrolled user 200 includes a text-dependent enrolled speaker vector. For instance, the text-dependent enrolled speaker vector may be extracted from one or more audio samples of the respective enrolled user 200 speaking a predetermined term such as the hotword 110 (e.g., “Ok computer”) used for invoking the AED 104 to wake-up from a sleep state.) [the context of enrolled speaker vector implies it is stored in memory] Also see para 0031-0032 and 0036 which disclose hardware memory and storage.
and comprises a speaker embedding extracted from one or more enrollment utterances spoken by the particular user. ([0040] In some examples, the enrolled speaker vector 154 for an enrolled user 200 includes a text-dependent enrolled speaker vector. For instance, the text-dependent enrolled speaker vector may be extracted from one or more audio samples of the respective enrolled user 200 speaking a predetermined term such as the hotword 110 (e.g., “Ok computer”) used for invoking the AED 104 to wake-up from a sleep state.)
Regarding Claim 9, Sharifi and Oktem disclose all of claim 1,
wherein the operations further comprise, in response to determining that the utterance comprises the keyword, processing, using a speech recognition model, the streaming audio. ([0034] When the hotword detector 108 determines that the audio data 402 that corresponds to the utterance 106 includes the hotword 110, the AED 104 may trigger a wake-up process to initiate speech recognition on the audio data 402 that corresponds to the utterance 106. For example, a speech recognizer 116 running on the AED 104 may perform speech recognition or semantic interpretation on the audio data 402 that corresponds to the utterance 106. The speech recognizer 116 may perform speech recognition on the portion of the audio data 402 that follows the hotword 110. In this example, the speech recognizer 116 may identify the words “play music” in the command 118.)
Regarding Claim 10, Sharifi and Oktem disclose all of claim 1,
wherein the utterance comprises a keyword followed by one or more other terms corresponding to a voice command. ([0034] When the hotword detector 108 determines that the audio data 402 that corresponds to the utterance 106 includes the hotword 110, the AED 104 may trigger a wake-up process to initiate speech recognition on the audio data 402 that corresponds to the utterance 106. For example, a speech recognizer 116 running on the AED 104 may perform speech recognition or semantic interpretation on the audio data 402 that corresponds to the utterance 106. The speech recognizer 116 may perform speech recognition on the portion of the audio data 402 that follows the hotword 110. In this example, the speech recognizer 116 may identify the words “play music” in the command 118.)
Regarding Claim 11, Sharifi discloses: 11. A system comprising: data processing hardware; ([0009] Another aspect of the disclosure provides a system for activating speaker-dependent warm words. The system includes data processing hardware and memory hardware in communication with the data processing hardware.)
and memory hardware in communication with the data processing hardware, the memory hardware storing instruction that when executed on the data processing hardware cause the data processing hardware to perform operations comprising: ([0009] The memory hardware stores instructions that when executed by the data processing hardware cause the data processing hardware to perform operations that include receiving audio data corresponding to an utterance spoken by a user and captured by an assistant-enabled device associated with the user.)
As for the rest of the claim, they claim the elements of claim 1, therefor the rationale applied in the rejection of claim 1 is equally applicable.
Claim 12 recites system claims that corresponds to the method of claim 2 is therefore rejected under the same grounds as claim 2 above.
Claim 14 recites system claims that corresponds to the method of claim 4 is therefore rejected under the same grounds as claim 4 above.
Claim 16 recites system claims that corresponds to the method of claim 6 is therefore rejected under the same grounds as claim 6 above.
Claim 17 recites system claims that corresponds to the method of claim 7 is therefore rejected under the same grounds as claim 7 above.
Claim 18 recites system claims that corresponds to the method of claim 8 is therefore rejected under the same grounds as claim 8 above.
Claim 19 recites system claims that corresponds to the method of claim 9 is therefore rejected under the same grounds as claim 9 above.
Claim 20 recites system claims that corresponds to the method of claim 10 is therefore rejected under the same grounds as claim 10 above.
Claims 3 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Sharifi, in view of Oktem, and further in view of Rikhye, R., Wang, Q., Liang, Q., He, Y., & McGraw, I. (2022). Closing the gap between single-user and multi-user voicefilter-lite. arXiv preprint arXiv:2202.12169.
Regarding Claim 3, Shaifi and Oktem disclose all the elements of Claim 2,
However, Sharifi and Oktem do not disclose the Feature-wise linear Modulation.
Rikhye discloses: wherein: the keyword detection model comprises a Feature-wise Linear Modulation (FiLM) layer; ([Abstract] In this paper, we devised a series of experiments to improve the multi-user VoiceFilter-Lite model. By incorporating a dua learning rate schedule and by using feature-wise linear modulation (FiLM) to condition the model with the attended speaker embedding,)
and conditioning the speaker-agnostic keyword detection model comprises: generating, using a FILM generator, FiLM parameters based on the speaker embedding; ([Abstract] By incorporating a dual learning rate schedule and by using feature-wise linear modulation (FiLM) to condition the model with the attended speaker embedding, we successfully closed the performance gap between multi-user and single-user VoiceFilter-Lite models on single-speaker evaluations. At the same time, the new model can also be easily extended to support any number of users, and significantly outperforms our previously published model on multi-speaker evaluations.) [implicit in conditioned modulation]
and modulating the FILM layer of the speaker-agnostic keyword detection model using the FiLM parameters to provide the keyword detection model personalized for the particular user. ([Abstract] By incorporating a dual learning rate schedule and by using feature-wise linear modulation (FiLM) to condition the model with the attended speaker embedding, we successfully closed the performance gap between multi-user and single-user VoiceFilter-Lite models on single-speaker evaluations. At the same time, the new model can also be easily extended to support any number of users, and significantly outperforms our previously published model on multi-speaker evaluations.)
Sharifi, Oktem and Rikhye are considered analogous art. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Sharifi and Oktem to combine the teaching of Rikhye, because FiLM only requires two parameters, making it computationally more efficient conditioned method (Rikhye, [sect 2.3]).
Claim 13 recites system claims that corresponds to the method of claim 3 is therefore rejected under the same grounds as claim 3 above.
Claims 5 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Sharifi, in view of Oktem, and further in view of Higuchil, T., Gupta, A., & Dhir, C. (2021, December). Multi-task learning with cross attention for keyword spotting. In 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) (pp. 571-578). IEEE.
Regarding Claim 5, Sharifi and Oktem disclose all the elements of claim 2,
Oktem further discloses: and conditioning the speaker-agnostic keyword detection model comprises processing the one or more enrollment utterances using the([0093] While the use of speaker identification features is generally described here, audio recordings may similarly be used. For example, the speech-enabled device 125 may store four audio recordings corresponding to the known user “Dad” saying “OK Computer” and then use the four audio recordings to generate a hotword detection model that can be later used to detect the known user “Dad” speaking the hotword. The hotword detection model may even be generated based on speaker identification features extracted from the four audio recordings. Accordingly, the description herein of the system 100 storing and using speaker identification features to detect a known user speaking a hotword may similarly apply to storing and using audio recordings to detect a known user speaking a hotword, and vice versa.)
However, Sharifi and Oktem do not explicitly disclose keyword detection model comprises of a stack of cross-attention layers and conditioning the detection model processing enrollment utterance using the stack of cross-attention layer to provide keyword detection.
Higuchil discloses: wherein: the speaker-agnostic keyword detection model comprises a stack of cross-attention layers; ([sect 3.3] Cross attention decoder – our cross attention decoder is based on Transformer blocks with attention layers. The attention layer for a query matrix Q, a
key matrix K and a value matrix V can be written as …)
Sharifi, Oktem and Higuchil are considered analogous art. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Sharifi and Oktem to combine the teaching of Higuchil, because performing cross attention can improve accuracy of keyword detection (Higuchil, [Abstract]).
Claim 15 recites system claims that corresponds to the method of claim 5 is therefore rejected under the same grounds as claim 5 above.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: Yang, S., Kim, B., Chung, I., & Chang, S. (2022). Personalized keyword spotting through multi-task learning. arXiv preprint arXiv:2206.13708. – discloses method/system for keyword spotting using multi-task learning and task adaptation. See Abstract and fig.2 for details.
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Philip H Lam whose telephone number is (571)272-1721. The examiner can normally be reached 9 AM-3 PM Pacific time.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bhavesh Mehta can be reached on 571-272-7453. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PHILIP H LAM/ Examiner, Art Unit 2656