))DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 3/5/2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim 1-20 are rejected under 35 U.S.C. 101 because:
Claims 1 (method), 10 (“device” (apparatus)), 18 (“non-transitory computer-readable medium”) are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claims recite a “wake-up model” used to “detect[]” a “wake-up command” (e.g., “OK Google” “Alexa” (Sp. ¶ 0003)) “included in a voice input”. Upon detection, a “speech recognition” is performed “on the voice input received from a user”. If a “confidence score” of the said “recognition” is “above a threshold value”, then the “voice input” and a “result of” its “speech recognition” is used as “user specific training data” (e.g., “personal information of the user, such as an accent of the user” “daily environment noise encountered by the user” (Sp. ¶ 0065 S. before last)). Other examples of “user-specific data” include “ e.g., “amount of user-specific data” that “has been collected” such as “battery usage history” “current battery charge” (Sp. ¶ 0065 S3)). This information is used to e.g. do training when it is least intrusive with the user using his electronic device.
Then this “training data” is used to further “train[]” “the wake-up model”. So the “wake-up model” is used to obtain “user specific training data”, which in turn is used to further “train” “the wake-up model”. However, the claim limitations as drafted are silent on any active wakeup or wakeword model and their effect on the “electronic device”. As such the “wake-up command” could correspond to a series of words that a user anticipates from another user and then performs a subsequent action associated with them.
Therefore, other than reciting using “at least one processor” (claim 10), or “by at least one processor” (claim 18), nothing in the claims limitations precludes their steps from practically being performed in the mind. For example, suppose a user is trying to catch a taxi at an airport in DC area. If he utters “white house”, the taxi driver more than likely recognizes the series of words and will function to what amounts as a wake-up command. If for some reason the taxi driver failed to understand due e.g., environmental noise, and/or accent of the user (i.e., his recognition fell below a threshold of comprehensibility), he most likely will request the user to either speak louder (to overcome the noise) and/or perhaps write on a sheet of paper or use his mobile phone the destination due to his accent interfering the taxi driver’s comprehension, until the Taxi driver feels his comprehension of the user’s speech is above a minimum comprehension threshold before he pursues for the trip to the destination (i.e., the white house). Furthermore, for the taxi driver following this experience, he could comprehend the same user should he encounter him again for the same or even a different destination by getting accustomed to the user’s accent (i.e., his mind gets trained following the original experience).
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claims recite an abstract idea.
The judicial exception is not integrated into a practical application. In particular, the claims only recite one additional element i.e., the “processor” to be responsible for the limitations of “detect …”, “Perform” “speech recognition”, “determine a confidence score …”, “obtain user-specific training data ..”, and “perform user-specific training …”. The “processor” is therefore recited at a high level of generality such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claims are thus directed to an abstract idea.
The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of using a processor to perform all the limitations of “detect …”, “Perform” “speech recognition”, “determine a confidence score …”, “obtain user-specific training data ..”, and “perform user-specific training …” amounts to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. The claims are thus not patent eligible.
Regarding claims 2-4 (11-13), and 19, the taxi driver upon hearing “white house” (wake up command), uses his knowledge of vocabulary in the English language to detect” (functioning as a key word detection model) and verify (functioning as a key word verification model) each word before taking any further action.
Regarding claim 5 (14), recognition of “white house” by the taxi driver amounts to both a speech recognition by the taxi driver as well as wake-up model recognition, and a positive response amounts to passing both speech recognition as well as wakeup model recognition above an acceptable level of comprehension which qualifies it as passing both speech recognition as well as wake-up scores receiving a passing score of “1” in a binary system assigning “1” (passing) and “0” (not passing).
Regarding claims 6 (15) and 7 (16), the taxi driver could use information about weather and traffic (additional user-specific data) in order to decide whether or not at a given time to provide service to a passenger (e.g., if there is bad weather in DC to avoid a trip to that destination).
Regarding claim 8, the taxi driver can do all the comprehension (i.e., training) of the phrase uttered by the passenger user without any help from anyone else based entirely on his own knowledge of vocabulary and grammar in English.
Regarding claims 9 (17), and 20, the taxi driver may hear in addition to “white house” (a first wake-up command) other words (wake-up words) such as “the Mall” (i.e., next to the white house), and will recognize it the same way, and if somehow he could not hear “the Mall” he could request the user to speak that phrase again (obtain new user specific training data) and perform better understanding (perform additional user-specific training) to fully grasp it to for instance decode an input like “the Mall next to the white house”.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claim(s) 1-5, 8-14, 17-20 is/are rejected under 35 U.S.C. 102(a)(2) as being anticipated by CHEN et al. (US 2024/0177707).
Regarding claim 1, CHEN et al. do teach a method of training a user-specific wake-up model, the method being performed by an electronic device (Abstract; and ¶ 0054 S2: “the voice apparatus” (an electronic device) “can have an initial wake-up model and an initial confidence level threshold when first powered on, and be trained” (training a wake-up model) “and updated during subsequent use to make it compatible with usage scenarios” (based on a user specific training data) “of the user”),
and comprising:
detecting, using a wake-up model, a wake-up command included in a voice input received from a user (¶ 0043 S2: “the user may say the wake-up word first, and then say the rest of a voice instruction” (a voice input including a wake-up command received) e.g., ¶ 0044 last S: “Xiaomei Xiaomei, turn up the temperature” (receiving from a user to a wake-up model, including “Xiaomei Xiaomei” (a wake up word) as part of an “instruction” (a command) included in a “user” input);
based on the detecting of the wake-up command, performing a speech recognition operation based on the voice input; determining a confidence score based on a result of the speech recognition operation (¶ 0047 S1: “At S102, the to-be-recognized audio” (e.g., the detected wake-up command) “is processed” (is speech recognized) “to obtain at least two confidence levels” (e.g., as disclosed in ¶ 0060 last S a “first confidence level” (to determine a confidence score));
based on the confidence score being above a threshold value, obtaining user- specific training data based on the voice input and a result of the speech recognition operation (¶ 0060 last S: “triggering the wake-up event of the voice apparatus when the first confidence level is greater than or equal to the first confidence level threshold” (determining the confidence score being above a threshold based on the voice input and a result of the speech recognition for the “input” including the “wake-up word” (i.e., a user specific training data is obtained associated with the recognition of the wake-up word)); also according to ¶ 0046 last 7 lines: “the wake-up model is trained with noise” (based on a user-specific training data) “a corresponding confidence level threshold can be adjusted” (the said threshold (another user-specific training data) was obtained)) ; and
performing user-specific training on the wake-up model based on the user- specific training data to obtain a user-specific wake-up model that is trained to respond to the user (to respond to the user “input” above, according to ¶ 0064 S1: “a confidence level corresponding to “Xiaomei Xiaomei” and a confidence level corresponding to “Assistant Xiaomei” are obtained by the wake-up model based on training data corresponding to the wake-up word “Xiaomei Xiaomei” and training data” (a user specific training based on data) “corresponding to the wake-up word” (associated with recognized user specific recognition of the wake-up word) “Assistant Xiaomei”; ¶ 0046 S2: “subsequent to an extraction of the environmental sound information from the to-be-recognized audio, the environmental sound information” (another user specific) “may be transmitted to a server as training data”( training data on the wake-up model); ¶ 0046 last 7 lines: “the wake-up model is trained” (performing training of the wake-up model) “with noise” (based on a user-specific training data) “a corresponding confidence level threshold” (and also another user-specific training data) “can be adjusted”, because according to ¶ 0009 S1: “training data includes a model parameter and a confidence level threshold”)).
Regarding claim 2, CHEN et al. do teach the method of claim 1, wherein the wake-up model comprises a key word detector (KWD) model trained to detect a wake-up command, and a key word verifier (KWV) model trained to verify the wake-up command (¶ 0053 lines 6+: “Then, the audio text information is processed for matching by means of text matching” (detection) “or semantic matching” (and verification) “to determine at least one keyword” (of a key word) “or key character. The at least one keyword or key character is then processed separately using the wake-up model” (in the wake-up model for e.g., the “wake-up word” “Xiaomei Xiaomei” (a keyword which is a wakeup command)) “and the at least two groups of training data”).
Regarding claim 3, CHEN et al. do teach the method of claim 2, wherein the performing of the user-specific training comprises training the KWV model using the user-specific training data to obtain a user-specific KWV model (¶ 0064 S1: “a confidence level corresponding to “Xiaomei Xiaomei” and a confidence level corresponding to “Assistant Xiaomei” are obtained by the wake-up model” (the wakeup model functioning also as key word verifier as it recognizes the key word “Xiaomei”) “based on training data” (which is also the user specific training data) “corresponding to the wake-up word “Xiaomei Xiaomei””).
Regarding claim 4, CHEN et al. do teach the method of claim 3, wherein the user-specific wake-up model comprises the KWD model and the user-specific KWV model ((¶ 0053 lines 6+: “Then, the audio text information is processed for matching by means of text matching” (key word detection) “or semantic matching” (and key word verification) “to determine at least one keyword” “or key character. The at least one keyword or key character is then processed separately using the wake-up model” (as part of the wake-up model) “and the at least two groups of training data”).
Regarding claim 5, CHEN et al. do teach the method of claim 1, wherein the determining of the confidence score comprises:
determining a wake-up score based on an output of the wake-up model (¶ 0010 lines 12-15: “using the wake-up model” (using a wake-up model) “and the first model parameter in the first group of training data to obtain a first confidence level” (to determine a wake-up score));
determining a speech recognition score based on the output of the speech recognition model (¶ 0010 lines 12-15: “using the wake-up model” (using a “speech recognition module” (speech recognition model (¶ 0081 lines 11+)) “and the first model parameter in the first group of training data to obtain a first confidence level” (to determine a speech recognition score)); ¶ 0081 lines 11+: “The wake-up model may have a built-in speech recognition module, through which the to-be-recognized audio may be recognized to output the wake-up event corresponding to the to-be-recognized audio”) ; and
determining the confidence score based on the wake-up score and the speech recognition score (the “first confidence level” (the confidence score) is both the wake-up score as well as the speech recognition score).
Regarding claim 8, CHEN et al. do teach the method of claim 1, wherein the user-specific training is performed by the electronic device (¶ 0054 S2: “the voice apparatus” (the electronic device) “can have an initial wake-up model and an initial confidence level threshold when first powered on, and be trained” (performs training the wake-up model) “and updated during subsequent use to make it compatible with usage scenarios” (based on the user specific training data) “of the user”).
Regarding claim 9, CHEN et al. do teach the method of claim 1, further comprising: detecting, using the user-specific wake-up model, a new wake-up command included in a new voice input received from the user (¶ 0057 S1: “hypothetically, wake-up word A and wake-up word B” (a new wake-up command included in a new voice input) “are available”);
based on the detecting of the new wake-up command, performing a new speech recognition operation based on the voice input (¶ 0057 S2+: “Subsequent to obtaining training data of wake-up word A and training data of wake-up word B, the to-be-recognized audio may be processed using the wake-up model and the training data of wake-up word A to obtain a confidence level of wake-up word A and a confidence level threshold corresponding to wake-up word A and processed using the wake-up model and the training data of wake-up word B to obtain a confidence level” (performing new speech recognition based on the “wake-up word B” (new wake-up command) “of wake-up word B and a confidence level threshold corresponding to wake-up word B”); and
based on a result of the new speech recognition operation indicating that a performance of the user-specific wake-up model is below a threshold performance, obtaining new user-specific training data based on the new voice input and the result of the new speech recognition operation, and performing additional user-specific training on the user-specific wake-up model based on the new user-specific training data ( ¶ 0063: “It should be noted that for a case where both the first confidence level and the second confidence level are greater than or equal to respective confidence level thresholds corresponding to the first confidence level and the second confidence level, the first value by which the first confidence level exceeds the first confidence level threshold and the second value by which the second confidence level exceeds the second confidence level threshold need to be calculated” (obtaining new user-specific training data based on the new voice input) “In some embodiments, that “triggering a target wake-up event of the voice apparatus based on the first value and the second value” may include: when the first value is greater than or equal to the second value” (when the new speech recognition is below a threshold performance compared to the first recognition results) “determining the target wake-up event as the first wake-up event” “and triggering” (and performing additional training) “the first wake-up event; or when the first value is smaller than the second value, determining the target wake-up event as the second wake-up event and triggering the second wake-up event.
Regarding claim 10, CHEN et al. do teach an electronic device for training a user-specific wake-up model (Abstract; and ¶ 0054 S2: “the voice apparatus” (an electronic device) “can have an initial wake-up model and an initial confidence level threshold when first powered on, and be trained” (training a wake-up model) “and updated during subsequent use to make it compatible with usage scenarios” (based on a user specific training data) “of the user”),
The electronic device comprising:
At least one memory configured to store instructions; and at least one processor configured to execute the instructions (¶ 0018: “In a third aspect, an embodiment of the present disclosure provides a voice apparatus. The voice apparatus includes a memory and one or more processors. The memory is configured to store a computer program or computer instructions executable by the processor. The one or more processors is configured to perform, when executing the computer program, the method according to any embodiment in the first aspect”)
To:
detect, using a wake-up model, a wake-up command included in a voice input received from a user (¶ 0043 S2: “the user may say the wake-up word first, and then say the rest of a voice instruction” (a voice input including a wake-up command received) e.g., ¶ 0044 last S: “Xiaomei Xiaomei, turn up the temperature” (receiving from a user to a wake-up model, including “Xiaomei Xiaomei” (a wake up word) as part of an “instruction” (a command) included in a “user” input);
based on the detecting of the wake-up command, perform a speech recognition operation based on the voice input; determine a confidence score based on a result of the speech recognition operation (¶ 0047 S1: “At S102, the to-be-recognized audio” (e.g., the detected wake-up command) “is processed” (is speech recognized) “to obtain at least two confidence levels” (e.g., as disclosed in ¶ 0060 last S a “first confidence level” (to determine a confidence score));
based on the confidence score being above a threshold value, obtain user- specific training data based on the voice input and a result of the speech recognition operation (¶ 0060 last S: “triggering the wake-up event of the voice apparatus when the first confidence level is greater than or equal to the first confidence level threshold” (determining the confidence score being above a threshold based on the voice input and a result of the speech recognition for the “input” including the “wake-up word” (i.e., a user specific training data is obtained associated with the recognition of the wake-up word)); also according to ¶ 0046 last 7 lines: “the wake-up model is trained with noise” (based on a user-specific training data) “a corresponding confidence level threshold can be adjusted” (the said threshold (another user-specific training data) was obtained)) ; and
perform user-specific training on the wake-up model based on the user- specific training data to obtain a user-specific wake-up model that is trained to respond to the user (to respond to the user “input” above, according to ¶ 0064 S1: “a confidence level corresponding to “Xiaomei Xiaomei” and a confidence level corresponding to “Assistant Xiaomei” are obtained by the wake-up model based on training data corresponding to the wake-up word “Xiaomei Xiaomei” and training data” (a user specific training based on data) “corresponding to the wake-up word” (associated with recognized user specific recognition of the wake-up word) “Assistant Xiaomei”; ¶ 0046 S2: “subsequent to an extraction of the environmental sound information from the to-be-recognized audio, the environmental sound information” (another user specific) “may be transmitted to a server as training data”( training data on the wake-up model); ¶ 0046 last 7 lines: “the wake-up model is trained” (performing training of the wake-up model) “with noise” (based on a user-specific training data) “a corresponding confidence level threshold” (and also another user-specific training data) “can be adjusted”, because according to ¶ 0009 S1: “training data includes a model parameter and a confidence level threshold”)).
Regarding claim 11, CHEN et al. do teach the electronic device of claim 10, wherein the wake-up model comprises a key word detector (KWD) model trained to detect a wake-up command, and a key word verifier (KWV) model trained to verify the wake-up command (¶ 0053 lines 6+: “Then, the audio text information is processed for matching by means of text matching” (detection) “or semantic matching” (and verification) “to determine at least one keyword” (of a key word) “or key character. The at least one keyword or key character is then processed separately using the wake-up model” (in the wake-up model for e.g., the “wake-up word” “Xiaomei Xiaomei” (a keyword which is a wakeup command)) “and the at least two groups of training data”).
Regarding claim 12, CHEN et al. do teach the electronic device of claim 11, wherein to perform the user-specific training, the at least one processor is further configured to execute the instructions to:
Train the KWV model using the user-specific training data to obtain a user-specific KWV model (¶ 0064 S1: “a confidence level corresponding to “Xiaomei Xiaomei” and a confidence level corresponding to “Assistant Xiaomei” are obtained by the wake-up model” (the wakeup model functioning also as key word verifier as it recognizes the key word “Xiaomei”) “based on training data” (which is also the user specific training data) “corresponding to the wake-up word “Xiaomei Xiaomei””).
Regarding claim 13, CHEN et al. do teach the electronic device of claim 12, wherein the user-specific wake-up model comprises the KWD model and the user-specific KWV model (¶ 0053 lines 6+: “Then, the audio text information is processed for matching by means of text matching” (key word detection) “or semantic matching” (and key word verification) “to determine at least one keyword” “or key character. The at least one keyword or key character is then processed separately using the wake-up model” (as part of the wake-up model) “and the at least two groups of training data”).
Regarding claim 14, CHEN et al. do teach the electronic device of claim 10, wherein the at least one processor is further configured to execute the instruction to:
determine a wake-up score based on an output of the wake-up model (¶ 0010 lines 12-15: “using the wake-up model” (using a wake-up model) “and the first model parameter in the first group of training data to obtain a first confidence level” (to determine a wake-up score));
determine a speech recognition score based on the output of the speech recognition model (¶ 0010 lines 12-15: “using the wake-up model” (using a “speech recognition module” (speech recognition model (¶ 0081 lines 11+)) “and the first model parameter in the first group of training data to obtain a first confidence level” (to determine a speech recognition score)); ¶ 0081 lines 11+: “The wake-up model may have a built-in speech recognition module, through which the to-be-recognized audio may be recognized to output the wake-up event corresponding to the to-be-recognized audio”) ; and
determine the confidence score based on the wake-up score and the speech recognition score (the “first confidence level” (the confidence score) is both the wake-up score as well as the speech recognition score).
Regarding claim 17, CHEN et al. do teach the electronic device of claim 10, wherein the at least one processor is further configured to execute the instructions to:
detect, using the user-specific wake-up model, a new wake-up command included in a new voice input received from the user (¶ 0057 S1: “hypothetically, wake-up word A and wake-up word B” (a new wake-up command included in a new voice input) “are available”);
based on the detecting of the new wake-up command, performing a new speech recognition operation based on the voice input (¶ 0057 S2+: “Subsequent to obtaining training data of wake-up word A and training data of wake-up word B, the to-be-recognized audio may be processed using the wake-up model and the training data of wake-up word A to obtain a confidence level of wake-up word A and a confidence level threshold corresponding to wake-up word A and processed using the wake-up model and the training data of wake-up word B to obtain a confidence level” (performing new speech recognition based on the “wake-up word B” (new wake-up command) “of wake-up word B and a confidence level threshold corresponding to wake-up word B”); and
based on a result of the new speech recognition operation indicating that a performance of the user-specific wake-up model is below a threshold performance, obtaining new user-specific training data based on the new voice input and the result of the new speech recognition operation, and performing additional user-specific training on the user-specific wake-up model based on the new user-specific training data ( ¶ 0063: “It should be noted that for a case where both the first confidence level and the second confidence level are greater than or equal to respective confidence level thresholds corresponding to the first confidence level and the second confidence level, the first value by which the first confidence level exceeds the first confidence level threshold and the second value by which the second confidence level exceeds the second confidence level threshold need to be calculated” (obtaining new user-specific training data based on the new voice input) “In some embodiments, that “triggering a target wake-up event of the voice apparatus based on the first value and the second value” may include: when the first value is greater than or equal to the second value” (when the new speech recognition is below a threshold performance compared to the first recognition results) “determining the target wake-up event as the first wake-up event” “and triggering” (and performing additional training) “the first wake-up event; or when the first value is smaller than the second value, determining the target wake-up event as the second wake-up event and triggering the second wake-up event.
Regarding claim 18, CHEN et al. do teach a non-transitory computer-readable medium storing instructions which, when executed by at least one processor of a device (page 12 2nd column last paragraph: “A computer-readable storage medium, having one or more computer programs stored thereon, wherein the one or more computer programs, when executed by at least one processor”)
for training a user-specific wake-up model (Abstract; and ¶ 0054 S2: “the voice apparatus” (an electronic device) “can have an initial wake-up model and an initial confidence level threshold when first powered on, and be trained” (training a wake-up model) “and updated during subsequent use to make it compatible with usage scenarios” (based on a user specific training data) “of the user”),
cause the device to:
detect, using a wake-up model, a wake-up command included in a voice input received from a user (¶ 0043 S2: “the user may say the wake-up word first, and then say the rest of a voice instruction” (a voice input including a wake-up command received) e.g., ¶ 0044 last S: “Xiaomei Xiaomei, turn up the temperature” (receiving from a user to a wake-up model, including “Xiaomei Xiaomei” (a wake up word) as part of an “instruction” (a command) included in a “user” input);
based on the detecting of the wake-up command, perform a speech recognition operation based on the voice input; determine a confidence score based on a result of the speech recognition operation (¶ 0047 S1: “At S102, the to-be-recognized audio” (e.g., the detected wake-up command) “is processed” (is speech recognized) “to obtain at least two confidence levels” (e.g., as disclosed in ¶ 0060 last S a “first confidence level” (to determine a confidence score));
based on the confidence score being above a threshold value, obtain user- specific training data based on the voice input and a result of the speech recognition operation (¶ 0060 last S: “triggering the wake-up event of the voice apparatus when the first confidence level is greater than or equal to the first confidence level threshold” (determining the confidence score being above a threshold based on the voice input and a result of the speech recognition for the “input” including the “wake-up word” (i.e., a user specific training data is obtained associated with the recognition of the wake-up word)); also according to ¶ 0046 last 7 lines: “the wake-up model is trained with noise” (based on a user-specific training data) “a corresponding confidence level threshold can be adjusted” (the said threshold (another user-specific training data) was obtained)) ; and
perform user-specific training on the wake-up model based on the user- specific training data to obtain a user-specific wake-up model that is trained to respond to the user (to respond to the user “input” above, according to ¶ 0064 S1: “a confidence level corresponding to “Xiaomei Xiaomei” and a confidence level corresponding to “Assistant Xiaomei” are obtained by the wake-up model based on training data corresponding to the wake-up word “Xiaomei Xiaomei” and training data” (a user specific training based on data) “corresponding to the wake-up word” (associated with recognized user specific recognition of the wake-up word) “Assistant Xiaomei”; ¶ 0046 S2: “subsequent to an extraction of the environmental sound information from the to-be-recognized audio, the environmental sound information” (another user specific) “may be transmitted to a server as training data”( training data on the wake-up model); ¶ 0046 last 7 lines: “the wake-up model is trained” (performing training of the wake-up model) “with noise” (based on a user-specific training data) “a corresponding confidence level threshold” (and also another user-specific training data) “can be adjusted”, because according to ¶ 0009 S1: “training data includes a model parameter and a confidence level threshold”)).
Regarding claim 19, CHEN et al. do teach the non-transitory computer-readable medium of claim 18, wherein the wake-up model comprises a key word detector (KWD) model trained to detect a wake-up command, and a key word verifier (KWV) model trained to verify the wake-up command (¶ 0053 lines 6+: “Then, the audio text information is processed for matching by means of text matching” (detection) “or semantic matching” (and verification) “to determine at least one keyword” (of a key word) “or key character. The at least one keyword or key character is then processed separately using the wake-up model” (in the wake-up model for e.g., the “wake-up word” “Xiaomei Xiaomei” (a keyword which is a wakeup command)) “and the at least two groups of training data”);
wherein to perform the user-specific training comprises training, the instructions further cause the device to train the KWV model using the user-specific training data to obtain a user-specific KWV model (¶ 0064 S1: “a confidence level corresponding to “Xiaomei Xiaomei” and a confidence level corresponding to “Assistant Xiaomei” are obtained by the wake-up model” (the wakeup model functioning also as key word verifier as it recognizes the key word “Xiaomei”) “based on training data” (which is also the user specific training data) “corresponding to the wake-up word “Xiaomei Xiaomei””),
and wherein the user-specific wake-up model comprises the KWD model and the user-specific KWV model ((¶ 0053 lines 6+: “Then, the audio text information is processed for matching by means of text matching” (key word detection) “or semantic matching” (and key word verification) “to determine at least one keyword” “or key character. The at least one keyword or key character is then processed separately using the wake-up model” (as part of the wake-up model) “and the at least two groups of training data”).
Regarding claim 20, CHEN et al. do teach the non-transitory computer-readable medium of claim 18, the instructions further cause the device to:
detect, using the user-specific wake-up model, a new wake-up command included in a new voice input received from the user (¶ 0057 S1: “hypothetically, wake-up word A and wake-up word B” (a new wake-up command included in a new voice input) “are available”);
based on the detecting of the new wake-up command, performing a new speech recognition operation based on the voice input (¶ 0057 S2+: “Subsequent to obtaining training data of wake-up word A and training data of wake-up word B, the to-be-recognized audio may be processed using the wake-up model and the training data of wake-up word A to obtain a confidence level of wake-up word A and a confidence level threshold corresponding to wake-up word A and processed using the wake-up model and the training data of wake-up word B to obtain a confidence level” (performing new speech recognition based on the “wake-up word B” (new wake-up command) “of wake-up word B and a confidence level threshold corresponding to wake-up word B”); and
based on a result of the new speech recognition operation indicating that a performance of the user-specific wake-up model is below a threshold performance, obtaining new user-specific training data based on the new voice input and the result of the new speech recognition operation, and performing additional user-specific training on the user-specific wake-up model based on the new user-specific training data ( ¶ 0063: “It should be noted that for a case where both the first confidence level and the second confidence level are greater than or equal to respective confidence level thresholds corresponding to the first confidence level and the second confidence level, the first value by which the first confidence level exceeds the first confidence level threshold and the second value by which the second confidence level exceeds the second confidence level threshold need to be calculated” (obtaining new user-specific training data based on the new voice input) “In some embodiments, that “triggering a target wake-up event of the voice apparatus based on the first value and the second value” may include: when the first value is greater than or equal to the second value” (when the new speech recognition is below a threshold performance compared to the first recognition results) “determining the target wake-up event as the first wake-up event” “and triggering” (and performing additional training) “the first wake-up event; or when the first value is smaller than the second value, determining the target wake-up event as the second wake-up event and triggering the second wake-up event.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 6-7, 15-16 is/are rejected under 35 U.S.C. 103 as being unpatentable over CHEN et al., and further in view of Kim et al. (EP 3067884 A1).
Regarding claim 6, CHEN et al. do teach the method of claim 1, further comprising: collecting additional user-specific training data; obtaining a user-specific training dataset comprising the user-specific training data and the additional user-specific training data (¶ 0064 S1: “a confidence level corresponding to “Xiaomei Xiaomei” and a confidence level corresponding to “Assistant Xiaomei” are obtained by the wake-up model based on training data corresponding to the wake-up word “Xiaomei Xiaomei” and training data” (the user specific training based on data) “corresponding to the wake-up word” (associated with recognized user specific recognition of the wake-up word) “Assistant Xiaomei”; ¶ 0046 S2: “subsequent to an extraction of the environmental sound information from the to-be-recognized audio, the environmental sound information” (additional user specific) “may be transmitted to a server as training data”( training data on the wake-up model); ¶ 0046 last 7 lines: “the wake-up model is trained” (performing training of the wake-up model) “with noise” (based on a user-specific training data) “a corresponding confidence level threshold” (another additional user-specific training data) “can be adjusted”; together these all are characterized as “training data” (the user-specific training data) which according to ¶ 0096 are: “training data of wake-up word A” “and” “B are stored in a speech module” (user specific training dataset) “for wake-up recognition”);
CHEN et al. do not specifically disclose:
And selecting a time to perform the user-specific training based on at least one parameter corresponding to the electronic device.
Kim et al. do teach:
And selecting a time to perform the user-specific training based on at least one parameter corresponding to the electronic device (¶ 0097 last S: “The device 100 may register a different wake-up keyword model” (a user specific wake-up model training data to be used) “according to the schedule” (e.g., by selecting “6:00 a.m.” (a specific time of start)) “of the user” (based on at least one parameter) “101 that is detected by the device 100” (corresponding to a user electronic device); i.e., ¶ 0097 lines 32-34: “The device 100 may differently register the wake-up keyword model when the time detected by the device 100 is 6 a.m.”).
It would have therefore been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the “schedule[ing]” of “wake-up model[s]” of Kim et al. into the “wake-up model” of CHEN et al. would enable the combined systems and their associated methods to perform in combination as they do separately and to further enable CHEN et al. to invoke its “wake-up keyword model” according to “user” specific “location”, “time” and “weather” as disclosed in Kim et al. ¶ 0097.
Regarding claim 7, CHEN et al. do not specifically disclose the method of claim 6, wherein the at least one parameter comprises at least one from among a computation power of the electronic device, battery usage information of the electronic device, an amount of the user-specific training data collected by the electronic device, and usage pattern information regarding usage patterns of the user.
Kim et al. do teach the method of claim 6, wherein the at least one parameter comprises at least one from among a computation power of the electronic device, battery usage information of the electronic device, an amount of the user-specific training data collected by the electronic device, and usage pattern information regarding usage patterns of the user (¶ 0097 last S: “The device 100 may register a different wake-up keyword model” (a user specific wake-up model training data to be used) “according to the schedule” (e.g., by selecting “6:00 a.m.” (a specific time of start)) “of the user” (based on at least one parameter based on a usage pattern of the user) “101 that is detected by the device 100” (corresponding to a user electronic device)).
For obviousness to combine CHEN et al. and Kim et al. see claim 6.
Regarding claim 15, CHEN et al. do not specifically disclose the electronic device of claim 10, wherein the at least one processor is further configured to execute the instructions to:
select a time to perform the user-specific training based on at least one parameter corresponding to the electronic device.
Kim et al. do teach:
select a time to perform the user-specific training based on at least one parameter corresponding to the electronic device (¶ 0097 last S: “The device 100 may register a different wake-up keyword model” (a user specific wake-up model training data to be used) “according to the schedule” (e.g., by selecting “6:00 a.m.” (a specific time of start)) “of the user” (based on at least one parameter) “101 that is detected by the device 100” (corresponding to a user electronic device); i.e., ¶ 0097 lines 32-34: “The device 100 may differently register the wake-up keyword model when the time detected by the device 100 is 6 a.m.”).
It would have therefore been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the “schedule[ing]” of “wake-up model[s]” of Kim et al. into the “wake-up model” of CHEN et al. would enable the combined systems and their associated methods to perform in combination as they do separately and to further enable CHEN et al. to invoke its “wake-up keyword model” according to “user” specific “location”, “time” and “weather” as disclosed in Kim et al. ¶ 0097.
Regarding claim 16, CHEN et al. do not specifically disclose the electronic device of claim 15, wherein the at least one parameter comprises at least one from among a computation power of the electronic device, battery usage information of the electronic device, an amount of the user-specific training data collected by the electronic device, and usage pattern information regarding usage patterns of the user.
Kim et al. do teach the electronic device of claim 15, wherein the at least one parameter comprises at least one from among a computation power of the electronic device, battery usage information of the electronic device, an amount of the user-specific training data collected by the electronic device, and usage pattern information regarding usage patterns of the user (¶ 0097 last S: “The device 100 may register a different wake-up keyword model” (a user specific wake-up model training data to be used) “according to the schedule” (e.g., by selecting “6:00 a.m.” (a specific time of start)) “of the user” (based on at least one parameter based on a usage pattern of the user) “101 that is detected by the device 100” (corresponding to a user electronic device)).
For obviousness to combine CHEN et al. and Kim et al. see claim 6.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Chang et al. (CN111312222A): “Abstract” teaches “obtaining a wake-up voice” “voice recognition model trained by the wake up speech as input parameter” (i.e., training based on “speech” “input parameters”); “identifying the wake-up voice is contained in the preset wake-up word to obtain the wake-up voice is contained in the preset probability score wake-up word” (determining a confidence score of the “wake-up word” (a wake-up command); “wherein the speech recognition model” (resulting from the speech recognition) “according to the speech sample obtained by training”).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to FARZAD KAZEMINEZHAD whose telephone number is (571)270-5860. The examiner can normally be reached 10:30 am to 11:30 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Paras D. Shah can be reached at (571) 270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Farzad Kazeminezhad/
Art Unit 2653
August 8th 2026.