DETAILED ACTION
This communication is in response to the Amendments and Arguments filed on . Claims 1, 3-13, and 15-31 are pending and have been examined.
Any previous objections/rejections not mentioned in this Office Action has been withdrawn by the examiner.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Change of Examiner
The Examiner of Record has changed from Cameron Young to SPE Paras Shah.
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 04/30/2026 has been entered.
Response to Amendments and Arguments
With respect to the 35 USC 103 rejections of all the claims, the Applicant’s arguments are directed towards the newly added limitations as now recited in the independent claims. Thus, these arguments are moot in view of new grounds for rejection. The primary reference of Cilingir has been retained where two new references have not been incorporated to teach the amended limitations and to replace the prior reference(s).
Further, upon reconsideration, a 35 USC 101 abstract has also been applied to the claims indicated below.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
In this case, the limitations containing the word “means” interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph are present in claim 30. The limitations are:
“means for receiving voice data,”
“means for, … generating a user verification score,”
“means for, … determining a quality of the voice data,” and
“means for, … updating a second user verification ML model.”
“means for determining the voice data includes an utterance…”
“means for confirming that the voice data includes the utterance…”
These limitations are being interpreted in view of the specification and are being treated under the broadest reasonable interpretation in view of the instant disclosure.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1, 3-11, 13, 15-19, and 21-31 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The independent claims 1, 13, 19 and 30 relate to a method, method, system, and system relating to a statutory category. The claims further recite per claim 1 “receiving voice data from a first user; determining that the voice data includes an utterance of a defined keyword based on processing the voice data using a first keyword identification machine learning (ML) model; confirming that the voice data includes the utterance of the defined keyword by processing the voice data using a second keyword identification ML model and generating a keyword score indicating a likelihood that the voice data includes the utterance of the defined keyword, wherein the second keyword identification ML model is more computationally expensive and more accurate than the first keyword identification ML model; in response to confirming that the voice data includes the utterance of the defined keyword: generating a user verification score by processing the voice data using a first user verification ML model, the user verification score indicating a probability that the voice data belongs to the first user; and determining a quality of the voice data; and in response to determining that the user verification score, the keyword score, and the determined quality satisfy one or more defined criteria, updating a second user verification ML model based on the voice data” as recited in claims 1, 19, and 30. Claim 13 recites similar limitations as in claims 1, 19, and 30 with the exception of the last limitation “storing the voice data as a training exemplar”.
The limitation of claims 1, 19, and 30 of “receiving…”, “determining…”, “confirming...”, “generating…”, and “determining…” as drafted covers mental activities. More specifically, a human receiving speech from another user, where the human determines the presence of a specific keyword. Once determined, the human forwards that keyword to their supervisor to determine if the keyword is an authorized keyword to be used by the user, where the supervisor knows the keywords from plural users compared to the human. Then, once the keyword has been confirmed confirming if the user is the correct user that indicated the keyword based on the voice and determining the quality of the received audio. Once this is determined, providing, the keyword, description of the voice and user’s name to a co-worker so they can remember the user next time. With respect to claim 13, all of the limitations are the same except for the last limitation where a human can remember the data received as part of an example.
This judicial exception is not integrated into a practical application. In particular, claims 1 and 13 recites additional limitations of “machine learning model” (first and second) and claims 19 and 30 recite “memory”, “processor” and “machine learning model”. Each of these elements are being used as a tool on which the stated steps operate on. Paragraph [0208] describe the processor and CRM can be implemented using general purpose circuity, processor etc. Para [0163] notes that the machine learning model can be ANN, DNN, “the like” thus signifying any well-known machine learning algorithms can be used. Accordingly, since there are no additional elements then the abstract idea cannot be integrated into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claims are not patent eligible.
With respect to claim 3 and 21, the claim relates to “wherein determining the quality of the voice data comprises at least one of: determining a signal-to-noise (SNR) ratio of the voice data; determining a clipping ratio of the voice data; or determining a duration of the voice data..” This reads on a human determining the duration of the speech that was spoken to determine quality, for example. No additional limitations are present. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception other than those already mention in the independent claims.
With respect to claim 4, 15, and 22, the claims relate to “storing the voice data as a training exemplar, wherein updating the second user verification ML model is performed based further in response to determining that a number of stored training exemplars satisfies one or more defined criteria.” This relates to a human determining whether to keep the training data based on number of training data present or closeness to another voice keyword. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
With respect to claim 5, 16, and 23, the claim relates to “subsequent to updating the second user verification ML model, deleting the stored training exemplars.” This relates to the human deleting the training exemplar after noting the patterns. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
With respect to claim 6, 17, and 24, the claim relates to “subsequent to updating the second user verification ML model, storing the training exemplars in a storage location that satisfies one or more defined security criteria. ” This relates to a human storing and providing relevant coworkers the training exemplars that are authorized at a specific security level. No additional elements are present. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
With respect to claims 7, 18, and 25, the claim relates to “using the second user verification ML model to process subsequent voice data.” This relates the coworker using the training data to determine and verify the user. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
With respect to claim 8 and 26, the claim relates to “wherein updating the second user verification ML model based on the voice data comprises: extracting one or more features of the voice data; labeling the one or more features of the voice data based on the user verification score; and storing the one or more features and the label as a training exemplar.” This relates to a human extracting specific features of the voice such as tone, pitch, and labelling such sounds on a piece of paper with a known spectrum and then storing them in memory/paper based on the annotations. No additional elements are present. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
With respect to claim 9 and 27, the claim relates to “wherein updating the second user verification ML model based on the voice data is performed based further in response to determining that a number of stored training exemplars satisfies one or more defined criteria, wherein the one or more defined criteria indicate at least one of:a minimum number of stored positive exemplars corresponding to utterances made by the first user;a minimum number of stored negative exemplars corresponding to utterances not made by the first user; or a ratio of stored positive exemplars to stored negative exemplars.” This relates to a human determining number of examples for a specific used and determining if min number of examples are present. No additional elements are present. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
With respect to claim 10 and 28, the claim relates to “wherein updating the second user verification ML model is performed using a federated learning operation.” The human can get help from another human to perform learning of the data. The claim do not provide how the federated learning is different from conventional systems as outlined in [0070]. No additional elements are present. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
With respect to claim 11 and 29, the claim relates to “wherein the federated learning operation comprises transmitting the training exemplar to a host system that performs the updating of the second user verification ML model.” This relates to a human sending the training example to their coworker for updating their understanding. No additional elements are present. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
With respect to claim 31, the claim relates to “wherein the one or more defined criteria are associated with one or more thresholds that are higher than one or more threshold used for non-training processing of the voice data.” This relates to a human where the threshold is set higher during training than non-training. No additional elements are present. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
These claims further do not remedy the judicial exception being integrated into a practical application and further fail to include additional elements that are sufficient to amount to significantly more than the judicial exception.
NOTE: Claim 12 has not been included as part of this rejection.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim 13 is rejected under 35 U.S.C. 103 as being unpatentable over U.S. Patent Application Publication No. 2018/0366124 A1 to Gokcen Cilingir et al. (hereinafter Cilingir) in view of US Patent 12,525.250 to Yiteng Huang (hereinafter Huang) in view of U.S. Patent Application Publication No. 2016/0071516 to Minsuh Lee (US 2016/0071516).
Regarding claim 13, Cilingir teaches a computer-implemented method for performing user verification using machine learning, comprising: (Cilingir teaches training a text independent speaker recognition model (i.e., machine learning model) to identify a user. (i.e., verify a user) Cilingir at ¶ [0013]. Further, Cilingir teaches training the speaker recognition model (i.e., the speaker recognition model learns as does a machine learning model.))
receiving voice data from a first user; (Cilingir teaches collecting speech utterances (i.e., voice data) from a user. Cilingir at ¶ [0050]. Further, Cilingir teaches the speech utterances may be collected over a period of time and may be represented as feature vectors. Cilingir at ¶ [0050].)
in response to confirming that the voice data includes the utterance of the defined keyword: generating a user verification score by processing the voice data using a first user verification ML model, the user verification score indicating a probability that the voice data belongs to the user; (Cilingir teaches after recognizing a pre-defined utterance such as "hello computer" or "wake up," generating a confidence value (i.e., a verification score) by processing the audio using the speaker recognition model (i.e., user verification machine learning model). Cilingir at ¶¶ [0016] - [0017]. Cilingir teaches generating a speaker ID and associated confidence value based on processed utterances for a plurality of users. (i.e., generating a user verification score indicating a probability that the voice belongs to the user associated with the speaker ID.) Cilingir at ¶¶ [0016] - [0017]. Further, A person of ordinary skill in the art would have understood that a confidence value associated with a speaker ID determined by processing an utterance amounts to an indication of a probability that the utterance belongs to that user. (e.g., probability is a base element of confidence, and confidence values are probabilities.))
and determining a quality of the voice data; (Cilingir teaches evaluating the quality of the speech utterances and the state of the speaker (i.e., determining a quality of the voice data). Cilingir at ¶ [0014].)
and in response to determining that the user verification score and the determined quality satisfy one or more defined criteria, storing the voice data as a training exemplar. (Cilingir teaches storing speech utterances as training data in a training database (i.e., storing the voice data as training exemplars (i.e., training examples or training data)) Cilingir at ¶ [0014].)
Cilingir, however, does not teach determining that the voice data includes an utterance of a defined keyword based on processing the voice data using a first keyword identification machine learning (ML) model; and confirming that the voice data includes the utterance of the defined keyword by processing the voice data using a second keyword identification ML model, wherein the second keyword identification ML model is more accurate than the first keyword identification ML model.
In a similar field of endeavor (e.g., natural language processing of audio data from a speaking individual and multi-step machine learning processes), Huang teaches determining that the voice data includes an utterance of a defined keyword based on processing the voice data using a first keyword identification machine learning (ML) model; (see col. 8, lines 30-57, where the first stage hotword detector is a smaller model size for coarse screening and the second state hotword detector is of larger size and see col. 6, lines 41-45, where hotword detectors implemented using neural networks, where the detection of a hotword is determined)
confirming that the voice data includes the utterance of the defined keyword by processing the voice data using a second keyword identification ML model and generating a keyword score indicating a likelihood that the voice data includes the utterance of the define keyword, wherein the second keyword identification ML model is more computationally expensive and more accurate than the first keyword identification ML model; (see col 8, lines 30-57, where the first stage hotword detector is smaller than the second and where the second hotword detector is of larger size providing more accurate detection and see col. 11, lines 1-8, where second hotword detector detects presence of a hotword base d on a probability score)
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date to have modified the wakeword detection of Cilingir with the multiple model hotword detection of Huang in order to optimize power consumption, latency, and noise robustness (see Huang col. 7, lines 61-col. 8, lines 5).
However, Cilingir in view of Huang do not specifically teach in response to determining that the … keyword score…satisfy one or more defined criteria, storing the voice data as a training exemplar. The Examiner notes that Cilingir already provides a teaching of storing training data based on satisfaction of criteria (see above) but not with respect to a keyword score.
Lee does teach in response to determining that the … keyword score…satisfy one or more defined criteria, storing the voice data as a training exemplar (see [0068], where based on confidence score of the keyword, determination is made to replace the keyword model with another speaker independent model for the keyword).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date to have modified the wakeword detection of Cilingir in view of Huang with the adding/updating of a model based on new information as taught by Lee in order to overcome issues due to differing voice characteristics or pronunciation of users (see Lee [0004]).
Claims 1, 4, 7 – 8, 10 – 12, 19, 22, 25, 26, and 28 – 30, is/are rejected under 35 U.S.C. 103 as being unpatentable over Cilingir in view of U.S. Patent Application Publication No. 2022/0293093 A1 to Francoise Beaufays et al. (hereinafter Beaufays) and in further view of Huang (US 12525250) in view of Lee (US 20160071516).
Regarding claim 1, Cilingir teaches a computer-implemented method for training a machine learning model for user verification, comprising: (Cilingir teaches training a text independent speaker recognition model (i.e., machine learning model) to identify a user. (i.e., verify a user) Cilingir at ¶ [0013]. Further, Cilingir teaches training the speaker recognition model (i.e., the speaker recognition model learns as does a machine learning model.))
receiving voice data from a first user; (Cilingir teaches collecting speech utterances (i.e., voice data) from a user. Cilingir at ¶ [0050]. Further, Cilingir teaches the speech utterances may be collected over a period of time and may be represented as feature vectors. Cilingir at ¶ [0050].)
in response to confirming that the voice data includes an utterance of the defined keyword: generating a user verification score by processing the voice data using a first user verification ML model, the user verification score indicating a probability that the voice data belongs to the user; (Cilingir teaches after recognizing a pre-defined utterance such as "hello computer" or "wake up," generating a confidence value (i.e., a verification score) by processing the audio using the speaker recognition model (i.e., user verification machine learning model). Cilingir at ¶¶ [0016] - [0017]. Cilingir teaches generating a speaker ID and associated confidence value based on processed utterances for a plurality of users. (i.e., generating a user verification score indicating a probability that the voice belongs to the user associated with the speaker ID.) Cilingir at ¶¶ [0016] - [0017]. Further, A person of ordinary skill in the art would have understood that a confidence value associated with a speaker ID determined by processing an utterance amounts to an indication of a probability that the utterance belongs to that user. (e.g., probability is a base element of confidence, and confidence values are probabilities.))
and determining a quality of the voice data; (Cilingir teaches evaluating the quality of the speech utterances and the state of the speaker (i.e., determining a quality of the voice data). Cilingir at ¶ [0014].)
and in response to determining that the user verification score and determined quality satisfy one or more defined criteria, updating a … ML model based on the voice data. (Cilingir teaches updating the speaker recognition model based on the speech quality analysis and a training merit analysis exceeding a certain threshold. (i.e., if the quality is high enough and the user was identified previously (the user is identified before quality analysis) then the model is trained based on the speech utterances.) Cilingir at ¶ [0014].)
Cilingir, however, does not teach updating a second user verification ML model based on the voice data.
In a similar field of endeavor (e.g., updating and training machine learning models used in keyword detection.) Beaufays teaches updating a second user verification ML model based on the voice data. (Beaufays teaches updating a global machine learning model using outputs from a first (or local) machine learning model. (i.e., Cilingir's speaker recognition model in view of Beaufays training of global models based on results of a local ml model is a second user verification model updated by the outputs of a first user verification model.) Beaufays at ¶¶ [0024] – [0040].)
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date to combine the teachings of Cilingir with the teachings of Beaufays to provide updating a second user verification ML model based on the voice data. Doing so would have provided a noticeable improvement to the machine learning models as recognized by Beaufays at ¶ [0076].
Further, Cilingir in view of Beaufays (hereinafter Cilingir-Beaufays) does not teach determining that the voice data includes an utterance of a defined keyword based on processing the voice data using a first keyword identification machine learning (ML) model; and confirming that the voice data includes the utterance of the defined keyword by processing the voice data using a second keyword identification ML model, wherein the second keyword identification ML model is more computationally expensive and more accurate than the first keyword identification ML model;
In a similar field of endeavor (e.g., natural language processing of audio data from a speaking individual and multi-step machine learning processes), Huang teaches determining that the voice data includes an utterance of a defined keyword based on processing the voice data using a first keyword identification machine learning (ML) model; (see col. 8, lines 30-57, where the first stage hotword detector is a smaller model size for coarse screening and the second state hotword detector is of larger size and see col. 6, lines 41-45, where hotword detectors implemented using neural networks, where the detection of a hotword is determined)
confirming that the voice data includes the utterance of the defined keyword by processing the voice data using a second keyword identification ML model, wherein the second keyword identification ML model is more computationally expensive and more accurate than the first keyword identification ML model; (Huang teaches confirming the results of the first biometric process using a second machine learning voice biometric process. (see col 8, lines 30-57, where the first stage hotword detector is smaller than the second and where the second hotword detector is of larger size providing more accurate detection and see col. 11, lines 1-8, where second hotword detector detects presence of a hotword base d on a probability score)
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date to have modified the wakeword detection of Cilingir with the multiple model hotword detection of Huang in order to optimize power consumption, latency, and noise robustness (see Huang col. 7, lines 61-col. 8, lines 5).
However, Cilingir in view of Huang do not specifically teach in response to determining that the … keyword score…satisfy one or more defined criteria, storing the voice data as a training exemplar. The Examiner notes that Cilingir already provides a teaching of storing training data based on satisfaction of criteria (see above) but not with respect to a keyword score.
Lee does teach in response to determining that the … keyword score…satisfy one or more defined criteria, storing the voice data as a training exemplar (see [0068], where based on confidence score of the keyword, determination is made to replace the keyword model with another speaker independent model for the keyword).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date to have modified the wakeword detection of Cilingir in view of Huang with the adding/updating of a model based on new information as taught by Lee in order to overcome issues due to differing voice characteristics or pronunciation of users (see Lee [0004]).
Regarding claim 4, Cilingir-Beaufays-Huang-Lee teaches all the limitations of claim 1 as laid out above. Further, Cilingir teaches the computer-implemented method of Claim 1, further comprising storing the voice data as a training exemplar, (Cilingir teaches storing speech utterances as training data in a training database (i.e., storing the voice data as training exemplars (i.e., training examples or training data)) Cilingir at ¶ [0014].)
wherein updating the second user verification ML model is performed based further in response to determining that a number of stored training exemplars satisfies one or more defined criteria. (Cilingir teaches determining a level of sufficiency (i.e., the stored training materials satisfy one or more defined criteria) before beginning training. Cilingir at ¶ [0014].)
Regarding claim 7, Cilingir-Beaufays-Huang-Lee teaches all the limitations of claim 1 as laid out above. Further, Beaufays teaches the computer-implemented method of Claim 1, further comprising using the second user verification ML model to process subsequent voice data. (Beaufays teaches transferring the updated global machine learning model to the local device and replacing the first model with the updated global model to further process data using the updated machine learning model (i.e., the second machine learning model is used to process subsequent data.) Beaufays at ¶¶ [0026] - [0030].)
Regarding claim 8, Cilingir-Beaufays-Huang-Lee teaches all the limitations of claim 1 as laid out above. Further, Cilingir teaches the computer-implemented method of Claim 1, wherein updating the second user verification ML model based on the voice data comprises: extracting one or more features of the voice data; (Cilingir teaches the speech utterances that are stored as training data may be represented as feature vectors (i.e., the speech data is processed to remove a feature vector which represents the utterance.) Cilingir at ¶ [0017].)
labeling the one or more features of the voice data based on the user verification score; (Cilingir teaches the speech utterances are stored indexed by user identity (i.e., labeled based on the user verification score). Cilingir at ¶ [0031].)
and storing the one or more features and the label as a training exemplar. (Cilingir teaches storing utterances as training data (i.e., storing the voice data as training exemplars (i.e., training examples or training data)) Cilingir at ¶ [0014]. Further, because the labels are stored indexed by user identity (i.e., based on the user verification score) and the utterances may be represented as a feature vector, then storing the utterances indexed by user identity is storing the one or more features and the label as a training exemplar.)
Regarding claim 10, Cilingir-Beaufays-Huang-Lee teaches all the limitations of claim 8 as laid out above. Further, Beaufays teaches the computer-implemented method of Claim 8, wherein updating the second user verification ML model is performed using a federated learning operation. (Beaufays discloses training machine learning models using federated learning. (i.e., using federated learning to update the recognition model of Cilingir.) Beaufays at Abstract, ¶¶ [0003] and [0043].)
Regarding claim 11, Cilingir-Beaufays-Huang-Lee teaches all the limitations of claim 10 as laid out above. Further, Beaufays teaches the computer-implemented method of Claim 10, wherein the federated learning operation comprises transmitting the training exemplar to a host system that performs the updating of the second user verification ML model. (Beaufays teaches a remote training engine that uses a plurality of gradients (i.e., aggregated parameters) to update the weights of a global machine learning model (i.e., a global version of the machine learning model.) Beaufays at ¶¶ [0025] - [0026].)
Regarding claim 12, Cilingir-Beaufays- Huang-Lee teaches all the limitations of claim 10 as laid out above. Further, Beaufays teaches the computer-implemented method of Claim 10, wherein the federated learning operation comprises transmitting updated parameters of the second user verification ML model to a host system, wherein the host system aggregates updated parameters to update a global version of the second user verification ML model. (Beaufays teaches a remote training engine that uses a plurality of gradients (i.e., aggregated parameters) to update the weights of a global machine learning model (i.e., a global version of the machine learning model.) Beaufays at ¶¶ [0025] - [0026].)
Regarding claim 19, Cilingir teaches a processing system, comprising: (Cilingir teaches a platform (i.e., a processing system) implemented on computer hardware. Cilingir at ¶ [0052].)
a memory comprising computer-executable instructions; (Cilingir teaches the platform comprising memory (i.e., memory for storing computer-executable instructions.) Cilingir at ¶ [0052].)
and one or more processors configured to execute the computer-executable instructions and cause the processing system to perform an operation comprising: (Cilingir teaches the platform comprising a processor (i.e., one or more processors). Cilingir at ¶ [0050].)
receiving voice data from a first user; (Cilingir teaches collecting speech utterances (i.e., voice data) from a user. Cilingir at ¶ [0050]. Further, Cilingir teaches the speech utterances may be collected over a period of time and may be represented as feature vectors. Cilingir at ¶ [0050].)
in response to confirming that the voice data includes an utterance of a defined keyword: generating a user verification score by processing the voice data using a first user verification ML model, the user verification score indicating a probability that the voice data belongs to the first user; (Cilingir teaches after recognizing a pre-defined utterance such as "hello computer" or "wake up," generating a confidence value (i.e., a verification score) by processing the audio using the speaker recognition model (i.e., user verification machine learning model). Cilingir at ¶¶ [0016] - [0017]. Cilingir teaches generating a speaker ID and associated confidence value based on processed utterances for a plurality of users. (i.e., generating a user verification score indicating a probability that the voice belongs to the user associated with the speaker ID.) Cilingir at ¶¶ [0016] - [0017]. Further, A person of ordinary skill in the art would have understood that a confidence value associated with a speaker ID determined by processing an utterance amounts to an indication of a probability that the utterance belongs to that user. (e.g., probability is a base element of confidence, and confidence values are probabilities.))
and determining a quality of the voice data; (Cilingir teaches evaluating the quality of the speech utterances and the state of the speaker (i.e., determining a quality of the voice data). Cilingir at ¶ [0014].)
and in response to determining that the user verification score and the determined quality satisfy one or more defined criteria, updating … model based on the voice data. (Cilingir teaches updating the speaker recognition model based on the speech quality analysis and a training merit analysis exceeding a certain threshold. (i.e., if the quality is high enough and the user was identified previously (the user is identified before quality analysis) then the model is trained based on the speech utterances.) Cilingir at ¶ [0014].)
Cilingir, however, does not teach updating a second user verification ML model based on the voice data.
In a similar field of endeavor (e.g., updating and training machine learning models used in keyword detection.) Beaufays teaches updating a second user verification ML model based on the voice data. (Beaufays teaches updating a global machine learning model using outputs from a first (or local) machine learning model. (i.e., Cilingir's speaker recognition model in view of Beaufays training of global models based on results of a local ml model is a second user verification model updated by the outputs of a first user verification model.) Beaufays at ¶¶ [0024] – [0040].)
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date to combine the teachings of Cilingir with the teachings of Beaufays to provide updating a second user verification ML model based on the voice data. Doing so would have provided a noticeable improvement to the machine learning models as recognized by Beaufays at ¶ [0076].
Further, Cilingir-Beaufays does not teach determining that the voice data includes an utterance of a defined keyword based on processing the voice data using a first keyword identification machine learning (ML) model; and confirming that the voice data includes the utterance of the defined keyword by processing the voice data using a second keyword identification ML model, wherein the second keyword identification ML model is more computationally expensive and more accurate than the first keyword identification ML model;
In a similar field of endeavor (e.g., natural language processing of audio data from a speaking individual and multi-step machine learning processes), Huang teaches determining that the voice data includes an utterance of a defined keyword based on processing the voice data using a first keyword identification machine learning (ML) model; (see col. 8, lines 30-57, where the first stage hotword detector is a smaller model size for coarse screening and the second state hotword detector is of larger size and see col. 6, lines 41-45, where hotword detectors implemented using neural networks, where the detection of a hotword is determined)
confirming that the voice data includes the utterance of the defined keyword by processing the voice data using a second keyword identification ML model, wherein the second keyword identification ML model is more computationally expensive and more accurate than the first keyword identification ML model; (Huang teaches confirming the results of the first biometric process using a second machine learning voice biometric process. (see col 8, lines 30-57, where the first stage hotword detector is smaller than the second and where the second hotword detector is of larger size providing more accurate detection and see col. 11, lines 1-8, where second hotword detector detects presence of a hotword base d on a probability score)
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date to have modified the wakeword detection of Cilingir-Beaufays with the multiple model hotword detection of Huang in order to optimize power consumption, latency, and noise robustness (see Huang col. 7, lines 61-col. 8, lines 5).
However, Cilingir-Beaufays-Huang do not specifically teach in response to determining that the … keyword score…satisfy one or more defined criteria, storing the voice data as a training exemplar. The Examiner notes that Cilingir already provides a teaching of storing training data based on satisfaction of criteria (see above) but not with respect to a keyword score.
Lee does teach in response to determining that the … keyword score…satisfy one or more defined criteria, storing the voice data as a training exemplar (see [0068], where based on confidence score of the keyword, determination is made to replace the keyword model with another speaker independent model for the keyword).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date to have modified the wakeword detection of Cilingir-Beaufays-Huang with the adding/updating of a model based on new information as taught by Lee in order to overcome issues due to differing voice characteristics or pronunciation of users (see Lee [0004]).
Regarding claim 22, Cilingir in view of Beaufays in view of Huang in view of Lee (hereinafter Cilingir-Beaufays-Huang-Lee) teaches all the limitations of claim 19 as laid out above. Further, Cilingir teaches the processing system of claim 19, the operation further comprising storing the voice data as a training exemplar, (Cilingir teaches storing speech utterances as training data in a training database (i.e., storing the voice data as training exemplars (i.e., training examples or training data)) Cilingir at ¶ [0014].)
wherein updating the second user verification ML model is performed based further in response to determining that a number of stored training exemplars satisfies one or more defined criteria. (Cilingir teaches determining a level of sufficiency (i.e., the stored training materials satisfy one or more defined criteria) before beginning training. Cilingir at ¶ [0014].)
Regarding claim 25, Cilingir-Beaufays-Huang-Lee teaches all the limitations of claim 19 as laid out above. Further, Beaufays teaches the processing system of Claim 19, the operation further comprising using the second user verification ML model to process subsequent voice data. (Beaufays teaches transferring the updated global machine learning model to the local device and replacing the first model with the updated global model to further process data using the updated machine learning model (i.e., the second machine learning model is used to process subsequent data.) Beaufays at ¶¶ [0026] - [0030].)
Regarding claim 26, Cilingir-Beaufays-Huang-Lee teaches all the limitations of claim 19 as laid out above. Further, Cilingir teaches the processing system of Claim 19, wherein updating the second user verification ML model based on the voice data comprises: extracting one or more features of the voice data; (Cilingir teaches the speech utterances that are stored as training data may be represented as feature vectors (i.e., the speech data is processed to remove a feature vector which represents the utterance.) Cilingir at ¶ [0017].)
labeling the one or more features of the voice data based on the user verification score; (Cilingir teaches the speech utterances are stored indexed by user identity (i.e., labeled based on the user verification score). Cilingir at ¶ [0031].)
and storing the one or more features and the label as a training exemplar. (Cilingir teaches storing utterances as training data (i.e., storing the voice data as training exemplars (i.e., training examples or training data)) Cilingir at ¶ [0014]. Further, because the labels are stored indexed by user identity (i.e., based on the user verification score) and the utterances may be represented as a feature vector, then storing the utterances indexed by user identity is storing the one or more features and the label as a training exemplar.)
Regarding claim 28, Cilingir-Beaufays-Huang-Lee teaches all the limitations of claim 26 as laid out above. Further, Beaufays teaches the processing system of Claim 26, wherein updating the second user verification ML model is performed using a federated learning operation. (Beaufays discloses training machine learning models using federated learning. (i.e., using federated learning to update the recognition model of Cilingir.) Beaufays at Abstract, ¶¶ [0003] and [0043].)
Regarding claim 29, Cilingir-Beaufays-Huang-Lee teaches all the limitations of claim 28 as laid out above. Further, Beaufays teaches the processing system of Claim 28, wherein the federated learning operation comprises transmitting the training exemplar to a host system that performs the updating of the second user verification ML model. (Beaufays teaches a remote training engine that uses a plurality of gradients (i.e., aggregated parameters) to update the weights of a global machine learning model (i.e., a global version of the machine learning model.) Beaufays at ¶¶ [0025] - [0026].)
Regarding claim 30, Cilingir teaches A processing system, comprising: (Cilingir teaches a platform (i.e., a processing system) implemented on computer hardware. Cilingir at ¶ [0050].)
means for receiving voice data from a first user; (Cilingir teaches collecting speech utterances (i.e., voice data) from a user. Cilingir at ¶ [0050]. Further, Cilingir teaches the speech utterances may be collected over a period of time and may be represented as feature vectors. Cilingir at ¶ [0050].)
means for, in response to determining that the voice data includes an utterance of a defined keyword: generating a user verification score by processing the voice data using a first user verification machine learning (ML) model, the user verification score indicating a probability that the voice data belongs to the first user; (Cilingir teaches after recognizing a pre-defined utterance such as "hello computer" or "wake up," generating a confidence value (i.e., a verification score) by processing the audio using the speaker recognition model (i.e., user verification machine learning model). Cilingir at ¶¶ [0016] - [0017]. Cilingir teaches generating a speaker ID and associated confidence value based on processed utterances for a plurality of users. (i.e., generating a user verification score indicating a probability that the voice belongs to the user associated with the speaker ID.) Cilingir at ¶¶ [0016] - [0017]. Further, A person of ordinary skill in the art would have understood that a confidence value associated with a speaker ID determined by processing an utterance amounts to an indication of a probability that the utterance belongs to that user. (e.g., probability is a base element of confidence, and confidence values are probabilities.))
and determining a quality of the voice data; (Cilingir teaches evaluating the quality of the speech utterances and the state of the speaker (i.e., determining a quality of the voice data). Cilingir at ¶ [0014].)
and means for, in response to determining that the user verification score and determined quality satisfy one or more defined criteria, updating a … model based on the voice data. (Cilingir teaches updating the speaker recognition model based on the speech quality analysis and a training merit analysis exceeding a certain threshold. (i.e., if the quality is high enough and the user was identified previously (the user is identified before quality analysis) then the model is trained based on the speech utterances.) Cilingir at ¶ [0014].)
Cilingir, however, does not teach updating a second user verification ML model based on the voice data.
In a similar field of endeavor (e.g., updating and training machine learning models used in keyword detection.) Beaufays teaches updating a second user verification ML model based on the voice data. (Beaufays teaches updating a global machine learning model using outputs from a first (or local) machine learning model. (i.e., Cilingir's speaker recognition model in view of Beaufays training of global models based on results of a local ml model is a second user verification model updated by the outputs of a first user verification model.) Beaufays at ¶¶ [0024] – [0040].)
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date to combine the teachings of Cilingir with the teachings of Beaufays to provide updating a second user verification ML model based on the voice data. Doing so would have provided a noticeable improvement to the machine learning models as recognized by Beaufays at ¶ [0076].
Further, Cilingir-Beaufays does not teach means for determining that the voice data includes an utterance of a defined keyword based on processing the voice data using a first keyword identification machine learning (ML) model; and means for confirming that the voice data includes the utterance of the defined keyword by processing the voice data using a second keyword identification ML model, wherein the second keyword identification ML model is more computationally expensive and more accurate than the first keyword identification ML model.
In a similar field of endeavor (e.g., natural language processing of audio data from a speaking individual and multi-step machine learning processes), Huang teaches determining that the voice data includes an utterance of a defined keyword based on processing the voice data using a first keyword identification machine learning (ML) model; (see col. 8, lines 30-57, where the first stage hotword detector is a smaller model size for coarse screening and the second state hotword detector is of larger size and see col. 6, lines 41-45, where hotword detectors implemented using neural networks, where the detection of a hotword is determined)
confirming that the voice data includes the utterance of the defined keyword by processing the voice data using a second keyword identification ML model, wherein the second keyword identification ML model is more computationally expensive and more accurate than the first keyword identification ML model; (Huang teaches confirming the results of the first biometric process using a second machine learning voice biometric process. (see col 8, lines 30-57, where the first stage hotword detector is smaller than the second and where the second hotword detector is of larger size providing more accurate detection and see col. 11, lines 1-8, where second hotword detector detects presence of a hotword base d on a probability score)
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date to have modified the wakeword detection of Cilingir-Beaufays with the multiple model hotword detection of Huang in order to optimize power consumption, latency, and noise robustness (see Huang col. 7, lines 61-col. 8, lines 5).
However, Cilingir-Braufays-Huang do not specifically teach in response to determining that the … keyword score…satisfy one or more defined criteria, storing the voice data as a training exemplar. The Examiner notes that Cilingir already provides a teaching of storing training data based on satisfaction of criteria (see above) but not with respect to a keyword score.
Lee does teach in response to determining that the … keyword score…satisfy one or more defined criteria, storing the voice data as a training exemplar (see [0068], where based on confidence score of the keyword, determination is made to replace the keyword model with another speaker independent model for the keyword).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date to have modified the wakeword detection of Cilingir-Beaufays-Huang with the adding/updating of a model based on new information as taught by Lee in order to overcome issues due to differing voice characteristics or pronunciation of users (see Lee [0004]).
Claims 15 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Cilingir in view of Huang in view of Lee and in further view of U.S. Patent Application Publication No. 2022/0293093 A1 to Francoise Beaufays et al. (hereinafter Beaufays).
Regarding claim 15, Cilingir-Huang-Lee teaches all the limitations of claim 13 as laid out above. Further, Cilingir teaches the computer-implemented method of Claim 13, further comprising updating a … model based further in response to determining that a number of stored training exemplars satisfies one or more defined criteria. (Cilingir teaches determining a level of sufficiency (i.e., the stored training materials satisfy one or more defined criteria) before beginning training of a speaker recognition model (i.e., updating a second model). Cilingir at ¶ [0014].)
Cilingir-Huang-Lee, however, does not teach updating a second user verification ML model.
In a similar field of endeavor (e.g., updating and training machine learning models used in keyword detection.) Beaufays teaches updating a second user verification ML model. (Beaufays teaches updating a global machine learning model using outputs from a first (or local) machine learning model. (i.e., Cilingir's speaker recognition model in view of Beaufays training of global models based on results of a local ml model is a second user verification model updated by the outputs of a first user verification model.) Beaufays at ¶¶ [0024] – [0040].)
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date to combine the teachings of Cilingir-Huang-Lee with the teachings of Beaufays to provide updating a second user verification ML model. Doing so would have provided a noticeable improvement to the machine learning models as recognized by Beaufays at ¶ [0076].
Regarding claim 18, Cilingir-Huang-Lee-Beaufays teaches all the limitations of claim 15 as laid out above. Further, Beaufays teaches the computer-implemented method of Claim 15, further comprising using the second user verification ML model to process subsequent voice data. (Beaufays teaches transferring the updated global machine learning model to the local device and replacing the first model with the updated global model to further process data using the updated machine learning model (i.e., the second machine learning model is used to process subsequent data.) Beaufays at ¶¶ [0026] - [0030].)
Claims 3 and 21 are rejected under 35 U.S.C. 103 as being unpatentable over Cilingir-Beaufays-Huang-Lee as applied to claims 1 and 19 above, and further in view of U.S. Patent Application Publication No. 2023/0048401 A1 to William E., Sherwood et al. (hereinafter Sherwood).
Regarding claim 3, Cilingir-Beaufays-Huang-Lee teaches all the limitations of claim 1 as laid out above Cilingir-Beaufays-Huang-Lee, however, do not teach the limitations of claim 3.
Sherwood teaches the computer-implemented method of Claim 1, wherein determining the quality of the voice data comprises at least one of: determining a signal-to-noise (SNR) ratio of the voice data; determining a clipping ratio of the voice data; or determining a duration of the voice data. (Sherwood teaches determining the quality of an audio signal including signal to noise ratio, presence of clipping, and positive or negative durations of the audio waveform (i.e., duration of the voice data.) Sherwood at ¶¶ [0104] – [0111]. Further, Sherwood teaches authentication a user. Sherwood at ¶ [0073].)
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date to combine the teachings of Cilingir-Beaufays-Huang-Lee with the teachings of Sherwood to provide the limitations of claim 3. Doing so would have aided in determining whether or not the voice data should be used to recognize the user as recognized by Sherwood at ¶¶ [0104] – [0111].
Regarding claim 21, Cilingir-Beaufays-Huang-Lee teaches all the limitations of claim 19 as laid out above. Cilingir-Beaufays-Huang-Lee, however, do not teach the limitations of claim 21.
Sherwood teaches the processing system of Claim 19, wherein determining the quality of the voice data comprises at least one of: determining a signal-to-noise (SNR) ratio of the voice data; determining a clipping ratio of the voice data; or determining a duration of the voice data. (Sherwood teaches determining the quality of an audio signal including signal to noise ratio, presence of clipping, and positive or negative durations of the audio waveform (i.e., duration of the voice data.) Sherwood at ¶¶ [0104] – [0111]. Further, Sherwood teaches authentication a user. Sherwood at ¶ [0073].)
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date to combine the teachings of Cilingir-Beaufays-Huang-Lee with the teachings of Sherwood to provide the limitations of claim 21. Doing so would have aided in determining whether or not the voice data should be used to recognize the user as recognized by Sherwood at ¶¶ [0104] – [0111].
Claims 5, 9, 23, and 27 is/are rejected under 35 U.S.C. 103 as being unpatentable over Cilingir-Beaufays-Huang-Lee as applied to claims 4, 8, 15, 22, and 26 above, and further in view of U.S. Patent Application Publication No. 2021/0326421 A1 to Elie Khoury et al. (hereinafter Khoury).
Regarding claim 5, Cilingir-Beaufays-Huang-Lee teaches all the limitations of claim 1 as laid out above. Cilingir-Beaufays-Huang-Lee, however, does not teach the computer-implemented method of Claim 4, further comprising, subsequent to updating the second user verification ML model, deleting the stored training exemplars.
Khoury teaches the computer-implemented method of Claim 4, further comprising, subsequent to updating the second user verification ML model, deleting the stored training exemplars. (Khoury teaches training models using stored voiceprints (i.e., training exemplars). Khoury at ¶¶ [0048] – [0049]. Then Khoury discloses deleting stored voiceprints for training models (i.e., training exemplars) from a server. Khoury at ¶ [0178]. Thus, Khoury teaches deleting the voiceprints used to train models after the training has been completed.).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date to combine the teachings of Cilingir-Beaufays-Huang-Lee with the teachings of Khoury to provide subsequent to updating the second user verification ML model, deleting the stored training exemplars. Doing so would have improved the accuracy of the speaker recognition models as recognized by Khoury at ¶ [0060].
Regarding claim 9, Cilingir-Beaufays-Huang-Lee teaches all the limitations of claim 8 as laid out above. Further, Beaufays teaches the computer-implemented method of Claim 8, wherein updating the second user verification ML model based on the voice data is performed based further in response to determining that a number of stored training exemplars satisfies one or more defined criteria, wherein the one or more defined criteria indicate at least one of: (Cilingir-Beaufays teaches updating a global machine learning model using outputs from a first (or local) machine learning model. (i.e., Cilingir's speaker recognition model in view of Beaufays training of global models based on results of a local ml model is a second user verification model updated by the outputs of a first user verification model.) Beaufays at ¶¶ [0024] - [0040].)
Cilingir-Beaufays-Huang-Lee, however, does not teach the one or more defined criteria indicates at least one of. a minimum number of stored positive exemplars corresponding to utterances made by the first user; a minimum number of stored negative exemplars corresponding to utterances not made by the first user; or a ratio of stored positive exemplars to stored negative exemplars.
Khoury teaches a minimum number of stored positive exemplars corresponding to utterances made by a first user; (Khoury teaches processing a large number of utterances using a machine learning architecture. Khoury at ¶ [0233]. Further, Khoury teaches determining a similarity score for utterances by comparing the processed utterances to voiceprints stored in memory. Khoury at ¶¶ [0230] - [0241]. Further, Khoury teaches comparing the utterances to generate a similarity score for current users (users already in the system with stored voiceprints i.e., positive exemplars corresponding to utterances of a first user) and users not currently in the system (i.e., negative exemplars corresponding to utterances not made by the first user). Khoury at ¶¶ [0230] - [0241] and Fig. 5.)
a minimum number of stored negative exemplars corresponding to utterances not made by the first user; (Khoury teaches comparing the utterances to generate a similarity score for current users (users already in the system with stored voiceprints i.e., positive exemplars corresponding to utterances of a first user) and users not currently in the system (i.e., negative exemplars corresponding to utterances not made by the first user). Khoury at ¶¶ [0230] - [0241] and Fig. 5.)
or a ratio of stored positive exemplars to stored negative exemplars. (Further, Khoury teaches the similarity scores exceeding a certain threshold (i.e., a ratio) for determining the user. Khoury at ¶¶ [0230] - [0241] and Fig. 5.)
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date to combine the teachings of Cilingir-Beaufays-Huang-Lee with the teachings of Khoury to provide the limitations of claim 9. Doing so would have improved the accuracy of the speaker recognition models as recognized by Khoury at ¶ [0060].
Regarding claim 23, Cilingir-Beaufays-Huang-Lee teaches all the limitations of claim 22 as laid out above. Cilingir-Beaufays-Huang-Lee, however, does not teach the processing system of Claim 22, further comprising, subsequent to updating the second user verification ML model, deleting the stored training exemplars.
Khoury teaches the computer-implemented method of Claim 22, the operation further comprising, subsequent to updating the second user verification ML model, deleting the stored training exemplars. (Khoury teaches training models using stored voiceprints (i.e., training exemplars). Khoury at ¶¶ [0048] – [0049]. Then Khoury discloses deleting stored voiceprints for training models (i.e., training exemplars) from a server. Khoury at ¶ [0178]. Thus, Khoury teaches deleting the voiceprints used to train models after the training has been completed.).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date to combine the teachings of Cilingir-Beaufays-Huang-Lee with the teachings of Khoury to provide subsequent to updating the second user verification ML model, deleting the stored training exemplars. Doing so would have improved the accuracy of the speaker recognition models as recognized by Khoury at ¶ [0060].
Regarding claim 27, Cilingir-Beaufays-Huang-Lee teaches all the limitations of claim 26 as laid out above. Further, Beaufays teaches the processing system of Claim 26, wherein updating the second user verification ML model based on the voice data is performed based further in response to determining that a number of stored training exemplars satisfies one or more defined criteria, wherein the one or more defined criteria indicate at least one of: (Cilingir-Beaufays teaches updating a global machine learning model using outputs from a first (or local) machine learning model. (i.e., Cilingir's speaker recognition model in view of Beaufays training of global models based on results of a local ml model is a second user verification model updated by the outputs of a first user verification model.) Beaufays at ¶¶ [0024] - [0040].)
Cilingir-Beaufays-Huang-Lee, however, does not teach the one or more defined criteria indicates at least one of. a minimum number of stored positive exemplars corresponding to utterances made by the first user; a minimum number of stored negative exemplars corresponding to utterances not made by the first user; or a ratio of stored positive exemplars to stored negative exemplars.
Khoury teaches a minimum number of stored positive exemplars corresponding to utterances made by a first user; (Khoury teaches processing a large number of utterances using a machine learning architecture. Khoury at ¶ [0233]. Further, Khoury teaches determining a similarity score for utterances by comparing the processed utterances to voiceprints stored in memory. Khoury at ¶¶ [0230] - [0241]. Further, Khoury teaches comparing the utterances to generate a similarity score for current users (users already in the system with stored voiceprints i.e., positive exemplars corresponding to utterances of a first user) and users not currently in the system (i.e., negative exemplars corresponding to utterances not made by the first user). Khoury at ¶¶ [0230] - [0241] and Fig. 5.)
a minimum number of stored negative exemplars corresponding to utterances not made by the first user; (Khoury teaches comparing the utterances to generate a similarity score for current users (users already in the system with stored voiceprints i.e., positive exemplars corresponding to utterances of a first user) and users not currently in the system (i.e., negative exemplars corresponding to utterances not made by the first user). Khoury at ¶¶ [0230] - [0241] and Fig. 5.)
or a ratio of stored positive exemplars to stored negative exemplars. (Further, Khoury teaches the similarity scores exceeding a certain threshold (i.e., a ratio) for determining the user. Khoury at ¶¶ [0230] - [0241] and Fig. 5.)
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date to combine the teachings of Cilingir-Beaufays-Huang-Lee with the teachings of Khoury to provide the limitations of claim 27. Doing so would have improved the accuracy of the speaker recognition models as recognized by Khoury at ¶ [0060].
Claims 16 is rejected under 35 U.S.C. 103 as being unpatentable over Cilingir-Huang-Lee in view of Beaufays as applied to claims 15 above, and further in view of U.S. Patent Application Publication No. 2021/0326421 A1 to Elie Khoury et al. (hereinafter Khoury).
Regarding claim 16, Cilingir-Huang-Lee-Beaufays teaches all the limitations of claim 15 as laid out above. Cilingir-Huang-Lee-Beaufays, however, does not teach the computer-implemented method of Claim 15, further comprising, subsequent to updating the second user verification ML model, deleting the stored training exemplars.
Khoury teaches the computer-implemented method of Claim 4, further comprising, subsequent to updating the second user verification ML model, deleting the stored training exemplars. (Khoury teaches training models using stored voiceprints (i.e., training exemplars). Khoury at ¶¶ [0048] – [0049]. Then Khoury discloses deleting stored voiceprints for training models (i.e., training exemplars) from a server. Khoury at ¶ [0178]. Thus, Khoury teaches deleting the voiceprints used to train models after the training has been completed.).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date to combine the teachings of Cilingir-Huang-Lee-Beaufays with the teachings of Khoury to provide subsequent to updating the second user verification ML model, deleting the stored training exemplars. Doing so would have improved the accuracy of the speaker recognition models as recognized by Khoury at ¶ [0060].
Claims 6 and 24 are rejected under 35 U.S.C. 103 as being unpatentable over Cilingir-Beaufays-Huang-Lee as applied to claims 4 and 22 above, and further in view of U.S. Patent Application Publication No. 2020/0175961 A1 to David Thomson et al. (hereinafter Thomson).
Regarding claim 6, Cilingir-Beaufays-Huang-Lee teaches all the limitations of claim 4 as laid out above. Cilingir-Beaufays-Huang-Lee, however, does not teach the computer-implemented method of Claim 4, further comprising, subsequent to updating the second user verification ML model, storing the training exemplars in a storage location that satisfies one or more defined security criteria.
Thomson teaches the computer-implemented method of Claim 4, further comprising, subsequent to updating the second user verification ML model, storing the training exemplars in a storage location that satisfies one or more defined security criteria. (Thomson teaches storing transcripts and session audio in a secure location (the transcripts and audio are used for training speech recognition systems) (i.e., training exemplars stored in a location that satisfies one or more defined security criteria). Thomson at ¶ [0300].)
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date to combine the teachings of Cilingir-Beaufays-Huang-Lee with the teachings of Thomson to provide the limitations of claim 6. Doing so would have improved the accuracy of automatic speech recognition as recognized by Thomson at ¶¶ [0096] – [0098].
Regarding claim 24, Cilingir-Beaufays-Huang-Lee teaches all the limitations of claim 22 as laid out above. Cilingir-Beaufays-Huang-Lee, however, does not teach the processing system of Claim 22, the operation further comprising, subsequent to updating the second user verification ML model, storing the training exemplars in a storage location that satisfies one or more defined security criteria.
Thomson teaches the processing system of Claim 22, the operation further comprising, subsequent to updating the second user verification ML model, storing the training exemplars in a storage location that satisfies one or more defined security criteria. (Thomson teaches storing transcripts and session audio in a secure location (the transcripts and audio are used for training speech recognition systems) (i.e., training exemplars stored in a location that satisfies one or more defined security criteria). Thomson at ¶ [0300].)
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date to combine the teachings of Cilingir-Beaufays-Huang-Lee with the teachings of Thomson to provide the limitations of claim 24. Doing so would have improved the accuracy of automatic speech recognition as recognized by Thomson at ¶¶ [0096] – [0098].
Claims 17 is rejected under 35 U.S.C. 103 as being unpatentable over Cilingir-Huang-Lee-Beaufays as applied to claims 15 above, and further in view of U.S. Patent Application Publication No. 2020/0175961 A1 to David Thomson et al. (hereinafter Thomson).
Regarding claim 17, Cilingir-Huang-Lee-Beaufays teaches all the limitations of claim 15 as laid out above. Cilingir-Huang-Lee-Beaufays, however, does not teach the computer-implemented method of Claim 15, further comprising, subsequent to updating the second user verification ML model, storing the training exemplars in a storage location that satisfies one or more defined security criteria.
Thomson teaches the computer-implemented method of Claim 15, further comprising, subsequent to updating the second user verification ML model, storing the training exemplars in a storage location that satisfies one or more defined security criteria. (Thomson teaches storing transcripts and session audio in a secure location (the transcripts and audio are used for training speech recognition systems) (i.e., training exemplars stored in a location that satisfies one or more defined security criteria). Thomson at ¶ [0300].)
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date to combine the teachings of Cilingir-Haung-Lee-Beaufays with the teachings of Thomson to provide the limitations of claim 17. Doing so would have improved the accuracy of automatic speech recognition as recognized by Thomson at ¶¶ [0096] – [0098].
Claim 31 is rejected under 35 U.S.C. 103 as being unpatentable over Cilingir-Beaufays-Huang-Lee as applied to claims 1 above, and further in view of U.S. Patent 11,507,876 to Shiun-Zu Kuo et al. (hereinafter Kuo).
Regarding claim 31, Cilingir-Beaufays-Huang-Lee teaches all the limitations of claim 1 as laid out above.
However, Cilingir-Beaufays-Huang-Lee do not specifically teach wherein the one or more defined criteria are associated with one or more thresholds that are higher than one or more thresholds used for non-training processing of the voice data.
Kuo does teach wherein the one or more defined criteria are associated with one or more thresholds that are higher than one or more thresholds used for non-training processing of the voice data (see col. 11, lines 40-65, where training confidence threshold can be set high so that high precision is achieved).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date to modify the training as taught by Cilingir-Beaufays-Haung-Lee with the high threshold as taught by Kuo in order to be able to produce a machine learning model with high precision and low recall which directly affects the accuracy of the model in prediction tasks (see Kuo col. 11, lines 40-65).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Cui (CN 110718212A) is cited to disclose determining confidence if specific words using a low complexity model followed by a high complexity model.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PARAS D SHAH whose telephone number is (571)270-1650. The examiner can normally be reached Monday-Thursday 7:30AM-2:30PM, 5PM-7PM (EST), Friday 8AM-noon (EST).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, PARAS D SHAH can be reached at 571-270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Paras D Shah/ Supervisory Patent Examiner, Art Unit 2653
08/21/2026