DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-3, 6, 9, 11, 13-15, 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Feng U.S. PAP 2025/0329334 A1 in view of Li 2021/0050020 A1.
Regarding claim 1 Feng teaches a method of training an artificial intelligence (AI) neural network in authenticating a voice fingerprint, the method comprising:
extracting data from a training dataset of voice fingerprints with the AI neural network, the AI neural network comprising a residual convolutional neural network (ResCNN) (, the voiceprint vector extraction model is constructed based on improved pretrained audio neural networks (PANNs) and a transformer network. The voiceprint vector extraction model is fully trained using an open-source large-scale speaker dataset, see par. [0031]; or convolutional layer structures (for example, vgg (a convolutional network) in a residual neural network (ResNet)) of some existing networks and corresponding trained weight files may be used for training calculation, see par. [0060]);
generating processed voice fingerprint data from the extracted data based at least in part on a softmax function implemented in accordance with a softmax layer of the ResCNN (the trained voiceprint vector extraction model has a capability of fully expressing a voiceprint characteristic of an object, see par. [0031]; linear transformation (Linear) and classification (softmax) are performed on an output feature of the decoder side, to obtain a voiceprint representation vector of the specified object , see par. [0103]);
training the AI neural network based at least in part on the training dataset to produce a trained AI neural network (vector embedding mechanism allows the speech segmentation model for additional training without relying on historical speech data of any object, but only requires the extraction of a voiceprint representation vector from a small amount of reference speech data, see par. [0058]).
However Feng does not teach preparing a training dataset and a validation dataset of voice fingerprints based at least in part on the processed voice fingerprint data; and validating the trained AI neural network based at least in part on the validation dataset.
In the same field of endeavor Li teaches a voiceprint recognition method performed by a computer, see abstract. Li teaches to verify an application effect of the voiceprint recognition method provided by the embodiments of this application, verification and comparison are performed on a large data set. The data set is divided into three parts: a training set, a validation set, and a test set, see par. [0145]. To ensure normal training of the network, a 120-dimensional log-mel feature is used as an input of the network. FIG. 9 is a schematic diagram of comparison between accuracy rates of a validation set applied to different networks according to an embodiment of this application. As shown in the figure, using two loss functions for optimizing network training at the same time is better than individually using the softmax loss function, and a network structure of this embodiment of this application may achieve the highest accuracy rate on the verification set when an inputted feature has a smaller dimension and a minimum duration of inputted voice is the shortest, see par. [00146].
It would have been obvious to one of ordinary skill in the art to combine the Feng invention with the teachings of Li for the benefit of achieving the highest accuracy rate on the verification set, see par. [0146].
Regarding claim 2 Feng teaches the method of claim 1, wherein extracting data from the training dataset comprises extracting, from a hidden layer of the ResCNN, speaker embeddings associated with the voice fingerprints encoded in a d-vector representation (obtain a voiceprint representation vector configured for representing an identity of the specified object, and then embed the voiceprint representation vector into a speech segmentation model in the system, see par. [0038]).
Regarding claim 3 Feng teaches the method of claim 1, further comprising pre-processing the training dataset based on one or more of: labeling the voice fingerprints with categories and identities of speakers, removing outliers, and one of inputting and providing missing values associated with at least a subset of the voice fingerprints (identify an identity of a specified object for segmentation, and extract an identity semantic vector of the specified object. The identity semantic vector herein is referred to as a voiceprint representation vector, see par. [0030]).
Regarding claim 6 Feng teaches the method of claim 1, wherein validating the AI neural network further comprises determining an accuracy of the AI neural network in accordance with determining a cumulative accuracy profile of the AI neural network (thereby implementing integration of features in different scales and improving the result accuracy of the model, see par. [0062]).
Regarding claim 9 Feng teaches the method of claim 1 wherein training the AI neural network produces a trained AI neural network, and further comprising deploying the trained AI neural network in a remote examination session (The terminal is deployed with a system shown in FIG. 1, an application (or a plug-in) providing the system, or the like. The server may be an independent physical server, or may be a server cluster or a distributed system formed by a plurality of physical servers, see par. [0037]), the deploying comprising:
receiving a voice sample purportedly associated with an examination candidate, the examination candidate being remotely located relative to an examination proctoring computing device (When the capture entry 4021 is triggered, it indicates that the user intends to input reference speech data of a specified object (where the specified object is the user, or the specified object and the user are in the same physical environment) through real-time capture, see par. [0053]);
generating a voice fingerprint in accordance with the voice sample (a reference speech signal in the physical environment in which the user is located can be captured in real time, to generate the reference speech data, see par. [0053]);
and authenticating the examination candidate based at least in part on the generated voice fingerprint (Extract a voiceprint representation vector of the specified object from the reference speech data, and input the overlapping speech data and the voiceprint representation vector into a preset speech segmentation model, the speech segmentation model being configured to: segment, based on an attention mechanism, the overlapping speech data to obtain a speech signal matching a voiceprint characteristic of the specified object, see par. [0055]).
Regarding claim 11 Feng teaches an examination proctoring computer system comprising: one or more processors; and a memory storing instructions executable in the one or more processors, the instructions when executed causing the one or more processors to implement operations (The computer device includes: a processor, configured to load and execute a computer program; and a non-transitory computer-readable storage medium, having the computer program stored therein, the computer program, when executed by the processor, implementing the foregoing speech processing method., see par. [0008]) comprising:
receiving, from a candidate computing device that is interconnected with the examination proctoring computing system within a distributed network computing system, a voice sample purportedly associated with an examination candidate remotely located relative to the examination proctoring computing device (When the capture entry 4021 is triggered, it indicates that the user intends to input reference speech data of a specified object (where the specified object is the user, or the specified object and the user are in the same physical environment) through real-time capture, see par. [0053]); generating a voice fingerprint in accordance with the voice sample (a reference speech signal in the physical environment in which the user is located can be captured in real time, to generate the reference speech data, see par. [0053]); and authenticating the examination candidate, in accordance with a trained AI neural network, based at least in part on the generated voice fingerprint (Extract a voiceprint representation vector of the specified object from the reference speech data, and input the overlapping speech data and the voiceprint representation vector into a preset speech segmentation model, the speech segmentation model being configured to: segment, based on an attention mechanism, the overlapping speech data to obtain a speech signal matching a voiceprint characteristic of the specified object, see par. [0055]).
Regarding claim 13 Feng teaches a computer-readable non-transitory memory having instructions stored thereon, the instructions being executable to cause one or more processors to implement operations (a non-transitory computer-readable storage medium, having a computer program stored therein. The computer program is configured to be loaded and executed by a processor and perform the foregoing speech processing method, see par. [0009]) comprising:
extracting data from a training dataset of voice fingerprints with the AI neural network, the AI neural network comprising a residual convolutional neural network (ResCNN) (, the voiceprint vector extraction model is constructed based on improved pretrained audio neural networks (PANNs) and a transformer network. The voiceprint vector extraction model is fully trained using an open-source large-scale speaker dataset, see par. [0031]; or convolutional layer structures (for example, vgg (a convolutional network) in a residual neural network (ResNet)) of some existing networks and corresponding trained weight files may be used for training calculation, see par. [0060]);
generating processed voice fingerprint data from the extracted data based at least in part on a softmax function implemented in accordance with a softmax layer of the ResCNN (the trained voiceprint vector extraction model has a capability of fully expressing a voiceprint characteristic of an object, see par. [0031]; linear transformation (Linear) and classification (softmax) are performed on an output feature of the decoder side, to obtain a voiceprint representation vector of the specified object , see par. [0103]);
training the AI neural network based at least in part on the training dataset to produce a trained AI neural network (vector embedding mechanism allows the speech segmentation model for additional training without relying on historical speech data of any object, but only requires the extraction of a voiceprint representation vector from a small amount of reference speech data, see par. [0058]).
However Feng does not teach preparing a training dataset and a validation dataset of voice fingerprints based at least in part on the processed voice fingerprint data; and validating the trained AI neural network based at least in part on the validation dataset.
In the same field of endeavor Li teaches a voiceprint recognition method performed by a computer, see abstract. Li teaches to verify an application effect of the voiceprint recognition method provided by the embodiments of this application, verification and comparison are performed on a large data set. The data set is divided into three parts: a training set, a validation set, and a test set, see par. [0145]. To ensure normal training of the network, a 120-dimensional log-mel feature is used as an input of the network. FIG. 9 is a schematic diagram of comparison between accuracy rates of a validation set applied to different networks according to an embodiment of this application. As shown in the figure, using two loss functions for optimizing network training at the same time is better than individually using the softmax loss function, and a network structure of this embodiment of this application may achieve the highest accuracy rate on the verification set when an inputted feature has a smaller dimension and a minimum duration of inputted voice is the shortest, see par. [00146].
It would have been obvious to one of ordinary skill in the art to combine the Feng invention with the teachings of Li for the benefit of achieving the highest accuracy rate on the verification set, see par. [0146].
Regarding claim 14 Feng teaches the computer-readable non-transitory memory of claim 13, the instructions being executable in the one or more processors to cause operations comprising extracting data from the training dataset comprises extracting, from a hidden layer of the ResCNN, speaker embeddings associated the voice fingerprints encoded in a d-vector representation (obtain a voiceprint representation vector configured for representing an identity of the specified object, and then embed the voiceprint representation vector into a speech segmentation model in the system, see par. [0038]).
Regarding claim 15 Feng teaches the computer-readable non-transitory memory of claim 13, the instructions being executable in the one or more processors to cause operations comprising pre-processing the training dataset based on one or more of: labeling the voice fingerprints with categories and identities of speakers, removing outliers, and one of inputting and providing missing values associated with at least a subset of the voice fingerprints (identify an identity of a specified object for segmentation, and extract an identity semantic vector of the specified object. The identity semantic vector herein is referred to as a voiceprint representation vector, see par. [0030]).
Regarding claim 18 Feng teaches the computer-readable non-transitory memory of claim 13, wherein the instructions cause the one or more processors to implement operations comprising validating the AI neural network further comprises determining an accuracy of the AI neural network in accordance with determining a cumulative accuracy profile of the AI neural network (thereby implementing integration of features in different scales and improving the result accuracy of the model, see par. [0062])..
Claim(s) 4, 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Feng U.S. PAP 2025/0329334 A1 in view of Li 2021/0050020 A1 further in view of Fatemi U.S. PAP 2023/0104662 A1.
Regarding claim 4 Feng in view of Li does not teach the method of claim 3 wherein the training dataset is subjected to removal of gender bias based at least in part on the softmax function.
IN a similar field of endeavor Fatemi teaches a training framework for reducing gender bias in a pre-trained language model, see abstract. In some embodiments the self-attention layer produces output by calculating a softmax function where Q is a query, K is a key, and V is a value, and Q, K, V are the hidden representations outputted from the previous layer of the BERT model to subsequent layer, and d is a dimension of the hidden representations, see par. [0028]. Compared with the base BERT, the scores generated by the language model 136 are significantly closer to zero, indicating that the language model 136 is more effective at removing gender bias from such biased professions compared with the base BERT, see par. [0052].
It would have been obvious to one of ordinary skill in the art top combine the Feng in view of Li invention with the teachings of Fatemi for the benefit of effectively removing gender bias from biased professions, see par. [0052].
Regarding claim 16 Feng in view of Li does not teach the computer-readable non-transitory memory of claim 15, the instructions being executable in the one or more processors to cause operations comprising the training dataset is subjected to removal of gender bias based at least in part on the softmax function.
IN a similar field of endeavor Fatemi teaches a training framework for reducing gender bias in a pre-trained language model, see abstract. In some embodiments the self-attention layer produces output by calculating a softmax function where Q is a query, K is a key, and V is a value, and Q, K, V are the hidden representations outputted from the previous layer of the BERT model to subsequent layer, and d is a dimension of the hidden representations, see par. [0028]. Compared with the base BERT, the scores generated by the language model 136 are significantly closer to zero, indicating that the language model 136 is more effective at removing gender bias from such biased professions compared with the base BERT, see par. [0052].
It would have been obvious to one of ordinary skill in the art top combine the Feng in view of Li invention with the teachings of Fatemi for the benefit of effectively removing gender bias from biased professions, see par. [0052].
Claim(s) 5, 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Feng U.S. PAP 2025/0329334 A1 in view of Li 2021/0050020 A1 Further in view of Novoselov “Triplet Loss Based Cosine Similarity Metric Learning for Text-Independent Speaker Recognition”.
Regarding claim 5 Feng in view of Li does not teach the method of claim 1, wherein training the AI neural network further comprises training the AI neural network in accordance with a triplet loss function that minimizes the distance between embedding pairs from a same speaker and maximizes the distance between embedding pairs from different speakers in the training dataset of voice fingerprints.
IN a similar field of endeavor Novoselov teaches a DNN speaker embedding extractor is usually trained discriminatively in the closed set classification scenario using softmax. The problem we addressed in the paper is choosing a dnn based speaker embedding backend solution for the speaker verification scoring. There are several options to perform speaker verification in the dnn embedding space. One of them is using a simple heuristic speaker similarity metric for scoring (e.g. co sine metric). Similarly with i-vector based systems, the standard Linear Discriminant Analisys (LDA) followed by the Probabilistic Linear Discriminant Analisys (PLDA) can be used for segregating speaker information. As an alternative, the discriminative metric learning approach can be considered. This work demonstrates that performance of deep speaker embeddings-based systems can be improved by using Cosine Similarity Metric Learning (CSML) with the triplet loss training scheme. Results obtained on Speakers in the Wild and NIST SRE 2016 evaluation sets demonstrate superiority and robust ness of CSML based systems, see abstract. The performance of deep speaker embeddings based systems can be improved by using CSML with the triplet loss training scheme in both” clean” and ”in the-wild” conditions, see conclusion.
It would have been obvious to one of ordinary skill in the art to combine the Feng in view of Li invention with the teachings of Novoselov for the benefit of improving the performance of deep speaker embeddings, see conclusion.
Regarding claim 17 Feng in view of Li does not teach the computer-readable non-transitory memory of claim 13, wherein the instructions cause the one or more processors to implement operations comprising training the AI neural network in accordance with a triplet loss function that minimizes the distance between embedding pairs from a same speaker and maximizes the distance between embedding pairs from different speakers in the training dataset of voice fingerprints.
IN a similar field of endeavor Novoselov teaches a DNN speaker embedding extractor is usually trained discriminatively in the closed set classification scenario using softmax. The problem we addressed in the paper is choosing a dnn based speaker embedding backend solution for the speaker verification scoring. There are several options to perform speaker verification in the dnn embedding space. One of them is using a simple heuristic speaker similarity metric for scoring (e.g. co sine metric). Similarly with i-vector based systems, the standard Linear Discriminant Analisys (LDA) followed by the Probabilistic Linear Discriminant Analisys (PLDA) can be used for segregating speaker information. As an alternative, the discriminative metric learning approach can be considered. This work demonstrates that performance of deep speaker embeddings-based systems can be improved by using Cosine Similarity Metric Learning (CSML) with the triplet loss training scheme. Results obtained on Speakers in the Wild and NIST SRE 2016 evaluation sets demonstrate superiority and robust ness of CSML based systems, see abstract. The performance of deep speaker embeddings based systems can be improved by using CSML with the triplet loss training scheme in both” clean” and ”in the-wild” conditions, see conclusion.
It would have been obvious to one of ordinary skill in the art to combine the Feng in view of Li invention with the teachings of Novoselov for the benefit of improving the performance of deep speaker embeddings, see conclusion.
Claim(s) 8, 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Feng U.S. PAP 2025/0329334 A1 in view of Li 2021/0050020 A1 further in view of Wang U.S. PAP 2020/0258527 A1.
Regarding claim 8 Feng in view of Li does not teach the method of claim 1, wherein the AI neural network further comprises a gated recurrent unit (GRU).
IN the same field of endeavor Wang teaches a speaker recognition system includes a non-transitory computer readable medium configured to store instructions. The speaker recognition system further includes a processor connected to the non-transitory computer readable medium. The processor is configured to execute the instructions for extracting acoustic features from each frame of a plurality of frames in input speech data. The processor is configured to execute the instructions for calculating a saliency value for each frame of the plurality of frames using a first neural network (NN) based on the extracted acoustic features, wherein the first NN is a trained NN using speaker posteriors, see abstract. Any type of neural network is usable for the speaker-discriminative NN trainer 104, e.g., a Time-Delay Neural Network (TDNN), Convolutional Neural Network (CNN), LSTM, or Gated Recurrent Unit (GRU). In operation B04, speaker-discriminative NN trainer 104 trains a speaker-discriminative NN. The speaker-feature discriminative NN trainer 104 trains the speaker-discriminative NN by determining parameters for nodes with the speaker-discriminative NN based on the read speaker IDs as well as the extracted acoustic features from the speech data. In at least one embodiment, the speaker-discriminative NN is a TDNN, a CNN, an LSTM, a GRU, or another suitable NN, see par. [0040].
It would have been obvious to one of ordinary skill in the art to combine the Feng in view of Li invention with the teachings of Wang for the benefit of outputting a speaker identity in a speaker identification scheme or genuine/imposter results in a speaker verification scheme, see par. [0001].
Regarding claim 20 Feng in view of Li do not teach the computer-readable non-transitory memory of claim 13, wherein the AI neural network further comprises a gated recurrent unit (GRU).
IN the same field of endeavor Wang teaches a speaker recognition system includes a non-transitory computer readable medium configured to store instructions. The speaker recognition system further includes a processor connected to the non-transitory computer readable medium. The processor is configured to execute the instructions for extracting acoustic features from each frame of a plurality of frames in input speech data. The processor is configured to execute the instructions for calculating a saliency value for each frame of the plurality of frames using a first neural network (NN) based on the extracted acoustic features, wherein the first NN is a trained NN using speaker posteriors, see abstract. Any type of neural network is usable for the speaker-discriminative NN trainer 104, e.g., a Time-Delay Neural Network (TDNN), Convolutional Neural Network (CNN), LSTM, or Gated Recurrent Unit (GRU). In operation B04, speaker-discriminative NN trainer 104 trains a speaker-discriminative NN. The speaker-feature discriminative NN trainer 104 trains the speaker-discriminative NN by determining parameters for nodes with the speaker-discriminative NN based on the read speaker IDs as well as the extracted acoustic features from the speech data. In at least one embodiment, the speaker-discriminative NN is a TDNN, a CNN, an LSTM, a GRU, or another suitable NN, see par. [0040].
It would have been obvious to one of ordinary skill in the art to combine the Feng in view of Li invention with the teachings of Wang for the benefit of outputting a speaker identity in a speaker identification scheme or genuine/imposter results in a speaker verification scheme, see par. [0001].
Claim(s) 7, 10, 12 and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Feng U.S. PAP 2025/0329334 A1 in view of Li 2021/0050020 A1 further in view of Khoury U.S. PAP 2024/036125 A1 .
Regarding claim 7 Feng in view of Li do not teach the method of claim 1, wherein validating the AI neural network further comprises determining an accuracy of the AI neural network in accordance with an equal error rate algorithm.
In the same field of endeavor Khoury teaches systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. Embodiments include systems and methods for detecting fraudulent presentation attacks using multiple functional engines that implement various fraud-detection techniques, to produce calibrated scores and/or fused scores. A computer may, for example, evaluate the audio quality of speech signals within audio signals, where speech signals contain the speech portions having speaker utterances, see abstract. Khoury teaches improved voice biometric systems to check whether the voice received is from a live person who is speaking into a microphone. This is called “voice liveness detection.”, see par. [0007]. At the backend of the machine-learning architecture, a content verification engine of the content verifier includes one or more machine-learning models trained classify and/or score the spoken response content in the input speech signal. The content verification engine ingests the challenge content representation and the response content representation and generates a content verification score (S.sub.c). The content verifier may determine the content verification by computing or executing various techniques. Non-limiting examples of the operations for determining the content verification score may include: computing an error rate, see par. [0074].
It would have been obvious to one of ordinary skill in the art to combine the Feng in view of Li invention with the teachings of Khoury for the benefit of improved voice biometric systems to check whether the voice received is from a live person who is speaking into a microphone. This is called “voice liveness detection.”, see par. [0007].
Regarding claim 10 Feng in view of Li do not teach the method of claim 9 further comprising authenticating the examination candidate based at least in part on a liveness detection measurement and a threshold percentage match with a pre-existing registration voice sample associated with the examination candidate.
In the same field of endeavor Khoury teaches ystems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. Embodiments include systems and methods for detecting fraudulent presentation attacks using multiple functional engines that implement various fraud-detection techniques, to produce calibrated scores and/or fused scores. A computer may, for example, evaluate the audio quality of speech signals within audio signals, where speech signals contain the speech portions having speaker utterances, see abstract. Khoury teaches improved voice biometric systems to check whether the voice received is from a live person who is speaking into a microphone. This is called “voice liveness detection.”, see par. [0007]. generating, by the computer, a fused liveness score based upon the content verification score and the passive liveness score; and identifying, by the computer, the inbound audio signal as genuine or fraudulent based upon comparing the fused liveness score against an overall risk threshold, see par. [0017].
It would have been obvious to one of ordinary skill in the art to combine the Feng in view of Li invention with the teachings of Khoury for the benefit of improved voice biometric systems to check whether the voice received is from a live person who is speaking into a microphone. This is called “voice liveness detection.”, see par. [0007].
Regarding claim 12 Feng in view of Li do not teach the examination proctoring computing system of claim 11 wherein the instructions further cause the one or more processors to implement operations comprising authenticating the examination candidate based at least in part on a liveness detection measurement and a threshold percentage match with a pre-existing registration voice sample associated with the examination candidate.
In the same field of endeavor Khoury teaches ystems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. Embodiments include systems and methods for detecting fraudulent presentation attacks using multiple functional engines that implement various fraud-detection techniques, to produce calibrated scores and/or fused scores. A computer may, for example, evaluate the audio quality of speech signals within audio signals, where speech signals contain the speech portions having speaker utterances, see abstract. Khoury teaches improved voice biometric systems to check whether the voice received is from a live person who is speaking into a microphone. This is called “voice liveness detection.”, see par. [0007]. generating, by the computer, a fused liveness score based upon the content verification score and the passive liveness score; and identifying, by the computer, the inbound audio signal as genuine or fraudulent based upon comparing the fused liveness score against an overall risk threshold, see par. [0017].
It would have been obvious to one of ordinary skill in the art to combine the Feng in view of Li invention with the teachings of Khoury for the benefit of improved voice biometric systems to check whether the voice received is from a live person who is speaking into a microphone. This is called “voice liveness detection.”, see par. [0007].
Regarding claim 19 Feng in view of Li do not teach the computer-readable non-transitory memory of claim 13, wherein the instructions cause the one or more processors to implement operations comprising validating the AI neural network further comprises determining an accuracy of the AI neural network in accordance with an equal error rate algorithm.
In the same field of endeavor Khoury teaches systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. Embodiments include systems and methods for detecting fraudulent presentation attacks using multiple functional engines that implement various fraud-detection techniques, to produce calibrated scores and/or fused scores. A computer may, for example, evaluate the audio quality of speech signals within audio signals, where speech signals contain the speech portions having speaker utterances, see abstract. Khoury teaches improved voice biometric systems to check whether the voice received is from a live person who is speaking into a microphone. This is called “voice liveness detection.”, see par. [0007]. At the backend of the machine-learning architecture, a content verification engine of the content verifier includes one or more machine-learning models trained classify and/or score the spoken response content in the input speech signal. The content verification engine ingests the challenge content representation and the response content representation and generates a content verification score (S.sub.c). The content verifier may determine the content verification by computing or executing various techniques. Non-limiting examples of the operations for determining the content verification score may include: computing an error rate, see par. [0074].
It would have been obvious to one of ordinary skill in the art to combine the Feng in view of Li invention with the teachings of Khoury for the benefit of improved voice biometric systems to check whether the voice received is from a live person who is speaking into a microphone. This is called “voice liveness detection.”, see par. [0007].
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Chen ‘541 teaches systems and methods for implementing a neural network architecture for spoof detection in audio signals. The neural network architecture contains a layers defining embedding extractors that extract embeddings from input audio signals. Spoofprint embeddings are generated for particular system enrollees to detect attempts to spoof the enrollee's voice, see abstract.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Michael Ortiz-Sanchez whose telephone number is (571)270-3711. The examiner can normally be reached Monday- Friday 9AM-6PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bhavesh Mehta can be reached at 571-272-7453. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MICHAEL ORTIZ-SANCHEZ/Primary Examiner, Art Unit 2656