DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Drawings
The drawings were submitted on 03/05/2025. These drawings are reviewed and accepted by the examiner.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim(s) 1-19 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Khoury et al. (US 20240363125 A1).
Regarding claims 1 and 8, Khoury teaches:
“obtaining, by an electronic device, a training data set of genuine and replay audio data files” (par. 0094; ‘With reference to FIG. 2A, in the training phase, the server feeds the training audio signals 203a into the input layers 204, where the training audio signals 203a may include any number of genuine and fraudulent audio signals, as indicated by training labels 223 associated with the training audio signals 203a.’);
“partitioning the replay audio data files into data sets based on intrinsic properties of the replay audio data files” (par. 0094, feature extraction; ‘The input layers 204 may perform one or more pre-processing operations on the training audio signals 203a. The input layers 204 extract certain features from the training audio signals 203a and perform various pre-processing and/or data augmentation operations on the training audio signals 203a.’); and
“training a different machine learning model for each data set using the genuine audio data files and the replay audio data files in the respective data set” (par. 0107; ‘The training audio signals 303c may be fed to the spoken content verifier 302 for training the machine-learning models of the speech recognizer 306 or content verification engine 308.’).
Regarding claims 2 (dep. on claim 1), 9 (dep. on claim 8), and 14 (dep. on claim 13), Khoury further teaches:
“said partitioning step comprising: determining features of each obtained replay audio data file” (par. 0126; ‘Using the features extracted from the input audio signal 403, the fakeprint extractor 408 extracts and outputs an inbound fakeprint 405 as mathematical representation of fraud artifacts in the input audio signal 403.’); and
“clustering together, using a clustering technique and based on feature similarity, the obtained replay audio data files, wherein each cluster is a different one of the data sets of obtained replay audio data files” (par. 0127; ‘The scoring layers and/or the fraud classifier 410 perform a distance scoring operation that determines the distance (e.g., similarities, differences) between the inbound fakeprint 405 and a centroid or feature vector previously generated as fraud-detection cluster using the training fakeprints 405 extracted from the training audio signal 403a.’).
Regarding claims 3 (dep. on claim 2), 10 (dep. on claim 9), and 15 (dep. on claim 14), Khoury further teaches:
“wherein the clustering technique is an unsupervised or supervised technique” (par. 0088; ‘In some embodiments, the analytics server 102 employs supervised training to train the machine-learning models of the machine-learning architecture, where the analytics database 104 includes labels associated with the training audio signals that indicate, for example, the characteristics or features of the training audio signals.’).
Regarding claims 4 (dep. on claim 1), 11 (dep. on claim 8), and 16 (dep. on claim 13), Khoury further teaches:
“creating a set of measurable types of artifacts based on the intrinsic properties of the replay audio data files” (par. 0120; ‘The passive liveness detector 402 ingests input audio signals 403a-403c (generally referred to as input audio signals 403), extracts features related to or indicative of fraud artifacts and a fakeprint vector embedding 405, and executes the fraud classifier 410 or other scoring layers of the fraud classifier 410 to generate a liveness score (S.sub.P) 407.’);
“determining at least one artifact pattern, wherein any combination of the types of artifacts is an artifact pattern” par. 0120; ‘The passive liveness detector 402 ingests input audio signals 403a-403c (generally referred to as input audio signals 403), extracts features related to or indicative of fraud artifacts and a fakeprint vector embedding 405, and executes the fraud classifier 410 or other scoring layers of the fraud classifier 410 to generate a liveness score (S.sub.P) 407.’); and
“partitioning, based on the determined at least one artifact pattern, replay audio data files into a number of the data sets, wherein the number of data sets matches the number of artifact patterns in the at least one artifact pattern” (par. 0126; ‘Using the features extracted from the input audio signal 403, the fakeprint extractor 408 extracts and outputs an inbound fakeprint 405 as mathematical representation of fraud artifacts in the input audio signal 403.’).
Regarding claims 5 (dep. on claim 1), 12 (dep. on claim 8), and 17 (dep. on claim 13), Khoury further teaches:
“training a machine learning model using the genuine audio data files and the replay audio data files in one of the data sets” (par. 0078; ‘The passive liveness detector includes one or more machine-learning models trained to analyze and detect audio signals containing instances of a presentation attack, such as replayed pre-recorded human speech or instances of synthetic speech (e.g., deepfakes, TTS-generated speech).’); and
“training a different machine learning model using the genuine audio data files and the replay audio data files in a different one of the data sets” (par. 0078; ‘The passive liveness detector includes one or more machine-learning models trained to analyze and detect audio signals containing instances of a presentation attack, such as replayed pre-recorded human speech or instances of synthetic speech (e.g., deepfakes, TTS-generated speech).’).
Regarding claim 6 (dep. on claim 4), Khoury further teaches:
“wherein the measurable artifacts are created by low quality hardware in replay devices” (par. 0077; ‘Additionally or alternatively, the fakeprint features include fraud-related artifacts of, for example, the speaker (e.g., speaker speech characteristics, speaker patterns) and/or the end-user device 114 or network (e.g., DTMF tones, background noise, codecs, packet loss).’).
Regarding claim 7 (dep. on claim 4), Khoury further teaches:
“wherein the replay audio data files free of any artifact pattern are created by replay devices including high quality hardware” (par. 0122, clean audio signal; ‘In the training or enrollment phase, the passive liveness detector 402 or other software component of the server may generate simulated instances of the training audio signals 403a or enrollment audio signals 403d using one or more types of data augmentation operations that manipulate the audio features or metadata of a “clean” or “genuine” training audio signal or enrollment audio signal.’).
Regarding claim 13, Khoury further teaches:
“obtaining a training data set of genuine and replay audio data files” (see claim 1);
“partitioning the replay audio data files into data sets based on intrinsic properties of measurable artifacts in the replay audio data files” (par. 0076; ‘These operations and data structures are similar to the voiceprint embedding and the voiceprint embedding extractor used by the speaker verifier, but the fakeprint embedding and the fakeprint features represent fraud-related artifacts of speech signals that are typically present in the replayed speech or synthetic speech of presentation attacks.’; see also par. 0126); and
“training a different machine learning model for each data set using the genuine audio data files and the replay audio data files in the respective data set” (see claim 1).
Regarding claim 18, Khoury further teaches:
“receiving, by an electronic device, audio data of a speaker” (par. 0047; ‘For instance, the operations described herein could be implemented by any call center system 110 that receives speaker audio inputs via one or more types of communications channels.’)
“calculating, by one of a plurality of trained machine learning models operated by the electronic device, an artifact pattern voice replay score reflecting the likelihood that the received audio data includes an artifact pattern” (par. 0076; ‘The liveness detector takes input audio signals as input and generates and outputs a liveness score (S.sub.P), indicating a probability or likelihood that the audio signal includes a fraudulent speech signal for a presentation attack.’ ‘These operations and data structures are similar to the voiceprint embedding and the voiceprint embedding extractor used by the speaker verifier, but the fakeprint embedding and the fakeprint features represent fraud-related artifacts of speech signals that are typically present in the replayed speech or synthetic speech of presentation attacks.’);
“calculating, by a different one of the plurality of trained machine learning models operated by the electronic device, a free voice replay score reflecting the likelihood that the received audio data is free of the artifact pattern” (par. 0120; ‘The passive liveness detector 402 ingests input audio signals 403a-403c (generally referred to as input audio signals 403), extracts features related to or indicative of fraud artifacts and a fakeprint vector embedding 405, and executes the fraud classifier 410 or other scoring layers of the fraud classifier 410 to generate a liveness score (S.sub.P) 407.’);
“combining the artifact pattern and free voice replay scores into a replay detection score” (par. 0142; ‘In the example embodiment of FIG. 7, a server executes the various functions and features for generating the various output scores for an input audio signal 703, which may be ingested and analyzed by the combined liveness detector 702 to generate a combined liveness score 710 (S.sub.L) for the input audio signal 703.’)
“comparing the replay detection score against a threshold value” (par. 0127; ‘ For instance, the scoring layers or other component of the passive liveness detector 402 determine whether the distance score or other outputted values satisfy threshold values.’) and
“in response to determining the replay detection score satisfies the threshold value, determining the received audio data is of a live person” (par. 0127; ‘The passive liveness score 407 (S.sub.P) indicates the likelihood that the input audio signal 403 is fraudulent’).
Regarding claim 19 (dep. on claim 18), Khoury further teaches:
“the step of determining the received audio data is fraudulent in response to determining the replay detection score fails to satisfy the threshold value” (par. 0018; ‘identify the inbound audio signal as genuine or fraudulent based upon comparing the fused liveness score against an overall risk threshold.’).
Conclusion
Other pertinent prior art are cited in the PTO-892 for the applicant's consideration.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MARK VILLENA whose telephone number is (571)270-3191. The examiner can normally be reached 10 am - 6pm EST Monday through Friday.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richemond Dorvil can be reached at (571) 272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
MARK . VILLENA
Examiner
Art Unit 2658
/MARK VILLENA/Examiner, Art Unit 2658