Prosecution Insights
Last updated: September 26, 2026
Application No. 19/070,716

METHODS AND SYSTEMS FOR TRAINING MACHINE LEARNING MODELS TO ENHANCE THE DETECTION OF FRAUDULENT AUDIO DATA

Non-Final OA §102
Filed
Mar 05, 2025
Examiner
VILLENA, MARK
Art Unit
2658
Tech Center
2600 — Communications
Assignee
Daon Tehcnology
OA Round
1 (Non-Final)
71%
Grant Probability
Favorable
1-2
OA Rounds
2y 1m
Est. Remaining
86%
With Interview

Examiner Intelligence

Grants 71% — above average
71%
Career Allowance Rate
356 granted / 500 resolved
+9.2% vs TC avg
Moderate +15% lift
Without
With
+14.8%
Interview Lift
resolved cases with interview
Typical timeline
3y 8m
Avg Prosecution
18 currently pending
Career history
513
Total Applications
across all art units

Statute-Specific Performance

§101
15.3%
-24.7% vs TC avg
§103
52.0%
+12.0% vs TC avg
§102
19.1%
-20.9% vs TC avg
§112
4.8%
-35.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 500 resolved cases

Office Action

§102
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Drawings The drawings were submitted on 03/05/2025. These drawings are reviewed and accepted by the examiner. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claim(s) 1-19 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Khoury et al. (US 20240363125 A1). Regarding claims 1 and 8, Khoury teaches: “obtaining, by an electronic device, a training data set of genuine and replay audio data files” (par. 0094; ‘With reference to FIG. 2A, in the training phase, the server feeds the training audio signals 203a into the input layers 204, where the training audio signals 203a may include any number of genuine and fraudulent audio signals, as indicated by training labels 223 associated with the training audio signals 203a.’); “partitioning the replay audio data files into data sets based on intrinsic properties of the replay audio data files” (par. 0094, feature extraction; ‘The input layers 204 may perform one or more pre-processing operations on the training audio signals 203a. The input layers 204 extract certain features from the training audio signals 203a and perform various pre-processing and/or data augmentation operations on the training audio signals 203a.’); and “training a different machine learning model for each data set using the genuine audio data files and the replay audio data files in the respective data set” (par. 0107; ‘The training audio signals 303c may be fed to the spoken content verifier 302 for training the machine-learning models of the speech recognizer 306 or content verification engine 308.’). Regarding claims 2 (dep. on claim 1), 9 (dep. on claim 8), and 14 (dep. on claim 13), Khoury further teaches: “said partitioning step comprising: determining features of each obtained replay audio data file” (par. 0126; ‘Using the features extracted from the input audio signal 403, the fakeprint extractor 408 extracts and outputs an inbound fakeprint 405 as mathematical representation of fraud artifacts in the input audio signal 403.’); and “clustering together, using a clustering technique and based on feature similarity, the obtained replay audio data files, wherein each cluster is a different one of the data sets of obtained replay audio data files” (par. 0127; ‘The scoring layers and/or the fraud classifier 410 perform a distance scoring operation that determines the distance (e.g., similarities, differences) between the inbound fakeprint 405 and a centroid or feature vector previously generated as fraud-detection cluster using the training fakeprints 405 extracted from the training audio signal 403a.’). Regarding claims 3 (dep. on claim 2), 10 (dep. on claim 9), and 15 (dep. on claim 14), Khoury further teaches: “wherein the clustering technique is an unsupervised or supervised technique” (par. 0088; ‘In some embodiments, the analytics server 102 employs supervised training to train the machine-learning models of the machine-learning architecture, where the analytics database 104 includes labels associated with the training audio signals that indicate, for example, the characteristics or features of the training audio signals.’). Regarding claims 4 (dep. on claim 1), 11 (dep. on claim 8), and 16 (dep. on claim 13), Khoury further teaches: “creating a set of measurable types of artifacts based on the intrinsic properties of the replay audio data files” (par. 0120; ‘The passive liveness detector 402 ingests input audio signals 403a-403c (generally referred to as input audio signals 403), extracts features related to or indicative of fraud artifacts and a fakeprint vector embedding 405, and executes the fraud classifier 410 or other scoring layers of the fraud classifier 410 to generate a liveness score (S.sub.P) 407.’); “determining at least one artifact pattern, wherein any combination of the types of artifacts is an artifact pattern” par. 0120; ‘The passive liveness detector 402 ingests input audio signals 403a-403c (generally referred to as input audio signals 403), extracts features related to or indicative of fraud artifacts and a fakeprint vector embedding 405, and executes the fraud classifier 410 or other scoring layers of the fraud classifier 410 to generate a liveness score (S.sub.P) 407.’); and “partitioning, based on the determined at least one artifact pattern, replay audio data files into a number of the data sets, wherein the number of data sets matches the number of artifact patterns in the at least one artifact pattern” (par. 0126; ‘Using the features extracted from the input audio signal 403, the fakeprint extractor 408 extracts and outputs an inbound fakeprint 405 as mathematical representation of fraud artifacts in the input audio signal 403.’). Regarding claims 5 (dep. on claim 1), 12 (dep. on claim 8), and 17 (dep. on claim 13), Khoury further teaches: “training a machine learning model using the genuine audio data files and the replay audio data files in one of the data sets” (par. 0078; ‘The passive liveness detector includes one or more machine-learning models trained to analyze and detect audio signals containing instances of a presentation attack, such as replayed pre-recorded human speech or instances of synthetic speech (e.g., deepfakes, TTS-generated speech).’); and “training a different machine learning model using the genuine audio data files and the replay audio data files in a different one of the data sets” (par. 0078; ‘The passive liveness detector includes one or more machine-learning models trained to analyze and detect audio signals containing instances of a presentation attack, such as replayed pre-recorded human speech or instances of synthetic speech (e.g., deepfakes, TTS-generated speech).’). Regarding claim 6 (dep. on claim 4), Khoury further teaches: “wherein the measurable artifacts are created by low quality hardware in replay devices” (par. 0077; ‘Additionally or alternatively, the fakeprint features include fraud-related artifacts of, for example, the speaker (e.g., speaker speech characteristics, speaker patterns) and/or the end-user device 114 or network (e.g., DTMF tones, background noise, codecs, packet loss).’). Regarding claim 7 (dep. on claim 4), Khoury further teaches: “wherein the replay audio data files free of any artifact pattern are created by replay devices including high quality hardware” (par. 0122, clean audio signal; ‘In the training or enrollment phase, the passive liveness detector 402 or other software component of the server may generate simulated instances of the training audio signals 403a or enrollment audio signals 403d using one or more types of data augmentation operations that manipulate the audio features or metadata of a “clean” or “genuine” training audio signal or enrollment audio signal.’). Regarding claim 13, Khoury further teaches: “obtaining a training data set of genuine and replay audio data files” (see claim 1); “partitioning the replay audio data files into data sets based on intrinsic properties of measurable artifacts in the replay audio data files” (par. 0076; ‘These operations and data structures are similar to the voiceprint embedding and the voiceprint embedding extractor used by the speaker verifier, but the fakeprint embedding and the fakeprint features represent fraud-related artifacts of speech signals that are typically present in the replayed speech or synthetic speech of presentation attacks.’; see also par. 0126); and “training a different machine learning model for each data set using the genuine audio data files and the replay audio data files in the respective data set” (see claim 1). Regarding claim 18, Khoury further teaches: “receiving, by an electronic device, audio data of a speaker” (par. 0047; ‘For instance, the operations described herein could be implemented by any call center system 110 that receives speaker audio inputs via one or more types of communications channels.’) “calculating, by one of a plurality of trained machine learning models operated by the electronic device, an artifact pattern voice replay score reflecting the likelihood that the received audio data includes an artifact pattern” (par. 0076; ‘The liveness detector takes input audio signals as input and generates and outputs a liveness score (S.sub.P), indicating a probability or likelihood that the audio signal includes a fraudulent speech signal for a presentation attack.’ ‘These operations and data structures are similar to the voiceprint embedding and the voiceprint embedding extractor used by the speaker verifier, but the fakeprint embedding and the fakeprint features represent fraud-related artifacts of speech signals that are typically present in the replayed speech or synthetic speech of presentation attacks.’); “calculating, by a different one of the plurality of trained machine learning models operated by the electronic device, a free voice replay score reflecting the likelihood that the received audio data is free of the artifact pattern” (par. 0120; ‘The passive liveness detector 402 ingests input audio signals 403a-403c (generally referred to as input audio signals 403), extracts features related to or indicative of fraud artifacts and a fakeprint vector embedding 405, and executes the fraud classifier 410 or other scoring layers of the fraud classifier 410 to generate a liveness score (S.sub.P) 407.’); “combining the artifact pattern and free voice replay scores into a replay detection score” (par. 0142; ‘In the example embodiment of FIG. 7, a server executes the various functions and features for generating the various output scores for an input audio signal 703, which may be ingested and analyzed by the combined liveness detector 702 to generate a combined liveness score 710 (S.sub.L) for the input audio signal 703.’) “comparing the replay detection score against a threshold value” (par. 0127; ‘ For instance, the scoring layers or other component of the passive liveness detector 402 determine whether the distance score or other outputted values satisfy threshold values.’) and “in response to determining the replay detection score satisfies the threshold value, determining the received audio data is of a live person” (par. 0127; ‘The passive liveness score 407 (S.sub.P) indicates the likelihood that the input audio signal 403 is fraudulent’). Regarding claim 19 (dep. on claim 18), Khoury further teaches: “the step of determining the received audio data is fraudulent in response to determining the replay detection score fails to satisfy the threshold value” (par. 0018; ‘identify the inbound audio signal as genuine or fraudulent based upon comparing the fused liveness score against an overall risk threshold.’). Conclusion Other pertinent prior art are cited in the PTO-892 for the applicant's consideration. Any inquiry concerning this communication or earlier communications from the examiner should be directed to MARK VILLENA whose telephone number is (571)270-3191. The examiner can normally be reached 10 am - 6pm EST Monday through Friday. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richemond Dorvil can be reached at (571) 272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. MARK . VILLENA Examiner Art Unit 2658 /MARK VILLENA/Examiner, Art Unit 2658
Read full office action

Prosecution Timeline

Mar 05, 2025
Application Filed
Sep 10, 2026
Non-Final Rejection mailed — §102 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12744037
SYSTEM AND METHOD FOR CHANGE POINT DETECTION IN MULTI-MEDIA MULTI-PERSON INTERACTIONS
2y 9m to grant Granted Sep 22, 2026
Patent 12743586
METHOD AND SYSTEM OF CONTEXT WINDOW ENGINEERING FOR LARGE LANGUAGE MODELS FINE-TUNED FOR CONVERSATIONS
2y 3m to grant Granted Sep 22, 2026
Patent 12731598
SPEECH RECOGNITION OF AUDIO
3y 8m to grant Granted Sep 08, 2026
Patent 12718809
Method And Device For Voice Operated Control
2y 8m to grant Granted Aug 25, 2026
Patent 12720395
ENHANCED WIRELESS COMMUNICATION HANDOVER MANAGEMENT SYSTEM
2y 6m to grant Granted Aug 25, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
71%
Grant Probability
86%
With Interview (+14.8%)
3y 8m (~2y 1m remaining)
Median Time to Grant
Low
PTA Risk
Based on 500 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month