DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Applicant’s claim for the benefit of a prior-filed application under 35 U.S.C. 119(e) or under 35 U.S.C. 120, 121, 365(c), or 386(c) is acknowledged. Accordingly, the claims under examination are granted the benefit of the earlier filing date of PRO 63/551023, filed February 7th, 2024.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claims recite a mental process that can be performed in the human mind or with the aid of pen and paper. This judicial exception is not integrated into a practical application because a computer is invoked merely as a tool to execute an abstract idea. The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because an abstract idea is merely applied on a generic computer without any element that would otherwise preclude performance of the abstrac.
Regarding claim 1, the claim recites “A method for training a machine learning model to detect a fraudulent attempt to mimic an employee using a synthetic voice copy, the method comprising:receiving a voice sample of an employee of an enterprise;normalizing the voice sample;generating a synthetic voice sample using the normalized voice sample;training the machine learning model to identify whether received audio includes a synthetically generated voice sample using the voice sample and the synthetic voice sample; andoutputting the trained machine learning model.”
The operable limitation of “identify whether received audio includes a synthetically generated voice sample” as drafted cover mental activities which can be performed in the mind or with the aid of pen and paper. The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the steps of the claimed invention can be performed mentally, and the additional limitations of “normalizing the voice sample,” “training the machine learning model,” and “outputting the trained machine learning model” describe method steps well-understood in the art. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible.
Regarding claim 2, the claim depends from claim 1, and thus recites the limitations of claim 1, “wherein the voice sample is positively weighted for training the machine learning model.”
Weighting of machine learning model training data is well-understood and readily available to a person having ordinary skill in the art of machine learning. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible.
Regarding claim 3, the claim depends from claim 1, and thus recites the limitations of claim 1, “wherein the synthetic voice sample is negatively weighted for training the machine learning model.”
Weighting of machine learning model training data is well-understood and readily available to a person having ordinary skill in the art of machine learning. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible.
Regarding claim 4, the claim depends from claim 1, and thus recites the limitations of claim 1, “further comprising receiving a second voice sample from a person other than the employee, and determining whether the second voice sample is a match using the trained machine learning model.”
The limitation of “determining whether the second voice sample is a match” as drafted covers mental activities which can be performed in the mind or with the aid of pen and paper. Taken individually, or as a whole with claim 1, these limitations describe acts which are equivalent to human mental work of detecting voice fakes. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible.
Regarding claim 5, the claim depends from claim 1, and thus recites the limitations of claim 1, “further comprising testing a second voice sample from the employee, and determining whether the second voice sample is a match using the trained machine learning model.”
The limitation of “determining whether the second voice sample is a match” as drafted covers mental activities which can be performed in the mind or with the aid of pen and paper. Taken individually, or as a whole with claim 1, these limitations describe acts which are equivalent to human mental work of detecting voice fakes. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible.
Regarding claim 6, the claim depends from claim 1, and thus recites the limitations of claim 1, “wherein the voice sample includes a set of voice samples including at least two voice samples of differing duration.”
The limitation describes pre-solution activity that is well known and imposes no meaningful limitation on the method of the independent claim, as in data gathering or data selection. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible.
Regarding claim 7, the claim depends from claim 1, and thus recites the limitations of claim 1, “wherein the voice sample includes a low quality verification test sample with background noise.”
The limitation describes pre-solution activity that is well known and imposes no meaningful limitation on the method of the independent claim, as in data gathering or data selection. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible.
Regarding claims 8-14, computer-readable medium claims 8-14 and method claims 1-7 are related as method and computer-readable medium for performing the same, with each computer-readable medium element’s function corresponding to the method step. Accordingly, claims 8-14 are similarly rejected under the same rationale as applied to claims 1-7.
Regarding claims 15-20, system claims 15-20 and method claims 1-6 are related as a method and system of using the same, with each system element’s function corresponding to the method step. Accordingly, claims 15-20 are similarly rejected under the same rationale as applied to claims 1-6.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claim 1-20 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by U.S. Patent 12,288,562 to Havdan et al. (hereinafter, "Havdan").
Regarding claims 1, 8 and 15, Havdan teaches a method, system and computer-readable medium that includes receiving a voice sample of an employee of an enterprise (column 7, line 24, "DL NN model 300 may obtain a training dataset including labeled samples. A labeled sample may include sample data 380 and a label 370. Label 370 of a sample may indicate or describe whether the sample pertains to a first class or a second class.");
normalizing the voice sample (column 7, line 37, "According to embodiments of the invention, loss function 360 may include a regulation factor, where the regulation factor is set to zero for data samples labeled as pertaining or belonging to the first class and is proportional to the prediction 350 of DL NN model 300 for data samples labeled as pertaining to the second class.");
generating a synthetic voice sample using the normalized voice sample (column 6, line 44, "Using embodiments of the invention, a voiced sample with as little as three seconds of voice, or other durations, may be sufficient for detecting spoofing calls. Silence removal block 200 may generate net-speech samples of three seconds or longer from the voice sample."));
training the machine learning model to identify whether received audio includes a synthetically generated voice sample using the voice sample and the synthetic voice sample (column 7, line 11, "Reference is now made to FIG. 3, which depicts a DL NN model 300 during a training phase, according to some embodiments of the invention. The model presented in FIG. 3 is an example only, other models may be used. DL NN model 300 may be trained to classify samples as pertaining to or belonging to a first class or a second class."); and
outputting the trained machine learning model (column 7, line 16, "For example, DL NN model 300 may be trained to classify voice samples, or feature sets of voice samples, as genuine or spoof, and may be used by classifier 220 for this purpose. DL NN model 300 may be trained and used for other classification tasks.").
Regarding claims 2, 9 and 16, Havdan further teaches a method, system and computer readable medium wherein the voice sample is positively weighted for training the machine learning model (column 7, line 37, "According to embodiments of the invention, loss function 360 may include a regulation factor, where the regulation factor is set to zero for data samples labeled as pertaining or belonging to the first class and is proportional to the prediction 350 of DL NN model 300 for data samples labeled as pertaining to the second class.").
Regarding claims 3, 10 and 17, Havdan further teaches a method, system and computer readable medium wherein the synthetic voice sample is negatively weighted for training the machine learning model (column 7, line 37, "According to embodiments of the invention, loss function 360 may include a regulation factor, where the regulation factor is set to zero for data samples labeled as pertaining or belonging to the first class and is proportional to the prediction 350 of DL NN model 300 for data samples labeled as pertaining to the second class.").
Regarding claims 4, 11 and 18, Havdan teaches a method, system and computer readable medium further comprising receiving a second voice sample from a person other than the employee, and determining whether the second voice sample is a match using the trained machine learning model (column 6, line 66, "Classifier block 220 may be or may include a trained deep learning neural network (DL NN). Classifier block 220 may obtain the features from feature extraction block 210 and perform inference on the neural network. Inference on the neural network may provide a score, indicative of the chances of the call or interaction being genuine or spoof. The score may be compared to a threshold. If the score satisfies a threshold (e.g., is above or below the threshold, depending on the labeling scheme), it may be determined or concluded that the call is genuine, otherwise, it may be determined that spoofing is detected. Typically, the higher the score the more likely that audio is real or genuine speech.").
Regarding claims 5, 12 and 19, Havdan teaches a method, system and computer readable medium further comprising testing a second voice sample from the employee, and determining whether the second voice sample is a match using the trained machine learning model (column 6, line 66, "Classifier block 220 may be or may include a trained deep learning neural network (DL NN). Classifier block 220 may obtain the features from feature extraction block 210 and perform inference on the neural network. Inference on the neural network may provide a score, indicative of the chances of the call or interaction being genuine or spoof. The score may be compared to a threshold. If the score satisfies a threshold (e.g., is above or below the threshold, depending on the labeling scheme), it may be determined or concluded that the call is genuine, otherwise, it may be determined that spoofing is detected. Typically, the higher the score the more likely that audio is real or genuine speech.").
Regarding claims 6, 13 and 20, Havdan further teaches a method, system and computer readable medium wherein the voice sample includes a set of voice samples including at least two voice samples of differing duration (column 6, line 36, "Reference is now made to FIG. 2, which is a schematic illustration of spoofing detection engine 40, according to embodiments of the present invention. Spoofing detection engine 40 may include silence removal block 200, feature extractor 210 and classifier 220, which may be a NN. Silence removal block 200 may remove silent periods, or periods that do not contain speech, from the voice sample, for example using voice activity detector (VAD), as known in the art. Using embodiments of the invention, a voiced sample with as little as three seconds of voice, or other durations, may be sufficient for detecting spoofing calls. Silence removal block 200 may generate net-speech samples of three seconds or longer from the voice sample. Silence removal block 200 may provide an indication, e.g., to analyst terminal 14, in case the voice sample does not include enough speech, e.g., in case the net-speech sample is shorter than a threshold duration, e.g., three seconds or other durations. Other threshold durations may be used.").
Regarding claims 7 and 14, Havdan further teaches a method, system and computer readable medium wherein the voice sample includes a low quality verification test sample with background noise (column 6, line 38, "Spoofing detection engine 40 may include silence removal block 200, feature extractor 210 and classifier 220, which may be a NN. Silence removal block 200 may remove silent periods, or periods that do not contain speech, from the voice sample, for example using voice activity detector (VAD), as known in the art.").
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
U.S. Patent 11,790,884 to Shakeri et al. teaches creating synthetic voice samples based on voiceprint features.
U.S. Patent 12,254,864 to Lajszczak et al. teaches creating synthetic voice samples for training machine learning models.
U.S. Patent 12,555,567 to Bromand teaches augmenting machine learning model training sets with synthetic voice samples.
U.S. Patent Application Publication 2021/0233541 to Chen et al. teaches implementing a neural network for voiceprint spoof detection.
U.S. Patent Application Publication 2022/0070207 to Simonchik et al. teaches implementing a neural network for speech spoofing prediction.
U.S. Patent Application Publication 2022/0108702 to Sheu et al. teaches speaker recognition with spoof detection.
U.S. Patent Application Publication 2022/0116388 to Johnson et al. teaches speak authentication based on audio data features.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SEAN T SMITH whose telephone number is (571)272-6643. The examiner can normally be reached Monday - Friday 8:00am - 5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, PIERRE-LOUIS DESIR can be reached at (571) 272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SEAN THOMAS SMITH/Examiner, Art Unit 2659
/PIERRE LOUIS DESIR/Supervisory Patent Examiner, Art Unit 2659