Prosecution Insights
Last updated: October 02, 2026
Application No. 19/076,895

DIFFUSION-BASED AUDIO PURIFICATION FOR DEFENDING AGAINST ADVERSARIAL DEEPFAKE ATTACKS

Non-Final OA §103
Filed
Mar 11, 2025
Priority
Mar 12, 2024 — provisional 63/564,449
Examiner
NIPA, WASIKA
Art Unit
Tech Center
Assignee
Pindrop Security Inc.
OA Round
1 (Non-Final)
76%
Grant Probability
Favorable
1-2
OA Rounds
1y 3m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 76% — above average
76%
Career Allowance Rate
237 granted / 314 resolved
+15.5% vs TC avg
Strong +30% interview lift
Without
With
+30.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
9 currently pending
Career history
325
Total Applications
across all art units

Statute-Specific Performance

§101
14.2%
-25.8% vs TC avg
§103
55.8%
+15.8% vs TC avg
§102
2.5%
-37.5% vs TC avg
§112
15.1%
-24.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 314 resolved cases

Office Action

§103
Detailed Action The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This is the initial office action that has been issued in response to patent application, 19/076,895, filed on 03/11/2025. Claims 1-20 are currently pending and have been considered below. Claim 1, 10 and 19 are independent claim. Claim 1 has been cancelled. Information Disclosure Statement The information disclosure statements (IDS's) submitted on 06/09/2025 are in compliance with provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Priority This application has PRO 63/564,449 filed on 03/12/2024. Drawings The drawings filed on 03/11/2025 are accepted by the examiner. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Phatak (US Patent Application No 2021/0280171 A1) in view of Xiao (US Patent Application Publication No 2023/0095092 A1). Regarding Claim 1, Phatak discloses a computer-implemented method for detecting fraudulent calls based on adversarial noise indicating adversarial attacks, the method comprising: extracting, by a computer, a plurality of input features for an input audio signal (Phatak, ¶[0036], components of a system for receiving and analyzing audio signals from end users. ¶[0037], a smartphone may execute the deep-phoneprinting software when receiving an inbound call from another end-user to perform certain downstream operations such as verifying the identity of the other end user or indicating whether the other end-user is using a spoofing service. ¶[0043], extracting speaker independent embeddings, extracting DP vectors); identifying, by the computer, an instance of adversarial noise in the input audio signal based upon the plurality of input features using a diffusion model of a machine-learning architecture, the diffusion model trained to identify instances of adversarial noise features extracted in audio signals (Phatak, ¶[0030], determining whether an identifier associated with the audio source is spoofed, determining a spoofing service that may be used to change the source identifier associated with the audio. ¶[0032], the DP vector is employed in various downstream operations. ¶[0044], the analytic server may perform the pre-processing operations and data augmentation operations when executing certain neural network layers. ¶[0049]- ¶[0052]. ¶[0072], the server places the neural network architecture and the task specific models into the enrollment phase to extract enrolled embeddings for an enrolled audio source); generating, by the computer, a plurality of purified features corresponding to the plurality of input features according to the instance of the adversarial noise as identified in the input audio signal using the diffusion model (Phatak, ¶[0056], the analytics server feeds each enrollment audio signal to parse the audio signal into speech portions and non-speech portions. ¶[0081], data augmentation operations may generate or retrieve certain training audio signals, including clean audio signals and noise samples); generating, by the computer, a deepfake score for the input audio signal indicating a likelihood that the input audio signal is fraudulent using a deepfake detector of the machine- learning architecture based upon the plurality of purified features (Phatak, ¶[0054], predicted similarity scores. ¶[0060], the analytics server may determine a similarity score based upon the distance, differences/similarities, between the enrollment DP vector and the inbound DP vector, where the similarity score indicates the likelihood that the enrollment DP vector originated from the same audio source as the inbound DP vector. ¶[0077], the neural network architecture generates one or more similarity score); and identifying, by the computer, the input audio signal as genuine or fraudulent based upon the deepfake score (Phatak, ¶[0134], the training phase, the exclusion list modeling layers generate predicted outputs (predicted classifications, predicted similarity score) for training audio signals, which are used to determine a level of error. Fig-12, ¶[0159]- ¶[0163], neural network detects spoofing in inbound audio signal. Fig-13, ¶[0169], the spoof detection layers define a binary classifier trained to determine a spoof detection likelihood score based on classification score). Phatak does not explicitly teach the following limitation that Xiao teaches: diffusion model (Xiao, ¶[0057], each denoising step is modeled with a conditional generative adversarial network (GAN), where the GAN is given a noisy observation and tries to produce a less noisy sample. ¶[0104]- ¶[0106], Fig-4A, 4B). Phatak in view of Xiao are analogous art because they are from the “same field of endeavor” and are from the same “problem solving area”. Namely, they pertain to the field of “denoising adversarial network and generative deep learning network”. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the invention of Phatak in view of Xiao to include the idea of providing a denoising diffusion generative adversarial network that directly reduces a number of denoising steps during a reverse process (Xiao, ¶[0055]). Regarding Claim 2, Phatak in view of Xiao discloses the method according to claim 1, further comprising: generating, by the computer, a loss for the diffusion model using a loss function, the loss indicating a distance between the plurality of purified features for the input audio signal and a plurality of expected purified features indicated by a training label associated with the input audio signal (Phatak, ¶[0054], loss layers perform various types of loss functions to evaluate the distances between predicted outputs. ¶[0070], the loss function evaluates the level of error for each task specific model based upon the relative distances. ¶[0088]- ¶[0092], ¶[0101], the loss function may be a sum, concatenation or other combination of the several output layers processing the speech portions. Also Xiao, ¶[0057], each denoising step is modeled with a conditional generative adversarial network (GAN), where the GAN is given a noisy observation and tries to produce a less noisy sample. ¶[0104]- ¶[0106], Fig-4A, 4B); and updating, by the computer, one or more diffusion parameters of the diffusion model based upon the loss (Phatak, ¶[0054], loss layers perform various types of loss functions to evaluate the distances between predicted outputs. ¶[0070], the loss function evaluates the level of error for each task specific model based upon the relative distances. ¶[0088]- ¶[0092], ¶[0101], the loss function may be a sum, concatenation or other combination of the several output layers processing the speech portions. Also Xiao, ¶[0057], each denoising step is modeled with a conditional generative adversarial network (GAN), where the GAN is given a noisy observation and tries to produce a less noisy sample. ¶[0104]- ¶[0106], Fig-4A, 4B). Regarding Claim 3, Phatak in view of Xiao discloses the method according to claim 1, further comprising: generating, by the computer, a loss for the deepfake detector using a loss function, the loss indicating a distance between the deepfake score as generated for the input audio signal and an expected deepfake score indicated by a training label associated with the input audio signal (Phatak, ¶[0054], loss layers perform various types of loss functions to evaluate the distances between predicted outputs. ¶[0070], the loss function evaluates the level of error for each task specific model based upon the relative distances. ¶[0088]- ¶[0092], ¶[0101], the loss function may be a sum, concatenation or other combination of the several output layers processing the speech portions. Also Xiao, ¶[0057], each denoising step is modeled with a conditional generative adversarial network (GAN), where the GAN is given a noisy observation and tries to produce a less noisy sample. ¶[0104]- ¶[0106], Fig-4A, 4B); and updating, by the computer, one or more diffusion parameters of the diffusion model based upon the loss (Xiao, ¶[0064], ¶[0076], the denoising diffusion GAN may be trained to generate an image according to those parameters); and updating, by the computer, one or more detection parameters of the deepfake detector based upon the loss (Xiao, ¶[0064], ¶[0076], the denoising diffusion GAN may be trained to generate an image according to those parameters). Regarding Claim 4, Phatak in view of Xiao discloses the method according to claim 1, further comprising: receiving, by the computer, the input audio signal having the plurality of features (Phatak, ¶[0036], components of a system for receiving and analyzing audio signals from end users. ¶[0037], a smartphone may execute the deep-phoneprinting software when receiving an inbound call from another end-user to perform certain downstream operations such as verifying the identity of the other end user or indicating whether the other end-user is using a spoofing service. ¶[0043], extracting speaker independent embeddings, extracting DP vectors); and executing, by the computer, a transformation function on the input audio signal to convert the input audio signal from a time domain to a transformed domain, wherein the computer extracts the plurality of features from the transformed domain of the input audio signal (Phatak, ¶[0099], ¶[0104], extract various types of features from the portions and transform one or more extracted features from a time-domain representation into a frequency-domain representation by performing an SFT or FFT operation). Regarding Claim 5, Phatak in view of Xiao discloses the method according to claim 4, wherein the transformed domain includes at least one of a Gaussian space, a frequency domain, or a time-frequency domain (Phatak, ¶[0099], ¶[0104], extract various types of features from the portions and transform one or more extracted features from a time-domain representation into a frequency-domain representation by performing an SFT or FFT operation). Regarding Claim 6, Phatak in view of Xiao discloses the method according to claim 1, further comprising generating, by the computer, a clean version of the input audio signal based upon the plurality of purified features using a transform function (Phatak, ¶[0045], transforming the extracted features from a time-domain representation into a frequency-domain representation by performing Short-time Fourier Transforms (SFT) and Fast Fourier Transforms (FFT) operations, among other pre- processing operations. ¶[0097]). Regarding Claim 7, Phatak in view of Xiao discloses the method according to claim 1, further comprising extracting, by the computer, a fakeprint feature vector embedding based upon the plurality of purified features, wherein the deepfake detector generates the deepfake score using the fakeprint feature vector embedding (Phatak, ¶[0032]- ¶[0033], ¶[0134], the training phase, the exclusion list modeling layers generate predicted outputs (predicted classifications, predicted similarity score) for training audio signals, which are used to determine a level of error. Fig-12, ¶[0159]- ¶[0163], neural network detects spoofing in inbound audio signal. Fig-13, ¶[0169], the spoof detection layers define a binary classifier trained to determine a spoof detection likelihood score based on classification score). Regarding Claim 8, Phatak in view of Xiao discloses the method according to claim 1, wherein the computer identifies the input audio signal fraudulent in response to determining that the deepfake scores satisfies a fraud detection threshold score (Phatak, ¶[0054], predicted similarity scores. ¶[0060], the analytics server may determine a similarity score based upon the distance, differences/similarities, between the enrollment DP vector and the inbound DP vector, where the similarity score indicates the likelihood that the enrollment DP vector originated from the same audio source as the inbound DP vector. ¶[0077], the neural network architecture generates one or more similarity score). Regarding Claim 9, Phatak in view of Xiao discloses the method according to claim 8, further comprising: generating, by the computer, an alert notification for display at a user interface indicating that the input audio signal has been identified as fraudulent in response to determining that the deepfake score satisfied the fraud detection threshold score (Phatak, ¶[0054], predicted similarity scores. ¶[0060], the analytics server may determine a similarity score based upon the distance, differences/similarities, between the enrollment DP vector and the inbound DP vector, where the similarity score indicates the likelihood that the enrollment DP vector originated from the same audio source as the inbound DP vector. ¶[0077], the neural network architecture generates one or more similarity score). Regarding Claim 10, Phatak discloses a system for detecting fraudulent calls based on adversarial noise indicating adversarial attacks, the system comprising: a computer comprising at least one processor, the computer configured to (Phatak, Fig-1): extract a plurality of input features for an input audio signal (Phatak, ¶[0036], components of a system for receiving and analyzing audio signals from end users. ¶[0037], a smartphone may execute the deep-phoneprinting software when receiving an inbound call from another end-user to perform certain downstream operations such as verifying the identity of the other end user or indicating whether the other end-user is using a spoofing service. ¶[0043], extracting speaker independent embeddings, extracting DP vectors); identify an instance of adversarial noise in the input audio signal based upon the plurality of input features using a diffusion model of a machine-learning architecture, the diffusion model trained to identify instances of adversarial noise features extracted in audio signals (Phatak, ¶[0030], determining whether an identifier associated with the audio source is spoofed, determining a spoofing service that may be used to change the source identifier associated with the audio. ¶[0032], the DP vector is employed in various downstream operations. ¶[0044], the analytic server may perform the pre-processing operations and data augmentation operations when executing certain neural network layers. ¶[0049]- ¶[0052]. ¶[0072], the server places the neural network architecture and the task specific models into the enrollment phase to extract enrolled embeddings for an enrolled audio source); generate a plurality of purified features corresponding to the plurality of input features according to the instance of the adversarial noise as identified in the input audio signal using the diffusion model (Phatak, ¶[0056], the analytics server feeds each enrollment audio signal to parse the audio signal into speech portions and non-speech portions. ¶[0081], data augmentation operations may generate or retrieve certain training audio signals, including clean audio signals and noise samples); generate a deepfake score for the input audio signal indicating a likelihood that the input audio signal is fraudulent using a deepfake detector of the machine-learning architecture based upon the plurality of purified features (Phatak, ¶[0054], predicted similarity scores. ¶[0060], the analytics server may determine a similarity score based upon the distance, differences/similarities, between the enrollment DP vector and the inbound DP vector, where the similarity score indicates the likelihood that the enrollment DP vector originated from the same audio source as the inbound DP vector. ¶[0077], the neural network architecture generates one or more similarity score); and identify the input audio signal as genuine or fraudulent based upon the deepfake score (Phatak, ¶[0134], the training phase, the exclusion list modeling layers generate predicted outputs (predicted classifications, predicted similarity score) for training audio signals, which are used to determine a level of error. Fig-12, ¶[0159]- ¶[0163], neural network detects spoofing in inbound audio signal. Fig-13, ¶[0169], the spoof detection layers define a binary classifier trained to determine a spoof detection likelihood score based on classification score). Phatak does not explicitly teach the following limitation that Xiao teaches: diffusion model (Xiao, ¶[0057], each denoising step is modeled with a conditional generative adversarial network (GAN), where the GAN is given a noisy observation and tries to produce a less noisy sample. ¶[0104]- ¶[0106], Fig-4A, 4B). Phatak in view of Xiao are analogous art because they are from the “same field of endeavor” and are from the same “problem solving area”. Namely, they pertain to the field of “denoising adversarial network and generative deep learning network”. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the invention of Phatak in view of Xiao to include the idea of providing a denoising diffusion generative adversarial network that directly reduces a number of denoising steps during a reverse process (Xiao, ¶[0055]). Regarding Claim 11, Phatak in view of Xiao discloses the system according to claim 10, wherein the computer is further configured to: generate a loss for the diffusion model using a loss function, the loss indicating a distance between the plurality of purified features for the input audio signal and a plurality of expected purified features indicated by a training label associated with the input audio signal (Phatak, ¶[0054], loss layers perform various types of loss functions to evaluate the distances between predicted outputs. ¶[0070], the loss function evaluates the level of error for each task specific model based upon the relative distances. ¶[0088]- ¶[0092], ¶[0101], the loss function may be a sum, concatenation or other combination of the several output layers processing the speech portions. Also Xiao, ¶[0057], each denoising step is modeled with a conditional generative adversarial network (GAN), where the GAN is given a noisy observation and tries to produce a less noisy sample. ¶[0104]- ¶[0106], Fig-4A, 4B); and update one or more diffusion parameters of the diffusion model based upon the loss (Phatak, ¶[0054], loss layers perform various types of loss functions to evaluate the distances between predicted outputs. ¶[0070], the loss function evaluates the level of error for each task specific model based upon the relative distances. ¶[0088]- ¶[0092], ¶[0101], the loss function may be a sum, concatenation or other combination of the several output layers processing the speech portions. Also Xiao, ¶[0057], each denoising step is modeled with a conditional generative adversarial network (GAN), where the GAN is given a noisy observation and tries to produce a less noisy sample. ¶[0104]- ¶[0106], Fig-4A, 4B). Regarding Claim 12, Phatak in view of Xiao discloses the system according to claim 10, wherein the computer is further configured to: generate a loss for the deepfake detector using a loss function, the loss indicating a distance between the deepfake score as generated for the input audio signal and an expected deepfake score indicated by a training label associated with the input audio signal (Phatak, ¶[0054], loss layers perform various types of loss functions to evaluate the distances between predicted outputs. ¶[0070], the loss function evaluates the level of error for each task specific model based upon the relative distances. ¶[0088]- ¶[0092], ¶[0101], the loss function may be a sum, concatenation or other combination of the several output layers processing the speech portions. Also Xiao, ¶[0057], each denoising step is modeled with a conditional generative adversarial network (GAN), where the GAN is given a noisy observation and tries to produce a less noisy sample. ¶[0104]- ¶[0106], Fig-4A, 4B); and update one or more diffusion parameters of the diffusion model based upon the loss (Xiao, ¶[0064], ¶[0076], the denoising diffusion GAN may be trained to generate an image according to those parameters); and update one or more detection parameters of the deepfake detector based upon the loss (Xiao, ¶[0064], ¶[0076], the denoising diffusion GAN may be trained to generate an image according to those parameters). Regarding Claim 13, Phatak in view of Xiao discloses the system according to claim 10, wherein the computer is further configured to: receive the input audio signal having the plurality of features (Phatak, ¶[0036], components of a system for receiving and analyzing audio signals from end users. ¶[0037], a smartphone may execute the deep-phoneprinting software when receiving an inbound call from another end-user to perform certain downstream operations such as verifying the identity of the other end user or indicating whether the other end-user is using a spoofing service. ¶[0043], extracting speaker independent embeddings, extracting DP vectors); and execute a transformation function on the input audio signal to convert the input audio signal from a time domain to a transformed domain, wherein the computer extracts the plurality of features from the transformed domain of the input audio signal (Phatak, ¶[0099], ¶[0104], extract various types of features from the portions and transform one or more extracted features from a time-domain representation into a frequency-domain representation by performing an SFT or FFT operation). Regarding Claim 14, Phatak in view of Xiao discloses the system according to claim 13, wherein the transformed domain includes at least one of a Gaussian space, a frequency domain, or a time-frequency domain (Phatak, ¶[0099], ¶[0104], extract various types of features from the portions and transform one or more extracted features from a time-domain representation into a frequency-domain representation by performing an SFT or FFT operation). Regarding Claim 15, Phatak in view of Xiao discloses the system according to claim 10, wherein the computer is further configured to generate a clean version of the input audio signal based upon the plurality of purified features using a transform function (Phatak, ¶[0045], transforming the extracted features from a time-domain representation into a frequency-domain representation by performing Short-time Fourier Transforms (SFT) and Fast Fourier Transforms (FFT) operations, among other pre- processing operations. ¶[0097]). Regarding Claim 16, Phatak in view of Xiao discloses the system according to claim 10, wherein the computer is further configured to extract a fakeprint feature vector embedding based upon the plurality of purified features, and wherein the deepfake detector generates the deepfake score using the fakeprint feature vector embedding (Phatak, ¶[0032]- ¶[0033], ¶[0134], the training phase, the exclusion list modeling layers generate predicted outputs (predicted classifications, predicted similarity score) for training audio signals, which are used to determine a level of error. Fig-12, ¶[0159]- ¶[0163], neural network detects spoofing in inbound audio signal. Fig-13, ¶[0169], the spoof detection layers define a binary classifier trained to determine a spoof detection likelihood score based on classification score). Regarding Claim 17, Phatak in view of Xiao discloses the system according to claim 10, wherein the computer identifies the input audio signal fraudulent in response to determining that the deepfake scores satisfies a fraud detection threshold score (Phatak, ¶[0054], predicted similarity scores. ¶[0060], the analytics server may determine a similarity score based upon the distance, differences/similarities, between the enrollment DP vector and the inbound DP vector, where the similarity score indicates the likelihood that the enrollment DP vector originated from the same audio source as the inbound DP vector. ¶[0077], the neural network architecture generates one or more similarity score). Regarding Claim 18, Phatak in view of Xiao discloses the system according to claim 17, wherein the computer is further configured to generate an alert notification for display at a user interface indicating that the input audio signal has been identified as fraudulent in response to determining that the deepfake score satisfied the fraud detection threshold score (Phatak, ¶[0054], predicted similarity scores. ¶[0060], the analytics server may determine a similarity score based upon the distance, differences/similarities, between the enrollment DP vector and the inbound DP vector, where the similarity score indicates the likelihood that the enrollment DP vector originated from the same audio source as the inbound DP vector. ¶[0077], the neural network architecture generates one or more similarity score). Regarding Claim 19, Phatak discloses a non-transitory computer readable medium configured to stored executable instructions for detecting fraudulent calls based on adversarial noise indicating adversarial attacks that when executed by one or more processors, cause the one or more processor to: extract a plurality of input features for an input audio signal (Phatak, ¶[0036], components of a system for receiving and analyzing audio signals from end users. ¶[0037], a smartphone may execute the deep-phoneprinting software when receiving an inbound call from another end-user to perform certain downstream operations such as verifying the identity of the other end user or indicating whether the other end-user is using a spoofing service. ¶[0043], extracting speaker independent embeddings, extracting DP vectors); identify an instance of adversarial noise in the input audio signal based upon the plurality of input features using a diffusion model of a machine-learning architecture, the diffusion model trained to identify instances of adversarial noise features extracted in audio signals (Phatak, ¶[0030], determining whether an identifier associated with the audio source is spoofed, determining a spoofing service that may be used to change the source identifier associated with the audio. ¶[0032], the DP vector is employed in various downstream operations. ¶[0044], the analytic server may perform the pre-processing operations and data augmentation operations when executing certain neural network layers. ¶[0049]- ¶[0052]. ¶[0072], the server places the neural network architecture and the task specific models into the enrollment phase to extract enrolled embeddings for an enrolled audio source); generate a plurality of purified features corresponding to the plurality of input features according to the instance of the adversarial noise as identified in the input audio signal using the diffusion model (Phatak, ¶[0056], the analytics server feeds each enrollment audio signal to parse the audio signal into speech portions and non-speech portions. ¶[0081], data augmentation operations may generate or retrieve certain training audio signals, including clean audio signals and noise samples); generate a deepfake score for the input audio signal indicating a likelihood that the input audio signal is fraudulent using a deepfake detector of the machine-learning architecture based upon the plurality of purified features (Phatak, ¶[0054], predicted similarity scores. ¶[0060], the analytics server may determine a similarity score based upon the distance, differences/similarities, between the enrollment DP vector and the inbound DP vector, where the similarity score indicates the likelihood that the enrollment DP vector originated from the same audio source as the inbound DP vector. ¶[0077], the neural network architecture generates one or more similarity score); and identify the input audio signal as genuine or fraudulent based upon the deepfake score (Phatak, ¶[0134], the training phase, the exclusion list modeling layers generate predicted outputs (predicted classifications, predicted similarity score) for training audio signals, which are used to determine a level of error. Fig-12, ¶[0159]- ¶[0163], neural network detects spoofing in inbound audio signal. Fig-13, ¶[0169], the spoof detection layers define a binary classifier trained to determine a spoof detection likelihood score based on classification score). Phatak does not explicitly teach the following limitation that Xiao teaches: diffusion model (Xiao, ¶[0057], each denoising step is modeled with a conditional generative adversarial network (GAN), where the GAN is given a noisy observation and tries to produce a less noisy sample. ¶[0104]- ¶[0106], Fig-4A, 4B). Phatak in view of Xiao are analogous art because they are from the “same field of endeavor” and are from the same “problem solving area”. Namely, they pertain to the field of “denoising adversarial network and generative deep learning network”. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the invention of Phatak in view of Xiao to include the idea of providing a denoising diffusion generative adversarial network that directly reduces a number of denoising steps during a reverse process (Xiao, ¶[0055]). Regarding Claim 20, Phatak in view of Xiao discloses the computer-readable medium of claim 19, wherein the instructions further instruct the one or more processors to: generate a loss for the deepfake detector using a loss function, the loss indicating a distance between the deepfake score as generated for the input audio signal and an expected deepfake score indicated by a training label associated with the input audio signal (Phatak, ¶[0054], loss layers perform various types of loss functions to evaluate the distances between predicted outputs. ¶[0070], the loss function evaluates the level of error for each task specific model based upon the relative distances. ¶[0088]- ¶[0092], ¶[0101], the loss function may be a sum, concatenation or other combination of the several output layers processing the speech portions. Also Xiao, ¶[0057], each denoising step is modeled with a conditional generative adversarial network (GAN), where the GAN is given a noisy observation and tries to produce a less noisy sample. ¶[0104]- ¶[0106], Fig-4A, 4B); and update one or more diffusion parameters of the diffusion model based upon the loss (Phatak, ¶[0054], loss layers perform various types of loss functions to evaluate the distances between predicted outputs. ¶[0070], the loss function evaluates the level of error for each task specific model based upon the relative distances. ¶[0088]- ¶[0092], ¶[0101], the loss function may be a sum, concatenation or other combination of the several output layers processing the speech portions. Also Xiao, ¶[0057], each denoising step is modeled with a conditional generative adversarial network (GAN), where the GAN is given a noisy observation and tries to produce a less noisy sample. ¶[0104]- ¶[0106], Fig-4A, 4B); and update one or more detection parameters of the deepfake detector based upon the loss (Phatak, ¶[0054], loss layers perform various types of loss functions to evaluate the distances between predicted outputs. ¶[0070], the loss function evaluates the level of error for each task specific model based upon the relative distances. ¶[0088]- ¶[0092], ¶[0101], the loss function may be a sum, concatenation or other combination of the several output layers processing the speech portions. Also Xiao, ¶[0057], each denoising step is modeled with a conditional generative adversarial network (GAN), where the GAN is given a noisy observation and tries to produce a less noisy sample. ¶[0104]- ¶[0106], Fig-4A, 4B). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure (see PTO-Form 892). Any inquiry concerning this communication or earlier communications from the examiner should be directed to WASIKA NIPA whose telephone number is (571)272-8923. The examiner can normally be reached on M-F, 8 am to 5 pm. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jeffrey Pwu can be reached on 571-272-6798. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, Applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /WASIKA NIPA/ Primary Examiner, Art Unit 2433
Read full office action

Prosecution Timeline

Mar 11, 2025
Application Filed
Jun 15, 2026
Applicant Interview (Telephonic)
Jun 16, 2026
Examiner Interview Summary
Aug 10, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12750674
SECURE COMMUNICATION FOR COMMISSIONING AND DECOMMISSIONING CIRCUIT BREAKERS AND PANEL SYSTEM
1y 9m to grant Granted Sep 29, 2026
Patent 12744818
IDENTIFYING SERVERLESS FUNCTIONS WITH OVER-PERMISSIVE ROLES
2y 4m to grant Granted Sep 22, 2026
Patent 12732819
DETECTING CELL SITE SIMULATOR
2y 9m to grant Granted Sep 08, 2026
Patent 12726359
SYSTEM AND METHOD FOR PROVIDING INFORMATION USABLE TO IDENTIFY TRUST IN DATA
3y 4m to grant Granted Sep 01, 2026
Patent 12712863
SYSTEM FOR SECURE DATA TRANSMISSION
1y 12m to grant Granted Aug 18, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
76%
Grant Probability
99%
With Interview (+30.0%)
2y 10m (~1y 3m remaining)
Median Time to Grant
Low
PTA Risk
Based on 314 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month