Detailed Action
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This is the initial office action that has been issued in response to patent application, 19/076,895, filed on 03/11/2025. Claims 1-20 are currently pending and have been considered below. Claim 1, 10 and 19 are independent claim. Claim 1 has been cancelled.
Information Disclosure Statement
The information disclosure statements (IDS's) submitted on 06/09/2025 are in compliance with provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Priority
This application has PRO 63/564,449 filed on 03/12/2024.
Drawings
The drawings filed on 03/11/2025 are accepted by the examiner.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Phatak (US Patent Application No 2021/0280171 A1) in view of Xiao (US Patent Application Publication No 2023/0095092 A1).
Regarding Claim 1, Phatak discloses a computer-implemented method for detecting fraudulent calls based on adversarial noise indicating adversarial attacks, the method comprising:
extracting, by a computer, a plurality of input features for an input audio signal (Phatak, ¶[0036], components of a system for receiving and analyzing audio signals from end users. ¶[0037], a smartphone may execute the deep-phoneprinting software when receiving an inbound call from another end-user to perform certain downstream operations such as verifying the identity of the other end user or indicating whether the other end-user is using a spoofing service. ¶[0043], extracting speaker independent embeddings, extracting DP vectors);
identifying, by the computer, an instance of adversarial noise in the input audio signal based upon the plurality of input features using a diffusion model of a machine-learning architecture, the diffusion model trained to identify instances of adversarial noise features extracted in audio signals (Phatak, ¶[0030], determining whether an identifier associated with the audio source is spoofed, determining a spoofing service that may be used to change the source identifier associated with the audio. ¶[0032], the DP vector is employed in various downstream operations. ¶[0044], the analytic server may perform the pre-processing operations and data augmentation operations when executing certain neural network layers. ¶[0049]- ¶[0052]. ¶[0072], the server places the neural network architecture and the task specific models into the enrollment phase to extract enrolled embeddings for an enrolled audio source);
generating, by the computer, a plurality of purified features corresponding to the plurality of input features according to the instance of the adversarial noise as identified in the input audio signal using the diffusion model (Phatak, ¶[0056], the analytics server feeds each enrollment audio signal to parse the audio signal into speech portions and non-speech portions. ¶[0081], data augmentation operations may generate or retrieve certain training audio signals, including clean audio signals and noise samples);
generating, by the computer, a deepfake score for the input audio signal indicating a likelihood that the input audio signal is fraudulent using a deepfake detector of the machine- learning architecture based upon the plurality of purified features (Phatak, ¶[0054], predicted similarity scores. ¶[0060], the analytics server may determine a similarity score based upon the distance, differences/similarities, between the enrollment DP vector and the inbound DP vector, where the similarity score indicates the likelihood that the enrollment DP vector originated from the same audio source as the inbound DP vector. ¶[0077], the neural network architecture generates one or more similarity score); and
identifying, by the computer, the input audio signal as genuine or fraudulent based upon the deepfake score (Phatak, ¶[0134], the training phase, the exclusion list modeling layers generate predicted outputs (predicted classifications, predicted similarity score) for training audio signals, which are used to determine a level of error. Fig-12, ¶[0159]- ¶[0163], neural network detects spoofing in inbound audio signal. Fig-13, ¶[0169], the spoof detection layers define a binary classifier trained to determine a spoof detection likelihood score based on classification score).
Phatak does not explicitly teach the following limitation that Xiao teaches:
diffusion model (Xiao, ¶[0057], each denoising step is modeled with a conditional generative adversarial network (GAN), where the GAN is given a noisy observation and tries to produce a less noisy sample. ¶[0104]- ¶[0106], Fig-4A, 4B).
Phatak in view of Xiao are analogous art because they are from the “same field of endeavor” and are from the same “problem solving area”. Namely, they pertain to the field of “denoising adversarial network and generative deep learning network”. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the invention of Phatak in view of Xiao to include the idea of providing a denoising diffusion generative adversarial network that directly reduces a number of denoising steps during a reverse process (Xiao, ¶[0055]).
Regarding Claim 2, Phatak in view of Xiao discloses the method according to claim 1, further comprising:
generating, by the computer, a loss for the diffusion model using a loss function, the loss indicating a distance between the plurality of purified features for the input audio signal and a plurality of expected purified features indicated by a training label associated with the input audio signal (Phatak, ¶[0054], loss layers perform various types of loss functions to evaluate the distances between predicted outputs. ¶[0070], the loss function evaluates the level of error for each task specific model based upon the relative distances. ¶[0088]- ¶[0092], ¶[0101], the loss function may be a sum, concatenation or other combination of the several output layers processing the speech portions. Also Xiao, ¶[0057], each denoising step is modeled with a conditional generative adversarial network (GAN), where the GAN is given a noisy observation and tries to produce a less noisy sample. ¶[0104]- ¶[0106], Fig-4A, 4B); and
updating, by the computer, one or more diffusion parameters of the diffusion model based upon the loss (Phatak, ¶[0054], loss layers perform various types of loss functions to evaluate the distances between predicted outputs. ¶[0070], the loss function evaluates the level of error for each task specific model based upon the relative distances. ¶[0088]- ¶[0092], ¶[0101], the loss function may be a sum, concatenation or other combination of the several output layers processing the speech portions. Also Xiao, ¶[0057], each denoising step is modeled with a conditional generative adversarial network (GAN), where the GAN is given a noisy observation and tries to produce a less noisy sample. ¶[0104]- ¶[0106], Fig-4A, 4B).
Regarding Claim 3, Phatak in view of Xiao discloses the method according to claim 1, further comprising:
generating, by the computer, a loss for the deepfake detector using a loss function, the loss indicating a distance between the deepfake score as generated for the input audio signal and an expected deepfake score indicated by a training label associated with the input audio signal (Phatak, ¶[0054], loss layers perform various types of loss functions to evaluate the distances between predicted outputs. ¶[0070], the loss function evaluates the level of error for each task specific model based upon the relative distances. ¶[0088]- ¶[0092], ¶[0101], the loss function may be a sum, concatenation or other combination of the several output layers processing the speech portions. Also Xiao, ¶[0057], each denoising step is modeled with a conditional generative adversarial network (GAN), where the GAN is given a noisy observation and tries to produce a less noisy sample. ¶[0104]- ¶[0106], Fig-4A, 4B); and
updating, by the computer, one or more diffusion parameters of the diffusion model based upon the loss (Xiao, ¶[0064], ¶[0076], the denoising diffusion GAN may be trained to generate an image according to those parameters); and
updating, by the computer, one or more detection parameters of the deepfake detector based upon the loss (Xiao, ¶[0064], ¶[0076], the denoising diffusion GAN may be trained to generate an image according to those parameters).
Regarding Claim 4, Phatak in view of Xiao discloses the method according to claim 1, further comprising:
receiving, by the computer, the input audio signal having the plurality of features (Phatak, ¶[0036], components of a system for receiving and analyzing audio signals from end users. ¶[0037], a smartphone may execute the deep-phoneprinting software when receiving an inbound call from another end-user to perform certain downstream operations such as verifying the identity of the other end user or indicating whether the other end-user is using a spoofing service. ¶[0043], extracting speaker independent embeddings, extracting DP vectors); and
executing, by the computer, a transformation function on the input audio signal to convert the input audio signal from a time domain to a transformed domain, wherein the computer extracts the plurality of features from the transformed domain of the input audio signal (Phatak, ¶[0099], ¶[0104], extract various types of features from the portions and transform one or more extracted features from a time-domain representation into a frequency-domain representation by performing an SFT or FFT operation).
Regarding Claim 5, Phatak in view of Xiao discloses the method according to claim 4, wherein the transformed domain includes at least one of a Gaussian space, a frequency domain, or a time-frequency domain (Phatak, ¶[0099], ¶[0104], extract various types of features from the portions and transform one or more extracted features from a time-domain representation into a frequency-domain representation by performing an SFT or FFT operation).
Regarding Claim 6, Phatak in view of Xiao discloses the method according to claim 1, further comprising generating, by the computer, a clean version of the input audio signal based upon the plurality of purified features using a transform function (Phatak, ¶[0045], transforming the extracted features from a time-domain representation into a frequency-domain representation by performing Short-time Fourier Transforms (SFT) and Fast Fourier Transforms (FFT) operations, among other pre- processing operations. ¶[0097]).
Regarding Claim 7, Phatak in view of Xiao discloses the method according to claim 1, further comprising extracting, by the computer, a fakeprint feature vector embedding based upon the plurality of purified features,
wherein the deepfake detector generates the deepfake score using the fakeprint feature vector embedding (Phatak, ¶[0032]- ¶[0033], ¶[0134], the training phase, the exclusion list modeling layers generate predicted outputs (predicted classifications, predicted similarity score) for training audio signals, which are used to determine a level of error. Fig-12, ¶[0159]- ¶[0163], neural network detects spoofing in inbound audio signal. Fig-13, ¶[0169], the spoof detection layers define a binary classifier trained to determine a spoof detection likelihood score based on classification score).
Regarding Claim 8, Phatak in view of Xiao discloses the method according to claim 1, wherein the computer identifies the input audio signal fraudulent in response to determining that the deepfake scores satisfies a fraud detection threshold score (Phatak, ¶[0054], predicted similarity scores. ¶[0060], the analytics server may determine a similarity score based upon the distance, differences/similarities, between the enrollment DP vector and the inbound DP vector, where the similarity score indicates the likelihood that the enrollment DP vector originated from the same audio source as the inbound DP vector. ¶[0077], the neural network architecture generates one or more similarity score).
Regarding Claim 9, Phatak in view of Xiao discloses the method according to claim 8, further comprising: generating, by the computer, an alert notification for display at a user interface indicating that the input audio signal has been identified as fraudulent in response to determining that the deepfake score satisfied the fraud detection threshold score (Phatak, ¶[0054], predicted similarity scores. ¶[0060], the analytics server may determine a similarity score based upon the distance, differences/similarities, between the enrollment DP vector and the inbound DP vector, where the similarity score indicates the likelihood that the enrollment DP vector originated from the same audio source as the inbound DP vector. ¶[0077], the neural network architecture generates one or more similarity score).
Regarding Claim 10, Phatak discloses a system for detecting fraudulent calls based on adversarial noise indicating adversarial attacks, the system comprising:
a computer comprising at least one processor, the computer configured to (Phatak, Fig-1):
extract a plurality of input features for an input audio signal (Phatak, ¶[0036], components of a system for receiving and analyzing audio signals from end users. ¶[0037], a smartphone may execute the deep-phoneprinting software when receiving an inbound call from another end-user to perform certain downstream operations such as verifying the identity of the other end user or indicating whether the other end-user is using a spoofing service. ¶[0043], extracting speaker independent embeddings, extracting DP vectors);
identify an instance of adversarial noise in the input audio signal based upon the plurality of input features using a diffusion model of a machine-learning architecture, the diffusion model trained to identify instances of adversarial noise features extracted in audio signals (Phatak, ¶[0030], determining whether an identifier associated with the audio source is spoofed, determining a spoofing service that may be used to change the source identifier associated with the audio. ¶[0032], the DP vector is employed in various downstream operations. ¶[0044], the analytic server may perform the pre-processing operations and data augmentation operations when executing certain neural network layers. ¶[0049]- ¶[0052]. ¶[0072], the server places the neural network architecture and the task specific models into the enrollment phase to extract enrolled embeddings for an enrolled audio source);
generate a plurality of purified features corresponding to the plurality of input features according to the instance of the adversarial noise as identified in the input audio signal using the diffusion model (Phatak, ¶[0056], the analytics server feeds each enrollment audio signal to parse the audio signal into speech portions and non-speech portions. ¶[0081], data augmentation operations may generate or retrieve certain training audio signals, including clean audio signals and noise samples);
generate a deepfake score for the input audio signal indicating a likelihood that the input audio signal is fraudulent using a deepfake detector of the machine-learning architecture based upon the plurality of purified features (Phatak, ¶[0054], predicted similarity scores. ¶[0060], the analytics server may determine a similarity score based upon the distance, differences/similarities, between the enrollment DP vector and the inbound DP vector, where the similarity score indicates the likelihood that the enrollment DP vector originated from the same audio source as the inbound DP vector. ¶[0077], the neural network architecture generates one or more similarity score); and
identify the input audio signal as genuine or fraudulent based upon the deepfake score (Phatak, ¶[0134], the training phase, the exclusion list modeling layers generate predicted outputs (predicted classifications, predicted similarity score) for training audio signals, which are used to determine a level of error. Fig-12, ¶[0159]- ¶[0163], neural network detects spoofing in inbound audio signal. Fig-13, ¶[0169], the spoof detection layers define a binary classifier trained to determine a spoof detection likelihood score based on classification score).
Phatak does not explicitly teach the following limitation that Xiao teaches:
diffusion model (Xiao, ¶[0057], each denoising step is modeled with a conditional generative adversarial network (GAN), where the GAN is given a noisy observation and tries to produce a less noisy sample. ¶[0104]- ¶[0106], Fig-4A, 4B).
Phatak in view of Xiao are analogous art because they are from the “same field of endeavor” and are from the same “problem solving area”. Namely, they pertain to the field of “denoising adversarial network and generative deep learning network”. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the invention of Phatak in view of Xiao to include the idea of providing a denoising diffusion generative adversarial network that directly reduces a number of denoising steps during a reverse process (Xiao, ¶[0055]).
Regarding Claim 11, Phatak in view of Xiao discloses the system according to claim 10, wherein the computer is further configured to:
generate a loss for the diffusion model using a loss function, the loss indicating a distance between the plurality of purified features for the input audio signal and a plurality of expected purified features indicated by a training label associated with the input audio signal (Phatak, ¶[0054], loss layers perform various types of loss functions to evaluate the distances between predicted outputs. ¶[0070], the loss function evaluates the level of error for each task specific model based upon the relative distances. ¶[0088]- ¶[0092], ¶[0101], the loss function may be a sum, concatenation or other combination of the several output layers processing the speech portions. Also Xiao, ¶[0057], each denoising step is modeled with a conditional generative adversarial network (GAN), where the GAN is given a noisy observation and tries to produce a less noisy sample. ¶[0104]- ¶[0106], Fig-4A, 4B); and
update one or more diffusion parameters of the diffusion model based upon the loss (Phatak, ¶[0054], loss layers perform various types of loss functions to evaluate the distances between predicted outputs. ¶[0070], the loss function evaluates the level of error for each task specific model based upon the relative distances. ¶[0088]- ¶[0092], ¶[0101], the loss function may be a sum, concatenation or other combination of the several output layers processing the speech portions. Also Xiao, ¶[0057], each denoising step is modeled with a conditional generative adversarial network (GAN), where the GAN is given a noisy observation and tries to produce a less noisy sample. ¶[0104]- ¶[0106], Fig-4A, 4B).
Regarding Claim 12, Phatak in view of Xiao discloses the system according to claim 10, wherein the computer is further configured to:
generate a loss for the deepfake detector using a loss function, the loss indicating a distance between the deepfake score as generated for the input audio signal and an expected deepfake score indicated by a training label associated with the input audio signal (Phatak, ¶[0054], loss layers perform various types of loss functions to evaluate the distances between predicted outputs. ¶[0070], the loss function evaluates the level of error for each task specific model based upon the relative distances. ¶[0088]- ¶[0092], ¶[0101], the loss function may be a sum, concatenation or other combination of the several output layers processing the speech portions. Also Xiao, ¶[0057], each denoising step is modeled with a conditional generative adversarial network (GAN), where the GAN is given a noisy observation and tries to produce a less noisy sample. ¶[0104]- ¶[0106], Fig-4A, 4B); and
update one or more diffusion parameters of the diffusion model based upon the loss (Xiao, ¶[0064], ¶[0076], the denoising diffusion GAN may be trained to generate an image according to those parameters); and
update one or more detection parameters of the deepfake detector based upon the loss (Xiao, ¶[0064], ¶[0076], the denoising diffusion GAN may be trained to generate an image according to those parameters).
Regarding Claim 13, Phatak in view of Xiao discloses the system according to claim 10, wherein the computer is further configured to:
receive the input audio signal having the plurality of features (Phatak, ¶[0036], components of a system for receiving and analyzing audio signals from end users. ¶[0037], a smartphone may execute the deep-phoneprinting software when receiving an inbound call from another end-user to perform certain downstream operations such as verifying the identity of the other end user or indicating whether the other end-user is using a spoofing service. ¶[0043], extracting speaker independent embeddings, extracting DP vectors); and
execute a transformation function on the input audio signal to convert the input audio signal from a time domain to a transformed domain, wherein the computer extracts the plurality of features from the transformed domain of the input audio signal (Phatak, ¶[0099], ¶[0104], extract various types of features from the portions and transform one or more extracted features from a time-domain representation into a frequency-domain representation by performing an SFT or FFT operation).
Regarding Claim 14, Phatak in view of Xiao discloses the system according to claim 13, wherein the transformed domain includes at least one of a Gaussian space, a frequency domain, or a time-frequency domain (Phatak, ¶[0099], ¶[0104], extract various types of features from the portions and transform one or more extracted features from a time-domain representation into a frequency-domain representation by performing an SFT or FFT operation).
Regarding Claim 15, Phatak in view of Xiao discloses the system according to claim 10, wherein the computer is further configured to generate a clean version of the input audio signal based upon the plurality of purified features using a transform function (Phatak, ¶[0045], transforming the extracted features from a time-domain representation into a frequency-domain representation by performing Short-time Fourier Transforms (SFT) and Fast Fourier Transforms (FFT) operations, among other pre- processing operations. ¶[0097]).
Regarding Claim 16, Phatak in view of Xiao discloses the system according to claim 10, wherein the computer is further configured to extract a fakeprint feature vector embedding based upon the plurality of purified features, and wherein the deepfake detector generates the deepfake score using the fakeprint feature vector embedding (Phatak, ¶[0032]- ¶[0033], ¶[0134], the training phase, the exclusion list modeling layers generate predicted outputs (predicted classifications, predicted similarity score) for training audio signals, which are used to determine a level of error. Fig-12, ¶[0159]- ¶[0163], neural network detects spoofing in inbound audio signal. Fig-13, ¶[0169], the spoof detection layers define a binary classifier trained to determine a spoof detection likelihood score based on classification score).
Regarding Claim 17, Phatak in view of Xiao discloses the system according to claim 10, wherein the computer identifies the input audio signal fraudulent in response to determining that the deepfake scores satisfies a fraud detection threshold score (Phatak, ¶[0054], predicted similarity scores. ¶[0060], the analytics server may determine a similarity score based upon the distance, differences/similarities, between the enrollment DP vector and the inbound DP vector, where the similarity score indicates the likelihood that the enrollment DP vector originated from the same audio source as the inbound DP vector. ¶[0077], the neural network architecture generates one or more similarity score).
Regarding Claim 18, Phatak in view of Xiao discloses the system according to claim 17, wherein the computer is further configured to generate an alert notification for display at a user interface indicating that the input audio signal has been identified as fraudulent in response to determining that the deepfake score satisfied the fraud detection threshold score (Phatak, ¶[0054], predicted similarity scores. ¶[0060], the analytics server may determine a similarity score based upon the distance, differences/similarities, between the enrollment DP vector and the inbound DP vector, where the similarity score indicates the likelihood that the enrollment DP vector originated from the same audio source as the inbound DP vector. ¶[0077], the neural network architecture generates one or more similarity score).
Regarding Claim 19, Phatak discloses a non-transitory computer readable medium configured to stored executable instructions for detecting fraudulent calls based on adversarial noise indicating adversarial attacks that when executed by one or more processors, cause the one or more processor to:
extract a plurality of input features for an input audio signal (Phatak, ¶[0036], components of a system for receiving and analyzing audio signals from end users. ¶[0037], a smartphone may execute the deep-phoneprinting software when receiving an inbound call from another end-user to perform certain downstream operations such as verifying the identity of the other end user or indicating whether the other end-user is using a spoofing service. ¶[0043], extracting speaker independent embeddings, extracting DP vectors);
identify an instance of adversarial noise in the input audio signal based upon the plurality of input features using a diffusion model of a machine-learning architecture, the diffusion model trained to identify instances of adversarial noise features extracted in audio signals (Phatak, ¶[0030], determining whether an identifier associated with the audio source is spoofed, determining a spoofing service that may be used to change the source identifier associated with the audio. ¶[0032], the DP vector is employed in various downstream operations. ¶[0044], the analytic server may perform the pre-processing operations and data augmentation operations when executing certain neural network layers. ¶[0049]- ¶[0052]. ¶[0072], the server places the neural network architecture and the task specific models into the enrollment phase to extract enrolled embeddings for an enrolled audio source);
generate a plurality of purified features corresponding to the plurality of input features according to the instance of the adversarial noise as identified in the input audio signal using the diffusion model (Phatak, ¶[0056], the analytics server feeds each enrollment audio signal to parse the audio signal into speech portions and non-speech portions. ¶[0081], data augmentation operations may generate or retrieve certain training audio signals, including clean audio signals and noise samples);
generate a deepfake score for the input audio signal indicating a likelihood that the input audio signal is fraudulent using a deepfake detector of the machine-learning architecture based upon the plurality of purified features (Phatak, ¶[0054], predicted similarity scores. ¶[0060], the analytics server may determine a similarity score based upon the distance, differences/similarities, between the enrollment DP vector and the inbound DP vector, where the similarity score indicates the likelihood that the enrollment DP vector originated from the same audio source as the inbound DP vector. ¶[0077], the neural network architecture generates one or more similarity score); and
identify the input audio signal as genuine or fraudulent based upon the deepfake score (Phatak, ¶[0134], the training phase, the exclusion list modeling layers generate predicted outputs (predicted classifications, predicted similarity score) for training audio signals, which are used to determine a level of error. Fig-12, ¶[0159]- ¶[0163], neural network detects spoofing in inbound audio signal. Fig-13, ¶[0169], the spoof detection layers define a binary classifier trained to determine a spoof detection likelihood score based on classification score).
Phatak does not explicitly teach the following limitation that Xiao teaches:
diffusion model (Xiao, ¶[0057], each denoising step is modeled with a conditional generative adversarial network (GAN), where the GAN is given a noisy observation and tries to produce a less noisy sample. ¶[0104]- ¶[0106], Fig-4A, 4B).
Phatak in view of Xiao are analogous art because they are from the “same field of endeavor” and are from the same “problem solving area”. Namely, they pertain to the field of “denoising adversarial network and generative deep learning network”. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the invention of Phatak in view of Xiao to include the idea of providing a denoising diffusion generative adversarial network that directly reduces a number of denoising steps during a reverse process (Xiao, ¶[0055]).
Regarding Claim 20, Phatak in view of Xiao discloses the computer-readable medium of claim 19, wherein the instructions further instruct the one or more processors to:
generate a loss for the deepfake detector using a loss function, the loss indicating a distance between the deepfake score as generated for the input audio signal and an expected deepfake score indicated by a training label associated with the input audio signal (Phatak, ¶[0054], loss layers perform various types of loss functions to evaluate the distances between predicted outputs. ¶[0070], the loss function evaluates the level of error for each task specific model based upon the relative distances. ¶[0088]- ¶[0092], ¶[0101], the loss function may be a sum, concatenation or other combination of the several output layers processing the speech portions. Also Xiao, ¶[0057], each denoising step is modeled with a conditional generative adversarial network (GAN), where the GAN is given a noisy observation and tries to produce a less noisy sample. ¶[0104]- ¶[0106], Fig-4A, 4B); and
update one or more diffusion parameters of the diffusion model based upon the loss (Phatak, ¶[0054], loss layers perform various types of loss functions to evaluate the distances between predicted outputs. ¶[0070], the loss function evaluates the level of error for each task specific model based upon the relative distances. ¶[0088]- ¶[0092], ¶[0101], the loss function may be a sum, concatenation or other combination of the several output layers processing the speech portions. Also Xiao, ¶[0057], each denoising step is modeled with a conditional generative adversarial network (GAN), where the GAN is given a noisy observation and tries to produce a less noisy sample. ¶[0104]- ¶[0106], Fig-4A, 4B); and
update one or more detection parameters of the deepfake detector based upon the loss (Phatak, ¶[0054], loss layers perform various types of loss functions to evaluate the distances between predicted outputs. ¶[0070], the loss function evaluates the level of error for each task specific model based upon the relative distances. ¶[0088]- ¶[0092], ¶[0101], the loss function may be a sum, concatenation or other combination of the several output layers processing the speech portions. Also Xiao, ¶[0057], each denoising step is modeled with a conditional generative adversarial network (GAN), where the GAN is given a noisy observation and tries to produce a less noisy sample. ¶[0104]- ¶[0106], Fig-4A, 4B).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure (see PTO-Form 892).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to WASIKA NIPA whose telephone number is (571)272-8923. The examiner can normally be reached on M-F, 8 am to 5 pm. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jeffrey Pwu can be reached on 571-272-6798. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, Applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/WASIKA NIPA/ Primary Examiner, Art Unit 2433