DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Objections
Claim 1 is objected to because of the following informalities: “extracting, a computer, input features …” should be “extracting, by a computer, input features …” Appropriate correction is required.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
1. Claims 1-6 and 11-16 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Suthokumar “Spoofing Countermeasures for Voice Biometric System: Feature Extraction, Modelling and Compensation” (“Suthokumar”)
Per claim 1, Suthokumar discloses a computer-implemented method for detecting fraudulent speech based on voiced-speech and unvoiced-speech, the method comprising:
extracting, a computer, input features for an input audio signal including speech audio data having voiced-speech portions and unvoiced-speech portions (fig. 3.1; Figure 3.1 shows an overview of the proposed approach in which speech frames are initially identified as either high energy (HE) or low energy (LE) frames using a voice activity detector (VAD) …, sec. 3.1.2);
identifying, by the computer, a voiced-speech portion of the speech audio signal and an unvoiced-speech portion of the speech audio signal using a segmentation engine of a machine-learning architecture, the segmentation engine trained to identify instances of at least one of voiced-speech portions or unvoiced-speech portions according to the input features (fig. 2.17; fig. 3.1; Figure 3.1 shows an overview of the proposed approach in which speech frames are initially identified as either high energy (HE) or low energy (LE) frames using a voice activity detector (VAD)…., sec. 3.1.2; sec. 3.1.3.3);
generating, by the computer, a voiced-speech segment containing the voiced-speech portion from the speech audio data and an unvoiced-speech segment containing the unvoiced- speech portion from the speech audio data (Low Energy Features, High Energy Features, fig. 3.1; Here it is expected that the LE frames will contain unvoiced speech …, sec. 3.1.2);
generating, by the computer, a first risk score for the voiced-speech segment indicating a first likelihood that the input audio signal is fraudulent using a first deepfake detector of the machine-learning architecture based upon a set of voiced features for the voiced-speech segment (the log-likelihood ratios from both HE and LE systems …, sec. 3.1.2);
generating, by the computer, a second risk score for the unvoiced-speech segment of the input audio signal indicating a second likelihood that the input audio signal is fraudulent using a second deepfake detector of the machine-learning architecture based upon a set of unvoiced features for the unvoiced-speech segment (the log-likelihood ratios from both HE and LE systems …, sec. 3.1.2);
generating, by the computer, an overall risk score for the input audio signal based upon the first risk score and second risk score, the overall risk indicating a third likelihood that the input audio signal is fraudulent (Score fusion, Genuine/Spoofed, fig. 3.1; the log-likelihood ratios from both HE and LE systems are linearly fused to obtain the final decision, sec. 3.1.2); and
identifying, by the computer, the input audio signal as genuine or fraudulent based upon overall risk score (Genuine/Spoofed, fig. 3.1; the log-likelihood ratios from both HE and LE systems are linearly fused to obtain the final decision, sec. 3.1.2).
Per claim 2, Suthokumar discloses the method according to claim 1, further comprising extracting, by the computer, the set of voiced features for the voiced-speech segment and the set of unvoiced features for the unvoiced- speech segment (fig. 3.1).
Per claim 3, Suthokumar discloses the method according to claim 2, further comprising: extracting, by the computer, a first fakeprint feature vector embedding using the set of voiced features for the voiced-speech segment (GMM Classifier, Spoof, fig. 3.1); and
extracting, by the computer, a second fakeprint feature vector embedding using the set of unvoiced features for the unvoiced-speech segment (GMM Classifier, Spoof, fig. 3.1).
Per claim 4, Suthokumar discloses the method according to claim 1, further comprising detecting, by the computer, the voiced-speech portion based upon a pitch frequency indicative of the voiced-speech using a pitch detector of the segmentation engine (fig. 1.1; fig. 3.1; Long term features, fig. 2.10; Fundamental Frequency Variation (FFV) Feature: The use of FFV features to capture information related to the pitch contours …, sec. 2.3.1.4).
Per claim 5, Suthokumar discloses the method according to claim 1, further comprising detecting, by the computer, the unvoiced-speech based upon a pitch frequency indicative of the unvoiced-speech using a pitch detector of the segmentation engine (fig. 1.1; fig. 3.1; Long term features, fig. 2.10; Fundamental Frequency Variation (FFV) Feature: The use of FFV features to capture information related to the pitch contours …, sec. 2.3.1.4).
Per claim 6, Suthokumar discloses the method according to claim 1, further comprising: detecting, by the computer, a non-speech portion of the input audio signal using a Speech Activity Detection (SAD) engine trained to trained to identify instances of non-speech portions according to the input features (sec. 3.1.3.1);
generating, by the computer, a non-speech segment containing the non-speech portion from the input audio signal (sec. 3.1.3.1); and
filtering, by the computer, the non-speech segment from the input audio signal (Ideally, the LE frames should comprise of unvoiced frames and frames that capture start and end points of voiced speech (V-UV and UV-V) only, and not incorporate silence frames containing no speech…., sec. 3.1.3.1)
Per claim 11, Suthokumar discloses a system for detecting fraudulent speech based on voiced-speech and unvoiced-speech, the system comprising:
a computer comprising at least one processor, the computer configured to: extract input features for an input audio signal including speech audio data having voiced-speech portions and unvoiced-speech portions (fig. 3.1; Figure 3.1 shows an overview of the proposed approach in which speech frames are initially identified as either high energy (HE) or low energy (LE) frames using a voice activity detector (VAD) …, sec. 3.1.2);
identify a voiced-speech portion of the speech audio signal and an unvoiced-speech portion of the speech audio signal using a segmentation engine of a machine-learning architecture, the segmentation engine trained to identify instances of at least one of voiced-speech portions or unvoiced-speech portions according to the input features (fig. 2.17; fig. 3.1; Figure 3.1 shows an overview of the proposed approach in which speech frames are initially identified as either high energy (HE) or low energy (LE) frames using a voice activity detector (VAD)…., sec. 3.1.2; sec. 3.1.3.3);
generate a voiced-speech segment containing the voiced-speech portion from the speech audio data, and an unvoiced-speech segment containing the unvoiced-speech portion from the speech audio data (Low Energy Features, High Energy Features, fig. 3.1; Here it is expected that the LE frames will contain unvoiced speech …, sec. 3.1.2);
generate a first risk score for the voiced-speech segment indicating a first likelihood that the input audio signal is fraudulent using a first deepfake detector of the machine-learning architecture based upon a set of voiced features for the voiced-speech segment (the log-likelihood ratios from both HE and LE systems …, sec. 3.1.2);
generate a second risk score for the unvoiced-speech segment of the input audio signal indicating a second likelihood that the input audio signal is fraudulent using a second deepfake detector of the machine-learning architecture based upon a set of unvoiced features for the unvoiced-speech segment (the log-likelihood ratios from both HE and LE systems …, sec. 3.1.2);
generate an overall risk score for the input audio signal based upon the first risk score and second risk score, the overall risk indicating a third likelihood that the input audio signal is fraudulent (Score fusion, Genuine/Spoofed, fig. 3.1; the log-likelihood ratios from both HE and LE systems are linearly fused to obtain the final decision, sec. 3.1.2); and
identify the input audio signal as genuine or fraudulent based upon overall risk score (Genuine/Spoofed, fig. 3.1; the log-likelihood ratios from both HE and LE systems are linearly fused to obtain the final decision, sec. 3.1.2).
Per claim 12, Suthokumar discloses the system according to claim 11, wherein the computer is further configured to extract the set of voiced features for the voiced-speech segment and the set of unvoiced features for the unvoiced-speech segment (fig. 3.1).
Per claim 13, Suthokumar discloses the system according to claim 12, wherein the computer is further configured to: extract a first fakeprint feature vector embedding using the set of voiced features for the voiced-speech segment (GMM Classifier, Spoof, fig. 3.1); and
extract a second fakeprint feature vector embedding using the set of unvoiced features for the unvoiced-speech segment (GMM Classifier, Spoof, fig. 3.1).
Per claim 14, Suthokumar discloses the system according to claim 11, wherein the computer is further configured to detect the voiced-speech portion based upon a pitch frequency indicative of the voiced-speech using a pitch detector of the segmentation engine (fig. 1.1; fig. 3.1; Long term features, fig. 2.10; Fundamental Frequency Variation (FFV) Feature: The use of FFV features to capture information related to the pitch contours …, sec. 2.3.1.4).
Per claim 15, Suthokumar discloses the system according to claim 11, wherein the computer is further configured to detect the unvoiced-speech based upon a pitch frequency indicative of the unvoiced-speech using a pitch detector of the segmentation engine (fig. 1.1; fig. 3.1; Long term features, fig. 2.10; Fundamental Frequency Variation (FFV) Feature: The use of FFV features to capture information related to the pitch contours …, sec. 2.3.1.4).
Per claim 16, Suthokumar discloses the system according to claim 11, wherein the computer is further configured to: detect a non-speech portion of the input audio signal using a Speech Activity Detection (SAD) engine trained to trained to identify instances of non-speech portions according to the input features (sec. 3.1.3.1);
generate a non-speech segment containing the non-speech portion from the input audio signa (sec. 3.1.3.1); and
filter the non-speech segment from the input audio signal (Ideally, the LE frames should comprise of unvoiced frames and frames that capture start and end points of voiced speech (V-UV and UV-V) only, and not incorporate silence frames containing no speech…., sec. 3.1.3.1).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
2. Claims 7 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Suthokumar in view of Wang et al US 2022/0165297 A1 (“Wang”)
Per claim 7, Suthokumar discloses the method according to claim 6,
Suthokumar discloses generating, by the computer, a risk score for the input audio signal indicating a likelihood that the input audio signal is fraudulent using a third deepfake detector of the machine-learning architecture based upon a set of features for the input audio signal having the voiced-speech segment and unvoiced-speech segment, wherein the computer generates the overall risk score based upon the first risk score and the second risk score, and the third risk score (fig. 3.1)
Suthokumar does not explicitly disclose generating, by the computer, a third risk score for the input audio signal indicating a fourth likelihood that the input audio signal is fraudulent using a third deepfake detector of the machine-learning architecture based upon a set of features for the input audio signal having the voiced-speech segment, unvoiced-speech segment, and non-speech segment, wherein the computer generates the overall risk score based upon the first risk score, the second risk score, and the third risk score
However, this feature is suggested by Wang that discloses separating the audio into silence, unvoiced and voiced segments and calculating an overall score based on the score of the silence and unvoiced segments as well as the score of the voiced segment using deepfake detector models (fig. 5A)
It would have been obvious to one of ordinary skill in the art before the effective filing of the instant invention to try to implement the invention according to Wangs teachings, because such implementation would have resulted in providing higher discriminative information used in identifying replay channels/deepfake audio channels.
Per claim 17, Suthokumar discloses the system according to claim 11,
System claim 17 and method claim 7 are related as system and the method of using same, with each claimed element's function corresponding to the claimed method step. Accordingly claim 17 is similarly rejected under the same rationale as applied above with respect to claim 7.
3. Claims 8-10 and 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over Suthokumar in view of Shaaban et al “Audio Deepfake Approaches” (“Shaaban”)
Per claim 8, Suthokumar discloses the method according to claim 1,
Suthokumar does not explicitly disclose generating, by the computer, a loss for the machine-learning architecture using a loss function, the loss indicating a distance between the overall risk score as generated for the input audio signal and an expected overall risk score indicated by a training label associated with the input audio signal or updating, by the computer, one or more parameters of at least one of the first deepfake detector or the second deepfake detector based upon the loss
However, these features are taught by Shaaban:
generating, by the computer, a loss for the machine-learning architecture using a loss function, the loss indicating a distance between the overall risk score as generated for the input audio signal and an expected overall risk score indicated by a training label associated with the input audio signal (pg. 132662, sec. 2); and
updating, by the computer, one or more parameters of at least one of the first deepfake detector or the second deepfake detector based upon the loss (pg. 132662, sec. 2; pg. 132673, training as involving updating weights of models)
It would have been obvious to one of ordinary skill in the art before the effective filing of the instant invention to combine the teachings of Shaaban with the method of Suthokumar in arriving at the missing features of Suthokumar, because such combination would have resulted in facilitating a model’s ability to differentiate between authentic and counterfeit audio (Shaaban, pg. 132662)
Per claim 9, Suthokumar discloses the e method according to claim 1,
Suthokumar does not explicitly disclose generating, by the computer, a loss for a first deepfake detector using a loss function, the loss indicating a distance between the first risk score as generated for the voiced-speech segment and an expected first risk score indicated by a training label associated with the input audio signal or updating, by the computer, one or more parameters of the first deepfake detector based upon the loss
However, these features are suggested by Shaaban:
generating, by the computer, a loss for a first deepfake detector using a loss function, the loss indicating a distance between the first risk score as generated for the voiced-speech segment and an expected first risk score indicated by a training label associated with the input audio signal (There are three different varieties of speech: voiced, unvoiced, and silent. Voiced speech contains a limited amount of energy and a periodic sequence of impulses, whereas unvoiced speech consists of random, no periodic noise-like patterns…., sec. II; pg. 132662, sec. 2); and
updating, by the computer, one or more parameters of the first deepfake detector based upon the loss (pg. 132662, sec. 2; pg. 132673, training as involving updating weights of models)
It would have been obvious to one of ordinary skill in the art before the effective filing of the instant invention to combine the teachings of Shaaban with the method of Suthokumar in arriving at the missing features of Suthokumar, because such combination would have resulted in facilitating a model’s ability to differentiate between authentic and counterfeit audio (Shaaban, pg. 132662)
Per claim 10, Suthokumar discloses the method according to claim 1,
Suthokumar does not explicitly disclose generating, by the computer, a loss for a second deepfake detector using a loss function, the loss indicating a distance between the second risk score as generated for the unvoiced-speech segment and an expected second risk score indicated by a training label associated with the input audio signal or updating, by the computer, one or more parameters of the second deepfake detector based upon the loss
However, these features are suggested by Shaaban:
generating, by the computer, a loss for a second deepfake detector using a loss function, the loss indicating a distance between the second risk score as generated for the unvoiced-speech segment and an expected second risk score indicated by a training label associated with the input audio signal (There are three different varieties of speech: voiced, unvoiced, and silent. Voiced speech contains a limited amount of energy and a periodic sequence of impulses, whereas unvoiced speech consists of random, no periodic noise-like patterns…., sec. II; pg. 132662, sec. 2); and
updating, by the computer, one or more parameters of the second deepfake detector based upon the loss (pg. 132662, sec. 2; pg. 132673, training as involving updating weights of models)
It would have been obvious to one of ordinary skill in the art before the effective filing of the instant invention to combine the teachings of Shaaban with the method of Suthokumar in arriving at the missing features of Suthokumar, because such combination would have resulted in facilitating a model’s ability to differentiate between authentic and counterfeit audio (Shaaban, pg. 132662).
Per claim 18, Suthokumar discloses the system according to claim 11,
System claim 18 and method claim 8 are related as system and the method of using same, with each claimed element's function corresponding to the claimed method step. Accordingly claim 18 is similarly rejected under the same rationale as applied above with respect to claim 8.
Per claim 19, Suthokumar discloses the system according to claim 11,
System claim 19 and method claim 9 are related as system and the method of using same, with each claimed element's function corresponding to the claimed method step. Accordingly claim 19 is similarly rejected under the same rationale as applied above with respect to claim 9.
Per claim 20, Suthokumar discloses the system according to claim 11,
System claim 20 and method claim 10 are related as system and the method of using same, with each claimed element's function corresponding to the claimed method step. Accordingly claim 20 is similarly rejected under the same rationale as applied above with respect to claim 10.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. See PTO 892 form
Any inquiry concerning this communication or earlier communications from the examiner should be directed to OLUJIMI A ADESANYA whose telephone number is (571)270-3307. The examiner can normally be reached Monday-Friday 8:30-5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richemond Dorvil can be reached at 571-272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/OLUJIMI A ADESANYA/Primary Examiner, Art Unit 2658