Prosecution Insights
Last updated: October 02, 2026
Application No. 18/661,313

DETECTING SYNTHETIC SPEECH

Non-Final OA §103
Filed
May 10, 2024
Priority
May 11, 2023 — provisional 63/465,740
Examiner
HE, JIALONG
Art Unit
2659
Tech Center
2600 — Communications
Assignee
Sri International
OA Round
3 (Non-Final)
81%
Grant Probability
Favorable
3-4
OA Rounds
7m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 81% — above average
81%
Career Allowance Rate
759 granted / 932 resolved
+19.4% vs TC avg
Strong +33% interview lift
Without
With
+32.9%
Interview Lift
resolved cases with interview
Typical timeline
3y 0m
Avg Prosecution
22 currently pending
Career history
948
Total Applications
across all art units

Statute-Specific Performance

§101
14.2%
-25.8% vs TC avg
§103
41.7%
+1.7% vs TC avg
§102
15.2%
-24.8% vs TC avg
§112
20.9%
-19.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 932 resolved cases

Office Action

§103
DETAILED ACTION The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Request for Continued Examination A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 05/12/2026 has been entered. Response to Amendments and Arguments Regarding outstanding rejections under 35 U.S.C. §103, applicant amended independent claims by adding more limitations into independent claims. Applicant also filed a declaration under 37 C.F.R. 1.130(a) to disqualify a cited Rahman reference (co-authored by inventors) by stating: PNG media_image1.png 772 1354 media_image1.png Greyscale Applicant stated that remaining authors contributed to other portions of the Rahman reference not relied upon in making the rejection. The Rahman reference was cited for teaching “segment scores” (in dependent claim 7) and “an utterance level score” (in a previously presented, now cancelled, claim 8). Applicant amended independent claims by adding limitations in an alternative language using “include at least ONE of … OR…”. By reviewing previously cited primary reference (Chen et al., US PG Pub. 2021/0233541) and the secondary reference to Zhang, the examiner believed that Chen in view of Zhang meets the added limitations recited in alternative language. For example, Chen discloses detecting if an inbound audio contains spoofed voice (i.e., synthesized voice) by comparing a score with a threshold (Chen, [0010], [0045], Fig. 7). Chen discloses one alternative: “an utterance level score representing a likelihood the audio clip includes the injected synthetic speech” (Chen, [0045]). In addition, Zhang discloses detecting partial spoof speeches (i.e., a portion of an audio is synthesized speech) by comparing frame level scores and utterance level scores (Zhang, section IV, pages 4-5, computing utterance-level score and segment-level score). Zhang meets both alternative limitations recited using “ONE of … OR”. In the following rejection, the Rahman reference is no longer relied upon because Chen in view Zhang meets the added claim limitations in alternative language. Applicant arguments regarding amended claims are considered. The arguments are not persuasive. In addition, applicant added a new dependent claim 22 to further limit “the utterance level score”. Since independent claims recite limitations in alternative language “at least ONE of … OR…”, the cited references one need to meet ONE alternative. Further limiting an none-addressed alternative does not affect rejection to a claimed scope. Information Disclosure Statement The information disclosure statement (IDS) submitted on 06/10/2026 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim Rejections - 35 USC § 103 Claims 1-4, 7, 9-12, 15 and 17-22 are rejected under 35 U.S.C. 103 as being unpatentable over Chen et al. (US PG Pub. 2021/0233541, referred to as Chen) in view of Zhang (“The PartialSpoof Database and Countermeasures for the Detection of Short Fake Speech Segments Embedded in an Utterance”, published in April, 2022, a reference submitted by the applicant in an IDS, referred to Zhang). Chen discloses a neural network based spoofing detection technique to detect deepfake speech generated by using speech synthesis technique (Chen, [0009-0011], [0036-0039], Fig. 6). Chen discloses generating feature embedding (Chen, [0009], [0026], Fig. 6, #606, Fig. 7). Chen further discloses detecting whether an inbound call is from a genuine speaker or spoofed by using speech synthesis technique (Chen, [0045], [0059-0061], [0070], Fig. 7, #718). Zhang discloses detecting partially spoofed speech generated by inserting synthesized speech into a genuine speech (Zhang, section I and section III, inserting a synthesized word “NOT” to change a meaning of the original utterance). Regarding claims 1, 10 and 18, Chen discloses a method, a system and a computer readable storage medium for detecting synthetic speech in frames of an audio clip (Chen [0025], [0107], Fig. 6, a computer implemented method / system for detecting deepfake synthesized audios using neural network), comprising: processing, by a machine learning system, an audio clip to generate a plurality of speech artifact embeddings based on a plurality of synthetic speech artifact features (Chen, [0026-0028], [0037], [0053], Fig. 6, #606; generating spoofing embeddings from artifacts of audio/speech frames); computing, by the machine learning system, one or more scores based on the plurality of speech artifact embeddings (Chen, [0010], [0045], Fig. 7, #718; calculating spoofing scores using neural network models); wherein the one or more scores include at least ONE of: a segment score for a speech artifact embedding of the plurality of speech artifact embeddings, the segment score representing a likelihood that a corresponding frame of the audio clip includes the injected synthetic speech, OR an utterance level score representing a likelihood the audio clip includes the injected synthetic speech (Chen, [0045], [0059-0061], generating score from inbound audio and comparing the score with a threshold to determine if inbound audio is spoofed audio / synthesized audio; Note, the reference only need to teach ONE alternative recited using ONE of … OR); determining, by the machine learning system, based on the one or more scores, whether one or more frames of the audio clip include synthetic speech (Chen, [0045], [0070], determining whether audio signals are genuine or spoofed by from synthesis based on spoofing or similarity scores); and outputting an indication of whether the one or more frames of the audio clip include synthetic speech (Chen, [0009], [0087], indicating whether audio is from genuine speaker or spoofed from synthesized speech). Chen discloses detecting spoofed speech generated by speech synthesizers based on detected spoof characteristics (Chen, Summary of the invention, Fig. 4). Chen does not explicitly discloses detecting a partially spoofed speech. Therefore, Chen does not explicitly disclose the newly added limitations: “wherein the plurality of synthetic speech artifact features have been created by one or more synthetic speech generators injecting synthetic speech into the audio clip”. Zhang discloses detecting partially spoofed speech that were generated by inserting certain words generated by using a speech synthesizer to change meaning of original utterance (Zhang, section 1, Introduction, section III and Fig. 1). In addition, Zhang further discloses calculating utterance level score and segment level score for detecting partial spoofed audio (Zhang, page 4): PNG media_image2.png 274 652 media_image2.png Greyscale Both Chen and Zhang are dealing with detecting spoofed audio using a neural network, it would have been obvious to a person having ordinary skill in the art at the time the invention was filed to combine Chen’s teaching with Zhang’s teaching to detect partially spoofed speech and obtain a training database by labelling which section / frames are real speech and which sections / frames are spoofed speech from speech synthesis, and calculating utterance level score / segment level score to detect partial spoof audio. One having ordinary skill in the art would have been motivated to make such a modification to improve accuracy of spoof detection when using speaker verification (Zhang, Abstract, Section V). In addition, all the claimed elements were known in the prior art and one skilled in the art could have combined the elements as claimed by known methods, and in the combination each element merely would have performed the same function as it did separately. “A combination of familiar elements according to known methods is likely to be obvious when it does no more than yield predictable results.” KSR, 550 U.S. ___, 82 USPQ2d at 1395 (2007). One of ordinary skill in the art would have recognized that the results of the combination were predictable. Regarding claims 2, 11 and 19, Chen in view of Zhang further discloses: extracting the plurality of synthetic speech artifact features from frames of the audio clip (Chen, [0009-0011], [0026], [0038], processing audio frames to extract artifacts and training neural network based spoofing audio detection), wherein the synthetic speech artifact features include at least ONE of artifacts, distortions, or degradations that are associated with one or more synthetic speech generators and that are included in the audio clip (Chen, [0026], [0037-0038], spoofing artifact features such as degradation, distortion in synthesized spoofing speech). Regarding claims 3-4 and 12, limitations recited in these dependent claims are related to preparing training data for training a neural network-based spoofing speech detection. The claimed “determining one or more boundaries in the mapping … based on label” is related to labelling training data. Chen discloses training a neural network for detecting synthesized / spoofed speech (Chen, [0038-0039], Fig. 2 and Fig. 3). Chen further discloses preparing a training audio with labels (Chen, [0047], [0070], labelling speech as genuine or spoofed audio data). Although Chen implicitly discloses all features in dependent claims 3-4 and 12, the examiner further cites Zhang, which discloses more details about labelling training audio data to indicate segments related real speech or spoofed synthesized speech (Zhang, section III, creating partial spoof database, Fig. 1). Regarding claims 7 and 15, Chen in view of Zhang further discloses: wherein each speech artifact embedding of the plurality of speech artifact embeddings corresponds to a different frame of the audio clip (Chen, [0037-0038], [0053], [0065], parsing an audio signal into frames / sub-frames, applying short-time Fourier transform SFT; Fig. 5 shows embedding), Zhang discloses detecting partial spoofed speech by calculating both utterance level scores and segment level scores. Zhang discloses the one or more scores further include corresponding segment score for each of the plurality of speech artifact embeddings representing likelihoods that corresponding frames f the audio clip includes at least a portion of the injected synthetic speech (Zhang, Zhang, page 4). Regarding claims 9, 17 and 20, Chen in view of Zhang further discloses: wherein outputting the indication comprises: responsive to determining a score of the one or more scores satisfies a threshold, outputting an indication that a frame of the one or more frames that corresponds to the score includes synthetic speech (Chen, [0045], [0059-0061], [0087], comparing scores with spoofing threshold to determine whether audio is from genuine speaker or from spoofed synthesized speech). Regarding claim 21, Chen in view of Zhang further discloses wherein the audio clip originally includes authentic speech, and wherein the one or more frames of the audio clip include the injected synthetic speech (Zhang, section I, Introduction, inserting a synthesized “NOT” by a speech synthesizer to change meaning of original utterance; See Fig. 1). Regarding claim 22, this limitation further limit ONE alternative “utterance level score” using “OR”. The cited references only need to teach ONE alternative. Further imitating an not addressed different alternative does not affect rejection to a claimed scope. Claims 5-6 and 13-14 are rejected under 35 U.S.C. §103 as being unpatentable over Chen in view of Zhang and further in view of Castan (“Speaker-targeted Synthetic Speech Detection”, published in June, 2022, a reference submitted by the applicant in an IDS, referred to as Castan). Regarding claims 5 and 13, Chen in view of Zhang discloses training a neural network model to detect spoofed audio (Chen, [0023], [0038-0039], Fig. 2 and Fig. 3). Chen further discloses labelling data to indicate which portion containing speech (Chen, [0047], [0069], [0092]). Chen does not explicitly disclose removing non-speech information from training data. Castan discloses a neural network structure for detecting spoofing speech in multimedia data (Castan, Abstract, xResNet-PLDA system). Castan further discloses discarding silence frames in audio data (Castan, Section 4.1). Chen in view of Zhang and Castan are dealing with detecting spoofed / synthesized audio using a neural network, it would have been obvious to a person having ordinary skill in the art at the time the invention was filed to modify Chen’s teaching with Castan’s teaching to discard silence frames from speech data. One having ordinary skill in the art would have been motivated to make such a modification to reduce error and improve performance (Castan, section 5, using a network structure with xResNet and PLDA outperforms the baseline system for detecting fake speech). In addition, all the claimed elements were known in the prior art and one skilled in the art could have combined the elements as claimed by known methods, and in the combination each element merely would have performed the same function as it did separately. “A combination of familiar elements according to known methods is likely to be obvious when it does no more than yield predictable. Regarding claims 6 and 14, Chen in view of Zhang and Castan further discloses wherein computing the one or more scores comprises computing one or more log-likelihood ratios by at least comparing the plurality of speech artifact embeddings to a plurality of enrollment embeddings (Castan, section 2.2.2, we applied a common and simple solution using a discriminatively trained affine transformation from scores to log-likelihood ratios (LLRs.), see Fig. 2., comparing log likelihood ration), wherein each of the plurality of enrollment embeddings are associated with authentic speech (Chen, [0036], [0038-0039], Fig. 4, Castan, Section 3.3). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to JIALONG HE whose telephone number is (571)270-5359. The examiner can normally be reached on Monday-Thursday, 7:00AM-4:30PM, ALT. Fridays, EST. If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Pierre Desir can be reached on (571) 272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JIALONG HE/Primary Examiner, Art Unit 2659
Read full office action

Prosecution Timeline

Show 7 earlier events
Apr 16, 2026
Interview Requested
Apr 30, 2026
Applicant Interview (Telephonic)
Apr 30, 2026
Examiner Interview Summary
May 12, 2026
Response after Non-Final Action
Jun 10, 2026
Request for Continued Examination
Jun 15, 2026
Response after Non-Final Action
Jul 14, 2026
Examiner Interview (Telephonic)
Aug 12, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749476
SENTIMENT-BASED ADAPTATION OF DIGITAL HUMAN RESPONSES
2y 5m to grant Granted Sep 29, 2026
Patent 12738264
Semantic Segmentation With Language Models For Long-Form Automatic Speech Recognition
2y 6m to grant Granted Sep 15, 2026
Patent 12731579
COMMAND GENERATION SYSTEM AND METHOD OF ISSUING COMMANDS
2y 11m to grant Granted Sep 08, 2026
Patent 12730967
TRANSFORMER-BASED HYBRID RECOMMENDATION MODEL WITH CONTEXTUAL FEATURE SUPPORT
2y 10m to grant Granted Sep 08, 2026
Patent 12718013
METHOD AND ELECTRONIC DEVICE FOR PROCESSING USER UTTERANCE BASED ON LANGUAGE MODEL
2y 7m to grant Granted Aug 25, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
81%
Grant Probability
99%
With Interview (+32.9%)
3y 0m (~7m remaining)
Median Time to Grant
High
PTA Risk
Based on 932 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month