Prosecution Insights
Last updated: October 01, 2026
Application No. 18/431,211

NEURAL NETWORKS TO GENERATE SPEECH

Non-Final OA §102§103
Filed
Feb 02, 2024
Examiner
CAUDLE, PENNY LOUISE
Art Unit
2657
Tech Center
2600 — Communications
Assignee
NVIDIA Corporation
OA Round
3 (Non-Final)
70%
Grant Probability
Favorable
3-4
OA Rounds
3m
Est. Remaining
85%
With Interview

Examiner Intelligence

Grants 70% — above average
70%
Career Allowance Rate
59 granted / 84 resolved
+8.2% vs TC avg
Moderate +14% lift
Without
With
+14.5%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
18 currently pending
Career history
97
Total Applications
across all art units

Statute-Specific Performance

§101
21.3%
-18.7% vs TC avg
§103
47.1%
+7.1% vs TC avg
§102
15.0%
-25.0% vs TC avg
§112
15.9%
-24.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 84 resolved cases

Office Action

§102 §103
DETAILED ACTION This examination is in response to the communication filed on 07/20/2026. Claims 1-20 are currently pending, wherein claims 1-5 and 7-19 have been amended. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendments/Arguments Applicant's arguments with respect to the rejection of claims 1-20 have been fully considered but they are not persuasive. Applicant's arguments fail to comply with 37 CFR 1.111(b) because they amount to a general allegation that the claims define a patentable invention without specifically pointing out how the language of the claims patentably distinguishes them from the references. Applicant argues on page 9 of the Response that independent claim 1, and for similar reasons independent claims 9 and 15, are allowable over Fejzo because Fejzo fails to disclose that the key parameters “separately generation (i) features from first-environment speech audio and (ii) characteristics from second-environment reference audio data, and then using one or more neural networks to modify the generated features based on the generated characteristics”. In addition, Applicant argues that using a neural network to generate key parameters that are used to transform an input audio signal to output an audio signal that has target audio characteristics as taught by Fejzo is different from “generating one or more features from the audio data…generating one or more characteristics from the reference audio…using one or more neural networks to modify the one or more features…” as recited in amended independent claims 1, 9, and 15. The Examiner, respectfully disagrees. First, under a broadest reasonable interpretation “generat[ing] one or more features” includes downsampling (see detailed rejection below). Therefore, Fejzo’s teaching of pre-processing the input signal by downsampling and then modifying the pre-processed input signal is equivalent to “generating one or more features” as claimed (see detailed rejection below). Second, under a broadest reasonable interpretation “generating one or more characteristics” includes determining characteristics of a desired signal. Therefore, using a neural network to determine key parameters for transforming an input speech signal to match reference signal as taught by Fejzo is equivalent to using a neural network to modify features as recited independent claim 1, 9 and 15. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1-3, 5-17 and 19-20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Fejzo et al. (WO 2022/025922 A1; herein “Fejzo”). Regarding claims 1, 9 and 15, Fejzo teaches one or more processors (Fig. 9, processor 916), comprising circuitry, a method (Fig. 8) and system comprising one or more processors (Fig. 9, processor 916) to: obtain audio data of speech associated with a first environment and reference audio data associated with a second environment (Fig. 8, step 802 teaches the system “receive[s] input audio and target audio having target audio characteristics”; under a broadest reasonable interpretation, obtaining includes receiving and the input audio is interpreted as first audio of speech and the target audio is interpreted as a reference audio); generate one or more features from the audio data of the speech associated with the first environment (Under a broadest reasonable interpretation “generate one or more features” includes down-sampling and/or low pass filtering the input signal. This interpretation is supported by ¶¶[0096]-0097] of the Specification which teaches that the waveform encoder 210 is a module that generates spectral features and that the waveform encoder include a downsampling module 222 that applies filters and performs resampling. Fejzo ¶[0024] teaches “the input signal and the target signal may each be pre-processed … Example pre-processing operations that may be performed on the input signal…include one or more of: resampling (e.g., down-sampling or up-sampling); direct current (DC) filtering to remove low frequencies…pre-emphasis filtering to compensate for a spectral tilt in the input signal; and/or adjusting gain such that the input signal is normalized before its subsequent signal transformation”); generate one or more characteristics from the reference audio data (¶[0020] teaches “…key parameters KP configure the audio synthesizer ML model to perform the signal transformation such that spectral or temporal characteristics of the output signal match corresponding desired/target spectral or temporal characteristics of the target signal” and ¶[0043] teaches “key estimating operation 306 performs temporal analysis (i.e., time-domain analysis) of a least one of the target signal…The temporal analysis produces target key parameters as parameters that compactly represent temporal evolution in a given frame…generally referred to as ‘temporal amplitude’ characteristics”); use one or more neural networks (Fig. 1, Key Generator Model 102 and Audio Synthesis Model 104 ) to modify the one or more features generated from the audio data of the speech associated with the first environment (Figs. 6 and 7 show that the audio synthesis ML model modifies that the pre-processed input signal, e.g., downsampled signal) based, at least in part, on the one or more characteristics from the reference audio data (Fig. 8, Key frame parameters KP); and generate modified audio data of speech having the one or more characteristics associated with the second environment based, at least in part, on the modified one or more features (¶[0018] teaches “System 100 includes trained key generator ML model 102…and a trained audio synthesis ML model 104…key generator 102 receives key generation data that may include at least an input signal and/or a target or desired signal…generates a set of transform parameters KP…Audio synthesizer 104 performs a desired signal transformation of the input signal based on key parameters KP, to produce an output signal having an output signal characteristic similar to or that matches the desired/target signal characteristic of the target signal”). Regarding claims 2, 10 and 16, Fejzo teaches all of the elements of claims 1, 9 and 15 (see detailed element mapping above). In addition, Fejzo further teaches the reference audio data includes one or more speech signals (¶[0023] teaches “In the audio context, the target signal may be a speech or audio signal”). Regarding claims 3, 11 and 17, Fejzo teaches all of the elements of claims 1, 9 and 15 (see detailed element mapping above). In addition, Fejzo further teaches the one or more circuits are further to use the one or more neural networks to generate one or more spectral features based, at least in part, on reference audio data associated with the second environment (¶[0018] teaches “Key parameters KP parameterize or represent a desired/target signal characteristic of the target signal, such as a spectral/frequency-based characteristic…” ). Regarding claims 5, 13 and 19, Fejzo teaches all of the elements of claims 1, 9 and 15 (see detailed element mapping above). In addition, Fejzo further teaches the one or more circuits are further configured to use the one or more neural networks to generate context information based, at least in part, on the reference audio data associated with the second environment (Based on ¶[0079] of the Specification, context information is interpreted as information that “indicates desired characteristics of a high-quality audio”; ¶[0018] teaches “Key parameters KP parameterize or represent a desired/target signal characteristic of the target signal, such as a spectral/frequency-based characteristic…” and ¶[0015] teaches “given a low resolution/bandwidth representation and a high resolution/bandwidth representation of an audio signal, the key generator ML model generates a size-constrained set of metadata, e.g., key parameters, for guiding the audio synthesis ML model.” Thus the target/desired signal from which the key parameters are generated is a higher resolution/quality signal). Regarding claims 6 and 20, Fejzo teaches all of the elements of claims 1 and 15 (see detailed element mapping above). In addition, Fejzo further teaches the one or more circuits are further configured to use the one or more neural networks to fuse one or more spectral features and one or more waveform features (¶[0039] teaches “Examples of target key parameters include a line spectral frequency, (LSF) key, a harmonic key, and a temporal envelope key, as described below…such that the key generator learns to generate key parameters KPT that approximate the target key parameters” Thus, Fejzo teaches fusing/combining both spectral (e.g., LSF) and waveform (e.g., temporal envelope) features). Regarding claims 7 and 14, Fejzo teaches all of the elements of claims 1, 9 (see detailed element mapping above). In addition, Fejzo further teaches the one or more circuits are further configured to use the one or more neural networks fuse one or more spectral features and one or more waveform features generated from the audio data of the speech associated with the first environment (¶[0027] teaches the “input signal pre-processing operations include: resampling; DC filtering to remove low frequencies, e.g., below 50 HZ; pre-emphasis filtering to compensate for a spectral tilt in the input signal; and/or adjusting gain such that the input signal is normalized before a subsequent signal transformation” Thus, Fejzo teaches fusing/combining both spectral (removing low frequencies) and waveform (adjusting gain) features); and use context information generated from the reference audio data to modify the one or more fused features (¶[0018] teaches “System 100 includes trained key generator ML model 102…and a trained audio synthesis ML model 104…key generator 102 receives key generation data that may include at least an input signal and/or a target or desired signal…generates a set of transform parameters KP…Audio synthesizer 104 performs a desired signal transformation of the input signal based on key parameters KP, to produce an output signal having an output signal characteristic similar to or that matches the desired/target signal characteristic of the target signal” the key parameters are interpreted as context information generated from the reference audio. Furthermore, as shown in Figs. 6 and 7, the audio synthesis ML model operations on the fused, i.e., pre-processed, input signal). Regarding claim 8, Fejzo teaches all of the elements of claim 1 (see detailed element mapping above). In addition, Fejzo further teaches the one or more circuits are further configured to determine a down sampling rate to perform down sampling of reference audio data associated with the second environment (¶[0024] teaches “…the input signal and the target signal may be pre-processed to produce…a pre-processed target signal…Example pre-processing operations that may be performed…include one or more of: resampling (e.g., down-sampling or up-sampling)…”) Regarding claim 12, Fejzo teaches all of the elements of claim 9 (see detailed element mapping above). In addition, Fejzo further teaches generating, using the one or more neural networks, one or more waveform features based, at least in part, on the reference audio data associated with the second environment (¶[0018] teaches “Key parameters KP parameterize or represent a desired/target signal characteristic of the target signal, such as … temporal/time-based characteristic of the target signal”). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or non-obviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 4 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Fejzo as applied to claims 1 and 15 above, and further in view of Buskies et al. (US 2015/0013528 A1; herein “Buskies”). Regarding claims 4 and 18, Fejzo teaches all of the elements of claims 1 and 15 (see detailed element mapping above). In addition, Fejzo further teaches the one or more circuits are further to use the one or more neural networks to generate one or more waveform features based, at least in part, on the reference audio data associated with the second environment (¶[0018] teaches “Key parameters KP parameterize or represent a desired/target signal characteristic of the target signal, such as … temporal/time-based characteristic of the target signal”). In addition, ¶[0039] of Fejzo teaches fusing/combining spectral and waveform features. However, Fejzo fails to specifically disclose the one or more waveform features include phase information. Buskies teaches a system and method for audio signal enhancement/processing that analyzes an audio signal waveform to determine basic accent patterns in the waveform. More specifically, Buskies teaches analyzing the waveform characteristics, including “audio peaks, patterns, frequency characteristics, phase characteristics, timing characteristics (e.g., tempo), and the like” (Buskies, ¶[0113]). In addition, Buskies teaches the one or more waveform features include phase information (¶[0111] teaches “The various types of audio waveforms and their pertinent characteristics (e.g., frequency, amplitude, phase etc.) would be understood by those of ordinary skill in the art”). Fejzo differs for the claimed invention, as defined by claims 4 and 18, in that Fejzo fails to explicitly disclose that the temporal/time-based characteristic of the target signal includes phase information. Phase information is a well-known pertinent characteristic of audio waveform signals as taught in Buskies. Therefore, it would have been obvious to one having ordinary skill in the art to modify the temporal/time-based key parameters extracted for the target signal as taught by Fejzo to include phase information as suggested by Buskies as it merely constitutes the combination of known processes to achieve the predictable result to determining pertinent characteristics of the target audio signal for audio enhancement. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to PENNY L CAUDLE whose telephone number is (703)756-1432. The examiner can normally be reached M-Th 8:00 am to 5:00 pm eastern. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at 571-272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /PENNY L CAUDLE/Examiner, Art Unit 2657 /DANIEL C WASHBURN/Supervisory Patent Examiner, Art Unit 2657
Read full office action

Prosecution Timeline

Show 4 earlier events
Jan 27, 2026
Applicant Interview (Telephonic)
Mar 09, 2026
Response Filed
Apr 21, 2026
Final Rejection mailed — §102, §103
Jun 29, 2026
Examiner Interview Summary
Jun 29, 2026
Applicant Interview (Telephonic)
Jul 20, 2026
Request for Continued Examination
Jul 23, 2026
Response after Non-Final Action
Jul 28, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12731587
VISUAL SPEECH RECOGNITION FOR DIGITAL VIDEOS UTILIZING GENERATIVE ADVERSARIAL LEARNING
4y 7m to grant Granted Sep 08, 2026
Patent 12724812
AUTOMATIC GENERATION OF HANDOUTS FROM MULTI-MODAL DOCUMENTS
2y 8m to grant Granted Sep 01, 2026
Patent 12718829
SYSTEMS AND METHODS FOR VOICE RECEPTION AND DETECTION
3y 10m to grant Granted Aug 25, 2026
Patent 12711313
HUMAN-MACHINE COLLABORATIVE CONVERSATION INTERACTION SYSTEM AND METHOD
3y 2m to grant Granted Aug 18, 2026
Patent 12706091
PRONUNCIATION-AWARE EMBEDDING GENERATION FOR CONVERSATIONAL AI SYSTEMS AND APPLICATIONS
2y 6m to grant Granted Aug 11, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
70%
Grant Probability
85%
With Interview (+14.5%)
2y 11m (~3m remaining)
Median Time to Grant
High
PTA Risk
Based on 84 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month