Prosecution Insights
Last updated: August 17, 2026
Application No. 18/457,221

UNSUPERVISED ALIGNMENT FOR TEXT TO SPEECH SYNTHESIS USING NEURAL NETWORKS

Non-Final OA §101
Filed
Aug 28, 2023
Priority
Oct 07, 2021 — continuation of 11/769,481 +2 more
Examiner
PATEL, SHREYANS A
Art Unit
2659
Tech Center
2600 — Communications
Assignee
NVIDIA Corporation
OA Round
3 (Non-Final)
89%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
97%
With Interview

Examiner Intelligence

Grants 89% — above average
89%
Career Allowance Rate
364 granted / 411 resolved
+26.6% vs TC avg
Moderate +8% lift
Without
With
+8.5%
Interview Lift
resolved cases with interview
Fast prosecutor
2y 0m
Avg Prosecution
34 currently pending
Career history
457
Total Applications
across all art units

Statute-Specific Performance

§101
26.4%
-13.6% vs TC avg
§103
41.1%
+1.1% vs TC avg
§102
21.4%
-18.6% vs TC avg
§112
1.6%
-38.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 411 resolved cases

Office Action

§101
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Applicant's arguments with respect to 35 U.S.C. 103 rejection of claims 1, 9 and 16 have been considered and found persuasive, and the rejection has been withdrawn. However, the Applicant must overcome the 101 Abstract Idea set forth. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101. Claims 1, 9 and 16, under step1/step 2A, prong one, claim is directed to the abstract idea of managing and using information to train and operate a TTS system. The claimed steps involve updating model parameters using two sets of data, removing one of the data distributions after training, and then using the remaining distribution during inference. These steps describe collecting, organizing, modifying, and using information according to a set of rules for improving a model. Such data evaluation, manipulation, and decision making are mental processes that can be performed conceptually and therefore fall within the category of abstract ideas. Although the claim is implemented on a computer, the focus of the claim is on the logical process of managing training and inference data rather than on any specific technological improvement to computer functionality. Under step 2A, prong two, the claim does not integrate the abstract idea into a practical application. The additional elements do not integrate the abstract idea into a practical application. The claim generically recites “a computer implemented method” and “one or more text to speech systems,” but it does not disclose any particular machine architecture, specialized hardware, or specific improvement to the operation of the TTS system itself. The claim does not describe how the parameters are updated, how the distributions are represented or removed, or how the inferencing distribution is technically implemented. Instead, the computer merely performs the claimed data processing operations as a tool to carry out the abstract idea. As a result, the claim does not improve the functioning of a computer or another technology and instead applies conventional computing components to perform information processing. Under Step 2B, the claim does not include an inventive concept. Viewed individually and as an ordered combination, the additional claim elements do not amount to significantly more than the abstract idea itself. The recited functions of updating parameters, removing a distribution, and performing inference are stated at a high level of generality without specifying any unconventional algorithm or technological mechanism for accomplishing them. The claim therefore merely instructs a generic computer and a generic TTS system to perform routine data processing operations in a particular sequence. Because the claim does not recite any non-conventional technical solution or other inventive concept that transforms the abstract idea into patent eligible subject matter, it is directed to ineligible subject matter under 101 abstract idea. The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the claims are (i) mere instructions to implement the idea on a computer, and/or (ii) recitation of generic computer structure that serves to perform generic computer functions that are well-understood, routine, and conventional activities previously known to the pertinent industry. Viewed as a whole, these additional claim element(s) do not provide meaningful limitation(s) to transform the abstract idea into a patent eligible application of the abstract idea such that the claim(s) amounts to significantly more than the abstract idea itself. Therefore, the claim(s) are rejected under 35 U.S.C. 101 as being directed to non-statutory subject matter. There is further no improvement to the computing device. Dependent claims 2-8, 10-15 and 17-20 further recite an abstract idea performable by a human and do not amount to significantly more than the abstract idea as they do not provide steps other than what is conventionally known. Claims 2-8, 10-15 and 17-20 are rejected under 101 abstract idea because they merely add further abstract data processing steps to the abstract idea recited in the claims. The additional limitations, such as generating synthesized audio samples, modifying audio features, including original audio samples, assigning identifiers, generating an audio clip during inference, using a phoneme distribution, and determining phoneme durations, are all forms of creating, organizing, labeling, selecting, or processing data using generic computer functions. These limitations do not improve the operation of a computer or TTS system, nor do they recite any unconventional technological solution or inventive concept. Instead, they simply refine or apply the underlying abstract idea using routine and conventional computer operations. Allowable Subject Matter Claims 1-20 would be allowable if the claimed can overcome the 101 Abstract Idea. The following is a statement of reasons for the indication of allowable subject matter: Kim et al. (“Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech”; 2021) teaches updating the parameters of a TTS system using both real audio samples and synthesized audio samples through adversarial learning. Specifically, the discriminator distinguishes between generated waveforms and ground-truth waveforms, while the generator is trained using adversarial and feature-matching losses based on those two distributions (see eqs. 8-11, sections 2.3 and 2.4). Kim further explains that the posterior encoder and discriminator are used only during training, while inference uses the trained model (see Figs. 1a-b, Section 2.5). Binkowski et al. (“High Fidelity Speech Synthesis with Adversarial Networks”, 2019) teaches updating TTS model parameters using adversarial training with a first distribution of real speech and a second distribution of synthesized speech. Binkowski describes a feed-forward generator trained against an ensemble of discriminators that distinguish real and generated audio, with the discriminators evaluating both realism and correspondence to the input text (see sections 3.2-3.4). During inference, only the trained generator is retained while the discriminators are discarded. The difference between the prior art and the claimed invention is that Kim nor Binkowski explicitly teach removing, after the updating of the one or more parameters of the one or more TTS systems, the second distribution to form an inferencing distribution. Therefore, it would not have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the teachings of Kim and Binkowski to include removing, after the updating of the one or more parameters of the one or more TTS systems, the second distribution to form an inferencing distribution. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHREYANS A PATEL whose telephone number is (571)270-0689. The examiner can normally be reached Monday-Friday 8am-5pm PST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre Desir can be reached at 571-272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. SHREYANS A. PATEL Primary Examiner Art Unit 2653 /SHREYANS A PATEL/Examiner, Art Unit 2659
Read full office action

Prosecution Timeline

Show 1 earlier event
Oct 28, 2025
Non-Final Rejection mailed — §101
Jan 15, 2026
Applicant Interview (Telephonic)
Jan 15, 2026
Examiner Interview Summary
Jan 26, 2026
Response Filed
Apr 22, 2026
Final Rejection mailed — §101
Jun 11, 2026
Request for Continued Examination
Jun 15, 2026
Response after Non-Final Action
Aug 04, 2026
Non-Final Rejection mailed — §101 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12659658
ACOUSTIC ECHO CANCELLATION SYSTEM AND ASSOCIATED METHOD
2y 10m to grant Granted Jun 16, 2026
Patent 12646496
METHODS AND SYSTEMS OF TEXT-CONDITIONED AUDIO-VISUAL SPEECH GENERATION WITH MULTI-MODAL LATENT DIFFUSION MODELS
2y 2m to grant Granted Jun 02, 2026
Patent 12608559
METHOD AND SYSTEM FOR ENHANCING A MUTIMODAL INPUT CONTENT
3y 0m to grant Granted Apr 21, 2026
Patent 12609128
METHOD FOR IMPROVING FAR-FIELD SPEECH INTERACTION PERFORMANCE, AND FAR-FIELD SPEECH INTERACTION SYSTEM
2y 0m to grant Granted Apr 21, 2026
Patent 12586597
ENHANCED AUDIO FILE GENERATOR
3y 6m to grant Granted Mar 24, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
89%
Grant Probability
97%
With Interview (+8.5%)
2y 0m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 411 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month