Prosecution Insights
Last updated: October 01, 2026
Application No. 18/823,176

USER-CUSTOMIZED SYNTHETIC VOICE

Final Rejection §101
Filed
Sep 03, 2024
Priority
Sep 29, 2022 — continuation of 12/087,270
Examiner
PATEL, SHREYANS A
Art Unit
2659
Tech Center
2600 — Communications
Assignee
Amazon Technologies Inc.
OA Round
2 (Final)
89%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
97%
With Interview

Examiner Intelligence

Grants 89% — above average
89%
Career Allowance Rate
368 granted / 415 resolved
+26.7% vs TC avg
Moderate +8% lift
Without
With
+8.5%
Interview Lift
resolved cases with interview
Fast prosecutor
2y 0m
Avg Prosecution
30 currently pending
Career history
464
Total Applications
across all art units

Statute-Specific Performance

§101
28.5%
-11.5% vs TC avg
§103
40.4%
+0.4% vs TC avg
§102
20.4%
-19.6% vs TC avg
§112
1.5%
-38.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 415 resolved cases

Office Action

§101
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Applicant's arguments with respect to 35 U.S.C. 101 in regard to claims 1-20 have been considered, however are not found to be persuasive due to the following reasons. With regard to claims 1 and 11, the claims do not recite a specific technical improvements to speech synthesis technology. The claims describe the desired results – using NL input and feedback to generate more satisfactory voice – but do not identify a particular model, voice embedding structure, signal processing technique, or algorithm that improves how a computer generates speech. Improving the accuracy of a voice relative to a user’s subjective preference and making customization easier are improvement to user experience, not improvements to computer or speech synthesis functionality. The asserted privacy improvement is also not reflected in the claims. The claims merely associate the accepted voice with a profile; it does not prevent another user from receiving the same voice or recite any privacy or access control mechanism. Considered as a whole, the claims use generic speech synthesis, audio presentation, user feedback, and profile storage to automate the abstract process of proposing, evaluating, revising, and saving a preferred option. These conventional functions do not integrate the abstract idea into a practical application or provide significantly more under Step 2B. therefore, the 101 rejection is maintained. Applicant's arguments with respect to 35 U.S.C. 103 rejection of claims 1 and 11 have been considered and found persuasive due to amendments, and the rejection has been withdrawn. See detailed reason for allowance below. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 Abstract Idea. Claims 1 and 11, under Step 2A, prong one, the claims recite the abstract mental process of receiving a person’s preferences, proposing an option, evaluating feedback, revising the option, receiving approval, and recording the approved selection. A person could perform the underlying evaluation and selection by reviewing a description of a desired voice, proposing a voice, considering feedback, selecting a different voice, and associating the accepted voice with a profile. Under step 2A, prong two, the abstract idea is not integrated into a practical application. The claim merely requires speech synthesis processing to generate and play audio so the user can evaluate the proposed voices. It does not recite a particular speech synthesis model, voice embedding structure, signal processing technique, or other technical mechanism that improves speech synthesis or computer performance. Generating and presenting ordinary synthesis speech merely applies to the preference and feedback process in a TTS environment. Under step 2B, the remaining limitations do not provide significantly more than the abstract idea. Receiving user inputs, processing text, generating synthetic audio, presenting the audio, receiving feedback, and storing an association with a profile are described only at a functional level and amount to conventional computer and TTS operations. Considered individually and as an ordered combination, the limitations merely automate an iterative human selection process without supplying an inventive technical concept. The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the claims are (i) mere instructions to implement the idea on a computer, and/or (ii) recitation of generic computer structure that serves to perform generic computer functions that are well-understood, routine, and conventional activities previously known to the pertinent industry. Viewed as a whole, these additional claim element(s) do not provide meaningful limitation(s) to transform the abstract idea into a patent eligible application of the abstract idea such that the claim(s) amounts to significantly more than the abstract idea itself. Therefore, the claim(s) are rejected under 35 U.S.C. 101 as being directed to non-statutory subject matter. There is further no improvement to the computing device. Dependent claims 2-10 and 12-20 further recite an abstract idea performable by a human and do not amount to significantly more than the abstract idea as they do not provide steps other than what is conventionally known. Claims 2 and 12: receiving the inputs through a graphical user interface only adds a generic way to enter information, so the claim still covers the same abstract voice-selection process on a computer. Claims 3 and 13: using a screen element to change a voice characteristic just adds routine interface controls to the same abstract idea of getting feedback and adjusting a choice. Claims 4 and 14: playing the first speech in response to a GUI input is only routine computer interaction and output activity, not a specific technological improvement. Claims 5 and 15: encoding the feedback and using a machine-learning model to pick the next voice still just automates the same abstract choice with a generic tool, without explaining a specific technical improvement in the model or speech synthesis itself. Claims 6 and 16: showing different GUI elements for different voice characteristics merely puts the same abstract preference-selection process onto a screen. Claims 7 and 17: using profile data to help choose the next voice only adds more information to the same abstract decision-making process. Claims 8 and 18: representing the proposed voice and the feedback as data and then using that data to choose a new voice is still just collecting and analyzing information to make a decision. Claims 9 and 19: letting a second user ask for a similar voice and then associating that voice with the second user is just reusing and storing a preference, not a technological solution. Claims 10 and 20: turning the request into a speech label before choosing a voice is only categorizing information to carry out the same abstract selection step. Allowable Subject Matter Claims 1-20 would be allowable if the Applicant can overcome the 101 Abstract Idea set forth. The following is a statement of reasons for the indication of allowable subject matter: for claims 1 and 11: Wang et al. (US 2010/0312565) teaches an interactive TTS system that receives text to be converted into speech, performs a first synthesis pass using default or user preferred parameters, and audibly presents the synthesized speech. The user may then modify pitch, duration, energy, pronunciation, or acoustic units, listen to the revised speech, and save the acceptable result after indicating completion (see [0033-0034] [0045-0049] [0065-0067] [Fig. 12, steps 1210-2140] [claims 1 and 6]). Kapilow et al. (US 2009/0063153) teaches generating a customized synthetic voice by receiving a user’s selection of a TTS voice and desired voice characteristics, such as accent, gender, pitch, emotion, or friendliness, and generating and presenting a blended voice for review. After review, the system receives user selected adjustments, generates and presents a revised voice, and presents a final blended voice when no further adjustments are received. Kapilow also describes a “voice profile” containing speaker specific parameters (see [0018-0024] [0032-0035] [Figs. 3A-3B] [claims 1-2, 11 and 15]). The difference between the prior art and the claimed invention is that Wang nor Kapilow explicitly teach determining, based at least in part on the third user input and the first proposed synthetic voice, a second proposed synthetic voice different from the first proposed synthetic voice; and after receiving the fourth user input, associating the second proposed synthetic voice with a first profile. Therefore, it would not have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the teachings of Wang and Kapilow to include determining, based at least in part on the third user input and the first proposed synthetic voice, a second proposed synthetic voice different from the first proposed synthetic voice; and after receiving the fourth user input, associating the second proposed synthetic voice with a first profile. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHREYANS A PATEL whose telephone number is (571)270-0689. The examiner can normally be reached Monday-Friday 8am-5pm PST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre Desir can be reached at 571-272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. SHREYANS A. PATEL Primary Examiner Art Unit 2653 /SHREYANS A PATEL/Examiner, Art Unit 2659
Read full office action

Prosecution Timeline

Sep 03, 2024
Application Filed
Apr 15, 2026
Non-Final Rejection mailed — §101
Jul 14, 2026
Response Filed
Sep 04, 2026
Final Rejection mailed — §101 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12744030
Method and System for a Parametric Speech Synthesis
2y 3m to grant Granted Sep 22, 2026
Patent 12738291
AUDIO PROCESSING APPARATUS, AUDIO PROCESSING METHOD, AND RECORDING MEDIUM
2y 4m to grant Granted Sep 15, 2026
Patent 12711948
EFFICIENT ADAPTATION OF SPOKEN LANGUAGE UNDERSTANDING BASED ON AUTOMATIC SPEECH RECOGNITION USING MULTI-TASK LEARNING
2y 5m to grant Granted Aug 18, 2026
Patent 12659658
ACOUSTIC ECHO CANCELLATION SYSTEM AND ASSOCIATED METHOD
2y 10m to grant Granted Jun 16, 2026
Patent 12646496
METHODS AND SYSTEMS OF TEXT-CONDITIONED AUDIO-VISUAL SPEECH GENERATION WITH MULTI-MODAL LATENT DIFFUSION MODELS
2y 2m to grant Granted Jun 02, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
89%
Grant Probability
97%
With Interview (+8.5%)
2y 0m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 415 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month