Prosecution Insights
Last updated: October 04, 2026
Application No. 18/474,484

NATURAL LANGUAGE GENERATION

Final Rejection §103
Filed
Sep 26, 2023
Examiner
MEIS, JON CHRISTOPHER
Art Unit
2654
Tech Center
2600 — Communications
Assignee
Amazon Technologies Inc.
OA Round
3 (Final)
33%
Grant Probability
At Risk
4-5
OA Rounds
0m
Est. Remaining
86%
With Interview

Examiner Intelligence

Grants only 33% of cases
33%
Career Allowance Rate
11 granted / 33 resolved
-28.7% vs TC avg
Strong +52% interview lift
Without
With
+52.4%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
17 currently pending
Career history
60
Total Applications
across all art units

Statute-Specific Performance

§101
21.7%
-18.3% vs TC avg
§103
55.9%
+15.9% vs TC avg
§102
12.5%
-27.5% vs TC avg
§112
9.2%
-30.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 33 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION Claims 1-20 are pending. Claims 1, 5, and 13 are independent. This Application was published as US 20250104693. Apparent priority is 9/26/2023. The instant Application is directed to a method of generating synthetic speech with prosody characteristics based on the natural language input. Applicant’s amendments and arguments are considered but are either unpersuasive or moot in view of the new grounds of rejection that, if presented, were necessitated by the amendments to the Claims. This action is Final. Response to Arguments 35 USC 101 Applicant’s amendments/arguments have been fully considered and are persuasive. Therefore, the rejections under 35 USC 101 are withdrawn. 35 USC 103 Applicant’s arguments with respect to 35 USC 103 have been considered but are not persuasive. Applicant argues (pg. 12) that because Liu teaches multiple layers configured for related text-classification tasks, it would not have been obvious to use multiple layers configured for natural language generation and prosody prediction. Examiner disagrees. Bonar discloses multiple layers ([0005] discloses an LLM which would have multiple layers) which are configured for natural language and prosody prediction (Fig. 3 shows Large Language Model 302 outputs Text Response 318 and Style Cues 322). As noted previously, Bonar does not explicitly disclose that one set of the layers is specific to natural language generation while another set of the layers is specific to prosody prediction. Liu teaches the application of a multi-task architecture with a specific example of text classification. However, Section 3 discloses shared layer architecture generically. Classification is not mentioned in this section; rather Liu refers only to “tasks” and presents an architecture that could be applied to any neural machine learning model. Furthermore, Liu teaches multi-task learning for a model that classifies sentiment from text. (Section 5.1 lists two sentiment classification tasks.) Bonar also teaches a model that classifies sentiment from text (“[0038]…Similarly, the style cue 410 can be generated by the large language model to augment the text response 408 with one or more selected sentiments 412 to express various emotions (e.g., friendly, inquisitive).”) then further uses the sentiment (style cue) as a prosody prediction in speech synthesis (see [0040]). Examiner argues that given the teaching of Bonar to use a model for both natural language generation and sentiment prediction, and the teaching of Liu to use task specific layers for sentiment prediction, it would have been obvious to one of ordinary skill in the art to use task specific layers for sentiment prediction as specifically taught by Liu, and use other layers for the other task, as taught generically by Liu. This would have been beneficial as Liu teaches that multi-task learning can improve the performance of a task (Abstract) and help deal with limited training data (pg. 2 last para). Therefore, the rejection is maintained. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 1-9, 11-17, and 19-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Bonar in view of Liu et al. (“Recurrent Neural Network for Text Classification with Multi-Task Learning”). Regarding claim 1, Bonar discloses: 1. A computer-implemented method comprising: receiving first input data representing a first natural language input; ("Audio Input 110" Fig. 1) receiving first context data associated with the first natural language input; ("Conversation Prompt 308" Fig. 3) generating first prompt data including the first input data and the first context data, wherein the first prompt data represents a first request for a first language model to determine a first output responsive to the first natural language input; ("Large Language Model Input 218" Fig. 2) processing, using a first set of layers of the first language model, the first prompt data to generate first natural language data responsive to the first natural language input, ("Text Response(s) 318" Fig. 3) wherein the first set of layers is configured for natural language generation; (not explicitly disclosed ) processing, using a second set of layers of the first language model, the first prompt data to generate first prosody data representing at least a first synthetic voice characteristic, ("Style Cue(s) 322" Fig. 3) wherein the second set of layers is configured for prosody prediction; (not explicitly disclosed ) processing the first natural language data and the first prosody data to generate first output audio data representing first synthetic speech corresponding to the first synthetic voice characteristic and responsive to the first natural language input; and causing presentation of the first output audio data. ("Audio Output 406" Fig. 4) Bonar does not explicitly disclose multiple task specific layers. Liu discloses multiple task specific layers. (“Model-II: Coupled-Layer Architecture In Model-II, we assign a LSTM layer for each task, which can use the information for the LSTM layer of the other task.”) PNG media_image1.png 287 539 media_image1.png Greyscale Liu Fig. 2 Bonar and Liu are considered analogous art to the claimed invention because they disclose neural networks for NLP. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of Bonar with a multi-layer shared architecture for Multi-Task learning as taught by Liu for natural language and prosody tasks. Doing so would have been beneficial because features learned from a task may be useful for other tasks. (Liu pg. 3 para 1) Regarding claim 2, Bonar discloses: 2. The computer-implemented method of claim 1, wherein processing the first prompt data to generate the first prosody data comprises: receiving, from a first layer of the first set of layers, first embedding data representing the first prompt data; and processing, by the second set of layers, the first prompt data and the first embedding data to generate the first prosody data. (See claim 1.) Bonar does not disclose: embedding data that is processed by shared layers. Liu discloses: embedding data (“In all of our experiments, the word embeddings are trained using word2vec [Mikolov et al., 2013] on the Wikipedia corpus (1B words).” Pg. 4, Section 5.2) Liu also discloses a shared architecture where information is shared between task layers. (See claim 1) Bonar and Liu are considered analogous art to the claimed invention because they disclose deep neural networks for NLP. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of Bonar with word embedding as taught by Liu. This combination falls under combining prior art elements according to known methods to yield predictable results or use of known technique to improve similar devices (methods, or products) in the same way. See MPEP 2141, KSR, 550 U.S. at 418, 82 USPQ2d at 1396. Regarding claim 3, Bonar does not disclose a shared layer architecture. Liu discloses: 3. The computer-implemented method of claim 2, wherein the first embedding data is received at a second layer of the second set of layers, (“In Model-II, we assign a LSTM layer for each task, which can use the information for the LSTM layer of the other task.” Pg. 3, section Model-II)) and the method further comprises: processing, by the second layer of the second set of layers, the first prompt data and the first embedding data to generate second embedding data representing the first prompt data and the first embedding data; and (Fig. 2, (b) shows that each layer has an output) processing the second embedding data using a third layer of the second set of layers to generate third embedding data representing the first prompt data, (Fig. 2(b) shows at least 4 layers for each task.) wherein the first prosody data is generated based at least in part on the third embedding data. (Fig. 2(b) shows the task specific outputs are based on at least a third embedding data.) Bonar and Liu are considered analogous art to the claimed invention because they disclose deep neural networks for NLP. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of Bonar with a shared architecture for Multi-Task learning as taught by Liu for natural language and prosody tasks. Doing so would have been beneficial because features learned from a task may be useful for other tasks. (Liu pg. 3 para 1) Note that the specification of the instant application discloses: “[0064] … data representing a latent representation (e.g., embedding data) representing synthesized speech…” Based on this, “embedding data” is interpreted to mean any latent representation of data that has been embedded. Regarding claim 4, Bonar does not disclose a shared layer architecture. Liu discloses: 4. The computer-implemented method of claim 1, wherein processing the first prompt data to generate the first natural language data comprises: receiving, from a first layer of the second set of layers and at a second layer of the first set of layers, first embedding data representing the first prompt data, and processing, by the first set of layers, the first prompt data and the first embedding data to generate the first natural language data. (Fig. 2, (b) discloses that embeddings from both tasks are shared with the layers for both tasks.) See claim 3 for motivation statement. Regarding claim 5, Bonar discloses: 5. A computer-implemented method comprising: receiving first input data corresponding to a user input; ("Audio Input 110" Fig. 1) determining first prompt data including the first input data, wherein the first prompt data represents a first request for a first language model to determine an output responsive to the user input; ("Large Language Model Input 218" Fig. 2) processing, using a first portion of the first language model, the first prompt data to generate first natural language data responsive to the user input, wherein the first portion of the first language model is configured for natural language generation; ("Text Response(s) 318" Fig. 3) processing, using a second portion of the first language model, the first prompt data to generate first prosody data representing at least a first voice characteristic, wherein the second portion of the first language model is configured for prosody prediction; ("Style Cue(s) 322" Fig. 3) using the first natural language data and the first prosody data to generate first output audio data representing first synthetic speech corresponding to the at least first voice characteristic; and causing presentation of the first output audio data. ("Audio Output 406" Fig. 4; see also “Translate the text response and the style cue to generate an audio output response to the user input 514” Fig. 5) Bonar does not explicitly disclose different portions of the language model configured for different tasks. Liu discloses multiple task specific portions of a language model (layers). (“Model-II: Coupled-Layer Architecture In Model-II, we assign a LSTM layer for each task, which can use the information for the LSTM layer of the other task.”) Bonar and Liu are considered analogous art to the claimed invention because they disclose neural networks for NLP. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of Bonar with a multi-layer shared architecture for Multi-Task learning as taught by Liu for natural language and prosody tasks. Doing so would have been beneficial because features learned from a task may be useful for other tasks. (Liu pg. 3 para 1) Regarding claim 6, Bonar discloses: 6. The computer-implemented method of claim 5, wherein: the first portion of the first language model corresponds to a first set of layers of the first language model, ; and the second portion of the first language model corresponds to a second set of layers of the first language model. Bonar does not explicitly disclose multiple task specific layers. Liu discloses multiple task specific layers. (“Model-II: Coupled-Layer Architecture In Model-II, we assign a LSTM layer for each task, which can use the information for the LSTM layer of the other task.”) Bonar and Liu are considered analogous art to the claimed invention because they disclose neural networks for NLP. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of Bonar with a multi-layer shared architecture for Multi-Task learning as taught by Liu for natural language and prosody tasks. Doing so would have been beneficial because features learned from a task may be useful for other tasks. (Liu pg. 3 para 1) Regarding claim 7, Bonar does not disclose a shared layer architecture with embeddings. Liu discloses: 7. The computer-implemented method of claim 6, wherein processing, using the second set of layers of the first language model, the first prompt data comprises: receiving, from a first layer of the first set of layers and at a second layer of the second set of layers, first embedding data representing the first prompt data, wherein the first prosody data is generated based at least in part by processing the first embedding data. (Fig. 2, (b) discloses that embeddings from both tasks are shared with the layers for both tasks.) See claim 3 for motivation statement. Regarding claim 8, Bonar does not disclose a shared layer architecture with embeddings. Liu discloses: 8. The computer-implemented method of claim 6, wherein processing, using the first set of layers of the first language model, the first prompt data comprises: receiving, from a first layer of the second set of layers and at a second layer of the first set of layers, first embedding data representing the first prompt data, and processing, by the first set of layers, the first prompt data and the first embedding data to generate the first natural language data. (Fig. 2, (b) discloses that embeddings from both tasks are shared with the layers for both tasks.) See claim 3 for motivation statement. Regarding claim 9, Bonar does not disclose a shared layer architecture with embeddings. Liu discloses: 9. The computer-implemented method of claim 7, further comprising: processing, by the second layer of the second set of layers, the first prompt data and the first embedding data to generate second embedding data representing the first prompt data and the first embedding data; and processing the second embedding data using a third layer of the second set of layers to generate third embedding data representing the first prompt data, wherein the first prosody data is generated based at least in part on the third embedding data. (Fig. 2(b) shows at least 4 layers for each task, and each layer receives the data from the previous data as well as shared data.) See claim 3 for motivation statement. Regarding claim 11, Bonar discloses: 11. The computer-implemented method of claim 5, wherein the first prosody data represents a natural language description of the at least first voice characteristic. (“Style Cue(s) 322” (“Friendly”) Fig. 3- “friendly” is a natural language description of the characteristic.) Regarding claim 12, Bonar discloses: 12. The computer-implemented method of claim 5, further comprising: receiving context data associated with the user input, wherein the first prompt data represents a further instruction for the first language model to generate prosody data associated with the first natural language data based on the first input data and the context data. (See Fig. 3 – the “Conversation Prompt 308” (“Be friendly and good at conversation”) reads on context data and is included in the prompt to the LLM. See also: "[0031] Furthermore, the large language model 302 can be configured with a conversational profile 312 which can enable the large language model 302 to not only respond to individual inputs but rather carry on a conversation in which context can persist and change over time. Consequently, what constitutes an appropriate response can be nebulous and depend heavily on implications of previous statements, the current mood, and other indefinite factors. ... As such, the large language model 302 can appropriately respond to user inputs while accounting for conversational history, mood, and other context clues." See also: “[0034] The word selection and phrasing of the text response 318 can be determined by the large language model 302 based on a context derived from the speech-to-text translation 306 in combination with the instructions of the conversation prompt 308 and/or the style prompt 310…” Claim 13 is a system claim with limitations corresponding to the limitations of Claim 5 and is rejected under similar rationale. Additionally, at least one processor; and at least one memory including instructions of the Claim are taught by Bonar. ( Processing Unit(s) 602; Memory 604, Fig. 6) Claim 14 is a system claim with limitations corresponding to the limitations of Claim 6 and is rejected under similar rationale. Claim 15 is a system claim with limitations corresponding to the limitations of Claim 7 and is rejected under similar rationale. Claim 16 is a system claim with limitations corresponding to the limitations of Claim 8 and is rejected under similar rationale. Claim 17 is a system claim with limitations corresponding to the limitations of Claim 9 and is rejected under similar rationale. Claim 19 is a system claim with limitations corresponding to the limitations of Claim 11 and is rejected under similar rationale. Claim 20 is a system claim with limitations corresponding to the limitations of Claim 12 and is rejected under similar rationale. Claim(s) 10 and 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Bonar in view of Calapodescu et al. (US 20230215421 A1). Regarding claim 10, Bonar discloses: 10. The computer-implemented method of claim 5, wherein the first prosody data corresponds to a spectrogram, ([0040] discloses that style cues can include a sentiment) and generating the first output audio data comprises: processing, using a vocoder, the spectrogram. ( Vocal Synthesizer 422, Fig. 4 ) Bonar does not disclose that the prosody data is a spectrogram. Calapodescu discloses: wherein the first prosody data corresponds to a spectrogram, and generating the first output audio data comprises: processing, using a vocoder, the spectrogram. ("[0045] Speech representations may additionally or alternatively be embodied in a speech signal that can be processed downstream by a voice synthesizer, vocoder, etc. to generate speech. An example of such speech signals is a spectrogram, such as a Mel-spectrogram." ) Bonar and Calapodescu are considered analogous art to the claimed invention because they disclose methods of TTS with prosody control. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of Bonar to use a spectrogram for the prosody data. Doing so would have been beneficial to achieve fast speed with comparable voice quality. (Calapodescu [0063].) This combination falls under simple substitution of one known element for another to obtain predictable results or use of known technique to improve similar devices (methods, or products) in the same way. See MPEP 2141, KSR, 550 U.S. at 418, 82 USPQ2d at 1396. Claim 18 is a system claim with limitations corresponding to the limitations of Claim 10 and is rejected under similar rationale. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to JON C MEIS whose telephone number is (703)756-1566. The examiner can normally be reached Monday - Thursday, 8:30 am - 5:30 pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Hai Phan can be reached at 571-272-6338. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JON CHRISTOPHER MEIS/Examiner, Art Unit 2654 /HAI PHAN/Supervisory Patent Examiner, Art Unit 2654
Read full office action

Prosecution Timeline

Sep 26, 2023
Application Filed
Sep 15, 2025
Non-Final Rejection mailed — §103
Dec 11, 2025
Response Filed
Mar 24, 2026
Non-Final Rejection mailed — §103
Jun 29, 2026
Response Filed
Aug 25, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12711961
SYSTEM AND METHOD FOR DIGITAL VOICE DATA PROCESSING AND AUTHENTICATION
3y 4m to grant Granted Aug 18, 2026
Patent 12603087
VOICE RECOGNITION USING ACCELEROMETERS FOR SENSING BONE CONDUCTION
3y 8m to grant Granted Apr 14, 2026
Patent 12579975
Detecting Unintended Memorization in Language-Model-Fused ASR Systems
2y 11m to grant Granted Mar 17, 2026
Patent 12482487
MULTI-SCALE SPEAKER DIARIZATION FOR CONVERSATIONAL AI SYSTEMS AND APPLICATIONS
3y 0m to grant Granted Nov 25, 2025
Patent 12475312
FOREIGN LANGUAGE PHRASES LEARNING SYSTEM BASED ON BASIC SENTENCE PATTERN UNIT DECOMPOSITION
2y 9m to grant Granted Nov 18, 2025
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

4-5
Expected OA Rounds
33%
Grant Probability
86%
With Interview (+52.4%)
2y 10m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 33 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month