Prosecution Insights
Last updated: August 18, 2026
Application No. 17/631,695

Paragraph synthesis with cross utterance features for neural TTS

Non-Final OA §101
Filed
Jan 31, 2022
Priority
Sep 12, 2019 — CN 201910864208.0 +2 more
Examiner
MASTERS, KRISTEN MICHELLE
Art Unit
2659
Tech Center
2600 — Communications
Assignee
Microsoft Technology Licensing, LLC
OA Round
5 (Non-Final)
65%
Grant Probability
Favorable
5-6
OA Rounds
0m
Est. Remaining
86%
With Interview

Examiner Intelligence

Grants 65% — above average
65%
Career Allowance Rate
32 granted / 49 resolved
+3.3% vs TC avg
Strong +21% interview lift
Without
With
+21.2%
Interview Lift
resolved cases with interview
Typical timeline
3y 0m
Avg Prosecution
22 currently pending
Career history
84
Total Applications
across all art units

Statute-Specific Performance

§101
37.9%
-2.1% vs TC avg
§103
48.1%
+8.1% vs TC avg
§102
7.8%
-32.2% vs TC avg
§112
3.4%
-36.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 49 resolved cases

Office Action

§101
Detailed Action This communication is in response to the Request for Continued Examination filed on 5/26/2026. Claims 1-6, 8-12 and 15-21 are pending and have been examined. Claims 7, 13 and 14 have been cancelled. Any previous objection/rejection not mentioned in this Office Action has been withdrawn by the Examiner. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55. Response to Amendment The Applicant has amended the claims to include “by repeating vectors in the average embedding vector sequence to align with the phone sequence;” “multi-level” “ input, the multi-level context features comprising a current semantic feature based on the word sequence of the text input and a historical acoustic feature generated by converting acoustic features of at least one sentence before the text input using at least one acoustic encoder to provide acoustic state continuity across sentences; updating the phone sequence by adding a begin token or an end token to the phone sequence, a length of the begin token or a length of the end token being determined based on the multi-level context features such that a pause duration between sentences varies based on the multi-level context features;” “updated” “generated from the updated phone sequence,” “multi-level” Regarding the 35 U.S. C. 101 rejection, The Applicant notes The amended claims now recite three notable technical elements. Applicant notes technological improvement alignment is performed "by repeating vectors in the average embedding vector sequence to align with the phone sequence," which recites a concrete computational mechanism rather than a merely functional recitation of alignment. Examiner notes applicant repeated the claim wording and added “repeating”. Examiner considers this claim language to be high level, functional language. Applicant notes technological improvement context features are now "multi-level context features" comprising a current semantic feature and a historical acoustic feature generated by converting acoustic features of at least one previous sentence using an acoustic encoder. This is an unconventional combination that departs from traditional TTS systems which use only phone features of the current sentence. Examiner notes unconventional combination alone without sufficient structure or constraints is insufficient to transform an abstract data processing claim into a patent eligible technological improvement. Applicant must show that the recited components and their combination are unconventional and that the claimed steps produce a technical improvement in TTS systems. Examiner notes the claim language of acoustic state continuity is not adequately defined or constrained. Examiner notes if the “historical acoustic feature” is defined as “at least one previous sentence” this was possible in conventional TTS systems where the system passed features from one sentence to the next. For example, in HMM-based TTS with previous sentence prosody context. Applicant notes technological improvement phone sequence is updated with begin or end tokens whose length is determined by the multi-level context features such that pause durations between sentences vary based on the context. This last element directly addresses the technical problem identified in paragraphs [0027]-[0029] (i.e., that traditional TTS systems use fixed pause durations resulting in unnatural, repetitive rhythm) by producing speech with variable, context-dependent pause durations that result in more natural sounding speech. Examiner notes that conventional TTS systems can use word sequence information for pauses and conventional TTS systems already use pause durations that vary based on sentence boundaries, grammatical structure, commas, question marks, periods and other linguistic multi-context features. The applicants’ arguments and amendments do not overcome the 35 U.S. C. 101 rejection. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-6, 8-12 and 15-21 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Independent Claim 1 recites, “1. A method for generating speech through neural text-to-speech (TTS) synthesis, comprising: obtaining a text input; generating, by an encoder, a phone feature of the text input; generating a word embedding vector sequence based on a word sequence using a word embedding model comprising a sequence-to-sequence encoder-decoder framework; generating an average embedding vector sequence corresponding to at least one sentence based on the word embedding vector sequence; aligning the average embedding vector sequence with a phone sequence of the text input by repeating vectors in the average embedding vector sequence to align with the phone sequence; generating multi-level context features of the text input based on a set of sentences associated with the text input; the multi-level context features comprising a current semantic feature based on the word sequence of the text input and a historical acoustic feature generated by converting acoustic features of at least one sentence before the text input using at least one acoustic encoder to provide acoustic state continuity across sentences; updating the phone sequence by adding a begin token or an end token to the phone sequence, a length of the begin token or a length of the end token being determined based on the multi-level context features such that a pause duration between sentences varies based on the multi-level context features; and generating, by a vocoder, a speech waveform corresponding to the text input based on an updated phone feature generated from the updated phone sequence, the multi-level context features, and the aligned average embedding vector sequence.” The limitations of “obtaining …”, “generating …”, “generating …”, “generating …”, “aligning …”, “generating …”, “updating…”, “generating …”, as drafted covers a mental activity or human process. More specifically, a human is capable of obtaining a text input using the human visual system by observing the text that is written, inscribed or displayed visually in some way, a human can use this visual information in further downstream cognitive processes. A human is capable of generating the context features based on the word sequence comprises: generating a word embedding vector sequence based on the word sequence; This relates to a human using cognitive processes to generate context features A human is capable of generating, using the logic and reasoning powers of the human mind a word embedding vector based on the word sequence, and using the human physical processes, typing or writing out the word embedding vector based on the word sequence. A human is capable of generating an average embedding vector sequence corresponding to the at least one sentence based on the word embedding vector sequence This relates to a human using cognitive processes to generate using the logic and reasoning powers of the human mind an average embedding vector sequence corresponding to the at least one sentence based on the word embedding vector sequence, and using the human physical processes, typing or writing out the word embedding vector based on the word sequence. A human is capable of aligning the average embedding vector sequence with a phone sequence of the text input; This relates to a human using cognitive processes and logic and reasoning to aligned the average embedding vector sequence with a phone sequence of the text input. The claim relates to generating the context features based on the aligned average embedding vector sequence. This relates to a human using cognitive processes and logic and reasoning to generate the context features based on the aligned average embedding vector sequence. A human is capable of generating context features of the text input based on a set of sentences associated with the text input using the cognitive processes of the human mind such as natural language understanding and linguistic knowledge to map the text to an element or elements of context that resides in memory and then interpret the context of the given sentence, thereby creating context features. A human is capable of updating the phone sequence by adding a begin token or an end token to the phone sequence, a length of the begin token or a length of the end token being determined based on the multi-level context features such that a pause duration between sentences varies based on the multi-level context features using pen and paper. A human is capable of generating a speech waveform corresponding to the text input based on the phone feature and the context features using the natural process of speech production through the vocal tract to form and articulate words and sentences. A human is capable aligning average embedding vector sequence using pen and paper. The claim is directed to an abstract idea. No additional elements are present in the claim. Regarding Independent Claim 15, Claim 15 is an apparatus claim with limitations similar to that of claim 1 and is rejected under the same rationale. Regarding Independent Claim 19, Claim 19 is storage medium claim with limitations similar to that of claim 1 and is rejected under the same rationale. With respect to Claims 2, 16 and 20 the claim relates to generating the context features comprises: obtaining acoustic features corresponding to at least one sentence of the set of sentences before the text input This relates to a human using auditory processes to listen to a given spoken sentence and picking out the acoustic features before and after a sentence. The claim relates to generating the context features based on the acoustic features. This relates to a human using auditory processes to listen to a sentence with acoustic features and picking out context features. No additional limitations are present. With respect to Claim 3, 17 and 21 the claim relates to aligning the context features with a phone sequence of the text input. This relates to a human using the context features in combination with a phone sequence as described above in claim 1. No additional limitations are present. With respect to Claim 4 and 18, the claim relates to identifying a word sequence from at least one sentence of the set of sentences; This relates to a human using cognitive processes to listen or observe a sentence and identify a word sequence in a given set of sentences. The claim relates to generating the context features based on the word sequence. This relates to a human using cognitive processes as described in claims 1 and 15 to create context features for a given word sequence. No additional limitations are present. With respect to Claim 5, the claim relates to at least one sentence comprises at least one of: a sentence corresponding to the text input, sentences before the text input, and sentences after the text input. This relates to a human using cognitive processes to create or identify a suitable sentence before and after a given sentence text input. No additional limitations are present. With respect to Claim 6, the claim relates to the at least one sentence represents content of the set of sentences. This relates to a human using cognitive processes to create or identify a sentence that represent content of a set of sentences. No additional limitations are present. With respect to Claim 8, the claim relates to determining a position of the text input in the set of sentences. This relates to a human using cognitive processes and logic and reasoning to determining a position of the text input in the set of sentences generate the context features based on the location. No additional limitations are present. With respect to Claim 9 the claim relates to generating the context features based on the location comprises: generating a position embedding vector sequence based on the location. This relates to a human using cognitive processes and logic and reasoning to generate a position embedding vector sequence based on the location. The claim relates to aligning the position embedding vector sequence with a phone sequence of the text input. This relates to a human using cognitive processes and logic and reasoning to align the position embedding vector sequence with a phone sequence of the text input. The claim relates to generating the context features based on the aligned position embedding vector sequence. This relates to a human using cognitive processes and logic and reasoning to generate the context features based on the aligned position embedding vector sequence. No additional limitations are present. With respect to Claim 10, the claim relates to combining the phone feature and the context features into mixed features. This relates to a human using cognitive processes and logic and reasoning to combine the phone feature and the context features into mixed features. The claim relates to applying an attention mechanism on the mixed features to obtain attended mixed features. This relates to a human using cognitive processes and logic and reasoning to apply an attention mechanism on the mixed features to obtain attended mixed features. The claim relates to generate the speech waveform based on the attended mixed features. This relates to a human using cognitive processes and logic and reasoning to generate the speech waveform based on the attended mixed features. No additional limitations are present. With respect to Claim 11, the claim relates to generating the speech waveform comprises: combining the phone feature and the context features into first mixed features. This relates to a human using cognitive processes and logic and reasoning to combine the phone feature and the context features into first mixed features. The claim relates to applying a first attention mechanism on the first mixed features to obtain first attended mixed features. This relates to a human using cognitive processes and logic and reasoning to apply a first attention mechanism on the first mixed features to obtain first attended mixed features. The claim relates to applying a second attention mechanism on at least one context feature of the context features to obtain at least one attended context feature. This relates to a human using cognitive processes and logic and reasoning to apply a second attention mechanism on at least one context feature of the context features to obtain at least one attended context feature. The claim relates to combining the first attended mixed features and the at least one attended context feature into second mixed features. This relates to a human using cognitive processes and logic and reasoning to combine the first attended mixed features and the at least one attended context feature into second mixed features. The claim relates to generating the speech waveform based on the second mixed features. This relates to a human using cognitive processes and logic and reasoning to generate the speech waveform based on the second mixed features. No additional limitations are present. With respect to Claim 12 the claim relates to combining the phone feature and the context features into first mixed features. This relates to a human using cognitive processes and logic and reasoning to combine the phone feature and the context features into first mixed features. The claim relates to applying an attention mechanism on the first mixed features to obtain first attended mixed features. This relates to a human using cognitive processes and logic and reasoning to apply an attention mechanism on the first mixed features to obtain first attended mixed features. The claim relates to performing averaging pooling on at least one context feature of the context features to obtain at least one average context feature. This relates to a human using cognitive processes and logic and reasoning to perform averaging pooling on at least one context feature of the context features to obtain at least one average context feature. The claim relates to combining the first attended mixed features and the at least one average context feature into second mixed features. This relates to a human using cognitive processes and logic and reasoning to combine the first attended mixed features and the at least one average context feature into second mixed features. The claim relates to generating the speech waveform based on the second mixed features. This relates to a human using cognitive processes and logic and reasoning to generate the speech waveform based on the second mixed features. No additional limitations are present. With respect to Claim 13 the claim relates to identifying a phone sequence from the text input. This relates to a human using cognitive processes and logic and reasoning to identify a phone sequence from the text input. The claim relates to updating the phone sequence by adding a begin token and/or an end token to the phone sequence. This relates to a human using cognitive processes and logic and reasoning to update the phone sequence by adding a begin token and/or an end token to the phone sequence. The claim relates to determining the length of the begin token and the length of the end token according to the context features. This relates to a human using cognitive processes and logic and reasoning to determine the length of the begin token and the length of the end token according to the context features. The claim relates to generating the phone feature based on the updated phone sequence. This relates to a human using cognitive processes and logic and reasoning to generate the phone feature based on the updated phone sequence. No additional limitations are present. Allowable Subject Matter Claims 1-6, 8-12 and 15-21 are rejected due to 35 USC § 101 indicated above, but would be allowable if rewritten or amended to overcome the rejections under 35 USC § 101. The following is a statement of reasons for the indication of allowable subject matter: None of the found prior arts, either alone or in combination, disclose the subject matter as claimed. Any comments considered necessary by applicant must be submitted no later than the payment of the issue fee and, to avoid processing delays, should preferably accompany the issue fee. Such submissions should be clearly labeled “Comments on Statement of Reasons for Allowance.” Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to KRISTEN MICHELLE MASTERS whose telephone number is (703)756-1274. The examiner can normally be reached M-F 8:30 AM - 5:00 PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre Louis Desir can be reached at 571-272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /KRISTEN MICHELLE MASTERS/Examiner, Art Unit 2659 /PIERRE LOUIS DESIR/Supervisory Patent Examiner, Art Unit 2659
Read full office action

Prosecution Timeline

Show 8 earlier events
Nov 17, 2025
Response Filed
Mar 23, 2026
Final Rejection mailed — §101
May 11, 2026
Interview Requested
May 20, 2026
Examiner Interview Summary
May 20, 2026
Applicant Interview (Telephonic)
May 26, 2026
Request for Continued Examination
May 28, 2026
Response after Non-Final Action
Jul 31, 2026
Non-Final Rejection mailed — §101 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12707198
ACOUSTIC ECHO CANCELLATION SYSTEM AND ASSOCIATED METHOD
3y 0m to grant Granted Aug 11, 2026
Patent 12694889
PROFANITY FILTER FOR COLLABORATION SESSIONS IN HETEROGENOUS COMPUTING PLATFORMS
3y 10m to grant Granted Jul 28, 2026
Patent 12664366
CROSS-DOMAIN LABEL-ADAPTIVE STANCE DETECTION
3y 11m to grant Granted Jun 23, 2026
Patent 12592219
Hearing Device User Communicating With a Wireless Communication Device
4y 5m to grant Granted Mar 31, 2026
Patent 12548569
METHOD AND SYSTEM OF DETECTING AND IMPROVING REAL-TIME MISPRONUNCIATION OF WORDS
3y 2m to grant Granted Feb 10, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
65%
Grant Probability
86%
With Interview (+21.2%)
3y 0m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 49 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month