Prosecution Insights
Last updated: September 17, 2026
Application No. 19/060,542

MULTI-LINGUAL TEXT-TO-SPEECH CONTROLLING

Non-Final OA §102§103
Filed
Feb 21, 2025
Priority
Feb 23, 2024 — provisional 63/556,991
Examiner
AGAHI, DARIOUSH
Art Unit
Tech Center
Assignee
Murf Inc.
OA Round
1 (Non-Final)
84%
Grant Probability
Favorable
1-2
OA Rounds
1y 0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 84% — above average
84%
Career Allowance Rate
154 granted / 183 resolved
+24.2% vs TC avg
Strong +29% interview lift
Without
With
+29.2%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
23 currently pending
Career history
208
Total Applications
across all art units

Statute-Specific Performance

§101
23.9%
-16.1% vs TC avg
§103
54.6%
+14.6% vs TC avg
§102
10.9%
-29.1% vs TC avg
§112
6.7%
-33.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 183 resolved cases

Office Action

§102 §103
DETAILED ACTION This office action is in response to Applicant’s submission filed on 2/21/2025. Claims 1-20 are pending in the application of which Claims 1, 9, and 17 are independent and have been examined. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Applicant’s claim for the benefit of a prior-filed application under 35 U.S.C. 119(e) or under 35 U.S.C. 120, 121, or 365 is acknowledged. The prior-filed application (Provisional application No. 63/556991 Filed on 2/23/2024) is acknowledged. Information Disclosure Statement The information disclosure statement(s)(IDS) submitted on 6/11/2025 has been considered by the examiner. Claim Objections Listed claims are objected to for the informalities shown and may be addressed with suggested amendments: Claim 20, line 1 recite: ”The one or more non-transitory computer-readable media of claim 15, wherein…”. It is recommended to change it to “The one or more non-transitory computer-readable media of claim 17, wherein …”. Applicant is advised to review all claims for any potential claim objection issues. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1-5, 7, 9-13, 15, and 17-20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Zhao et al. (US20230335107A1)(herein "Zhao"). Regarding claims 1, 9, and 17, Zhao teaches [A computer device implemented method comprising: - claim 1], [A system comprising: one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising: - claim 9], and [One or more non-transitory computer-readable media storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising: - claim 17] (Zhao, Par. 0011:” … computer system comprises at least one processor, at least one memory in communication with the processor, and at least one network connection. A plurality of trainable models in communication with the processor is configured to convert input utterances from a non-native (L2) speaker to native-like sounding output utterances of the one or more languages. A software toolkit comprises a library of algorithms tangibly stored in at least one memory and in communication with at least one processor and with the plurality of models which when said algorithms are executed by the processor train the plurality of models to convert the input L2 utterances.”) receiving input text in a first language to convert to a desired speech in a second language; (Zhao, Par. 0038:” ... a reference-free foreign accent conversion computer system, comprising at least one processor; at least one memory in communication with the processor; at least one network connection; a plurality of trainable models in communication with the processor configured to convert input utterances from a non-native (L2) speaker learning one or more languages ..."). receiving one or more criteria for modifying the desired speech; (Zhao, Par. 0036:” In this embodiment the plurality of models may be trained to create the golden-speaker using a set of utterances from a reference L1 speaker, which are discarded thereafter, and the L2 speaker learning the at least one language; and convert the L2 speaker utterances to match the golden speaker utterances. Further to this embodiment the plurality of models are trained to convert new utterances from the L2 speaker to match a new golden speaker utterances. ”) converting the input text to a desired text in the second language; (Zhao, Par. 0004: "In addition to pronunciation training, FAC finds applications in movie dubbing (5), personalized Text-To-Speech (TTS) synthesis (6, 7), and improving automatic speech recognition (ASR) performance (8). ", para [0008], "The speaker encoder and the TTS model are trained with L1 speech only, and the ASR encoder is trained on speech data from L1 speakers and the target L2 speaker. During testing, they use the speaker encoder and ASR encoder to extract speaker embeddings and linguistic representations from the input L2 testing utterance, respectively. Then, they concatenate the two and feed them to the multi-speaker TTS model, which then generates the accent-converted utterance.") generating audio representations of the desired text in the second language; (Zhao, Par. 0037: "In another aspect the L2 speaker speech synthesizer may be trained to re-create the L2 speech from the speaker independent embeddings. In yet another aspect the speaker independent acoustic model may be trained to transform L1 speech into L1 speaker independent embeddings which are passed through the L2 speaker speech synthesizer to generate the golden speaker utterances.") predicting, for each of the audio representations of the desired text and using the one or more received criteria, a pitch value and a duration value; (Zhao, Par. 0054: "The attention mechanism allows the decoder to decide which parts of the hidden representation sequence contain useful information to make the predictions. At each output time step, the attention mechanism computes an attention context vector (a weighted sum of the hidden representation sequence) to summarize the contextual information. The decoder RNN reads the attention context vectors and predicts the output sequence in an autoregressive manner.", and Par. 0064: "Unlike conventional frame-by-frame VC systems (e.g., GMM, feedforward neural networks), which need time-alignment between the source and target speakers to generate the training frame pairs, seq2seq systems use an attention mechanism to produce learnable alignments between the input and output sequences. Therefore, they can also adjust for prosodic differences, for example, pitch, duration, and stressing) between the input and output sequences. This is crucial since prosody errors also contribute to foreign accentedness.") generating the desired speech using the predicted pitch value and the predicted duration value for each of the audio representations; and (Zhao, Par. 0091: "F0 RMSE: the F0 RMSE between the L2-GS and L1-GS speech on voiced frames. Lower F0 RMSE represents better pitch conversion performance. The F0 and voicing features were extracted by the WORLD vocoder with the Harvest pitch tracker (65).", and Par. 0092: "DDUR: the absolute difference in duration between the L2-GS and L1-GS speech. Lower DDUR implies better duration conversion performance.") providing the desired speech for output. (Zhao, Par. 0070: "A WaveGlow vocoder (15) is used to convert the output of the speech synthesizer back into a speech waveform. WaveGlow is a flow-based (54) network capable of generating high-quality speech from Mel-spectrograms. It takes samples from a zero mean spherical Gaussian (with variance a) with the same number of dimensions as the desired output and passes those samples through a series of layers that transform the simple distribution to one that has the desired distribution.", and Par. 0090: "MCD: the Mel-Cepstral Distortion (28) between the L2-GS (actual output) and L1-GS speech (desired output). It was computed on time-aligned (Dynamic Time Warping) Mel-cepstra between the L2-GS and the L1-GS audio. Lower MCD correlates with better spectral predictions. SPTK (63) and the, WORLD vocoder (64) were used to extract the Mel-cepstra with a shift size of 10 ms."). Regarding claims 2, 10, and 18, Zhao teaches the computer device implemented method, the system, and the readable media of claims 1, 9, and 17, respectively. Zhao further teaches wherein the first language and the second language are different. (Zhao, Par. 0004: "In the context of computer-assisted pronunciation training (1-4 ), this synthetic voice is often referred to as a "golden speaker" for the L2 speaker or a second language (L2) learner. The rationale is that the golden speaker is a better target for the L2 learner to imitate than an arbitrary native speaker, because the only difference between the golden speaker and the L2 learner's own voice is the accent, which makes mispronunciations more salient.", and Par. 0038: "A plurality of trainable models in communication with the processor configured to convert input utterances from a non-native (L2) speaker learning one or more languages to native-like sounding output utterances of the one or more languages; and a software toolkit comprising a library of algorithms tangibly stored in the at least one memory and in communication with the at least one processor and with the plurality of models which when said algorithms are executed by the processor train the plurality of models to convert the input L2 utterances.") Regarding claims 3, 11, and 19, Zhao teaches the computer device implemented method, the system, and the readable media of claims 1, 9, and 17, respectively. Zhao further teaches wherein receiving the one or more criteria for modifying the desired speech comprises receiving (i) a user specified pause rate, (ii) a user specified speaking rate, (iii) a user specified pitch, and (iv) a user specified sentence, word, or phoneme duration, for modifying the desired speech. (Zhao, Par. 0045: "The plurality of trainable models comprises a speech-independent acoustic model to extract speaker independent speech embeddings from an L1 speaker input utterance and/or the L2 speaker, a speech synthesizer to generate an L1 speaker reference-based golden-speaker utterances; and a pronunciation correction model to generate an L2 speaker reference-free golden speaker utterances. The reference-free foreign accent conversion system also uses transfer learning to reduce the amount of training data needed for the golden-speaker generation process.", para [0063], "The rationale behind using a VC system as the pronunciation-correction model is that VC can convert both the voice identity and the accent to match the target speaker. The L2 speaker and the L1-GS are treated as the source and target speakers in a VC task, respectively. Since the two speakers already share the same voice identity, the VC model only needs to match the accent of the target speaker, i.e., the golden speaker. During the inference stage, L2 speech is directly inputted into the pronunciation correction model, and the output will share similar pronunciation patterns as the L1-GS. The difficulty of this procedure is that L2 speakers tend to have disfluencies, hesitations, and inconsistent pronunciations, making the conversion much harder than converting between two native speakers, as discussed in prior literature (11 ). ", and Par. 0077: "To further evaluate the three L1-GS systems, formal listening tests were conducted to rate three perceptual attributes of the synthesized speech: accentedness, acoustic quality, and voice similarity. All listening tests were conducted through the Amazon Mechanical Turk platform (mturk.com). Instructions were given in each test to help the participants focus on the target speech attribute.") Note: pause rate maps to disfluencies and hesitations, speaking rate maps to fluency, hesitations, and accentedness. Furthermore, pitch maps to accent, acoustic quality, and intonation patterns. Also, word and phoneme Duration maps to pronunciation accuracy, accentedness, and disfluencies. Regarding claims 4, 12, and 20, Zhao teaches the computer device implemented method, the system, and the readable media of claims 1, 9, and 17, respectively. Zhao further teaches wherein generating the audio representations of the desired text in the second language comprises generating phoneme representations of the desired text in the second language. (Zhao, Par. 0067:” In addition, the baseline system uses multi-task learning (52, 53) to make the synthesized pronunciations more stable. Two independent phoneme classifiers, each containing one fully-connected layer and a softmax operation, are added to predict the input and output phoneme sequences ŶinP=[ŷ1inP, . . . , ŷTininP] and ŶoutP=[ŷ1outP, . . . , ŷToutoutP], respectively. These phoneme classifiers are only used during training and are discarded in inference. ci and qi are defined in the same manner as in equations (1) and (3).”, and Par. 0105:” The phoneme prediction ground-truth labels were per-frame phoneme labels (with word positions) that were produced by force-aligning the audio to its orthographic transcriptions. It is noted that the phoneme predictions were only required in training, not testing.”) Regarding claims 5, and 13, Zhao teaches the computer device implemented method, and the system of claims 4, and 12, respectively. Zhao further teaches wherein predicting, for each of the audio representations of the desired text and using the one or more received criteria, the pitch value and the duration value comprises predicting, for each of the generated phoneme representations of the desired text and using the one or more received criteria, the pitch value and the duration value for the generated phoneme representations using a pitch predictor and a duration predictor of a text-to-speech model. (Zhao, Par.0064: "FIG. 4 shows an overview of the baseline system. Unlike conventional frame-by-frame VC systems (e.g., GMM, feedforward neural networks), which need time alignment between the source and target speakers to generate the training frame pairs, seq2seq systems use an attention mechanism to produce learnable alignments between the input and output sequences. Therefore, they can also adjust for prosodic differences, for example, pitch, duration, and stressing) between the input and output sequences. This is crucial since prosody errors also contribute to foreign accentedness.", and Par. 0068: The backward decoder, like its forward counterpart, also predicts its own set of stop tokens Ŷstopbwd, output phoneme labels ŶoutPbwd, and uses the shared PostNet to predict a refined Mel-spectrogram Ŷmel-PostNetbwd.”, and Par. 0094: "Results are summarized in Table 4. For all measures, the scores between the original L2 speech and the L1-GS speech also were computed as a reference. In addition, the WER of the L1-GS speech was included as an upper-bound. By definition, the other three measures on the L1-GS speech are all zero. For Baseline 2, the WER were only computed since the system was not trained to predict L1-GS, which makes computing the other objective scores ill-defined."). Regarding claims 7, and 15, Zhao teaches the computer device implemented method, and the system of claims 1, and 9, respectively. Zhao further teaches wherein generating the desired speech using the predicted pitch value and the predicted duration value for each of the audio representations comprises generating the desired speech by concatenating the predicted pitch value and the predicted duration value for each of the audio representations. (Zhao, Par. 0049: "TDNNF achieves performance on Large Vocabulary Continuous Speech Recognition (LVCSR) tasks that is comparable to that of AMs based on recurrent structures (e.g., Bi-LSTMs), but is more efficient during training and inference due to its feedforward nature ( 42). To produce an SI speech embedding, each acoustic feature vector (40-dim MFCC) is concatenated with an i-vector (100-dim) of the corresponding speaker (44) and used them to the AM, which then is trained on a large corpus from a few thousand native speakers (Librispeech (45)).", and Par. 0062: "The predicted Mel-spectrograms are converted back to audio waveforms using a WaveGlow neural vocoder trained on the L2 utterances. The L2 synthesizer is then driven with a set of utterances from the reference L1 speaker, to produce the L1-GS utterances that are used in Step 2.", and Par. 0064: "FIG. 4 shows an overview of the baseline system. Unlike conventional frame-by-frame VC systems (e.g., GMM, feedforward neural networks), which need time-alignment between the source and target speakers to generate the training frame pairs, seq2seq systems use an attention mechanism to produce learnable alignments between the input and output sequences. Therefore, they can also adjust for prosodic differences, for example, pitch, duration, and stressing) between the input and output sequences. This is crucial since prosody errors also contribute to foreign accentedness.", and Par. 0065: "Specifically, let xi be the i-th feature vector in the sequence, the input X=[x1, . . . , xTin] to the conversion system is the concatenation of the bottleneck features, i.e., BNFs, and Mel-spectrogram computed from the L2 utterance.”, and Par. 0091:”F0 RMSE: the F0 RMSE between the L2-GS and L1-GS speech on voiced frames. Lower F0 RMSE represents better pitch conversion performance. The F0 and voicing features were extracted by the WORLD vocoder with the Harvest pitch tracker (65).", and Par. 0092: "DDUR: the absolute difference in duration between the L2-GS and L1-GS speech. Lower DDUR implies better duration conversion performance.") Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 6, and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Zhao, and in further view of Elias et al. (US 20220301543 A1)(herein " Elias "), and Lovelace et al. (US 20250104692 A1)(herein “Lovelace”). Regarding claims 6, and 14, Zhao teaches the computer device implemented method, and the system of claims 5, and 13, respectively. Zhao, does not teach, however, Elias teaches wherein the text-to-speech model is trained using phoneme durations obtained [[using a residual vector quantization (RVQ)]] based aligner model. (Elias, Par. 0010:” … Here, training the TTS model is further based on the global phoneme duration loss. In some examples, training the TTS model based on the final spectrogram loss and the global phoneme duration loss includes training the duration model network to predict the phoneme duration for each phoneme without using supervised phoneme duration labels extracted from an external aligner.”) Elias is considered to be analogous to the claimed invention because it is in the same field of endeavor. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Zhao further in view of Elias to wherein the text-to-speech model is trained using phoneme durations. Motivation to do so would allow manually adjust the length of specific sounds for expressive, precise speech synthesis. Zhao, as modified above, does not teach, however, Lovelace teaches a residual vector quantization (RVQ). (Lovelace, Par. 0115:” … Simple-TTS is capable of generating speech using only the transcript at inference time. The viability of end-to-end latent diffusion for text-to-speech synthesis is demonstrated, paving the way for further scaling and improvements of generative speech models.”, and Par. 0126:”… The pre-trained EnCodec model may be used to map raw audio waveforms to a sequence of continuous vectors. EnCodec, like other neural audio codecs such as SoundStream, applies residual vector quantization to map each continuous vector to a variable number of discrete tokens that capture increasingly fine details. The number of quantizers may be adjusted to trade off compression rates and quality.”) Lovelace is considered to be analogous to the claimed invention because it is in the same field of endeavor. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Zhao, as modified above, further in view of Lovelace to employ a residual vector quantization. Motivation to do so would improve audio quality by using multiple codebooks in a chain. Claims 8, and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Zhao, and in further view of Ittycheriah et al. (US6041300A)(herein " Ittycheriah"). Regarding claims 8, and 16, Zhao teaches the computer device implemented method, and the system of claims 1, and 9, respectively. Zhao, further teaches generating a linguistic context aligner that is configured to align TTS input phoneme sequences with large language model (LLM)-based linguistic context features, wherein the generating comprises: providing, to the linguistic context aligner, the phoneme sequences as input; (Zhao, Par. 0054:” At each output time step, the attention mechanism computes an attention context vector (a weighted sum of the hidden representation sequence) to summarize the contextual information. The decoder RNN reads the attention context vectors and predicts the output sequence in an autoregressive manner.”, and Par. 0100:” Additionally, the instant method appears to have used a broader window to compute the attention context compared with Baseline 1, as reflected by the width of the attention alignment path. Therefore, the instant system utilized more contextual information during the decoding process.”, and Par. 0067:” In addition, the baseline system uses multi-task learning (52, 53) to make the synthesized pronunciations more stable. Two independent phoneme classifiers, each containing one fully-connected layer and a softmax operation, are added to predict the input and output phoneme sequences ŶinP=[ŷ1inP, . . . , ŷTininP] and ŶoutP=[ŷ1outP, . . . , ŷToutoutP], respectively. These phoneme classifiers are only used during training and are discarded in inference. … Where YinP, YoutP are the ground-truth input and output phoneme sequence, respectively.”, and Par. 0105:” The phoneme prediction ground-truth labels were per-frame phoneme labels (with word positions) that were produced by force-aligning the audio to its orthographic transcriptions.”, and Par. 0076: "In a first experiment, the word error rate (WER) of L1-GS utterances synthesized was computed using each of the three speaker embeddings. In this case, the speech recognizer consisted of the TDNN-F acoustic model combined with an unpruned 3-gram language model trained on the Librispeech transcripts. As a reference, WERs on test utterances also were computed from the L1 speaker (BDL) and the two L2 speakers (YKWK, TXHC). Results are summarized in Table 1 L1-GS utterances from the three systems achieve lower WERs than the corresponding utterances from the L2 speakers.") providing, to the linguistic context aligner, the linguistic context features as the input as a target; (Zhao, Par. 0054:” At each output time step, the attention mechanism computes an attention context vector (a weighted sum of the hidden representation sequence) to summarize the contextual information. The decoder RNN reads the attention context vectors and predicts the output sequence in an autoregressive manner.”, and Par. 0100:” … Qualitatively, it is observed that the attention weights of the Baseline 1 system contained an abnormal jump towards the end of the synthesis, while the instant system produced smooth alignments at the same time steps. Additionally, the instant method appears to have used a broader window to compute the attention context compared with Baseline 1, as reflected by the width of the attention alignment path. Therefore, the instant system utilized more contextual information during the decoding process.”) encoding, by the linguistic context aligner, the phoneme sequences as first embeddings; (Zhao, Par. 0054:” The model follows a general encoder-decoder (or seq2seq) paradigm with an attention mechanism. Conceptually, an encoder-decoder architecture uses an encoder (usually a recurrent neural network; RNN) to “consume” input sequences and generate a high-level hidden representation sequence.”, and Par. 0055:” The speech synthesizer takes the speech embeddings as input.”, and Par. 0059:” The original Tacotron 2 was designed to accept character sequences as input, which are significantly shorter than our speech embedding sequences. For example, each sentence in our corpus contains 41 characters on average, whereas the corresponding speech embedding sequence has a few hundred frames.”, and Par. 0067:” In addition, the baseline system uses multi-task learning (52, 53) to make the synthesized pronunciations more stable. Two independent phoneme classifiers, each containing one fully-connected layer and a softmax operation, are added to predict the input and output phoneme sequences ŶinP=[ŷ1inP, . . . , ŷTininP] and ŶoutP=[ŷ1outP, . . . , ŷToutoutP], respectively.”) encoding, by the linguistic context aligner, the linguistic context features as second embeddings; (Zhao, Par. 0054:” At each output time step, the attention mechanism computes an attention context vector (a weighted sum of the hidden representation sequence) to summarize the contextual information. The decoder RNN reads the attention context vectors and predicts the output sequence in an autoregressive manner.”, and Par. 0055: "The speech embeddings are then passed through multiple 1-D convolutional layers, which model longer-term context. Next, an encoder (one Bi-LSTM) converts the convolutions into a hidden linguistic representation sequence. Finally, the hidden linguistic representation sequence is passed to the decoder, which consists of a location-sensitive attention mechanism (47) and a decoder LSTM, to predict the raw Mel-spectrogram.", and Par. 0056:” … the attention context vector ci is the weighted sum of h …”, and Par. 0072: "The objectives of this experiment were to determine the optimal speech embedding, and more importantly, to establish that L1-GS utterances captured the native accent and the L2 speaker identity, which is critical since they would be used as targets for the reference-free FAC task.") extracting, by the linguistic context aligner, features from the first embeddings and the second embeddings; and (Zhao, Par. 0008:” … speaker encoder and ASR encoder to extract speaker embeddings and linguistic representations from the input L2 testing utterance, respectively.”, and Par. 0046: "The speech embeddings are extracted using an acoustic model trained on a large corpus of native speech, so they are speaker-independent (10, 11). The L2 synthesizer is then driven with speech embeddings extracted from the L1 utterances. This results in a set of golden-speaker utterances that have the voice identity of the L2 learner since they are generated from the L2 synthesizer and the pronunciation patterns of the L1 speaker since the input is obtained from an L1 utterance.", and Par. 0075: "To generate the L1-GS utterances for testing, the three speech embeddings were extracted from speaker BDL's test set and drove the systems with their respective input. The output Mel-spectrograms were then converted to speech through the WaveGlow vocoders.") determining, using an attention mechanism and a [[Viterbi]] decoder, an alignment between hidden representations of the phoneme sequences as the first embeddings and hidden representations of the linguistic context feature as the second embeddings. (Zhao, Par. 0054: "Then, a decoder (an RNN with an attention mechanism) processes the hidden representation sequence. The attention mechanism allows the decoder to decide which parts of the hidden representation sequence contain useful information to make the predictions. At each output time step, the attention mechanism computes an attention context vector ( a weighted sum of the hidden representation sequence) to summarize the contextual information. The decoder RNN reads the attention context vectors and predicts the output sequence in an autoregressive manner.", and Par. 0055: "Finally, the hidden linguistic representation sequence is passed to the decoder, which consists of a location-sensitive attention mechanism ( 47) and a decoder LSTM, to predict the raw Mel-spectrogram. It is noted that the input and output sequences of the speech synthesizer have the same length, and thus, the speech synthesizer only models the speaker identity and retains the phonetic and prosodic cues carried by the input speech embeddings. In a similar conversion model in a recent study ( 48) it was observed that if the temporal structure, such as the length, of the input and output sequences were the same, then removing the attention module did not hurt performance, which suggests a potential path to further simplify the model structure of the speech synthesizer built herein.", and Par. 0065: "The down-sampling effectively reduces the sequence length of the input, which speeds up the encoder computation by a factor of two and makes it easier for the attention mechanism to learn a meaningful alignment between the input and output sequences.") Zhao, does not teach, however, Ittycheriah teaches employ a viterbi decoder. (Ittycheriah, claim 13:” A speech recognition and synthesis system, as recited in claim 12 wherein said decoder is a Viterbi decoder.”) Ittycheriah is considered to be analogous to the claimed invention because it is in the same field of endeavor. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Zhao further in view of Ittycheriah to employ a viterbi decoder. Motivation to do so would provide ability to efficiently find the globally optimal sequence of hidden states, such as phonemes, units, or acoustic features. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. Chen et al. (US11990117B2) teaches in ABS:” … a multilingual text-to-speech (TTS) model. The method also includes generating a native synthesized speech representation for an input text sequence in a first language that is conditioned on speaker characteristics of a native speaker of the first language. The method also includes generating a cross-lingual synthesized speech representation for the input text sequence in the first language that is conditioned on speaker characteristics of a native speaker of a different second language. The method also includes generating a first speech recognition result for the native synthesized speech representation and a second speech recognition result for the cross-lingual synthesized speech representation. …” Examiner's Note: Examiner has cited particular columns and line numbers and/or paragraph numbers in the references applied to the claims above for the convenience of the applicant. Although the specified citations are representative of the teachings of the art and are applied to specific limitations within the individual claim, other passages and figures may apply as well. It is respectfully requested from the applicant in preparing responses, to fully consider the references in entirety as potentially teaching all or part of the claimed invention, as well as the context of the passage as taught by the prior art or disclosed by the Examiner. In the case of amending the Claimed invention, Applicant is respectfully requested to indicate the portion(s) of the specification which dictate(s) the structure relied on for proper interpretation and also to verify and ascertain the metes and bounds of the claimed invention. Any inquiry concerning this communication or earlier communications from the examiner should be directed to DARIOUSH AGAHI whose telephone number is (408)918-7689. The examiner can normally be reached Monday - Thursday and alternate Fridays, 7:30-4:30 PT. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bhavesh Mehta can be reached on 571-272-7453. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. DARIOUSH AGAHI, P.E. Primary Examiner /DARIOUSH AGAHI/Primary Examiner, Art Unit 2656
Read full office action

Prosecution Timeline

Feb 21, 2025
Application Filed
Sep 11, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12718030
LARGE LANGUAGE MODELS PROVIDING EVIDENCE MAPPINGS FOR GENERATED OUTPUT
3y 1m to grant Granted Aug 25, 2026
Patent 12718017
SYSTEM AND METHOD FOR PROVIDING LARGE LANGUAGE MODEL FOR SANCTIONS ARTIFICIAL INTELLIGENCE ASSISTED AUTOMATION
2y 11m to grant Granted Aug 25, 2026
Patent 12718819
INTERRUPTION DETECTION AND HANDLING BY DIGITAL ASSISTANTS
2y 3m to grant Granted Aug 25, 2026
Patent 12710915
VOICE MODIFICATION FOR WEARABLE DEVICE
3y 9m to grant Granted Aug 18, 2026
Patent 12712972
System and method for generating and managing a workflow using webhook technology
2y 4m to grant Granted Aug 18, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
84%
Grant Probability
99%
With Interview (+29.2%)
2y 7m (~1y 0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 183 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month