Prosecution Insights
Last updated: October 02, 2026
Application No. 18/854,360

PERSONALIZED AND DYNAMIC TEXT TO SPEECH VOICE CLONING USING INCOMPLETELY TRAINED TEXT TO SPEECH MODELS

Final Rejection §102§103
Filed
Oct 04, 2024
Priority
Apr 13, 2022 — nonprovisional of PCTCN2022086591
Examiner
WEAVER, ADAM MICHAEL
Art Unit
2658
Tech Center
2600 — Communications
Assignee
Microsoft Technology Licensing, LLC
OA Round
2 (Final)
88%
Grant Probability
Favorable
3-4
OA Rounds
6m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 88% — above average
88%
Career Allowance Rate
14 granted / 16 resolved
+25.5% vs TC avg
Strong +31% interview lift
Without
With
+31.2%
Interview Lift
resolved cases with interview
Typical timeline
2y 6m
Avg Prosecution
21 currently pending
Career history
53
Total Applications
across all art units

Statute-Specific Performance

§101
29.6%
-10.4% vs TC avg
§103
52.6%
+12.6% vs TC avg
§102
14.6%
-25.4% vs TC avg
§112
1.6%
-38.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 16 resolved cases

Office Action

§102 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement (IDS) submitted on 07/12/2026 is being considered by the examiner. Response to Amendment The Amendment filed 05/20/2026 has been entered. Claims 1-15 remain pending in the application. Response to Arguments Applicant’s arguments, filed 05/20/2026, have been fully considered. With respect to the 35 U.S.C. 112 rejections, Applicant’s arguments have been fully considered and are persuasive. Therefore, this rejection has been withdrawn. With respect to the 35 U.S.C. 102 rejection, on pages 9-10, of claims 1-3, 5-7, and 9-11, under Azizah et al. ("Transfer Learning, Style Control, and Speaker Reconstruction Loss for Zero-Shot Multilingual Multi-Speaker Text-to-Speech on Low-Resource Languages", 01/18/2022), hereinafter referred to as Azizah, and the 35 U.S.C. 103 rejection, on pages 10-11, of claim 4, under Azizah, in view of Wu et al. ("AdaSpeech 4: Adaptive Text to Speech in Zero-Shot Scenarios", 04/01/2022), hereinafter referred to as Wu, claims 8 and 12 under Azizah, in view of Germain et al. ("Speech Denoising with Deep Feature Losses", 09/14/2018), hereinafter referred to as Germain, and claims 13-15 under Azizah, in view of Beaufays et al. (US Patent Application Publication No. 2023/0177382), hereinafter referred to as Beaufays, the Applicant asserts that Azizah fails to disclose the serial architecture and data routing of disentangled inputs processed by a text-to-speech module and shown in the amended claims. The Applicant also asserts that the prior Office Action treated Azizah’s style encoder as the claimed feature extractor and Azizah’s style embedding as the claimed prosodic features. They also state that Azizah’s style encoder is not upstream of the speaker encoder and does not provide acoustic features to the speaker encoder. The Applicant further asserts that Azizah does not disclose that the prosodic features are provided directly to the TTS module as separate input from the speaker embedding, in a disentangled fashion. They also state that Azizah does not disclose the prosodic features comprising at least fundamental frequency and energy extracted from the reference speech. They also state that the further cited references fail to remedy the deficiencies of Azizah. The Examiner respectfully disagrees. In response to Applicant’s argument that Azizah fails to disclose the serial architecture and data routing of disentangled inputs, Azizah Figure 2 pg. 5900 shows serial architecture, as it is required that the mel-spectrogram be extracted first, then the speaker encoder and style encoder output their respective embeddings next, which is then fed into the TTS model. This shows that there is a definite, serial processing structure to the model disclosed by Azizah. Concerning the disentangled inputs, Azizah pg. 5900 states “Attention network and autoregressive decoder process the hidden representation H = (h_1,...,h_Tx) output from text encoder concatenated with language embedding l, style embedding g, and speaker embedding q to generate predicted mel-spectrogram Y = y_1,...,y_TY and stop token Z = (z_1, . . . ,z_TY).” These are initially separate vectors and therefore it would be obvious to provide as such. Combining the vectors is therefore indistinguishable from providing them separately. Furthermore, Azizah pg. 5900 states “Speaker embedding q is not only fed into attention network, it is also fed into pre-net.” This further shows that speaker embedding q is fed into the TTS in a “disentangled fashion”, as not only is it provided in a concatenation into the attention network, but it is also fed by itself into the pre-net. In response to Applicant’s argument that the prior Office Action treated Azizah’s style encoder as the claimed feature extractor and Azizah’s style embedding as the claimed prosodic features, the feature extractor is mentioned in Azizah pg. 5903 Section IV A.: “For each utterance, we extract mel-spectrogram as the acoustic feature with 80 mel channels, Hann typed window size 1024, hop size = 256, and 1024-point FFT,” which inherently implies the use of a feature extractor to generate and extract the mel-spectrogram. In response to the argument that Azizah’s style encoder is not upstream of the speaker encoder and does not provide acoustic features to the speaker encoder, this is now moot considering that it pulls dependency on assuming that Azizah’s style encoder was claimed as the feature extract, which was touched upon in the previous sentence. Both the speaker encoder and style encoders are upstream from the feature extractor as referenced in Azizah pg. 5903 Section IV A. In response to Applicant’s argument that Azizah does not disclose that the prosodic features are provided directly to the TTS module as separate input from the speaker embedding, in a disentangled fashion, this was touched on prior in this section. The concatenation of vectors that were previously singular is an indistinguishable difference, as it therefore would be obvious to provide them to the TTS model as single vectors. Yet still, Azizah discloses that the speaker embedding is fed, alone, into the pre-net section of the TTS, thus even more so disclosing the “disentangled fashion” into which the features are provided to the TTS. In response to Applicant’s argument that Azizah does not disclose the prosodic features comprising at least fundamental frequency and energy extracted from the reference speech, this argument has been considered but is moot because the new ground of rejection of claim 3 does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Hence, Azizah cures all deficiencies set forth in Applicant’s arguments Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claim(s) 1-2, 5-7, and 9-11 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Azizah et al. ("Transfer Learning, Style Control, and Speaker Reconstruction Loss for Zero-Shot Multilingual Multi-Speaker Text-to-Speech on Low-Resource Languages", 01/18/2022), hereinafter referred to as Azizah. Regarding claim 1, Azizah discloses a computing system configured to instantiate a machine learning model that is capable of generating a personalized voice for a new target speaker in response to applying the machine learning model to target reference speech from a new target speaker (Azizah Figure 2 pg. 5900 and Azizah pg. 5899 Section III B.), the machine learning model having not been previously applied to any labeled training data associated with the new target speaker, the computing system ("There are three speaker adaptation approaches in neural TTS: 1) Using speaker encoder network that is jointly trained with the entire TTS network [12]–[14] or neural vocoder [15]; 2) Using speaker embedding generated by the pre-trained speaker encoder and fine-tuning the TTS model [14]–[18]; 3) Using speaker embedding generated by the pre-trained speaker encoder without fine-tuning the TTS model [20]–[22]. Among these approaches only the third one is able to handle zero-shot speaker adaptation," Azizah pg. 5897 Section II B.) comprising: one or more processors (Azizah Abstract Pg. 5895 speaks about deep neural network based systems, which inherently need one or more processors); and one or more storage devices storing computer-executable instructions (Azizah Abstract Pg. 5895 speaks about deep neural network based systems, which inherently need memory) which are executable by the one or more processors for instantiating a machine learning model comprising a feature extractor (“For each utterance, we extract mel-spectrogram as the acoustic feature with 80 mel channels, Hann typed window size 1024, hop size = 256, and 1024-point FFT,” Azizah pg. 5903 Section IV A., inherently implies a feature extractor was used to extract the mel-spectrogram), a speaker encoder (Azizah Figure 2 pg. 5900 shows the speaker encoder), and a text-to-speech module arranged in a serial architecture (Azizah Figure 2 pg. 5900 shows a serial architecture, as the mel-spectrogram is extracted first, the speaker encoder and style encoder output their respective embeddings next, etc.), wherein the machine learning model by the feature extractor (“For each utterance, we extract mel-spectrogram as the acoustic feature with 80 mel channels, Hann typed window size 1024, hop size = 256, and 1024-point FFT,” Azizah pg. 5903 Section IV A., inherently implies a feature extractor was used to extract the mel-spectrogram), extract acoustic features and prosodic features from new target reference speech (Azizah Figure 2 pg. 5900 shows the speaker encoder outputting speaker embeddings q and the style encoder outputting prosodic style embeddings g from the input of target speaker mel-spectrogram M); by the speaker encoder, receive the acoustic features extracted by the feature extractor (Azizah Figure 2 pg. 5900 shows the mel-spectrogram being input into the speaker encoder to generate speaker embeddings q) and generate a speaker embedding corresponding to the new target speaker based on the extracted acoustic features (Azizah Figure 2 pg. 5900 shows the speaker encoder outputting speaker embeddings q from the input of target speaker mel-spectrogram M); by the text-to-speech module (Azizah Figure 2 pg. 5900 shows the TTS model), and generate the personalized voice corresponding for the new target speaker based on: (i) the speaker embedding generated by the speaker encoder (Azizah Figure 2 pg. 5900 shows the speaker embedding q being generated by the speaker encoder), and (ii) the prosodic features extracted by the feature extractor (Azizah Figure 2 pg. 5900 shows the prosodic style embedding g being generated by the style encoder), wherein the prosodic features are provided directly to the text-to-speech module as a separate input from the speaker embedding (Azizah Figure 2 pg. 5900 shows the prosodic style embeddings g being input into H to be concatenated, combining, or concatenating, versus keeping separate is indistinguishable, as from a concatenation, it would be obvious to instead provide them separately), and such that the prosodic features are provided to the text-to-speech module in a disentangled fashion from the speaker embedding (Azizah Figure 2 pg. 5900 shows the prosodic style embeddings g being input into H to be concatenated, combining, or concatenating, versus keeping separate is indistinguishable, as from a concatenation, it would be obvious to instead provide them separately), and wherein the machine learning model generates the personalized voice (Azizah Figure 2 pg. 5900 shows output mel-spectrogram Y' being created from the inputs of the speaker embeddings q and the style embeddings g, created from the input of the target speaker mel-spectrogram M input to both the speaker encoder and the style encoder) ("There are three speaker adaptation approaches in neural TTS: 1) Using speaker encoder network that is jointly trained with the entire TTS network [12]–[14] or neural vocoder [15]; 2) Using speaker embedding generated by the pre-trained speaker encoder and fine-tuning the TTS model [14]–[18]; 3) Using speaker embedding generated by the pre-trained speaker encoder without fine-tuning the TTS model [20]–[22]. Among these approaches only the third one is able to handle zero-shot speaker adaptation," Azizah pg. 5897 Section II B.); and use the extracted acoustic features to generate the speaker embedding and to utilize both (i) the extracted prosodic features and (ii) the speaker embedding to generate the personalized voice for the new target speaker as output in response to applying the machine learning model to input comprising the new target reference speech (Azizah Figure 2 pg. 5900 shows output mel-spectrogram Y' being created from the inputs of the speaker embeddings q and the style embeddings g, created from the input of the target speaker mel-spectrogram M input to both the speaker encoder and the style encoder). Regarding claim 2, Azizah discloses all of the limitations of claim 1. Azizah further discloses wherein the acoustic features include a Mel-spectrogram (Azizah Figure 2 pg. 5900 shows target speaker Mel-spectrogram). Regarding claim 5, Azizah discloses all of the limitations of claim 1. Azizah further discloses wherein the machine learning model is further configured to capture residual prosodic features and generate a style token (Azizah Figure 2 pg. 5900 shows the style encoder outputting prosodic style embeddings g). Regarding claim 6, Azizah discloses all of the limitations of claim 5. Azizah further discloses wherein the machine learning model is further configured to capture a speaking rate associated with new target speaker ("The style encoder is used to explicitly control the prosody from the target speaker. It produces high-level style representation such as speaker style, pitch range, and speaking rate from a collection of voice data that is jointly trained with the overall TTS system," Azizah pg. 5896 Section I). Regarding claim 7, Azizah discloses all of the limitations of claim 1. Azizah further discloses wherein the machine learning model is further configured to generate the personalized voice corresponding for the new target speaker based on the speaker embedding, the prosodic features, and a language embedding (Azizah Figure 2 pg. 5900 shows the speaker encoder outputting speaker embeddings q, the style encoder outputting prosodic style embeddings g, and the language encoder outputting language embedding l), such that the machine learning model is configured as a cross-lingual personalized text-to-speech model capable of generating speech in a second language that is different than a first language corresponding to the new target reference speech by using the personalized voice associated with the new target speaker ("The TTS model in this study is an extension of Tacotron 2 [7] with additional networks to handle multilingual and zero-shot multi-speaker adaptation," Azizah pg. 5899 Section III B.). Regarding claim 9, Azizah discloses a method for generating a personalized voice for a new target speaker using a zero-shot personalized text-to-speech model, the method comprising (Azizah Figure 2 pg. 5900 and Azizah pg. 5899 Section III B.): accessing a personalized text-to-speech model that is configured to generate a personalized voice corresponding for a new target speaker based on speaker embeddings and prosodic features extracted from new target reference speech of the new target speaker (Azizah Figure 2 pg. 5900 shows the speaker encoder outputting speaker embeddings q and the style encoder outputting prosodic style embeddings g and Azizah pg. 5899 Section III B.), and without having to first fine-tune the text-to-speech model based on new labeled training data associated with the new target speaker ("There are three speaker adaptation approaches in neural TTS: 1) Using speaker encoder network that is jointly trained with the entire TTS network [12]–[14] or neural vocoder [15]; 2) Using speaker embedding generated by the pre-trained speaker encoder and fine-tuning the TTS model [14]–[18]; 3) Using speaker embedding generated by the pre-trained speaker encoder without fine-tuning the TTS model [20]–[22]. Among these approaches only the third one is able to handle zero-shot speaker adaptation," Azizah pg. 5897 Section II B.); receiving the new target reference speech associated with the new target speaker (Azizah Figure 2 pg. 5900 shows target speaker Mel-spectrogram); extracting (Azizah Figure 2 pg. 5900 shows the speaker encoder outputting speaker embeddings q and the style encoder outputting prosodic style embeddings g and Azizah pg. 5899-5900 Section III B. (3) Style Encoder and pg. 5900 Section III B. (4) Speaker Encoder); generating a speaker embedding corresponding to the new target speaker based on the extracted acoustic features (Azizah pg. 5900 Section III B. (4) Speaker Encoder); and generating, by the text-to-speech module, the personalized voice corresponding to the new target speaker based on (Azizah Figure 2 pg. 5900 shows the output Mel-spectrogram and "Our model is an end-to end TTS attention-based encoder decoder that predicts mel-spectrogram Y = y1,...,yTY directly from the input grapheme-level text sequences X = (x1,...,xTx), language identity LangID, and the mel-spectrogram of the target speaker speech sample M = (m1,...,mTM)," Azizah pg. 5899 Section III B.): (i) the speaker embedding and (ii) the prosodic features, wherein the prosodic features are provided directly to the text-to-speech module as a separate input from the speaker embedding (Azizah Figure 2 pg. 5900 shows the prosodic style embeddings g being input into H to be concatenated, combining, or concatenating, versus keeping separate is indistinguishable, as from a concatenation, it would be obvious to instead provide them separately), such that the prosodic features are provided to the text-to-speech module in a disentangled fashion from the speaker embedding (Azizah Figure 2 pg. 5900 shows the prosodic style embeddings g being input into H to be concatenated, combining, or concatenating, versus keeping separate is indistinguishable, as from a concatenation, it would be obvious to instead provide them separately), and wherein generating the personalized voice is performed without first applying the text-to-speech model to any labeled training data associated with the new target speaker ("There are three speaker adaptation approaches in neural TTS: 1) Using speaker encoder network that is jointly trained with the entire TTS network [12]–[14] or neural vocoder [15]; 2) Using speaker embedding generated by the pre-trained speaker encoder and fine-tuning the TTS model [14]–[18]; 3) Using speaker embedding generated by the pre-trained speaker encoder without fine-tuning the TTS model [20]–[22]. Among these approaches only the third one is able to handle zero-shot speaker adaptation," Azizah pg. 5897 Section II B.). Regarding claim 10, Azizah discloses all of the limitations of claim 9. Azizah further discloses further comprising: receiving new input text (Azizah Figure 2 pg. 5900 shows Text Sequence X being input into the Text Encoder); and generating synthesized speech in the personalized voice based on the new input text (Azizah Figure 2 pg. 5900 shows output Mel-Spectrogram Y' based on the input Text Sequence X). Regarding claim 11, Azizah discloses all of the limitations of claim 10. Azizah further discloses wherein the new target reference speech comprises spoken language utterances in a first language and the new input text comprises text-based language utterances in a second language, the method further comprising (Azizah Figure 3 pg. 5901 Stage 2 inputs the speaker encoder embeddings from one language into the speaker encoder of Stage 3, and Stage 1 inputs the text encoder embeddings from another language into the text encoder of Stage 3): identifying a new target language based on the second language associated with the new input text (Azizah Figure 3 pg. 5901 Stage 1 inputs the text encoder embeddings from another language into the text encoder of Stage 3); accessing a language embedding configured to control language information for the synthesized speech (Azizah Figure 3 pg. 5901 Stage 3 shows Language Encoder outputting language embeddings l); and generating the synthesized speech in the second language using the language embedding (Azizah Figure 3 pg. 5901 Stage 3 shows Output Mel-Spectrogram Y'). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 3 is/are rejected under 35 U.S.C. 103 as being unpatentable over Azizah, in view of Howard (US Patent No. 11,004,461). Regarding claim 3, Azizah discloses all of the limitations of claim 1. However, Azizah fails to disclose wherein the prosodic features comprise at least a fundamental frequency and an energy extracted from the new target reference speech Howard teaches a system and method for real-time vocal features extraction. Howard teaches wherein the prosodic features comprise at least a fundamental frequency and an energy extracted from the new target reference speech (Howard col. 14 lines 42-60). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Azizah’s disclosure of a zero-shot multilingual text-to-speech reconstruction system by including Howard’s teaching of extracting energy and fundamental frequency as acoustic/prosodic features. This would ensure that only the necessary information is kept from the waveform, saving storage and computational time and power in the process. Energy and fundamental frequency carry any speech characteristics along with them, further assuring that the reconstructed speech would be as accurate as possible. Claim(s) 4 is/are rejected under 35 U.S.C. 103 as being unpatentable over Azizah, in view of Wu et al. ("AdaSpeech 4: Adaptive Text to Speech in Zero-Shot Scenarios", 04/01/2022), hereinafter referred to as Wu. Regarding claim 4, Azizah discloses all of the limitations of claim 1. However, Azizah fails to disclose wherein the machine learning model is further configured to: generate phoneme representations in response to receiving phonemes; predict phoneme duration and phone-level fundamental frequency in response to receiving the speaker embedding; and decode the speaker embedding along with encoder output and other input features. Wu teaches an adaptive text-to-speech system and method for zero-shot synthesis of new voices. Wu teaches wherein the machine learning model is further configured to: generate phoneme representations in response to receiving phonemes (Wu Figure 1(a) pg. 2 shows a phoneme encoder); predict phoneme duration and phone-level fundamental frequency in response to receiving the speaker embedding ("Besides, as the output of phoneme encoder is used to predict variance information (e.g., pitch, duration) related to speaker identity through variance adaptor, it is also required to have strong generalization ability on speaker characteristics," Wu pg. 3 Section 2.2); and decode the speaker embedding along with encoder output and other input features (Wu Figure 1(a) pg. 2 shows a decoder taking in the encoder output and other features). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Azizah’s disclosure of a zero-shot multilingual text-to-speech reconstruction system by including Wu’s teaching of utilizing a phoneme encoder. Using a phoneme encoder with embeddings of phonemes obtained from speech would allow for the final, generated speech to have increased accuracy and prosody associated with it. This would allow for the generation/reconstruction of much more natural sounding speech. Claim(s) 8 and 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Azizah, in view of Germain et al. ("Speech Denoising with Deep Feature Losses", 09/14/2018), hereinafter referred to as Germain. Regarding claim 8, Azizah discloses all of the limitations of claim 1. However, Azizah fails to disclose wherein the machine learning model is further configured to denoise the new target reference speech. Germain teaches a deep learning approach to denoising speech signals. Germain teaches wherein the machine learning model is further configured to denoise the new target reference speech ("We present an end-to-end deep learning approach to denoising speech signals by processing the raw waveform directly. Given input audio containing speech corrupted by an additive background signal, the system aims to produce a processed signal that contains only the speech content," Germain pg. 1 Abstract). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Azizah’s disclosure of a zero-shot multilingual text-to-speech reconstruction system by including Germain’s teaching of denoising speech signals. Denoising is well known within the art of speech processing, as it is commonly used to remove the noise and background information in order to isolate the speech signal. This would have been an obvious inclusion. Regarding claim 12, Azizah discloses all of the limitations of claim 9. However, Azizah fails to disclose wherein the method further comprises denoising Germain teaches wherein the method further comprises denoising ("We present an end-to-end deep learning approach to denoising speech signals by processing the raw waveform directly. Given input audio containing speech corrupted by an additive background signal, the system aims to produce a processed signal that contains only the speech content," Germain pg. 1 Abstract). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Azizah’s disclosure of a zero-shot multilingual text-to-speech reconstruction system by including Germain’s teaching of denoising speech signals. Denoising is well known within the art of speech processing, as it is commonly used to remove the noise and background information in order to isolate the speech signal. This would have been an obvious inclusion. Claim(s) 13-15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Azizah, in view of Beaufays et al. (US Patent Application Publication No. 2023/0177382), hereinafter referred to as Beaufays. Regarding claim 13, Azizah discloses a system configured for facilitating a creation of a zero-shot personal text-to-speech model, the system comprising: at least one hardware processor; and at least one hardware storage device storing (Azizah Abstract Pg. 5895 speaks about deep neural network based systems, which inherently need memory): access a feature extractor configured to extract acoustic features and prosodic features from new target reference speech associated with a new target speaker (Azizah Figure 2 pg. 5900 shows the speaker encoder outputting speaker embeddings q and the style encoder outputting prosodic style embeddings g from the input of target speaker mel-spectrogram M), access a speaker encoder configured to generate a speaker embedding corresponding to the new target speaker based on the acoustic features extracted from the new target reference speech (Azizah Figure 2 pg. 5900 shows the speaker encoder outputting speaker embeddings q from the input of target speaker mel-spectrogram M), access a text-to-speech module configured to generate a personalized voice corresponding to the new target speaker based on (i) the speaker embedding and (ii) the prosodic features extracted from the new target reference speech (Azizah Figure 2 pg. 5900 shows output mel-spectrogram Y' being created from the inputs of the speaker embeddings q and the style embeddings g, created from the input of the target speaker mel-spectrogram M input to both the speaker encoder and the style encoder), wherein the prosodic features are provided to the text-to-speech module as a separate input from the speaker embedding, and (Azizah Figure 2 pg. 5900 shows the prosodic style embeddings g being input into H to be concatenated, combining, or concatenating, versus keeping separate is indistinguishable, as from a concatenation, it would be obvious to instead provide them separately) without applying the text-to-speech module on new labeled training data associated with the new target speaker ("There are three speaker adaptation approaches in neural TTS: 1) Using speaker encoder network that is jointly trained with the entire TTS network [12]–[14] or neural vocoder [15]; 2) Using speaker embedding generated by the pre-trained speaker encoder and fine-tuning the TTS model [14]–[18]; 3) Using speaker embedding generated by the pre-trained speaker encoder without fine-tuning the TTS model [20]–[22]. Among these approaches only the third one is able to handle zero-shot speaker adaptation," Azizah pg. 5897 Section II B.), and generate the zero-shot personal text-to-speech model by compiling the feature extractor, the speaker encoder, and the text-to-speech module in a serial architecture (Azizah Figure 2 pg. 5900 shows a serial architecture, as the mel-spectrogram is extracted first, the speaker encoder and style encoder output their respective embeddings next, etc.) such separate inputs (Azizah Figure 2 pg. 5900 shows the prosodic style embeddings g being input into H to be concatenated, combining, or concatenating, versus keeping separate is indistinguishable, as from a concatenation, it would be obvious to instead provide them separately) to the text-to-speech module, thereby configuring the zero-shot personal text-to-speech model to generate the personalized voice for the new target speaker as model output in response to applying the zero-shot personal text-to-speech model (Azizah Figure 2 pg. 5900 shows the TTS model compiled with the feature extractors, the speaker encoder, and the TTS module so that, based on the input target mel-spectrogram, the speaker encoder creates speaker embedding q, the style encoder creates prosodic style embedding g, and these are then input into the decoder of the TTS model in order to output the mel-spectrogram Y'). However, Azizah fails to disclose (a) a first set of computer-executable instructions that are executable by one or more processors of a remote computing system for causing the remote computing system to at least and (b) a second set of computer-executable instructions that are executable by the at least one hardware processor for causing the system to send the first set of computer-executable instructions to the remote computing system. Beaufays teaches a method and system for improved efficiency in federated learning of machine learning models. Beaufays teaches (a) a first set of computer-executable instructions that are executable by one or more processors of a remote computing system for causing the remote computing system to at least (Beaufays Fig. 7 reference character 752 states that model data is stored remotely and updated), and (b) a second set of computer-executable instructions that are executable by the at least one hardware processor for causing the system to send the first set of computer-executable instructions to the remote computing system (Beaufays Fig. 7 reference character 760 shows transmitting the updated global model to the client (remote) devices). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Azizah’s disclosure of a zero-shot multilingual text-to-speech reconstruction system by including Beaufays’ teaching of transmitting instructions and the model to a remote computing system. This would allow for the TTS system to be more easily distributed and distilled to remote, or client, devices. It would allow the remote devices and systems to create the same structured TTS system on their respective devices and systems, facilitating the improvement and increase accuracy of a locally ran system over that of a decentralized model. Regarding claim 14, Azizah, in view of Beaufays, discloses all of the limitations of claim 13. However, Azizah fails to disclose wherein the first set of computer-executable instructions further include instructions for the remote computing system to execute the first set of computer-executable instructions for generating the zero-shot personal text-to-speech model. Beaufays teaches wherein the first set of computer-executable instructions further include instructions for the remote computing system to execute the first set of computer-executable instructions for generating the zero-shot personal text-to-speech model (Beaufays Fig. 7 reference character 760 shows transmitting the updated global model to the client (remote) devices, i.e. showing that the model can be transmitted from the global to the client). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Azizah’s disclosure of a zero-shot multilingual text-to-speech reconstruction system by including Beaufays’ teaching of transmitting instructions and the model to a remote computing system. This would allow for the TTS system to be more easily distributed and distilled to remote, or client, devices. It would allow the remote devices and systems to create the same structured TTS system on their respective devices and systems, facilitating the improvement and increase accuracy of a locally ran system over that of a decentralized model. Regarding claim 15, Azizah, in view of Beaufays, discloses all of the limitations of claim 14. Azizah further discloses wherein the first set of computer-executable instructions further include instructions for causing the remote system to, prior to generating the zero-shot personal text-to-speech model, apply the text-to-speech module to a multi-speaker multi-lingual training corpus to train the text-to-speech module ("To train the monolingual single-speaker TTS as the first source model we use LJSpeech for English [68], a 24 hours English transcribed speech corpus consisting of clips from a female with a sampling rate of 22050 Hz," Azizah pg. 5902-5903 Section IV) using a speaker cycle consistency training loss ("Some other works used cycle consistency training using automatic speech recognition (ASR) models to train TTS [38]–[40], where in the training process, ASR was used to find transcripts of speech sounds and TTS reconstructed transcripts into speech sounds," Azizah pg. 5897 Section II). Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to ADAM MICHAEL WEAVER whose telephone number is (571)272-7062. The examiner can normally be reached Monday-Friday, 8AM-5PM EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richemond Dorvil can be reached at (571) 272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ADAM MICHAEL WEAVER/ Examiner, Art Unit 2658 /RICHEMOND DORVIL/ Supervisory Patent Examiner, Art Unit 2658
Read full office action

Prosecution Timeline

Oct 04, 2024
Application Filed
Apr 08, 2026
Non-Final Rejection mailed — §102, §103
May 19, 2026
Applicant Interview (Telephonic)
May 20, 2026
Response Filed
May 22, 2026
Examiner Interview Summary
Sep 01, 2026
Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12664978
FEDERATED KNOWLEDGE DISTILLATION ON AN ENCODER OF A GLOBAL ASR MODEL AND/OR AN ENCODER OF A CLIENT ASR MODEL
3y 6m to grant Granted Jun 23, 2026
Patent 12657219
INFORMATION PROCESSING DEVICE, COMPUTER PROGRAM PRODUCT, AND INFORMATION PROCESSING METHOD
2y 3m to grant Granted Jun 16, 2026
Patent 12651117
METHODS AND SYSTEMS FOR VERIFICATION OF PLANT PROCEDURES' COMPLIANCE TO WRITING MANUALS
4y 0m to grant Granted Jun 09, 2026
Patent 12651266
SYSTEMS AND METHODS FOR RANKING CALL INTENT PROBABILITY
2y 3m to grant Granted Jun 09, 2026
Patent 12639355
IDENTIFYING HALLUCINATIONS IN LARGE LANGUAGE MODEL OUTPUT
2y 9m to grant Granted May 26, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
88%
Grant Probability
99%
With Interview (+31.2%)
2y 6m (~6m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 16 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month