Prosecution Insights
Last updated: October 02, 2026
Application No. 19/039,252

METHOD OF GENERATING SPEECH BASED ON NORMALIZING FLOW MODEL THAT GENERATES TIMBRE FROM TEXT

Non-Final OA §103§112
Filed
Jan 28, 2025
Priority
Feb 26, 2024 — RE 10-2024-0027572
Examiner
AZIZ, SHEZA ABDUL
Art Unit
Tech Center
Assignee
POSTECH Research and Business Development Foundation
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
17 currently pending
Career history
16
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Applicant claims the benefit of Korean Patent Application No. 10-2024-0027572, filed on February 26, 2024. Claims 1-11 have been afforded the benefit of February 26, 2024 filing date. Information Disclosure Statement The information disclosure statement(s) (IDS) submitted on 06 May 2025 and 28 January 2025 are being considered by the examiner. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims [5, 9, 10, 11 ] are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 5 recites “ a low-dimensional latent vector”. Low dimensional is a relative term, and the claim does not identify the dimensional threshold, the vector to which it is comparatively low dimensional, or an objective method for determining whether a vector is low-dimensional. Claim 9 introduces “a reference timbre embedding vector”. However, Claim 11 recites “the reference timbre embedding data”. It is unclear whether “the reference timbre embedding data” refers to the previously recited “a reference timbre embedding vector” or to a different data. Claim 10 recites “timbre that appears in the training speech data” renders the scope unclear because it is uncertain whether the text must describe the identity, vocal quality, acoustic characteristics, or perceived timbre of the speaker represented by the training speech data. “Corresponding to” or “represented in” would be clearer. Claim 11 recites “the reference timbre embedding data and the timbre embedding vector are passed through a projection layer”. It is unclear whether both embeddings are passed through a projection layer or each passes through a respective projection layer. This matters because the two input embeddings may originate in different modalities. Appropriate correction is required. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim [1, 3, 4, 7, 8 ] are rejected under 35 U.S.C. 103 as being unpatentable over Ezzerg (US12100383B1, hereinafter Ezzerg) in view of Zhang (Zhang, Yongmao, et al. "Promptspeaker: Speaker generation based on text descriptions." 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). IEEE, 2023., hereinafter Zhang). Regarding claim 1, Ezzerg teaches a method of generating synthetic speech data, comprising: [Abstract, lines 6-8 " For example, speech may be received at a first device and sent to a second device for output as audio"]; [Column 3, lines 10-13 "The machine learning model may use the speech attributes to transform the speech data to generate the latent representation"]; [Column 10, lines 64-66 "The method 400 may include receiving (430) speech attribute data corresponding to desired voice characteristics for synthesized speech."]; [Column 11, lines 2- 13 "The method 400 may include processing (440) the data representing the latent representation using a flow model and the speech attribute data to generate spectrogram data. The flow model may be, for example, the inverse flow model 160 described previously with reference to FIGS. 1 and 2 . If the speech attribute data used for generating the spectrogram data is the same as the speech attribute data used to generate the latent representation, then the voice characteristics of the resulting speech may be the same as those of the source speech. If the speech attribute data is different, the resulting speech may have different voice characteristics." where the speech attribute data represents the desired voice characteristic. ]; [Column 11, lines 38-48 "The method 400 may include processing (450) the second spectrogram data with a vocoder to generate an audio signal representing synthesized speech having the desired voice characteristics. The vocoder may receive the spectrogram data in the form of, for example, Mel-spectrograms. The vocoder may convert the spectrogram data into an audio signal representing a digitized analog signal (e.g., encoded in a pulse code modulation format). The vocoder (or other downstream component) may convert the audio data to an analog audio signal suitable for amplification and/or output from a loudspeaker."]. However, Ezzerg does not teach acquiring, by a data processing device, text data extracting, by the data processing device, timbre information from the text data timbre information But Zhang teaches acquiring, by a data processing device, text data extracting, by the data processing device, timbre information from the text data timbre information [Page 5, 4.Conculsions "PromptSpeaker encodes text descriptions as semantic representations and transforms the semantic representations into a novel speaker representation" where text descriptions are mapped to text data and encoding the description as a semantic representation and transforming it into a speaker representation in a speaker timbre space which builds a speaker voice profile from that mapped data (extracted timbre info)]; [Page 5, 3.3.2 Subjective evaluation results "As shown in Table 4, we can see that PromptSpeaker performs significantly better than baseline in gender accuracy and speaker timbre MOS. This also proves that there is a one-to-many mapping between text description and speaker timbre, and it is reasonable to establish a mapping between speaker representation and text prior distribution for speaker generation. We notice that there is still a gap between ground-truth and generated speaker in speaker timbre matching MOS. In other word, the generalization ability of the learned semantic-speaker-timbre space is limited due to limited training data on speaker text-prompt pairs." where the mapped distribution builds the speaker timbre]; It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Ezzerg with Zhang because both control the voice characteristics of generated speech. The combination would advantageously allow the timbre of input speech to be customized using a natural-language description, providing simpler and more intuitive voice control. Regarding claim 3, the rejection of claim 1 is incorporated Ezzerg does not teach the method of generating synthetic speech data of claim wherein the timbre information includes a timbre embedding vector, the timbre embedding vector is extracted from the text data using a context information extractor and a normalizing flow-training model, the context information extractor extracts a sentence embedding vector from the text data, and the normalizing flow-training model extracts the timbre embedding vector from the sentence embedding vector. However, Zhang teaches The method of generating synthetic speech data of claim 1, wherein the timbre information includes a timbre embedding vector, the timbre embedding vector is extracted from the text data using a context information extractor and a normalizing flow-training model, the context information extractor extracts a sentence embedding vector from the text data, and the normalizing flow-training model extracts the timbre embedding vector from the sentence embedding vector. ["Page 2, 1.1 Prompt Encoder "The prompt encoder extracts semantic information from speaker prompt with BERT [13] to predict the distribution of semantic prior representations. As shown in Fig.1 (c), the prompt encoder consists of a pre-trained BERT model, FFT blocks [14], a GRU layer, a token layer [15] and a linear layer. The text-based speaker prompt is first fed into the BERT model to extract semantic features. Then, we feed the semantic features into a GRU layer to compress the sequence features into a vector to obtain a global-level speaker timbre representation. " where text data is the speaker text based speaker prompt, the sentence embedding vector is mapped to sequence features into a vector, and global level speaker timbre representation is the timbre vector]; [Page 2, lines 1-4 "Then the Glow model transforms the semantic representation into a speaker representation, and finally the zero-shot TTS system synthesizes the generated speaker’s speech based on the generated speaker representation." where Glow is a type of Normalizing flow model]; [Page 3, 1.3 Glow "In the inference process, the semantic representations sampled from the prompt prior distribution is transformed by the inverse of Glow to obtain the speaker representation."]; [Page 5, 4.Conculsions "PromptSpeaker encodes text descriptions as semantic representations and transforms the semantic representations into a novel speaker representation" where text descriptions are mapped to text data and encoding the description as a semantic representation and transforming it into a speaker representation in a speaker timbre space which builds a speaker voice profile from that mapped data (extracted timbre info)]; [Page 5, 3.3.2 Subjective evaluation results "As shown in Table 4, we can see that PromptSpeaker performs significantly better than baseline in gender accuracy and speaker timbre MOS. This also proves that there is a one-to-many mapping between text description and speaker timbre, and it is reasonable to establish a mapping between speaker representation and text prior distribution for speaker generation. We notice that there is still a gap between ground-truth and generated speaker in speaker timbre matching MOS. In other word, the generalization ability of the learned semantic-speaker-timbre space is limited due to limited training data on speaker text-prompt pairs." where the mapped distribution builds the speaker timbre]; It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Ezzerg’s normalizing flow model with Zhang’s context information extractor to convert a textual description into timbre embedding data for controlling synthesized speech. This would advantageously provide flexible and accurate timbre customization from natural language input. Regarding claim 4, the rejection of claim 3 is incorporated Ezzerg does not teach the method of generating synthetic speech data of claim 3, wherein the context information extractor includes a pre-trained natural language processing (NLP) model. However, Zhang teaches the method of generating synthetic speech data of claim 3, wherein the context information extractor includes a pre-trained natural language processing (NLP) model. ["Page 2, 1.1 Prompt Encoder "The prompt encoder extracts semantic information from speaker prompt with BERT [13] to predict the distribution of semantic prior representations. As shown in Fig.1 (c), the prompt encoder consists of a pre-trained BERT model, FFT blocks [14], a GRU layer, a token layer [15] and a linear layer. The text-based speaker prompt is first fed into the BERT model to extract semantic features."]; It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Ezzerg’s system with Zhang’s pretrained BERT model to extract timbre-related semantic information from textual descriptions. This would advantageously improve understanding of natural -language voice descriptions without requiring the language model to be trained from scratch. Regarding claim 7, the rejection of claim 1 is incorporated Ezzerg teaches the method of generating synthetic speech data of claim 1, further comprising: outputting, by the data processing device, the generated synthetic speech data. [Column 11, lines 38-41"The method 400 may include processing (450) the second spectrogram data with a vocoder to generate an audio signal representing synthesized speech having the desired voice characteristics. "]. Regarding claim 8, Ezzerg teaches A data processing device comprising an input unit configured to an operation unit configured to and an output unit configured to output the generated speech data. [Abstract, lines 6-8 " For example, speech may be received at a first device and sent to a second device for output as audio"]; [Column 3, lines 10-13 "The machine learning model may use the speech attributes to transform the speech data to generate the latent representation"]; [Column 10, lines 64-66 "The method 400 may include receiving (430) speech attribute data corresponding to desired voice characteristics for synthesized speech."]; [Column 11, lines 2- 13 "The method 400 may include processing (440) the data representing the latent representation using a flow model and the speech attribute data to generate spectrogram data. The flow model may be, for example, the inverse flow model 160 described previously with reference to FIGS. 1 and 2 . If the speech attribute data used for generating the spectrogram data is the same as the speech attribute data used to generate the latent representation, then the voice characteristics of the resulting speech may be the same as those of the source speech. If the speech attribute data is different, the resulting speech may have different voice characteristics. " where the speech attribute data represents the desired voice characteristic.]; [Column 11, lines 38-48 "The method 400 may include processing (450) the second spectrogram data with a vocoder to generate an audio signal representing synthesized speech having the desired voice characteristics. The vocoder may receive the spectrogram data in the form of, for example, Mel-spectrograms. The vocoder may convert the spectrogram data into an audio signal representing a digitized analog signal (e.g., encoded in a pulse code modulation format). The vocoder (or other downstream component) may convert the audio data to an analog audio signal suitable for amplification and/or output from a loudspeaker."]. However, Ezzerg does not teach acquire text data an operation unit configured to extract timbre information from the text data But Zhang teaches acquire text data an operation unit configured to extract timbre information from the text data [Page 5, 4.Conculsions "PromptSpeaker encodes text descriptions as semantic representations and transforms the semantic representations into a novel speaker representation" where text descriptions are mapped to text data and encoding the description as a semantic representation and transforming it into a speaker representation in a speaker timbre space which builds a speaker voice profile from that mapped data (extracted timbre info)]; [Page 5, 3.3.2 Subjective evaluation results "As shown in Table 4, we can see that PromptSpeaker performs significantly better than baseline in gender accuracy and speaker timbre MOS. This also proves that there is a one-to-many mapping between text description and speaker timbre, and it is reasonable to establish a mapping between speaker representation and text prior distribution for speaker generation. We notice that there is still a gap between ground-truth and generated speaker in speaker timbre matching MOS. In other word, the generalization ability of the learned semantic-speaker-timbre space is limited due to limited training data on speaker text-prompt pairs." where the mapped distribution builds the speaker timbre]; It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Ezzerg with Zhang because both control the voice characteristics of generated speech. The combination would advantageously allow the timbre of input speech to be customized using a natural-language description, providing simpler and more intuitive voice control. Claim [2 ] are rejected under 35 U.S.C. 103 as being unpatentable over Ezzerg (US12100383B1, hereinafter Ezzerg) in view of Zhang (Zhang, Yongmao, et al. "Promptspeaker: Speaker generation based on text descriptions." 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). IEEE, 2023., hereinafter Zhang) and Yang (Yang, Dongchao, et al. "Instructtts: Modelling expressive tts in discrete latent space with natural language style prompt." IEEE/ACM Transactions on Audio, Speech, and Language Processing 32 (2024): 2913-2925, hereinafter Yang). Regarding claim 2, the rejection of claim 1 is incorporated. Ezzerg in view of Zhang does not teach The method of generating synthetic speech data of claim 1, wherein the text data includes text that embodies a speaker's mood through timbre. However, Yang does teach The method of generating synthetic speech data of claim 1, wherein the text data includes text that embodies a speaker's mood through timbre. ["Abstract "In this study, we attempt to use natural language as a style prompt to control the styles in the synthetic speech, e.g., “Sigh tone in full of sad mood with some helpless feeling"… we first construct a speech corpus whose speech samples are annotated with not only content transcriptions but also style descriptions in natural language.” where "sad mood" and "helpless feeling" is mapped to text expressing a speaker's mood .]; [Abstract "…to obtain a robust sentence embedding model that can effectively capture semantic information from the style prompts and control the speaking style in the generated speech."]. It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Ezzerg’s in view of Zhang’s system with Yang’s mood based textual description because Yang teaches using natural language mood and vocal tone descriptions to control synthesized speech. This would advantageously allow users to control both the timbre and emotional expression of the generated speech through text. Claim [ 5 ] are rejected under 35 U.S.C. 103 as being unpatentable over Ezzerg (US12100383B1, hereinafter Ezzerg) in view of Zhang (Zhang, Yongmao, et al. "Promptspeaker: Speaker generation based on text descriptions." 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). IEEE, 2023., hereinafter Zhang) and in further view of Kingma (Kingma, Durk P., and Prafulla Dhariwal. "Glow: Generative flow with invertible 1x1 convolutions." Advances in neural information processing systems 31 (2018), hereinafter Kingma). Regarding claim 5, the rejection of claim 3 is incorporated Ezzerg in view of Zhang do teach The method of generating synthetic speech data of claim 3, wherein Zhang teaches ["Page 2, 1. Method "As shown in Fig. 1, PromptSpeaker consists of a prompt encoder, a zero-shot TTS system and a Glow model. The prompt encoder produces the mean and variance for the semantic prior distribution."]; [Page 2, 1. Method "To generate a new speaker, we first sample from the semantic prior distribution and transform it to speaker representations by the Glow model. "]; [Page 2, 1.1 Prompt Encoder "The prompt encoder extracts semantic information from speaker prompt with BERT [13] to predict the distribution of semantic prior representations…The text-based speaker prompt is first fed into the BERT model to extract semantic features. Then, we feed the semantic features into a GRU layer to compress the sequence features into a vector to obtain a global-level speaker timbre representation."]; [Page 2, 1.1 Prompt Encoder "Subsequently a token layer is followed to further simplify the semantic features. The last linear layer finally produces the mean and variance of the semantic prior distribution."]; [Page 3, Prompt Speaker "The dimensions of both semantic representation and speaker representation are 256. "]; [Page 3, 1.3 Glow "In the inference process, the semantic representations sampled from the prompt prior distribution is transformed by the inverse of Glow to obtain the speaker representation."]; It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Ezzerg’s with Zhang’s system because using a normalizing flow model to generate a timbre embedding from a text-derived latent representation and supply that embedding as speech attribute data would advantageously enable flexible and intuitive timbre customization from natural language descriptions. Ezzerg in view of Zhang do not teach normalizing flow-training model is a model based on the change of variable theorem. However, Kingma teaches normalizing flow-training model is a model based on the change of variable theorem [Page 3, 2. Background: Flow-based Generative Models lines 5-6 " Such a sequence of invertible transformations is also called a (normalizing) flow (Rezende and Mohamed, 2015). Under the change of variables of eq. (4), the probability density function (pdf) of the model given a datapoint can be written as: PNG media_image1.png 105 411 media_image1.png Greyscale Where normalizing flow model is a model based on the change of variable theorem"]. It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Ezzerg in view of Zhang’ system with Kingma’s change of variables formulation because Kingma provides the established mathematical basis for training Glow models. This would advantageously enable exact likelihood calculation and reliable transformation between the text-derived latent vector and the timbre embedding. Claim [ 6 ] are rejected under 35 U.S.C. 103 as being unpatentable over Ezzerg (US12100383B1, hereinafter Ezzerg) in view of Zhang (Zhang, Yongmao, et al. "Promptspeaker: Speaker generation based on text descriptions." 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). IEEE, 2023., hereinafter Zhang) and in further view of Tang (Tang, Huaizhen, et al. "Tgavc: Improving autoencoder voice conversion with text-guided and adversarial training." 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). IEEE, 2021, hereinafter Tang). Regarding claim 6, the rejection of claim 3 is incorporated. Ezzerg in view of Zhang do teach The method of generating synthetic speech data of claim 3, wherein the generating of the synthetic speech data includes generating the synthetic speech data Zhang teaches ["Page 2, 1.1 Prompt Encoder "The prompt encoder extracts semantic information from speaker prompt with BERT [13] to predict the distribution of semantic prior representations. As shown in Fig.1 (c), the prompt encoder consists of a pre-trained BERT model, FFT blocks [14], a GRU layer, a token layer [15] and a linear layer. The text-based speaker prompt is first fed into the BERT model to extract semantic features. Then, we feed the semantic features into a GRU layer to compress the sequence features into a vector to obtain a global-level speaker timbre representation. " where text data is the speaker text based speaker prompt, the sentence embedding vector is mapped to sequence features into a vector, and global level speaker timbre representation is the timbre vector]; It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Ezzerg’s with Zhang’s system by using Zhang’s text generated timbre representation as Ezzerg’s speech attribute data when generating synthetic speech output would allow the output speech’s timbre to be controlled through a natural language description. Ezzerg in view of Zhang do not teach inputting the However, Tang teaches [Page 2, 2.1 AutoVC "The framework is composed of three modules: a content encoder Ec that outputs a content embedding Hc from source speech xi, a style encoder Es that outputs a style embedding Hs from target speech xj, and a decoder that produces the converted speech xi→j from the content and style embeddings."]; [Page 3, Column 1 Section 2.2 AutoVC "We input the source speech into the trained content encoder and the target speech into the trained style encoder, then we would get the content embedding of the source speech and the voice characteristic of the target speech, and the decoder could produce a converted speech with the linguistic information from source speech and the voice characteristic from target speech." ]; ["Page 4, 2.2.2 Network, Column 2 "The style embedding is first copied to the same length as the content embedding, and then concatenated with it in the channel dimension. The concatenated embedding is passed into the decoder to generate the speech."]. It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Ezzerg in view of Zhang with Tang’s decoder by substituting Zhang’s text generated timbre embedding vector for Tang’s style embedding when generating synthetic speech. Both embeddings represent voice characteristics used to condition speech generation. This substitution would advantageously allow the timbre of Ezzerg’s input speech to be controlled through a natural language description. Claim [ 9, 10, 11 ] are rejected under 35 U.S.C. 103 as being unpatentable over Zhang (Zhang, Yongmao, et al. "Promptspeaker: Speaker generation based on text descriptions." 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). IEEE, 2023., hereinafter Zhang) and in view of Elizalde (Elizalde, Benjamin, et al. "Clap learning audio concepts from natural language supervision." ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2023, hereinafter, Elizalde). Regarding claim 9, Zhang does teach A method of training a normalizing flow-training model, the method comprising: acquiring, by a data processing device, training speech data and training text data; [Abstract "We verify that PromptSpeaker can generate speakers new from the training set by objective metrics, and the synthetic speaker voice has reasonable subjective matching quality with the speaker prompt. "]; [2. Dataset "We built a multi-speaker Mandarin dataset with natural language descriptions of speaker characteristics." where Mandarin dataset is the training data]; [2. Datatset “Specifically, we first manually annotate the high-quality internal data from 74 stylistic speakers, as described in Table 1. “]. generating, by the data processing device, a reference timbre embedding vector from the training speech data using a timbre information extractor; [1.3 Glow " We use a speaker encoder pre-trained on the speaker classification task to extract speaker representations, but the distribution of speaker representations" where speaker representations are the speech derived reference embeddings]; [1.4 "During the training process, since there is a one-to-many mapping between text description and speaker timbre, we use the speaker representations extracted by the speaker encoder as the input of the TTS system. generating, by the data processing device, a sentence embedding vector from the training text data using a context information extractor; [1.1 Prompt Encoder , "The prompt encoder extracts semantic information from speaker prompt with BERT [13] to predict the distribution of semantic prior representations. As shown in Fig.1 (c), the prompt encoder consists of a pre-trained BERT model, FFT blocks [14], a GRU layer, a token layer [15] and a linear layer.The text-based speaker prompt is first fed into the BERT model to extract semantic features. Then, we feed the semantic features into a GRU layer to compress the sequence features into a vector to obtain a global-level speaker timbre representation. " where text data is the speaker text based speaker prompt, the sentence embedding vector is mapped to sequence features into a vector, and global level speaker timbre representation is the timbre vector]; generating, by the data processing device, a timbre embedding vector from the sentence embedding vector using a normalizing flow-training model; [1.3 Glow "In the training process, the Glow model transforms the speaker representation into a semantic representation that matches the prompt prior distribution. In the inference process, the semantic representations sampled from the prompt prior distribution is transformed by the inverse of Glow to obtain the speaker representation."]; [1.4 "During the training process, since there is a one-to-many mapping between text description and speaker timbre, we use the speaker representations extracted by the speaker encoder as the input of the TTS system. In this way, we can ensure that the input of the TTS system contains complete speaker information. In the inference process, we obtain the prior distribution based on the text description of a desired speaker and sample from the prior distribution. The Glow model transforms the semantic representation into the speaker representation. The zero-shot TTS system synthesizes the generated speaker’s speech based on the input text and the generated speaker representation."]; Zhang does not teach calculating, by the data processing device, a loss value between the reference timbre embedding vector and the timbre embedding vector; and updating, by the data processing device, a parameter of the normalizing flow- training model so that the calculated loss value is minimized, . However, Elizalde teaches calculating a loss value between the reference timbre embedding vector and the timbre embedding vector; and updating, by the data processing device, a parameter of the ["Page 2, Section 2.1 Contrastive Language Audio Pretraining "Now that the audio and text embeddings (Ea, Et) are comparable, we can measure similarity:'….We used this symmetric cross entropy loss (L) over the similarity matrix to jointly train the audio encoder and the text encoder along with their linear projections." where training models are updated based on the loss]; It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Zhang with Elizalde by using Elizalde’s cross modal similarity loss to train Zhang’s normalizing flow model based on Ezzerg’s training speech data and corresponding text data. This would advantageously improve alignment between the text generated timbre embedding and the timbre represented in the corresponding speech. Regarding claim 10, the rejection of claim 9 is incorporated. Zhang teaches the method of claim 9, wherein the training text data includes text expressing timbre that appears in the training speech data. [2. Dataset "We built a multi-speaker Mandarin dataset with natural language descriptions of speaker characteristics." where Mandarin dataset is the training data]; [2.Dataset " Specifically, we first manually annotate the high-quality internal data from 74 stylistic speakers as described in Table 1. Different people may have different descriptions of the same speaker, so each speaker is annotated by 13 annotators, reflecting the one-to many mapping between text description and speaker timbre." where the descriptions are annotations of the speakers whole recordings constitute the training speech data and therefore express timbre appearing in that training speech]; [Table 1: Example is "a husky voice from a middle-aged man"]; Regarding claim 11, the rejection of claim 9 is incorporated. Zhang does teach the method of claim 10, wherein the reference timbre embedding data and the timbre embedding vector [Page 3, Prompt Speaker "The dimensions of both semantic representation and speaker representation are 256. " where both embedding vectors have the same dimensions.]; Zhang does not teach they are passed through a projection layer But, Elizalde teaches they are passed through a projection layer [2. Method "The input is audio and text pairs passed to an audio encoder and a text encoder. Both representations are connected in joint multimodal space with linear projections."]; [2.1 Contrastive Language-Audio Pretraining "We brought audio and text representations, ˆXa and ˆXt, into a joint multimodal space of dimension d by using a learnable linear projection:.. where Ea ∈ RN× d, Et ∈ RN× d, La and Lt are the linear projections for audio and text respectively."] It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Zhang with Elizalde by projecting the speech derived reference timbre embedding and the text generated timbre embedding into the same dimensional space. This would advantageously make the embeddings directly compatible for calculating the similarity loss and training the model. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHEZA ABDUL AZIZ whose telephone number is (571)272-9610. The examiner can normally be reached Monday-Friday 7:30am-5pm Alternate Fridays off. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at (571) 272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SHEZA ABDUL AZIZ/Examiner, Art Unit 2657 /DANIEL C WASHBURN/Supervisory Patent Examiner, Art Unit 2657
Read full office action

Prosecution Timeline

Jan 28, 2025
Application Filed
Aug 26, 2026
Non-Final Rejection mailed — §103, §112 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month