Prosecution Insights
Last updated: October 01, 2026
Application No. 19/073,736

AUDIO TRANSLATION WITH PRESERVED SPEAKER CHARACTERISTICS

Non-Final OA §102§103
Filed
Mar 07, 2025
Priority
Mar 08, 2024 — provisional 63/563,066 +2 more
Examiner
SWAMY, ARJUN RAJ
Art Unit
Tech Center
Assignee
Roblox Corporation
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
14 currently pending
Career history
11
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§102 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless –(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claim(s) 1, 8, 9, 16 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Jia(US PGPub 20230013777). Regarding Claim 1, Jia teaches a computer-implemented method of audio translation comprising: receiving an audio stream from a first user associated with a first client device(source speaker 104, device 110[Figure 1]), wherein the audio stream is spoken in a first language by the first user(At operation 502, the method 500 includes receiving an input speech representation 102 that corresponds to an utterance 108 spoken by a source speaker 104 in a first language.[0038]); retrieving translation data associated with a second user, wherein the translation data includes at least a language preference associated with the second user(target language[0024], user 118 that natively speaks English[0024]), and wherein the second user is associated with a second user device(recipient user 118, device 116 [Figure 1]); converting a first portion of the audio stream received from the first user into a plurality of phonemes of a second language, wherein the second language is defined by the language preference(The operations also include predicting, by the decoder, a phoneme representation that corresponds to a translation of the utterance in a second different language.[0014]); predicting a respective duration of each of the phonemes in the plurality of phonemes(the synthesizer includes a duration model network configured to predict a duration of each phoneme in a sequence of phonemes represented by the phoneme representation[0012]); outputting, by a synthesizer, a first portion of output speech that includes the plurality of phonemes where each of the phonemes in the plurality of phonemes has the respective duration(synthesizer may be configured to generate the translated synthesized speech representation by upsampling the sequence of phonemes based on the predicted duration of each phoneme. The translated synthesized speech representation may be configured to a speaking style/prosody of the source speaker.[0012]); and providing the first portion of output speech to the second user device(output audio data (e.g., mel-spectrogram) 106 corresponding to a translated synthesized speech representation of a translated utterance 114 spoken in a different second language[0022]); wherein additional portions of the output speech are output based on subsequent portions of the audio stream(output audio data (e.g., mel-spectrogram) 106 corresponding to a translated synthesized speech representation of a translated utterance 114 spoken in a different second language[0022], autoregressive and generates, at each output step, the probability distribution of possible phonemes for the given output step based on each previous phoneme in the phoneme sequence 245 selected by the Softmax 240 during each of the previous output steps[0032]). Claim 16 recites similar limitations and is rejected under the same rationale. Regarding Claim 8, Jia, as in Claim 1, teaches outputting the first portion of output speech by the synthesizer includes outputting hidden states corresponding to the plurality of phonemes(The operations also include encoding the input speech representation into a hidden feature representation by an encoder of the S2ST model. The operation also include generating, by a decoder of the S2ST model, a context vector that attends to hidden feature representation encoded by the encoder.[0014]), and further comprising, before providing the first portion of output speech to the second user device, transforming the hidden states corresponding to the plurality of phonemes into audio using a vocoder(a vocoder configured to receive the translated synthesized speech representation and synthesize the translated synthesized speech representation into an audible output of the translated synthesized speech representation.[0013]). Regarding Claim 9, Jia, as in Claim 1, teaches the audio stream is associated with a voice chat function of a virtual experience(Figure 1 discloses voice chat of a virtual experience, In this example, the source speaker 104 and the user 118 are speaking with each other through their respective computing devices 110, 116, such as over an audio/video call (e.g., video meeting/chat) telephone call or other type of voice communication protocol, for example, voice over internet protocol.[0026]). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 2-4, 6, 17-19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Jia(US PGPub 20230013777) in view of Shen(Non-Attentive Tacotron: Robust and Controllable Neural TTS Synthesis Including Unsupervised Duration Modeling). Regarding Claim 2, Jia teaches (Here, the input audio data 102 includes a sequence of input spectrograms that correspond to the utterance 108 spoken by the source speaker 104 in the source/first language (e.g., Spanish). The sequence of input phonemes may include an 80-channel mel-spectrogram sequence.[0031]). Additionally Jia teaches outputting, with the encoder, a vector representation of the audio stream, wherein the first portion of the audio stream that is converted is the vector representation of the audio stream(The encoder 210 is configured to encode the input audio data 102 into a hidden feature representation (e.g., a series of vectors) 215[0031]). Jia does not teach providing the first portion the audio stream as input to a tokenizer; However, Shen teaches (a feature generation network that transforms input tokens (e.g., grapheme or phoneme ids) into acoustic features (e.g., mel-spectrogram)[3. Model]). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention with the teachings of Jia to incorporate the use of a tokenizer to generate a mel spectrogram because it would lead to significantly better robustness with naturalness matching recorded natural speech(Shen Introduction). Claim 17 recites similar limitations as Claim 2 and is rejected under the same rationale. Regarding Claim 3, Jia teaches outputting, with a phoneme decoder, a first predicted phoneme of the plurality of phonemes(autoregressive and generates, at each output step, the probability distribution of possible phonemes for the given output step based on each previous phoneme in the phoneme sequence 245 selected by the Softmax 240 during each of the previous output steps[0032]); generating a first query from the first predicted phoneme(Attention Module 220); providing the first query, and a first key and a first value based on the vector representation of the audio stream(hidden feature representation 215 encoded by the encoder 210[0032]) to the phoneme decoder(In some implementations, the decoder 230 includes a stack of long short-term memory (LSTM) cells assisted by the attention module 220[0032]); outputting, with the phoneme decoder, a subsequent predicted phoneme of the plurality of phonemes(autoregressive and generates, at each output step, the probability distribution of possible phonemes for the given output step based on each previous phoneme in the phoneme sequence 245 selected by the Softmax 240 during each of the previous output steps[0032]); generating a subsequent query from the subsequent predicted phoneme(Attention Module 220); providing the subsequent query, and a subsequent key and a subsequent value based on the vector representation of the audio stream(hidden feature representation 215 encoded by the encoder 210[0032]) to the phoneme decoder((In some implementations, the decoder 230 includes a stack of long short-term memory (LSTM) cells assisted by the attention module 220[0032])); and continuing to predict phonemes with the phoneme decoder until a remaining vector representation of the audio stream is processed(autoregressive and generates, at each output step, the probability distribution of possible phonemes for the given output step based on each previous phoneme in the phoneme sequence 245 selected by the Softmax 240 during each of the previous output steps[0032]). Claim 18 recites similar limitations as Claim 3 and is rejected under the same rationale. Regarding Claim 4, Jia as in Claim 2 teaches receiving, at an attention layer(Attention Module 220) and from a phoneme decoder(Decoder 230), a plurality of vectors that correspond to the plurality of phonemes(phoneme representation 235); converting, by the attention layer, the plurality of vectors into respective queries(Attention Module 220, Note: converting vectors to queries is what an attention layer performs); and generating, by the attention layer, output feature vectors based on the plurality of vectors that correspond to the plurality of phonemes and the vector representation of the audio stream(the synthesizer 300 may receive the phoneme representation 235 and the context vector 225, Note: vector 225 comes from the attention block and is based on the plurality of phenomes). Claim 19 recites similar limitations as Claim 4 and is rejected under the same rationale. Regarding Claim 6, Jia teaches providing the output feature vectors(225 [Figure 3]) to a duration predictor(Duration Predictor 310 [Figure 3]), wherein predicting the respective duration of each of the phonemes in the plurality of phonemes is performed using the duration predictor(the duration modeling network 310 is tasked with predicting a duration 315 for each phoneme in the phoneme representation 235 corresponding to the output audio data 106 that represents the translated synthesized speech representation in the target/second language.[0034]). Claim(s) 5, 20 are rejected under 35 U.S.C. 103 as being unpatentable over Jia(US PGPub 20230013777) in view of Shen(Non-Attentive Tacotron: Robust and Controllable Neural TTS Synthesis Including Unsupervised Duration Modeling) as applied to claim 4 above, and further in view of Fang(DASpeech: Directed Acyclic Transformer for Fast and High-quality Speech-to-Speech Translation). Regarding Claim 5, neither Jia nor Shen teach providing the output feature vectors to a variance predictor that predicts a variance of each phoneme of the plurality of phonemes. However, Fang teaches a variance predictor that predicts a variance of each phoneme of the plurality of phonemes(The variance adaptor contains three variance predictors including duration predictor(Jia teaches the duration predictor 310 takes input from attention module via 225), pitch predictor, and energy predictor, which are used to reduce the information gap between input phoneme sequences and output mel-spectrograms[2.2 FastSpeech 2]). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention with the teachings of Jia and Shen to further incorporate the variance predictor of Fang because it would reduce the information gap between input phoneme sequences and output mel-spectrograms[Fang 2.2] Claim 20 recites similar limitations as Claim 5 and is rejected under the same rationale. Claim(s) 7 is rejected under 35 U.S.C. 103 as being unpatentable over Jia(US PGPub 20230013777) in view of Shen(Non-Attentive Tacotron: Robust and Controllable Neural TTS Synthesis Including Unsupervised Duration Modeling) as applied to claim 4 above, and further in view of Chen(WaveGrad 2: Iterative Refinement for Text-to-Speech Synthesis). Regarding Claim 7, Jia in view of Shen teaches performing Gaussian upsampling(Introduction of Gaussian upsampling significantly improving the naturalness compared to vanilla upsampling through Repetition[1. Introduction]) of the output feature vectors. Neither Jia nor Shen teach the upsampling is performed to an input rate of the synthesizer. However, Chen teaches the upsampling is performed to an input rate of the synthesizer (decoder gradually upsamples the hidden representations to match the waveform resolution. In our case, the waveform is sampled at 24 kHz and we need to upsample by 300 times[3.4 Decoder]). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention with the teachings of Jia and Shen to further incorporate the rate of upsampling of Chen because it would generate high fidelity audio(Chen Abstract). Claim(s) 10 is rejected under 35 U.S.C. 103 as being unpatentable over Jia(US PGPub 20230013777) in view of Shen(Non-Attentive Tacotron: Robust and Controllable Neural TTS Synthesis Including Unsupervised Duration Modeling) and further in view of Zhang(A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI). Regarding Claim 10, Jia does not teach training a first synthesizer using a synthesizer loss; replacing the first synthesizer with a diffusion synthesizer; and fine-tuning the diffusion synthesizer. However Shen teaches training a first synthesizer using a synthesizer loss(mel-spectrogram reconstruction loss is a L1 +L2 loss between the predicted and the groundtruth mel-spectrogram[3. Model]). Additionally, Zhang teaches a diffusion synthesizer(diffusion model in speech synthesis[2.2 Background on Diffusion Model]); and fine-tuning the diffusion synthesizer(Moreover, Guided-TTS 2 [38] adapts the pretrained diffusion model to target speakers with classifier-free guidance and also finetunes the pretrained diffusion model with a short reference speech of the target speaker directly[3.2.3 Adaptive Modeling for multi-speaker setting]). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention with the teachings of Jia to incorporate the use of synthesizer loss of Shen and the diffusion synthesizer of Zhang because it would lead to significantly better robustness with naturalness matching recorded natural speech(Shen Introduction) and because it would improve the quality of audio(Zhang 4. Speech Enhancement). Claim(s) 11 is rejected under 35 U.S.C. 103 as being unpatentable over Jia(US PGPub 20230013777) in view of Fujita(Speech Rhythm-Based Speaker Embeddings Extraction from Phonemes and Phoneme Duration for Multi-Speaker Speech Synthesis). Regarding Claim 11, Jia teaches a trained machine-learning system comprising: an encoder implemented by one or more processors, the encoder trained to perform operations comprising receiving an audio stream spoken in a first language and outputting encoded audio(an encoder configured to receive an input speech representation that to an utterance spoken by a source speaker in a first language and encode the input speech representation into a hidden feature representation.[Abstract]); a phoneme decoder implemented by the one or more processors, the phoneme decoder trained to perform operations comprising receiving the encoded audio from the encoder and converting a first portion of the encoded audio into a plurality of phonemes of a second language(a decoder configured to receive the context vector generated by the attention module and predict a phoneme representation that corresponds to a translation of the utterance in a second different language.[Abstract]); a duration predictor implemented by the one or more processors, that is trained to perform operations comprising receiving the plurality of phonemes from the phoneme decoder and predicting a respective duration of respective phonemes in the plurality of phonemes(a duration model network configured to predict a duration of each phoneme in a sequence of phonemes represented by the phoneme representation.[0012]); and a synthesizer implemented by the one or more processors, the synthesizer trained to perform operations comprising outputting the first portion of output speech that includes the plurality of phonemes where each of the phonemes in the plurality of phonemes has the respective duration(synthesizer may be configured to generate the translated synthesized speech representation by upsampling the sequence of phonemes based on the predicted duration of each phoneme. The translated synthesized speech representation may be configured to a speaking style/prosody of the source speaker.[0012]). Jia does not teach a duration predictor including a transformer encoder. However Fujita teaches a duration predictor including a transformer encoder(The Transformer encoder extracts features by looking at the entire input sequence. This feature extraction is suitable for speech rhythm because the speech rhythm feature derives from a long speech context.[3.2.3 Transformer Encoder Block], speech rhythm representing features, i.e., pairs of phonemes and their durations[3.2.1 Input Features]). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention with the teachings of Jia to incorporate the transformer encoder of Fujita because it would derive the features from a long speech context[Fujita 3.2.3]. Claim(s) 12, 13 are rejected under 35 U.S.C. 103 as being unpatentable over Jia(US PGPub 20230013777) in view of Fujita(Speech Rhythm-Based Speaker Embeddings Extraction from Phonemes and Phoneme Duration for Multi-Speaker Speech Synthesis) as applied to claim 11 above, and further in view of Fang(DASpeech: Directed Acyclic Transformer for Fast and High-quality Speech-to-Speech Translation). Regarding Claim 12, Jia teaches the duration predictor using a per-phoneme L2 duration loss(During training, a loss term (e.g., L2 loss term) is then determined between the predicted phoneme durations and the target average duration[0034]). Jia also teaches training using an overall loss(Implementations herein are directed toward a robust direct S2ST model that is trained end-to-end[0020]) Neither Jia nor Fujita teach training the decoder using a decoder cross-entropy loss, the synthesizer using a synthesizer loss. However, Fang teaches training the decoder using a decoder cross-entropy loss(minimizing the negative log-likelihood loss: LDAT[2.1 Directed Acyclic Transformer]), the synthesizer using a synthesizer loss(LL1 measures the L1 distance between the predicted and ground truth mel-spectrograms[2.2 FastSpeech 2]). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention with the teachings of Jia and Fujita to incorporate the cross-entropy and synthesizer loss of Fang because it would achieve comparable performance while being faster[Fang Abstract]. Regarding Claim 13, Jia in view of Fang teaches overall loss is a weighted sum of the decoder loss, the duration loss, and the synthesizer loss(training objective of DASpeech is as follows: LDASpeech = LDAT +µ·LTTS, where µ is the weight of TTS loss[3.2 Training]). Claim(s) 14 is rejected under 35 U.S.C. 103 as being unpatentable over Jia(US PGPub 20230013777) in view of Fujita(Speech Rhythm-Based Speaker Embeddings Extraction from Phonemes and Phoneme Duration for Multi-Speaker Speech Synthesis) in view of Fang(DASpeech: Directed Acyclic Transformer for Fast and High-quality Speech-to-Speech Translation) as applied to claim 12 above, and further in view of Popuri(Enhanced Direct Speech-to-Speech Translation Using Self-supervised Pre-training and Data Augmentation). Regarding Claim 14, Jia teaches training the phoneme decoder, the duration predictor, and the synthesizer together includes training on speech to speech translation tasks(Implementations herein are directed toward a robust direct S2ST model that is trained end-to-end[0020], the S2ST model is trained on pairs of parallel source language and target language utterances[0016]). Neither Jia, Fujita nor Fang teach training on speech to text translation tasks. However Popuri teaches training(train the supervised S2UT models with multitasks[A.3 Supervised S2UT Baselines]) on speech to text translation tasks and speech to speech translation tasks(For auxiliary tasks, we have a Transformer decoder on the sixth layer of the encoder for source character prediction and a Transformer decoder on the eighth layer of the encoder for target character prediction[A.3 Supervised S2UT Baselines]). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention with the teachings of Jia, Fujita and Fang to incorporate the auxillary task training of Popuri because it would improve the model’s performance on translation tasks[Popuri Abstract] Claim(s) 15 is rejected under 35 U.S.C. 103 as being unpatentable over Jia(US PGPub 20230013777) in view of Fujita(Speech Rhythm-Based Speaker Embeddings Extraction from Phonemes and Phoneme Duration for Multi-Speaker Speech Synthesis) in view of Fang(DASpeech: Directed Acyclic Transformer for Fast and High-quality Speech-to-Speech Translation) as applied to claim 12 above, in view of Popuri(Enhanced Direct Speech-to-Speech Translation Using Self-supervised Pre-training and Data Augmentation) and in further view of Lin(Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken Conversations). Regarding Claim 15, neither Jia, Fujita nor Fang teach encoder and the phoneme decoder are trained using synthetic training data that is generated by: generating text in a style associated with a virtual experience from a chatbot; translating the text to source audio in one or more different languages; and using the text as ground truth data. However, Popuri teaches translating the text to source audio in one or more different languages; and using the text as ground truth data(we apply MT and TTS to prepare weakly supervised S2ST data from speech in the source language). Additionally, Lin teaches generating text in a style associated with a virtual experience(we prompt the GPT-4 with 17 common daily dialogue topics: school, work, family, health, entertainment, travel, food, sports, finance, technology, music, movies, books, games, beauty, shopping, and weather.[2.2.1 LLM for Data Generation]) from a chatbot(leveraging GPT-4 (OpenAI, 2023) to generate spoken dialogue set consisting of a dialogue context, the same sentence presented in three different speaking styles, and three corresponding responses.[2.2.1 LLM for Data Generation]) It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention with the teachings of Jia, Fujita and Fang to incorporate the auxillary task training using synthetic data of Popuri and the style-based text generation of Lin because it would improve the model’s performance on translation tasks[Popuri Abstract] and because it would help the models learn to use speaking styles[Lin 2.1 Overview]. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to ARJUN R SWAMY whose telephone number is (571)272-9763. The examiner can normally be reached Mon-Fri 8-5. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Hai Phan can be reached at (571) 272-6338. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ARJUN SWAMY/Examiner, Art Unit 2654 /HAI PHAN/Supervisory Patent Examiner, Art Unit 2654
Read full office action

Prosecution Timeline

Mar 07, 2025
Application Filed
Sep 10, 2026
Non-Final Rejection mailed — §102, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month