Prosecution Insights
Last updated: August 17, 2026
Application No. 18/958,501

VOICE-PRESERVING MULTI-LINGUAL SPEECH AUDIO TRANSLATION

Non-Final OA §102§103§112
Filed
Nov 25, 2024
Examiner
MARLOW, ALEXANDER G
Art Unit
2658
Tech Center
2600 — Communications
Assignee
The Trustees of Princeton University
OA Round
1 (Non-Final)
79%
Grant Probability
Favorable
1-2
OA Rounds
1y 0m
Est. Remaining
97%
With Interview

Examiner Intelligence

Grants 79% — above average
79%
Career Allowance Rate
66 granted / 84 resolved
+16.6% vs TC avg
Strong +18% interview lift
Without
With
+18.1%
Interview Lift
resolved cases with interview
Typical timeline
2y 8m
Avg Prosecution
7 currently pending
Career history
90
Total Applications
across all art units

Statute-Specific Performance

§101
16.9%
-23.1% vs TC avg
§103
50.3%
+10.3% vs TC avg
§102
16.2%
-23.8% vs TC avg
§112
10.2%
-29.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 84 resolved cases

Office Action

§102 §103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Introduction This office action is in response to communications filed 11/25/2024. Claims 1-20 are pending and likewise have been examined. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 3, 10 and 17 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claims 3, 10 and 17 recites the limitation "the plurality of translated segments" in lines 3, 4 and 3, respectively. There is insufficient antecedent basis for this limitation in the claim. The limitations seems to be referring to “generating a translation for each source segment” in claim 2 and duplicates. However, claims 3, 10 and 17 also recites “generating translated segments” before “the plurality of translated segments”, which causes uncertainty as to which previous limitation is being referred to. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claim(s) 1, 4-5, 8, 11-12, 15 and 18-19 is/are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Waibel (US 20250315631 A1). Regarding Claim 1: Waibel teaches a method comprising: receiving an input audio sequence, the input audio sequence including speech audio having a source vocal identity in a first language(Para [0009], Ln 1-15, system comprises a remote source for capturing input audio by the output speaker in a first language that is different from the target language and converting speech by the output speaker in the input audio into text in the first language. Para [0010], Ln 1-13, Generating the adapted speech comprises adapting the speech in to target language to voice characteristics of the first speaker in the input audio); translating a first transcription of the speech audio to a second transcription of the speech audio in a second language(Para [0012], Ln 1-20, generating, by a translation module, trained through machine learning, of the computer system, a textual translation into the target language from the text in the first language from the remote source); generating, by a text-to-speech model, initial translated speech audio in the second language using the second transcription of the speech audio, the initial translated speech audio having a default vocal identity(Para [0054], Ln 1-16, Subsequently, the TTS module 18 is given this translated text and the resulting Mel spectrogram is turned into a waveform file by the HiFi-GAN vocoder. The final audio is now created by the voice conversion module 20, which gets the waveform of German speech that the vocoder produced as input and uses the original English audio of the input video as target speaker. Abstract, Ln 1-14, The translated text is then synthesized by a Text-to-Speech model that recreates the original emphases in the target sentence. The resulting synthetic speech is then converted back to the original speakers' voice using a voice conversion model); processing the initial translated speech audio to generate a translated content embedding and translated intonation data(Para [0045], Ln 1-13, For voice conversion 20, VQMIVC (Vector quantization mutual information voice conversion) can be used, which uses a straightforward autoencoder architecture to solve the voice conversion issue. The framework consists of four modules: a content encoder that produces a content embedding from speech, a speaker encoder that produces a speaker embedding (D-vector) from speech, a pitch encoder that produces prosody embedding from speech, and a decoder that generates from content, prosody, and speaker embeddings. Para [0046], Ln 1-15, To extract target speaker embedding, the target speech is sent into the speaker encoder. Finally, the decoder reconstructs the converted speech using the source speech's content embedding and prosody embedding and the target speech's speaker embedding); processing the input audio sequence to generate a source speaker embedding representing the source vocal identity of the speech audio(Para [0045], Ln 1-13, For voice conversion 20, VQMIVC (Vector quantization mutual information voice conversion) can be used, which uses a straightforward autoencoder architecture to solve the voice conversion issue. The framework consists of four modules: a content encoder that produces a content embedding from speech, a speaker encoder that produces a speaker embedding (D-vector) from speech, a pitch encoder that produces prosody embedding from speech, and a decoder that generates from content, prosody, and speaker embeddings. Para [0046], Ln 1-15, To extract target speaker embedding, the target speech is sent into the speaker encoder. Finally, the decoder reconstructs the converted speech using the source speech's content embedding and prosody embedding and the target speech's speaker embedding); and generating final translated speech audio using the source speaker embedding, the translated content embedding, and the translated intonation data, the final translated speech audio having the source vocal identity in the second language(Para [0045], Ln 1-13, For voice conversion 20, VQMIVC (Vector quantization mutual information voice conversion) can be used, which uses a straightforward autoencoder architecture to solve the voice conversion issue. The framework consists of four modules: a content encoder that produces a content embedding from speech, a speaker encoder that produces a speaker embedding (D-vector) from speech, a pitch encoder that produces prosody embedding from speech, and a decoder that generates from content, prosody, and speaker embeddings. Para [0046], Ln 1-15, To extract target speaker embedding, the target speech is sent into the speaker encoder. Finally, the decoder reconstructs the converted speech using the source speech's content embedding and prosody embedding and the target speech's speaker embedding. See also Claim 1 of Waibel, Ln 15-18, wherein the voice characteristics comprise pitch). Regarding Claim 4: Waibel teaches the method of claim 1, wherein processing the initial translated speech audio to generate a translated content embedding and translated intonation data further comprises: generating, by an encoder model, the translated content embedding representing speech content of the initial translated speech audio(Para [0045], Ln 1-13, For voice conversion 20, VQMIVC (Vector quantization mutual information voice conversion) can be used, which uses a straightforward autoencoder architecture to solve the voice conversion issue. The framework consists of four modules: a content encoder that produces a content embedding from speech, a speaker encoder that produces a speaker embedding (D-vector) from speech, a pitch encoder that produces prosody embedding from speech, and a decoder that generates from content, prosody, and speaker embeddings); and generating, by a pitch detector, the translated intonation data(Para [0045], Ln 1-13, For voice conversion 20, VQMIVC (Vector quantization mutual information voice conversion) can be used, which uses a straightforward autoencoder architecture to solve the voice conversion issue. The framework consists of four modules: a content encoder that produces a content embedding from speech, a speaker encoder that produces a speaker embedding (D-vector) from speech, a pitch encoder that produces prosody embedding from speech, and a decoder that generates from content, prosody, and speaker embeddings. Para [0046], Ln 1-15, To extract target speaker embedding, the target speech is sent into the speaker encoder. Finally, the decoder reconstructs the converted speech using the source speech's content embedding and prosody embedding and the target speech's speaker embedding). Regarding Claim 5: Waibel teaches the method of claim 1, wherein processing the input audio sequence to generate a source speaker embedding representing the source vocal identity of the speech audio further comprises: generating, by an encoder model, the source speaker embedding representing the source vocal identity of the speech audio by processing a mel-spectrogram of the speech audio(Para [0045], Ln 1-13, For voice conversion 20, VQMIVC (Vector quantization mutual information voice conversion) can be used, which uses a straightforward autoencoder architecture to solve the voice conversion issue. The framework consists of four modules: a content encoder that produces a content embedding from speech, a speaker encoder that produces a speaker embedding (D-vector) from speech. Para [00051], Ln 10-14, The audio data can be send to audio encoder after a mel-spectrogram representation is obtained of the corresponding audio sequence). Regarding Claim 8: Waibel teaches a non-transitory computer-readable medium storing executable instructions, which when executed by a processing device, cause the processing device to perform operations comprising(Para [0133], Ln 1-7, computer system according to the present invention comprises one or more processor cores and a memory in communication with the one or more processor cores. The memory stores instructions that when executed by the one or more processor cores, cause the one or more processor cores. Para [0079], Ln 1-10, RAM): receiving an input audio sequence, the input audio sequence including speech audio having a source vocal identity in a first language(Para [0009], Ln 1-15, system comprises a remote source for capturing input audio by the output speaker in a first language that is different from the target language and converting speech by the output speaker in the input audio into text in the first language. Para [0010], Ln 1-13, Generating the adapted speech comprises adapting the speech in to target language to voice characteristics of the first speaker in the input audio); translating a first transcription of the speech audio to a second transcription of the speech audio in a second language(Para [0012], Ln 1-20, generating, by a translation module, trained through machine learning, of the computer system, a textual translation into the target language from the text in the first language from the remote source); generating, by a text-to-speech model, initial translated speech audio in the second language using the second transcription of the speech audio, the initial translated speech audio having a default vocal identity(Para [0054], Ln 1-16, Subsequently, the TTS module 18 is given this translated text and the resulting Mel spectrogram is turned into a waveform file by the HiFi-GAN vocoder. The final audio is now created by the voice conversion module 20, which gets the waveform of German speech that the vocoder produced as input and uses the original English audio of the input video as target speaker. Abstract, Ln 1-14, The translated text is then synthesized by a Text-to-Speech model that recreates the original emphases in the target sentence. The resulting synthetic speech is then converted back to the original speakers' voice using a voice conversion model); processing the initial translated speech audio to generate a translated content embedding and translated intonation data(Para [0045], Ln 1-13, For voice conversion 20, VQMIVC (Vector quantization mutual information voice conversion) can be used, which uses a straightforward autoencoder architecture to solve the voice conversion issue. The framework consists of four modules: a content encoder that produces a content embedding from speech, a speaker encoder that produces a speaker embedding (D-vector) from speech, a pitch encoder that produces prosody embedding from speech, and a decoder that generates from content, prosody, and speaker embeddings. Para [0046], Ln 1-15, To extract target speaker embedding, the target speech is sent into the speaker encoder. Finally, the decoder reconstructs the converted speech using the source speech's content embedding and prosody embedding and the target speech's speaker embedding); processing the input audio sequence to generate a source speaker embedding representing the source vocal identity of the speech audio(Para [0045], Ln 1-13, For voice conversion 20, VQMIVC (Vector quantization mutual information voice conversion) can be used, which uses a straightforward autoencoder architecture to solve the voice conversion issue. The framework consists of four modules: a content encoder that produces a content embedding from speech, a speaker encoder that produces a speaker embedding (D-vector) from speech, a pitch encoder that produces prosody embedding from speech, and a decoder that generates from content, prosody, and speaker embeddings. Para [0046], Ln 1-15, To extract target speaker embedding, the target speech is sent into the speaker encoder. Finally, the decoder reconstructs the converted speech using the source speech's content embedding and prosody embedding and the target speech's speaker embedding); and generating final translated speech audio using the source speaker embedding, the translated content embedding, and the translated intonation data, the final translated speech audio having the source vocal identity in the second language(Para [0045], Ln 1-13, For voice conversion 20, VQMIVC (Vector quantization mutual information voice conversion) can be used, which uses a straightforward autoencoder architecture to solve the voice conversion issue. The framework consists of four modules: a content encoder that produces a content embedding from speech, a speaker encoder that produces a speaker embedding (D-vector) from speech, a pitch encoder that produces prosody embedding from speech, and a decoder that generates from content, prosody, and speaker embeddings. Para [0046], Ln 1-15, To extract target speaker embedding, the target speech is sent into the speaker encoder. Finally, the decoder reconstructs the converted speech using the source speech's content embedding and prosody embedding and the target speech's speaker embedding. See also Claim 1 of Waibel, Ln 15-18, wherein the voice characteristics comprise pitch). Regarding Claim 11: Claim 11 contains similar limitations as Claim 4, and is therefore rejected for the same reasons. Regarding Claim 12: Claim 12 contains similar limitations as Claim 5, and is therefore rejected for the same reasons. Regarding Claim 15: Waibel teaches a system comprising: a memory component; and a processing device coupled to the memory component, the processing device to perform operations comprising(Para [0133], Ln 1-7, computer system according to the present invention comprises one or more processor cores and a memory in communication with the one or more processor cores. The memory stores instructions that when executed by the one or more processor cores, cause the one or more processor cores. Para [0079], Ln 1-10, RAM): receiving an input audio sequence, the input audio sequence including speech audio having a source vocal identity in a first language(Para [0009], Ln 1-15, system comprises a remote source for capturing input audio by the output speaker in a first language that is different from the target language and converting speech by the output speaker in the input audio into text in the first language. Para [0010], Ln 1-13, Generating the adapted speech comprises adapting the speech in to target language to voice characteristics of the first speaker in the input audio); translating a first transcription of the speech audio to a second transcription of the speech audio in a second language(Para [0012], Ln 1-20, generating, by a translation module, trained through machine learning, of the computer system, a textual translation into the target language from the text in the first language from the remote source); generating, by a text-to-speech model, initial translated speech audio in the second language using the second transcription of the speech audio, the initial translated speech audio having a default vocal identity(Para [0054], Ln 1-16, Subsequently, the TTS module 18 is given this translated text and the resulting Mel spectrogram is turned into a waveform file by the HiFi-GAN vocoder. The final audio is now created by the voice conversion module 20, which gets the waveform of German speech that the vocoder produced as input and uses the original English audio of the input video as target speaker. Abstract, Ln 1-14, The translated text is then synthesized by a Text-to-Speech model that recreates the original emphases in the target sentence. The resulting synthetic speech is then converted back to the original speakers' voice using a voice conversion model); processing the initial translated speech audio to generate a translated content embedding and translated intonation data(Para [0045], Ln 1-13, For voice conversion 20, VQMIVC (Vector quantization mutual information voice conversion) can be used, which uses a straightforward autoencoder architecture to solve the voice conversion issue. The framework consists of four modules: a content encoder that produces a content embedding from speech, a speaker encoder that produces a speaker embedding (D-vector) from speech, a pitch encoder that produces prosody embedding from speech, and a decoder that generates from content, prosody, and speaker embeddings. Para [0046], Ln 1-15, To extract target speaker embedding, the target speech is sent into the speaker encoder. Finally, the decoder reconstructs the converted speech using the source speech's content embedding and prosody embedding and the target speech's speaker embedding); processing the input audio sequence to generate a source speaker embedding representing the source vocal identity of the speech audio(Para [0045], Ln 1-13, For voice conversion 20, VQMIVC (Vector quantization mutual information voice conversion) can be used, which uses a straightforward autoencoder architecture to solve the voice conversion issue. The framework consists of four modules: a content encoder that produces a content embedding from speech, a speaker encoder that produces a speaker embedding (D-vector) from speech, a pitch encoder that produces prosody embedding from speech, and a decoder that generates from content, prosody, and speaker embeddings. Para [0046], Ln 1-15, To extract target speaker embedding, the target speech is sent into the speaker encoder. Finally, the decoder reconstructs the converted speech using the source speech's content embedding and prosody embedding and the target speech's speaker embedding); and generating final translated speech audio using the source speaker embedding, the translated content embedding, and the translated intonation data, the final translated speech audio having the source vocal identity in the second language(Para [0045], Ln 1-13, For voice conversion 20, VQMIVC (Vector quantization mutual information voice conversion) can be used, which uses a straightforward autoencoder architecture to solve the voice conversion issue. The framework consists of four modules: a content encoder that produces a content embedding from speech, a speaker encoder that produces a speaker embedding (D-vector) from speech, a pitch encoder that produces prosody embedding from speech, and a decoder that generates from content, prosody, and speaker embeddings. Para [0046], Ln 1-15, To extract target speaker embedding, the target speech is sent into the speaker encoder. Finally, the decoder reconstructs the converted speech using the source speech's content embedding and prosody embedding and the target speech's speaker embedding. See also Claim 1 of Waibel, Ln 15-18, wherein the voice characteristics comprise pitch). Regarding Claim 18: Claim 18 contains similar limitations as Claim 4, and is therefore rejected for the same reasons. Regarding Claim 19: Claim 19 contains similar limitations as Claim 5, and is therefore rejected for the same reasons. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 2-3, 9-10 and 16-17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Waibel as applied to claim 1 above, and further in view of Di Gangi et al. (US 20250118336 A1). Regarding Claim 2: Waibel teaches the method of claim 1, but does not specifically teaches wherein translating the first transcription of the speech audio to the second transcription of the speech audio in the second language further comprises: generating a first transcription of the speech audio by: segmenting the speech audio into a plurality of source segments; generating a translation for each source segment of the plurality of source segments; and storing timestamps for each source segment of the plurality of source segments with a corresponding generated translation. In the same field of Speech-to-Speech translation, Di Gangi teaches wherein translating the first transcription of the speech audio to the second transcription of the speech audio in the second language further comprises: generating a first transcription of the speech audio by: segmenting the speech audio into a plurality of source segments(Para [0063], Ln 1-7, The information from ASR, punctuation and diarization is then merged into a data structure that includes the transcripts with the initial and final time stamps of each word, divided into segments, each segment with their timing and speaker. Para [0075], Ln 1-11, The speech placement system then determines the start and stop time of each segment, as well as a time correction factor for each word in order to expand or shrink its duration according to the time needs.); generating a translation for each source segment of the plurality of source segments(Para [0079], Ln 1-14, MT model is trained to output sequences of phonemes and their respective duration….output of MT is a sequence of pairs (phoneme, duration). The output sequence is then postprocessed to extract the phoneme and duration sequences separately and these sequences are fed as input to the TTS systems after speech placement and N-best selection from the previous points); and storing timestamps for each source segment of the plurality of source segments with a corresponding generated translation(Para [0079], Ln 1-14, MT model is trained to output sequences of phonemes and their respective duration….output of MT is a sequence of pairs (phoneme, duration). The output sequence is then postprocessed to extract the phoneme and duration sequences separately and these sequences are fed as input to the TTS systems after speech placement and N-best selection from the previous points). It would have been obvious for one skilled in the art, at the effective time of filling, to modify Waibel, with the speech timing information of Di Gangi, as it improves the accuracy of the generated speech(Para [0014], Ln 1-7). Regarding Claim 3: The combination of Waibel and Di Gangi teaches the method of claim 2, but does not teach wherein generating the initial translated speech audio in the second language using the second transcription of the speech audio further comprises: generating translated segments for each of the plurality of translated segments; and modifying a speaking rate for each of the translated segments based on the timestamps for a corresponding source segment. In the same field of Speech-to-Speech translation, Di Gangi wherein generating the initial translated speech audio in the second language using the second transcription of the speech audio further comprises: generating translated segments for each of the plurality of translated segments(Para [0079], Ln 1-14, MT model is trained to output sequences of phonemes and their respective duration….output of MT is a sequence of pairs (phoneme, duration). The output sequence is then postprocessed to extract the phoneme and duration sequences separately and these sequences are fed as input to the TTS systems after speech placement and N-best selection from the previous points); and modifying a speaking rate for each of the translated segments based on the timestamps for a corresponding source segment(Para [0079], Ln 1-14, MT model is trained to output sequences of phonemes and their respective duration….output of MT is a sequence of pairs (phoneme, duration). The output sequence is then postprocessed to extract the phoneme and duration sequences separately and these sequences are fed as input to the TTS systems after speech placement and N-best selection from the previous points. Para [0080], Ln 1-11, FIG. 15 depicts a combination of TTS with speech placement according to an embodiment of the invention. There is a component 1520 for predicting the utterance duration, usually the TTS encoder itself. These values are taken as input, together with the original durations, to compute in 1540 the tempo factors for each sentence). It would have been obvious for one skilled in the art, at the effective e time of filling, to modify the combination of Waibel and Di Gangi, with the speech timing information of Di Gangi, as it improves the accuracy of the generated speech(Para [0014], Ln 1-7). Regarding Claim 9: Claim 9 contains similar limitations as Claim 2, and is therefore rejected for the same reasons. Regarding Claim 10: Claim 10 contains similar limitations as Claim 3, and is therefore rejected for the same reasons. Regarding Claim 16: Claim 16 contains similar limitations as Claim 2, and is therefore rejected for the same reasons. Regarding Claim 17: Claim 17 contains similar limitations as Claim 3, and is therefore rejected for the same reasons. Claim(s) 6-7, 13-14 and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Waibel as applied to claim 1 above, and further in view of Federico et al. (US 11545134 B1). Regarding Claim 6: Waibel teaches the method of claim 1, but does not teach further comprising: obtaining the speech audio from the input audio sequence by extracting background noise and impulse response data from the input audio sequence. In the same field of Speech to Speech translation, Federico teaches further comprising: obtaining the speech audio from the input audio sequence by extracting background noise and impulse response data from the input audio sequence(Col 3, Ln 35-42, The speech segments may then be processed to separate speech from possible background noise. Col 3, Ln 62-67, A rendering step may be applied which adds to the clean signal the reverberation and background noise present in the original utterance. Col 7, Ln 37-42, A ML-based environment modeler 317 extracts acoustic-response information (reverberation information) of the environment from the speech and background exported from the audio segmentator 301, so that it can be re-inserted in the target speech. Col 8, Ln 16-28, An environmental modeler 317 estimate the environment reverberation from the original audio. Unfortunately, estimating the room impulse response (RIR) from a reverberated signal requires solving an ill-posed blind deconvolution problem. In some embodiments, a blind estimation of the reverberation time (RT) is made which is commonly used to assess the amount of room reverberation or its effects…...In some embodiments, the estimated RT used to generate a synthetic RIR using a RIR generator). It would have been obvious for one skilled in the art, at the effective time of filling, to modify Waibel, with the noise and reverberation system of Federico, as it improves the quality of the resulting speech(Col 2, Ln 6-17). Regarding Claim 7: The combination of Waibel and Federico teaches the method of claim 6, but does not teach wherein generating the final translated speech audio using the source speaker embedding, the translated content embedding, and the translated intonation data further comprises: generating an output audio sequence by convolving the background noise and impulse response data from the input audio sequence with the final translated speech audio. In the same field of Speech to Speech translation, Federico teaches wherein generating the final translated speech audio using the source speaker embedding, the translated content embedding, and the translated intonation data further comprises: generating an output audio sequence by convolving the background noise and impulse response data from the input audio sequence with the final translated speech audio(Col 3, Ln 35-42, The speech segments may then be processed to separate speech from possible background noise. Col 3, Ln 62-67, A rendering step may be applied which adds to the clean signal the reverberation and background noise present in the original utterance. Col 7, Ln 37-42, A ML-based environment modeler 317 extracts acoustic-response information (reverberation information) of the environment from the speech and background exported from the audio segmentator 301, so that it can be re-inserted in the target speech. Col 8, Ln 16-28, An environmental modeler 317 estimate the environment reverberation from the original audio. Unfortunately, estimating the room impulse response (RIR) from a reverberated signal requires solving an ill-posed blind deconvolution problem. In some embodiments, a blind estimation of the reverberation time (RT) is made which is commonly used to assess the amount of room reverberation or its effects…...In some embodiments, the estimated RT used to generate a synthetic RIR using a RIR generator. Col 8, ln 29-33, speech signal is subjected to a speech renderer 319 to re-introduce background noise and environmental reverberation in the generated speech signal. In particular, one or more of the synthetic RIR and/or background noise are applied to the clean speech audio). It would have been obvious for one skilled in the art, at the effective time of filling, to modify the combination of Waibel and Federico, with the noise and reverberation system of Federico, as it improves the quality of the resulting speech(Col 2, Ln 6-17). Regarding Claim 13: Claim 13 contains similar limitations as Claim 6, and is therefore rejected for the same reasons. Regarding Claim 14: Claim 14 contains similar limitations as Claim 7, and is therefore rejected for the same reasons. Regarding Claim 20: Waibel teaches the system of claim 15, but does not teach wherein the operations of generating the final translated speech audio using the source speaker embedding, the translated content embedding, and the translated intonation further comprise: extracting background noise and impulse response data from the input audio sequence; and generating an output audio sequence by convolving the background noise and impulse response data from the input audio sequence with the final translated speech audio. In the same field of Speech to Speech translation, Federico teaches wherein the operations of generating the final translated speech audio using the source speaker embedding, the translated content embedding, and the translated intonation further comprise: extracting background noise and impulse response data from the input audio sequence(Col 3, Ln 35-42, The speech segments may then be processed to separate speech from possible background noise. Col 3, Ln 62-67, A rendering step may be applied which adds to the clean signal the reverberation and background noise present in the original utterance. Col 8, Ln 16-28, An environmental modeler 317 estimate the environment reverberation from the original audio. Unfortunately, estimating the room impulse response (RIR) from a reverberated signal requires solving an ill-posed blind deconvolution problem. In some embodiments, a blind estimation of the reverberation time (RT) is made which is commonly used to assess the amount of room reverberation or its effects…...In some embodiments, the estimated RT used to generate a synthetic RIR using a RIR generator); and generating an output audio sequence by convolving the background noise and impulse response data from the input audio sequence with the final translated speech audio(Col 3, Ln 35-42, The speech segments may then be processed to separate speech from possible background noise. Col 3, Ln 62-67, A rendering step may be applied which adds to the clean signal the reverberation and background noise present in the original utterance. Col 7, Ln 37-42, A ML-based environment modeler 317 extracts acoustic-response information (reverberation information) of the environment from the speech and background exported from the audio segmentator 301, so that it can be re-inserted in the target speech. Col 8, Ln 16-28, An environmental modeler 317 estimate the environment reverberation from the original audio. Unfortunately, estimating the room impulse response (RIR) from a reverberated signal requires solving an ill-posed blind deconvolution problem. In some embodiments, a blind estimation of the reverberation time (RT) is made which is commonly used to assess the amount of room reverberation or its effects…...In some embodiments, the estimated RT used to generate a synthetic RIR using a RIR generator. Col 8, ln 29-33, speech signal is subjected to a speech renderer 319 to re-introduce background noise and environmental reverberation in the generated speech signal. In particular, one or more of the synthetic RIR and/or background noise are applied to the clean speech audio). It would have been obvious for one skilled in the art, at the effective time of filling, to modify Waibel, with the noise and reverberation system of Federico, as it improves the quality of the resulting speech(Col 2, Ln 6-17). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Hu et al. (US 20230343319 A1). Speech to speech translation with vocal characteristics of source speaker. Li et al. (US 20250349282 A1) Text to speech synthesis including speaker characteristics and embeddings Iyer et al. (US 20240331681 A1) Speech to Speech translation with adaption of speech output based on source speaker voice characteristics. Federico et al. “From Speech-to-Speech Translation to Automatic Dubbing”. Speech to Speech translation with recreation of impulse response and background noise in output speech. Wang et al. “CONTROLLABLE SPEECH REPRESENTATION LEARNING VIA VOICE CONVERSIONAND AICLOSS”. Speech conversion based on voice characteristics. Shares inventors with instant application. Is prior art. Disong Wang et al. “VQMIVC:Vector Quantization and Mutual Information-Based Unsupervised Speech Representation Disentanglement for One-shot Voice Conversion” Speech conversion model cited in Waibel. Uses content encoder, speaker encoder and pitch extractor to recreate speech. Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALEXANDER G MARLOW whose telephone number is (571)272-4536. The examiner can normally be reached Monday - Thursday 10:00 am - 8:00 pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richmond Dorvil can be reached at (571)272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ALEXANDER G MARLOW/ Assistant Examiner, Art Unit 2658 /RICHEMOND DORVIL/ Supervisory Patent Examiner, Art Unit 2658
Read full office action

Prosecution Timeline

Nov 25, 2024
Application Filed
Jul 24, 2026
Non-Final Rejection mailed — §102, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12694204
SYSTEM AND METHOD FOR COMPARING DOCUMENTS USING AN ARTIFICIAL INTELLIGENCE (AI) MODEL
2y 4m to grant Granted Jul 28, 2026
Patent 12682903
Voice Query QoS based on Client-Computed Content Metadata
2y 9m to grant Granted Jul 14, 2026
Patent 12670896
INFORMATION PROCESSING METHOD, NON-TRANSITORY RECORDING MEDIUM, INFORMATION PROCESSING APPARATUS, AND INFORMATION PROCESSING SYSTEM
3y 6m to grant Granted Jun 30, 2026
Patent 12664992
SPATIAL AUDIO PARAMETER ENCODING AND ASSOCIATED DECODING
3y 3m to grant Granted Jun 23, 2026
Patent 12646515
SELECTIVELY PROVIDING ENHANCED CLARIFICATION PROMPTS IN AUTOMATED ASSISTANT INTERACTIONS
2y 8m to grant Granted Jun 02, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
79%
Grant Probability
97%
With Interview (+18.1%)
2y 8m (~1y 0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 84 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month