Prosecution Insights
Last updated: August 17, 2026
Application No. 18/110,141

AUDIO SIGNAL GENERATION USING NEURAL NETWORKS

Non-Final OA §103
Filed
Feb 15, 2023
Examiner
BECKER, TYLER JUSTIN
Art Unit
2657
Tech Center
2600 — Communications
Assignee
NVIDIA Corporation
OA Round
3 (Non-Final)
74%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
80%
With Interview

Examiner Intelligence

Grants 74% — above average
74%
Career Allowance Rate
17 granted / 23 resolved
+11.9% vs TC avg
Moderate +6% lift
Without
With
+6.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 8m
Avg Prosecution
15 currently pending
Career history
45
Total Applications
across all art units

Statute-Specific Performance

§101
19.2%
-20.8% vs TC avg
§103
51.1%
+11.1% vs TC avg
§102
14.3%
-25.7% vs TC avg
§112
15.4%
-24.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 23 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on April 2nd, 2026 has been entered. Response to Amendment The amendment filed April 2nd, 2026 has been entered. Claims 1, 5, 8, 12, 15, and 18 have been amended. Claims 1-20 are pending and have been examined. Response to Arguments Applicant’s arguments, see page 6 of the applicant's remarks, filed April 2nd, 2026, with respect to the objection to the specification have been fully considered and are persuasive. The objection of October 2nd, 2025 has been withdrawn. Applicant’s amendments and arguments, see pages 10-12 of the applicant's remarks, filed April 2nd, 2026, with respect to the rejection of the claims under 35 U.S.C. 101 have been fully considered and are persuasive. In particular, the added limitation detailing a generator portion that is updated based on the output of one or more discriminator networks is seen as an additional element that amounts to significantly more than the judicial exception, and is not a generic computing component. As such, the rejection of October 2nd, 2025 has been withdrawn. Applicant’s arguments with respect to the rejection of claim(s) 1-20 under 35 U.S.C. 103 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Claim Objections Claims 5, 12, and 18 objected to because of the following informalities: Each of these claims recites the limitation “wherein one or more first encodings of the one or more first audio features and one or more second encodings of the one or more second features.” This is an incomplete phrase, and should be completed to read “wherein one or more first encodings are of the one or more first audio features and one or more second encodings are of the one or more second features”, “wherein one or more first encodings of the one or more first audio features and one or more second encodings of the one or more second features are generated”, or in another way deemed appropriate by the applicant. Appropriate correction is required. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-3, 5, 8-10, 12, and 15-18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Trueba et al. (US Pat. No. 11,735,156 B1 hereinafter Trueba), in view of Gupta et al. (US Pat. No. 11,605,388 B1 hereinafter Gupta) and Lubin et al. (US Pat. Pub. No. 2024/0355346 A1 hereinafter Lubin). Regarding claim 1, Trueba discloses a processor, comprising: one or more circuits to cause one or more neural networks (Trueba, Col. 2, lines 7-11: “The processing component(s), referred to herein as a voice-transfer component, may include one or more neural-network models configured as one or more encoders and one or more neural-network models configured as one or more decoders.”) to generate, from an input speech and a reference speech, an audio signal, wherein the one or more neural networks comprise: one or more portions to obtain: one or more first audio features corresponding to a first voice signal of the input speech; and one or more second features different from the one or more first audio features corresponding to a second voice signal of the reference speech (Trueba, Fig. 1B; Col. 4, lines 27-65: "The user device 110 and/or remote system 120 processes (134) the first audio data to determine first encoded data corresponding to phoneme characteristics of the first speech."; "The user device 110 and/or remote system 120 may also process (136) the first audio data to determine second encoded data corresponding to a phrase corresponding to the first speech."; "The user device 110 and/or remote system 120 processes (138) the second audio data (e.g., the target input data 152) to determine third encoded data corresponding to vocal characteristics of the second speech (e.g., the target speech)."; "The user device 110 and/or remote system 120 may then process (140) the first encoded data, the second encoded data, and the third encoded data to determine third audio data (e.g., the output data 162) that corresponds to the phrase encoded data, the phoneme characteristics encoded data, and the vocal characteristics encoded data."). However, Trueba fails to expressly recite an audio signal that maintains prosody of the input speech, wherein the one or more neural networks comprise: one or more portions to obtain: one or more first audio features corresponding to a first voice signal of the input speech; and one or more second features different from the one or more first audio features corresponding to a second voice signal of the reference speech; and a generator portion to generate the audio signal from the one or more first audio features and the one or more second audio features based, at least in part, on one or more updates to the generator determined according to respective outputs of one or more discriminator networks to one or more prior audio signals generated by the one or more neural networks. Gupta teaches an audio signal that maintains prosody of the input speech, wherein the one or more neural networks comprise: one or more portions to obtain: one or more first audio features corresponding to a first voice signal of the input speech; and one or more second features different from the one or more first audio features corresponding to a second voice signal of the reference speech (Gupta, Col. 3, lines 37-41: “The methods and systems described in this specification enable speech audio to be generated in a target speaker's voice, while maintaining the performance (e.g. speech prosody) and timing of source speech audio from which the acoustic features relating to a source speaker are derived.”). Trueba and Gupta are analogous arts because they both belong to the field of audio processing. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the synthetic speech processing system of Trueba to incorporate the teachings of Gupta to maintain the prosody of the input speech during speech generation. This allows the sound of a person’s voice to be modified without changing the original speaker’s performance and timing (Gupta, Col. 3). This helps retain quality in the original speech even when it is modified to sound different, resulting in higher quality output audio. However, Trueba, in view of Gupta, fails to expressly recite a generator portion to generate the audio signal from the one or more first audio features and the one or more second audio features based, at least in part, on one or more updates to the generator determined according to respective outputs of one or more discriminator networks to one or more prior audio signals generated by the one or more neural networks. Lubin teaches a generator portion to generate the audio signal from the one or more first audio features and the one or more second audio features based, at least in part, on one or more updates to the generator determined according to respective outputs of one or more discriminator networks to one or more prior audio signals generated by the one or more neural networks (Lubin, [0006]: “the disclosure describes a method comprising: processing, with an encoder of a machine learning system, an input audio waveform comprising first utterances by a speaker to generate an encoder output, the first utterances having a first accent, processing, with a decoder of the machine learning system, the encoder output to generate an output audio waveform comprising second utterances, the second utterances having a second accent different from the first accent, computing, with a signal loss discriminator of the machine learning system and based on the input audio waveform and the output audio waveform, a signal loss for the output audio waveform, computing, with an identification loss discriminator of the machine learning system, based on the input audio waveform and the output audio waveform, an identification loss for the output audio waveform, computing, with a text loss discriminator of the machine learning system, based on the input audio waveform and the output audio waveform, a text loss for the output audio waveform; and training the decoder using the signal loss, the identification loss, and the text loss.”). Trueba, Gupta, and Lubin are analogous arts because they each belong to the field of audio processing. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the synthetic speech processing system of Trueba, as modified by the speaker conversion system of Gupta, to incorporate the teachings of Lubin to update a generator portion based on one or more discriminator networks. Using a generator portion updated in this way ensures the enhanced speech output sounds like a specific speaker and is intelligible (Lubin, [0016]). As such, the system can produce a high quality output and provide a better experience for the user. Regarding claim 2, the rejection of claim 1 is incorporated. Trueba, in view of Gupta and Lubin, discloses all of the elements of the current invention as stated above. Trueba further discloses wherein the one or more second features comprises a timbre of the second voice signal (Trueba, Col. 3, lines 63-65: "the vocal characteristics may represent features of the voice of a particular speaker, such as tone, resonance, timbre, pitch, and/or frequency."), and the one or more neural networks are to generate the audio signal such that the audio signal comprises the one or more first audio features and the timbre corresponding to the second voice signal (Trueba, Col. 4, line 65- col. 5, line 3: "The output data 162 thus may include a representation of the phrase and/or phoneme characteristics corresponding to the source input data 150, while the representation further corresponds to the vocal characteristics represented in the target input data 152."; Col. 3, lines 50-51: "The user device 110 and/or other device may output audio 14 corresponding to the output data 162."). Regarding claim 3, the rejection of claim 1 is incorporated. Trueba, in view of Gupta and Lubin, discloses all of the elements of the current invention as stated above. Trueba further discloses wherein the one or more first audio features comprise at least one of: pitch, amplitude, and linguistic content (Trueba, Col. 3, lines 58-65: "The first audio data may further correspond to phoneme characteristics and vocal characteristics; the phoneme characteristics may represent pronunciation of the first speech that is independent of a voice of a particular speaker, such as syllable breaks, cadence, and/or emphasis, while the vocal characteristics may represent features of the voice of a particular speaker, such as tone, resonance, timbre, pitch, and/or frequency."). Regarding claim 5, the rejection of claim 1 is incorporated. Trueba, in view of Gupta and Lubin, discloses all of the elements of the current invention as stated above. Trueba further discloses wherein one or more first encodings of the one or more first audio features and one or more second encodings of the one or more second features (Trueba, Col. 2, lines 7-19: “The processing component(s), referred to herein as a voice-transfer component, may include one or more neural-network models configured as one or more encoders and one or more neural-network models configured as one or more decoders. A first encoder may process first input data corresponding to a source voice to determine first encoded data representing phoneme characteristics of the source voice, and a second encoder may process the first input data to determine second encoded data representing a phrase represented in the first input data. A third encoder may process second input data corresponding to a target voice to determine vocal characteristic data representing vocal characteristics of the target voice.”). Regarding claim 8, Trueba discloses a method, comprising: causing one or more neural networks (Trueba, Col. 2, lines 7-11: “The processing component(s), referred to herein as a voice-transfer component, may include one or more neural-network models configured as one or more encoders and one or more neural-network models configured as one or more decoders.”) to generate, from an input speech and a reference speech, an audio signal, wherein the one or more neural networks comprise: one or more portions to obtain: one or more first audio features corresponding to a first voice signal of the input speech; and one or more second features different from the one or more first audio features corresponding to a second voice signal of the reference speech (Trueba, Fig. 1B; Col. 4, lines 27-65: "The user device 110 and/or remote system 120 processes (134) the first audio data to determine first encoded data corresponding to phoneme characteristics of the first speech."; "The user device 110 and/or remote system 120 may also process (136) the first audio data to determine second encoded data corresponding to a phrase corresponding to the first speech."; "The user device 110 and/or remote system 120 processes (138) the second audio data (e.g., the target input data 152) to determine third encoded data corresponding to vocal characteristics of the second speech (e.g., the target speech)."; "The user device 110 and/or remote system 120 may then process (140) the first encoded data, the second encoded data, and the third encoded data to determine third audio data (e.g., the output data 162) that corresponds to the phrase encoded data, the phoneme characteristics encoded data, and the vocal characteristics encoded data."). However, Trueba fails to expressly recite an audio signal that maintains prosody of the input speech, wherein the one or more neural networks comprise: one or more portions to obtain: one or more first audio features corresponding to a first voice signal of the input speech and one or more second features different from the one or more first audio features corresponding to a second voice signal of the reference speech; and a generator portion to generate the audio signal from the one or more first audio features and the one or more second audio features based, at least in part, on one or more updates to the generator determined according to respective outputs of one or more discriminator networks to one or more prior audio signals generated by the one or more neural networks. Gupta teaches an audio signal that maintains prosody of the input speech, wherein the audio signal is generated based, at least in part, on one or more first audio features corresponding to a first voice signal of the input speech and one or more second features different from the one or more first audio features corresponding to a second voice signal of the reference speech (Gupta, Col. 3, lines 37-41: “The methods and systems described in this specification enable speech audio to be generated in a target speaker's voice, while maintaining the performance (e.g. speech prosody) and timing of source speech audio from which the acoustic features relating to a source speaker are derived.”). Trueba and Gupta are analogous arts because they both belong to the field of audio processing. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the synthetic speech processing system of Trueba to incorporate the teachings of Gupta to maintain the prosody of the input speech during speech generation. This allows the sound of a person’s voice to be modified without changing the original speaker’s performance and timing (Gupta, Col. 3). This helps retain quality in the original speech even when it is modified to sound different, resulting in higher quality output audio. However, Trueba, in view of Gupta, fails to expressly recite a generator portion to generate the audio signal from the one or more first audio features and the one or more second audio features based, at least in part, on one or more updates to the generator determined according to respective outputs of one or more discriminator networks to one or more prior audio signals generated by the one or more neural networks. Lubin teaches a generator portion to generate the audio signal from the one or more first audio features and the one or more second audio features based, at least in part, on one or more updates to the generator determined according to respective outputs of one or more discriminator networks to one or more prior audio signals generated by the one or more neural networks (Lubin, [0006]: “the disclosure describes a method comprising: processing, with an encoder of a machine learning system, an input audio waveform comprising first utterances by a speaker to generate an encoder output, the first utterances having a first accent, processing, with a decoder of the machine learning system, the encoder output to generate an output audio waveform comprising second utterances, the second utterances having a second accent different from the first accent, computing, with a signal loss discriminator of the machine learning system and based on the input audio waveform and the output audio waveform, a signal loss for the output audio waveform, computing, with an identification loss discriminator of the machine learning system, based on the input audio waveform and the output audio waveform, an identification loss for the output audio waveform, computing, with a text loss discriminator of the machine learning system, based on the input audio waveform and the output audio waveform, a text loss for the output audio waveform; and training the decoder using the signal loss, the identification loss, and the text loss.”). Trueba, Gupta, and Lubin are analogous arts because they each belong to the field of audio processing. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the synthetic speech processing system of Trueba, as modified by the speaker conversion system of Gupta, to incorporate the teachings of Lubin to update a generator portion based on one or more discriminator networks. Using a generator portion updated in this way ensures the enhanced speech output sounds like a specific speaker and is intelligible (Lubin, [0016]). As such, the system can produce a high quality output and provide a better experience for the user. Regarding claim 9, the rejection of claim 1 is incorporated. Trueba, in view of Gupta and Lubin, discloses all of the elements of the current invention as stated above. Trueba further discloses wherein the one or more second features comprises a timbre of the second voice signal (Trueba, Col. 3, lines 63-65: "the vocal characteristics may represent features of the voice of a particular speaker, such as tone, resonance, timbre, pitch, and/or frequency."), and the one or more neural networks are to generate the audio signal such that the audio signal comprises the one or more first audio features and the timbre corresponding to the second voice signal (Trueba, Col. 4, line 65- col. 5, line 3: "The output data 162 thus may include a representation of the phrase and/or phoneme characteristics corresponding to the source input data 150, while the representation further corresponds to the vocal characteristics represented in the target input data 152."; Col. 3, lines 50-51: "The user device 110 and/or other device may output audio 14 corresponding to the output data 162."). Regarding claim 10, the rejection of claim 1 is incorporated. Trueba, in view of Gupta and Lubin, discloses all of the elements of the current invention as stated above. Trueba further discloses wherein the one or more first audio features comprise at least one of: pitch, amplitude, and linguistic content (Trueba, Col. 3, lines 58-65: "The first audio data may further correspond to phoneme characteristics and vocal characteristics; the phoneme characteristics may represent pronunciation of the first speech that is independent of a voice of a particular speaker, such as syllable breaks, cadence, and/or emphasis, while the vocal characteristics may represent features of the voice of a particular speaker, such as tone, resonance, timbre, pitch, and/or frequency."). Regarding claim 12, the rejection of claim 1 is incorporated. Trueba, in view of Gupta and Lubin, discloses all of the elements of the current invention as stated above. Trueba further discloses wherein one or more first encodings of the one or more first audio features, and one or more second encodings of the one or more second features (Trueba, Col. 2, lines 7-19: “The processing component(s), referred to herein as a voice-transfer component, may include one or more neural-network models configured as one or more encoders and one or more neural-network models configured as one or more decoders. A first encoder may process first input data corresponding to a source voice to determine first encoded data representing phoneme characteristics of the source voice, and a second encoder may process the first input data to determine second encoded data representing a phrase represented in the first input data. A third encoder may process second input data corresponding to a target voice to determine vocal characteristic data representing vocal characteristics of the target voice.”). Regarding claim 15, Trueba discloses a system, comprising: one or more processors comprising one or more circuits to use one or more neural networks (Trueba, Col. 2, lines 7-11: “The processing component(s), referred to herein as a voice-transfer component, may include one or more neural-network models configured as one or more encoders and one or more neural-network models configured as one or more decoders.”) to generate, from an input speech and a reference speech, an audio signal, wherein the one or more neural networks comprise: one or more portions to obtain: one or more first audio features corresponding to a first voice signal of the input speech; and one or more second features different from the one or more first audio features corresponding to a second voice signal of the reference speech (Trueba, Fig. 1B; Col. 4, lines 27-65: "The user device 110 and/or remote system 120 processes (134) the first audio data to determine first encoded data corresponding to phoneme characteristics of the first speech."; "The user device 110 and/or remote system 120 may also process (136) the first audio data to determine second encoded data corresponding to a phrase corresponding to the first speech."; "The user device 110 and/or remote system 120 processes (138) the second audio data (e.g., the target input data 152) to determine third encoded data corresponding to vocal characteristics of the second speech (e.g., the target speech)."; "The user device 110 and/or remote system 120 may then process (140) the first encoded data, the second encoded data, and the third encoded data to determine third audio data (e.g., the output data 162) that corresponds to the phrase encoded data, the phoneme characteristics encoded data, and the vocal characteristics encoded data."). However, Trueba fails to expressly recite an audio signal that maintains prosody of the input speech, wherein the one or more neural networks comprise: one or more portions to obtain: one or more first audio features corresponding to a first voice signal of the input speech and one or more second features different from the one or more first audio features corresponding to a second voice signal of the reference speech; and a generator portion to generate the audio signal from the one or more first audio features and the one or more second audio features based, at least in part, on one or more updates to the generator determined according to respective outputs of one or more discriminator networks to one or more prior audio signals generated by the one or more neural networks. Gupta teaches an audio signal that maintains prosody of the input speech, wherein the audio signal is generated based, at least in part, on one or more first audio features corresponding to a first voice signal of the input speech and one or more second features different from the one or more first audio features corresponding to a second voice signal of the reference speech (Gupta, Col. 3, lines 37-41: “The methods and systems described in this specification enable speech audio to be generated in a target speaker's voice, while maintaining the performance (e.g. speech prosody) and timing of source speech audio from which the acoustic features relating to a source speaker are derived.”). Trueba and Gupta are analogous arts because they both belong to the field of audio processing. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the synthetic speech processing system of Trueba to incorporate the teachings of Gupta to maintain the prosody of the input speech during speech generation. This allows the sound of a person’s voice to be modified without changing the original speaker’s performance and timing (Gupta, Col. 3). This helps retain quality in the original speech even when it is modified to sound different, resulting in higher quality output audio. However, Trueba, in view of Gupta, fails to expressly recite a generator portion to generate the audio signal from the one or more first audio features and the one or more second audio features based, at least in part, on one or more updates to the generator determined according to respective outputs of one or more discriminator networks to one or more prior audio signals generated by the one or more neural networks. Lubin teaches a generator portion to generate the audio signal from the one or more first audio features and the one or more second audio features based, at least in part, on one or more updates to the generator determined according to respective outputs of one or more discriminator networks to one or more prior audio signals generated by the one or more neural networks (Lubin, [0006]: “the disclosure describes a method comprising: processing, with an encoder of a machine learning system, an input audio waveform comprising first utterances by a speaker to generate an encoder output, the first utterances having a first accent, processing, with a decoder of the machine learning system, the encoder output to generate an output audio waveform comprising second utterances, the second utterances having a second accent different from the first accent, computing, with a signal loss discriminator of the machine learning system and based on the input audio waveform and the output audio waveform, a signal loss for the output audio waveform, computing, with an identification loss discriminator of the machine learning system, based on the input audio waveform and the output audio waveform, an identification loss for the output audio waveform, computing, with a text loss discriminator of the machine learning system, based on the input audio waveform and the output audio waveform, a text loss for the output audio waveform; and training the decoder using the signal loss, the identification loss, and the text loss.”). Trueba, Gupta, and Lubin are analogous arts because they each belong to the field of audio processing. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the synthetic speech processing system of Trueba, as modified by the speaker conversion system of Gupta, to incorporate the teachings of Lubin to update a generator portion based on one or more discriminator networks. Using a generator portion updated in this way ensures the enhanced speech output sounds like a specific speaker and is intelligible (Lubin, [0016]). As such, the system can produce a high quality output and provide a better experience for the user. Regarding claim 16, the rejection of claim 1 is incorporated. Trueba, in view of Gupta and Lubin, discloses all of the elements of the current invention as stated above. Trueba further discloses wherein the one or more second features comprise a timbre of the second voice signal (Trueba, Col. 3, lines 63-65: "the vocal characteristics may represent features of the voice of a particular speaker, such as tone, resonance, timbre, pitch, and/or frequency."), and the one or more neural networks are to generate the audio signal such that the audio signal comprises the one or more first audio features and the timbre corresponding to the second voice signal (Trueba, Col. 4, line 65- col. 5, line 3: "The output data 162 thus may include a representation of the phrase and/or phoneme characteristics corresponding to the source input data 150, while the representation further corresponds to the vocal characteristics represented in the target input data 152."; Col. 3, lines 50-51: "The user device 110 and/or other device may output audio 14 corresponding to the output data 162."). Regarding claim 17, the rejection of claim 1 is incorporated. Trueba, in view of Gupta and Lubin, discloses all of the elements of the current invention as stated above. Trueba further discloses wherein the one or more first audio features comprise at least one of: pitch, amplitude, and linguistic content (Trueba, Col. 3, lines 58-65: "The first audio data may further correspond to phoneme characteristics and vocal characteristics; the phoneme characteristics may represent pronunciation of the first speech that is independent of a voice of a particular speaker, such as syllable breaks, cadence, and/or emphasis, while the vocal characteristics may represent features of the voice of a particular speaker, such as tone, resonance, timbre, pitch, and/or frequency."). Regarding claim 18, the rejection of claim 1 is incorporated. Trueba, in view of Gupta and Lubin, discloses all of the elements of the current invention as stated above. Trueba further discloses wherein one or more first encodings of the one or more first audio features and one or more second encodings of the one or more second features (Trueba, Col. 2, lines 7-19: “The processing component(s), referred to herein as a voice-transfer component, may include one or more neural-network models configured as one or more encoders and one or more neural-network models configured as one or more decoders. A first encoder may process first input data corresponding to a source voice to determine first encoded data representing phoneme characteristics of the source voice, and a second encoder may process the first input data to determine second encoded data representing a phrase represented in the first input data. A third encoder may process second input data corresponding to a target voice to determine vocal characteristic data representing vocal characteristics of the target voice.”). Claim(s) 4 and 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Trueba, in view of Gupta and Lubin, as applied to claims 1-3, 5, 8-10, 12, and 15-18 above, and further in view of Carmiel et al. (US Pat. Pub. No. 2023/0352001 A1 hereinafter Carmiel). Regarding claim 4, the rejection of claim 3 is incorporated. Trueba, in view of Gupta and Lubin, discloses all of the elements of the current invention as stated above. However, Trueba, in view of Gupta and Lubin, fails to expressly recite wherein the linguistic content is represented by one or more phoneme posteriorgrams. Carmiel teaches wherein the linguistic content is represented by one or more phoneme posteriorgrams (Carmiel, [0055]: "In another example, another approach is based on converting speech using phonetic posteriorgrams (PPGs). Such prior approach is limited to voice conversion, whereas at least some embodiments described herein enable converting other and/or selected voice attributes such as accent. Such prior approach is are limited to a “many-to-one” approach, i.e., different input voices are mapped to a single output voice. At least some embodiments described herein provide a “many-to-many” approach where the input voice may be converted to different sets of output voice attributes."). Trueba, Gupta, Lubin, and Carmiel are analogous arts because they all belong to the field of audio processing. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the synthetic speech processing system of Trueba, as modified by the speaker conversion system of Gupta and the voice modification system of Lubin, to incorporate the teachings of Carmiel to represent linguistic content using phoneme posteriorgrams. Using phoneme posteriorgrams is a known technique for speech recognition in many-to-one voice conversion systems (Carmiel, [0082]). Using this technique allows for a voice conversion system to effectively recognize the input speech. Regarding claim 11, the rejection of claim 10 is incorporated. Trueba, in view of Gupta and Lubin, discloses all of the elements of the current invention as stated above. However, Trueba, in view of Gupta and Lubin, fails to expressly recite wherein the linguistic content is represented by one or more phoneme posteriorgrams. Carmiel teaches wherein the linguistic content is represented by one or more phoneme posteriorgrams (Carmiel, [0055]: "In another example, another approach is based on converting speech using phonetic posteriorgrams (PPGs). Such prior approach is limited to voice conversion, whereas at least some embodiments described herein enable converting other and/or selected voice attributes such as accent. Such prior approach is are limited to a “many-to-one” approach, i.e., different input voices are mapped to a single output voice. At least some embodiments described herein provide a “many-to-many” approach where the input voice may be converted to different sets of output voice attributes."). Trueba, Gupta, Lubin, and Carmiel are analogous arts because they all belong to the field of audio processing. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the synthetic speech processing system of Trueba, as modified by the speaker conversion system of Gupta and the voice modification system of Lubin, to incorporate the teachings of Carmiel to represent linguistic content using phoneme posteriorgrams. Using phoneme posteriorgrams is a known technique for speech recognition in many-to-one voice conversion systems (Carmiel, [0082]). Using this technique allows for a voice conversion system to effectively recognize the input speech. Claim(s) 6, 13, and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Trueba, in view of Gupta and Lubin, as applied to claims 1-3, 5, 8-10, 12, and 15-18 above, and further in view of Jia et al. (US Pat. Pub. No. 2022/0068256 A1 hereinafter Jia). Regarding claim 6, the rejection of claim 5 is incorporated. Trueba, in view of Gupta and Lubin, discloses all of the elements of the current invention as stated above. However, Trueba, in view of Gupta and Lubin, fails to expressly recite wherein the generator is trained based, at least in part, on one or more audio signals including voices not included in the second voice signal. Jia teaches wherein the generator is trained based, at least in part, on one or more audio signals including voices not included in the second voice signal (Jia, [0003]: "The method includes receiving, at data processing hardware, a first plurality of recorded speech samples from an assortment of speakers and a second plurality of recorded speech samples from a target speaker where the assortment of speakers does not include the target speaker. The method further includes training, at the data processing hardware, a text-to-speech (TTS) model using the first plurality of recorded speech samples from the assortment of speakers."). Trueba, Gupta, Lubin, and Jia are analogous arts because they all belong to the field of audio processing. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the synthetic speech processing system of Trueba, as modified by the speaker conversion system of Gupta and the voice modification system of Lubin, to incorporate the teachings of Jia to train the audio generation model based on audio that is different from a specified audio. The accuracy and/or robustness of a neural network for audio generation depends on the training data set (Jia, [0002]). As such, it is important to have a varied set of data for training a neural network. Regarding claim 13, the rejection of claim 12 is incorporated. Trueba, in view of Gupta and Lubin, discloses all of the elements of the current invention as stated above. However, Trueba, in view of Gupta and Lubin, fails to expressly recite wherein the generator is trained based, at least in part, on one or more audio signals including voices not included in the second voice signal. Jia teaches wherein the generator is trained based, at least in part, on one or more audio signals including voices not included in the second voice signal (Jia, [0003]: "The method includes receiving, at data processing hardware, a first plurality of recorded speech samples from an assortment of speakers and a second plurality of recorded speech samples from a target speaker where the assortment of speakers does not include the target speaker. The method further includes training, at the data processing hardware, a text-to-speech (TTS) model using the first plurality of recorded speech samples from the assortment of speakers."). Trueba, Gupta, Lubin, and Jia are analogous arts because they all belong to the field of audio processing. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the synthetic speech processing system of Trueba, as modified by the speaker conversion system of Gupta and the voice modification system of Lubin, to incorporate the teachings of Jia to train the audio generation model based on audio that is different from a specified audio. The accuracy and/or robustness of a neural network for audio generation depends on the training data set (Jia, [0002]). As such, it is important to have a varied set of data for training a neural network. Regarding claim 19, the rejection of claim 18 is incorporated. Trueba, in view of Gupta and Lubin, discloses all of the elements of the current invention as stated above. However, Trueba, in view of Gupta and Lubin, fails to expressly recite wherein the generator is trained based, at least in part, on one or more audio signals including voices not included in the second voice signal. Jia teaches wherein the generator is trained based, at least in part, on one or more audio signals including voices not included in the second voice signal (Jia, [0003]: "The method includes receiving, at data processing hardware, a first plurality of recorded speech samples from an assortment of speakers and a second plurality of recorded speech samples from a target speaker where the assortment of speakers does not include the target speaker. The method further includes training, at the data processing hardware, a text-to-speech (TTS) model using the first plurality of recorded speech samples from the assortment of speakers."). Trueba, Gupta, Lubin, and Jia are analogous arts because they both belong to the field of audio processing. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the synthetic speech processing system of Trueba, as modified by the speaker conversion system of Gupta and the voice modification system of Lubin, to incorporate the teachings of Jia to train the audio generation model based on audio that is different from a specified audio. The accuracy and/or robustness of a neural network for audio generation depends on the training data set (Jia, [0002]). As such, it is important to have a varied set of data for training a neural network. Claim(s) 7, 14, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Trueba, in view of Gupta and Lubin, as applied to claims 1-3, 5, 8-10, 12, and 15-18 above, and further in view of Prenger et al. (US Pat. Pub. No. 2020/0394994 A1 hereinafter Prenger). Regarding claim 7, the rejection of claim 5 is incorporated. Trueba, in view of Gupta and Lubin, discloses all of the elements of the current invention as stated above. However, Trueba, in view of Gupta and Lubin, fails to expressly recite wherein the generator comprises one or more residual blocks, and each of the one or more residual blocks receive the one or more of the first encodings as input. Prenger teaches wherein the generator comprises one or more residual blocks, and each of the one or more residual blocks receive the one or more of the first encodings as input (Prenger, [0022]: “In at least one embodiment, conditioning a model on other input variables, an audio transform function generation of audio with one or more characteristics encoded as a set of parameters for different properties of audio including but not limited to: pitch; loudness; intonation; rate of speech; rhythm; stress; articulation; and more. In at least one embodiment, in a multi-speaker sets of parameters are encoded for different speaker profiles and a speaker identifier is fed to a model as an extra input to condition speech.”; [0023]: "In at least one embodiment, WN( ) is an audio transformation that uses layers of dilated convolutions with gated-tan h nonlinearities, as well as residual connections and skip connections."). Trueba, Gupta, Lubin, and Prenger are analogous arts because they all belong to the field of audio processing. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the synthetic speech processing system of Trueba, as modified by the speaker conversion system of Gupta and the voice modification system of Lubin, to incorporate the teachings of Prenger to use residual blocks in an audio generation neural network. Residual blocks can be used to create a system that generates “high quality audio without sacrificing quality audio without sacrificing quality at rates that may even exceed real-time requirements” (Prenger, [0002]). Generating high quality audio is important to ensure a good user experience. Regarding claim 14, the rejection of claim 12 is incorporated. Trueba, in view of Gupta and Lubin, discloses all of the elements of the current invention as stated above. However, Trueba, in view of Gupta and Lubin, fails to expressly recite wherein the generator comprises one or more residual blocks, and each of the one or more residual blocks receive the one or more of the first encodings as input. Prenger teaches wherein the generator comprises one or more residual blocks, and each of the one or more residual blocks receive the one or more of the first encodings as input (Prenger, [0022]: “In at least one embodiment, conditioning a model on other input variables, an audio transform function generation of audio with one or more characteristics encoded as a set of parameters for different properties of audio including but not limited to: pitch; loudness; intonation; rate of speech; rhythm; stress; articulation; and more. In at least one embodiment, in a multi-speaker sets of parameters are encoded for different speaker profiles and a speaker identifier is fed to a model as an extra input to condition speech.”; [0023]: "In at least one embodiment, WN( ) is an audio transformation that uses layers of dilated convolutions with gated-tan h nonlinearities, as well as residual connections and skip connections."). Trueba, Gupta, Lubin, and Prenger are analogous arts because they all belong to the field of audio processing. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the synthetic speech processing system of Trueba, as modified by the speaker conversion system of Gupta and the voice modification system of Lubin, to incorporate the teachings of Prenger to use residual blocks in an audio generation neural network. Residual blocks can be used to create a system that generates “high quality audio without sacrificing quality audio without sacrificing quality at rates that may even exceed real-time requirements” (Prenger, [0002]). Generating high quality audio is important to ensure a good user experience. Regarding claim 20, the rejection of claim 18 is incorporated. Trueba, in view of Gupta and Lubin, discloses all of the elements of the current invention as stated above. However, Trueba, in view of Gupta and Lubin, fails to expressly recite wherein the generator comprises one or more residual blocks, and each of the one or more residual blocks receive the one or more of the first encodings as input. Prenger teaches wherein the generator comprises one or more residual blocks, and each of the one or more residual blocks receive the one or more of the first encodings as input (Prenger, [0022]: “In at least one embodiment, conditioning a model on other input variables, an audio transform function generation of audio with one or more characteristics encoded as a set of parameters for different properties of audio including but not limited to: pitch; loudness; intonation; rate of speech; rhythm; stress; articulation; and more. In at least one embodiment, in a multi-speaker sets of parameters are encoded for different speaker profiles and a speaker identifier is fed to a model as an extra input to condition speech.”; [0023]: "In at least one embodiment, WN( ) is an audio transformation that uses layers of dilated convolutions with gated-tan h nonlinearities, as well as residual connections and skip connections."). Trueba, Gupta, Lubin, and Prenger are analogous arts because they all belong to the field of audio processing. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the synthetic speech processing system of Trueba, as modified by the speaker conversion system of Gupta and the voice modification system of Lubin, to incorporate the teachings of Prenger to use residual blocks in an audio generation neural network. Residual blocks can be used to create a system that generates “high quality audio without sacrificing quality audio without sacrificing quality at rates that may even exceed real-time requirements” (Prenger, [0002]). Generating high quality audio is important to ensure a good user experience. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to TYLER J BECKER whose telephone number is (703)756-1271. The examiner can normally be reached M-Th, 7:15am-5:45pm PT. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at (571) 272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /TYLER BECKER/ Examiner, Art Unit 2657 /SAMUEL G NEWAY/ Primary Examiner, Art Unit 2657
Read full office action

Prosecution Timeline

Feb 15, 2023
Application Filed
Feb 24, 2025
Non-Final Rejection mailed — §103
Jul 24, 2025
Response Filed
Oct 02, 2025
Final Rejection mailed — §103
Apr 02, 2026
Request for Continued Examination
Apr 03, 2026
Response after Non-Final Action
Jun 30, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12694228
REAL-TIME USER COMMUNICATION SENTIMENT DETECTION FOR DYNAMIC ANOMALY DETECTION AND MITIGATION
3y 4m to grant Granted Jul 28, 2026
Patent 12682113
SYSTEMS, METHODS, AND APPARATUSES FOR GENERATING STRUCTURED DATA FROM UNSTRUCTURED DATA USING NATURAL LANGUAGE PROCESSING TO GENERATE A SECURE MEDICAL DASHBOARD
3y 0m to grant Granted Jul 14, 2026
Patent 12651592
SYSTEM, METHOD, AND COMPUTER PROGRAM FOR REAL-TIME LANGUAGE TRANSLATION USING GENERATIVE ARTIFICIAL INTELLIGENCE
3y 0m to grant Granted Jun 09, 2026
Patent 12632657
Joint Speech and Text Streaming Model for ASR
2y 10m to grant Granted May 19, 2026
Patent 12614560
REVERBERATION REMOVAL DEVICE, PARAMETER ESTIMATION DEVICE, REVERBERATION REMOVAL METHOD, PARAMETER ESTIMATION METHOD, AND PROGRAM
2y 9m to grant Granted Apr 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
74%
Grant Probability
80%
With Interview (+6.3%)
2y 8m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 23 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month