DETAILED ACTION
Introduction
Applicant's submission filed on 05/08/2026 has been entered. Claims 1-20 are pending in the application and have been examined.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Examiner Recommendation
It is the recommendation of the Examiner, to include style encoder style regularization loss during the training of the autoencoder to further promote compact prosecution and putting the application in condition for allowance as indicated to the Representative during the interview on 07/09/2026 and as indicated in the included Interview Summary.
Response to Amendment
The response filed on 05/08/2026 has been correspondingly accepted and considered in this Office Action. Claims 1-20 have been examined.
Response to Arguments
Applicant's arguments filed 05/08/2026 have been fully considered as follows:
Applicant’s arguments with respect to claim 1 (also representative of claim 11) state that
“While Garbacea discloses an encoder neural network for audio coding, Garbacea's encoder is designed for general audio coding and compression, not for disentangling speaking style from linguistic content.... Accordingly, Garbacea fails to disclose "the encoder trained to disentangle speaking style information from the latent representation of linguistic content" and "the VQ layer configured to discard speaking style variations in the input speech"”
Applicant’s arguments above with respect to claim 1 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
In response to the art rejection(s) of the remainder of dependent claims are rejected under 35 U.S.C 103, in case said claims are correspondingly discussed and/or argued for at least the same rationale presented in Remarks filed 05/08/2026, Examiner respectfully notes as follows. For completeness, should the mentioned claims be likewise traversed for similar reasons to independent claims 1 and 11 correspondingly, Examiner respectfully directs Applicant to the same previous supra reasons provided in the response directed towards claims 1 and 11 correspondingly discussed above. For at least the same supra provided reasons, Examiner likewise respectfully disagrees, and Applicant's arguments have been fully considered but they are not persuasive.
Claim Rejections - 35 USC § 103
The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action.
Claims 1-4 and 11-14 are rejected under 35 U.S.C. 103 as being unpatentable over Garbacea et. al. US PgPub. 2020/0234725 in view of Wu, Da-Yi, et. al. "Vqvc+: One-shot voice conversion by vector quantization and u-net architecture." arXiv preprint arXiv:2006.04154 (2020).
Regarding claim 1, Garbacea teaches computer-implemented method executed on data processing hardware that causes the data processing hardware to perform operations comprising: receiving input speech (see Garbacea, [0020] encoder system received input audio data 102); processing, using an encoder of an autoencoder model, the input speech to predict a latent representation of linguistic content for the input speech (see Garbacea, [0020] encodes the input audio data 102 to generate a discrete latent representation 122 of the input audio data 102, [0035]), wherein the encoder comprises: a plurality of convolutional layers configured to receive the input speech as input and generate an initial discrete per-timestep latent representation of the linguistic content (see Garbacea, [0031] discusses discrete latent representation fixed duration (per timestamp), [0036] the encoder neural network 110 can be a dilated convolutional neural network that receives the sequence of audio data and generates the sequence of encoded vectors); processing, using a decoder of the autoencoder model, the latent representation of linguistic content for the input speech to predict output speech, the output speech comprising a reconstruction of the input speech(see Garbacea, [0032] the decoder system 150 generates an estimate of the input audio data 102 based on the discrete latent representation 122 of the input audio data 102; Fig 3); determining a reconstruction loss between the input speech and the reconstruction of the input speech (see Garbacea, [0098-0099] discusses reconstruction loss).
However, Garbacea fails to teach the encoder trained to disentangle speaking style information from the latent representation of linguistic content; a vector quantization (VQ) layer configured to receive each initial discrete per-timestep latent representation and predict the latent representation of linguistic content as a sequence of latent variables representing the linguistic content from the input speech, the VQ layer configured to discard speaking style variations in the input speech; determining a VQ loss for the encoder based on the latent representation of linguistic content; and training the autoencoder model on the VQ loss and the reconstruction loss.
PNG
media_image1.png
280
350
media_image1.png
Greyscale
However, Wu teaches receiving input speech (see Wu, sect 2.1 We denote X ∈ X as an au dio segment, represented as a sequence of acoustic features); processing, using an encoder of an autoencoder model, the input speech to predict a latent representation of linguistic content for the input speech, the encoder trained to disentangle speaking style information from the latent representation of linguistic content(see Wu, sect 2.1, VQVC [17] is an one-shot voice conversion system with self-reconstruction loss. The core idea is: the content information can be represented by discrete codes [24, 25], and the speaker information can be viewed as the difference between the continuous representations and the discrete codes ), wherein the encoder comprises: a plurality of convolutional layers configured to receive the input speech as input and generate an initial discrete per-timestep latent representation of the linguistic content(see Wu, sect 2.2.1 Fig. 3 shows the VQ down convolution layers further processed by codebook VQ to generate C ( latent representation of content) ); and a vector quantization (VQ) layer configured to receive each initial discrete per-timestep latent representation and predict the latent representation of linguistic content as a sequence of latent variables representing the linguistic content from the input speech, the VQ layer configured to discard speaking style variations in the input speech (see Wu, sect 2.1 equation (1)); determining a VQ loss for the encoder based on the latent representation of linguistic content(see Wu, sect. 2.1 equation 4 ( latent loss: VQ loss)); processing, using a decoder of the autoencoder model, the latent representation of linguistic content for the input speech to predict output speech, the output speech comprising a reconstruction of the input speech (see Wu, sect 2.1Then, S is added back to C, passed through the decoder, and at last we get the reconstruction , equation (2) ); determining a reconstruction loss between the input speech and the reconstruction of the input speech(see Wu, sect. 2.1, equation 2 reconstruction loss); and training the autoencoder model on the VQ loss and the reconstruction loss(see Wu, sect 2.1In the training phase, equation 3 ( reconstruction loss) equation 4 ( latent loss) equation 5 ( whole loss) ).
Garbacea and Wu are considered to be analogous to the claimed invention because they relate to speech coding using neural networks. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Garbacea on reconstruction of audio using discrete latent representation using neural networks with the disentangle latent space teachings of Wu to provide VQVC+ system which can synthesize high-quality audio ( see Wu, sect 1).
Regarding claim 2, Garbacea in view of Wu teach a method of claim 1, Garbacea further teaches, wherein the decoder comprises a plurality of convolutional layers configured to receive the latent representation of linguistic content for the input speech (see Garbacea, [0056] the decoder neural network 170 is an auto-regressive neural network, e.g., a WaveNet or other auto-regressive convolutional neural network ).
Regarding claim 3, Garbacea in view of Wu teach a method of claim 2, Garbacea further teaches, a number of the plurality of convolutional layers of the decoder is equal to a number of the plurality of convolutional layers of the encoder (see Garbacea, [0055-0056] the decoder neural network 170 is the same type of neural network as the encoder neural network 110, but configured to generate a reconstruction from a decoder input rather than an encoder output (which is the same size as the decoder input) from an input audio data. the decoder neural network 170 is an auto-regressive neural network, e.g., a WaveNet or other auto-regressive convolutional neural network ).
Regarding claim 4, Garbacea in view of Wu teach a method of claim 1, Wu further teaches, wherein the VQ loss encourages the encoder to minimize a distance between an output and a nearest codebook (see Wu, sect 2.1 In addition, the latent loss L latent is added as Equation (4), which minimizes the distance between the discrete codes and the continuous embedding). The same motivation to combine as claim 1 applies here.
Regarding claim 11, is directed to a system claim corresponding to the method claim presented in claim 1 and is rejected under the same grounds stated above regarding claim 1.
Regarding claim 12, is directed to a system claim corresponding to the method claim presented in claim 2 and is rejected under the same grounds stated above regarding claim 2.
Regarding claim 13, is directed to a system claim corresponding to the method claim presented in claim 3 and is rejected under the same grounds stated above regarding claim 3.
Regarding claim 14, is directed to a system claim corresponding to the method claim presented in claim 4 and is rejected under the same grounds stated above regarding claim 4.
Claims 5-7 and 15-17 are rejected under 35 U.S.C. 103 as being unpatentable over Garbacea et. al. US PgPub. 2020/0234725 in view of Wu, Da-Yi, et. al. "Vqvc+: One-shot voice conversion by vector quantization and u-net architecture." arXiv preprint arXiv:2006.04154 (2020) further in view of Qian, Kaizhi, et al. "Autovc: Zero-shot voice style transfer with only autoencoder loss." International Conference on Machine Learning. PMLR, 2019 (cited in IDS).
Regarding claim 5, Garbacea in view of Wu teach a method of claim 1, Garbacea further teaches, wherein processing the input speech to predict the latent representation of linguistic content for the input speech comprises processing the input speech to generate the latent representation of linguistic content as a discrete per-timestep latent representation of linguistic content that discards speaking style variations in the input speech (see Garbacea, [0040, 0041] discusses content codebook to determine the latent embedding vectors and only if speaker codebook is used uses the speaker codebook( when no speaker codebook: discards speaking styles) ). Garbacea in view of Wu teaches latent representation of linguistic content, to further teach latent representation of linguistic content as a discrete per-timestep latent representation of linguistic content that discards speaking style variations in the input speech, Quan further teaches processing the input speech to predict the latent representation of linguistic content for the input speech comprises processing the input speech to generate the latent representation of linguistic content as a discrete per-timestep latent representation of linguistic content that discards speaking style variations in the
PNG
media_image2.png
360
356
media_image2.png
Greyscale
input speech (see Qian, sect 3.3 Therefore, as shown in Fig. 2(c), when the dimension of C1 is chosen such that the dimension reduction is just enough to get rid of all the speaker information(style variations) but no content information is harmed. See Qian, sect 4. As shown in Fig. 3, AUTOVC consists of three major modules: a speaker encoder, a content encoder, a decoder. AUTOVC works on the speech mel-spectrogram of size N-by T, where N is the number of mel-frequency bins and T is the number of time steps (frames)(discrete per timestep latent representation).
Garbacea, Wu and Qian are considered to be analogous to the claimed invention because they relate to speech coding using neural networks. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Garbacea in view of Wu on reconstruction of audio using discrete latent representation using neural networks with the a content and speaking style model that includes a content encoder, a style encoder, and a decoder teachings of Qian to provide a simpler and better voice conversion and general style transfer system ( see Qian, sect 1, pg. 2).
Regarding claim 6, Garbacea in view of Wu teach a method of claim 1, Garbacea further teaches, wherein the operations further comprise processing, using a style encoder, the same input speech to generate a latent representation of speaking style for the same input speech (see Garbacea, [0040, 0041] discusses content codebook to determine the latent embedding vectors and only if speaker codebook is used uses the speaker codebook( speaking style) ). Garbacea teaches latent representation of speaking style, to further teach using a style encoder, Quan further teaches wherein the operations further comprise processing, using a style encoder, the same input speech to generate a latent representation of speaking style for the same input speech (see Qian, Fig. 1, Es (.), the Es is the style/speaker encoder which receives speech X2 ; the Es outputs the speaker embedding/style of the utterance or speech style)). The same motivation to combine as claim 5 applies here.
Regarding claim 7, Garbacea in view of Wu further in view of Qian teach a method of claim 6. Qian further teaches wherein processing, using the decoder of the autoencoder model, the latent representation of linguistic content for the input speech to predict the output speech comprises processing, using the decoder, the latent representation of linguistic content for the input speech and the latent representation of speaking style for the same input speech to predict the output speech (see Qian, Fig. 1 , sect 3.2 The framework consists of three modules, a content encoder Ec(·) that produces a content embedding from speech, a speaker encoder Es(·) that produces a speaker embedding from speech, and a decoder D(·, ·) that produce speech from content and speaker embeddings ). The same motivation to combine as claim 5 applies here.
Regarding claim 15, is directed to a system claim corresponding to the method claim presented in claim 5 and is rejected under the same grounds stated above regarding claim 5.
Regarding claim 16, is directed to a system claim corresponding to the method claim presented in claim 6 and is rejected under the same grounds stated above regarding claim 6.
Regarding claim 17, is directed to a system claim corresponding to the method claim presented in claim 7 and is rejected under the same grounds stated above regarding claim 7.
Claims 8-10 and 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over Garbacea et. al. US PgPub. 2020/0234725 in view of Wu, Da-Yi, et. al. "Vqvc+: One-shot voice conversion by vector quantization and u-net architecture." arXiv preprint arXiv:2006.04154 (2020) further in view of Qian, Kaizhi, et al. "Autovc: Zero-shot voice style transfer with only autoencoder loss." International Conference on Machine Learning. PMLR, 2019 (cited in IDS) further in view of Poole et.al, US Patent 10,671,889(cited in IDS).
Regarding claim 8, Garbacea in view of Wu further in view of Qian teaches a method of claim 6, Qian further teaches one or more convolutional layers configured to receive the input speech as input(see Qian, sect 4.2, As shown in Fig. (3)(b), the speaker encoder consists of a stack of two LSTM layers with cell size 768); and a variational layer with Gaussian posterior configured to summarize an output from the one or more convolutional layers with a global average pooling operation across the time-axis to extract a global latent style variable that corresponds to the latent representation of speaking style ((see Qian, sect 4.2 Only the output of the last time is selected and projected down to dimension 256 with a fully connected layer. The resulting speaker embedding is a 256-by-1 vector. The speaker encoder is pre-trained on the GE2E loss (Wan et al., 2018) (the SoftMax loss version), which maximizes the embedding similarity among different utterances of the same speaker, and minimizes the similarity among different speakers; interpreted as the variational layer with Gaussian posterior to summarize a global averaging pooling operation). Qian teaches generating the speaker encoding, to further teach variational layer with Gaussian posterior, Poole is used to further teach a variational layer with Gaussian posterior configured to summarize an output from the one or more convolutional layers with a global average pooling operation across the time-axis to extract a global latent style variable that corresponds to the latent representation of speaking style (see Poole, col 1 lines 19-35 a variational layer with Gaussian posterior configured to summarize an output from the one or more convolutional layers with a global average pooling operation across the time-axis to extract a global latent style variable that corresponds to the latent representation of speaking style. See Poole, col. 8 lines 13-27 teaches the system is configured to sample values for a set of latent variables from the posterior distribution for a set of latent variables which may define values for a latent variable data structure such as a latent variable vector z; the vector z is interpreted as a latent representation of speaking style). .
Garbacea in view of Wu in view of Qian teaches a content and speaking style model that includes a content encoder, a style encoder, and a decoder. Using the known technique of training a variational autoencoder neural network system to keep the posterior distribution close to a prior p(z), using a standard Gaussian to generate an output data item teachings by Poole (see Poole, col 1 lines 19-35), to provide the variational layer with Gaussian posterior in speaking style encoder in the reference Garbacea in view of Wu further in view of Qian and to global latent style variable that corresponds to the latent representation of speaking style would have been obvious to one of ordinary skill in the art.
Regarding claim 9, Garbacea in view of Wu further in view of Qian further in view of Poole teaches a method of claim 8, Poole further teaches during training, the global style latent variable is sampled from a mean and variance of style latent variables predicted by the style encoder (see Poole, col. 8 lines 13-27 The data items are provided to an encoder neural network 104 which outputs a set of parameters 106 defining a posterior distribution of a set of latent variables, e.g., defining the mean and variance of a multivariate Gaussian distribution; encoder interpreted as style encoder); and during inference, the global style latent variable is sampled from the mean of the global latent style variables predicted by the style encoder(see Poole, col. 8 lines 13-27 & Poole, col 8 lines 38-48 The system is configured to sample values for a set of latent variables 108 from the posterior distribution; system is a VAE Neural network and the set of latent variables are interpreted as global style latent variable).
Regarding claim 10, Garbacea in view of Wu further in view of Qian further in view of Poole teaches a method of claim 8, Poole further teaches wherein the style encoder is trained using a style regularization loss based on a mean and variance of style latent variables predicted by the style encoder, the style encoder using the style regularization loss to minimize a Kullback-Leibler (KL) divergence between a Gaussian posterior with a unit Gaussian prior (see Poole, col 8 lines 38-48 and Poole, col 1 lines 60-66, the variational autoencoder neural network system may configured for training with an objective function. This may have a first term, such as a cross-entropy term, dependent upon a difference between the input data item and the output data item and a second term, such as a KL divergence term, dependent upon a difference between the posterior distribution and a second, prior distribution of the set of latent variables).
Regarding claim 18, is directed to a system claim corresponding to the method claim presented in claim 8 and is rejected under the same grounds stated above regarding claim 8.
Regarding claim 19, is directed to a system claim corresponding to the method claim presented in claim 9 and is rejected under the same grounds stated above regarding claim 9.
Regarding claim 20, is directed to a system claim corresponding to the method claim presented in claim 10 and is rejected under the same grounds stated above regarding claim 10.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Trueba et. al. US Patent 11,735,156 teaches a speech-processing system determines vocal characteristics of the second voice and determines output corresponding to the first audio data and the vocal characteristics (see Trueba, Fig. 1).
Zhang et. al. US PgPub 2020/0365166 teaches a zero-shot voice conversion with non-parallel data includes receiving source speaker speech data as input data into a content encoder of a style transfer autoencoder system (see Zhang, Fig. 3).
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to NANDINI SUBRAMANI whose telephone number is (571)272-3916. The examiner can normally be reached Monday - Friday 12:00pm - 5:00 pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bhavesh M Mehta can be reached at (571)272-7453. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/NANDINI SUBRAMANI/Examiner, Art Unit 2656
/BHAVESH M MEHTA/Supervisory Patent Examiner, Art Unit 2656