DETAILED ACTION
This office action is in response to Applicant’s Amendment/Request for Reconsideration, received on 06/18/2026. Claims 1 and 14 have been amended. Claims 1-26 are pending and have been considered.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant's arguments filed 06/18/2026, see pg. 13, with respect to “Objections to the Specification” have been fully considered but they are not persuasive.
Applicant’s representative asserts “Applicant has amended Paragraphs [0045] and [0046] to change instances of ‘contrastive loss’ designated as reference numeral 316 to reference numeral 315”.
In response, the examiner recognizes the amendments to the specification, but respectfully asserts that they have been not been correctly entered with regard to 37 C.F.R. 1.125: “The text of any deleted matter must be shown by strike-through except that double brackets placed before and after the deleted characters may be used to show deletion of five or fewer consecutive characters”. Applicant has not included double brackets around the incorrect reference numeral of 315. The specification remains objected to because these instances now assign two differing reference numerals to the same component in paragraphs [0045] and [0046]. See updated objections below.
Applicant’s arguments, see pgs. 13-14, filed 06/18/2026, with respect to the rejection(s) of claim(s) 1 and 14 under 35 U.S.C. 103 (Yang in view of Wang) have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of Zheng et al. (US-20230169281-A1), hereinafter Zheng. Zheng discloses “Representation learning for text and speech has improved many language-related tasks. However, existing methods only learn from one input modality, while a unified representation for both speech and text is needed for tasks such as end-to-end speech translation. Consequently, these methods cannot exploit various large-scale text and speech data and their performance is limited by the scarcity of parallel speech translation data. To address these problems, embodiments of a fused acoustic and text masked language model (FAT-MLM) are disclosed. FAT-MLM embodiments jointly learn a unified representation for both acoustic and text input from various types of corpora including parallel data for speech recognition and machine translation, and pure speech and text data. Within this cross-modal representation learning framework, an end-to-end model is further presented for fused acoustic and text speech translation. Experiments show that by fine-tuning from FAT-MLM, the speech translation model embodiments substantially improve translation quality” (abstract). See updated rejections below.
Applicant’s arguments with respect to claim(s) 2-13, 15-26 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Specification
The disclosure is objected to because of the following informalities: [0045] refers to a “contrastive loss 316 315”. The previous reference numeral of 316 has been improperly deleted. The contrastive loss module of Fig. 3A (and all other references to the contrastive loss) are given a reference numeral of 315. The examiner believes this instance should be amended to “contrastive loss [[316]] 315”.
[0046] refers to a “contrastive context vector 315 215”. The previous reference numeral of 315 has been improperly deleted. The examiner believes this instance should be amended to “contrastive loss [[315]] 215”.
Appropriate correction is required.
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Claims 1, 3-4, 7, 9, 12, 14, 16-17, 20, 22, 25 are provisionally rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1, 5, 9, 10, 12-13, 17, 21, 22, 24 of copending Application No. 18/494,324 (reference application). Although the claims at issue are not identical, they are not patentably distinct from each other, as the chart shows.
This is a provisional nonstatutory double patenting rejection because the patentably indistinct claims have not in fact been patented.
18823661 (instant app)
18494324
Examiner Notes
A computer-implemented method executed on data processing hardware that causes the data processing hardware to perform operations comprising:
A computer-implemented method that when executed on data processing hardware causes the data processing hardware to perform operations comprising:
receiving training data comprising a plurality of sets of training utterances, each set of training utterances associated with a respective language that is different than the respective language associated with each other set of the training utterances and comprising speech spoken in the respective language, each training utterance comprising a corresponding reference speech representation paired with a corresponding input text sequence;
receiving training data comprising a plurality of sets of text-to-speech (TTS) spoken utterances, each set of the TTS spoken utterances associated with a respective language from among a plurality of different languages that is different than the respective language associated with each other set of the TTS spoken utterances and comprising TTS utterances of synthetic speech spoken in the respective language, each TTS utterance of synthetic speech comprising a corresponding reference speech representation paired with a corresponding input text sequence;
for each training utterance in each set of training utterances of the received training data:
for each TTS utterance in each set of the TTS spoken utterances of the received training data:
generating, using a text encoder, a corresponding encoded textual representation for the corresponding input text sequence;
generating, using a text encoder, a corresponding TTS encoded textual representation for the corresponding input text sequence;
generating, using a speech encoder, a corresponding speech encoding for the corresponding reference speech representation;
generating, using a speech encoder, a corresponding speech encoding for the corresponding TTS utterance of synthetic speech;
generating, using a shared encoder configured to receive the corresponding encoded textual representation, a first shared encoder output;
generating, using a shared encoder configured to receive the corresponding TTS encoded textual representation or the corresponding speech encoding, a shared encoder output;
See below
generating, using the shared encoder configured to receive the corresponding speech encoding, a second shared encoder output:
A shared encoder which can receive both an encoded textual representation or a corresponding speech encoding (as seen in ‘324) indicates that step to disclose both of the separate ‘generating…’ elements of the instant app
generating, using a speech decoder configured to receive the shared encoder output, a predicted speech representation for the corresponding TTS utterance of synthetic speech;
generating, using an automatic speech recognition (ASR) decoder configured to receive the shared encoder output, a sequence of speech recognition hypotheses representing a candidate transcription for the corresponding TTS utterance of synthetic speech;
determining an ASR loss based on the sequence of speech recognition hypotheses and the corresponding input text sequence;
determining a text-to-speech (TTS) loss based on the corresponding encoded textual representation, the corresponding speech encoding, and the shared encoder output;
determining a reconstruction loss based on the predicted speech representation and the corresponding reference speech representation for the corresponding TTS utterance;
Reconstruction loss and TTS loss are synonymous in the context of speech generation, wherein the predicted speech representation of ‘324 is within the shared encoder output and a reference speech tracks to a corresponding speech encoding
training a TTS model based on the TTS losses determined for the training utterances in each set of the training utterances to teach the TTS model to learn how to synthesize speech in each of the respective languages.
training a TTS model based on the reconstruction losses and the ASR losses determined for the TTS utterances in each set of the TTS spoken training utterances to teach the TTS model to learn how to synthesize speech in each of the plurality of different languages
Underlined sections of the claims have been incorporated to indicate where similar subject matter of the claim sets is found within dependent claims. As the table above demonstrates, each limitation of claim 1 of the present application is found in claim 1 of copending Application 18/494,324, thus claim 1 of the present application is anticipated by claim 1 of the copending application. Independent claim 14 is similarly anticipated by claim 13 of the copending application. Dependent claims 3-4, 7, 9, and 12 of the present application are substantially similar to claims 1, 5, 9, 10, and 12 respectively (and associated equivalent claims for the system of claim 14) and are, therefore, also anticipated by copending application 18/494,324, as shown below.
18/823,661
18/494,324
3. generating, using an automatic speech recognition (ASR) decoder configured to receive the shared encoder output as input, a speech recognition hypothesis representing a candidate transcription for the corresponding training utterance; and determining an ASR loss based on the speech recognition hypothesis and the corresponding input text sequence, wherein the TTS loss comprises the ASR loss.
1. generating, using an automatic speech recognition (ASR) decoder configured to receive the shared encoder output, a sequence of speech recognition hypotheses representing a candidate transcription for the corresponding TTS utterance of synthetic speech; and determining an ASR loss based on the sequence of speech recognition hypotheses and the corresponding input text sequence;
4. The computer-implemented method of claim 3, wherein the ASR decoder comprises a recurrent neural network-transducer (RNN-T) architecture.
5. (Original) The computer-implemented method of claim 1, wherein the speech decoder comprises a recurrent neural network-transducer (RNN-T) architecture.
7. The computer-implemented method of claim 1, wherein: the training data further comprises unspoken textual utterances associated with a respective plurality of different languages, each unspoken textual utterance not paired with any corresponding spoken utterance; and the operations further comprise, for each unspoken textual utterance: generating, using the text encoder, a corresponding unspoken encoded textual representation for the corresponding unspoken textual utterance; and determining an aligned-text masked language modeling (MLM) loss for the corresponding unspoken encoded textual representation generated for the corresponding unspoken textual utterance, wherein the TTS loss comprises the aligned-text MLM loss.
9. (Original) The computer-implemented method of claim 1, wherein: the training data further comprises unspoken textual utterances in a respective plurality of different languages, each unspoken textual utterance not paired with any corresponding spoken utterance of synthetic speech; and the operations further comprise, for each unspoken textual utterance: generating, using the text encoder, a corresponding unspoken encoded textual representation for the corresponding unspoken textual utterance; and obtaining an aligned masked language modeling (MLM) loss for the corresponding unspoken encoded textual representation generated for the corresponding unspoken textual utterance, wherein training the TTS model further comprises training the TTS model based on the aligned MLM loss obtained for the unspoken encoded textual representation.
9. The computer-implemented method of claim 1,wherein: the training data further comprises unpaired spoken utterances spoken in a respective plurality of different languages, each unpaired spoken utterance not paired with any corresponding text; and the operations further comprise, for each unpaired spoken utterance. generating, using the speech encoder, a corresponding unpaired speech encoding for the corresponding unpaired spoken utterance; and determining an aligned-speech masked language modeling (MLM) loss for the corresponding unpaired speech encoding generated for the corresponding unpaired spoken utterance, wherein the TTS loss comprises the aligned-speech MLM loss.
10. (Original) The computer-implemented method of claim 1, wherein: the training data further comprises un-transcribed non-synthetic speech utterances in a respective plurality of different languages, each un-transcribed non-synthetic speech utterance not paired with a corresponding transcription; and the operations further comprise, for each un-transcribed non-synthetic speech utterance: generating, using the speech encoder, a corresponding speech encoding for the corresponding un-transcribed non-synthetic speech utterance; and obtaining a masked language modeling (MLM) loss for the corresponding speech encoding generated for the corresponding un-transcribed non-synthetic speech utterance, wherein training the TTS model further comprises training the TTS model based on the MLM loss obtained for the corresponding speech encoding.
12. The computer-implemented method of claim 1, wherein each corresponding input text sequence comprises a sequence of graphemes, word-piece-model units, phonemes, or bytes.
12. (Original) The computer-implemented method of claim 1, wherein each corresponding input text sequence comprises a sequence of graphemes, word-piece-model units, phonemes, or bytes.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-2, 5-6, 12, 14-15, 18-19, and 25 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yang et al. (US-20220246136-A1), hereinafter Yang, in view of Zheng et al. (US-20230169281-A1), hereinafter Zheng, further in view of Wang et al. (US-20240274122-A1), hereinafter Wang.
Regarding claim 1, Yang discloses: A computer-implemented method executed on data processing hardware that causes the data processing hardware to perform operations ([0133] Processors are described in connection with various apparatus and methods. These processors may be implemented using electronic hardware) comprising:
receiving training data comprising a plurality of sets of training utterances ([Fig. 11, Multilingual Corpus 1120 used for training Multilingual Neural TTS system 1110], [0086] [0086] Training data for any one or any combinations of the speaker encoder 1112, the language encoder 1114, the acoustic feature predictor 1116, and the neural vocoder 1118 may be obtained based on speech waveforms in the multilingual corpus 1120, [A speech waveform indicates a training utterance]), each set of training utterances associated with a respective language that is different than the respective language associated with each other set of the training utterances and comprising speech spoken in the respective language ([0053] The language embedding generator 600 may be trained with a corpus set of multiple languages, and is designed for language recognition that is independent of text or content, [In view of the plurality of languages defined in the previously cited corpus]), each training utterance comprising a corresponding reference speech representation paired with a corresponding input text sequence ([0086] various derived information may be obtained from the speech waveforms, e.g., text information obtained by applying any speech recognition technologies, acoustic features obtained by applying any acoustic feature extraction technologies, speaker embedding vectors obtained by applying any speaker recognition technologies, language embedding vectors obtained by applying any language recognition technologies, etc. The derived information along with the speech waveforms in the multilingual corpus 1120 may form various training data for any one or any combinations of the speaker encoder 1112, the language encoder 1114, the acoustic feature predictor 1116, and the neural vocoder 1118, [Providing derived information (including an input text sequence, see text input 1202 for training acoustic feature predictor 1210 of Fig. 12) along with the speech waveform (reference speech) indicates a pairing of these items for training]);
for each training utterance in each set of training utterances of the received training data:
generating, using a text encoder ([Fig. 2, Encoder 212 receiving Text Input 202]), a corresponding encoded textual representation for the corresponding input text sequence ([0068] A text input 902 may be provided to the encoder 910 which may correspond to the encoder 212 in FIG. 2); and,
generating, using a speech encoder ([Fig. 2, Speaker Encoder 230]), a corresponding speech encoding for the corresponding reference speech representation ([0041] The speaker encoder 230 may provide speaker latent space information 232 of a target speaker, [As previously disclosed via the training operation/data of Fig. 11, providing speech waveform data to train the speaker encoder 1112 of Fig. 11 indicates the speaker encodings to be generated using the training, i.e. reference, speech]).
Yang does not disclose:
generating, using a shared encoder configured to receive the corresponding encoded textual representation, a first shared encoder output; and,
generating, using the shared encoder configured to receive the corresponding speech encoding, a second shared encoder output.
Zheng discloses:
generating, using a shared encoder configured to receive the corresponding encoded textual representation ([Fig. 5, Transformer Encoder 510 which receives Text embeddings 504/506], [0059] A multimodal transformer encoder 510 encodes the concatenated embeddings h.sub.s,x,y into a unified representation ƒ(h.sub.s,x,y) 512 for speech, source language texts, and target language texts , [The examiner asserts that an embedding tracks to an encoding]), a first shared encoder output ([Fig. 5, Token 516 output], [0059] one or more reconstructed target tokens 516 corresponding to the one or more masked target tokens); and,
generating, using the shared encoder configured to receive the corresponding speech encoding ([Fig. 5, previously cited Transformer Encoder which receives Acoustic Embeddings 502]), a second shared encoder output ([Fig. 5, Acoustic features output from Transformer Model and sent to Speech Reconstruction Module], [0059] The unified representation ƒ(h.sub.s,x,y) may be used to reconstruct a reconstructed sequence of acoustic features using a speech reconstruction module 540).
Yang and Zheng are considered analogous art within multi-lingual speech synthesis. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Yang to incorporate the teachings of Zheng, because of the novel way to unify text and acoustic representations within a FAT-MLM, improving the quality of end-to-end speech translation models while still maintaining a smaller model size and faster decoding time (Zheng, [0037]-[0041]).
Yang in view of Zheng does not disclose:
for each training utterance in each set of training utterances of the received training data:
determining a text-to-speech (TTS) loss based on the corresponding encoded textual representation, the corresponding speech encoding, and the shared encoder output.
Wang discloses:
for each training utterance in each set of training utterances of the received training data:
determining a text-to-speech (TTS) loss based on the corresponding encoded textual representation, the corresponding speech encoding, and the shared encoder output ([Fig. 4, Training 480 Based on Transcript Embedding 175 (encoded textual representation), Output Audio Data 185 (shared encoder, i.e. that which receives a speech encoding, output) and Performance Representation 225 (which is a transformation of acoustic embedding data Z 215, i.e. speech encoding)], [0063] During training, the various models of the system may be trained based on a comparison of the transcript embedding data 175 and the performance representation f.sub.θ(Z) 225 generated by the invertible transformation 240. The spectrogram encoder 210, as condition by the speaker embedding data 135, may process the spectrogram data 125 representing the source speech to generate acoustic embedding data Z 215. The acoustic embedding decoder 220 may process the acoustic embedding data Z 215 to generate the output audio data 185. The output audio data 185 may be compared to the input audio data 111 (e.g., from which the spectrogram data 125 was generated) to calculate the result of a loss function (or multiple loss functions) with gradients propagated back through the spectrogram encoder 210 and the acoustic embedding decoder 220).
Yang, Zheng, and Wang are considered analogous art within multi-lingual speech synthesis. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Yang in view of Zheng to incorporate the teachings of Wang, because of the novel way to jointly encode performance features like phoneme duration, prosody, or pitch and determining a target language based on the features, allowing for more accurate reproduction of vocal performance characteristics of speech samples in a target language (Wang, [0023]).
Yang further discloses:
training a TTS model based on the TTS losses determined for the training utterances in each set of the training utterances to teach the TTS model to learn how to synthesize speech in each of the respective languages ([0052] Specifically, in addition to the loss adopted in the training of the conventional neural TTS systems, additional losses may be introduced to ensure that the embedding vectors generated for the same language are in proximity to each other, and the embedding vectors generated for different languages are far from each other. For example, an additional loss function may be defined to minimize the distance among the embedding vectors of the same language, maximize the distance among the embedding vectors of different languages, etc., [In view of the TTS loss function of Wang which determines loss based on output, i.e. from the vocoder of Yang (which itself depends upon a textual encoding and speech encoding as previously disclosed, indicating the elements of Yang could be substituted in the loss function of Wang without a change in functionality to either art), as applied to the plurality of languages of Yang]).
Regarding claim 2, Yang in view of Zheng, further in view of Wang discloses: the computer-implemented method of claim 1.
Yang further discloses:
wherein the operations further comprise, for each training utterance in each set of the training utterances of the received training data:
obtaining a corresponding speaker embedding characterizing speaker characteristics of a corresponding speaker that spoke the training utterance in the respective language ([Fig. 3, Target Speaker ID 302, Target Speaker Corpus 304], [0047] The speaker embedding selector 310 may attempt to retrieve a speaker embedding vector corresponding to the target speaker ID 302, [0048] The speaker embedding generator 320 may generate a speaker embedding vector corresponding to a target speaker based on a corpus 304 of the target speaker. For example, the corpus 304 of the target speaker may be obtained, which comprises multiple speech waveforms of the target speaker. Acoustic features may be extracted from the speech waveforms in the corpus 304); and
obtaining a corresponding language embedding identifying the respective language of the utterance ([Fig. 5, Reference Language ID 502], [Fig. 11, Multilingual Corpus 1120], [Having a multilingual corpus of training data with each language defined indicates the language ID identifying the language of the corpuses]),
wherein the text encoder is configured to receive a concatenation of the corresponding speaker embedding and the corresponding language embedding ([Fig. 14, Generate an acoustic feature 1420 based on text input 1402, target speaker embeddings 1412, and target language embedding 1414], [0104] an acoustic feature corresponding to the text input 1402 may be generated through taking the speaker latent space information corresponding to the target speaker and the language latent space information corresponding to the reference language, [Generation of an acoustic feature based on the combination of text, speaker, and language embeddings indicates a required concatenation of the embeddings to generated the acoustic feature, wherein the text encoder of Yang is within the acoustic feature predictor which receives the required embeddings, indicating the text encoder to also be receiving the required embeddings as part of the system 210]).
Regarding claim 5, Yang in view of Zheng, further in view of Wang discloses: the computer-implemented method of claim 1.
Yang further discloses:
wherein the operations further comprise, for each training utterance in each set of the training utterances of the received training data:
determining a feature loss between the encoded textual representation generated for the corresponding input text sequence using the text encoder and the speech encodings generated for the corresponding reference speech representation using the speech encoder ([0048] Specifically, in addition to the loss adopted in the training of the conventional neural TTS systems, additional losses may be introduced to ensure that the embedding vectors generated for the same speaker are in proximity to each other, and the embedding vectors generated for different speakers are far from each other. For example, an additional loss function may be defined to minimize the distance among the embedding vectors of the same speaker, maximize the distance among the embedding vectors of different speakers, etc., [In view of the multilingual training corpus of Fig. 11, indicating the determination as to distances of speaker embeddings are based on reference samples to know how to define proximity]),
wherein the TTS loss comprises the feature loss ([Defining additional losses of TTS systems, wherein the speaker encoder is within the multilingual neural TTS system, indicates the feature loss, i.e. distance, to be at least part of the TTS loss. See [0088] discussion of ground-truth acoustic features]).
Regarding claim 6, Yang in view of Zheng, further in view of Wang discloses: the computer-implemented method of claim 1.
Wang further discloses:
wherein the operations further comprise, for each training utterance in each set of the training utterances of the received training data:
obtaining a sequence representation of the corresponding input text sequence concatenated with a variational embedding ([0031] a representational encoder 170 configured to process the transcript data 165 and the performance embedding data 145 to generate transcript embedding data 175 that represents both the content (e.g., text) of the speech to be synthesized as well as the vocal performance characteristics of the source speech, [Vocal performance characteristics of how transcribed speech should be said track to variational embeddings which are concatenated with the transcript sequence resulting in the transcript embedding]);
using a duration model ([Fig. 2, Duration Prediction 250]):
predicting, based on the sequence representation, a duration of the input text sequence ([0054] The duration predictor 250 may use the transcript embedding data 175 and the language embedding data 195 for the target language to predict durations (e.g., corresponding to the words, phrases, segments, etc. of the input audio data 111)); and
upsampling, based on the duration of the input text sequence, the sequence representation into an upsampled output specifying a number of frames ([0054] An upsampler component 260 may use the scaling factor convert the transcript embedding data 175 into a data stream having a frame rate of audio data (e.g., with a vector corresponding to each 10 ms, 20 ms, or 30 ms, etc. of audio data). For example, representational embedding data corresponding to a first phoneme may be duplicated to generate a number of frames corresponding to the predicted duration of that phoneme); and
determining a duration loss based on the predicted duration of the input text sequence and a ground-truth duration ([Fig. 2, Duration Predictor 250 and Actual Durations 290], [0053] The system 100 may generate duration data 295, which may indicate a start time and duration of a given segment of speech. The duration scaling component 255 may use the duration data 295 to modify the transcript embedding data 175 to reflect the determined durations (e.g., by scaling the transcript embedding data 175 from the predicted durations determined by the duration predictor 250), [Scaling predicted duration data to match determined durations (which are based on ground truth duration data) indicates the scaling to minimize a loss between the two sequences]),
wherein the TTS loss comprises the duration loss ([As previously disclosed, the duration loss minimization of Wang is within a TTS system, indicating the duration loss to be part of the TTS loss]).
Regarding claim 12, Yang in view of Zheng, further in view of Wang discloses: the computer-implemented method of claim 1.
Wang further discloses:
wherein each corresponding input text sequence comprises a sequence of graphemes, word-piece-model units, phonemes, or bytes ([Fig. 7, Source Text 755 being passed through G2P 770], [0043] In some implementations, the system 100 may further include a grapheme-to-phoneme (G2P) component such as the G2P component 770 shown in FIG. 7).
Regarding claim 14, Yang discloses: A system ([0026] Embodiments of the present disclosure proposes a multilingual neural TTS system) comprising:
data processing hardware ([0133] Processors are described in connection with various apparatus and methods. These processors may be implemented using electronic hardware); and
memory hardware in communication with the data processing hardware ([0134] Computer readable medium may include, e.g., a memory), the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations ([0134] Software should be considered broadly to represent instructions, instruction sets, code, code segments, program code, programs, subroutines, software modules, applications, software applications, software packages, routines, subroutines, objects, running threads, processes, functions, etc. Software may reside on computer readable medium) comprising:
receiving training data comprising a plurality of sets of training utterances ([Fig. 11, Multilingual Corpus 1120 used for training Multilingual Neural TTS system 1110], [0086] [0086] Training data for any one or any combinations of the speaker encoder 1112, the language encoder 1114, the acoustic feature predictor 1116, and the neural vocoder 1118 may be obtained based on speech waveforms in the multilingual corpus 1120, [A speech waveform indicates a training utterance]), each set of training utterances associated with a respective language that is different than the respective language associated with each other set of the training utterances and comprising speech spoken in the respective language ([0053] The language embedding generator 600 may be trained with a corpus set of multiple languages, and is designed for language recognition that is independent of text or content, [In view of the plurality of languages defined in the previously cited corpus]), each training utterance comprising a corresponding reference speech representation paired with a corresponding input text sequence ([0086] various derived information may be obtained from the speech waveforms, e.g., text information obtained by applying any speech recognition technologies, acoustic features obtained by applying any acoustic feature extraction technologies, speaker embedding vectors obtained by applying any speaker recognition technologies, language embedding vectors obtained by applying any language recognition technologies, etc. The derived information along with the speech waveforms in the multilingual corpus 1120 may form various training data for any one or any combinations of the speaker encoder 1112, the language encoder 1114, the acoustic feature predictor 1116, and the neural vocoder 1118, [Providing derived information (including an input text sequence, see text input 1202 for training acoustic feature predictor 1210 of Fig. 12) along with the speech waveform (reference speech) indicates a pairing of these items for training]);
for each training utterance in each set of training utterances of the received training data:
generating, using a text encoder ([Fig. 2, Encoder 212 receiving Text Input 202]), a corresponding encoded textual representation for the corresponding input text sequence ([0068] A text input 902 may be provided to the encoder 910 which may correspond to the encoder 212 in FIG. 2); and,
generating, using a speech encoder ([Fig. 2, Speaker Encoder 230]), a corresponding speech encoding for the corresponding reference speech representation ([0041] The speaker encoder 230 may provide speaker latent space information 232 of a target speaker, [As previously disclosed via the training operation/data of Fig. 11, providing speech waveform data to train the speaker encoder 1112 of Fig. 11 indicates the speaker encodings to be generated using the training, i.e. reference, speech]).
Yang does not disclose:
generating, using a shared encoder configured to receive the corresponding encoded textual representation, a first shared encoder output; and,
generating, using the shared encoder configured to receive the corresponding speech encoding, a second shared encoder output.
Zheng discloses:
generating, using a shared encoder configured to receive the corresponding encoded textual representation ([Fig. 5, Transformer Encoder 510 which receives Text embeddings 504/506], [0059] A multimodal transformer encoder 510 encodes the concatenated embeddings h.sub.s,x,y into a unified representation ƒ(h.sub.s,x,y) 512 for speech, source language texts, and target language texts , [The examiner asserts that an embedding tracks to an encoding]), a first shared encoder output ([Fig. 5, Token 516 output], [0059] one or more reconstructed target tokens 516 corresponding to the one or more masked target tokens); and,
generating, using the shared encoder configured to receive the corresponding speech encoding ([Fig. 5, previously cited Transformer Encoder which receives Acoustic Embeddings 502]), a second shared encoder output ([Fig. 5, Acoustic features output from Transformer Model and sent to Speech Reconstruction Module], [0059] The unified representation ƒ(h.sub.s,x,y) may be used to reconstruct a reconstructed sequence of acoustic features using a speech reconstruction module 540).
Yang and Zheng are considered analogous art within multi-lingual speech synthesis. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Yang to incorporate the teachings of Zheng, because of the novel way to unify text and acoustic representations within a FAT-MLM, improving the quality of end-to-end speech translation models while still maintaining a smaller model size and faster decoding time (Zheng, [0037]-[0041]).
Yang in view of Zheng does not disclose:
for each training utterance in each set of training utterances of the received training data:
determining a text-to-speech (TTS) loss based on the corresponding encoded textual representation, the corresponding speech encoding, and the shared encoder output.
Wang discloses:
for each training utterance in each set of training utterances of the received training data:
determining a text-to-speech (TTS) loss based on the corresponding encoded textual representation, the corresponding speech encoding, and the shared encoder output ([Fig. 4, Training 480 Based on Transcript Embedding 175 (encoded textual representation), Output Audio Data 185 (shared encoder, i.e. that which receives a speech encoding, output) and Performance Representation 225 (which is a transformation of acoustic embedding data Z 215, i.e. speech encoding)], [0063] During training, the various models of the system may be trained based on a comparison of the transcript embedding data 175 and the performance representation f.sub.θ(Z) 225 generated by the invertible transformation 240. The spectrogram encoder 210, as condition by the speaker embedding data 135, may process the spectrogram data 125 representing the source speech to generate acoustic embedding data Z 215. The acoustic embedding decoder 220 may process the acoustic embedding data Z 215 to generate the output audio data 185. The output audio data 185 may be compared to the input audio data 111 (e.g., from which the spectrogram data 125 was generated) to calculate the result of a loss function (or multiple loss functions) with gradients propagated back through the spectrogram encoder 210 and the acoustic embedding decoder 220).
Yang, Zheng, and Wang are considered analogous art within multi-lingual speech synthesis. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Yang in view of Zheng to incorporate the teachings of Wang, because of the novel way to jointly encode performance features like phoneme duration, prosody, or pitch and determining a target language based on the features, allowing for more accurate reproduction of vocal performance characteristics of speech samples in a target language (Wang, [0023]).
Yang further discloses:
training a TTS model based on the TTS losses determined for the training utterances in each set of the training utterances to teach the TTS model to learn how to synthesize speech in each of the respective languages ([0052] Specifically, in addition to the loss adopted in the training of the conventional neural TTS systems, additional losses may be introduced to ensure that the embedding vectors generated for the same language are in proximity to each other, and the embedding vectors generated for different languages are far from each other. For example, an additional loss function may be defined to minimize the distance among the embedding vectors of the same language, maximize the distance among the embedding vectors of different languages, etc., [In view of the TTS loss function of Wang which determines loss based on output, i.e. from the vocoder of Yang (which itself depends upon a textual encoding and speech encoding as previously disclosed, indicating the elements of Yang could be substituted in the loss function of Wang without a change in functionality to either art), as applied to the plurality of languages of Yang]).
Regarding claim 15, Yang in view of Zheng, further in view of Wang discloses: the system of claim 14.
Yang further discloses:
wherein the operations further comprise, for each training utterance in each set of the training utterances of the received training data:
obtaining a corresponding speaker embedding characterizing speaker characteristics of a corresponding speaker that spoke the training utterance in the respective language ([Fig. 3, Target Speaker ID 302, Target Speaker Corpus 304], [0047] The speaker embedding selector 310 may attempt to retrieve a speaker embedding vector corresponding to the target speaker ID 302, [0048] The speaker embedding generator 320 may generate a speaker embedding vector corresponding to a target speaker based on a corpus 304 of the target speaker. For example, the corpus 304 of the target speaker may be obtained, which comprises multiple speech waveforms of the target speaker. Acoustic features may be extracted from the speech waveforms in the corpus 304); and
obtaining a corresponding language embedding identifying the respective language of the utterance ([Fig. 5, Reference Language ID 502], [Fig. 11, Multilingual Corpus 1120], [Having a multilingual corpus of training data with each language defined indicates the language ID identifying the language of the corpuses]),
wherein the text encoder is configured to receive a concatenation of the corresponding speaker embedding and the corresponding language embedding ([Fig. 14, Generate an acoustic feature 1420 based on text input 1402, target speaker embeddings 1412, and target language embedding 1414], [0104] an acoustic feature corresponding to the text input 1402 may be generated through taking the speaker latent space information corresponding to the target speaker and the language latent space information corresponding to the reference language, [Generation of an acoustic feature based on the combination of text, speaker, and language embeddings indicates a required concatenation of the embeddings to generated the acoustic feature, wherein the text encoder of Yang is within the acoustic feature predictor which receives the required embeddings, indicating the text encoder to also be receiving the required embeddings as part of the system 210]).
Regarding claim 18, Yang in view of Zheng, further in view of Wang discloses: the system of claim 14.
Yang further discloses:
wherein the operations further comprise, for each training utterance in each set of the training utterances of the received training data:
determining a feature loss between the encoded textual representation generated for the corresponding input text sequence using the text encoder and the speech encodings generated for the corresponding reference speech representation using the speech encoder ([0048] Specifically, in addition to the loss adopted in the training of the conventional neural TTS systems, additional losses may be introduced to ensure that the embedding vectors generated for the same speaker are in proximity to each other, and the embedding vectors generated for different speakers are far from each other. For example, an additional loss function may be defined to minimize the distance among the embedding vectors of the same speaker, maximize the distance among the embedding vectors of different speakers, etc., [In view of the multilingual training corpus of Fig. 11, indicating the determination as to distances of speaker embeddings are based on reference samples to know how to define proximity]),
wherein the TTS loss comprises the feature loss ([Defining additional losses of TTS systems, wherein the speaker encoder is within the multilingual neural TTS system, indicates the feature loss, i.e. distance, to be at least part of the TTS loss. See [0088] discussion of ground-truth acoustic features]).
Regarding claim 19, Yang in view of Zheng, further in view of Wang discloses: the system of claim 14.
Wang further discloses:
wherein the operations further comprise, for each training utterance in each set of the training utterances of the received training data:
obtaining a sequence representation of the corresponding input text sequence concatenated with a variational embedding ([0031] a representational encoder 170 configured to process the transcript data 165 and the performance embedding data 145 to generate transcript embedding data 175 that represents both the content (e.g., text) of the speech to be synthesized as well as the vocal performance characteristics of the source speech, [Vocal performance characteristics of how transcribed speech should be said track to variational embeddings which are concatenated with the transcript sequence resulting in the transcript embedding]);
using a duration model ([Fig. 2, Duration Prediction 250]):
predicting, based on the sequence representation, a duration of the input text sequence ([0054] The duration predictor 250 may use the transcript embedding data 175 and the language embedding data 195 for the target language to predict durations (e.g., corresponding to the words, phrases, segments, etc. of the input audio data 111)); and
upsampling, based on the duration of the input text sequence, the sequence representation into an upsampled output specifying a number of frames ([0054] An upsampler component 260 may use the scaling factor convert the transcript embedding data 175 into a data stream having a frame rate of audio data (e.g., with a vector corresponding to each 10 ms, 20 ms, or 30 ms, etc. of audio data). For example, representational embedding data corresponding to a first phoneme may be duplicated to generate a number of frames corresponding to the predicted duration of that phoneme); and
determining a duration loss based on the predicted duration of the input text sequence and a ground-truth duration ([Fig. 2, Duration Predictor 250 and Actual Durations 290], [0053] The system 100 may generate duration data 295, which may indicate a start time and duration of a given segment of speech. The duration scaling component 255 may use the duration data 295 to modify the transcript embedding data 175 to reflect the determined durations (e.g., by scaling the transcript embedding data 175 from the predicted durations determined by the duration predictor 250), [Scaling predicted duration data to match determined durations (which are based on ground truth duration data) indicates the scaling to minimize a loss between the two sequences]),
wherein the TTS loss comprises the duration loss ([As previously disclosed, the duration loss minimization of Wang is within a TTS system, indicating the duration loss to be part of the TTS loss]).
Regarding claim 25, Yang in view of Zheng, further in view of Wang discloses: the system of claim 14.
Wang further discloses:
wherein each corresponding input text sequence comprises a sequence of graphemes, word-piece-model units, phonemes, or bytes ([Fig. 7, Source Text 755 being passed through G2P 770], [0043] In some implementations, the system 100 may further include a grapheme-to-phoneme (G2P) component such as the G2P component 770 shown in FIG. 7).
Claim(s) 3-4, 16-17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yang in view of Zheng, further in view of Wang, further in view of Thomas et al. (US-20240371361-A1), hereinafter Thomas.
Regarding claim 3, Yang in view of Zheng, further in view of Wang discloses: the computer-implemented method of claim 1.
Wang further discloses:
wherein the operations further comprise, for each training utterance in each set of the training utterances of the received training data:
generating, using an automatic speech recognition (ASR) decoder configured to receive the shared encoder output as input ([Fig. 1B, ASR 150 receiving Spectrogram Data 125], [Fig. 9, Speech Recognition Engine 958], [The speech recognition engine is the ASR decoder receiving shared encoder output, i.e. that from joint network 930]), a speech recognition hypothesis representing a candidate transcription for the corresponding training utterance ([0090] The ASR model 950 may predict a probability (y|x) of labels y=(y.sub.1, . . . , y.sub.u) given acoustic features x=(x.sub.1, . . . , x.sub.t). During inference, the ASR model 950 can generate an N-best list using, for example, a beam search decoding algorithm, [0093] The ASR data 155 may include text, subword tokens, word tokens, and/or other character data representing a possible transcript of speech represented in the spectrogram data 125).
Yang in view of Zheng, further in view of Wang does not disclose:
determining an ASR loss based on the speech recognition hypothesis and the corresponding input text sequence,
wherein the TTS loss comprises the ASR loss.
Thomas discloses:
determining an ASR loss based on the speech recognition hypothesis and the corresponding input text sequence ([0035] ii) aligning, at a token level, one or more LLM based sentence embeddings with the one or more speech-based embeddings; iii) combining an alignment loss and an ASR loss),
wherein the TTS loss comprises the ASR loss ([0041] The cross-attention component 104 can output a corresponding vector having a one-to-one correspondence between the output of the large langue model LLM 208 and the speech encoder 110. Additionally, the distances between the two sequences (e.g., the speech-based embeddings 204 and the LLM based sentence embeddings 210) can be minimized by using a loss function, [Minimizing loss of sentence embeddings indicates a TTS loss based on a speech-based, i.e. ASR, embedding]).
Yang, Zheng, Wang, and Thomas are considered analogous art within speech recognition. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Yang in view of Zheng, further in view of Wang to incorporate the teachings of Thomas, because of the novel way to explicitly incorporate knowledge from an LLM acquired during ASR pretraining through token-wise knowledge transfer based on sentence embeddings to a deployment operation, reducing the amount of data required for training/pretraining for improved ASR performance (Thomas, [0008]).
Regarding claim 4, Yang in view of Zheng, further in view of Wang, further in view of Thomas discloses: the computer-implemented method of claim 3.
Thomas further discloses:
wherein the ASR decoder comprises a recurrent neural network-transducer (RNN-T) architecture ([0042] the ASR system 112(e.g., RNN-T)).
Regarding claim 16, Yang in view of Zheng, further in view of Wang discloses: the system of claim 14.
Wang further discloses:
wherein the operations further comprise, for each training utterance in each set of the training utterances of the received training data:
generating, using an automatic speech recognition (ASR) decoder configured to receive the shared encoder output as input ([Fig. 1B, ASR 150 receiving Spectrogram Data 125], [Fig. 9, Speech Recognition Engine 958], [The speech recognition engine is the ASR decoder receiving shared encoder output, i.e. that from joint network 930]), a speech recognition hypothesis representing a candidate transcription for the corresponding training utterance ([0090] The ASR model 950 may predict a probability (y|x) of labels y=(y.sub.1, . . . , y.sub.u) given acoustic features x=(x.sub.1, . . . , x.sub.t). During inference, the ASR model 950 can generate an N-best list using, for example, a beam search decoding algorithm, [0093] The ASR data 155 may include text, subword tokens, word tokens, and/or other character data representing a possible transcript of speech represented in the spectrogram data 125).
Yang in view of Zheng, further in view of Wang does not disclose:
determining an ASR loss based on the speech recognition hypothesis and the corresponding input text sequence,
wherein the TTS loss comprises the ASR loss.
Thomas discloses:
determining an ASR loss based on the speech recognition hypothesis and the corresponding input text sequence ([0035] ii) aligning, at a token level, one or more LLM based sentence embeddings with the one or more speech-based embeddings; iii) combining an alignment loss and an ASR loss),
wherein the TTS loss comprises the ASR loss ([0041] The cross-attention component 104 can output a corresponding vector having a one-to-one correspondence between the output of the large langue model LLM 208 and the speech encoder 110. Additionally, the distances between the two sequences (e.g., the speech-based embeddings 204 and the LLM based sentence embeddings 210) can be minimized by using a loss function, [Minimizing loss of sentence embeddings indicates a TTS loss based on a speech-based, i.e. ASR, embedding]).
Yang, Zheng, Wang, and Thomas are considered analogous art within speech recognition. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Yang in view of Zheng, further in view of Wang to incorporate the teachings of Thomas, because of the novel way to explicitly incorporate knowledge from an LLM acquired during ASR pretraining through token-wise knowledge transfer based on sentence embeddings to a deployment operation, reducing the amount of data required for training/pretraining for improved ASR performance (Thomas, [0008]).
Regarding claim 17, Yang in view of Zheng, further in view of Wang, further in view of Thomas discloses: the system of claim 16.
Thomas further discloses:
wherein the ASR decoder comprises a recurrent neural network-transducer (RNN-T) architecture ([0042] the ASR system 112(e.g., RNN-T)).
Claim(s) 7-11, 20-24 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yang in view of Zheng, further in view of Wang, further in view of Chen et al. (“MAESTRO: Matched Speech Text Representations through Modality Matching”), hereinafter Chen.
Regarding claim 7, Yang in view of Zheng, further in view of Wang discloses: the computer-implemented method of claim 1.
Yang in view of Zheng, further in view of Wang does not disclose:
wherein:
the training data further comprises unspoken textual utterances associated with a respective plurality of different languages, each unspoken textual utterance not paired with any corresponding spoken utterance; and the operations further comprise, for each unspoken textual utterance:
generating, using the text encoder, a corresponding unspoken encoded textual representation for the corresponding unspoken textual utterance; and
determining an aligned-text masked language modeling (MLM) loss for the corresponding unspoken encoded textual representation generated for the corresponding unspoken textual utterance,
wherein the TTS loss comprises the aligned-text MLM loss.
Chen discloses:
wherein:
the training data further comprises unspoken textual utterances associated with a respective plurality of different languages, each unspoken textual utterance not paired with any corresponding spoken utterance ([Fig. 1, Unspoken Text], [In view of the plurality of languages of Yang]); and
the operations further comprise, for each unspoken textual utterance:
generating, using the text encoder, a corresponding unspoken encoded textual representation for the corresponding unspoken textual utterance ([Fig. 1, Text Encoder receiving a text sequence comprising at least unspoken text]); and
determining an aligned-text masked language modeling (MLM) loss for the corresponding unspoken encoded textual representation generated for the corresponding unspoken textual utterance ([Section 3.2.3] We replace the original MLM/BERT loss used in prior work [1,11] with the aligned masked language model training objective (LA-MLM). This is the RNN-T loss applied over the masked, resampled text embeddings with masking in frequency and time domain similar to SpecAugment [28]. This new objective allows for the use of the same RNN-T objective on speech embedding or unspoken text with no associated speech embedding),
wherein the TTS loss comprises the aligned-text MLM loss ([Wherein the context of Chen is clearly within TTS (see Multilingual ASR pretraining data, Section 4.1), indicating the LA-MLM to be comprising the TTS loss]).
Yang, Zheng, Wang, and Chen are considered analogous art within multilingual speech recognition. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Yang in view of Zheng, further in view of Wang to incorporate the teachings of Chen, because of the novel way to unify text and speech representations of input that can transfer to complex downstream tasks such as ASR, improving multi-domain/language ASR tasks (Chen, Abstract).
Regarding claim 8, Yang in view of Zheng, further in view of Wang, further in view of Chen discloses: the computer implemented method of claim 7.
Yang further discloses:
wherein: each unspoken textual utterance is paired with a corresponding language identifier label ([Fig. 8, Text Input 702 joined with Language Embedding Vector 744 in Acoustic Feature Predictor 710], [0054] an acoustic feature predictor 710 may generate at least one acoustic feature 704 based at least on the text input 702, [0055] the acoustic feature predictor 710 may also use a speaker embedding vector of a target speaker and a language embedding vector of a reference language as global conditions, [Sending this information into the acoustic feature predictor to result in acoustic feature output indicates a pairing of text with language identifier to generate the acoustic features. See concatenation of Fig. 9]).
the operations further comprise, for each unspoken textual utterance:
generating, using a language identifier configured to receive the corresponding unspoken encoded textual representation generated for the corresponding unspoken textual utterance as input ([Fig. 5, Language Encoder 500], [0051] The language embedding vector database 512 may be established through collecting language embedding vectors of those languages in a multilingual corpus during the training of the multilingual neural TTS system, [Generating language embedding vectors based on a corpus used for TTS training (indicating unspoken text to be synthesized)]), a predicted language identifier ([Fig. 5, Language Embedding Vector]); and,
determining a text language identifier loss based on the predicted language identifier and the language identifier label ([Fig. 12, Discriminator 1220], [0088] The acoustic feature predictor 1210 may learn to predict or generate an acoustic feature 1214 based on a text input 1202 and global conditions 1212, so that the predicted acoustic feature may best approximate an acoustic feature in the training data, i.e., a ground-truth acoustic feature. The global conditions 1212 may comprise a speaker embedding vector of a target speaker and/or a language embedding vector of a reference language. The discriminator 1220 may learn to distinguish between the predicted acoustic feature 1214 output by the acoustic feature predictor 1210 and the ground-truth acoustic feature 1222 in the training data, and output a discrimination result 1220, e.g., true or false, [Wherein a reference language is a global condition as previously disclosed, indicating the predicted acoustic features to be in a language which is compared to a ground-truth feature in another language which may be the same and/or different]),
wherein the TTS loss comprises the text language identifier loss ([0088] This discrimination result may be further used for updating or improving the acoustic feature predictor 1210 and the discriminator 1220, [Wherein the figure associated with this operation is Fig. 12, defined to be a “training an acoustic feature predictor”, (1116) which is within the Multilingual TTS system 1110]).
Regarding claim 9, Yang in view of Zheng, further in view of Wang discloses: the computer-implemented method of claim 1.
Yang further discloses:
wherein: the training data further comprises unpaired spoken utterances spoken in a respective plurality of different languages ([Fig. 11, Multilingual Corpus 1120]), each unpaired spoken utterance not paired with any corresponding text ([0086] speech waveforms in the multilingual corpus 1120… various derived information may be obtained from the speech waveforms, e.g., text information obtained by applying any speech recognition technologies, [Optionally deriving text information from speech indicates that there is no required pairing of speech to text in Yang]).
the operations further comprise, for each unpaired spoken utterance:
generating, using the speech encoder ([in view of the previously disclosed speaker encoder]), a corresponding unpaired speech encoding for the corresponding unpaired spoken utterance ([0041] The speaker encoder 230 may provide speaker latent space information 232 of a target speaker, [0049] the speaker embedding generator 400 may be based on a neural network for generating a speaker embedding vector 404 based on an acoustic feature 402, [Wherein the acoustic features are gathered from the multilingual training corpus 1120 as required for the training operation of the speaker encoder 1112]).
Yang in view of Zheng, further in view of Wang does not disclose:
determining an aligned-speech masked language modeling (MLM) loss for the corresponding unpaired speech encoding generated for the corresponding unpaired spoken utterance,
wherein the TTS loss comprises the aligned-speech MLM loss.
Chen discloses:
determining an aligned-speech masked language modeling (MLM) loss for the corresponding unpaired speech encoding generated for the corresponding unpaired spoken utterance ([Section 3.2.3] We replace the original MLM/BERT loss used in prior work [1,11] with the aligned masked language model training objective (LA-MLM). This is the RNN-T loss applied over the masked, resampled text embeddings with masking in frequency and time domain similar to SpecAugment [28]. This new objective allows for the use of the same RNN-T objective on speech embedding or unspoken text with no associated speech embedding, [See “untranscribed speech” input of Fig. 1]),
wherein the TTS loss comprises the aligned-speech MLM loss ([Wherein the context of Chen is clearly within TTS (see Multilingual ASR pretraining data, Section 4.1), indicating the LA-MLM to be comprising the TTS loss]).
Yang, Zheng, Wang, and Chen are considered analogous art within multilingual speech recognition. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Yang in view of Zheng, further in view of Wang to incorporate the teachings of Chen, because of the novel way to unify text and speech representations of input that can transfer to complex downstream tasks such as ASR, improving multi-domain/language ASR tasks (Chen, Abstract).
Regarding claim 10, Yang in view of Zheng, further in view of Wang, further in view of Chen discloses: the computer-implemented method of claim 9.
Yang further discloses:
wherein the operations further comprise, for each unpaired spoken utterance:
generating, using the shared encoder further configured to receive the corresponding unpaired speech encoding, an unpaired shared encoder output ([Fig. 4, Speaker Embedding Vector 404]).
Wang further discloses:
generating, using an automatic speech recognition (ASR) decoder configured to receive the unpaired shared encoder output as input ([Fig. 1, ASR 150 receiving Spectrogram Data 125], [Fig. 9, Speech Recognition Engine 958], [The speech recognition engine is the ASR decoder receiving shared encoder output, i.e. that from joint network 930]), a pseudolabel representing a candidate transcription for the corresponding unpaired spoken utterance ([0090] The ASR model 950 may predict a probability (y|x) of labels y=(y.sub.1, . . . , y.sub.u) given acoustic features x=(x.sub.1, . . . , x.sub.t). During inference, the ASR model 950 can generate an N-best list using, for example, a beam search decoding algorithm, [0093] The ASR data 155 may include text, subword tokens, word tokens, and/or other character data representing a possible transcript of speech represented in the spectrogram data 125, [Generating labels for acoustic features, i.e. shared encoder output, indicates each label to be a pseudolabel representing a candidate transcription, i.e. the label itself, for a corresponding spoken utterance, wherein using labels to generate ASR data comprising transcriptions indicates the labels to be representing transcription candidates]),
wherein the training data further comprises unspoken textual utterances comprising the pseudolabels ([Fig. 4, Transcript Embedding 175 based on Transcript Data 165 for Training 480], [Wherein Fig. 4 is defined to be an example training operation for the expressive speech generator. The examiner asserts that a transcript embedding will be consisting of pseudolabels, i.e. embedding tags/labels]).
Regarding claim 11, Yang in view of Zheng, further in view of Wang, further in view of Chen discloses: the computer-implemented method of claim 9.
Yang further discloses:
wherein:
each unpaired spoken utterance is paired with a corresponding language identifier label ([Fig. 11, Multilingual Corpus 1120 consisting of Languages 1-N and Corpus 1-N], [The Corpus comprises spoken utterances which are paired with language identifiers]);
the operations further comprise, for each unpaired spoken utterance:
generating, using a language identifier configured to receive the corresponding unpaired speech encoding for the corresponding unpaired spoken utterance as input ([Fig. 5, Language Encoder 500], [0051] The language embedding vector database 512 may be established through collecting language embedding vectors of those languages in a multilingual corpus during the training of the multilingual neural TTS system, [0052] the corpus 504 of the reference language may be obtained, which comprises multiple speech waveforms in the reference language, [Generating language embedding vectors based on a corpus comprised of speech waveforms indicates the language identifier to be received only speech, i.e. unpaired with text]), a predicted language identifier ([Fig. 5, Language Embedding Vector]); and
determining a speech language identifier loss based on the predicted language identifier and the language identifier label ([Fig. 12, Discriminator 1220], [0088] The acoustic feature predictor 1210 may learn to predict or generate an acoustic feature 1214 based on a text input 1202 and global conditions 1212, so that the predicted acoustic feature may best approximate an acoustic feature in the training data, i.e., a ground-truth acoustic feature. The global conditions 1212 may comprise a speaker embedding vector of a target speaker and/or a language embedding vector of a reference language. The discriminator 1220 may learn to distinguish between the predicted acoustic feature 1214 output by the acoustic feature predictor 1210 and the ground-truth acoustic feature 1222 in the training data, and output a discrimination result 1220, e.g., true or false, [Wherein a reference language is a global condition as previously disclosed, indicating the predicted acoustic features to be in a language which is compared to a ground-truth feature in another language which may be the same and/or different]),
wherein the TTS loss comprises the speech language identifier loss ([0088] This discrimination result may be further used for updating or improving the acoustic feature predictor 1210 and the discriminator 1220, [Wherein the figure associated with this operation is Fig. 12, defined to be a “training an acoustic feature predictor”, (1116) which is within the Multilingual TTS system 1110]).
Regarding claim 20, Yang in view of Zheng, further in view of Wang discloses: the system of claim 14.
Yang in view of Zheng, further in view of Wang does not disclose:
wherein:
the training data further comprises unspoken textual utterances associated with a respective plurality of different languages, each unspoken textual utterance not paired with any corresponding spoken utterance; and the operations further comprise, for each unspoken textual utterance:
generating, using the text encoder, a corresponding unspoken encoded textual representation for the corresponding unspoken textual utterance; and
determining an aligned-text masked language modeling (MLM) loss for the corresponding unspoken encoded textual representation generated for the corresponding unspoken textual utterance,
wherein the TTS loss comprises the aligned-text MLM loss.
Chen discloses:
wherein:
the training data further comprises unspoken textual utterances associated with a respective plurality of different languages, each unspoken textual utterance not paired with any corresponding spoken utterance ([Fig. 1, Unspoken Text], [In view of the plurality of languages of Yang]); and
the operations further comprise, for each unspoken textual utterance:
generating, using the text encoder, a corresponding unspoken encoded textual representation for the corresponding unspoken textual utterance ([Fig. 1, Text Encoder receiving a text sequence comprising at least unspoken text]); and
determining an aligned-text masked language modeling (MLM) loss for the corresponding unspoken encoded textual representation generated for the corresponding unspoken textual utterance ([Section 3.2.3] We replace the original MLM/BERT loss used in prior work [1,11] with the aligned masked language model training objective (LA-MLM). This is the RNN-T loss applied over the masked, resampled text embeddings with masking in frequency and time domain similar to SpecAugment [28]. This new objective allows for the use of the same RNN-T objective on speech embedding or unspoken text with no associated speech embedding),
wherein the TTS loss comprises the aligned-text MLM loss ([Wherein the context of Chen is clearly within TTS (see Multilingual ASR pretraining data, Section 4.1), indicating the LA-MLM to be comprising the TTS loss]).
Yang, Zheng, Wang, and Chen are considered analogous art within multilingual speech recognition. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Yang in view of Zheng, further in view of Wang to incorporate the teachings of Chen, because of the novel way to unify text and speech representations of input that can transfer to complex downstream tasks such as ASR, improving multi-domain/language ASR tasks (Chen, Abstract).
Regarding claim 21, Yang in view of Zheng, further in view of Wang, further in view of Chen discloses: the system of claim 20.
Yang further discloses:
wherein: each unspoken textual utterance is paired with a corresponding language identifier label ([Fig. 8, Text Input 702 joined with Language Embedding Vector 744 in Acoustic Feature Predictor 710], [0054] an acoustic feature predictor 710 may generate at least one acoustic feature 704 based at least on the text input 702, [0055] the acoustic feature predictor 710 may also use a speaker embedding vector of a target speaker and a language embedding vector of a reference language as global conditions, [Sending this information into the acoustic feature predictor to result in acoustic feature output indicates a pairing of text with language identifier to generate the acoustic features. See concatenation of Fig. 9]).
the operations further comprise, for each unspoken textual utterance:
generating, using a language identifier configured to receive the corresponding unspoken encoded textual representation generated for the corresponding unspoken textual utterance as input ([Fig. 5, Language Encoder 500], [0051] The language embedding vector database 512 may be established through collecting language embedding vectors of those languages in a multilingual corpus during the training of the multilingual neural TTS system, [Generating language embedding vectors based on a corpus used for TTS training (indicating unspoken text to be synthesized)]), a predicted language identifier ([Fig. 5, Language Embedding Vector]); and,
determining a text language identifier loss based on the predicted language identifier and the language identifier label ([Fig. 12, Discriminator 1220], [0088] The acoustic feature predictor 1210 may learn to predict or generate an acoustic feature 1214 based on a text input 1202 and global conditions 1212, so that the predicted acoustic feature may best approximate an acoustic feature in the training data, i.e., a ground-truth acoustic feature. The global conditions 1212 may comprise a speaker embedding vector of a target speaker and/or a language embedding vector of a reference language. The discriminator 1220 may learn to distinguish between the predicted acoustic feature 1214 output by the acoustic feature predictor 1210 and the ground-truth acoustic feature 1222 in the training data, and output a discrimination result 1220, e.g., true or false, [Wherein a reference language is a global condition as previously disclosed, indicating the predicted acoustic features to be in a language which is compared to a ground-truth feature in another language which may be the same and/or different]),
wherein the TTS loss comprises the text language identifier loss ([0088] This discrimination result may be further used for updating or improving the acoustic feature predictor 1210 and the discriminator 1220, [Wherein the figure associated with this operation is Fig. 12, defined to be a “training an acoustic feature predictor”, (1116) which is within the Multilingual TTS system 1110]).
Regarding claim 22, Yang in view of Zheng, further in view of Wang discloses: the system of claim 14.
Yang further discloses:
wherein: the training data further comprises unpaired spoken utterances spoken in a respective plurality of different languages ([Fig. 11, Multilingual Corpus 1120]), each unpaired spoken utterance not paired with any corresponding text ([0086] speech waveforms in the multilingual corpus 1120… various derived information may be obtained from the speech waveforms, e.g., text information obtained by applying any speech recognition technologies, [Optionally deriving text information from speech indicates that there is no required pairing of speech to text in Yang]).
the operations further comprise, for each unpaired spoken utterance:
generating, using the speech encoder ([in view of the previously disclosed speaker encoder]), a corresponding unpaired speech encoding for the corresponding unpaired spoken utterance ([0041] The speaker encoder 230 may provide speaker latent space information 232 of a target speaker, [0049] the speaker embedding generator 400 may be based on a neural network for generating a speaker embedding vector 404 based on an acoustic feature 402, [Wherein the acoustic features are gathered from the multilingual training corpus 1120 as required for the training operation of the speaker encoder 1112]).
Yang in view of Zheng, further in view of Wang does not disclose:
determining an aligned-speech masked language modeling (MLM) loss for the corresponding unpaired speech encoding generated for the corresponding unpaired spoken utterance,
wherein the TTS loss comprises the aligned-speech MLM loss.
Chen discloses:
determining an aligned-speech masked language modeling (MLM) loss for the corresponding unpaired speech encoding generated for the corresponding unpaired spoken utterance ([Section 3.2.3] We replace the original MLM/BERT loss used in prior work [1,11] with the aligned masked language model training objective (LA-MLM). This is the RNN-T loss applied over the masked, resampled text embeddings with masking in frequency and time domain similar to SpecAugment [28]. This new objective allows for the use of the same RNN-T objective on speech embedding or unspoken text with no associated speech embedding, [See “untranscribed speech” input of Fig. 1]),
wherein the TTS loss comprises the aligned-speech MLM loss ([Wherein the context of Chen is clearly within TTS (see Multilingual ASR pretraining data, Section 4.1), indicating the LA-MLM to be comprising the TTS loss]).
Yang, Zheng, Wang, and Chen are considered analogous art within multilingual speech recognition. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Yang in view of Zheng, Wang to incorporate the teachings of Chen, because of the novel way to unify text and speech representations of input that can transfer to complex downstream tasks such as ASR, improving multi-domain/language ASR tasks (Chen, Abstract).
Regarding claim 23, Yang in view of Zheng, further in view of Wang, further in view of Chen discloses: the system of claim 22.
Yang further discloses:
wherein the operations further comprise, for each unpaired spoken utterance:
generating, using the shared encoder further configured to receive the corresponding unpaired speech encoding, an unpaired shared encoder output ([Fig. 4, Speaker Embedding Vector 404]).
Wang further discloses:
generating, using an automatic speech recognition (ASR) decoder configured to receive the unpaired shared encoder output as input ([Fig. 1, ASR 150 receiving Spectrogram Data 125], [Fig. 9, Speech Recognition Engine 958], [The speech recognition engine is the ASR decoder receiving shared encoder output, i.e. that from joint network 930]), a pseudolabel representing a candidate transcription for the corresponding unpaired spoken utterance ([0090] The ASR model 950 may predict a probability (y|x) of labels y=(y.sub.1, . . . , y.sub.u) given acoustic features x=(x.sub.1, . . . , x.sub.t). During inference, the ASR model 950 can generate an N-best list using, for example, a beam search decoding algorithm, [0093] The ASR data 155 may include text, subword tokens, word tokens, and/or other character data representing a possible transcript of speech represented in the spectrogram data 125, [Generating labels for acoustic features, i.e. shared encoder output, indicates each label to be a pseudolabel representing a candidate transcription, i.e. the label itself, for a corresponding spoken utterance, wherein using labels to generate ASR data comprising transcriptions indicates the labels to be representing transcription candidates]),
wherein the training data further comprises unspoken textual utterances comprising the pseudolabels ([Fig. 4, Transcript Embedding 175 based on Transcript Data 165 for Training 480], [Wherein Fig. 4 is defined to be an example training operation for the expressive speech generator. The examiner asserts that a transcript embedding will be consisting of pseudolabels, i.e. embedding tags/labels]).
Regarding claim 24, Yang in view of Zheng, further in view of Wang, further in view of Chen discloses: the system of claim 22.
Yang further discloses:
wherein:
each unpaired spoken utterance is paired with a corresponding language identifier label ([Fig. 11, Multilingual Corpus 1120 consisting of Languages 1-N and Corpus 1-N], [The Corpus comprises spoken utterances which are paired with language identifiers]);
the operations further comprise, for each unpaired spoken utterance:
generating, using a language identifier configured to receive the corresponding unpaired speech encoding for the corresponding unpaired spoken utterance as input ([Fig. 5, Language Encoder 500], [0051] The language embedding vector database 512 may be established through collecting language embedding vectors of those languages in a multilingual corpus during the training of the multilingual neural TTS system, [0052] the corpus 504 of the reference language may be obtained, which comprises multiple speech waveforms in the reference language, [Generating language embedding vectors based on a corpus comprised of speech waveforms indicates the language identifier to be received only speech, i.e. unpaired with text]), a predicted language identifier ([Fig. 5, Language Embedding Vector]); and
determining a speech language identifier loss based on the predicted language identifier and the language identifier label ([Fig. 12, Discriminator 1220], [0088] The acoustic feature predictor 1210 may learn to predict or generate an acoustic feature 1214 based on a text input 1202 and global conditions 1212, so that the predicted acoustic feature may best approximate an acoustic feature in the training data, i.e., a ground-truth acoustic feature. The global conditions 1212 may comprise a speaker embedding vector of a target speaker and/or a language embedding vector of a reference language. The discriminator 1220 may learn to distinguish between the predicted acoustic feature 1214 output by the acoustic feature predictor 1210 and the ground-truth acoustic feature 1222 in the training data, and output a discrimination result 1220, e.g., true or false, [Wherein a reference language is a global condition as previously disclosed, indicating the predicted acoustic features to be in a language which is compared to a ground-truth feature in another language which may be the same and/or different]),
wherein the TTS loss comprises the speech language identifier loss ([0088] This discrimination result may be further used for updating or improving the acoustic feature predictor 1210 and the discriminator 1220, [Wherein the figure associated with this operation is Fig. 12, defined to be a “training an acoustic feature predictor”, (1116) which is within the Multilingual TTS system 1110]).
Claim(s) 13, 26 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yang in view of Zheng, further in view of Wang, further in view of Maiti (“Speech Enhancement Using Speech Synthesis Techniques”).
Regarding claim 13, Yang in view of Zheng, further in view of Wang discloses: the computer-implemented method of claim 1.
Yang in view of Zheng, further in view of Wang does not disclose:
wherein generating the speech encoding for the corresponding reference speech representation comprises:
applying random projections to project the corresponding utterance using a random-projection quantizer; and
mapping the corresponding projected utterance to discrete labels.
Maiti discloses:
wherein generating the speech encoding for the corresponding reference speech representation comprises:
applying random projections to project the corresponding utterance using a random-projection quantizer ([Function 4.3], [pg. 55] The audio samples are quantized with m-law quantization with 256 possible values); and
mapping the corresponding projected utterance to discrete labels ([pg. 60] The waveform is quantized into discrete values similar to how it is in WaveNet, and the quantized values are used as a categorical output).
Yang, Zheng, Wang, and Maiti are considered analogous art within speech synthesis. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Yang in view of Zheng, further in view of Wang to incorporate the teachings of Maiti, because of the novel way to use non-linear quantization for audio encoding instead of traditional linear quantization, increasing the tractability of the system and quality of synthesized audio (Maiti, [pg. 56]).
Regarding claim 26, Yang in view of Zheng, further in view of Wang discloses: the system of claim 14.
Yang in view of Zheng, further in view of Wang does not disclose:
wherein generating the speech encoding for the corresponding reference speech representation comprises:
applying random projections to project the corresponding utterance using a random-projection quantizer; and
mapping the corresponding projected utterance to discrete labels.
Maiti discloses:
wherein generating the speech encoding for the corresponding reference speech representation comprises:
applying random projections to project the corresponding utterance using a random-projection quantizer ([Function 4.3], [pg. 55] The audio samples are quantized with m-law quantization with 256 possible values); and
mapping the corresponding projected utterance to discrete labels ([pg. 60] The waveform is quantized into discrete values similar to how it is in WaveNet, and the quantized values are used as a categorical output).
Yang, Zheng, Wang, and Maiti are considered analogous art within speech synthesis. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Yang in view of Zheng, further in view of Wang to incorporate the teachings of Maiti, because of the novel way to use non-linear quantization for audio encoding instead of traditional linear quantization, increasing the tractability of the system and quality of synthesized audio (Maiti, [pg. 56]).
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Kuo et al. (US-20210312906-A1) discloses “An illustrative embodiment includes a method for training an end-to-end (E2E) spoken language understanding (SLU) system. The method includes receiving a training corpus comprising a set of text classified using one or more sets of semantic labels but unpaired with speech and using the set of unpaired text to train the E2E SLU system to classify speech using at least one of the one or more sets of semantic labels. The method may include training a text-to-intent model using the set of unpaired text; and training a speech-to-intent model using the text-to-intent model. Alternatively or additionally, the method may include using a text-to-speech (TTS) system to generate synthetic speech from the unpaired text; and training the E2E SLU system using the synthetic speech.” (abstract). See entire document.
Sundararaman et al. (“PhonemeBERT: Joint Language Modelling of Phoneme Sequence and ASR Transcript”) discloses “a BERT-style language model, referred to as PhonemeBERT that learns a joint language model with phoneme sequence and ASR transcript to learn phonetic-aware representations that are robust to ASR errors. We show that PhonemeBERT leverages phoneme sequences as additional features that outperform word-only models on downstream tasks.” (abstract). See entire document.
Gonzales et al. (“Joint Speech-Text Embeddings with Disentangled Speaker Features”) discloses “a novel model architecture for speech processing that takes advantage of a joint speech-text embedding space and disentangled speaker features. Here unsupervised representation learning extracts latent features from the input without labels which results in task-agnostic, but information-entangled embeddings. On the other hand, a unified embedding space of speech and text aims to leverage acoustics and semantic knowledge from the two modalities, respectively.” (abstract). See entire document.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to THEODORE JOHN WITHEY whose telephone number is (703)756-1754. The examiner can normally be reached Monday - Friday, 8am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Flanders can be reached at (571) 272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/THEODORE WITHEY/Examiner, Art Unit 2655
/JESSE S PULLIAS/Primary Examiner, Art Unit 2655 08/14/26