DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1 is/are rejected under 35 U.S.C. 103 as being unpatentable over Sharma et al. (PGPUB 2022/0013105), hereinafter referenced as Sharma in view of Lakhotia et al. (On generative spoken language modeling from raw audio), hereinafter referenced as Lakhotia and in further view of Strelec et al. (Discrete acoustic space for an efficient sampling in neural text-to-speech), hereinafter referenced as Strong.
Regarding claim 1, Sharma discloses a computer-implemented method for generating a prediction of an audio signal, the method comprising:
receiving a request to generate an audio signal having a respective audio sample at each of a plurality of output time steps spanning a time window conditioned on an input (receiving a request to generate speech based on a text input and generating using a speech synthesis system, synthesized playback audio defining an audio signal corresponding to the text input and generating raw audio waveform as respective audio samples sequentially over a plurality of output time steps; p. 0015-0016), but does not specifically teach processing the input using an embedding neural network to map the input to one or more embedding tokens; generating a semantic representation of the audio signal that specifies a respective semantic token at each of a plurality of first time steps spanning the time window, each semantic token being selected from a vocabulary of semantic tokens conditioned on the embedding tokens and representing semantic content of the audio signal at the corresponding first time step; generating, using one or more generative neural networks and conditioned on at least the semantic representation and the embedding tokens, an acoustic representation of the audio signal, the acoustic representation specifying a set of one or more respective acoustic tokens at each of a plurality of second time steps spanning the time window, the one or more respective acoustic tokens at each second time step representing acoustic properties of the audio signal at the corresponding second time step; and processing at least the acoustic representation using a decoder neural network to generate the prediction of the audio signal.
Lakhotia discloses a method comprising generating a semantic representation of the audio signal that specifies a respective semantic token at each of a plurality of first time steps spanning the time window, each semantic token being selected from a vocabulary of semantic tokens and representing semantic content of the audio signal at the corresponding first time step (discrete text represent linguistic characteristics of speech and the generative language model operates on those discrete units wherein the discrete units are intended to abstract higher-level language information from fine-grained acoustic details; p. 1337); and
processing at least the acoustic representation using a decoder neural network to generate the prediction of the audio signal (speech decoder takes pseudo text units, outputs log Mel-spectrogram and then generates the time domain waveform; section 2), to represent linguistic information using discrete units
Therefore, it would have been obvious to one of ordinary skill of the art, before the effective filing date of the claimed invention, to modify the method as described above, to enable generation of speech based on sequences if such discrete linguistic units.
Strelec teaches a method comprising:
processing the input using an embedding neural network to map the input to one or more embedding tokens (embedding network operates in an audio embedding space; Introduction);
generating, using one or more generative neural networks and conditioned on at least the semantic representation and the embedding tokens, an acoustic representation of the audio signal, the acoustic representation specifying a set of one or more respective acoustic tokens at each of a plurality of second time steps spanning the time window, the one or more respective acoustic tokens at each second time step representing acoustic properties of the audio signal at the corresponding second time step (neural TTS architecture employing a vector quantized autoencoder having discretized latent acoustic space and vector quantization produces discrete representations selected from a finite quantized acoustic space; abstract, introduction and section 2.3), to assist with predicting for neural TTS synthesis.
Therefore, it would have been obvious to one of ordinary skill of the art, before the effective filing date of the claimed invention, to modify the method as described above, to improve representation of acoustic variability while permitting efficient prediction during speech synthesis.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. This information has been detailed in the PTO 892 attached (Notice of References Cited).
Zeghidour et al. teaches a fully convolutional encoder/decoder neural network with a residual vector quantizer and states that the system generates high-quality audio from the quantized embeddings.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JAKIEDA R JACKSON whose telephone number is (571)272-7619. The examiner can normally be reached Mon - Fri 6:30a-2:30p.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at 571.272.5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JAKIEDA R JACKSON/Primary Examiner, Art Unit 2657