DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant's arguments with respect to 35 U.S.C. 103 in regards to claim 1 has been considered but are moot due to new grounds of rejection necessitated by amendments. See detailed rejection below.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 3-4 and 6-7 is/are rejected under 35 U.S.C. 103 as being unpatentable over Qian et al. (US 2010/0066742) in view of Wei et al. (“Neural Network-Based Modeling of Phonetic Durations”; Sept. 6, 2019).
Claim 1,
Qian teaches a computer-implemented method, comprising ([0020] Qian teaches an HMM-based computer speech synthesis system having a training phase and synthesis phase. In synthesis, “input text is converted first into a sequence of contextual labels,” contextual HMMs are retrieved, and parameter generation is performed):
determining, from at least respective phoneme pitches or respective phoneme energies of the plurality of first audio segments, at least a pitch distribution, corresponding to the respective phoneme pitches, or an energy distribution, corresponding to the respective phoneme energies ([0019] Qian teaches probabilistic pitch modeling from training speech features: “An HMM training mechanism 230 inputs the log F0, LSP and Gain..” and the stream-dependent models “cluster the spectral, prosodic and duration features.” Qian further teaches that pitch/F0 is probabilistically modeled: “F0 are modeled by multi-space probability distribution HMM.”);
determining, for an input received at a text-to-speech model trained, as least in part, using the training data, a duration alignment between a sequence of text associated with the input and a total speech duration using the respective phoneme durations ([0020] [0023] Qian teaches converting input text to contextual labels and obtaining durations from a duration model: “input text is converted first into a sequence of contextual labels … and the duration of each state is obtained from a duration model.” Qian also teaches aligning/adjusting durations to a total duration: “T is the duration as modifiable by the user” and “Each state duration d(k) may be adjusted…”);
determining a speech alignment for a second audio segment, generated by the text-to- speech model, based at least on the duration alignment, the phoneme distribution, and at least one of the pitch distribution or the energy distribution ([0019-0020] [0023] Qian teaches generating duration-aligned speech trajectories after durations are obtained from the duration model: “The LSP, gain and F0 trajectories are generated …” These trajectories are based on Qian’s trained duration/prosody models, include “duration density” parameters and F0 probability modeling); and
generating a synthesized audio recitation, as an output audio signal, corresponding to the second audio segment ([0020] Qian teaches output speech synthesis: “A speech waveform is synthesized from the generated spectral and excitation parameters…”).
The difference between the prior art and the claimed invention is that Qian does not explicitly teach determining a probabilistically sampled phoneme distribution from respective phoneme durations of a plurality of first audio segments provided as training data.
Wei teaches determining a probabilistically sampled phoneme distribution from respective phoneme durations of a plurality of first audio segments provided as training data ([2.3] [2.4] [3.] [3.1] Wei teaches phoneme/phone duration distributions from forced-alignment-derived durations: “We obtain the reference durations from forced-alignment. We group them into 45 bins…”; Wei further teaches probabilistic duration modeling: “We use the value of the bin to which the reference duration belongs as the probability of the duration…”; Wei also evaluates “the model’s predicted duration distribution” and uses datasets of recorded utterances “recorded for TTS purposes,” with “forced alignment to get the duration for each phone.”).
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the teachings of Qian with teachings of Wei by modifying the stylized prosody for speech synthesis-based applications as taught by Qian to include determining a probabilistically sampled phoneme distribution from respective phoneme durations of a plurality of first audio segments provided as training data as taught by Wei for the benefit of indicating discrepancies between a transcription or script and what was actually spoken (Wei [Introduction]).
Claim 3,
Wei further teaches the computer-implemented method of claim 1, further comprising: applying a prior distribution to the duration alignment configured to exclude pairs of phonemes and durations from the plurality of first audio segments that are outside of a specified range ([2.3] [2.4] [5.] Wei teaches specified duration ranges/bins: “group them into 45 bins,” starting at “30 ms,” and “Duration larger than 670ms are all put into the 45 bins.”; Wei applies the predicted duration distribution to identify outliers: “ranking the probabilities” yields phonemes with “unlikely duration,” regarded as “outliers.”; Wei further teaches exclusion: anomalies in TTS training material can be used to “to exclude unsuitable speech from the TTS training set.”).
Claim 4,
Qian further teaches the computer-implemented method of claim 3, wherein the prior distribution is cigar-shaped ([0019] multi-space probability distribution HMM).
Claim 6,
Qian further teaches the computer-implemented method of claim 3, wherein the prior distribution is constructed from a beta-binomial distribution ([0019] multi-space probability distribution HMM).
Claim 7,
Qian further teaches the computer-implemented method of claim 1, wherein the synthesized audio recitation is generative such that a first synthesized recitation is different from a second synthesized recitation, each of the first synthesized recitation and the second synthesized recitation based on the sequence of text ([Abstract] [0005] [0025-0028] Qian teaches that synthesized speech is displayed with “corresponding text from which the speech was synthesized” and that the user may change “duration, pitch and/or loudness data” with respect to part or all of the speech.; Qian further teaches: “The changed speech can be played back to hear the change in prosody …”; Qian also taches specific changes: duration can be increase or decrease, “F0 trajectories are modifiable,” and loudness is changed by “modifying the gain trajectories.”).
Claim(s) 2 and 8 is/are rejected under 35 U.S.C. 103 as being unpatentable over Qian et al. (US 2010/0066742) in view of Wei et al. (“Neural Network-Based Modeling of Phonetic Durations”; Sept. 6, 2019) and further in view of Ren et al. (“FastSpeech 2: Fast and High-Quality End-To-End Text to Speech”; Mar. 4, 2021 (version 6)).
Claim 2,
Qian further teaches the computer-implemented method of claim 1, wherein the pitch distribution corresponds to frequency ([0019] Qian directly ties pitch to frequency: “the excitation feature is the log of the fundamental frequency (F0)” and “F0 are modeled by multi-space probability distribution HMM.”).
The difference between the prior art and the claimed invention is that Qian nor Wei explicitly teach the energy distribution corresponds to amplitude.
Ren teaches the energy distribution corresponds to amplitude ([2.3] Ren defines energy from amplitude: “We compute L2-norm of the amplitude of each short-time Fourier transform (STFT) frame as the energy.”; Ren further teaches quantizing energy values: “Then we quantize energy of each frame to 256 possible values…”).
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the teachings of Qian with teachings of Ren by modifying the stylized prosody for speech synthesis-based applications as taught by Qian to include the energy distribution corresponds to amplitude as taught by Ren for the benefit of simplifying the training pipeline and avoid the information loss (Ren [Introduction]).
Claim 8,
Qian and Wei teach all the limitations in claim 1. The difference between the prior art and the claimed invention is that Qian nor Wei explicitly teach aligning a plurality of text tokens, from the sequence of text, to respective mel frames, based on the duration alignment.
Ren teaches aligning a plurality of text tokens, from the sequence of text, to respective mel frames, based on the duration alignment ([2.2] [2.3] Ren teaches that the “duration predictor takes the phoneme hidden sequence as input and predicts the duration of each phoneme, which represents how many mel frames correspond to this phoneme.”; Ren also teaches that the “mel-spectrogram decoder converts the adapted hidden sequence into mel-spectrogram sequence in parallel.”; phoneme token derived from the text sequence are text tokens; Ren aligns each phoneme token to its corresponding mel frames using duration).
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the teachings of Qian with teachings of Ren by modifying the stylized prosody for speech synthesis-based applications as taught by Qian to include aligning a plurality of text tokens, from the sequence of text, to respective mel frames, based on the duration alignment as taught by Ren for the benefit of simplifying the training pipeline and avoid the information loss (Ren [Introduction]).
Claim(s) 9 is/are rejected under 35 U.S.C. 103 as being unpatentable over Qian et al. (US 2010/0066742) in view of Wei et al. (“Neural Network-Based Modeling of Phonetic Durations”; Sept. 6, 2019) in view of Ren et al. (“FastSpeech 2: Fast and High-Quality End-To-End Text to Speech”; Mar. 4, 2021 (version 6)) and further in view of Ping et al. (“Deep Voice 3: Scaling Text-to-Speech with Convolutional Sequence Learning”; Feb. 22, 2018).
Claim 9,
Qian, Wei and Ren teach all the limitations in claim 8. The difference between the prior art and the claimed invention is that Qian, Wei nor Ren explicitly teach normalizing probability distributions for the plurality of text tokens and the respective mel frames.
Ping teaches normalizing probability distributions for the plurality of text tokens and the respective mel frames ([Fig. 3] [3.4] [3.5] [3.6] Ping teaches that the encoder “converts characters or phonemes into trainable vector representations” and that the key vectors are used “to compute attention weights.”; Ping further teaches that the decoder uses “mel-band log-magnitude spectrogram” as the audio frame representation and predicts groups of “audio frames”; Ping’s attention block uses “softmax OR monotonic attention,” and later teaches computing “softmax” over the input/a fixed window of input tokens; Softmax normalizes attention weights into probability distribution over text tokens for each decoder timestep/mel-frame generation step).
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the teachings of Ren with teachings of Ping by modifying the fast and high-quality end-to-end text to speech system as taught by Ren to include normalizing probability distributions for the plurality of text tokens and the respective mel frames as taught by Ping for the benefit of identifying common errors modes of attention-based speech synthesis networks and mitigate them (Ping [Abstract]).
Allowable Subject Matter
Claims 10-20 are allowed.
Claim 5 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
The following is a statement of reasons for the indication of allowable subject matter:
For Claim 10:
Mohammadi (US 10,186,252) in view of Ping et al. (“Deep Voice 3: Scaling Text-to-Speech with Convolutional Sequence Learning”; Feb. 22, 2018) teach all the limitations. The difference between the prior art and the claimed invention is the Mohammadi nor Ping explicitly teach the synthetic alignment is based on probabilistically sampling phoneme distributions for a plurality of audio segments used to train the text-to-speech model, at inference, and a first alignment between the sequence of text and a total speech duration.
For Claim 17:
Mohammadi (US 10,186,252) in view of Ping et al. (“Deep Voice 3: Scaling Text-to-Speech with Convolutional Sequence Learning”; Feb. 22, 2018) teach all the limitations. The difference between the prior art and the claimed invention is the Mohammadi nor Ping explicitly teach determine one or more vectors corresponding to one or more speaker characteristics based on a concentrated probability distribution of a text sequence from the text and mel-frames of the alignment distribution across the duration of the plurality of audio samples.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHREYANS A PATEL whose telephone number is (571)270-0689. The examiner can normally be reached Monday-Friday 8am-5pm PST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre Desir can be reached at 571-272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
SHREYANS A. PATEL
Primary Examiner
Art Unit 2653
/SHREYANS A PATEL/Examiner, Art Unit 2659