DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 Abstract Idea.
Claims, 1, 11 and 20, under step 1, fall within the statutory category of process because it recites a computer implemented method for generating speech and training a speech synthesis model.
Under step 2A, the claims are directed to collecting and analyzing information and using the result to make a training decision. In particular, the method evaluates text and generated speech, determines quality indexes, compares and evaluation index with a preset condition, identifies abnormal sentences and their current ending positions, stores that information, and uses it to train a model. These steps amount to evaluating, organizing, and applying information according to rules, which can be characterized as mental process or mathematical/data-analysis concept. The claimed processor, storage medium, embedding layer, speech synthesis layer, position layer, and database merely provide a generic computer environment for carrying out that analysis. The claims do not recite a specific improvement to computer operation or a particular technical mechanism that changes how speech synthesis is performed; it states the desired result of improved training in functional terms. It therefore, the claims do not integrate the abstract idea into a practical application.
Under step 2B, the additional elements do not provide an inventive concept. Using a processor and memory, generating speech with model layers, storing selected data in a database, comparing an index with a threshold, and training a model when the threshold is met are conventional computer and machine-learning activities applied at a high level of generality. Considered individually and as an ordered combination, these elements merely instruct a computer to perform the abstract evaluation and training process. The claims do not require a new model architecture, a specific training algorithm, a particular calculation of the quality indexes, or a concrete technical modification to the synthesis system.
The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the claims are (i) mere instructions to implement the idea on a computer, and/or (ii) recitation of generic computer structure that serves to perform generic computer functions that are well-understood, routine, and conventional activities previously known to the pertinent industry. Viewed as a whole, these additional claim element(s) do not provide meaningful limitation(s) to transform the abstract idea into a patent eligible application of the abstract idea such that the claim(s) amounts to significantly more than the abstract idea itself. Therefore, the claim(s) are rejected under 35 U.S.C. 101 as being directed to non-statutory subject matter. There is further no improvement to the computing device.
Dependent claims 2-10 and 12-19 further recite an abstract idea performable by a human and do not amount to significantly more than the abstract idea as they do not provide steps other than what is conventionally known.
Claim 2, merely selecting a particular machine-learning model and does not add a technological improvement beyond the abstract idea.
Claims 3 and 12, mathematical analysis and data processing used to evaluate information using mathematical analysis and data manipulation performed on existing information.
Claims 4 and 13, merely evaluating information using mathematical relationships and does not improve computer technology.
Claims 5 and 14, mathematical analysis and data manipulation performed on existing information.
Claims 6 and 15, simply analyzing performance metrics to make a training decision.
Claims 7 and 16, evaluating and classifying information according to rules without improving computer functionality.
Claims 8 and 17, merely using analyzed information to make a determination.
Claims 9 and 18, a mathematical comparison used to trigger a decision and constitutes abstract idea data evaluation.
Claims 10 and 19, merely applies the abstract idea using conventional AI components without a specific technological improvement.
Allowable Subject Matter
Claims 1-20 would be allowable if the Applicant can overcome the 101 Abstract Idea rejection set forth.
The following is a statement of reasons for the indication of allowable subject matter:
Arik et al. (US 10,796,686) teaches a processor-implemented neural TTS system that generates synthesized model having an embedding layer, decoder (speech synthesis layer), and positional encoding layer. Specifically, the encoder beings with “an embedding layer, which converts characters or phonemes into trainable vector representations,” processes those representations through convolution blocks, and provides attention key and value representations to an attention-based decoder that generates low-dimensional audio representation, which are then converted into synthesized speech by a converter or vocoder. Arik further teaches adding positional encoding to attention keys and queries, as well as a binary final-frame of the utterance has been synthesized. The reference also identifies abnormal TTS outputs, including repeated words, skipped words, and mispronunciations, and evaluates those failures using a challenging sentence test set (see [col. 5 line 11 to col. 6 line 30] [col. 8 line 20 to col. 9 line 53] [col. 11 line 57 to col. 12 line 33] [Table 1] [col. 14 line 1 to col. 17 line 12] [Figs. 1, 2, 4 and 7]).
Reber et al. (US 2019/0304435) teaches improving a neural TTS system by evaluating synthesized speech quality and retraining the neural network when a quality criterion is satisfied. The system converts text to phoneme sequences, durations, and pitch information, generates speech through a neural network based signal generation unit, compares the generated speech with reference speech using psychoacoustic analysis, calculates a quality indicator (QI) from audio errors in the generated speech, and trains the neural network when the QI exceeds a preset threshold. Specifically, Reber explains that “the neural network is trained with the QI 442 is above a non-zero quality threshold,” and that the QI is determined from time domain and frequency domain audible errors based on both text derived phoneme information and generated speech (see [Summary] [0028-0030] [0045-0063] [Figs. 2 and 4A-4C]).
Ren et al. (“FastSpeech: Fast, Robust and Controllable Text to Speech”; 2019) teaches a transformer-based neural speech synthesis model that generates speech from text using phoneme embeddings, feed forward transformer blocks with positional encoding, a length regulator, and duration predictor trained from ground truth alignments extracted from a teacher TTS model. The reference explains that FastSpeech addresses abnormal synthesis behaviors, including skipped words and repeated words, by predicting phoneme durations and enforcing hard alignment between text and speech. During training, the duration predictor is trained using ground truth phoneme durations extracted from teacher model attention alignments, and the authors evaluate robustness using 50 sentences that are particularly hard for TTS systems, including single letters, repeated numbers, spellings, long sentences. Ren teaches identifying abnormal sentences and using correct alignment information during training (see [Sections 3.1-3.3] [section 4.1 and 4.3] [Section 5] [Fig. 1] [Table 3]).
The difference between the prior art and the claimed invention is that Arik, Reber nor Ren explicitly teach wherein the training the speech synthesis model when the evaluation index meets the preset condition includes: training the speech synthesis model based on an abnormal training database, wherein the abnormal training database is configured to store abnormal sentences and correct ending positions of the abnormal sentences.
Therefore, it would not have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the teachings of Arik, Reber and Ren to include wherein the training the speech synthesis model when the evaluation index meets the preset condition includes: training the speech synthesis model based on an abnormal training database, wherein the abnormal training database is configured to store abnormal sentences and correct ending positions of the abnormal sentences.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Zhang et al. (US 12,148,415) teaches a method for synthesizing a speech. The method includes generating the speech based on a text with a speech synthesis model, wherein the speech synthesis model includes an embedding layer, a speech synthesis layer, and a position layer; and training the speech synthesis model when an evaluation index meets a preset condition, wherein the evaluation index includes one or more quality indexes determined based on at least a part of the text and at least a part of the speech.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHREYANS A PATEL whose telephone number is (571)270-0689. The examiner can normally be reached Monday-Friday 8am-5pm PST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre Desir can be reached at 571-272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
SHREYANS A. PATEL
Primary Examiner
Art Unit 2653
/SHREYANS A PATEL/Examiner, Art Unit 2659