Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Specification
The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claim(s) 1-6 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Zhao et al (20220310058).
As per claim 1, Zhao et al (20220310058) teaches a training apparatus comprising: processing circuity configured to:
generate data related to a synthesized speech from first embedded data representing a characteristic of an utterance of a speaker and first text data by using a first model (as generating a first set of training data to generate synthesized speech data with a personalized voice – para 0086; to train a model – para 0085);
generate second text data from data related to the synthesized speech by using a second model (as using a second set of data which refines the initial TTS machine learning model – para 0086) ;
and update a parameter of the first model and a parameter of the second model such that the first embedded data is similar to second embedded data representing a characteristic of an utterance of the synthesized speech, and the first text data is similar to the second text data (as, the finalized TTS model when well trained, will mimic the voice characteristic of the target voice – para 0003, para 0041- 0042).
As per claim 2, Zhao et al (20220310058) teaches the training apparatus according to claim 1, wherein the processing circuitry is further configured to generate, in a first stage, data related to a synthesized speech from third embedded data representing a characteristic of an utterance of a speaker and third text data by using the first model (as generating data, also based on, personal content of the user – para 0049, in the voice of the user, by operating on prosody information, as an example – para 0081; as applied to claim 1 above, the first data is toward the training of the first model – para 0085),
in the first stage, fourth text data from data related to the synthesized speech by using the second model, update, in the first stage, the parameter of the first model such that the third text data and the fourth text data are similar to each other (as updating the model with a second set of synthesized speech generated by the TTS machine learning model – seeFigure 5, steps 540 to 550C),
generate, in a second stage after the first stage, data related to a synthesized speech from the first embedded data representing the characteristic of the utterance of the speaker and the first text data by using the first model whose parameter has been updated in the first stage, generate, in the second stage, the second text data from data related to the synthesized speech by using the second model (as updating/refining the TTS machine learning model after initial personalization of the model – see figure 5, subblocks 550A to 550C), and
update, in the second stage, the parameter of the first model and the parameter of the second model such that the first embedded data is similar to the second embedded data representing the characteristic of the utterance of the synthesized speech, and the first text data is similar to the second text data (as, updating/training the TTS based on the modified training data – and generating second data as the output of the personal content from a user using the neural TTS model – para 0049; the first embedded data is generated from the first model and the second model – see para 0003, 0041-0042 – the finalized TTS model mimics the voice characteristic of the target voice; see also Figure 5, subblocks 510-560C, last step of the generated speech data).
As per claim 3, Zhao et al (20220310058) teaches the training apparatus according to claim 1, wherein the processing circuitry is further configured to update the parameter of the first model and the parameter of the second model such that similarity between the first embedded data that is a vector and the second embedded data that is a vector increases (as, performing a vectorized measure between the target speaker vector, which represents characteristics of the target speaker – para 0078, and measure the accuracy of the model – para 0064).
Claims 4,5 are method/nontransitory computer readable medium claims, respectively, whose steps are performed throughout the apparatus claims 1-3 above and as such, claims 4,5, are similar in scope and content to these claim features found in claims 1-3 above; therefore, claims 4,5 are rejected under similar rationale as presented against claims 1-3 above.
As per claim 6, Zhao et al (20220310058) teaches a synthesis apparatus comprising:
processing circuity configured to generate data related to a synthesized speech from third embedded data representing a characteristic of an utterance of a speaker and a third text data by using a first model whose parameter has been updated by processing of generating data related to a synthesized speech from first embedded data representing a characteristic of an utterance of a speaker and first text data by using the first model (as generating data, also based on, personal content of the user – para 0049, in the voice of the user, by operating on prosody information, as an example – para 0081; as applied to claim 1 above, the first data is toward the training of the first model – para 0085);
generating second text data from data related to the synthesized speech by using a second model, and updating the parameter of the first model and a parameter of the second model such that the first embedded data is similar to second embedded data representing a characteristic of an utterance of the synthesized speech, and the first text data is similar to the second text data (as, generating second data as the output of the personal content from a user using the neural TTS model – para 0049; the first embedded data is generated from the first model and the second model – see para 0003, 0041-0042 – the finalized TTS model mimics the voice characteristic of the target voice; see also Figure 5, subblocks 510-560C, last step of the generated speech data).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Please see related art listed on the PTO-892 form.
Furthermore, the following references contain features that are commonly found in applicants spec/claims:
Pan et al (20230081659) teaches neural networks trained for target speakers and speaker style – fig. 6, as well as changing the target voice based on the text being spoken – para 0025.
Wang et al (20240274122) teaches target speaker characteristics, as well as language and vocal performance features – see fig. 1a, para 0028 – 0030.
Gabrys et al (20230260502) teaches concatenation of multiple sources/types of data for speech synthesis, including targeted spectrogram, speaker embedding, and a differential between the spectrograms – fig. 2, para 0025-0027.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Michael Opsasnick, telephone number (571)272-7623, who is available Monday-Friday, 9am-5pm.
If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Mr. Richemond Dorvil, can be reached at (571)272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free).
/Michael N Opsasnick/Primary Examiner, Art Unit 2658 09/01/2026