DETAILED ACTION
Notice of Pre-AIA or AIA Status
1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 112
2. The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
3. Claims 2-3, 11-12 and 20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor, or for pre-AIA the applicant regards as the invention.
Claim 2 recites the limitation “the number of repetitions” in line 4. There is insufficient antecedent basis for this limitation in the claim. Claim 3 depends on Claim 2, thus Claim 3 is rejected as the same ground by virtue of dependency.
Claim 11 recites the limitation “the number of repetitions” in line 4. There is insufficient antecedent basis for this limitation in the claim. Claim 12 depends on Claim 11, thus Claim 12 is rejected as the same ground by virtue of dependency.
Claim 20 recites the limitation “the number of repetitions” in line 5. There is insufficient antecedent basis for this limitation in the claim.
Claim Rejections - 35 USC § 103
4. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
5. Claims 1, 10 and 19 are rejected under 35 U.S.C.103 as being unpatentable over Nishiike et al. (US 2008/0319754 A1) in view of Vaughan et al. (US 2023/0178069 A1.)
With respect to Claim 1, Nishiike et al. disclose
A method of speech generation, comprising:
determining, based on a target text, a plurality of phoneme feature representations corresponding to a sequence of phonemes in the target text and respective phoneme durations for the plurality of phoneme feature representations (Nishiike et al. Fig. 1 elements 4 output a plurality of phoneme feature, element 14 the phoneme length setting unit 14 for setting a phoneme length for each phoneme received from the linguistic processor, [0037] describes receiving input text and words in the text are analyzed with reference to the word dictionary 6, readings, accents, and intonations are determined, and a string of phonetic characters (interlanguage) is output. The types (for example, parts of speech), readings, positions of accents, and the like of words are stored in the word dictionary 6, [0039] describes setting the duration of each phoneme received from the linguistic processor, [0043] describes phoneme duration for each of phoneme, [0063] describes multiply the length of a corresponding phoneme by a constant factor, Fig. 2 element 28.);
extending the plurality of phoneme feature representations based on the respective phoneme durations, to obtain an extended sequence of phoneme feature representations (Nishiike et al. [0039] describes setting the duration of each phoneme received from the linguistic processor, [0043] describes the phoneme length adjusting unit 24 adjusts the length of each phoneme, [0068] describes extending a phoneme duration);
Nishiike et al. fail to explicitly teach
masking at least one phoneme feature representation in the extended sequence of phoneme feature representations, to obtain a sequence of masked phoneme feature representations; and
generating a target speech corresponding to the target text at least based on the sequence of masked phoneme feature representations.
However, Vaughan et al. teach
masking at least one phoneme feature representation in the extended sequence of phoneme feature representations, to obtain a sequence of masked phoneme feature representations (Vaughan et al. [0200] The mask Mij is multiplied (element wise) by the difference; The purpose of the mask is to focus only on the part of the difference matrix that corresponds to the phoneme(s) of interest. The phoneme of interest are those for which it is intended to guide the attention. The mask is configured to set to zero all other parts of the difference matrix, other than those that correspond to the phoneme(s) of interest); and
generating a target speech corresponding to the target text at least based on the sequence of masked phoneme feature representations (Vaughan et al. Fig. 6 element 67 input text, 21 Prediction network, and 69 Output speech, Fig. 6 describes generating speech corresponding to the input text based on the mask phoneme feature representation, [0105] describes the attention network 26. The function of the attention network 26 may be understood to be to act as a mask that focusses on the important features of the encoded features 25 output by the encoder 23.)
Nishiike et al. and Vaughan et al. are analogous art because they are from a similar field of endeavor in the Signal Processing techniques and applications. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the steps of extending the length of the phoneme in synthesizing speech as taught by Nishiike et al., using teaching of masking as taught by Vaughan et al. for the benefit of focusing on the important features of the encoded features (Vaughan et al. [0105] describes the attention network 26. The function of the attention network 26 may be understood to be to act as a mask that focusses on the important features of the encoded features 25 output by the encoder 23.)
Claim 10 recites the similar features as Claim 1, thus Claim 10 is rejected as the same ground as Claim 1.
Claim 19 recites the similar features as Claim 1, thus Claim 19 is rejected as the same ground as Claim 1.
6. Claims 2, 11 and 20 are rejected under 35 U.S.C.103 as being unpatentable over Nishiike et al. (US 2008/0319754 A1) in view of Vaughan et al. (US 2023/0178069 A1) and Venkataramani et al. (US 2025/0273194 A1.)
With respect to Claim 2, Nishiike et al. in view of Vaughan et al. teach all the limitations of Claim 1 upon which Claim 2 depends. Nishiike et al. in view of Vaughan et al. fail to explicitly teach
wherein extending the plurality of phoneme feature representations, to obtain the extended sequence of phoneme feature representations comprises:
for a phoneme feature representation in the plurality of phoneme feature representations,
repeating the phoneme feature representation based on the number of repetitions indicated by a phoneme duration corresponding to the phoneme feature representation; and
concatenating respective repeated phoneme feature representations for the plurality of phoneme feature representations in an order of the plurality of phoneme feature representations, to obtain the extended sequence of phoneme feature representations.
However, Venkataramani et al. teach
wherein extending the plurality of phoneme feature representations, to obtain the extended sequence of phoneme feature representations comprises:
for a phoneme feature representation in the plurality of phoneme feature representations,
repeating the phoneme feature representation based on the number of repetitions indicated by a phoneme duration corresponding to the phoneme feature representation (Venkataramani et al. [0069] describes repeating a phoneme as many time as the corresponding predicted duration); and
concatenating respective repeated phoneme feature representations for the plurality of phoneme feature representations in an order of the plurality of phoneme feature representations, to obtain the extended sequence of phoneme feature representations (Venkataramani et al. [0068-0069] describes concatenating the repeated phonemes and pitches to synthesize speech in the TTS model.)
Nishiike et al., Vaughan et al. and Venkataramani et al. are analogous art because they are from a similar field of endeavor in the Signal Processing techniques and applications. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the steps of extending the length of the phoneme in synthesizing speech as taught by Nishiike et al., using teaching of masking as taught by Vaughan et al. for the benefit of focusing on the important features of the encoded features, using teaching of repeating a phoneme as many times as the corresponding predicted duration as taught by Venkataramani et al. for the benefit of satisfying the desired speaking rate (Fig. 1 element 118, 120, [0069] describes repeating a phoneme as many time as the corresponding predicted duration.)
Claim 11 recites the similar features as Claim 2, thus Claim 11 is rejected as the same ground as Claim 2.
Claim 20 recites the similar features as Claim 2, thus Claim 20 is rejected as the same ground as Claim 2.
7. Claims 4, 13 are rejected under 35 U.S.C.103 as being unpatentable over Nishiike et al. (US 2008/0319754 A1) in view of Vaughan et al. (US 2023/0178069 A1) and Adashcik et al. (US 2026/0171075 A1.)
With respect to Claim 4, Nishiike et al. in view of Vaughan et al. teach all the limitations of Claim 1 upon which Claim 4 depends. Nishiike et al. in view of Vaughan et al. fail to explicitly teach
wherein generating the target speech further comprises:
extracting an acoustic prompt feature representation from a prompt speech of a target speaker; and
generating the target speech based on the sequence of masked phoneme feature representations and the acoustic prompt feature representation.
However, Adashcik et al. teach
wherein generating the target speech further comprises:
extracting an acoustic prompt feature representation from a prompt speech of a target speaker (Adashcik et al.[0003] describes extracting the pitch, energy, duration, and acoustic feature from the speech prompt. See paragraph [0013]); and
generating the target speech based on the sequence of masked phoneme feature representations and the acoustic prompt feature representation (Adashcik et al [0003] This information is put into a transformer architecture configured to generate the target speech with similar acoustic characteristics by using a self-attention mechanism. By supplying this information to the target sentence, alongside the phoneme input representation, the disclosed model is altogether able to synthesize speech of high-quality that sounds similar to the speaking style of the speech prompt. See paragraph [0013].)
Nishiike et al., Vaughan et al. and Adashcik et al. are analogous art because they are from a similar field of endeavor in the Signal Processing techniques and applications. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the steps of extending the length of the phoneme in synthesizing speech as taught by Nishiike et al., using teaching of masking as taught by Vaughan et al. for the benefit of focusing on the important features of the encoded features, using teaching of the speech prompt as taught by Adashcik et al. for the benefit of generating speech of high-quality that sounds similar to the speaking style of the speech prompt (Adashcik et al. [0003] The present disclosure describes an approach for reproducing a voice by taking a speech prompt of a few seconds and learning how to continue speaking in a similar style and with the same speaker identity. The ability to do this is obtained by extracting the pitch, energy, duration, and acoustic features from the speech prompt. This information is put into a transformer architecture configured to generate the target speech with similar acoustic characteristics by using a self-attention mechanism. By supplying this information to the target sentence, alongside the phoneme input representation, the disclosed model is altogether able to synthesize speech of high-quality that sounds similar to the speaking style of the speech prompt.)
With respect to Claim 13, Claim 13 recites the similar features as Claim 4, thus Claim 13 is rejected as the same ground as Claim 4.
Allowable Subject Matter
8. Claims 3, 5-9, 12, 14-18 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. However, claims 2-3, 11-12 and 20 are rejected under 112(b), and for the application to pass to allowance these rejections need to be overcome. Any amendments to overcome the 112(b) rejection that results in any change in scope require further search and/or consideration in order to determine it allowability.
The following is a statement of reasons for the indication of allowable subject matter: the prior art(s) taken alone or in combination fail(s) to teach the following element(s) in combination with the other recited elements in the claim(s).
“wherein masking the at least one phoneme feature representation in the extended sequence of phoneme feature representations comprises:
for a phoneme feature representation in the plurality of phoneme feature representations,
masking one or more of repeated phoneme feature representations for the phoneme feature representation in the extended sequence of phoneme feature representations, to retain one of the repeated phoneme feature representations for the phoneme feature representation.” as recited in Claim 3.
Claim 12 recites the similar features as Claim 3.
“wherein generating the target speech based on the sequence of masked phoneme feature representations and the acoustic prompt feature representation comprises:
generating a first intermediate speech feature representation without condition information;
generating a second intermediate speech feature representation with a condition of the plurality of phoneme feature representations;
generating a third intermediate speech feature representation with a condition of the plurality of phoneme feature representations and a condition of the acoustic prompt feature representation of the prompt feature; and
determining the target speech based on the first intermediate speech feature representation, the second intermediate speech feature representation and the third intermediate speech feature representation.” as recited in Claim 5.
Claim 14 recites the similar features as Claim 5.
“wherein the target speech is generated at least based on the sequence of masked phoneme feature representations using a trained diffusion model, and wherein the diffusion model is trained at least by:
determining, using a language model, a plurality of sample phoneme feature representations corresponding to a sequence of sample phonemes in a sample text and respective sample phoneme durations for the plurality of sample phoneme feature representations based on the sample text and a first sample speech;
determining, using a trained acoustic encoder, a first speech feature representation based on the first sample speech;
extending the plurality of sample phoneme feature representations based on the respective sample phoneme durations, to obtain an extended sequence of sample phoneme feature representations;
masking at least one phoneme feature representation in the extended sequence of sample phoneme feature representations, to obtain a sequence of masked sample phoneme feature representations;
determining, using the diffusion model under training, a reconstructed speech feature representation based on the sequence of masked sample phoneme feature representations and the first speech feature representation;
generating, using a trained acoustic decoder, a reconstructed speech for the first sample speech based on the reconstructed speech feature representation; and
training the diffusion model based on a difference between the first sample speech and the reconstructed speech.” as recited in Claim 6.
Claims 7-9 depends on Claim 6, thus Claims 7-9 are objected to as the same ground by virtue of dependency.
Claim 15 recites the similar features as Claim 6. Claims 16-18 depends on Claim 15, thus Claims 16-18 are objected to as the same ground by virtue of dependency.
Conclusion
9. The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. See PTO-892.
a. Kim et al. (US 2025/0149023 A1.) In this reference, Kim et al. disclose speech synthesis system and method with adjustable utterance length.
b. Morrison et al. (US 2023/0197093 A1.) In this reference, Morrison et al. disclose a method and a system for adjusting the phoneme durations or pitch values.
c. Reber et al. (US 2019/0304434 A1.) In this reference, Reber et al. disclose a method and a system for synthesizing speech.
10. Any inquiry concerning this communication or earlier communications from the examiner should be directed to THUYKHANH LE whose telephone number is (571)272-6429. The examiner can normally be reached Mon-Fri: 9am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew C. Flanders can be reached on 571-272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/THUYKHANH LE/Primary Examiner, Art Unit 2655