DETAILED ACTION
Notice of Pre-AIA or AIA Status
1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments/Amendments
2. With respect to Claim Rejection 35 U.S.C § 102/§ 103, Applicant’s arguments have been considered but are moot because the new ground to rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenge in the argument.
With respect to 101 rejection towards claim 15, Applicant cancelled claim 15. Thus, 101 rejection towards claim 15 is withdrawn.
Claims Objections
3. Claim 16 is objected to because of the following informalities: typographical errors. Claim 16 ends the sentence with double periods. Appropriate correction is required.
Claim Rejections - 35 USC § 103
4. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
5. Claims 1-2, 4-6, 13-14, 16 are rejected under 35 U.S.C.103 as being unpatentable over Yi Ren et al. (‘FastSpeech 2: Fast and High-Quality End-to-End Text to Speech’, 16 Oct 2020) in view of Kim et al. (US 2023/0206896 A1).
With respect to Claim 1, Yi Ren et al. disclose
An apparatus for end-to-end text-to-speech synthesis, comprising:
input interface circuitry configured to receive:
first input data indicative of a phoneme (Yi Ren et al. Fig. 1(a): “Phoneme”, p.4, section 2.2, l. 2-5); and
second input data indicative of a first target duration for the phoneme (Yi Ren et al. Fig. 1(a), “Variance Adaptor” and Fig. 1(b): “Duration predictor; p.4, section 2.3 “Variance Adaptor”, 1st and 2nd paragraph);
processing circuitry configured to, using a trained machine-learning model:
map the phoneme indicated by the first input data to a state using an encoder sub- model of the trained machine-learning model (Yi Ren et al. p. 3, section 2.2, l 1-5: “phoneme embedding”);
estimate a second target duration for the phoneme indicated by the first input data based on the state (Yi Ren et al. p.3 section 2.2, l 1-5);
determine an attention weight based on the first target duration indicated by the second input data for the phoneme indicated by the first input data and the second target duration for the phoneme indicated by the first input data estimated based on the state (Yi Ren et al. Fig. 1(a), output of variance adaptor; p.5, section 2.5, l 3-6); and
map the state to audio data based on the attention weight using a decoder sub- model of the trained machine-learning model (Yi Ren et al. Fig. 1(a): “Waveform Decoder),
wherein the audio data are indicative of an audio waveform representing speech (Yi Ren et al. Fig. 1(a): “Waveform Decoder).
Yi Ren et al. fail to explicitly teach input interface circuitry to perform converting text to speech. However, Kim et al. teach
input interface circuitry configured to perform converting text to speech (Kim et al. [0037] CPU, ASIC, FPGA, converting text to speech at paragraphs [0087, 0033-0034, 0124-0125 and Fig. 9 element 930 decoder, 720 vocoder, Fig. 10 element 930 decoder.)
Yi Ren et al. and Kim et al. are analogous art because they are from a similar field of endeavor in the Speech Processing techniques and applications. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the steps of converting text to speech as taught by Yi Ren et al., using teaching of the circuitry as taught by Kim et al. for the benefit of perform the method of converting text to speech (Kim et al. [0037] CPU, ASIC, FPGA, converting text to speech at paragraphs [0087, 0033-0034, 0124-0125 and Fig. 9 element 930 decoder, 720 vocoder, Fig. 10 element 930 decoder.)
With respect to Claim 2, Yi Ren et al. in view of Kim et al. teach
wherein the first input data are indicative of a text to be converted into the speech, and wherein the processing circuitry is further configured to determine the phoneme based on the text (Kim et al. [0087] describes dividing the input text into the syllable, character and/or phoneme.)
With respect to Claim 4, Yi Ren et al. in view of Kim et al. teach
wherein the second input data are indicative of a video depicting a speaking person, and wherein the processing circuitry is further configured to determine the first target duration based on the video to synchronize the speech to the video (Kim et al. [0007] describes determining the number of frames including a speaker’s mouth shape, and [0048] describes the speaker’s mouth shape and the synthesis voice included in each frame included in the video content may be synchronized.)
With respect to Claim 5, Yi Ren et al. in view of Kim et al. teach
wherein, for determining the first target duration, the processing circuitry is configured to:
determine a gesture of the speaking person matching the phoneme based on the video (Kim et al. [0082] determining an image frame including the speaker’s mouth shape); and
determine the first target duration based on the determined gesture (Kim et al. [0082] determining the number of frames including the speaker’s mouth shape.)
With respect to Claim 6, Yi Ren et al. in view of Kim et al. teach
wherein the processing circuitry is configured to:
determine a second phoneme matching a shape of lips of the speaking person based on the video (Kim et al. [0007-0008] describes determining a phoneme corresponding with the speaker’s mouth shape);
determine a correlation between the phoneme and the second phoneme (Kim et al. [0047-0048] determining a relationship between the phoneme in the input text and the phoneme corresponding with the speaker’s mouth shape); and
determine the first target duration based on the correlation (Kim et al. [0053 and 0071] determining the target duration in the synthesized voice.)
With respect to Claim 13, Claim 13 recites similar features as Claim 1, thus Claim 13 is rejected as the same ground as Claim 1.
With respect to Claim 14, Claim 14 is a non-transitory machine-readable medium claim for performing the method of Claim 13, thus Claim 14 is rejected as the same ground as Claim 13.
With respect to Claim 16, Yi Ren et al. in view of Kim et al. teach
wherein
the phoneme indicated by the first input data is one of a plurality of phonemes,
the first input data are indicative of the plurality of phonemes (Yi Ren et al. Fig. 1(a): “Phoneme”, p.4, section 2.2, l. 2-5),
the second input data are indicative of a respective first target duration for each phoneme of the plurality of phonemes (Yi Ren et al. Fig. 1(a), “Variance Adaptor” and Fig. 1(b): “Duration predictor; p.4, section 2.3 “Variance Adaptor”, 1st and 2nd paragraph), and
the processing circuitry is configured to:
map each phoneme of the plurality of phonemes to a respective state using the encoder sub-model (Yi Ren et al. p. 3, section 2.2, line 1-5: “phoneme embedding”);
estimate a respective second target duration for each phoneme of the plurality of phonemes based on the respective state of that phoneme Yi Ren et al. p.3 section 2.2, lines 1-5); and
determine a respective attention weight for each phoneme of the plurality of phonemes based on the respective first target duration for that phoneme indicated by the second input data and the respective second target duration for that phoneme estimated based on the respective state of that phoneme (Yi Ren et al. Fig. 1(a), output of variance adaptor; p.5, section 2.5, lines 3-6, Fig. 1(a): “Waveform Decoder. Examiner notes that Yi Ren et al. provides the duration of each phoneme to train a duration predictor and generated mel-spectrograms for knowledge distillation. This implied that the method in Yi Ren et al. could be applied for any phoneme in the text)..
6. Claim 3 is rejected under 35 U.S.C.103 as being unpatentable over Yi Ren et al. (‘FastSpeech 2: Fast and High-Quality End-to-End Text to Speech’, 16 Oct 2020) in view of Kim et al. (US 2023/0206896 A1) and Bruckert (6,029,131).
With respect to Claim 3, Yi Ren et al. in view of Kim et al. teach all the limitations of Claim 2 upon which Claim 3 depends. Yi Ren et al. in view of Kim et al. fail to explicitly teach
wherein the second input data are indicative of a stress to be given to the text or parts thereof, and
wherein the processing circuitry is further configured to determine the first target duration based on the stress.
However, Bruckert teaches
wherein the second input data are indicative of a stress to be given to the text or parts thereof (Bruckert Fig. 6A and 6C, Col. 7 lines 1-16 and lines 28-34 describes adjusting the time durations of the phonemes based on the stress time intervals), and
wherein the processing circuitry is further configured to determine the first target duration based on the stress (Bruckert Fig. 6A and 6C, Col. 7 lines 1-16 and lines 28-34 describes adjusting the time durations of the phonemes based on the stress time intervals)
Yi Ren et al., Kim et al. and Bruckert are analogous art because they are from a similar field of endeavor in the Speech Processing techniques and applications. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the steps of converting text to speech as taught by Yi Ren et al., using teaching of the circuitry as taught by Kim et al. for the benefit of perform the method of converting text to speech, using teaching of determining the stress time intervals as taught by Bruckert for the benefit of adjusting the time durations of the phonemes in the synthetic speech (Bruckert Fig. 6A and 6C, Col. 7 lines 1-16 and lines 28-34 describes adjusting the time durations of the phonemes based on the stress time intervals.)
7. Claim 7 is rejected under 35 U.S.C.103 as being unpatentable over Yi Ren et al. (‘FastSpeech 2: Fast and High-Quality End-to-End Text to Speech’, 16 Oct 2020) in view of Kim et al. (US 2023/0206896 A1) and Brochu (10,276,189)
With respect to Claim 7, Yi Ren et al. in view of Kim et al. teach all the limitations of Claim 6 upon which Claim 7 depends. Yi Ren et al. in view of Kim et al. fail to explicitly teach
wherein the processing circuitry is configured to determine the correlation based on dynamic time warping.
However, Brochu teaches
wherein the processing circuitry is configured to determine the correlation based on dynamic time warping (Brochu Col. 10 lines 66-67, Col. 11 lines 1-19 describes using the dynamic time warping to measure similarity and/or distance between two time domain sequences in the video files and the audio files.)
Yi Ren et al., Kim et al. and Brochu are analogous art because they are from a similar field of endeavor in the Speech Processing techniques and applications. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the steps of converting text to speech as taught by Yi Ren et al., using teaching of the circuitry as taught by Kim et al. for the benefit of perform the method of converting text to speech, using teaching of the dynamic time warping as taught by Brochu for the benefit of measuring similarity between two time domain sequences in the video files and the audio files (Brochu Col. 10 lines 66-67, Col. 11 lines 1-19 describes using the dynamic time warping to measure similarity and/or distance between two time domain sequences in the video files and the audio files.)
8. Claim 8 is rejected under 35 U.S.C.103 as being unpatentable over Yi Ren et al. (‘FastSpeech 2: Fast and High-Quality End-to-End Text to Speech’, 16 Oct 2020) in view of Kim et al. (US 2023/0206896 A1) and Saino (US 2023/0098145 A1.)
With respect to Claim 8, Yi Ren et al. in view of Kim et al. teach all the limitations of Claim 1 upon which Claim 8 depends. Yi Ren et al. in view of Kim et al. fail to explicitly teach
wherein the processing circuitry is configured to estimate the second target duration based on at least one of a differentiable function and a stochastic process.
However, Saino et al. teach
wherein the processing circuitry is configured to estimate the second target duration based on at least one of a differentiable function and a stochastic process (Saino et al. [0038] describes using a statistical estimation model to estimate the duration of each phoneme.)
Yi Ren et al., Kim et al. and Saino are analogous art because they are from a similar field of endeavor in the Speech Processing techniques and applications. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the steps of converting text to speech as taught by Yi Ren et al., using teaching of the circuitry as taught by Kim et al. for the benefit of perform the method of converting text to speech, using teaching of the statistical estimation model as taught by Saino et al. for the benefit of estimation the duration of each phoneme (Saino et al. [0038] describes using a statistical estimation model to estimate a duration of each phoneme.)
9. Claim 9 is rejected under 35 U.S.C.103 as being unpatentable over Yi Ren et al. (‘FastSpeech 2: Fast and High-Quality End-to-End Text to Speech’, 16 Oct 2020) in view of Kim et al. (US 2023/0206896 A1) and Elias et al. (US 2022/0301543 A1.)
With respect to Claim 9, Yi Ren et al. in view of Kim et al. teach all the limitations of Claim 1 upon which Claim 9 depends. Yi Ren et al. in view of Kim et al. fail to explicitly teach
wherein the processing circuitry is configured to map the state to the audio data by resampling the state based on the attention weight.
However, Elias et al. teach
wherein the processing circuitry is configured to map the state to the audio data by resampling the state based on the attention weight (Elias et al. [0004] describes upsampling the sequence representation of an encoded text into an upsampled output specifying a number of frames using the interval representation matrix and the auxiliary attention context representation.)
Yi Ren et al., Kim et al. and Elias et al. are analogous art because they are from a similar field of endeavor in the Speech Processing techniques and applications. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the steps of converting text to speech as taught by Yi Ren et al., using teaching of the circuitry as taught by Kim et al. for the benefit of perform the method of converting text to speech, using teaching of the upsampling as taught by Elias et al. for the benefit of generating one or more predicted mel-frequency spectrogram sequences for the encoded text sequence (Elias et al. [0004] describes upsampling the sequence representation of an encoded text into an upsampled output specifying a number of frames using the interval representation matrix and the auxiliary attention context representation and generating one or more predicted mel-frequency spectrogram sequences for the encoded text sequence.)
10. Claim 10 is rejected under 35 U.S.C.103 as being unpatentable over Yi Ren et al. (‘FastSpeech 2: Fast and High-Quality End-to-End Text to Speech’, 16 Oct 2020) in view of Kim et al. (US 2023/0206896 A1) and Zhang (US 2022/0059072 A1.)
With respect to Claim 10, Yi Ren et al. in view of Kim et al. teach all the limitations of Claim 1 upon which Claim 10 depends. Yi Ren et al. in view of Kim et al. fail to explicitly teach
wherein the processing circuitry is configured to determine the attention weight by estimating a probability that the state is aligned with a predefined frame for the audio waveform.
However, Zhang et al. teach
wherein the processing circuitry is configured to determine the attention weight by estimating a probability that the state is aligned with a predefined frame for the audio waveform (Zhang et al. [0048, 0053] and Fig. 4 describes estimating a probability that the speech frames of the speech are aligned with the characters of the text in order to determining an importance index of each weight in the first weight matrix, and generating a second weight matrix based on the first weight matrix.)
Yi Ren et al., Kim et al. and Zhang et al. are analogous art because they are from a similar field of endeavor in the Speech Processing techniques and applications. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the steps of converting text to speech as taught by Yi Ren et al., using teaching of the circuitry as taught by Kim et al. for the benefit of perform the method of converting text to speech, using teaching of the probability as taught by Zhang et al. for the benefit of estimating how similar between the speech frames of the speech and the characters of the text in order to evaluate the accuracy of the speech model (Zhang et al. [0048, 0053] and Fig. 4 describes estimating a probability that the speech frames of the speech are aligned with the characters of the text in order to determining an importance index of each weight in the first weight matrix, and generating a second weight matrix based on the first weight matrix. Zhang et al. indicates that the greater the probability is higher accuracy of the speech model.)
11. Claim 19 is rejected under 35 U.S.C.103 as being unpatentable over Yi Ren et al. (‘FastSpeech 2: Fast and High-Quality End-to-End Text to Speech’, 16 Oct 2020) in view of Kim et al. (US 2023/0206896 A1) and Nishiike et al. (US 2008/0319754 A1.)
With respect to Claim 19, Yi Ren et al. in view of Kim et al. teach all the limitations of Claim 1 upon which Claim 19 depends. Yi Ren et al. in view of Kim et al. fail to explicitly teach
wherein the first target duration indicated by the second input data is specific to the phoneme indicated by the first input data.
However, Nishiike et al. teach
wherein the first target duration indicated by the second input data is specific to the phoneme indicated by the first input data (Nishiike et al. [0041] The phoneme length table 16 is means for storing phoneme lengths at the normal speech rate, each in response to a corresponding phoneme and preceding and following phonemes. In exemplary setting of a phoneme length, phoneme lengths (values extracted from a database) at the normal speech rate, each in response to a corresponding phoneme and preceding and following phonemes, are stored in the phoneme length table 16 in advance, and a phoneme length is set with reference to the values of the phoneme lengths. The phoneme length may be corrected using another parameter element.)
Yi Ren et al., Kim et al. and Nishiike et al. are analogous art because they are from a similar field of endeavor in the Speech Processing techniques and applications. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the steps of converting text to speech as taught by Yi Ren et al., using teaching of the circuitry as taught by Kim et al. for the benefit of perform the method of converting text to speech, using teaching of the phoneme length table as taught by Nishiike et al. for the benefit of converting text to speech for each phoneme based on the length of each phoneme in the phoneme length table (Nishiike et al. [0041] The phoneme length table 16 is means for storing phoneme lengths at the normal speech rate, each in response to a corresponding phoneme and preceding and following phonemes. In exemplary setting of a phoneme length, phoneme lengths (values extracted from a database) at the normal speech rate, each in response to a corresponding phoneme and preceding and following phonemes, are stored in the phoneme length table 16 in advance, and a phoneme length is set with reference to the values of the phoneme lengths. The phoneme length may be corrected using another parameter element.)
12. Claim 20 is rejected under 35 U.S.C.103 as being unpatentable over Yi Ren et al. (‘FastSpeech 2: Fast and High-Quality End-to-End Text to Speech’, 16 Oct 2020) in view of Kim et al. (US 2023/0206896 A1) and Saito et al. (US 2005/0114137 A1.)
With respect to Claim 20, Yi Ren et al. in view of Kim et al. teach all the limitations of Claim 1 upon which Claim 20 depends. Yi Ren et al. in view of Kim et al. fail to explicitly teach
wherein the first target duration indicated by the second input data is an estimate of a duration of pronunciation of the phoneme indicated by the first input data.
However, Saito et al. teach
wherein the first target duration indicated by the second input data is an estimate of a duration of pronunciation of the phoneme indicated by the first input data (Saito et al. [0143] describes estimate of a duration of pronunciation of the phoneme.)
Yi Ren et al., Kim et al. and Saito et al. are analogous art because they are from a similar field of endeavor in the Speech Processing techniques and applications. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the steps of converting text to speech as taught by Yi Ren et al., using teaching of the circuitry as taught by Kim et al. for the benefit of perform the method of converting text to speech, using teaching of the phoneme duration estimation as taught by Saito et al. for the benefit of generating a duration of a phoneme string to be synthesized based on the phoneme information (Saito et al. [0144] The phoneme duration estimation unit 50 generates a duration (time arrangement) of a phoneme string to be synthesized based on the phoneme information received from the text analysis unit 10, and stores the generated duration in a predetermined region of the cache memory of the CPU 101 or the main memory 103. The duration is read out in the F0 pattern generation unit 60, the synthesis unit selection unit 70 and the speech generation unit 30 and is used for each processing. For the generation technique of the duration, a publicly known existing technology can be used.)
Allowable Subject Matter
13. Claims 11-12 and 17-18 are objected to as being dependent upon a rejected base claim but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
The following is a statement of reasons for the indication of allowable subject matter: the prior art(s) taken alone or in combination fail(s) to teach the following element(s) in combination with the other recited elements in the claim(s).
“wherein the processing circuitry is configured to determine a second attention weight based on the second target duration and determine the attention weight by modifying the second attention weight based on the first target duration.” as recited in Claim 11.
Claim 12 depends on Claim 11 and thus Claim 12 is objected to as being dependent upon an objected claim(s) by virtue of their dependency.
“compare the first target duration against a predefined threshold, and
in response to the first target duration exceeding the predefined threshold, determine the attention weight from the second target duration.” as recited in Claim 17.
“compute an intermediate target duration as one of an average of the first target duration and the second target duration, a weighted median of the first target duration and the second target duration, and a central tendency of the first target duration and the second target duration, and
determine the attention weight from the intermediate target duration.” as recited in Claim 18.
Conclusion
14. The prior art made of record and not relied upon is considered pertinent to application’s disclosure. See PTO-892.
a. Jia et al. (US 2023/0013777 A1.) In this reference, Jia et al. disclose a method for speech-to-speech translation. The S2ST in Jia et al. receives the context vector and the phoneme representation and generates a translated synthesized speech representation that corresponds to a translation of the utterance spoken in the different second language.
b. Zhang et al. (US 2022/0108680 A1.) In this reference, Zhang et al. disclose a method for text-to-speech using duration prediction. Zhang et al. controls the pace of synthesized audio on a per-word or per-phoneme level by modifying the predicted durations of each word or phoneme determined by a duration prediction neural network, while still maintaining the naturalness of the synthesized speech.
c. Gu et al. (US 2019/0371292 A1.) In this reference, Gu et al. discloses a method for synthesizing speech. In this reference, Gu et al. predicts a time length of a state of each phoneme corresponding to a target text and synthesize speech corresponding to the target text according to the predicted time length.
15. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action.
16. Any inquiry concerning this communication or earlier communications from the examiner should be directed to THUYKHANH LE whose telephone number is (571)272-6429. The examiner can normally be reached Mon-Fri: 9am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew C. Flanders can be reached on 571-272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/THUYKHANH LE/Primary Examiner, Art Unit 2655