Prosecution Insights
Last updated: October 01, 2026
Application No. 19/020,264

METHOD FOR SPEECH GENERATION AND RELATED DEVICE

Non-Final OA §102§103
Filed
Jan 14, 2025
Priority
Jul 15, 2022 — RU 2022119398 +1 more
Examiner
MCLEAN, IAN SCOTT
Art Unit
Tech Center
Assignee
Huawei Technologies Co., Ltd.
OA Round
1 (Non-Final)
43%
Grant Probability
Moderate
1-2
OA Rounds
1y 5m
Est. Remaining
75%
With Interview

Examiner Intelligence

Grants 43% of resolved cases
43%
Career Allowance Rate
26 granted / 60 resolved
-16.7% vs TC avg
Strong +32% interview lift
Without
With
+32.1%
Interview Lift
resolved cases with interview
Typical timeline
3y 1m
Avg Prosecution
25 currently pending
Career history
95
Total Applications
across all art units

Statute-Specific Performance

§101
4.5%
-35.5% vs TC avg
§103
70.3%
+30.3% vs TC avg
§102
22.6%
-17.4% vs TC avg
§112
1.6%
-38.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 60 resolved cases

Office Action

§102 §103
Notice of Pre-AIA or AIA Status 1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 102 2. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. 3. Claims 1, 3-5, 7, 9, 11-13 and 15 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Tang et al. “TGAVC: IMPROVING AUTOENCODER VOICE CONVERSION WITH TEXT-GUIDED AND ADVERSARIAL TRAINING” herein Tang. Regarding Claim 1: Tang discloses a method for speech generation (Tang: Abstract and Section 1 discloses a text guided AutoVC voice conversation method that generates converted speech containing pieces of linguistic content from source speech and the voice characteristics of target speech. This conversion is interpreted as speech generation), comprising: obtaining a first source data input to a speech generation model comprising multiple encoders and a decoder, wherein types of input data of the multiple encoders are different (Tang: Section 2.2.1 and Fig. 2 disclose obtaining source speech, as first source data input xi to the TGAVC speech generation model. The model includes a content encoder Ec, a text encoder Et, a style encoder Es and a decoder D. The speech input is input into the content and style encoders and the text transcription is input to the text encoder. Therefore, the multiple encoders accept different input data types comprising speech data and text data.); generating a first acoustic feature by a first encoder among the multiple encoders based on the first source data, wherein a type of the first source data is consistent with the type of the input data of the first encoder (Tang: Section 2.2.1 discloses that the source speech represents acoustic features such as a mel spectrogram and is provided to the contend encoder which then generates content embedding Hc. The content embedding is a first acoustic feature derived from the acoustic source speech. The first source data and the input data accepted by content encoder are therefore consistently speech and acoustic data); and converting a second acoustic feature determined from the first acoustic feature into a third acoustic feature by the decoder, wherein the third acoustic feature is configured to generate a speech with a target voice (Tang: Fig. 2 and Section 2.2.1 – 2.2.2 disclose that the conversion combines the content embedding Hci with the style embedding Hsj which is extracted from target speech Xj. The style embedding is copied to the same length as the content embedding and concatenated with the content embedding in the channel dimension. This concatenated representation is the second acoustic feature determined from the first acoustic feature The concatenated representation to decoder D, which generates converted Speech xi->j. This is the third acoustic feature as claimed). Regarding Claim 3: Tang further discloses the method according to claim 1, wherein the multiple encoders comprise at least two: a speech encoder, a text encoder, or a video encoder, wherein the first encoder is the speech encoder when the first source data is audio data, the first encoder is the text encoder when the first source data is text data, or the first encoder is the video encoder when the first source data is video data (Tang: Section 2.2.1 and Fig. 2 discloses content encoder which receives source speech and a text encoder which receives text transcriptions). Regarding Claim 4: Tang further discloses the method according to claim 1, wherein the multiple encoders and the decoder are trained, respectively (Section 2.2.1 and Section 2.2.2 discloses training the content encoder, text encoder, style decoder and decoders respectively). Regarding Claim 5: Tang further discloses the method according to claim 1, wherein the third acoustic feature is a target spectrogram and the first acoustic feature is a spectrogram-like feature corresponding to the first source data, and the spectrogram-like feature corresponding to the first source data is anyone of: a spectrogram corresponding to the first source data, an acoustic feature corresponding to the first source data aligned with the target spectrogram on a time axis, or concatenation of the spectrogram corresponding to the first source data and the acoustic feature corresponding to the first source data aligned with the target spectrogram on the time axis (Tang: Sectopms 2.2.1 – 2.2.2 and Figs. 2-3 discloses text encoder generating content embedding Hc. Tang explicitly discloses a length regulator introduced to align text embedding Hc with speech content embedding Ĥc with the required alignment between text and speech obtained using Montreal Forced Aligner. Tang trains the embeddings to correspond using the elementwise loss and provides the temporally aligned content embedding to the decoder D along with the target style embedding to generate converted speech represented by target voice Mel-spectrogram acoustic features. Therefore, Hc is an acoustic feature corresponding to the first source data and aligned with the target spectrogram on a time axis). Regarding Claim 7: Tang further discloses the method according to claim 1, further comprising: obtaining a second source data input to a speech generation model (Tang: Sections 2.1, 2.2.1 and Fig. 2(b) discloses obtaining target speech xj as second source data input to the voice conversion model, in addition to obtaining source speech xi); and generating a fourth acoustic feature by a second encoder among the multiple encoders based on the second source data, wherein the type of the second source data is consistent with the type of the input data of the second encoder, and the second acoustic feature is obtained by concatenating the fourth acoustic feature and the first acoustic feature (Tang: Sections 2.2.1 and 2.2.2 discloses providing target speech xj to the style encoder which generates target style embeddings. The target speech and the input data accepted by the style encoder are speech data. Style embedding Hsj is the fourth acoustic feature because it represents acoustic voice characteristics extracted from the second source data). Regarding Claim 9: Claim 9 has been analyzed with regards to claim 1 (see rejection above) and is rejected for the same reasons of anticipation set forth above. It is noted that Tang necessarily discloses a processor in a communication interface configured to receive or send data, and a memory coupled to the processor to execute the instructions at least at Section 3.1 ‘Configurations’ where the training and execution is done on an NVIDIA V100 GPU, which necessarily has memory (Video Random Access Memory) and processors (CUDA cores and Tensor Cores). Regarding Claim 11: Claim 11 has been analyzed with regards to claim 3 (see rejection above) and is rejected for the same reasons of anticipation set forth above. Regarding Claim 12: Claim 12 has been analyzed with regards to claim 4 (see rejection above) and is rejected for the same reasons of anticipation set forth above. Regarding Claim 13: Claim 13 has been analyzed with regards to claim 5 (see rejection above) and is rejected for the same reasons of anticipation set forth above. Regarding Claim 15: Claim 15 has been analyzed with regards to claim 7 (see rejection above) and is rejected for the same reasons of anticipation set forth above. Regarding Claim 17: Claim 17 has been analyzed with regards to claim 1 (see rejection above) and is rejected for the same reasons of anticipation set forth above. It is noted that Tang necessarily discloses non-transitory machine readable storage medium, processor in a communication interface configured to receive or send data, and a memory coupled to the processor to execute the instructions at least at Section 3.1 ‘Configurations’ where the training and execution is done on an NVIDIA V100 GPU, which necessarily has memory (Video Random Access Memory) and processors (CUDA cores and Tensor Cores). Regarding Claim 19: Claim 19 has been analyzed with regards to claim 3 (see rejection above) and is rejected for the same reasons of anticipation set forth above. Regarding Claim 20: Claim 20 has been analyzed with regards to claim 4 (see rejection above) and is rejected for the same reasons of anticipation set forth above. Claim Rejections - 35 USC § 103 4. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. 5. Claims 2, 8, 10 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Tang in view of Jeong et al. “Diff-TTS: A Denoising Diffusion Model for Text-to-Speech” herein Jeong. Regarding Claim 2: Tang further discloses the method according to claim 1, wherein wherein the converting the second acoustic feature determined from the first acoustic feature into the third acoustic feature by the comprises: converting the second acoustic feature determined from the first acoustic feature into the third acoustic feature by the decoder (Tang: Section 2.2.1, Fig. 2 and Eq. (5) disclose generating the desired content embedding Hc with the text encoder and supplying Hc and the style embedding Hs to decoder D to generate converted speech x̂i. 2.2.2 further discloses copying the style embedding to the length of the content embedding, concatenating the embeddings and passing the concatenated embedding to the decoder. Therefore, Tang discloses the claimed conversion operation) through Tang does not explicitly disclose: the decoder is a diffusion- based decoder; converting the second acoustic feature determined from the first acoustic feature into the third acoustic feature by the decoder through a reverse diffusion process. However, Jeong discloses: the decoder is a diffusion- based decoder (Jeong: Abstract and Section 1 discloses Diff-TTS, a denoising diffusion speech synthesis model that transforms a noise signal into a mel-spectrogram through diffusion time steps, Section 2.3 and Fig. 3 disclose a diffusion decoder that predicts Gaussian noise from a diffusion step latent variable conditioned on an encoder embedding and diffusion step embedding); converting the second acoustic feature determined from the first acoustic feature into the third acoustic feature by the decoder through a reverse diffusion process (Jeong: Section 2.1 and Eqs. (3)-(5) disclose that the reverse process is the mel-spectrogram generation procedure performed backward from the forward diffusion process. The reverse transitions gradually restore the latent variables to a mel spectrogram and during inference the decoder iteratively predicts and removes the noise introduced by the forward transition). Tang and Jeong are combinable because they are from the same field of endeavor, both disclose systems and methods for generating speech from encoded source information using encoder-decoder neural networks. Tang discloses generating converted speech from content and target-speaker style embeddings but employs a conventional decoder. Jeong teaches using a diffusion-based decoder conditioned on an encoder embedding to generate a high fidelity mel spectrogram through reverse diffusion. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to disclose Tang’s decoder. The motivation for doing so is to obtain the robust, high quality speech synthesis and stable training identified in Jeong Section 1: “the denoising diffusion models can be stably optimized according to maximum likelihood and enjoy the freedom of architecture choices,” while maintaining Tang’s concatenated content and style embedding as the decoder information. Regarding Claim 8: Tang further discloses the method according to claim 1, wherein the converting the second acoustic feature determined from the first acoustic feature into the third acoustic feature by the decoder comprises: converting the second acoustic feature determined from the first acoustic feature into the third acoustic feature by the decoder (Tang Section 2.2.1, Fig. 2 and Eqs. (4)-(5) disclose generating content embedding Hc from source data and style embedding Hs from target speech and supplying the content and style embeddings to decoder D to generate converted speech. Tang Section 2.2.2 discloses copying the style embedding to the length of the content embedding, concatenating the style embedding with the content embedding and passing the concatenated embedding into the decoder to generate the speech. Accordingly, Tang’s decoder converts the concatenated acoustic feature determined from the acoustic feature into the target voice acoustic output while conditioned on Hs.). Tang does not explicitly disclose: through a reverse diffusion process. However, Tang discloses through a reverse diffusion process (Jeong: Section 2.1 and Eqs. (3)-(5) disclose that the reverse process is the mel-spectrogram generation procedure performed backward from the forward diffusion process. The reverse transitions gradually restore the latent variables to a mel spectrogram and during inference the decoder iteratively predicts and removes the noise introduced by the forward transition). Tang and Jeong are combinable because they are from the same field of endeavor, both disclose systems and methods for generating speech from encoded source information using encoder-decoder neural networks. Tang discloses generating converted speech from content and target-speaker style embeddings but employs a conventional decoder. Jeong teaches using a diffusion-based decoder conditioned on an encoder embedding to generate a high fidelity mel spectrogram through reverse diffusion. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to disclose Tang’s decoder. The motivation for doing so is to obtain the robust, high quality speech synthesis and stable training identified in Jeong Section 1: “the denoising diffusion models can be stably optimized according to maximum likelihood and enjoy the freedom of architecture choices,” while maintaining Tang’s concatenated content and style embedding as the decoder information. Regarding Claim 10: Claim 10 has been analyzed with regards to claim 2 (see rejection above) and is rejected for the same reasons of obviousness set forth above. Regarding Claim 16: Claim 16 has been analyzed with regards to claim 8 (see rejection above) and is rejected for the same reasons of obviousness set forth above. 6. Claims 6 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Tang in view of Chen et al. “ADASPEECH: ADAPTIVE TEXT TO SPEECH FOR CUSTOM VOICE” herein Chen. Regarding Claim 6: Tang further discloses the method according to claim 5, except wherein the first acoustic feature is an average spectrogram corresponding to the first source data. However, Chen discloses wherein the first acoustic feature is an average spectrogram corresponding to the first source data (Chen: Section 2.1 and Fig. 2(c) disclose a phoneme level acoustic encoder supplied with “Phoneme-Level Mel,” wherein “phoneme-level mel means the mel-frames aligned to the same phoneme are averaged.” Chen further explains that the speech frames corresponding to each phoneme are averaged according to the alignment between). Tang and Chen are combinable because they are from the same field of endeavor, both disclose systems and methods for generating speech in a selected voice using encoder-decoder neural networks and phoneme-aligned acoustic representations. Tang uses Montreal forced alignment to align its text and speech content representations but does not average the aligned mel spectrogram frames. Chen teaches averaging the Mel spectrogram frames aligned with each phoneme to produce a phoneme-level average mel-spectrogram. A person of ordinary skill in the art would have been motivated to incorporate Chen’s frame averaging operation into Tang’s content encoder so that Tang’s first acoustic feature is an average mel spectrogram corresponding to the source data because Cheng explains that “phoneme level acoustic modeling can indeed help the learning of acoustic conditions and is critical to ensure the adaption quality” in Section 2.1. Regarding Claim 14: Claim 14 has been analyzed with regards to claim 6 (see rejection above) and is rejected for the same reasons of obviousness set forth above. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to IAN SCOTT MCLEAN whose telephone number is (703)756-4599. The examiner can normally be reached "Monday - Friday 8:00-5:00 EST, off Every 2nd Friday". Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Hai Phan can be reached at (571) 272-6338. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /IAN SCOTT MCLEAN/Examiner, Art Unit 2654 /HAI PHAN/Supervisory Patent Examiner, Art Unit 2654
Read full office action

Prosecution Timeline

Jan 14, 2025
Application Filed
Aug 26, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12744032
SYNTHETIC VIDEO MODEL FOR INSTRUCTING A USER TO IMPROVE SPEECH
3y 10m to grant Granted Sep 22, 2026
Patent 12711468
SYSTEMS AND METHODS TO GENERATE AN ENRICHED MEETING PLAYBACK TIMELINE
4y 1m to grant Granted Aug 18, 2026
Patent 12700484
NAMED-ENTITY RECOGNITION OF PROTECTED HEALTH INFORMATION
3y 4m to grant Granted Aug 04, 2026
Patent 12609127
NEUTRALIZING DISTORTION IN AUDIO DATA
2y 5m to grant Granted Apr 21, 2026
Patent 12602553
SPEECH TRANSLATION METHOD, DEVICE, AND STORAGE MEDIUM
3y 0m to grant Granted Apr 14, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
43%
Grant Probability
75%
With Interview (+32.1%)
3y 1m (~1y 5m remaining)
Median Time to Grant
Low
PTA Risk
Based on 60 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month