Prosecution Insights
Last updated: October 01, 2026
Application No. 19/004,897

METHOD AND SYSTEM FOR AUGMENTED SPEECH EMBEDDINGS BASED AUTOMATIC SPEECH RECOGNITION

Non-Final OA §103
Filed
Dec 30, 2024
Priority
Jan 03, 2024 — IN 202421000499
Examiner
WOZNIAK, JAMES S
Art Unit
Tech Center
Assignee
Tata Group
OA Round
1 (Non-Final)
59%
Grant Probability
Moderate
1-2
OA Rounds
1y 10m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 59% of resolved cases
59%
Career Allowance Rate
241 granted / 408 resolved
-0.9% vs TC avg
Strong +40% interview lift
Without
With
+39.7%
Interview Lift
resolved cases with interview
Typical timeline
3y 7m
Avg Prosecution
23 currently pending
Career history
434
Total Applications
across all art units

Statute-Specific Performance

§101
19.1%
-20.9% vs TC avg
§103
43.2%
+3.2% vs TC avg
§102
16.1%
-23.9% vs TC avg
§112
16.8%
-23.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 408 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Examiner Notes on Patent Subject Matter Eligibility under 35 U.S.C. 101 Independent Claims 1, 3 and 5 involve a process for automatic speech recognition ASR that processes log Mel spectrograms to generate text from input speech using a specifically trained random noise augmented encoder-decoder model. Under the step 2A prong 1 analysis of the 2019 Patent Subject Matter Eligibility Guidelines (2019 PEG), a human could not practically perform such a complex method related to the ASR filed of technology nor does the claim relate to pure mathematics or any judicially created category of organizing human behavior. Accordingly, the independent claims and their dependents that inherit such subject matter are found to be patent eligible under step 2A prong 1 of the 2019 PEG. Claim Objections Claims 2, 4, and 6 are objected to because of the following informalities: In Claim 2, "generating a plurality spectrally flattened speech segments" should be corrected to read --generating a plurality of spectrally flattened speech segments--. Also, in claim 2 "windowing the plurality of speech segments to hamming window" should be corrected to read --windowing the plurality of speech segments with a hamming window--. Claims 4 and 6 contain similar informalities that require similar correction. Appropriate correction is required. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 3, and 5 are rejected under 35 U.S.C. 103 as being unpatentable over Watanabe, et al. (U.S. PG Publication: 2018/0261225 A1) in view of Liu, et al. (“TERA: Self-Supervised Learning of Transformer Encoder Representation for Speech,” 2021) and further in view of Chen, et al. (U.S. Patent: 12,548,559). With respect to Claim 1, Watanabe discloses: A processor-implemented method (processor implementation that applies to the below recitations of “by one or more hardware processors,” Paragraph 0032), the method comprising: receiving, by one or more hardware processors, a speech waveform comprising a plurality of speech signals from an audio recording device (audio waveform received via microphones as an audio recording device, Paragraph 0039); generating, by the one or more hardware processors, a plurality of speech frames based on the speech waveform by preprocessing the speech waveform using a preprocessing technique (time-based frame sequence generation according to a frame index and sampling, Paragraphs 0042-0046); generating, by the one or more hardware processors, a transformed speech signal corresponding to each of the plurality of speech frames by applying Fourier Transform (FT) on each of the plurality of speech frames (STFT performed on the sequence of frames, Paragraphs 0074, 0080, and 0102); generating, by the one or more hardware processors, a log Mel spectrogram associated with each of the plurality of speech frames by passing the transformed speech signal corresponding to each of the plurality of speech frames through a Mel filter bank and thereby computing logarithm of each output of the Mel-filter bank (log Mel filter bank and log operation used to obtain a log Mel feature from the STFT coefficients, Paragraph 0074); and converting, by the one or more hardware processors, the speech waveform into a corresponding textual information based on the log Mel spectrogram associated with each of the plurality of speech frames using a random noise augmented trained encoder-decoder model ("encoder decoder ASR network" used to transform "noisy" speech signals into a text output, Paragraphs 0005, 0033, 0074, and 0076; since the end-to-end network deals with noisy speech, it constitutes a noise-augmented encoder-decoder). Watanabe does not teach the steps for training the random noise augmented encoder-decoder model. Liu, however, discloses: receiving a plurality of training log Mel spectrograms associated with a plurality of training speech waveforms by an encoder-decoder model (training data in the form of log Mel spectrum feature sequence corresponding to utterances in a speech corpus, Section III.A., Page 4; Fig. 1 showing input into transformer structure including encoder and predictor); extracting a plurality of speech embeddings based on the plurality of training log Mel spectrograms from a plurality of hidden layers associated with the encoder-decoder model (acoustic features sampled/extracted from the utterance form an input vector, Section III.A., Page 4); selecting a set of random speech embeddings from among the plurality of speech embeddings and generating a plurality of augmented speech embeddings by replacing the set of random speech embeddings with a randomly generated Gaussian noise (selected input vector sequence is replaced using sampled "Gaussian noise" with a "random magnitude matrix" in a magnitude-based "data augmentation," Section III.A., Page 5); and training the encoder-decoder model based on the plurality of augmented speech embeddings until a predefined threshold (the speech parameters are used to train the encoder-decoder/predictor network until a reconstruction loss value is minimized and/or a number of steps is met, Section III.B., Pages 5-6). Watanabe and Liu are analogous art because they are from a similar field of endeavor in automatic speech recognition using transformer networks. Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date to train the noise augmented encoder-decoder model taught by Watanabe using the data augmentation-based approach taught by Liu to provide a predictable result of achieving a more robust encoder-decoder network that is trained with an increased amount of training data resulting from the augmentation operation (Liu, Section III.A., Page 5). In the above combination of Watanabe and Liu, the plurality of speech embeddings that undergo data augmentation are extracted from an input training utterance sequence and not from a plurality of hidden layers associated with the encoder-decoder model. Chen, however, discloses that a hidden state result in the form of an encoder embedding/vector can be replaced using a stochastic term in the form of a random noise as an alternative to an input vector (Col. 7, Line 64- Col. 8, Line 13; Fig. 4, Elements 135 and 450). Watanabe, Liu, and Chen are analogous art because they are from a similar field of endeavor in automatic speech recognition using transformer networks. Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date to apply stochastic random noise to the encoder-based embedding taught by Chen to the data augmentation process utilized in training the encoder-decoder model of Watanabe in view of Liu to substitute similar elements in the form of a processing order in order to introduce the data augmentation operation taught by Liu. Claim 3 is directed towards a system embodiment including one or more hardware processors coupled to a memory to execute programmed instructions to carry out the method of claim 1, and thus, is rejected under similar rationale. Moreover, Watanabe teaches a system embodiment having I/O interfaces and a memory storing a program coupled to a processor (Paragraphs 0032 and 0097; Fig. 1). Claim 5 is directed towards an embodiment including one or more non-transitory machine-readable media storing computer hardware-executable instructions to carry out the method of claim 1, and thus, is rejected under similar rationale. Moreover, Watanabe teaches a processor-readable memory storing a program (Paragraphs 0032 and 0097; Fig. 1). While the instructions stored on the medium have been addressed by the prior art of record via the citations/rationale applied to claim 1, Applicant should be aware that addressing such limitations has only been made in the interest of compact prosecution. Specifically, under the BRI (broadest reasonable interpretation), the prior art need only address the computer-readable medium limitation to address the claim, because the BRI includes “instructions” as a series of steps (i.e., non-functional descriptive material). Applicant should amend the claims to specify that the instructions are, for example, computer or processor executable or “program instructions” in order to resolve this issue. See MPEP 2111.05 (III). Claims 2, 4, and 6 are rejected under 35 U.S.C. 103 as being unpatentable over Watanabe, et al. in view of Liu, et al. in view of Chen, et al. and further in view of Ibrahim, et al. ("Preprocessing technique in automatic speech recognition for human computer interaction: an overview," 2017). With respect to Claim 2, Watanabe in view of Liu and further in view of Chen teaches the encoder-decoder-based ASR method/system as applied to Claim 1. Although common in the ASR art for many years, claim 2 defines standard ASR pre-processing techniques that is not specifically described in the combination of Watanabe in view of Liu and further in view of Chen. Ibrahim, however, discloses these standard pre-processing techniques: generating a plurality of audio segments by segmenting the speech waveform using an audio segmentation technique (framing of the continuous stream of speech samples to facilitate block-wise processing of the signal, Section II(d), Pages 189-190); obtaining a plurality of speech segments from among the plurality of audio segments by removing non-speech signals from the plurality of audio segments (performing background/ambient noise removal on a series of frames, Section II(a), Page 187); generating a plurality spectrally flattened speech segments based on the plurality of speech segments using an audio flattening technique (pre-emphasis filtering that is a known pre-processing technique for spectral smoothing/flattening, Section II(c), Page 189); and generating the plurality of speech frames by windowing the plurality of speech segments to hamming window (windowing function applied to speech frames that may take the form of a Hamming window, Section II(e), Page 190). Watanabe, Liu, Chen, and Ibrahim are analogous art because they are from a similar field of endeavor in automatic speech recognition. Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date to perform the claimed sequence of pre-processing steps of frame/window-based processing for manageable data portions, removing noise to save processing on non-speech data, and compensating for spectral tilt of the speech data in the pre-processing described by Ibrahim in the ASR system taught by Watanabe in view of Liu and further in view of Chen to provide a predictable result of adjusting and modifying an input speech signal to make it acceptable for feature extraction (Ibrahim, Section II, Page 186). Claim 4 contains subject matter similar to claim 2, and thus, is rejected under similar rationale. Claim 6 contains subject matter similar to claim 2, and thus, is rejected under similar rationale. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: Wu, et al. (U.S. PG Publication: 2025/0022457 A1)- teaches the modification of speech data via an augmentation process by added random Gaussian noise (Paragraphs 0012 and 0028). Any inquiry concerning this communication or earlier communications from the examiner should be directed to JAMES S WOZNIAK whose telephone number is (571)272-7632. The examiner can normally be reached 7-3, off alternate Fridays. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant may use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Flanders can be reached at (571)272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. JAMES S. WOZNIAK Primary Examiner Art Unit 2655 /JAMES S WOZNIAK/Primary Examiner, Art Unit 2655
Read full office action

Prosecution Timeline

Dec 30, 2024
Application Filed
Aug 10, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12744028
GENERATING CONVERSATIONAL OUTPUT USING A LARGE LANGUAGE MODEL
2y 9m to grant Granted Sep 22, 2026
Patent 12738262
ENABLING TRAINING OF A MACHINE-LEARNING MODEL FOR TRIGGER-WORD DETECTION
3y 4m to grant Granted Sep 15, 2026
Patent 12738287
AUDIO CODING METHOD AND APPARATUS, AUDIO DECODING METHOD AND APPARATUS, ELECTRONIC DEVICE, COMPUTER-READABLE STORAGE MEDIUM, AND COMPUTER PROGRAM PRODUCT
2y 4m to grant Granted Sep 15, 2026
Patent 12725608
AUTOMATIC SPEECH RECOGNITION WITH MULTILINGUAL SCALABILITY AND LOW-RESOURCE ADAPTATION
2y 3m to grant Granted Sep 01, 2026
Patent 12718793
Sonifying Visual Content For Vision-Impaired Users
2y 10m to grant Granted Aug 25, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
59%
Grant Probability
99%
With Interview (+39.7%)
3y 7m (~1y 10m remaining)
Median Time to Grant
Low
PTA Risk
Based on 408 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month