Prosecution Insights
Last updated: August 17, 2026
Application No. 18/431,690

SPEECH PROCESSING TECHNIQUE

Final Rejection §101§102§103
Filed
Feb 02, 2024
Examiner
OPSASNICK, MICHAEL N
Art Unit
2655
Tech Center
2600 — Communications
Assignee
NVIDIA Corporation
OA Round
2 (Final)
82%
Grant Probability
Favorable
3-4
OA Rounds
7m
Est. Remaining
92%
With Interview

Examiner Intelligence

Grants 82% — above average
82%
Career Allowance Rate
753 granted / 919 resolved
+19.9% vs TC avg
Moderate +10% lift
Without
With
+10.1%
Interview Lift
resolved cases with interview
Typical timeline
3y 2m
Avg Prosecution
34 currently pending
Career history
961
Total Applications
across all art units

Statute-Specific Performance

§101
19.5%
-20.5% vs TC avg
§103
33.5%
-6.5% vs TC avg
§102
30.0%
-10.0% vs TC avg
§112
5.3%
-34.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 919 resolved cases

Office Action

§101 §102 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 for being directed towards a patent ineligible mental process under the broadest reasonable interpretation (BRI). Independent Claims 1, 9, and 15 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claims regard a process that, as drafted under its broadest reasonable interpretation, covers performance of the limitations as a mental process, but for the recitation of generic computer components and a generic neural network, In regards to the illustrative process of claim 1 that is also included in claims 8 and 15, the claimed functionality could be practiced as a mental process under the BRI in the following manner: (a human could rely upon previous transcription content or recall audio that was previously/currently heard to transcribe a certain (prior, current, or future) portion of an audio signal using pen and paper). This judicial exception is not integrated into a practical application. In particular, the claim only recites two additional elements – the use of a generic neural network and generic computer components. The recitation of computer processors amount to no more than mere instructions to implement an otherwise abstract idea using generic computer. The claimed processor is not being improved as a tool since it is only being used as a tool to carry out an otherwise abstract mental process. The neural network is recited at a high-level of generality such that it amounts no more than mere instructions to apply the exception using a generic computer component. See MPEP 2106.05(f). MPEP 2106.05(f) provides the following considerations for determining whether a claim simply recites a judicial exception with the words “apply it” (or an equivalent), such as mere instructions to implement an abstract idea on a computer: (1) whether the claim recites only the idea of a solution or outcome i.e., the claim fails to recite details of how a solution to a problem is accomplished; (2) whether the claim invokes computers or other machinery merely as a tool to 8 perform an existing process; and (3) the particularity or generality of the application of the judicial exception. In the instant claim, the use of a neural network only presents the idea of a solution (i.e., transcribing text based upon features from different time periods) while failing to describe how the neural network is used or structured to achieve the solution. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The above identified additional generic computer components are no more than mere instructions to apply the exception using generic computer components that are well-known, routine, and conventional as is evidenced by Bancorp Services v. Sun Life (Fed. Cir. 2012) and Alice Corp. v. CLS Bank (2014). Further the use of a generic machine learning/neural networks does not amount to an inventive concept- see Recentive Analytics, Inc. v. Fox Corp. (Fed. Cir. April 18, 2025)- “Machine learning is a burgeoning and increasingly important field and may lead to patent-eligible improvements in technology. Today, we hold only that patents that do no more than claim the application of generic machine learning to new data environments, without disclosing improvements to the machine learning models to be applied, are patent ineligible under § 101.” As for evidence that the claimed neural network is well-known, routine, and conventional activity that does not direct patent ineligible subject matter to significantly more than the abstract idea, see the following prior art: Huang, et al. (U.S. PG Publication: 2023/0088915 A1- "neural network including a self-attention layer belongs to the conventional technology," Paragraph 0071), Jung, et al. (U.S. PG Publication: 2020/0193735 A1- deep learning including a recurrent neural network and attention mechanism are "known in the art," Paragraph 0059), and Whitenack, et al. (U.S. Patent: 12,198,689- neural networks with combinations of layers including self-attention and convolutional layers are "known", Col. 14, Lines 14-32). Accordingly, claims 1, 8, and 15 are not directed towards patent eligible subject matter under 35 U.S.C. 101. The remaining dependent claims fail to add patent eligible subject matter to their respective parent claims: Claims 2, 9-10, and 16-17 involve attention weighting as already addressed in the independent claims as not amounting to an inventive concept. Claims 3-4, 11, and 18 involve the identification/consideration of context that is a human process in identifying a word that was spoken and its semantic meaning along with the generic neural network that was addressed in the independent claims. Claims 5, 12, and 19 involve different neural network layers that were identified as being "known" in the independent claim rejection. If the applicant has improved such layers, the claim is devoid of the description of such improvements. Claims 6, 13, and 20 regard a decoder in a transformer that is evidenced as "known in the art" per Lu, et al. (U.S. PG Publication: 2023/0124006 A1). If the applicant has improved or has invented further specifics of this generic machine learning structure, these limitations have not been included in the claimed invention. Claims 7 and 14 regard a human process where an animator could draw an illustration or a series of illustrations to depict a character saying a particular transcription along with the generic neural networks addressed in the independent claims. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claims 1-6, 8-13, 15-20 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Gulati et al (20220207321). As per claim 1, Gulati et al (20220207321) teaches a processor, comprising: one or more circuits to cause one or more neural networks to generate text from an audio signal, wherein the one or more neural networks comprise (as processor based neural network computations for speech recognition – para 0024, using processors in para 0053): one or more convolutional portions (as convolutional blocks – para 0092, 0094); and one or more multi-head self-attention portions to each identify one or more features of a corresponding time period of the audio signal (as, using multihead self attention blocks – para 0094, as part of a conformer model – para 0094) to extract as different contextual representations according to different kernel sizes corresponding to respective portions of the one or more convolutional portions (as using the conformer models for speech to text recognition – para 0096; as applied to multiple applications – para 0085, with the use of a context manager – para 0083 and the subparts for the conformer model operate on differing subparts of the audio signal – e.g., spectrograph data representing the speech – para 0041, 0042; the convolution process alternates with a pointset convolution with an expansion factor – para 0100; and varying kernel sizes to affect convolution depth, within the conformer model – para 0123), wherein the one or more convolutional portions use the different contextual representation to generate a portion of the text corresponding to one or more other time periods of the audio signal (as varying the kernel size and altering the word-error-rate – pp 11, table 7, reflecting back on para 0123; examiner notes that in this application (ie, speech/word recognition), altering the kernel size which changes the error rate on word recognition and context. As per claim 2, Gulati et al (20220207321) teaches the processor of claim 1, wherein the one or more multi-head self-attention portions generate attention weights to indicate importance of the one or more features usable to generate the text corresponding to the one or more other time periods of the audio signal (as, using residual weight in the feed forward blocks in the conformer model – para 0094). As per claim 3, Gulati et al (20220207321) teaches the processor of claim 1, wherein the different contextual representations comprise local information and global information, respectively extracted by the one or more convolutional portions and one or more squeeze and excite portions of the one or more neural networks (as, the disclosed conformer models are a combination of self attention models and convolutional models – para 0007, which is further explained in Gulati et al, as getting the advantage of long range global context (self attention conformer) and the advantage of convolutional short term models, with reference to the squeeze and excitation technique – para 0005). As per claim 4, Gulati et al (20220207321) teaches wherein the different contextual representations comprise respective speaker utterances in the one or more other time periods of the audio signal (as, Gulati et al (20220207321) the use of the conformer models for a myriad of applications, such as speech recognition, text message, email, and dictation – see para 0082; examiner notes that it is old and notoriously well known to use transformer/CNN models for speaker identification – which is used in dictation applications to not only provide user security/privacy, but also for speaker-based voice characteristics; as an example, see Lee et al, 20210005183, equating performing speech tasks with speaker identification/verification, as well – see para 0026). As per claim 5, Gulati et al (20220207321) teaches the processor of claim 1, wherein the different kernel sizes of one or more convolution portions comprise a kernel size of 1 and a range linearly increasing kernel sizes greater than 1 (as kernel sizes range up to 65 – see para 0123). As per claim 6, Gulati et al (20220207321) teaches the processor of claim 1, wherein the text is generated by one or more decoder portions of the one or more neural networks, the one or more decoder portions comprising a transformer portion (as transformer block contains self-attention, convolutional and feed forward blocks – para 0037, which were identified as conformer blocks – para 0036, as equivalents to transformer/rnn’s). Claims 8-13 are system claims that perform the processor steps in claims 1-6, above and as such, claims 8-13 are similar in scope and content to claims 1-6; therefore, claims 8-13 are rejected under similar rationale as presented against claims 1-6 above. Claims 15-20 are method claims whose steps are performed by processor claims 1-6 above and as such, claims 15-20 are similar in scope and content to claims 1-6; therefore, claims 15-20 are rejected under similar rationale as presented against claims 1-6 above. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 7 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Gulati et al (20220207321). in view of Chen, et al. (“Transformer-S2A: Robust and Efficient Speech-to-Animation,” 2022). With respect to Claim 7, Gulati et al (20220207321) teaches the system for decoding speech to text as applied to Claim 1. Gulati et al (20220207321) does not explicitly teach the use of one or more decoders to generate one or more graphical representation of a character speaking the generated text as is set forth in claim 7. Chen, however, discloses: The processor of claim 1, wherein the one or more neural networks comprise one or more decoders to generate one or more graphical representation of a character speaking the generated text (speech-to-animation decoder (Fig. 2) that generates a facial animation (e.g., Figs. 3) in human-machine interactions speaking recognized words comprised of phonemes (see discussion of PPGs), Abstract; Sections 2.1-2.2, Pages 2-3). Gulati et al (20220207321) and Chen are analogous art because they are from a similar field of endeavor in speech understanding using transformer models. Thus, it would have been obvious to one of ordinary skill in the art, before the effective filing date, to apply the speech-to-animation decoder taught by Chen in the speech recognition/speech task described in Gulati et al (20220207321) (para 0079) to provide a predictable result of more immersive and engaging human-machine communications. Claim 14 contains subject matter similar to Claim 7, and thus, is rejected under similar rationale. Response to Arguments Applicant’s arguments with respect to the claims have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Examiner notes the replacement prior art, Gulati et al (20220207321), which addresses the newly presented claim features, such as, variable kernel size, and global/local model advantages, and the combination conformer blocks of Gulati et al (20220207321) teaching such. See new mappings above. As to applicants arguments against the 101 rejection, examiner notes that the Desjardins decision requires explicit detailed explanation of how the particular steps provide an improvement in the process. The generic recitation of “prone to failure and can be improved” does not provide enough detail of improvement provided by the claim limitations. Ex Parte Desjardins, MPEP2106.04 (d) 1). Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: Coucheiro Limeres (U.S. PG Publication: 2023/0360646 A1)- teaches an automatic speech recognition system for predicting a transcription based upon preceding/prefix weighting for an attention mechanism to ensure that a current token fits a current transcription context (Abstract; Paragraph 0058). Song et al (20200043508) teaching multi-head attention model using convolutional layers with varying kernel sizes (para 0022). Zhang, et al. (U.S. PG Publication: 2023/0244879 A1)- teaches a Transformer encoder-decoder structure having a self-attention layer that utilizes semantic context for use in speech recognition (Paragraph 0065). Any inquiry concerning this communication or earlier communications from the examiner should be directed to Michael Opsasnick, telephone number (571)272-7623, who is available Monday-Friday, 9am-5pm. If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Mr. Richemond Dorvil, can be reached at (571)272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). /Michael N Opsasnick/Primary Examiner, Art Unit 2658 08/01/2026
Read full office action

Prosecution Timeline

Feb 02, 2024
Application Filed
Oct 14, 2025
Non-Final Rejection mailed — §101, §102, §103
Mar 16, 2026
Response Filed
Aug 05, 2026
Final Rejection mailed — §101, §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12707212
SYSTEM FOR FACILITATING IN-PERSON INTERACTION BETWEEN MULTIUSER VIRTUAL ENVIRONMENT USERS WHOSE AVATARS HAVE INTERACTED VIRTUALLY
4y 4m to grant Granted Aug 11, 2026
Patent 12699201
Audio rendering of an electromagnetic metal detection signal
3y 9m to grant Granted Aug 04, 2026
Patent 12676164
System and Method for Modulation Domain-Based Audio Signal Encoding
3y 0m to grant Granted Jul 07, 2026
Patent 12658172
COMPUTING SYSTEM FOR UNSUPERVISED EMOTIONAL TEXT TO SPEECH TRAINING
4y 3m to grant Granted Jun 16, 2026
Patent 12651607
INFORMATION PROCESSING APPARATUS, INFORMATION PROCESSING METHOD, AND PROGRAM
2y 2m to grant Granted Jun 09, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
82%
Grant Probability
92%
With Interview (+10.1%)
3y 2m (~7m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 919 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month