Prosecution Insights
Last updated: October 02, 2026
Application No. 18/969,060

STREAMING LANGUAGE AI SYSTEMS WITH AUDIO INTEGRATION

Non-Final OA §101§103
Filed
Dec 04, 2024
Examiner
BLANKENAGEL, BRYAN S
Art Unit
2658
Tech Center
2600 — Communications
Assignee
NVIDIA Corporation
OA Round
1 (Non-Final)
67%
Grant Probability
Favorable
1-2
OA Rounds
11m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 67% — above average
67%
Career Allowance Rate
262 granted / 390 resolved
+5.2% vs TC avg
Strong +33% interview lift
Without
With
+33.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 8m
Avg Prosecution
23 currently pending
Career history
418
Total Applications
across all art units

Statute-Specific Performance

§101
25.1%
-14.9% vs TC avg
§103
50.7%
+10.7% vs TC avg
§102
11.8%
-28.2% vs TC avg
§112
7.4%
-32.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 390 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-11 and 13-20 rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception (i.e., a law of nature, a natural phenomenon, or an abstract idea) without significantly more. Using the subject matter eligibility test from page 74621 of the Federal Register Notice titled “2014 Interim Guidance on Patent Subject Matter Eligibility,” a two-step process is performed. Under step 1, the claims are analyzed to determine if the claim is directed to a process, machine, article of manufacture, or composition of matter. In this case, claims 1-12 are directed to a method, which is a process; claims 13-20 are directed to a system, which is a machine or an article of manufacture. Step 2A (part 1 of the Mayo test), using the guidance from pages 50-57 of the Federal Register Vol. 84 No. 4 from Monday, January 7, 2019, requires applying a two-prong inquiry. In Prong One, examiners evaluate whether the claim recites a judicial exception, determining if the claim is directed to a law of nature, a natural phenomenon, or an abstract idea. In this case, claims 1 recites predicting text tokens, updating embeddings, obtaining cross-attention states, and generating a streaming text output, which are mental processes and mathematical calculations. In Prong Two, examiners evaluate whether the judicial exception is integrated into a practical application that imposes a meaningful limit on the judicial exception. In this case, additional limitations of providing and receiving data to/from a language model are mere extrasolution activity, while elements such as processor, cross-modality network, and language model are generic computing components, and do not integrate the abstract ideas into a practical application. Step 2B (part 2 of the Mayo test) requires analyzing the claims to determine if they recite additional elements that amount to significantly more than the judicial exception. In this case, the claims do not include additional elements that are sufficient to amount to significantly more than the abstract idea itself. Regarding claims 1, 13, and 20, predicting text tokens, updating embeddings, obtaining cross-attention states, and generating a streaming text output are mental processes and mathematical calculations. For example, a human could predict text in a streaming manner from received audio, while performing the processing and updating are mathematical calculations. Additional limitations of providing and receiving data to/from a language model are mere extrasolution activity, while elements such as processor, cross-modality network, and language model are generic computing components, and do not integrate the abstract ideas into a practical application or constitute significantly more. Regarding claims 2-3, 5, 7, 9-11, 14-15, and 18, the limitations are a further clarification of the above abstract ideas. Regarding claims 4 and 16, computing a score, comparing scores, and including or rejecting text tokens are mental processes and mathematical calculations, which are abstract ideas without integration into a practical application and without significantly more. Regarding claim 6, removing an embedding is a mental process or mathematical calculation, which are abstract ideas without integration into a practical application and without significantly more. Regarding claims 8 and 17, computing keys, values, weights, and performing weighting are mathematical calculations, which is an abstract idea, while obtaining a query is mere extrasolution activity, which does not integrate the abstract idea into a practical application or constitute significantly more. Regarding claim 19, the limitations include generic computing components, which do not integrate the abstract ideas into a practical application or constitute significantly more. The limitations of the claims, taken alone, do not amount to significantly more than the above-identified judicial exception (the abstract idea). Looking at the limitations as an ordered combination adds nothing that is not already present when looking at the elements individually. Applicable case law cited in the Federal Register includes, but is not limited to: Alice Corp., 134 S. Ct. at 2355-56, Digitech Image Tech., LLC v. Electronics for Imaging, Inc., 758 F.3d 1344 (Fed. Cir. 2014), Benson, 409 U.S. at 63. See "Preliminary Examination Instructions in view of the Supreme Court Decision in Alice Corporation Pty. Ltd. v. CLS Bank International, et al.," dated June 25, 2014, and the Federal Register notice titled "2014 Interim Guidance on Patent Subject Matter Eligibility" (79 FR 74618). Regarding claim 12, Examiner notes that the modifying parameters of models based on a training output evaluation does not appear to be a mental process or a mathematical calculation, and therefore claim 12 is not rejected under 35 U.S.C. 101. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1, 3, 6-7, 9-13, 15, 18, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Panagopoulou et al. (US 2024/0370718 A1), hereinafter referred to as Panagopoulou, in view of Liu et al. (US 12,531,056 B1), hereinafter referred to as Liu. Regarding claim 1, Panagopoulou teaches: A method comprising: updating a plurality of audio embeddings with one or more audio embeddings representative of the streaming audio input during the respective time interval (Fig. 1A element 104, para [0027], where an audio encoder extracts audio features); processing, using a cross-modality network, the plurality of audio embeddings and a plurality of text embeddings representative of a text input associated with the streaming audio input to obtain a plurality of cross-attention states (Fig. 1A element 108, para [0027], where the multimodal encoder processes the audio feature embedding and the text feature vector from the instruction input to generate a vector representation of the input); providing, to a language model (LM), a prompt comprising a plurality of output embeddings obtained based at least on the plurality of cross-attention states (Fig. 1A element 122, para [0027-28], where the input representation and instruction are used to prompt a language model); and receiving, from the LM, a text token, of the plurality of text tokens, predicted for the respective time interval (Fig. 1A element 122, 124, para [0028], where the language model generates an output text based on the prompt); and generating, using the plurality of text tokens, the streaming text output (Fig. 1A element 122, 124, para [0028], where the language model generates an output text based on the prompt). Panagopoulou does not teach: predicting, over a plurality of iterations, a plurality of text tokens of a streaming text output associated with a streaming audio input, an individual iteration of the plurality of iterations being associated with a respective time interval of a plurality of time intervals and comprising: Liu teaches: predicting, over a plurality of iterations, a plurality of text tokens of a streaming text output associated with a streaming audio input, an individual iteration of the plurality of iterations being associated with a respective time interval of a plurality of time intervals (col. 4 lines 23-67, where the audio input is received in a streaming fashion such as in frames of 10 ms each, where the frames are processed as they are available, and col. 5 lines 40-50, where the predicted text tokens are output) and comprising: It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Panagopoulou by using the iteration of Liu (Liu col. 4 lines 23-67) for the token generation process of Panagopoulou (Panagopoulou Fig. 1A) by performing the process frame by frame, so that the audio may be processed in a streaming fashion (Liu col. 4 lines 23-35). Regarding claim 3, Panagopoulou in view of Liu teaches: The method of claim 1, wherein the generating the streaming text output comprises: including the text token into the streaming text output (Liu col. 5 lines 40-50, where the predicted text tokens are output). Regarding claim 6, Panagopoulou in view of Liu teaches: The method of claim 1, wherein the updating the plurality of audio embeddings comprises: removing one or more oldest audio embeddings from the plurality of audio embeddings (Liu col. 4 lines 63-67, where frames are processed as they are available). Regarding claim 7, Panagopoulou in view of Liu teaches: The method of claim 1, wherein the cross-modality network comprises one or more transformer blocks (Panagopoulou para [0026], where the encoder comprises a transformer structure). Regarding claim 9, Panagopoulou in view of Liu teaches: The method of claim 1, wherein the text input comprises at least one of: a text context for the streaming audio input (Panagopoulou para [0027], where the text input is an instruction), or one or more previously predicted text tokens (Liu col. 5 lines 24-39, where previously predicted token data is used). Regarding claim 10, Panagopoulou in view of Liu teaches: The method of claim 1, wherein the streaming text output comprises at least one of: a conversational response to the streaming audio input (Panagopoulou para [0025], where the output is a response to the input), a transcription or diarization of the streaming audio input (Liu Fig. 1 element 150, col. 4 lines 63-67, where ASR is performed), or a translation of the streaming audio input (where another limitation is chosen). Regarding claim 11, Panagopoulou in view of Liu teaches: The method of claim 1, wherein the one or more audio embeddings are generated using a speech model (Panagopoulou para [0027], where an audio encoder generates an input embedding, where an audio encoder is a speech model as in para [0030] of Applicant's specification). Regarding claim 12, Panagopoulou in view of Liu teaches: The method of claim 1, further comprising: obtaining a training input (Panagopoulou para [0029], where a training input is used), wherein the training input comprises: a first portion comprising a training audio input (Panagopoulou para [0025], [0029], where a training input is used, such as audio), and a second portion comprising a training text context for the training audio input (Panagopoulou para [0029], where a training input includes an instruction associated with the input); processing, using a speech model, the first portion to generate a plurality of training audio embeddings (Panagopoulou para [0027], where an audio encoder generates an input embedding, where an audio encoder is a speech model as in para [0030] of Applicant's specification); processing, using the cross-modality network, the training text context and the plurality of training audio embeddings to generate a training prompt to the LM (Panagopoulou para [0027], [0029-30], where the encoder outputs vector representations for input to the LM); obtaining a training output generated by the LM in response to the training prompt (Panagopoulou para [0030], where the LM outputs text); and modifying, based at least on an evaluation of the training output, one or more parameters of at least one of the speech model, the cross-modality network, or an adapter neural network (Panagopoulou para [0030], where the cross-modality network is updated). Regarding claim 13, Panagopoulou teaches: A system comprising: one or more processors (Panagopoulou Fig. 4A element 410, para [0052], where processors are used) to: updating a plurality of audio embeddings with one or more audio embeddings representative of the streaming audio input during the respective time interval (Fig. 1A element 104, para [0027], where an audio encoder extracts audio features); processing, using a cross-modality network, the plurality of audio embeddings and a plurality of text embeddings representative of a text input associated with the streaming audio input to obtain a plurality of cross-attention states (Fig. 1A element 108, para [0027], where the multimodal encoder processes the audio feature embedding and the text feature vector from the instruction input to generate a vector representation of the input); providing, to a language model (LM), a prompt comprising a plurality of output embeddings obtained based at least on the plurality of cross-attention states (Fig. 1A element 122, para [0027-28], where the input representation and instruction are used to prompt a language model); and receiving, from the LM, a text token, of the plurality of text tokens, predicted for the respective time interval (Fig. 1A element 122, 124, para [0028], where the language model generates an output text based on the prompt); and generate, using the plurality of text tokens, the streaming text output (Fig. 1A element 122, 124, para [0028], where the language model generates an output text based on the prompt). Panagopoulou does not teach: predict, over a plurality of iterations, a plurality of text tokens of a streaming text output associated with a streaming audio input, an individual iteration of the plurality of iterations being associated with a respective time interval of a plurality of time intervals and comprising: Liu teaches: predict, over a plurality of iterations, a plurality of text tokens of a streaming text output associated with a streaming audio input, an individual iteration of the plurality of iterations being associated with a respective time interval of a plurality of time intervals (col. 4 lines 23-67, where the audio input is received in a streaming fashion such as in frames of 10 ms each, where the frames are processed as they are available, and col. 5 lines 40-50, where the predicted text tokens are output) and comprising: It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Panagopoulou by using the iteration of Liu (Liu col. 4 lines 23-67) for the token generation process of Panagopoulou (Panagopoulou Fig. 1A) by performing the process frame by frame, so that the audio may be processed in a streaming fashion (Liu col. 4 lines 23-35). Regarding claim 15, Panagopoulou in view of Liu teaches: The system of claim 13, wherein to generate the streaming text output, one or more processors are to: include the text token into the streaming text output (Liu col. 5 lines 40-50, where the predicted text tokens are output). Regarding claim 18, Panagopoulou in view of Liu teaches: The system of claim 13, wherein the streaming text output comprises at least one of: a conversational response to the streaming audio input (Panagopoulou para [0025], where the output is a response to the input), a transcription of the streaming audio input (Liu Fig. 1 element 150, col. 4 lines 63-67, where ASR is performed), or a translation of the streaming audio input (where another limitation is chosen). Regarding claim 20, Panagopoulou teaches: based at least on a language model processing a prompt (Fig. 1A element 122, 124, para [0027-28], where the input representation and instruction are used to prompt a language model, and where the language model generates an output text based on the prompt), the prompt generated based at least on one or more computed cross-attention scores between one or more units of the streaming speech input and one or more units of the streaming text output (Fig. 1A element 108, para [0027], where the multimodal encoder processes the audio feature embedding and the text feature vector from the instruction input to generate a vector representation of the input). Panagopoulou does not teach: one or more processors to iteratively generate a streaming text output for a streaming speech input, Liu teaches: one or more processors to iteratively generate a streaming text output for a streaming speech input (col. 4 lines 23-67, where the audio input is received in a streaming fashion such as in frames of 10 ms each, where the frames are processed as they are available, and col. 5 lines 40-50, where the predicted text tokens are output) It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Panagopoulou by using the iteration of Liu (Liu col. 4 lines 23-67) for the token generation process of Panagopoulou (Panagopoulou Fig. 1A) by performing the process frame by frame, so that the audio may be processed in a streaming fashion (Liu col. 4 lines 23-35). Claim(s) 2, 8, 14, and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Panagopoulou, in view of Liu, and further in view of Shin. Regarding claim 2, Panagopoulou in view of Liu teaches: The method of claim 1, Panagopoulou in view of Liu does not teach: wherein the one or more audio embeddings for a first iteration of the plurality of iterations comprise more audio embeddings than the one or more audio embeddings for a second iteration of the plurality of iterations. Shin teaches: wherein the one or more audio embeddings for a first iteration of the plurality of iterations comprise more audio embeddings than the one or more audio embeddings for a second iteration of the plurality of iterations (para [0083], where one or more audio embeddings are generated for received audio data, where one iteration may contain one audio embedding and another may contain more than one audio embedding). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Panagopoulou in view of Liu by using the cross-attention calculations of Shin (Shin para [0072]) using the embeddings of Panagopoulou in view of Liu (Panagopoulou para [0027]), in order to combine asymmetrically the two separate embedding sequences (Shin para [0005]). Regarding claim 8, Panagopoulou in view of Liu teaches: The method of claim 1, wherein an individual cross-attention state of the plurality of cross-attention states is computed, at least in part, by: Panagopoulou in view of Liu does not teach: obtaining a query associated with an individual text embedding of the plurality of text embeddings; computing a plurality of keys and a plurality of values, an individual key of the plurality of keys and an individual value of the plurality of values computed using a corresponding audio embedding of the plurality of audio embeddings; computing a plurality of weights, wherein an individual weight of the plurality of weights is computed using the query and a corresponding key of the plurality of keys; and weighting, using the plurality of weights, the plurality of values to obtain the individual cross-attention state. Shin teaches: obtaining a query associated with an individual text embedding of the plurality of text embeddings (para [0072], where features computed from the text embeddings are mapped to queries); computing a plurality of keys and a plurality of values, an individual key of the plurality of keys and an individual value of the plurality of values computed using a corresponding audio embedding of the plurality of audio embeddings (para [0072], where the audio embeddings are mapped into keys and values); computing a plurality of weights, wherein an individual weight of the plurality of weights is computed using the query and a corresponding key of the plurality of keys (para [0072], where a dot product between queries and keys is performed to acquire the weights); and weighting, using the plurality of weights, the plurality of values to obtain the individual cross-attention state (para [0072], where the values are multiplied by the weights to compute the model output). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Panagopoulou in view of Liu by using the cross-attention calculations of Shin (Shin para [0072]) using the embeddings of Panagopoulou in view of Liu (Panagopoulou para [0027]), in order to combine asymmetrically the two separate embedding sequences (Shin para [0005]). Regarding claim 14, Panagopoulou in view of Liu teaches: The system of claim 13, Panagopoulou in view of Liu does not teach: wherein the one or more audio embeddings for a first iteration of the plurality of iterations comprise more audio embeddings than the one or more audio embeddings for a second iteration of the plurality of iterations. Shin teaches: wherein the one or more audio embeddings for a first iteration of the plurality of iterations comprise more audio embeddings than the one or more audio embeddings for a second iteration of the plurality of iterations (para [0083], where one or more audio embeddings are generated for received audio data, where one iteration may contain one audio embedding and another may contain more than one audio embedding). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Panagopoulou in view of Liu by using the cross-attention calculations of Shin (Shin para [0072]) using the embeddings of Panagopoulou in view of Liu (Panagopoulou para [0027]), in order to combine asymmetrically the two separate embedding sequences (Shin para [0005]). Regarding claim 17, Panagopoulou in view of Liu teaches: The system of claim 13, wherein to obtain an individual cross-attention state of the plurality of cross-attention states, the one or more processors are to: Panagopoulou in view of Liu does not teach: obtain a query associated with an individual text embedding of the plurality of text embeddings; compute a plurality of keys and a plurality of values, an individual key of the plurality of keys and an individual value of the plurality of values computed using a corresponding audio embedding of the plurality of audio embeddings; compute a plurality of weights, wherein an individual weight of the plurality of weights is computed using the query and a corresponding key of the plurality of keys; and weight, using the plurality of weights, the plurality of values to obtain the individual cross-attention state. Shin teaches: obtain a query associated with an individual text embedding of the plurality of text embeddings (para [0072], where features computed from the text embeddings are mapped to queries); compute a plurality of keys and a plurality of values, an individual key of the plurality of keys and an individual value of the plurality of values computed using a corresponding audio embedding of the plurality of audio embeddings (para [0072], where the audio embeddings are mapped into keys and values); compute a plurality of weights, wherein an individual weight of the plurality of weights is computed using the query and a corresponding key of the plurality of keys (para [0072], where a dot product between queries and keys is performed to acquire the weights); and weight, using the plurality of weights, the plurality of values to obtain the individual cross-attention state (para [0072], where the values are multiplied by the weights to compute the model output). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Panagopoulou in view of Liu by using the cross-attention calculations of Shin (Shin para [0072]) using the embeddings of Panagopoulou in view of Liu (Panagopoulou para [0027]), in order to combine asymmetrically the two separate embedding sequences (Shin para [0005]). Allowable Subject Matter Claims 4-5, 16, and 19 would be allowable if rewritten to overcome the rejection(s) under 35 U.S.C. 101, set forth in this Office action and to include all of the limitations of the base claim and any intervening claims. The following is a statement of reasons for the indication of allowable subject matter: the closest prior art of Panagopoulou, Liu, and Shin do not teach the limitations of the claims. Specifically, none of the cited prior art teaches computing an attention score between the text token output from the language model with a subset of the audio embeddings from the streaming audio input, comparing the score with a second attention score, and determining whether to include or reject the text token from the streaming text output, in combination with the other limitations. Hence, none of the cited prior art, either alone or in combination thereof, teaches the combination of limitations found in the claims. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. US 12,548,559 B1 col. 6 lines 1-24 teaches performing cross-attention on speech and text data while performing a machine translation task. Any inquiry concerning this communication or earlier communications from the examiner should be directed to BRYAN S BLANKENAGEL whose telephone number is (571)270-0685. The examiner can normally be reached 8:00am-5:30pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richemond Dorvil can be reached at 571-272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /BRYAN S BLANKENAGEL/Primary Examiner, Art Unit 2658
Read full office action

Prosecution Timeline

Dec 04, 2024
Application Filed
Jun 29, 2026
Non-Final Rejection mailed — §101, §103
Sep 15, 2026
Examiner Interview Summary
Sep 15, 2026
Applicant Interview (Telephonic)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12724962
PRE-TRAINING LANGUAGE MODELS USING NATURAL LANGUAGE EXPRESSIONS EXTRACTED FROM STRUCTURED DATABASES
4y 3m to grant Granted Sep 01, 2026
Patent 12718826
AUDIO SIGNAL ENCODING AND DECODING METHOD AND APPARATUS
2y 7m to grant Granted Aug 25, 2026
Patent 12711947
METHOD FOR TRAINING A NEURAL NETWORK AND A DATA PROCESSING DEVICE
2y 8m to grant Granted Aug 18, 2026
Patent 12711969
METHODS AND APPARATUS FOR SUPPLEMENTING PARTIALLY READABLE AND/OR INACCURATE CODES IN MEDIA
2y 2m to grant Granted Aug 18, 2026
Patent 12711983
METHOD OF DETECTING SPEECH AND SPEECH DETECTOR FOR LOW SIGNAL-TO-NOISE RATIOS
2y 1m to grant Granted Aug 18, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
67%
Grant Probability
99%
With Interview (+33.3%)
2y 8m (~11m remaining)
Median Time to Grant
Low
PTA Risk
Based on 390 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month