Prosecution Insights
Last updated: August 15, 2026
Application No. 19/064,094

SYSTEMS AND METHODS FOR USING MACHINE-LEARNING TO EXTRACT AND PROCESS AUDIO DATA

Non-Final OA §101§102
Filed
Feb 26, 2025
Priority
Feb 27, 2024 — IN 202411014105 +1 more
Examiner
DESIR, PIERRE LOUIS
Art Unit
Tech Center
Assignee
Stats LLC
OA Round
1 (Non-Final)
61%
Grant Probability
Moderate
1-2
OA Rounds
2y 6m
Est. Remaining
94%
With Interview

Examiner Intelligence

Grants 61% of resolved cases
61%
Career Allowance Rate
178 granted / 291 resolved
+1.2% vs TC avg
Strong +33% interview lift
Without
With
+33.1%
Interview Lift
resolved cases with interview
Typical timeline
3y 11m
Avg Prosecution
7 currently pending
Career history
298
Total Applications
across all art units

Statute-Specific Performance

§101
12.5%
-27.5% vs TC avg
§103
49.7%
+9.7% vs TC avg
§102
19.0%
-21.0% vs TC avg
§112
12.2%
-27.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 291 resolved cases

Office Action

§101 §102
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. § 101 because the claimed invention is directed to a judicial exception without significantly more. This judicial exception is not integrated into a practical application because the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. Step 2A, Prong One: The claims are directed to an abstract idea The claims are directed to the abstract idea of mental processes, including receiving information, analyzing or transforming that information, and generating an output based on that information. More particularly, the claims recite receiving multimedia content, extracting audio, converting spoken language to text, translating or rephrasing the text into another language using a machine-learning model, and transmitting the resulting text or audio to a user interface. These limitations recite activities that may be performed in the human mind or with pen and paper, such as understanding speech, converting speech to text, and translating or rephrasing language. Accordingly, the claims are directed to an abstract idea. The claims focus on the manipulation of linguistic content and the presentation of that content, rather than on an improvement to computer functionality or another technological process. Step 2A, Prong Two: The claims do not integrate the abstract idea into a practical application The additional elements in the claims do not impose a meaningful limit on the abstract idea. The claims merely recite a “computing system,” “memory,” “processor,” “user interface,” and machine-learning models at a high level of generality. These elements are used in their ordinary capacity to perform generic data receiving, processing, and output operations. The dependent claims reciting closed captioning data, real-time receipt, stored retrieval, JSON/audio/video/text file formats, text-to-speech conversion, and merging translated audio with video similarly amount to routine data handling and presentation steps. These limitations do not reflect a specific technological improvement to speech processing, translation models, packet handling, or multimedia synchronization. Accordingly, the claims do not integrate the abstract idea into a practical application. Step 2B: The claims do not recite an inventive concept Individually and in combination, the additional claim elements do not amount to significantly more than the abstract idea itself. The recited computer components perform generic functions well-understood, routine, and conventional in the field, such as receiving data, converting data, and transmitting output. The machine-learning model limitations are recited functionally and do not describe a specific technical implementation or improvement. Therefore, the claims do not include an inventive concept sufficient to transform the abstract idea into patent-eligible subject matter. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claims 1-20 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Arkhangorodsky et al. (US 20230089902 A1) (Arkhangorodsky). Regarding claims 1 and 11, Arkhangorodsky discloses a system for extracting and processing audio data, the system comprising a memory storing instructions and a generative machine-learning model trained to translate first text data in a first language to a second language and generate second text data in the second language (see abstract and paragraph 92) and a processor operatively connected to the memory and configured to execute the instructions to perform operations (see abstract and paragraph 92) and method for extracting and processing audio data, the method comprising: receiving, by a computing system, one or more packets of multimedia content, wherein the one or more packets of multimedia content comprise audio data (i.e., “the system 100 may support verbal communications among multiple participants speaking different languages (see paragraph 37) “[0040]: “the end-user terminal 120 may input audio sequences into the server 110…” [0052]: “The audio data 304 can include raw, pre-mixed audio byte stream data or processed byte stream data.” [015]: “the receiving the first audio sequence comprises continuously receiving audio signals…”[0039]: “The backend interface 112 may then redirect the input to proper components for audio to text conversion…”; extracting, by the computing system, the audio data from the one or more packets of multimedia content (i.e., [0040]: “The backend interface 112 may then redirect the input to proper components for audio to text conversion…”[0042]: “the live caption view 124 may be configured to capture video conference audio, browser audio, system audio, or another form of audio input…”[0052]: “The caption module 210 is configured to analyze audio data, in raw form, as received (e.g., as a byte stream) by the audio mixer 208.”[0072]: “the audio mixer 208 can send the audio data 304 to the caption module 210..”, wherein the audio data comprises verbal speech in a first language (i.e., [0041]: “the ASR 114 processes the incoming audio sequences in a streaming mode, transcribes it in the speaker’s language (e.g., a first language)…” [0054]: “a participant’s speech in a first language is fed in a streaming fashion to the ASR subsystem…” [0037]: “speaker A speaking Spanish…” / “speaker B speaking English…”; converting, by the computing system, the audio data into first text data in the first language based on the verbal speech in the first language (i.e., [0041]: “the ASR 114 processes the incoming audio sequences in a streaming mode, transcribes it in the speaker’s language (e.g., a first language)” [0053]: “the backend services of ASR and MT subsystems are joined in a cascaded manner. A participant’s speech in a first language is fed in a streaming fashion to the ASR subsystem…” [0055]: “The ASR subsystem may include a single sequence-to-sequence ASR model…” [0082]: “the first generating component 530 may be configured to generate a first translated text in a second language by feeding the first audio sequence into a pipeline comprising an Automatic Speech Recognition (ASR) subsystem…”; providing, by the computing system, the first text data to a generative machine- learning model trained to translate the first text data in the first language to a second language and generate second text data in the second language (i.e., [0041]: “the ASR 114 processes the incoming audio sequences in a streaming mode, transcribes it in the speaker’s language (e.g., a first language) to be used as input to the MT 116 to decode in the listener’s language (e.g., a second language)” [0053]: “the output of the ASR subsystem may be fed into the MT subsystem for translation from the first language into a second language.” [0012]: “the MT subsystem comprises a multilingual neural machine translation model…” [0004]: “generating a first translated text in a second language by feeding the first audio sequence into a pipeline comprising an Automatic Speech Recognition (ASR) subsystem and a machine translation (MT) subsystem”; and transmitting, to a user interface by the computing system, the second text data in the second language (i.e., [0040]: “feed the output text back to the end-user terminal for displaying to the users.” [0042]: “the meeting view 122 may display closed captions in a user-selected language…” [0042]: “the live caption view 124 may be configured to… display closed captions in the user-selected language.” [0056]: “The output generated by the MT subsystem may be denoted as a first translated text in the second language…” [0082]: “the second displaying component 540 may be configured to display the first translated text on a second user interface…”). Regarding claims 2 and 12, Arkhangorodsky discloses a method and system as disclosed in claims 1 and 11 above, wherein the audio data is converted into first text data using closed captioning data included in the one or more packets of multimedia content (i.e., [0042]: “the meeting view 122 may display closed captions in a user-selected language…t the live caption view 124 may be configured to capture video conference audio, browser audio, system audio … display closed captions…” [0016]: “The displayed first translated text…” [0040]: “feed the output text back to the end-user terminal for displaying to the users” ”[0052]: “The caption module 210 is configured to analyze audio data, in raw form, as received (e.g., as a byte stream) by the audio mixer 208”). Regarding claims 3 and 13, Arkhangorodsky discloses a method and system as disclosed in claims 1 and 11, wherein the generative machine-learning model is an artificial intelligence model trained to translate text data into a plurality of languages (i.e., [0012]: “The MT subsystem comprises a multilingual neural machine translation model trained based on a joint set of corpora from a plurality of languages” [0013]: “The MT subsystem comprises a plurality of MT models respectively trained based on training samples from a plurality of languages.” [0055]: “the MT subsystem may include a single multilingual neural machine translation model trained based on a joint set of corpora from a plurality of languages…”). Regarding claims 4 and 14, Arkhangorodsky discloses a method and system as disclosed in claims 1 and 11, further comprising: converting, by the computing system, the second text data in the second language to translated audio data (i.e, [0037]: “The user interface may include a video display (e.g., displaying the live translation captions as text), an audio display (e.g., playing the live translation captions in audio).” [0042]: “the live caption view 124 may be configured to capture video conference audio…” [0037]: “the system 100 may also be used to provide live translation captions…”; and merging, by the computing system, the translated audio data with video data of the one or more packets of multimedia content (i.e., [0037]: “In other embodiments, the system 100 may also be used to provide live translation captions for audio inputs from other sources, such as movies…”[0036]: “the system 100 may include a server 110 side and an end-user terminal 120 side.”[0042]: “meeting view 122 may display closed captions in a user-selected language overlaid on the user’s video”). Regarding claims 5 and 15, Arkhangorodsky discloses a method and system as disclosed in claims 1 and 11, further comprising: providing, by the computing system, the first text data to a rephrasing machine- learning model trained to rephrase the first text data in the first language to one or more strings of text data in a second language and generate rephrased second text data in the second language (see paragraph 41 and 53); and transmitting, to a user interface by the computing system, the rephrased second text data in the second language (see paragraphs 40 and 42). Regarding claims 6 and 16, Arkhangorodsky discloses a method and system as disclosed in claims 1 and 11, wherein the one or more packets of multimedia content comprise at least one of audio data, video data, text data, story data, or live feed data (i.e., [0037]: “live translation captions for audio inputs from other sources, such as movies…”[0042]: “capture video conference audio, browser audio, system audio…”[0037]: “The user interface may include a video display…”[0015]: “continuously receiving audio signals… streaming the continuous audio signals into the pipeline…” also see [0044]). Regarding claim 7, Arkhangorodsky discloses a method as disclosed in claim 1, wherein the one or more packets of multimedia content are included within at least one of a JSON file, an audio file, a video file, a story file, or a text file (see [0039], [0040] and [0097]. Regarding claims 8 and 17, Arkhangorodsky discloses a method and system as disclosed in claims 1 and 11, wherein the one or more packets of multimedia content are received by the computing system in real time (i.e., [0015]: “receiving audio signals, and the generating the first translated text in the second language comprises streaming the continuous audio signals into the pipeline…”[0041]: “the ASR 114 processes the incoming audio sequences in a streaming mode…”[0054]: “the live translation captioning system 340 may implement translate-K method…”[0056]: “The output generated by the MT subsystem…”[0065]: “real-time matching score determination…”). Regarding claims 9 and 18, Arkhangorodsky discloses a method and system as disclosed in claims 1 and 11, wherein the one or more packets of multimedia content are stored in a data store (see [0038] and are retrieved by the computing system (see [0045] and [0078]). Regarding claims 10 and 19, Arkhangorodsky discloses a method and system as disclosed in claims 1 and 11 above, further comprising: providing, by the computing system, the first text data to a predictive machine- learning model trained to identify language patterns in the first text data in the first language (i.e., [0055]: “The ASR subsystem may include a single sequence-to-sequence ASR model…”[0010]: “The ASR subsystem comprises a sequence-to-sequence ASR model trained based on a joint set of corpora from a plurality of languages.”[0069]: “The audio-based multi-language classifier 450 may be trained based on a plurality of labeled audio clips in a plurality of languages”) and generate second text data in the second language based on the identified language patterns (i.e., [0041]: “transcribes it in the speaker’s language … to be used as input to the MT 116 to decode in the listener’s language…”[0053]: “the output of the ASR subsystem may be fed into the MT subsystem for translation from the first language into a second language.”[0012]: “multilingual neural machine translation model…”); and transmitting, to a user interface by the computing system, the second text data in the second language (i.e., [0040]: “feed the output text back to the end-user terminal for displaying to the users.” [0056]: “The output generated by the MT subsystem may be denoted as a first translated text in the second language…” [0082]: “the second displaying component 540 may be configured to display the first translated text on a second user interface…”). Regarding claim 20, Arkhangorodsky discloses a method for extracting and processing audio data, the method comprising: receiving, by a computing system, one or more packets of multimedia content, wherein the one or more packets of multimedia content comprise audio data (i.e., “the system 100 may receive audio input from speaker A speaking Spanish” ([0037])…“The server 110 and the end-user terminal 120 may communicate with each other over the internet, through a local network (e.g., LAN), or through direct communication” ([0039])..“the end-user terminal 120 may input audio sequences into the server 110 by triggering corresponding Application Programming Interfaces (APIs) in the backend interface 112” [0040]); extracting, by the computing system, the audio data from the one or more packets of multimedia content, wherein the audio data comprises verbal speech in a first language (i.e., “the system 100 may receive audio input from speaker A speaking Spanish, transcribe the audio input into text in Spanish” ([0037])…“the ASR 114 processes the incoming audio sequences in a streaming mode, transcribes it in the speaker’s language (e.g., a first language)” ([0041])…“At step 306, the verbal description from the first user in the first language, denoted as a first audio sequence, may be received by the live translation captioning system 340.” [0052]); converting, by the computing system, the audio data into first text data in the first language based on the verbal speech in the first language (i.e., “the ASR 114 processes the incoming audio sequences in a streaming mode, transcribes it in the speaker’s language (e.g., a first language) to be used as input to the MT 116” ([0041])…“generating a first text sequence by feeding the first audio sequence into the ASR subsystem” ([0005])…“the ASR subsystem may transcribe every new extended source sentence in the first language as the speaker speaks” [0053]); providing, by the computing system, the first text data to a rephrasing machine-learning model trained to rephrase the first text data in the first language to one or more strings of text data in a second language and generate rephrased second text data in the second language (i.e., “generating the first translated text in the second language by feeding the first text sequence into the MT subsystem corresponding to the second language” ([0005])…“the MT subsystem comprises a multilingual neural machine translation model trained based on a joint set of corpora from a plurality of languages” ([0012])…“the output of the ASR subsystem may be fed into the MT subsystem for translation from the first language into a second language” ([0053])…“At step 308, the first audio sequence may be fed into a natural language processing (NLP) pipeline for audio-to-text conversion and text-to-text translation.” [0053]); and transmitting, to a user interface by the computing system, the rephrased second text data in the second language (i.e., “translate the Spanish text into English and display the English text for speaker B speaking English” ([0037])…“At step 310, the first translated text may be displayed to the second user.” ([0056])…“the live caption view 124 may be configured to capture video conference audio, browser audio, system audio, or another form of audio input, and display closed captions in the user-selected language” ([0042])…“The backend interface 112 may then redirect the input to proper components for audio to text conversion and subsequent machine translation, and then feed the output text back to the end-user terminal for displaying to the users.” [0040]). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to PIERRE LOUIS DESIR whose telephone number is (571)272-7799. The examiner can normally be reached Monday-Friday 9AM-5:30PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /PIERRE LOUIS DESIR/Supervisory Patent Examiner, Art Unit 2659
Read full office action

Prosecution Timeline

Feb 26, 2025
Application Filed
Jul 29, 2026
Non-Final Rejection mailed — §101, §102 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12694863
MODIFYING AN AUDIO SIGNAL TO INCORPORATE A NATURAL-SOUNDING INTONATION
3y 2m to grant Granted Jul 28, 2026
Patent 12670323
Golden Prompt Generation based on Authoritative Publications
2y 9m to grant Granted Jun 30, 2026
Patent 12651126
INTENT DISCOVERY USING NATURAL LANGUAGE INFERENCE
3y 3m to grant Granted Jun 09, 2026
Patent 12632788
PROMPT AUGMENTED GENERATIVE REPLAY VIA SUPERVISED CONTRASTIVE TRAINING FOR LIFELONG INTENT DETECTION
2y 10m to grant Granted May 19, 2026
Patent 12609124
VOICE AGENT SYSTEM
2y 8m to grant Granted Apr 21, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
61%
Grant Probability
94%
With Interview (+33.1%)
3y 11m (~2y 6m remaining)
Median Time to Grant
Low
PTA Risk
Based on 291 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month