Prosecution Insights
Last updated: August 17, 2026
Application No. 19/001,149

SYSTEM AND METHOD FOR AUTOMATED MULTI-SPEAKER AND MULTI-LINGUAL SPEECH ANALYSIS

Non-Final OA §102
Filed
Dec 24, 2024
Examiner
OPSASNICK, MICHAEL N
Art Unit
2658
Tech Center
2600 — Communications
Assignee
Eresearch Technology Inc.
OA Round
1 (Non-Final)
82%
Grant Probability
Favorable
1-2
OA Rounds
1y 6m
Est. Remaining
92%
With Interview

Examiner Intelligence

Grants 82% — above average
82%
Career Allowance Rate
753 granted / 919 resolved
+19.9% vs TC avg
Moderate +10% lift
Without
With
+10.1%
Interview Lift
resolved cases with interview
Typical timeline
3y 2m
Avg Prosecution
34 currently pending
Career history
961
Total Applications
across all art units

Statute-Specific Performance

§101
19.5%
-20.5% vs TC avg
§103
33.5%
-6.5% vs TC avg
§102
30.0%
-10.0% vs TC avg
§112
5.3%
-34.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 919 resolved cases

Office Action

§102
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Specification The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claim(s) 1-24 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Munoz et al (20240220737). As per claim 1, Munoz et al (20240220737) teaches a system for multi-lingual speech analysis, the system comprising: memory configured to store program code for performing speech analysis; a processor configured to execute the program code, and upon execution of the program code (as processors executing stored instructions in memory – para 0149, 0151 – processors, and 0152 – program code), the processor being configured to generate one or more application modules and at least one trained neural network which further configure the processor to (as using machine learned models for training, including neural networks – para 0029): receive, by the one or more application modules, a natural language input; decompose, by the one or more application modules, the natural language input into plural segments; accumulate, by the one or more application modules, a sub-group of the plural segments in a buffer, each segment representing a period during which voice activity is detected (as operating in segments that contain user speech, until the user has stopped speaking – para 0094 – “after the speaker has finished their speech”, and reflecting back on para 0012); analyze, by the one or more application modules, at least one sub-group of segments to determine whether the voice activity includes speech generated by plural speakers (as, segments (as speaker identity based in the voice activity – as speaker information – para 0083, and within the call – para 0087; and measuring when the speaker is speaking, so as to complete the translation after the speaker has finished – para 0094 – hence, VAD is used); generate, by the one or more application modules, one or more text segments from the at least one sub-group of audio segments based on the plural speaker determination (as, separating segments based on which speaker - see the translation technique in which there are 3 speakers of different languages and voices – para 0130, and into para 131 – 0134); generate, by the one or more application modules and one trained neural network, a semantic vector for each text segment and store each semantic vector in vector memory (as calculating and storing, word vectors with similar meaning in a semantic space – para 0031; using trained neural networks as part of LLM’s – para 0029); retrieve, by one or more application modules, relevant data associated with each semantic vector from the vector memory; and generate, by the at least one trained neural network, a response including specified information extracted from the one or more text segments based on at least the relevant data (as accessing relevant data associated with the semantic vector – see para 0032, wherein the groups of words are determined to be a commonly used phrase or collection of words; using the trained neural network – para 0029, in the prediction module predicts a word or sequence of words). As per claim 2, Munoz et al (20240220737) teaches the system according to claim 1, wherein the natural language input is a data file that includes streaming audio data (as the audio data is in the form of audio streams – para 0047). As per claim 3, Munoz et al (20240220737) teaches the system according to claim 1, wherein the natural language input includes streaming audio data (as, the natural language input included streamed audio – para 0047). As per claim 4, Munoz et al (20240220737) teaches the system according to claim 1, to decompose the natural language input into plural segments, the processor is further configured to: determine, by the one or more application modules, whether voice activity is present in each segment (as detecting speech via VAD – see para 0023). As per claim 5, Munoz et al (20240220737) teaches the system according to claim 4, wherein to generate at least one sub-group of segments from the plural segments, the processor is further configured to: identify each speaker that generates speech in the voice activity of the at least one sub-group of segments (as speaker identity based in the voice activity – as speaker information – para 0083, and within the call – para 0087; and measuring when the speaker is speaking, so as to complete the translation after the speaker has finished – para 0094 – hence, VAD is used). As per claim 6, Munoz et al (20240220737) teaches the system according to claim 5, wherein the at least one sub-group of segments includes at least one segment for each identified speaker (as, separating segments based on which speaker - see the translation technique in which there are 3 speakers of different languages and voices – para 0130, and into para 131 – 0134). As per claim 7, Munoz et al (20240220737) teaches the system according to claim 6, wherein to generate one or more text segments, the processor is further configured to: determine a source language of the speech included in each segment (as source language identifier – para 0096). As per claim 8, Munoz et al (20240220737) teaches the system according to claim 7, wherein the processor is further configured to: translate the one or more text segments from the source language determined for the associated segment to a target language selected by a user ( as translating from an input language to an output language; wherein the output/target language is selected by the user profile – para 0035). As per claim 9, Munoz et al (20240220737) teaches the system according to claim 8, wherein to extract specified data from the one or more text segments, the processor is further configured to: determine whether all speech in the one or more text segments has been translated and vectorized, prior to the structured data being extracted (as operating on previously translation data that is stored as a sequence of vectors – para 0040). As per claim 10, Munoz et al (20240220737) teaches the system according to claim 9, wherein the specified data is extracted from the one or more text segments based on predefined prompts, each predefined prompt being associated with specified data domain (as data extraction from the text segments are based on the use of large language models generating translation text – para 0035; wherein the prompts to the translation module takes text/word vectors as input – end of para 0035, which can be tied to certain topics/domains – each model may be specific to a certain language – para 0033). As per claim 11, Munoz et al (20240220737) teaches the system according to claim 10, wherein to retrieve relevant data associated with each semantic vector from the vector memory, the processor is further configured to: search the vector memory based on the specified data to identify the relevant data associated with the one or more text segments that is stored in the vector memory (as storing and searching, vector representations of the semantic meaning of the text segments/words – see para 0031). Claims 12-22 are method claims whose steps are performed by the system claims 1-11 above and as such, claims 12-22 are similar in scope and content to claims 1-11 above; therefore, claims 12-22 are rejected under similar rationale as presented against claims 1-11 above. Claim 23 is a non-transitory computer readable medium contain code when executed, performs the steps found throughout in claims 1-11 above and as such, claim 23 is similar in scope and content to these commonly found elements in claims 1-11 above; therefore, claim 23 is rejected under similar rationale as presented against claims 1-11 above. Further to claim 23, Munoz et al (20240220737) teaches a text summary of the input speech – see para 0086 which output the text derived from the enunciation module. Claim 24 is a system claim that performs the steps found throughout system claims 1-11 and as such, claim 24 is similar in scope and content to these common features in claims 1-11; therefore, claim 24 is rejected under similar rationale as presented against claims 1-11 above. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Please see related art listed on the PTO-892 form. Furthermore, the following prior art references were found to have common elements with applicants claims/spec: Wang et al (20240274122) teaches language translation and speaker ID/embedding using ASR language models (para 0031, 0034) Matusov et al (20200226327) teaches VAD, language translation, and speaker diarization (para 00070 Lee et al (20230186035) teaches the modeling of multi-speaker target speech derived from multiple speakers (para 0016). Any inquiry concerning this communication or earlier communications from the examiner should be directed to Michael Opsasnick, telephone number (571)272-7623, who is available Monday-Friday, 9am-5pm. If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Mr. Richemond Dorvil, can be reached at (571)272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). /Michael N Opsasnick/Primary Examiner, Art Unit 2658 07/31/2026
Read full office action

Prosecution Timeline

Dec 24, 2024
Application Filed
Aug 04, 2026
Non-Final Rejection mailed — §102 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12707212
SYSTEM FOR FACILITATING IN-PERSON INTERACTION BETWEEN MULTIUSER VIRTUAL ENVIRONMENT USERS WHOSE AVATARS HAVE INTERACTED VIRTUALLY
4y 4m to grant Granted Aug 11, 2026
Patent 12699201
Audio rendering of an electromagnetic metal detection signal
3y 9m to grant Granted Aug 04, 2026
Patent 12676164
System and Method for Modulation Domain-Based Audio Signal Encoding
3y 0m to grant Granted Jul 07, 2026
Patent 12658172
COMPUTING SYSTEM FOR UNSUPERVISED EMOTIONAL TEXT TO SPEECH TRAINING
4y 3m to grant Granted Jun 16, 2026
Patent 12651607
INFORMATION PROCESSING APPARATUS, INFORMATION PROCESSING METHOD, AND PROGRAM
2y 2m to grant Granted Jun 09, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
82%
Grant Probability
92%
With Interview (+10.1%)
3y 2m (~1y 6m remaining)
Median Time to Grant
Low
PTA Risk
Based on 919 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month