Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Specification
The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claim(s) 1-24 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Munoz et al (20240220737).
As per claim 1, Munoz et al (20240220737) teaches a system for multi-lingual speech analysis, the system comprising:
memory configured to store program code for performing speech analysis; a processor configured to execute the program code, and upon execution of the program code (as processors executing stored instructions in memory – para 0149, 0151 – processors, and 0152 – program code),
the processor being configured to generate one or more application modules and at least one trained neural network which further configure the processor to (as using machine learned models for training, including neural networks – para 0029):
receive, by the one or more application modules, a natural language input; decompose, by the one or more application modules, the natural language input into plural segments; accumulate, by the one or more application modules, a sub-group of the plural segments in a buffer, each segment representing a period during which voice activity is detected (as operating in segments that contain user speech, until the user has stopped speaking – para 0094 – “after the speaker has finished their speech”, and reflecting back on para 0012);
analyze, by the one or more application modules, at least one sub-group of segments to determine whether the voice activity includes speech generated by plural speakers (as, segments (as speaker identity based in the voice activity – as speaker information – para 0083, and within the call – para 0087; and measuring when the speaker is speaking, so as to complete the translation after the speaker has finished – para 0094 – hence, VAD is used);
generate, by the one or more application modules, one or more text segments from the at least one sub-group of audio segments based on the plural speaker determination (as, separating segments based on which speaker - see the translation technique in which there are 3 speakers of different languages and voices – para 0130, and into para 131 – 0134);
generate, by the one or more application modules and one trained neural network, a semantic vector for each text segment and store each semantic vector in vector memory (as calculating and storing, word vectors with similar meaning in a semantic space – para 0031; using trained neural networks as part of LLM’s – para 0029);
retrieve, by one or more application modules, relevant data associated with each semantic vector from the vector memory; and generate, by the at least one trained neural network, a response including specified information extracted from the one or more text segments based on at least the relevant data (as accessing relevant data associated with the semantic vector – see para 0032, wherein the groups of words are determined to be a commonly used phrase or collection of words; using the trained neural network – para 0029, in the prediction module predicts a word or sequence of words).
As per claim 2, Munoz et al (20240220737) teaches the system according to claim 1, wherein the natural language input is a data file that includes streaming audio data (as the audio data is in the form of audio streams – para 0047).
As per claim 3, Munoz et al (20240220737) teaches the system according to claim 1, wherein the natural language input includes streaming audio data (as, the natural language input included streamed audio – para 0047).
As per claim 4, Munoz et al (20240220737) teaches the system according to claim 1, to decompose the natural language input into plural segments, the processor is further configured to: determine, by the one or more application modules, whether voice activity is present in each segment (as detecting speech via VAD – see para 0023).
As per claim 5, Munoz et al (20240220737) teaches the system according to claim 4, wherein to generate at least one sub-group of segments from the plural segments, the processor is further configured to: identify each speaker that generates speech in the voice activity of the at least one sub-group of segments (as speaker identity based in the voice activity – as speaker information – para 0083, and within the call – para 0087; and measuring when the speaker is speaking, so as to complete the translation after the speaker has finished – para 0094 – hence, VAD is used).
As per claim 6, Munoz et al (20240220737) teaches the system according to claim 5, wherein the at least one sub-group of segments includes at least one segment for each identified speaker (as, separating segments based on which speaker - see the translation technique in which there are 3 speakers of different languages and voices – para 0130, and into para 131 – 0134).
As per claim 7, Munoz et al (20240220737) teaches the system according to claim 6, wherein to generate one or more text segments, the processor is further configured to: determine a source language of the speech included in each segment (as source language identifier – para 0096).
As per claim 8, Munoz et al (20240220737) teaches the system according to claim 7, wherein the processor is further configured to: translate the one or more text segments from the source language determined for the associated segment to a target language selected by a user ( as translating from an input language to an output language; wherein the output/target language is selected by the user profile – para 0035).
As per claim 9, Munoz et al (20240220737) teaches the system according to claim 8, wherein to extract specified data from the one or more text segments, the processor is further configured to: determine whether all speech in the one or more text segments has been translated and vectorized, prior to the structured data being extracted (as operating on previously translation data that is stored as a sequence of vectors – para 0040).
As per claim 10, Munoz et al (20240220737) teaches the system according to claim 9, wherein the specified data is extracted from the one or more text segments based on predefined prompts, each predefined prompt being associated with specified data domain (as data extraction from the text segments are based on the use of large language models generating translation text – para 0035; wherein the prompts to the translation module takes text/word vectors as input – end of para 0035, which can be tied to certain topics/domains – each model may be specific to a certain language – para 0033).
As per claim 11, Munoz et al (20240220737) teaches the system according to claim 10, wherein to retrieve relevant data associated with each semantic vector from the vector memory, the processor is further configured to: search the vector memory based on the specified data to identify the relevant data associated with the one or more text segments that is stored in the vector memory (as storing and searching, vector representations of the semantic meaning of the text segments/words – see para 0031).
Claims 12-22 are method claims whose steps are performed by the system claims 1-11 above and as such, claims 12-22 are similar in scope and content to claims 1-11 above; therefore, claims 12-22 are rejected under similar rationale as presented against claims 1-11 above.
Claim 23 is a non-transitory computer readable medium contain code when executed, performs the steps found throughout in claims 1-11 above and as such, claim 23 is similar in scope and content to these commonly found elements in claims 1-11 above; therefore, claim 23 is rejected under similar rationale as presented against claims 1-11 above. Further to claim 23, Munoz et al (20240220737) teaches a text summary of the input speech – see para 0086 which output the text derived from the enunciation module.
Claim 24 is a system claim that performs the steps found throughout system claims 1-11 and as such, claim 24 is similar in scope and content to these common features in claims 1-11; therefore, claim 24 is rejected under similar rationale as presented against claims 1-11 above.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Please see related art listed on the PTO-892 form.
Furthermore, the following prior art references were found to have common elements with applicants claims/spec:
Wang et al (20240274122) teaches language translation and speaker ID/embedding using ASR language models (para 0031, 0034)
Matusov et al (20200226327) teaches VAD, language translation, and speaker diarization (para 00070
Lee et al (20230186035) teaches the modeling of multi-speaker target speech derived from multiple speakers (para 0016).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Michael Opsasnick, telephone number (571)272-7623, who is available Monday-Friday, 9am-5pm.
If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Mr. Richemond Dorvil, can be reached at (571)272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free).
/Michael N Opsasnick/Primary Examiner, Art Unit 2658 07/31/2026