DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 15-20 are rejected under 35 U.S.C. §112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor, or for pre-AIA the applicant regards as the invention.
Independent claim 15 recites:
One or more processors comprising processing circuitry to cause performance of operations comprising:
generating a final textual transcript corresponding to an audio segment based at least on a multi-lingual language model processing the audio segment to generate an initial textual transcript including one or more language indicators and one or more monolingual language models processing the initial textual transcript and the one or more language indicators to generate the final textual transcript.
Claim 15 has just one long limitation without any punctuations or indentations. The above limitation could be interpreted differently depending on how to break the long limitation into separate limitations (indentation / punctuations are added for two different interpretations):
#1 Interpretation
generating a final textual transcript corresponding to an audio segment based at least on a multi-lingual language model;
processing the audio segment to generate an initial textual transcript including one or more language indicators, and
one or more monolingual language models processing the initial textual transcript and the one or more language indicators to generate the final textual transcript.
#2 Interpretation
generating a final textual transcript corresponding to an audio segment based at least on a multi-lingual language model processing the audio segment to generate an initial textual transcript including one or more language indicators, and
one or more monolingual language models processing the initial textual transcript and the one or more language indicators to generate the final textual transcript.
By comparing a properly presented independent claim 1, it appears interpretation #2 is likely applicant intended interpretation. However, interpretation #1 is a valid and reasonable interpretation. Since the single long claim limitation recited in claim 15 could have different interpretations depending on how to break the long limitation into different sections, the claimed scope of claim 15 is ambiguous / unclear. Dependent claims 16-20 are also rejected because these dependent claims include the long limitation of claim 15.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1, 5-7, 9, 12-15 and 18-20 are rejected under 35 U.S.C. §102 (a)(1) as being anticipated by Hung (US PG Pub. 2022/0343893, referred to as Hung).
Hung discloses a method / a system of multilingual speech recognition using a first speech recognition engine (corresponding to a claimed “using a multilingual speech-to-text module”) to generate initial speech transcript. The input speech contains both a first portion in a first language and a second portion in a second language (Hung, [0006], [0008]). The initial recognition results are stored in a cache. When a language identification (LID) module detects that the second portion of the input speech is in the second language, a third transcript is generated by using a second speech recognition engine (corresponding to a claimed “a monolingual LM”) for recognizing a second language in the second language. The final transcript is generated by replacing a portion in the initial transcript generated by the first speech recognition engine with the third transcript generated by the second speech recognition engine to correct transcription errors (Hung, [0006], [0008-0009], Fig. 8).
Independent claim 1 is a method claim. Independent claim 9 is directed to a system. Independent claim 15 is directed to a processor. Claim 9 includes similar limitations as the method claim 1. Claim 15 is a much broader than claim 1 or claim 9 by using broad terms and omitting several limitations. In the following analysis, the examiner analyzes limitations recited in the narrower claim 1 as a representative claim. Claim 9 (a system) and claim 15 (a processor) are rejected based on the same rationale.
Regarding claims 1, 9 and 15, Hung discloses a method, a system and a processor (Hung, [0006], Fig. 1, a computer implemented multilingual speech recognition using two speech recognition engines), comprising:
determining, using a multilingual speech-to-text (STT) model of a multilingual automatic speech recognition (ASR) system and using an audio sample as input to the multilingual STT model (Hung, [0006], [0009], using a first speech recognition engine to generate an initial transcript for a speech input speech that contains a first language and a second language), a textual transcript associated with the audio sample and one or more language indicators each associated with a respective grammatical unit of one or more grammatical units of the textual transcript (Hung, [0033], [0070-0071], audio stream contains words, phonemes in different languages, e.g., English and Chinese);
identifying a monolingual language model (LM) of a plurality of monolingual LMs of the ASR system using a language indicator of the one or more language indicators (Hung, [0006], [0044-0047], [0070], identifying audio stream contains spoken words in English or in Chinese, language indicator “en-US”, “zh-CH”, selecting a second speech recognition for recognizing Chinese words that are indicated as “zh-CH”); and
causing the textual transcript associated with the audio sample to be refined using the identified LM and using a subset of the textual transcript as input to the identified LM (Huang, [0009], [0058], [0071], [0094], replacing / overwritten a portion of incorrect transcript from the first speech recognition engine with a third transcript generated by the second speech recognition engine).
Regarding claims 5 and 13 Hung further discloses a multilingual vocabulary of the multilingual STT model comprises the one or more language indicators (Hung, [0070-0071], a code switching sentence containing English words, indicates as “en-US” and Chinese words, indicated as “zh-CN”).
Regarding claims 6 and 14 Hung further discloses a model architecture of the multilingual STT model corresponds to a model architecture of a monolingual STT model (Hung, [0006], [0029], Fig. 2, both first speech recognition engine SR A and the second speech recognition engine SR B are machine learning models).
Regarding claims 7 and 19, Hung further discloses two language indicators of the one or more language indicators are each associated with a code-switched grammatical unit of the one or more grammatical units of the textual transcript (Hung, [0056], [0070-0071], Fig. 2, #210, an input audio contains mixed English words indicated as “en-US”, Chinese words indicated as “zh-CN”).
Regarding claim 12, Hung further discloses the one or more operations include at least ONE of:
causing presentation of at least a portion of the updated textual transcript using one or more display devices of the system (Hung, [0082-0084], Fig. 6A-6F, a graphical user interface showing updated transcriptions; Note prior art reference only need to teach ONE alternative recited using “at least ONE of” and “OR”);
Regarding claim 18, Hung further discloses the final textual transcript is ONE of stored on a device, visually presented on a device (Hung, Fig. 6A-6F), OR used to generate a synthetic audio output using a device (Prior art reference only need to teach ONE alternative recited using “at least ONE of” and “OR”).
Regarding claim 20, Hung further discloses the one or more processors are comprised in at least ONE of: a system implemented at least partially using cloud computing resources (Hung, [0027], clouding computing).
Claim Rejections - 35 USC § 103
In the event the deter mination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 2-3, 10-11, and 16-17 are rejected under 35 U.S.C. §103 as being unpatentable over Hung in view of Zhang et al. (US PG Pub. 2024/0135923, referred to as Zhang).
Zhang discloses a neural network implemented multilingual speech recognition system (Zhang, [0035], Fig. 2B). Zhang discloses the neural network implemented multilingual speech recognition system has an encoder to process multi-language input and identify a language (Zhang, [0004], [0013]). The multilingual speech recognition system also has monolingual output layer that process outputs from the high order features from the encoder and language ID indicator from LID to generate a final transcription.
Regarding claims 2, 10 and 16, Hung discloses generating a final speech transcript by replacing incorrect transcription with a transcription using a second speech recognition engine based on detected language ID ([0070], Chinese: zh-CH). Since the final output does not have a language ID: zh-CH, Huang implies that the language identifier is removed. To further show this feature, the examiner cites Zhang. Zhang discloses output speech recognition results using a monolingual output layer based on language prediction as an input (Zhang, [0008], Fig. 2B). The output does not have the language prediction which means the language indicator is removed.
Both Hung and Zhuang are dealing with speech recognition. It would have been obvious to a person having ordinary skill in the art at the time the invention was filed to combine Hung’s teaching with Zhang’s teaching to output a final result without having a language indicator. One having ordinary skill in the art would have been motivated to make such a modification because the speech transcript itself can show what language it is.
Regarding claims 3, 11 and 17, Huang further discloses Huang further discloses training the multilingual STT model using training data comprising one or more second audio samples, one or more second textual transcripts each associated with a respective audio sample of the one or more second audio samples, and one or more second language indicators each associated with a respective textual transcript of the one or more second textual transcripts (Hung, [0029-0033], retrieving various training data, such as language spoken utterance, audio data, video data, words to train machine learning model).
Claims 4 and 8 are rejected under 35 U.S.C. §103 as being unpatentable over Hung in view of Aguilar et al. (“LinCE: A Centralized Benchmark for Linguistic Code-switching Evaluation”, published in 2020, referred to as Aguilar).
Regarding claims 4 and 8, Hung discloses training machine learning speech recognition models to recognize multilingual speech inputs. Hung discloses identifying a current language when the speech contains dynamically switching between two languages (Hung, [0036], [0056], [0070-0071], an utterance contains both English and Chinese at different time segments, labelled as “en-US” time period T1, and “zh-CH” time period T2, T3).
Hung does not explicitly disclose the one or more language indicators are each located following a punctuation mark of a respective sentence.
Aguilar discloses various training corpora for training machine learning models in linguistic code-switching applications. The code-switching applications contain sentences in mixed language pairs (Aguila, section 1, Introduction). Augilar discloses each token of a sentence is labeled with a language ID (Lang 1, Lang 2) and other information (part of speech POS, named entities: NER) (Aguila, section 1, Introduction, Table 2). Aguilar implies each token has a language identifier located following a punctuation (e.g., a space, comma, or other type of separator).
Both Hung and Aguilar are dealing with processing multilingual language data. It would have been obvious to a person having ordinary skill in the art at the time the invention was filed to combine modify Hung’s teaching with Aguilar’s teaching to label words with a language ID and using a space “” or a comma “,” after the word. One having ordinary skill in the art would have been motivated to make such a modification so that it is easier to determine which word is from a sentence and which is added language ID.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. The examiner discovered several relevant prior art references that are related to one or more concepts disclosed by the instant application. These references are included in the attached PTO-892 form for completeness of the record.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jialong He, whose telephone number is (571) 270-5359. The examiner can normally be reached on Monday – Friday, 8:00AM – 4:30PM, EST.
If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Pierre Desir can be reached on (571) 272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JIALONG HE/Primary Examiner, Art Unit 2659