DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of Claims
This action is in reply to the claims filed on 22 July 2026.
Claims 1, 12 and 13 have been amended.
Claims 1-23 are currently pending and have been examined.
Response to Arguments
Applicant's arguments, see Page(s) 8-9, filed 22 July 2026, with respect to the 35 USC § 101 rejection(s) of claim(s) 1-23 have been fully considered but they are not persuasive. Applicant argues 1) the claims do not recite a mental process and 2) the claim are integrated into a practical application. The Examiner respectfully disagrees.
Regarding argument 1, the applicant argues the claims do not recite a mental process because the claims require processing audio content with ML models, generating ASR outputs, using outputs for verification, and controlling the audio-processing pipeline based on whether the best language prediction is verified. The Examiner respectfully disagrees. USPTO guidance uses the term ‘‘additional elements’’ to refer to claim features, limitations, and/or steps that are recited in the claim beyond the identified judicial exception. In the below 101 rejection, the Examiner identifies an LID model (claims 1, 12, and 13); ASR Models (claims 1, 12, and 13); a non-transitory computer-readable medium having stored thereon instructions for causing a processing circuitry to execute a process (claim 12); and a processing circuity and memory, the memory containing instructions that, when executed by the processing circuitry configures the system (claim 13) as additional elements. These additional elements are analyzed in step 2A, prong 2 and step 2B of the Alice/Mayo test for eligibility and in the response to argument 2. The remaining limitations are the abstract idea. MPEP 2106.04(a)(2)III. Recites:
Nor do the courts distinguish between claims that recite mental processes performed by humans and claims that recite mental processes performed on a computer. As the Federal Circuit has explained, "[c]ourts have examined claims that required the use of a computer and still found that the underlying, patent-ineligible invention could be performed via pen and paper or in a person’s mind." Versata Dev. Group v. SAP Am., Inc., 793 F.3d 1306, 1335, 115 USPQ2d 1681, 1702 (Fed. Cir. 2015). See also Intellectual Ventures I LLC v. Symantec Corp., 838 F.3d 1307, 1318, 120 USPQ2d 1353, 1360 (Fed. Cir. 2016) (‘‘[W]ith the exception of generic computer-implemented steps, there is nothing in the claims themselves that foreclose them from being performed by a human, mentally or with pen and paper.’’); Mortgage Grader, Inc. v. First Choice Loan Servs. Inc., 811 F.3d 1314, 1324, 117 USPQ2d 1693, 1699 (Fed. Cir. 2016) (holding that computer-implemented method for "anonymous loan shopping" was an abstract idea because it could be "performed by humans without a computer"). Mental processes recited in claims that require computers are explained further below with respect to point C.
The use of the LID and ASR models in the independent claims are recited at a high level. For example, the LID model is used to “obtain a set of LID results” and “to output a plurality of language predictions for the audio content and a respective language prediction output score for each language prediction of the plurality of language predictions”. There is no explanation in the independent claims describing the intermediate steps being performed by the LID model to create the output. As claimed, the process of the outputting a language prediction and score can be performed the same way by a human as by a computer/LID model. Therefore, the LID model is being used in its ordinary capacity as a generic computing element. Similarly, the ASR model is being used to “generate a set of ASR outputs” and the ASR model is configured to “analyze the audio content”. There is no explanation of what steps/operations the ASR model is performing to generate an output from the input. Therefore, the process of the outputting the ASR output can be performed the same way by a human as by a computer/ASR model. The courts do distinguish between claims that recite mental processes performed by humans and claims that recite mental processes performed on a computer (see MPEP 2106/04(a)(2)III). Therefore, the Examiner maintains that a mental process is being recited.
Regarding argument 2, the applicant argues the claims are integrated into a practical application because the claims improve the reliability of an audio-processing pipeline by verifying a language prediction using ASR outputs and using an alternative ranked language prediction when the best prediction is not verified. As explained in argument 1, the ASR and LID model is being used at a high level in a way that is considered generic (see MPEP 2106.05(f)).
Because the claim are merely using a computer as a tool to perform an abstract idea, as discussed in MPEP 2106.05(f) and argument 1, the use of the LID and ASR models do not integrate the judicial exception into a practical application. Therefore, the Examiner maintains that the claims are not eligible. Applicant argues the dependent claims are eligible due to their dependency on independent claims 1, 12, and 13. The Examiner has maintained the 101 rejection for claims 1, 12, and 13. Therefore, the dependent claims have also been rejected for similar reasons.
Applicant's arguments, see Page(s) 10-11, filed 22 July 2026, with respect to the 35 USC § 103 rejection(s) of claim(s) 1-23 have been fully considered but they are not persuasive. Applicant argues the amendments to the independent claims overcome the 103 rejections of Kamano in view of Apsingekar. The Examiner respectfully disagrees.
The Examiner maintains that prior art, Kamano, teaches the amended feature of “wherein, when the verifying indicates that a language prediction having a highest respective language prediction output score is not verified, a language prediction having a second-highest respective language prediction output score is utilized.” Paragraph [0109] of Kamano recites:
Further, in the speech translation server 20 according to the present embodiment, it is also possible to improve the accuracy in identification of a type of a language, as compared with the language identification system described above. FIG. 11 is a diagram illustrating an example of language identification results of an input voice. Regarding two speeches Nos. 1 and 2, FIG. 11 presents the sentence likelihood for English, the sentence likelihood for Chinese, the phoneme count in the speech recognition result by the English speech recognition engine, the phoneme count in the speech recognition result by the Chinese speech recognition engine, and the identification result.
Paragraph [0111] of Kamano recites:
Further, in the example of the speech No. 2 illustrated in FIG. 11, if the type of language is identified only based on the sentence likelihood, the Chinese speech is erroneously recognized as the English speech because the sentence likelihood for English is higher than the sentence likelihood for Chinese. On the other hand, if the type of language is identified using the phoneme count, the language used in the speech No. 2 is correctly identified as Chinese because the phoneme count in the speech recognition result by the Chinese speech recognition engine is larger than the phoneme count in the speech recognition result by the English speech recognition engine.
As shown in the second example of Fig. 11, Paragraph [0109] and Paragraph [0111] of Kamano, there are two language options, English and Chinese. The language with the highest language prediction output score (i.e. sentence likelihood) is English. However, because the phoneme count of English is lower than Chinese, English is not verified as correct and Chinese is selected instead. Therefore, the Examiner maintains the 103 rejection of claims 1, 12, and 13 using Kamano in view of Apsingekar. The Applicant argues the dependent claims are allowable due to their dependence on claims 1, 12 and 13. Claims 1, 12, and 13 have been rejected, therefore, the Examiner is maintaining the rejection of the dependent claims.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-23 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception (i.e., an abstract idea) without significantly more.
Independent claims 1, 12 and 13 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claims regard a process that, as drafted under its broadest reasonable interpretation, covers performance of the limitations as a mental process, but for the recitation of generic computer hardware (e.g., a LID model (claims 1, 12, and 13); ASR Models (claims 1, 12, and 13); a non-transitory computer-readable medium having stored thereon instructions for causing a processing circuitry to execute a process (claim 12); and a processing circuity and memory, the memory containing instructions that, when executed by the processing circuitry configures the system (claim 13)).
In regards to the processing of independent claims 1, 12, and 13, the claimed functionality could be practiced as a mental process in the following manner:
applying a language identification to audio content in order to obtain a set of LID results to output a plurality of language predictions for the audio content and a respective language prediction output score for each language prediction of the plurality of language predictions; (a human can mentally identify a language in audio content by listening and a human can mentally determine a score indicating a language prediction by listening to the audio and mentally making a decision.)
applying audio speech recognition (ASR) to the audio content based on the set of LID results to generate a set of ASR outputs, wherein the set of ASR outputs include a language score and a plurality of predicted words, to analyze the audio content with respect to at least one language prediction; (a human can listen to audio and translate the audio content to text on pen and paper and a human can mentally determine a language score/predicted words based on a language identification) and
verifying an audio processing result based on the set of ASR outputs, wherein, when the verifying indicates that a language prediction having a highest respective language prediction output score is not verified, a language prediction having a second-highest respective language prediction output score is utilized. (a human can listen to audio and mentally verify a result based on data collected about that language)
This judicial exception is not integrated into a practical application. Outside of the identified abstract idea, the claimed invention only includes an LID model (claims 1, 12, and 13); ASR Models (claims 1, 12, and 13); a non-transitory computer-readable medium having stored thereon instructions for causing a processing circuitry to execute a process (claim 12); and a processing circuity and memory, the memory containing instructions that, when executed by the processing circuitry configures the system (claim 13)), which amount to no more than mere instructions to implement an otherwise abstract idea using generic components. Note that the computing components here are being used for their ordinary purpose of executing a program to carry out a process (i.e., being used as a tool) instead of being improved as a tool.
The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. The above identified additional generic computer components are no more than mere instructions to apply the exception using generic computer components that are well-known, routine, and conventional as is evidenced by Bancorp Services V. Sun Life, 687 F.3d 1266, 1278, 103 USPQ2d 1425, 1433 (Fed. Cir. 2012), which teaches a generic computer; and Apsingekar (US 20200219492 A1) which teaches well known ASR and LID models (see at least Paragraph [0003] “Voice-based interfaces are being used more and more often […] These types of interfaces often include an automatic speech recognition (ASR) system,” and Paragraph [0056] “Various approaches have been developed for training models to recognize words and phrases in various languages, and additional approaches are sure to be developed in the future.” of Apsingekar).
Therefore, claims 1, 12, and 13 is not eligible subject matter under 35 USC 101.
The remaining dependent claims fail to add patent eligible subject matter to their respective parent claims:
Claims 2-3 and 14-15 add at least one model, however this model is never specified and is described at a high level and is therefore a generic computing element. This/these additional element(s) alone or in ordered combination does no more than merely use a computer as a tool to perform an abstract idea (see MPEP 2106.05(f)), which does not integrate the claim(s) into a practical application nor does it render a claim as being significantly more than the abstract idea. See paragraph [0005] of the instant specification to explains that it is common for LID output to feed into ASR models.
Claims 4 and 16 further recites the at least one model performing natural language processing. The model and NLP is described at a high level and is therefore a generic computing element. This/these additional element(s) alone or in ordered combination does no more than merely use a computer as a tool to perform an abstract idea (see MPEP 2106.05(f)), which does not integrate the claim(s) into a practical application nor does it render a claim as being significantly more than the abstract idea. See at least Paragraph [004] in the background of the instant specification which describes how ASR outputs are known to be used for NLP tasks.
Claims 5 and 17 further recites an acoustic model and a language model to the results. These models are described at a high level and is therefore a generic computing element. This/these additional element(s) alone or in ordered combination does no more than merely use a computer as a tool to perform an abstract idea (see MPEP 2106.05(f)), which does not integrate the claim(s) into a practical application nor does it render a claim as being significantly more than the abstract idea. See Paragraph [003] in the background of the instant specification explaining how ASRs commonly use acoustic and language models.
Claims 6 and 18 recite additional steps in the mental process of classifying speech and verifying the results. These steps can be performed in the human mind. Claims 6 and 18 further recite the additional element of a classifier. This/these additional element(s) alone or in ordered combination does no more than merely use a computer as a tool to perform an abstract idea (see MPEP 2106.05(f)), which does not integrate the claim(s) into a practical application nor does it render a claim as being significantly more than the abstract idea. See at least Paragraph [0030] of the instant specification that lists many known ways of realizing a classifier.
Claims 7 and 19 recite additional steps in the mental process of determining how likely a character is to occur. This can be performed in the human mind.
Claims 8 and 20 recite additional steps in the mental process of determining variables for classification of a language that could be performed in the human mind. Claims 8 and 20 further recite a first decoding model and second decoding model. This/these additional element(s) alone or in ordered combination does no more than merely use a computer as a tool to perform an abstract idea (see MPEP 2106.05(f)), which does not integrate the claim(s) into a practical application nor does it render a claim as being significantly more than the abstract idea. See Paragraph [0030] of the instant specification which recites several known decoding techniques.
Claims 9 and 21 further recite a variable that can be measured by the human mind.
Claims 10 and 22 further recite a step for determining the likelihood of a language prediction that could be performed in the human mind.
Claims 11 and 23 further recites selecting an ASR model based on an LID output. A human can mentally determine a proper tool based on a language. The ASR model and LID models are additional elements. This/these additional element(s) alone or in ordered combination does no more than merely use a computer as a tool to perform an abstract idea (see MPEP 2106.05(f)), which does not integrate the claim(s) into a practical application nor does it render a claim as being significantly more than the abstract idea. See paragraph [0005] of the instant specification to explains that it is common for LID output to feed into ASR models.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claim(s) 1-6, 10-18, and 22-23 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kamano (US 20200111476 A1) in view of Apsingekar (US 20200219492 A1).
Regarding claim 1, Kamano teaches a method for audio processing verification, comprising:
applying a language identification (LID) model to audio content in order to obtain a set of LID results, wherein the LID model is configured to output a plurality of language predictions for the audio content and a respective language prediction output score for each language prediction of the plurality of language predictions; (see at least Paragraph [0100] “the likelihood calculation units 21-1 to 21-M calculate the sentence likelihoods for the respective first to M-th languages based on the linguistic models from the speech texts output by the speech recognition units 12-1 to 12-M (step S201).”; Paragraph [0102] “when the sentence likelihood l.sub.k,s exceeds the threshold value T.sub.1 (Yes in step S203), the language identification unit 22 adds the k-th language to a candidate list held in an internal memory (not illustrated) (step S204).”; Examiner notes the sentence likelihood is a language prediction output score.)
applying at least one audio speech recognition (ASR) model to the audio content based on a set of languages (see at least Paragraph [0097] “As Illustrated in FIG. 10, the speech recognition units 12-1 to 12-M input the voice data of the speech segment to the speech recognition engines for the respective languages allocated to the systems, thereby performing speech recognition for the respective systems of the speech recognition units 12-1 to 12-M, for example, for the respective first to M-th languages (step S101).”; Paragraph [0098] “the phoneme string conversion units 13-1 to 13-M convert the speech texts obtained as the speech recognition results by the speech recognition units 12-1 to 12-M into phoneme strings expressed by phoneme symbols in accordance with the IPA (step S102).”; Paragraph [0099] “The phoneme count calculation units 14-1 to 14-M calculate the phoneme counts in the phoneme strings converted from the speech texts by the phoneme string conversion units 13-1 to 13-M (step S103).”; Fig. 9 and 10) and
verifying an audio processing result based on the set of ASR outputs, (Paragraph [0105] “when the languages exist in the candidate list (Yes in step S207), the language identification unit 15 identifies, as the speech language, the language having the largest phoneme count among the languages existing in the candidate list stored in the internal memory (step S208).”; Paragraph [0113] “a correct rate of “70%” in the case where the type of language is identified using only the sentence likelihood, […] and a correct rate of “90%” in the case where the type of language is identified by using the sentence likelihood and the phoneme count.”; Fig. 9-12)
wherein, when the verifying indicates that a language prediction having a highest respective language prediction output score is not verified, a language prediction having a second-highest respective language prediction output score is utilized. (see at least Paragraph [0109] “FIG. 11 is a diagram illustrating an example of language identification results of an input voice.”; Paragraph [0111] “in FIG. 11, if the type of language is identified only based on the sentence likelihood, the Chinese speech is erroneously recognized as the English speech because the sentence likelihood for English is higher than the sentence likelihood for Chinese. On the other hand, if the type of language is identified using the phoneme count, the language used in the speech No. 2 is correctly identified as Chinese because the phoneme count in the speech recognition result by the Chinese speech recognition engine is larger than the phoneme count in the speech recognition result by the English speech recognition engine.”; Fig. 11; Examiner notes in the example no. 2 of Fig. 11, English has the higher sentence likelihood (i.e. language prediction output score), but Chinese is still determined to be the language because the phoneme count does not verify English as the correct language.)
Kamano does not teach:
applying at least one audio speech recognition (ASR) model to the audio content based on the set of LID results to generate a set of ASR outputs, wherein the set of ASR outputs include a language score and a plurality of predicted words, wherein each ASR model is configured to analyze the audio content with respect to at least one language prediction;
However, Apsingekar teaches the known technique of applying at least one ASR model to the audio content based a set of LID results. (see at least Paragraph [0049] “Each language model 210 can therefore be used to calculate a probability for each phoneme associated with a specific language, and this can be performed for each portion (segment) of the audio-based input 202.”; Paragraph [0052] “The neural classification model 214 uses the identified language or languages that are associated with the audio-based input 202 to control an ASR engine 216, which generally operates to convert the audio-based input 202 (or specific portions thereof) into text. Here, the ASR engine 216 includes or is used in conjunction with multiple ASR models 218, which are associated with different languages.”; el. 208 and 216 of Fig. 2 of Apsingekar)
This step of Apsingekar is applicable to the method of Kamano as they both share characteristics and capabilities, namely, they are directed to determining a language of an audio sample. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to have modified the method of Kamano to incorporate the known technique of each ASR model analyzing audio content based on at least one language prediction output by a LID model as taught by Apsingekar. One of ordinary skill in the art before the effective filling date of the claimed invention would have been motivated to modify Kamano in order to improve the detection of spoken languages to provide better responses to user inputs and enable systems to serve a wider population of potential users (see paragraph [0028] of Apsingekar).
Additionally, in regard to claim 1, the Examiner further notes the recited “when” on line 11 does not move to distinguish the claimed invention from the cited art. This phrase is a conditional/contingent limitation with the noted “a language prediction having a second-highest respective language prediction output score is utilized” step(s) not necessarily performed. The broadest reasonable interpretation of a method (or process) claim having contingent limitations requires only those steps that must be performed and does not include steps that are not required to be performed because the condition(s) precedent are not met. Language that suggests or makes optional but does not require steps to be performed or does not limit a claim to a particular structure does not limit the scope of a claim or claim limitation. [See Ex parte Schulhauser, Appeal 2013-007847 (PTAB April 28, 2016) for an analysis of contingent claim limitations in the context of both method claims and system claims.; MPEP §2111.04 II].
Regarding claim 2, Kamano in view of Apsingekar teaches the method of claim 1. Kamano further teaches:
applying at least one model based on the verified audio processing result. (see at least Paragraph [0106] “the speech translation unit 16 converts the speech text for the speech language identified in step S208 into a translated text for Japanese or the foreign language (step S209). Subsequently, the output unit 17 generates, from the translated text obtained in step S209, a synthesized voice for reading aloud the translated text, outputs the voice data of the synthesized voice to the speech translation terminal 30 (step S210)”; Fig. 10)
Regarding claim 3, Kamano in view of Apsingekar teaches the method of claim 2. Kamano further teaches:
wherein the audio processing result includes the set of LID results, wherein applying the at least one model includes (see at least Paragraph [0102] “when the sentence likelihood l.sub.k,s exceeds the threshold value T.sub.1 (Yes in step S203), the language identification unit 22 adds the k-th language to a candidate list held in an internal memory (not illustrated) (step S204).”; Paragraph [0105] “when the languages exist in the candidate list (Yes in step S207), the language identification unit 15 identifies, as the speech language, the language having the largest phoneme count among the languages existing in the candidate list stored in the internal memory (step S208).”; Fig. 10)
Kamano does not teach:
wherein the audio processing result includes the set of LID results, wherein applying the at least one model includes
However, Apsingekar teaches:
wherein the audio processing result includes the set of LID results, wherein applying the at least one model includes performing ASR using the set of LID results. (see at least Paragraph [0049] “Each language model 210 can therefore be used to calculate a probability for each phoneme associated with a specific language, and this can be performed for each portion (segment) of the audio-based input 202.”; Paragraph [0052] “The neural classification model 214 uses the identified language or languages that are associated with the audio-based input 202 to control an ASR engine 216, which generally operates to convert the audio-based input 202 (or specific portions thereof) into text. Here, the ASR engine 216 includes or is used in conjunction with multiple ASR models 218, which are associated with different languages.”; el. 208 and 216 of Fig. 2 of Apsingekar)
The motivation for making this modification to the teachings of Kamano is the same as that set forth above, in the rejection of claim 1.
Regarding claim 4, Kamano in view of Apsingekar teaches the method of claim 2. Kamano further teaches:
wherein applying the at least one model includes performing natural language processing using the set of ASR outputs. (see at least Paragraph [0106] “the speech translation unit 16 converts the speech text for the speech language identified in step S208 into a translated text for Japanese or the foreign language (step S209). Subsequently, the output unit 17 generates, from the translated text obtained in step S209, a synthesized voice for reading aloud the translated text, outputs the voice data of the synthesized voice to the speech translation terminal 30 (step S210), and terminates the processing.”)
Regarding claim 5, Kamano in view of Apsingekar teaches the method of claim 1. Kamano further teaches:
wherein the at least one ASR model includes an acoustic model and a language model [acoustic and linguistic model], wherein applying the at least one ASR model further comprises:
applying the acoustic model and the language model using the set of languages (Paragraph [0038] “the linguistic model and the acoustic model are modeled for each of the multiple languages”; Paragraph [0052] “Thereafter, each of the speech recognition engines performs matching of the feature quantity string f.sub.0, f.sub.1, . . . , f.sub.12 with a phoneme acoustic model […] the speech recognition engine performs matching of a word string allocated using the word acoustic model with a linguistic model in which the existence probability of each word order is defined”; Paragraph [0092] “In FIG. 9, the same reference numerals are given to the functional units having the same functions as those of the speech translation server 10 illustrated in FIG. 7”; Fig. 7 and 9)
Kamano does not teach:
applying the acoustic model and the language model using the set of LID results.
However, Apsingekar teaches the known technique of an LID model producing a set of languages as an LID result. (see at least Paragraph [0049] “Each language model 210 can therefore be used to calculate a probability for each phoneme associated with a specific language, and this can be performed for each portion (segment) of the audio-based input 202.”; el. 208 of Fig. 2 of Apsingekar)
The motivation for making this modification to the teachings of Kamano is the same as that set forth above, in the rejection of claim 1.
Regarding claim 6, Kamano in view of Apsingekar teaches the method of claim 1. Kamano further teaches:
determining a plurality of language accuracy factors based on the set of ASR outputs. (see at least Paragraph [0074] “The phoneme count calculation unit 14 may weight each phoneme contained in the phoneme string and calculate the sum of the weights as the phoneme count in accordance with the following equation”)
Kamano does not teach:
determining a plurality of input features for a classifier based on the language accuracy factors; and
applying the classifier to the plurality of input features, wherein the classifier is configured to output at least one score, wherein each score output by the classifier indicates a likelihood that a language prediction of the at least one language prediction is correct, wherein the audio processing result is verified based further on the at least one score output by the classifier.
However, Apsingekar teaches:
determining a plurality of input features for a classifier based on the language accuracy factors; (see at least Paragraph [0061] “the feature concatenator 212 stacks the information generated by the feature extractor 206 and the scoring function 208 to form the feature vectors, which are then further processed by the neural classification model 214.”; Paragraph [0071] “the feature concatenator 212 may temporally accumulate or otherwise combine any suitable number of inputs over any suitable time period.”; Fig. 2 and 3 of Apsingekar) and
applying the classifier to the plurality of input features, wherein the classifier is configured to output at least one score, wherein each score output by the classifier indicates a likelihood that a language prediction of the at least one language prediction is correct, wherein the audio processing result is verified based further on the at least one score output by the classifier. (see at least Paragraph [0068] “The activation function of the neural classification model 214 is implemented in this example using a softmax function 322. The softmax function 322 receives the average convoluted values from the pooling layer 320 and operates to calculate final probabilities 324, which identify the probability of each language being associated with at least part of the audio-based input 202. […] For each portion of the audio-based input 202, the probabilities across all languages can be used to identify the ASR model 218 that should be used to process that portion of the audio-based input 202.”; Fig. 3 of Apsingekar)
The motivation for making this modification to the teachings of Kamano is the same as that set forth above, in the rejection of claim 1.
Regarding claim 10, Kamano in view of Apsingekar teaches the method of claim 1. Kamano further teaches:
wherein the set of LID results include at least one language prediction output score, wherein each language prediction output score indicates a likelihood for a respective language prediction of the at least one language prediction. (see at least paragraph [0100] “the likelihood calculation units 21-1 to 21-M calculate the sentence likelihoods for the respective first to M-th languages based on the linguistic models from the speech texts output by the speech recognition units 12-1 to 12-M (step S201).”)
Regarding claim 11, Kamano in view of Apsingekar teaches the method of claim 1. Kamano further teaches:
wherein the at least one ASR model is at least one first ASR model among a plurality of ASR models (Examiner notes Fig. 9 shows multiple speech recognition units, phoneme string conversion units and phoneme count calculation units.)
Kamano does not teach:
selecting the at least one first ASR model to be applied from among the plurality of ASR models based on the LID outputs.
However, Apsingekar teaches:
selecting the at least one first ASR model to be applied from among the plurality of ASR models based on the LID outputs. (see at least Paragraph [0049] “Each language model 210 can therefore be used to calculate a probability for each phoneme associated with a specific language, and this can be performed for each portion (segment) of the audio-based input 202.”; Paragraph [0052] “The neural classification model 214 uses the identified language or languages that are associated with the audio-based input 202 to control an ASR engine 216, which generally operates to convert the audio-based input 202 (or specific portions thereof) into text. Here, the ASR engine 216 includes or is used in conjunction with multiple ASR models 218, which are associated with different languages.”; el. 208 and 216 of Fig. 2 of Apsingekar)
The motivation for making this modification to the teachings of Kamano is the same as that set forth above, in the rejection of claim 1.
Claim 12:
Claim(s) 12 is/are directed to a non-transitory computer-readable medium. Claim(s) 12 recite limitations parallel in nature as those addressed above for claim(s) 1, which are directed towards a method. Claim(s) 12 is/are therefore rejected for the same reasons as set above for claim(s) 1. Claim 12 further recites “a non-transitory computer-readable medium having stored thereon instructions for causing a processing circuitry to execute a process” (see Paragraph [0005] “non-transitory computer readable recording medium” of Kamano)
Claim 13:
Claim(s) 13 is/are directed to a system. Claim(s) 13 recite limitations parallel in nature as those addressed above for claim(s) 1, which are directed towards a method. Claim(s) 13 is/are therefore rejected for the same reasons as set above for claim(s) 1. Claim 13 further recites “a processing circuitry; and a memory, the memory containing instructions that, when executed by the processing circuitry, configure the system” (see Paragraph [0064] “processor reads a speech translation program in addition to operating system (OS) from a storage device” of Kamano).
Claim 14:
Claim(s) 14 is/are directed to a system. Claim(s) 14 recite limitations parallel in nature as those addressed above for claim(s) 2, which are directed towards a method. Claim(s) 14 is/are therefore rejected for the same reasons as set above for claim(s) 2. Claim 14 further recites “a processing circuitry; and a memory, the memory containing instructions that, when executed by the processing circuitry, configure the system” (see Paragraph [0064] “processor reads a speech translation program in addition to operating system (OS) from a storage device” of Kamano).
Claim 15:
Claim(s) 15 is/are directed to a system. Claim(s) 15 recite limitations parallel in nature as those addressed above for claim(s) 3, which are directed towards a method. Claim(s) 15 is/are therefore rejected for the same reasons as set above for claim(s) 3. Claim 15 further recites “a processing circuitry; and a memory, the memory containing instructions that, when executed by the processing circuitry, configure the system” (see Paragraph [0064] “processor reads a speech translation program in addition to operating system (OS) from a storage device” of Kamano).
Claim 16:
Claim(s) 16 is/are directed to a system. Claim(s) 16 recite limitations parallel in nature as those addressed above for claim(s) 4, which are directed towards a method. Claim(s) 16 is/are therefore rejected for the same reasons as set above for claim(s) 4. Claim 16 further recites “a processing circuitry; and a memory, the memory containing instructions that, when executed by the processing circuitry, configure the system” (see Paragraph [0064] “processor reads a speech translation program in addition to operating system (OS) from a storage device” of Kamano).
Claim 17:
Claim(s) 17 is/are directed to a system. Claim(s) 17 recite limitations parallel in nature as those addressed above for claim(s) 5, which are directed towards a method. Claim(s) 17 is/are therefore rejected for the same reasons as set above for claim(s) 5. Claim 17 further recites “a processing circuitry; and a memory, the memory containing instructions that, when executed by the processing circuitry, configure the system” (see Paragraph [0064] “processor reads a speech translation program in addition to operating system (OS) from a storage device” of Kamano).
Claim 18:
Claim(s) 18 is/are directed to a system. Claim(s) 18 recite limitations parallel in nature as those addressed above for claim(s) 6, which are directed towards a method. Claim(s) 18 is/are therefore rejected for the same reasons as set above for claim(s) 6. Claim 18 further recites “a processing circuitry; and a memory, the memory containing instructions that, when executed by the processing circuitry, configure the system” (see Paragraph [0064] “processor reads a speech translation program in addition to operating system (OS) from a storage device” of Kamano).
Claim 22:
Claim(s) 22 is/are directed to a system. Claim(s) 22 recite limitations parallel in nature as those addressed above for claim(s) 10, which are directed towards a method. Claim(s) 22 is/are therefore rejected for the same reasons as set above for claim(s) 10. Claim 22 further recites “a processing circuitry; and a memory, the memory containing instructions that, when executed by the processing circuitry, configure the system” (see Paragraph [0064] “processor reads a speech translation program in addition to operating system (OS) from a storage device” of Kamano).
Claim 23:
Claim(s) 23 is/are directed to a system. Claim(s) 23 recite limitations parallel in nature as those addressed above for claim(s) 11, which are directed towards a method. Claim(s) 23 is/are therefore rejected for the same reasons as set above for claim(s) 11. Claim 23 further recites “a processing circuitry; and a memory, the memory containing instructions that, when executed by the processing circuitry, configure the system” (see Paragraph [0064] “processor reads a speech translation program in addition to operating system (OS) from a storage device” of Kamano).
Claim(s) 7-8 and 19-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kamano (US 20200111476 A1) in view of Apsingekar (US 20200219492 A1) in further view of Frey (US 9858340 B1).
Regarding claim 7, Kamano in view of Apsingekar teaches the method of claim 6. Kamano further teaches:
wherein the set of ASR outputs includes a plurality of likelihoods of a plurality of (see at least Paragraph [0072] “a phoneme string expressed by phoneme symbols […] the phoneme string conversion unit 13 converts the speech text obtained as a speech recognition result by the speech recognition unit 12 into a phoneme string […] the phoneme string conversion unit 13 identifies the phonemes by performing maximum likelihood estimation, […] the speech text is converted into the time-series data of phonemes.”) wherein determining the plurality of language accuracy factors further comprises:
Kamano does not teach:
generating a character likelihoods table, wherein the character likelihoods table includes a likelihood for each of the
However, Apsingekar teaches:
generating a character likelihoods table, wherein the character likelihoods table includes a likelihood for each of the (see at least Paragraph [0047] “The feature extractor 206 may use any suitable technique for extracting phoneme features or other features from audio-based input 202.”; Paragraph [0048] “The extracted features of the audio-based input 202 are provided from the feature extractor 206 to a language-specific acoustic model scoring function 208, which generally operates to identify the likelihood that various portions of the audio-based input 202 are from specific languages”; Paragraph [0050] “the feature concatenator 212 may accumulate the probabilities that are output from the scoring function 208 over time to create a matrix of values”; Paragraph [0062] “Each column 304 is associated with a different time and thereof a different portion of the audio-based input 202.” of Apsingekar)
The motivation for making this modification to the teachings of Kamano is the same as that set forth above, in the rejection of claim 1.
Kamano in view of Apsingekar does not teach:
a plurality of likelihoods for each of the plurality of observed characters.
However, Frey teaches:
a plurality of likelihoods for each of the plurality of observed characters. (Col. 17, ll. 31-51 “the last layer in a speech-to-text network may output to probabilities of a letter of the alphabet.” of Frey)
This step of Frey is applicable to the method of Kamano as they both share characteristics and capabilities, namely, they are directed to using Audio Speech Recognition to analyze speech patterns. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to have modified the character symbols of Kamano to observed characters as taught by Frey. One of ordinary skill in the art before the effective filling date of the claimed invention would have been motivated to modify Kamano in order to use neural networks in speech and audio classification (see Col. 17, ll. 31-51 of Frey).
Regarding claim 8, Kamano in view of Apsingekar in further view of Frey teaches the method of claim 7. Kamano further teaches:
wherein determining the plurality of language accuracy factors further comprises:
(see at least Paragraph [0052] “the speech recognition engine performs matching of the phonemes allocated using the phoneme acoustic model with a word acoustic model in which the existence probability of a combination of each phoneme string and the corresponding English word are modeled, and thereby allocates the word to the phonemes allocated by using the phoneme acoustic model. Furthermore, the speech recognition engine performs matching of a word string allocated using the word acoustic model with a linguistic model in which the existence probability of each word order is defined, and thereby evaluates the word order by a score such as a likelihood. […] text associated with the word string having the highest evaluation score is output as a speech recognition result.”; Paragraph [0072] “the phoneme string conversion unit 13 converts the speech text obtained as a speech recognition result by the speech recognition unit 12 into a phoneme string […] the phoneme string conversion unit 13 identifies the phonemes by performing maximum likelihood estimation, […] the speech text is converted into the time-series data of phonemes.”; Fig. 9)
Kamano does not teach:
applying a first decoding model and a second decoding model to the character likelihoods table.
However, Apsingekar teaches:
applying a first decoding model and a second decoding model to the character likelihoods table. (see at least Paragraph [0050] “the feature concatenator 212 may accumulate the probabilities that are output from the scoring function 208 over time to create a matrix of values” of Apsingekar)
The motivation for making this modification to the teachings of Kamano is the same as that set forth above, in the rejection of claim 1.
Claim 19:
Claim(s) 19 is/are directed to a system. Claim(s) 19 recite limitations parallel in nature as those addressed above for claim(s) 7, which are directed towards a method. Claim(s) 19 is/are therefore rejected for the same reasons as set above for claim(s) 7. Claim 19 further recites “a processing circuitry; and a memory, the memory containing instructions that, when executed by the processing circuitry, configure the system” (see Paragraph [0064] “processor reads a speech translation program in addition to operating system (OS) from a storage device” of Kamano).
Claim 20:
Claim(s) 20 is/are directed to a system. Claim(s) 20 recite limitations parallel in nature as those addressed above for claim(s) 8, which are directed towards a method. Claim(s) 20 is/are therefore rejected for the same reasons as set above for claim(s) 8. Claim 20 further recites “a processing circuitry; and a memory, the memory containing instructions that, when executed by the processing circuitry, configure the system” (see Paragraph [0064] “processor reads a speech translation program in addition to operating system (OS) from a storage device” of Kamano).
Claim(s) 9 and 21 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kamano (US 20200111476 A1) in view of Apsingekar (US 20200219492 A1) in further view of Yoon (US 20110270612 A1).
Regarding claim 9, Kamano in view of Apsingekar teaches the method of claim 6. Kamano in view of Apsingekar does not teach:
wherein the plurality of features includes a number of words per second, wherein the audio processing result is verified based further on the number of words per second.
However, Yoon teaches:
wherein the plurality of features includes a number of words per second, wherein the audio processing result is verified based further on the number of words per second. (see at least Paragraph [0023] “The word accuracy rate calculator 512 further considers speech recognizer output metrics that include words per second” of Yoon)
This step of Yoon is applicable to the method of Kamano as they both share characteristics and capabilities, namely, they are directed to audio speech recognition. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to have modified the method of Kamano to incorporate the plurality of features including a number of words per second, wherein the audio processing result is verified based further on the number of words per second as taught by Yoon. One of ordinary skill in the art before the effective filling date of the claimed invention would have been motivated to modify Kamano in order to provide meaningful values for content metrics (see paragraph [0017] of Yoon).
Claim 21:
Claim(s) 21 is/are directed to a system. Claim(s) 21 recite limitations parallel in nature as those addressed above for claim(s) 9, which are directed towards a method. Claim(s) 21 is/are therefore rejected for the same reasons as set above for claim(s) 9. Claim 21 further recites “a processing circuitry; and a memory, the memory containing instructions that, when executed by the processing circuitry, configure the system” (see Paragraph [0064] “processor reads a speech translation program in addition to operating system (OS) from a storage device” of Kamano).
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DANIELLE ELIZABETH ZEVITZ whose telephone number is (703)756-1070. The examiner can normally be reached Mo-Th 10am-6pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Lynda Jasmin can be reached at (571) 272-6782. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/DANIELLE ELIZABETH ZEVITZ/Examiner, Art Unit 3629 /JAMES S WOZNIAK/Primary Examiner, Art Unit 2655