Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
Claims 1, 3, 8, and 16-20 are amended. Claims 1-20 are presented for examination.
Information Disclosure Statement
The information disclosure statements (IDS) submitted on 06/25/2026 was filed before the mailing date of the fin office action. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Response to Arguments
Applicant’s arguments filed on 06/25/2026 have been reviewed. Following are the
responses to the amendments.
Rejection under 35 U.S.C. 101
Applicant arguments have been considered and are persuasive, hence the rejection under 35 U.S.C. 101 is withdrawn.
Rejection under 35 U.S.C. 103
(i). “The cited references fail to disclose or suggest a plurality of parallel processing paths that each include "a machine learning model configured to compare the input speech signal to each language in an associated subset of the plurality of languages and determine a match score for each language" and "output an identification of a candidate language [that is a closest match]" per path, as recited in claim 1.”
The examiner agrees that the previously applied combination of Sung and Abdelhady does not teach or suggest the newly amended limitation set forth by the applicant. Accordingly, the previous rejection has been withdrawn. However, upon further consideration of the amended claims, a new ground of rejection is made below based on Hennecke in view of Sung.
Additionally, the applicant argues that “Sung's speech recognizers 120a-120n are not individually configured to compare an input speech signal to "an associated subset of the plurality of languages," let alone determine match scores or output a candidate "best match language," as generally recited in claim 1.”
The examiner notes, that Sung expressly teaches in para 0049, “at least one of the language models may be for multiple, different languages. In this regard, some language models may include elements of more than one language”. Thus, Sung is not limited to one language per recognizer.
(ii). “The cited references fail to disclose or suggest a ‘generating, by a further processing path downstream of the plurality of speech recognition processing paths, similarity scores quantifying similarity between the input speech signal and each the plurality of candidate languages.’”
The examiner agrees that the previously applied combination of Sung and Abdelhady does not teach or suggest the claimed ‘further processing downstream…’ in the manner recited in the amended claims. Accordingly, the previous rejection has been withdrawn. However, upon further consideration of the amended claims, a new ground of rejection is made below based on Hennecke in view of Sung.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1-6, 8-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hennecke (US20060206331) in view of Sung (US 20130238336 A1).
Regarding claim 1, Hennecke teaches receiving an input speech signal in a particular language of a plurality of languages; (Para 0017, “Speech input … is input to a plurality of speech recognition me subword units 100 and configured to recognize subword unit strings for different languages.”); Concurrently processing the input speech signal by a plurality of speech recognition processing paths (Para 0017, “the speech recognition subword module 120, 122, 124, 126 and 128 may operate in parallel using separate recognition modules”), generating, by a further processing path downstream of the plurality of speech recognition processing paths, similarity scores quantifying similarity between the input speech signal and each of the plurality of candidate languages (Para 0030, “The second speech recognition module 104 is configured to recognize, from the same speech input 110, the best matching item among the items listed in the candidate list 114” and “The second speech recognition module 104 compares the speech input 110 with acoustic representations of the items in the candidate list 114 and calculates a measure of similarity between the acoustic representations of items in the candidate list 114 and the speech input 110.”, where the ‘measure of similarity’ corresponds to the claimed similarity score, because it quantified how the input speech matches each item in the candidate list); and outputting, based on the similarity scores, an indication of a closest match language selected from the plurality of candidate languages. (Para 0032, “The best matching item from the candidate list 114 is selected and corresponding information indicating the selected item is output from the second speech recognition module 104.”).
Hennecke does not teach each speech recognition processing path including a machine-learning model configured to: compare the input speech signal to an associated subset of the plurality of languages and determine a match score for each language; based on the match score computed for each language, identify a candidate language from the associated subset of the plurality of languages that is a closest match to the particular language of the input speech signal; and output an identification of the candidate language, wherein a plurality of candidate languages are identified within outputs generated by the plurality of speech recognition processing paths
However, Sung teaches each speech recognition processing path including a machine-learning model configured to: compare the input speech signal to an associated subset of the plurality of languages and determine a match score for each language (Para 0047, “The recognition candidates are associated with corresponding scores. The score for each recognition candidate is indicative of the statistical likelihood that the recognition candidate is an accurate match for the input audio.” And Para 0049, “at least one of the language models may be for multiple, different languages.”); based on the match score computed for each language, identify a candidate language from the associated subset of the plurality of languages that is a closest match to the particular language of the input speech signal (Para 0050, “The scores of the multiple recognition candidates and the language scores may be used by the speech recognizer 210 to identify one of the candidates as the correct recognition of the input audio”); and output an identification of the candidate language, wherein a plurality of candidate languages are identified within outputs generated by the plurality of speech recognition processing paths (Para 0050, “multiple language recognizers may each provide one or more recognition candidates for an utterance in their respective languages.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Hennecke in order to incorporate the teachings of Sung in order to permit each recognition path to evaluate multiple possible language, and improve flexibility (Para 0049).
Regarding claim 2, Hen teaches wherein each speech recognition processing path associated subset includes languages not in any other speech recognition processing path associated subset (Para 0017, “a plurality of speech recognition me subword units 100 and configured to recognize subword unit strings for different languages.”, under the broadest reasonable interpretation of the word subset, a grouping of 1 comprises a subset).
Hennecke does not teach wherein each recognizer can be configured to recognize a multitude of languages.
However, Sung teaches wherein each recognizer can be configured to recognize a multitude of languages (Para 0049, “at least one of the language models may be for multiple, different languages.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Hennecke in order to incorporate the teachings of Sung in order to reduce redundant processing and thus computational load (Para 0049).
Regarding claim 3, Hennecke teaches caching the input speech signal for input to the further processing path downstream of the plurality of speech recognition processing paths. (Para 0030, “The second speech recognition module 104 compares the speech input 110 with acoustic representations of the items in the candidate list 114”, by storing the speech input 110 for comparison in this second speech, recognition module, the system inherently caches the signal).
Regarding claim 4, Hennecke does not teach wherein each speech recognition processing path associated subset includes languages common with languages in another speech recognition processing path associated subset.
However, Sung teaches wherein each speech recognition processing path associated subset includes languages common with languages in another speech recognition processing path associated subset. (Para 0049, “For example, if an utterance is recognized as a word in an English language model, and the same word is recognized in a Mandarin language model that includes English elements”).
It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Hennecke in order to incorporate the teachings of Sung because redundant recognitions may increase the likelihood that the word was identified correctly (Para 0049).
Regarding claim 5, Hennecke teaches wherein the input speech signal is streamed simultaneously to each of the plurality of speech recognition processing paths (Para 0017, “Speech input 110 from a user for selecting an item from a list of items 112 is input to a plurality of speech recognition me subword units” and “the speech recognition subword module 120, 122, 124, 126 and 128 may operate in parallel using separate recognition modules” wherein the input is provided to each of the recognizers in parallel).
Regarding claim 6, Hennecke teaches wherein the plurality of speech recognition processing paths operate in parallel with each other (Para 0017, “the speech recognition subword module 120, 122, 124, 126 and 128 may operate in parallel using separate recognition modules”).
Claims 8-14 are analogous to claims 1-7, respectively as they recite the same abstract
idea in a different statutory form (manufacture vs. method), and therefore are rejected for similar reasons.
Claims 15-20 are analogous to claims 1-6, respectively as they recite the same abstract
idea in a different statutory form (system vs. method), and therefore are rejected for similar
reasons.
Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Hennecke (US20060206331) in view of Sung et al. (US 20130238336 A1) as applied to claims 1, 2, 3, 4, 5, 6 as above, and further in view of Snyder et al., ("Spoken Language Recognition Using X-vectors", In Odyssey, 2018).
Hennecke modified by Sung does not teach wherein the plurality of speech recognition processing paths comprise a deep neural network utilizing an x-vector framework.
However, Snyder teaches wherein the plurality of speech recognition processing paths comprise a deep neural network utilizing an x-vector framework (Abstract, “This framework
consists of a deep neural network that maps sequences of speech features to fixed -dimensional
embeddings, called x-vectors.”).
It would have been obvious to one of ordinary skill in the art to modify Hennecke before the effective filing date in order to incorporate the teachings of Snyder to gain the benefit of better performance compared to other similar frameworks (Abstract, “In the 2017 NIST language recognition evaluation, x-vectors achieved excellent results and outperformed our state-of the-art i-vector systems.”).
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHAEL ALAN FOSTER JR. whose telephone number is (571)272-8874. The examiner can normally be reached M - F 8:00am - 5:00pm, Alternate Fridays Off.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Hai Phan can be reached at (571) 272-6338. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MICHAEL A FOSTER JR/ Examiner, Art Unit 2654
/Richa Sonifrank/ Primary Examiner, Art Unit 2654