Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of the claims
Claims 1-20 are presented for examination.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 7/10/2024 was filed before the
mailing date of the first office action. The submission is in compliance with the provisions of 37
CFR 1.97. Accordingly, the information disclosure statement is being considered by the
examiner.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed
to an abstract idea without significantly more. The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception because as
explained below.
Claim 1 recites a computer-implemented method, executed on a computing device, comprising:
(a). receiving an input speech signal in a particular language of a plurality of languages;
(b). processing the input speech signal by a plurality of speech recognition processing paths, each speech recognition processing path being configured to recognize an associated subset of the plurality languages;
(c). each of the plurality of speech recognition processing paths processing the input speech signal using machine learning to identify a language in the associated subset of languages which is a closest match to the particular language of the input speech signal
(d). the processing of the input speech signal by the plurality of speech recognition processing paths resulting in a plurality of identified languages;
(e). receiving the input speech signal and an indication of each of the plurality of identified languages in a further speech recognition processing path
(f). processing, using machine learning, the input speech signal to recognize one of the plurality of identified languages as a closest match to the particular language of the input speech signal.
Step (a) comprises a mental process. This step can be performed by a human as a person can receive spoken input in a particular language.
Step (b) comprises a mental process. This step can be performed by a human as a person can consider multiple possible languages and narrow recognition to a subset of languages.
Step (c) comprises a mental process. This step can be performed by a human as a person can analyze speech and determine which language from a subset is the closest match.
Step (d) comprises a mental process. This step can be performed by a human as a person can receive an indication of possible identified languages and further consider them.
Step (e) comprises a mental process. This step can be performed by a human as a person can receive the spoken input and consider the previously identified possible languages for further evaluation.
Step (f) comprises a mental process. This step can be performed by a human as a person can evaluate the possible identified languages and select the closest match.
Step 1: This part of the eligibility analysis evaluates whether the claim falls within any
statutory category. See MPEP 2106.03. The claim recites at least system. Thus, the claim is a
machine, which is one of the statutory categories of invention. (Step 1: YES).
Step 2A, Prong One: This part of the eligibility analysis evaluates whether the
claim recites a judicial exception. As explained in MPEP 2106.04, subsection II, a claim
“recites” a judicial exception when the judicial exception is “set forth” or “described” in
the claim. As discussed above, the broadest reasonable interpretation of steps (a)-(f)
recites a mental process.
Specifically, step (a) can be performed by a human as a person can receive spoken input in a particular language.
Step (b) can be performed by a human as a person can consider multiple possible languages and narrow recognition to a subset of languages.
Step (c) can be performed by a human as a person can analyze speech and determine which language from a subset is the closest match.
Step (d) can be performed by a human as a person can consider multiple possible identified languages resulting from analysis.
Step (e) can be performed by a human as a person can receive an indication of possible identified languages and further consider them.
Step (f) can be performed by a human as a person can evaluate the possible identified languages and select the closest match.
Hence the claim encompasses mental processes practically performed in the human mind
by observation, evaluation, judgement, and opinion. See MPEP 2106.04(a)(2), subsection III.
(Step 2A, Prong One: YES).
Step 2A, Prong Two: This part of the eligibility analysis evaluates whether the claim as a whole integrates the recited judicial exception into a practical application of the exception or whether the claim is “directed to” the judicial exception. This evaluation is performed by (1) identifying whether there are any additional elements recited in the claim beyond the judicial exception, and (2) evaluating those additional elements individually and in combination to determine whether the claim as a whole integrates the exception into a practical application. See MPEP 2106.04(d).
The claim recites additional elements including at least one processor, at least one memory storing executable instructions, and the use of machine learning to process the input speech signal. However, the at least one processor and at least one memory are recited at a high level of generality and perform generic computer functions, such as executing instructions and storing data. The use of machine learning to process the input speech signal constitutes a generic computational technique that merely automates the mental processes described above. Such processing does not impose any meaningful limit on the judicial exception. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. (Step 2A, Prong Two: NO), and the claim is directed to the judicial exception. (Step 2A: YES).
Step 2B: This part of the eligibility analysis evaluates whether the claim as a whole amounts to significantly more than the recited exception i.e., whether any additional element, or combination of additional elements, adds an inventive concept to the claim. As explained with respect to Step 2A, Prong Two, the processor, memory, and machine learning processing comprise additional elements that do not contribute to the patentability of the claim as a whole. The additional element of the machine learning processing in limitations (b)-(f) is at best mere instructions to “apply” the abstract ideas, which cannot provide an inventive concept. See MPEP 2106.05(f). At Step 2B, the evaluation of the insignificant extra-solutional activity consideration takes into account whether or not the extra-solutional activity is well understood, routine, and conventional in the field. See MPEP 2106.05(g). As known in the art these elements are well routine and conventional. Even when considered in combination, these additional elements represent mere instructions to implement an abstract idea or other exception on a computer and insignificant extra-solutional activity which do not provide an inventive concept. The claim is not patent eligible.
Claim 2 constitutes categorization, this can be performed by a human as a person can assign languages to distinct groups such that each group is separate from others.
Claim 3 constitutes data gathering, this can be performed by a human as a person can retain or remember input information for later use.
Claim 4 constitutes categorization, this can be performed by a human as a person can assign overlapping groups of languages based on shared characteristics.
Claim 5 constitutes data processing, this can be performed by a human as a person can provide the same input to multiple evaluators.
Claim 6 constitutes organization of human activity, this can be performed by a human as multiple individuals can analyze the same input in parallel.
Claim 7 constitutes insignificant extra-solutional activity as it recites the use of a deep neural network, which is well-understood, routine, and conventional in the art.
Claims 8-14 are analogous to claims 1-7, respectively as they recite the same abstract idea in a different statutory form (manufacture vs. method), and therefore are rejected for similar reasons.
Claims 15-20 are analogous to claims 1-6, respectively as they recite the same abstract idea in a different statutory form (system vs. method), and therefore are rejected for similar reasons.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1, 2, 4, 5 & 6 is/are rejected under 35 U.S.C. 103 as being unpatentable over Sung et al. (US 20130238336 A1) in view of AbdelHady at al. (US 12175968 B1).
Regarding claim 1, Sung teaches a computer-implemented method, executed on a computing device (Fig. 1), comprising: receiving an input speech signal in a particular language of a plurality of languages (Para 0022, “receives input audio spoken by the user”); processing the input speech signal by a plurality of speech recognition processing paths (Para 0024, “forwards that audio data to the speech recognition system 104. The speech recognition system 104 provides the audio data 110 to a collection of language recognition modules”), each speech recognition processing path being configured to recognize an associated subset of the plurality languages (Para 0049, “In some implementations, at least one of the language models may be for multiple, different languages.”); each of the plurality of speech recognition processing paths processing the input speech signal using machine learning to identify a language in the associated subset of languages which is a closest match to the particular language of the input speech signal (Para 0025, “Each of the language recognition modules 120a-120n represents a speech recognizer that is tuned to recognize speech in a particular one of the languages.”), the processing of the input speech signal by the plurality of speech recognition processing paths resulting in a plurality of identified languages (Para 0026, “Each of the language recognition modules 120a-120n identifies one or more recognition candidates for each of the spoken words included in the audio data.”); and processing, using machine learning, the input speech signal to recognize one of the plurality of identified languages as a closest match to the particular language of the input speech signal. (Para 0025, “Each of the language recognition modules 120a-120n represents a speech recognizer that is tuned to recognize speech in a particular one of the languages.”)
Sung does not teach receiving the input speech signal and an indication of each of the plurality of identified languages in a further speech recognition processing path.
However, AbdelHady does teach receiving the input speech signal and an indication of each of the plurality of identified languages in a further speech recognition processing path ("Alternatively, the user recognition component 295 may output multiple user identifiers (e.g., in the form of an N-best list) with respective values representing likelihoods of respective users originating the natural language input. The output of the user recognition component 295 may be used to inform NLU processing, processing performed by a skill 225, as well as processing performed by other components of the system 120 and/or other systems”, Col 11, line 27-37; here the recognition is further processed by the selected skill candidate ( selected skill candidate is selected in fig 7 or 9 ) hence further processing (Abstract, “The system determines first and second sets of skills corresponding to the locale/first language and locale/second language, respectively.”))
It would have been obvious to one of ordinary skill in the art to modify Sung before
the effective filing date in order to incorporate the teachings of AbdelHady to improve human interaction ( Col 1, line 5-20, AbdelHady ).
Regarding claim 2, Sung teaches wherein each speech recognition processing path associated subset includes languages not in any other speech recognition processing path associated subset. (Para 0049, “In some examples, linguistic overlap among two or more language models may improve speech recognition.” The word overlap implies that they share some but not all languages as if two recognizers fully overlapped it would lead to redundancy).
Regarding claim 4, Sung teaches wherein each speech recognition processing path associated subset includes languages common with languages in another speech recognition processing path associated subset (Para 0049, “In some examples, linguistic overlap among two or more language models may improve speech recognition.” and “For example, if an utterance is recognized as a word in an English language model, and the same word is recognized in a Mandarin language model that includes English elements”).
Regarding claim 5, Sung teaches wherein the input speech signal is streamed simultaneously to each of the plurality of speech recognition processing paths (Para 0048, “the input audio may be provided to one or more language recognizer components and the language identifier module at substantially the same time”).
Regarding claim 6, Sung teaches wherein the plurality of speech recognition processing paths operate in parallel with each other (Para 0069, “the results of one or more language recognizer components (e.g., components 212a-212n), operating substantially in parallel”).
Claim 3 is rejected under 35 U.S.C. 103 as being unpatentable over Sung et al. (US 20130238336 A1) in view of AbdelHady at al. (US 12175968 B1), as applied to claims 1, 2, 4, 5, 6 above, and further in view of Carbune et al. (US 20220122610 A1).
Sung modified by AbdelHady does not teach caching the input speech signal for input to the further speech recognition processing path.
However, Carbune does teach caching the input speech signal for input to the further speech recognition processing path (Para 0011, “the user's utterance can be cached locally for further processing on the device”).
It would have been obvious to one of ordinary skill in the art to modify Sung before
the effective filing date in order to incorporate the teachings of Carbune to gain the benefit of allowing the speech input to be processed by both the first device and another later device, where both of the language recognizers comprise different devices (Para 0014, “user's utterance can be cached locally for further processing on the device where the second automated assistant run").
Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Sung et al. (US 20130238336 A1) in view of AbdelHady at al. (US 12175968 B1), as applied to claims 1, 2, 4, 5, 6 above, and further in view of Snyder et al., "Spoken Language Recognition Using X-vectors", In Odyssey, 2018.
Sung modified by AbdelHady does not teach wherein the plurality of speech recognition processing paths comprise a deep neural network utilizing an x-vector framework.
However, Snyder teaches wherein the plurality of speech recognition processing paths comprise a deep neural network utilizing an x-vector framework (Abstract, “This framework consists of a deep neural network that maps sequences of speech features to fixed-dimensional embeddings, called x-vectors.”).
It would have been obvious to one of ordinary skill in the art to modify Sung before
the effective filing date in order to incorporate the teachings of Snyder to gain the benefit of better performance compared to other similar frameworks (Abstract, “In the 2017 NIST language recognition evaluation, x-vectors achieved excellent results and outperformed our state-of the-art i-vector systems.”).
Claims 8-14 are analogous to claims 1-7, respectively as they recite the same abstract idea in a different statutory form (manufacture vs. method), and therefore are rejected for similar reasons.
Claims 15-20 are analogous to claims 1-6, respectively as they recite the same abstract idea in a different statutory form (system vs. method), and therefore are rejected for similar reasons.
Conclusion
Any inquiry concerning this communication or earlier communications from the
examiner should be directed to MICHAEL ALAN FOSTER JR. whose telephone number is
(571)272-8874. The examiner can normally be reached M - Th 8:00am - 6:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using
a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is
encouraged to use the USPTO Automated Interview Request (AIR) at
http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s
supervisor, Hai Phan can be reached at (571) 272-6338. The fax phone number for the
organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be
obtained from Patent Center. Unpublished application information in Patent Center is available
to registered users. To file and manage patent submissions in Patent Center, visit:
https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more
information about Patent Center and https://www.uspto.gov/patents/docx for information about
filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC)
at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service
Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/HAI PHAN/Supervisory Patent Examiner, Art Unit 2654