Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-10, 11-21 20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter without significantly more. The claims as whole, considering all claim elements both individually and in combination, do not amount to significantly more than an abstract idea.
The group of independent claims 1, 12, and 13 disclose the same inventive concept with slightly different claim language. In this analysis, claim 1 is analyzed as a representative of the other independent claims. The independent claim 1 recites: “ … obtaining a first speech; obtaining a first text corresponding to a previous segment of speech of the first speech; obtaining a first set, the first set comprising a plurality of text identifications and a text feature corresponding to each of the plurality of text identifications, the text feature being a feature associated with a plurality of subsequent texts of a text corresponding to the text identification, the text feature being associated with frequencies of the plurality of subsequent texts of the text in a text set, and the first set being determined based on the text set; and determining, based on the first text and the first set, text content associated with the first speech”. The aforementioned limitations, under its broadest reasonable interpretation, cover performance of the limitation in the mind but for the recitation of generic computer components. That is, other than reciting “processor”, nothing in the claim element precludes the step from practically being performed in the mind. For example, but for the “processor” language “,
obtaining a first speech… in the context of this claim it amounts to receiving a question from another person
Obtaining a first text… in the context of this claim a human writing the previous words in a notebook.
obtaining a first set… in the context of this claim a human writing the previous words in a notebook that are corrected with proper context
determining, based on the first…, in the context of this claim, the human predicting based on the previous words with contextual information, the next word after listening to it and correcting for context as necessary
All of these steps (as noted above) can be performed in the mind and/or using a pen and paper. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
This judicial exception is not integrated into a practical application. In particular, the claim only recites additional elements - “processor” to perform all of the above-mentioned steps. “processor” is recited at a high-level of generality (i.e., as a generic computer device performing a generic computer function) such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The only element mentioned is the usage of a “vision language model”, which due to lack of specificity can be considered as a generic processor. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of a processor is merely for the purpose of data gathering and/or insignificant extra-solution activity that amount to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. The claim is not patent eligible.
Similarly, dependent claims 2-10, 14-17 and 18-21 are also not patent eligible as they include additional steps that are directed towards an abstract idea as they can be practically performed in the mind without being integrated into a practical application or including any additional elements sufficient to amount to significantly more than the judicial exception.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1, 2, 7, 12, 13, 14 and 18 is (are) rejected under 35 U.S.C. 103 as being unpatentable over Huang (US 20220310067 A1) in further view of Choi (US 20120023433 A1)
With respect to claims 1, 12 and 13 Huang teaches (Claim 1.) (Original) A speech recognition method, comprising: (Claim 12 )An electronic device, comprising: a processor and a memory, wherein the memory stores computer-executable instructions [0050] These computer programs (also known as programs, software, software applications or code) include machine instructions for a programmable processor, and can be implemented in a high-level procedural and/or object-oriented programming language, and/or in assembly/machine language. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, non-transitory computer readable medium, apparatus and/or device (e.g., magnetic discs, optical disks, memory, Programmable Logic Devices (PLDs)) used to provide machine instructions and/or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and/or data to a programmable processor); and the processor executes the computer-executable instructions stored in the memory, to cause the processor to ([0050] These computer programs (also known as programs, software, software applications or code) include machine instructions for a programmable processor, and can be implemented in a high-level procedural and/or object-oriented programming language, and/or in assembly/machine language. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, non-transitory computer readable medium, apparatus and/or device (e.g., magnetic discs, optical disks, memory, Programmable Logic Devices (PLDs)) used to provide machine instructions and/or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and/or data to a programmable processor.): (Claim 13) A non-transitory computer-readable storage medium, storing computer- executable instructions that, when executed by a processor, cause the processor to ([0050] These computer programs (also known as programs, software, software applications or code) include machine instructions for a programmable processor, and can be implemented in a high-level procedural and/or object-oriented programming language, and/or in assembly/machine language. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, non-transitory computer readable medium, apparatus and/or device (e.g., magnetic discs, optical disks, memory, Programmable Logic Devices (PLDs)) used to provide machine instructions and/or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and/or data to a programmable processor):
obtaining a first speech (Huang ¶[0031] The first embedding table 310 is configured to generate a token embedding 312 for a current token 133, t.sub.i in the sequence of tokens 133. Here, t.sub.i denotes the current token 133 in the sequence of tokens 133. The first embedding table 310 determines a respective token embedding 312 for each token 133 independent from the rest of the tokens 133 in the sequence of tokens 133. On the other hand, the second embedding table 320 includes an n-gram embedding table that receives a previous [first speech] n-gram token sequence 133, t.sub.0, . . . , t.sub.n-1 at the current output step (e.g. time step) and generates a respective n-gram token embedding 322. The previous n-gram token sequence (e.g., t.sub.0, . . . , t.sub.n-1) provides context information about the previously generated tokens 133 at each time step. The previous n-gram token sequence t.sub.0, . . . , t.sub.n-1 grows exponentially with n at each time step. Thus, the n-gram token embedding 322 at each output steps assists with short-range dependencies such as spelling out rare words, and thereby improves the modeling of long tail tokens (e.g., words) in subword models.)
obtaining a first text corresponding to a previous segment of speech of the first speech (Huang ¶[0033] At a single output step in the example shown, the second embedding table 320 may generate the n-gram token embedding 322 to represent the entire sequence of three previous n-gram tokens (e.g., “driving directions to”) [first set ]from the current token 133 (e.g., “bourbon) represented by the respective token embedding 312 looked-up by the first embedding table 310. Thereafter, the concatenator 330 may concatenate the n-gram token embedding 322 and the respective token embedding 312 to generate the concatenated output 335 (e.g., “driving directions to bourbon”).The RNN 340 rescores the candidate transcription 132 by processing the concatenated output 335. For example, the RNN 340 may rescore the candidate transcription 132 [text set](e.g., driving directions to bourbon) such that the RNN 340 boosts the likelihood/probability score of “Beaubien” to now have a higher likelihood/probability score [feature based on context of subsequent text] for the fourth token 133d than “bourbon.” The RNN 340 boosts the likelihood/probability score of “Beaubien” based, in part, on the determination that “bourbon” is not likely the correct token 133 with the current context information (e.g., previous n-gram token sequence t.sub.0, . . . , t.sub.n-1) [index 0 to n-1 refers to precious text]. Accordingly, the RNN 340 generates the restored transcription 345 [text context associated with first speech] (e.g., driving directions to Beaubien). Notably, the fourth token 133 is now correctly recognized in the rescored transcription 345 even though the fourth token 133 was misrecognized in the candidate transcription 132);
obtaining a first set, the first set comprising a plurality of text identifications and a text feature corresponding to each of the plurality of text identifications, the text feature being a feature associated with a plurality of subsequent texts of a text corresponding to the text identification, the text feature [[being associated with frequencies of the plurality of subsequent texts of the text in a text set, and the first set being determined based on the text set]] (Huang ¶[0033] At a single output step in the example shown, the second embedding table 320 may generate the n-gram token embedding 322 to represent the entire sequence of three previous n-gram tokens (e.g., “driving directions to”) [first set ]from the current token 133 (e.g., “bourbon) represented by the respective token embedding 312 looked-up by the first embedding table 310. Thereafter, the concatenator 330 may concatenate the n-gram token embedding 322 and the respective token embedding 312 to generate the concatenated output 335 (e.g., “driving directions to bourbon”).The RNN 340 rescores the candidate transcription 132 by processing the concatenated output 335. For example, the RNN 340 may rescore the candidate transcription 132 [text set](e.g., driving directions to bourbon) such that the RNN 340 boosts the likelihood/probability score of “Beaubien” to now have a higher likelihood/probability score [feature based on context of subsequent text ]for the fourth token 133d than “bourbon.” The RNN 340 boosts the likelihood/probability score of “Beaubien” based, in part, on the determination that “bourbon” is not likely the correct token 133 with the current context information (e.g., previous n-gram token sequence t.sub.0, . . . , t.sub.n-1) [index 0 to n-1 refers to precious text]. Accordingly, the RNN 340 generates the restored transcription 345 [text context associated with first speech] (e.g., driving directions to Beaubien). Notably, the fourth token 133 is now correctly recognized in the rescored transcription 345 even though the fourth token 133 was misrecognized in the candidate transcription 132); and
determining, based on the first text and the first set, text content associated with the first speech (Huang ¶[0033] At a single output step in the example shown, the second embedding table 320 may generate the n-gram token embedding 322 to represent the entire sequence of three previous n-gram tokens (e.g., “driving directions to”) [first set ]from the current token 133 (e.g., “bourbon) represented by the respective token embedding 312 looked-up by the first embedding table 310. Thereafter, the concatenator 330 may concatenate the n-gram token embedding 322 and the respective token embedding 312 to generate the concatenated output 335 (e.g., “driving directions to bourbon”).The RNN 340 rescores the candidate transcription 132 by processing the concatenated output 335. For example, the RNN 340 may rescore the candidate transcription 132 [text set](e.g., driving directions to bourbon) such that the RNN 340 boosts the likelihood/probability score of “Beaubien” to now have a higher likelihood/probability score [feature based on context of subsequent text ]for the fourth token 133d than “bourbon.” The RNN 340 boosts the likelihood/probability score of “Beaubien” based, in part, on the determination that “bourbon” is not likely the correct token 133 with the current context information (e.g., previous n-gram token sequence t.sub.0, . . . , t.sub.n-1) [index 0 to n-1 refers to precious text]. Accordingly, the RNN 340 generates the restored transcription 345 [text context associated with first speech] (e.g., driving directions to Beaubien). Notably, the fourth token 133 is now correctly recognized in the rescored transcription 345 even though the fourth token 133 was misrecognized in the candidate transcription 132)
Huang does not explicitly disclose however Choi teaches text feature being associated with frequencies of the plurality of subsequent texts of the text in a text set (Choi claim ¶16. The apparatus of claim 12, further comprising a storage unit for storing a plurality of words, wherein the controller predicts the next input character based on a usage frequency of each of words corresponding to the input character.)
It would have been obvious to one of ordinary skill in the art prior to the effective filing date of the invention to modify text prediction of Huang to include the features of Choi in order to enhance contextual mapping and improved prediction.
With respect to claims 2, 14 and 18 Huang teaches determining, based on the first text and the first set, a next segment of second text of the first text ( ¶[0033] At a single output step in the example shown, the second embedding table 320 may generate the n-gram token embedding 322 to represent the entire sequence of three previous n-gram tokens (e.g., “driving directions to”) [first set ]from the current token 133 (e.g., “bourbon) represented by the respective token embedding 312 looked-up by the first embedding table 310. Thereafter, the concatenator 330 may concatenate the n-gram token embedding 322 and the respective token embedding 312 to generate the concatenated output 335 (e.g., “driving directions to bourbon”).The RNN 340 rescores the candidate transcription 132 by processing the concatenated output 335. For example, the RNN 340 may rescore the candidate transcription 132 [text set](e.g., driving directions to bourbon) such that the RNN 340 boosts the likelihood/probability score of “Beaubien” to now have a higher likelihood/probability score [feature based on context of subsequent text ]for the fourth token 133d than “bourbon.” The RNN 340 boosts the likelihood/probability score of “Beaubien” based, in part, on the determination that “bourbon” is not likely the correct token 133 with the current context information (e.g., previous n-gram token sequence t.sub.0, . . . , t.sub.n-1) [index 0 to n-1 refers to precious text]. Accordingly, the RNN 340 generates the restored transcription 345 [text context associated with first speech] (e.g., driving directions to Beaubien). Notably, the fourth token 133 [next segment of second text]is now correctly recognized in the rescored transcription 345 even though the fourth token 133 was misrecognized in the candidate transcription 132. [Examiner Note: at the next instance tn+1, the input the second embedding has the index 0, to n and thus the next segment for second text is processed and the output has concatenation of the next segment]); and
determining, based on the second text and the first speech, the text content associated with the first speech ( ¶[0033] At a single output step in the example shown, the second embedding table 320 may generate the n-gram token embedding 322 to represent the entire sequence of three previous n-gram tokens (e.g., “driving directions to”) [first set ]from the current token 133 (e.g., “bourbon) represented by the respective token embedding 312 looked-up by the first embedding table 310. Thereafter, the concatenator 330 may concatenate the n-gram token embedding 322 and the respective token embedding 312 to generate the concatenated output 335 (e.g., “driving directions to bourbon”).The RNN 340 rescores the candidate transcription 132 by processing the concatenated output 335. For example, the RNN 340 may rescore the candidate transcription 132 [text set](e.g., driving directions to bourbon) such that the RNN 340 boosts the likelihood/probability score of “Beaubien” to now have a higher likelihood/probability score [feature based on context of subsequent text ]for the fourth token 133d than “bourbon.” The RNN 340 boosts the likelihood/probability score of “Beaubien” based, in part, on the determination that “bourbon” is not likely the correct token 133 with the current context information (e.g., previous n-gram token sequence t.sub.0, . . . , t.sub.n-1) [index 0 to n-1 refers to precious text]. Accordingly, the RNN 340 generates the restored transcription 345 [text context associated with first speech] (e.g., driving directions to Beaubien). Notably, the fourth token 133 [next segment of second text]is now correctly recognized in the rescored transcription 345 even though the fourth token 133 was misrecognized in the candidate transcription 132. Examiner Note: at the next instance tn+1, the input the second embedding has the index 0, to n and thus the next segment for second text is processed and the output has concatenation of the next segment.
With respect to claim 7 Huang teaches wherein determining, based on the second text and the first speech, the text content associated with the first speech comprises: performing text recognition on the first speech to obtain a third text; and determining, based on the second text and the third text, the text content associated with the first speech( ¶[0033] At a single output step in the example shown, the second embedding table 320 may generate the n-gram token embedding 322 to represent the entire sequence of three previous n-gram tokens (e.g., “driving directions to”) [first set ]from the current token 133 (e.g., “bourbon) represented by the respective token embedding 312 looked-up by the first embedding table 310. Thereafter, the concatenator 330 may concatenate the n-gram token embedding 322 and the respective token embedding 312 to generate the concatenated output 335 (e.g., “driving directions to bourbon”).The RNN 340 rescores the candidate transcription 132 by processing the concatenated output 335. For example, the RNN 340 may rescore the candidate transcription 132 [text set](e.g., driving directions to bourbon) such that the RNN 340 boosts the likelihood/probability score of “Beaubien” to now have a higher likelihood/probability score [feature based on context of subsequent text ]for the fourth token 133d than “bourbon.” The RNN 340 boosts the likelihood/probability score of “Beaubien” based, in part, on the determination that “bourbon” is not likely the correct token 133 with the current context information (e.g., previous n-gram token sequence t.sub.0, . . . , t.sub.n-1) [index 0 to n-1 refers to precious text]. Accordingly, the RNN 340 generates the restored transcription 345 [text context associated with first speech] (e.g., driving directions to Beaubien). Notably, the fourth token 133 [next segment of second text]is now correctly recognized in the rescored transcription 345 even though the fourth token 133 was misrecognized in the candidate transcription 132. [Examiner Note: at the next instance tn+1, the input the second embedding has the index 0, to n and thus the next segment for second text is processed and the output has concatenation of the next segment])
Allowable Subject Matter
Claims 3-6, 8-10, 15-16 and 19-21 are objected to as being dependent upon a rejected base claim, but would be allowable pending overcoming 101 rejections set forth in this Office Action, if written in independent form including all of the limitations of the base claim and any intervening claims.
Claims 3, 15 and 19 recites “… wherein determining, based on the first text and the first set, the next segment of second text of the first text comprises: obtaining a first identification of the first text; obtaining, based on the first identification, a first text feature associated with a plurality of subsequent texts of the first text from the first set; and determining the second text based on the first text and the first text feature.
The closest prior art of record to the currently claimed limitations is Huang who teaches “¶[0033] At a single output step in the example shown, the second embedding table 320 may generate the n-gram token embedding 322 to represent the entire sequence of three previous n-gram tokens (e.g., “driving directions to”) [first set ]from the current token 133 (e.g., “bourbon) represented by the respective token embedding 312 looked-up by the first embedding table 310. Thereafter, the concatenator 330 may concatenate the n-gram token embedding 322 and the respective token embedding 312 to generate the concatenated output 335 (e.g., “driving directions to bourbon”).The RNN 340 rescores the candidate transcription 132 by processing the concatenated output 335. For example, the RNN 340 may rescore the candidate transcription 132 [text set](e.g., driving directions to bourbon) such that the RNN 340 boosts the likelihood/probability score of “Beaubien” to now have a higher likelihood/probability score [feature based on context of subsequent text ]for the fourth token 133d than “bourbon.” The RNN 340 boosts the likelihood/probability score of “Beaubien” based, in part, on the determination that “bourbon” is not likely the correct token 133 with the current context information (e.g., previous n-gram token sequence t.sub.0, . . . , t.sub.n-1) [index 0 to n-1 refers to precious text]. Accordingly, the RNN 340 generates the restored transcription 345 [text context associated with first speech] (e.g., driving directions to Beaubien). Notably, the fourth token 133 [next segment of second text]is now correctly recognized in the rescored transcription 345 even though the fourth token 133 was misrecognized in the candidate transcription 132”.
However, Huang, Choi individually or in combination do not teach, suggest or render obvious to one of ordinary skill in the art before the effective filing date, at least the specific limitations as recited above. Furthermore, it would not have been obvious to one of ordinary skill in the art to modify the prior art in order to arrive at the claimed invention. Claims 4-6, 16 and 20 are allowable based on dependency
Claim(s) 8, 17 and 21 is (are) objected to as being dependent upon a rejected base claim, but would be allowable if written in independent form including all of the limitations of the base claim and any intervening claims. Claims 8 and 21 recites “… wherein obtaining the first set comprises: obtaining sample identifications of a plurality of sample texts in the text set and sample text features corresponding to subsequent texts of the sample texts; determining an initial set based on the sample identifications and the sample text features, wherein the initial set comprises a plurality of sample identifications and a sample text feature corresponding to each sample identification; and updating, based on the plurality of sample texts, the plurality of sample text features in the initial set to obtain the first set.
The closest prior art of record to the currently claimed limitations is Huang who teaches “¶[0033] At a single output step in the example shown, the second embedding table 320 may generate the n-gram token embedding 322 to represent the entire sequence of three previous n-gram tokens (e.g., “driving directions to”) [first set ]from the current token 133 (e.g., “bourbon) represented by the respective token embedding 312 looked-up by the first embedding table 310. Thereafter, the concatenator 330 may concatenate the n-gram token embedding 322 and the respective token embedding 312 to generate the concatenated output 335 (e.g., “driving directions to bourbon”).The RNN 340 rescores the candidate transcription 132 by processing the concatenated output 335. For example, the RNN 340 may rescore the candidate transcription 132 [text set](e.g., driving directions to bourbon) such that the RNN 340 boosts the likelihood/probability score of “Beaubien” to now have a higher likelihood/probability score [feature based on context of subsequent text ]for the fourth token 133d than “bourbon.” The RNN 340 boosts the likelihood/probability score of “Beaubien” based, in part, on the determination that “bourbon” is not likely the correct token 133 with the current context information (e.g., previous n-gram token sequence t.sub.0, . . . , t.sub.n-1) [index 0 to n-1 refers to precious text]. Accordingly, the RNN 340 generates the restored transcription 345 [text context associated with first speech] (e.g., driving directions to Beaubien). Notably, the fourth token 133 [next segment of second text]is now correctly recognized in the rescored transcription 345 even though the fourth token 133 was misrecognized in the candidate transcription 132”.
Schroeder teaches “(Claims (h) otherwise continuing at step (c), wherein the initial character subset is statistically determined from sample text to be the most common initial characters of words appearing in such sample text and the initial character subset is periodically updated by analyzing the character frequencies of messages entered by a user over time.)”
However, Huang, Choi and Schroeder individually or in combination do not teach, suggest or render obvious to one of ordinary skill in the art before the effective filing date, at least the specific limitations as recited above. Furthermore, it would not have been obvious to one of ordinary skill in the art to modify the prior art in order to arrive at the claimed invention. Claims 9-10 are allowable based on dependency.
Any comments considered necessary by applicant must be submitted no later than the payment of the issue fee and, to avoid processing delays, should preferably accompany the issue fee. Such submissions should be clearly labeled “Comments on Statement of Reasons for Allowance.”
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ATHAR N PASHA whose telephone number is (408)918-7675. The examiner can normally be reached Monday-Thursday Alternate Fridays, 7:30-4:30 PT.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at (571)272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ATHAR N PASHA/Primary Examiner, Art Unit 2657