Prosecution Insights
Last updated: October 02, 2026
Application No. 18/231,164

SPEECH RECOGNITION BASED ON BACKGROUND FEATURES

Non-Final OA §101§103
Filed
Aug 07, 2023
Examiner
SWAMY, ARJUN RAJ
Art Unit
Tech Center
Assignee
International Business Machines Corporation
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
14 currently pending
Career history
11
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Examiner’s Note While Claims 8-14 do not recite non-transitory storage medium, the specification in Paragraph 0019 discloses “A computer readable storage medium, as that term is used in the present disclosure, is not to be construed as storage in the form of transitory signals per se”. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim(s) recite(s) elements which under their broadest reasonable interpretation are directed to mental processes. This judicial exception is not integrated into a practical application as explained below. The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception as explained below. Regarding Claim 15, the claim recites a system performing: receiving speech input from a user; receiving background inputs describing at least one of a webpage and an application from which the speech input was received; determining a language associated with the speech input; determining a weight and confidence level for each of the background inputs; calculating a score for each of the background inputs based on the corresponding weight and confidence level; determining a plurality of textual candidates for output in response to the speech input; ranking the plurality of textual candidates using a function that takes into account the scores and the confidence levels of the background inputs; and returning, to the user, the textual candidate with the highest ranking. Claim Interpretation: Under the broadest reasonable interpretation, the terms of the claim are presumed to have their plain meaning consistent with the specification as it would be interpreted by one of ordinary skill in the art. See MPEP 2111. receiving speech input from a user. A human can receive a speech input from a user receiving background inputs describing at least one of a webpage and an application from which the speech input was received. A human can receive a description of what applications and webpages a user has open determine a language from speech. Humans regularly determine a language from speech determine a weight and confidence. A human can make a determination of what weight and confidence to give to a value calculating a score. A mathematical process which a human can perform determine what to output in response. Humans regularly determine what to respond with given a speech input Ranking the candidates using a score. A mathematical process which a human can perform Returning to the user the candidate. A human can respond with the candidate in mind in response to the user. Additional elements are processor and memory device. Step 1: This part of the eligibility analysis evaluates whether the claim falls within any statutory category. See MPEP 2106.03. The claim recites at least one step or act, including receiving continuous training data. Thus, the claim is to a system, which is one of the statutory categories of invention. (Step 1: YES). Step 2A, Prong One: This part of the eligibility analysis evaluates whether the claim recites a judicial exception. As explained in MPEP 2106.04, subsection II, a claim “recites” a judicial exception when the judicial exception is “set forth” or “described” in the claim. As described above, the broadest reasonable interpretation of (a)-(h) is that those steps fall within the mental process groupings of abstract ideas because they cover concepts performed in the human mind, including observation, evaluation, judgment, and opinion. See MPEP 2106.04(a)(2), subsection III. Claim element (a) is directed to a mental step as a human can receive a speech input. Claim element (b) is directed to a mental step as a human can receive an input describing what a user is looking at. Element (c) is directed to a mental step as a human can determine what language a speech input was in. Element (d) is directed to a mental step as a human can determine the weight and confidence of values. Element (e) is directed to a mental step as calculating a score is a mathematical calculation that can be performed in a human mind. Element (f) is directed to a mental step as a human can determine options of what to respond with given a user’s speech input. Element (g) is directed to a mental step as a human can rank/weigh their options given a score or confidence they assigned prior. Element (h) is directed to a mental step as a human can return their option to the user in a variety of manners. Hence, these steps can be performed by human, using “observation, evaluation, judgment, [and] opinion,” because they involve making determinations and identifications, which are mental tasks humans routinely do,' ” and thus can practically be performed in the human mind, In re Killian, 45 F.4th 1373, 1379 (Fed. Cir. 2022). Therefore, these limitations are considered together as an abstract idea for further analysis. (Step 2A, Prong One: YES). Step 2A, Prong Two: This part of the eligibility analysis evaluates whether the claim as a whole integrates the recited judicial exception into a practical application of the exception or whether the claim is “directed to” the judicial exception. This evaluation is performed by (1) identifying whether there are any additional elements recited in the claim beyond the judicial exception, and (2) evaluating those additional elements individually and in combination to determine whether the claim as a whole integrates the exception into a practical application. See MPEP 2106.04(d). The additional elements are processor and memory device. These additional elements provide nothing more than mere instructions to implement an abstract idea on a generic computer. See MPEP 2106.05(f). MPEP 2106.05(f) provides the following considerations for determining whether a claim simply recites a judicial exception with the words “apply it” (or an equivalent), such as mere instructions to implement an abstract idea on a computer: (1) whether the claim recites only the idea of a solution or outcome i.e., the claim fails to recite details of how a solution to a problem is accomplished; (2) whether the claim invokes computers or other machinery merely as a tool to perform an existing process; and (3) the particularity or generality of the application of the judicial exception. Even when viewed in combination, these additional elements do not integrate the recited judicial exception into a practical application (Step 2A, Prong Two: NO), and the claim is directed to the judicial exception. (Step 2A: YES). Step 2B: This part of the eligibility analysis evaluates whether the claim as a whole amount to significantly more than the recited exception, i.e., whether any additional element, or combination of additional elements, adds an inventive concept to the claim. See MPEP 2106.05. At Step 2A, the additional elements of processor and memory device were found to represent no more than mere instructions to apply the judicial exception on a computer using generic computer components. Mere instructions to “apply” the abstract ideas, cannot provide an inventive concept. See MPEP 2106.05(f). The analysis under Step 2A, Prong Two is carried through to Step 2B. Even when considered in combination, these additional elements represent mere instructions to implement an abstract idea or other exception on a computer and insignificant extra-solution activity, which do not provide an inventive concept. (Step 2B: NO). As such Claim 15 is patent illegible The analysis above is applicable to Claims 1, 6, 8 and 13. Regarding Claims 2, 9 and 16, a human can determine a language by reading a language code listed on a website (e.g. en.wikipedia.org) Regarding Claims 3-5, 10-12 and 17-19, a human can parse descriptions related to input field, time and location and use that gained information to make determinations of confidence and weight. Regarding Claim 7, 14, 20 a human can return to a user their options with what rankings were assigned to them Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 1, 3, 6-8, 10, 13-15, 17, 20 are rejected under 35 U.S.C. 103 as being unpatentable over Ballinger(US PGPub 20220405046) in view of Mirkovic(US PGPub 20060271364). Regarding Claim 1, Ballinger teaches a method for improving the accuracy of speech recognition, the method comprising: receiving speech input from a user(spoken input from the user[Abstract]); receiving background inputs describing at least one of a webpage and an application from which the speech input was received(The context indicator can specify the context in which the user input is received, a webpage in which the user input is received, an application in which the user input is received[0012], the meta data may be used by the server system to identify a context in which the user is entering the spoken input[0031]); determining a language associated with the speech input(A language indicator can be provided to the remote server based on the determined language indicating the language of the spoken input[0012]); determining a weight and confidence level for each of the background inputs(a confidence level representing the likelihood or accuracy of the language model(the speech service 418 can map the context to broader categories or to the categories of the language models 414a-414n[0069]) to correctly identify the text in the utterance.[0088], Each referenced language model that is referenced in the data structure can have a weight.[0088]); determining a plurality of textual candidates for output in response to the speech input;(A list of candidates of text representing the spoken input[0011]); ranking the plurality of textual candidates; returning, to the user, the textual candidate with the highest ranking(The text 113, which can be a selected best candidate or can be a list of n-best candidates that correspond to the speech utterance, is provided back to the electronic device 104 over the network 112. The text 113 can be displayed to the user on a display 122 of the electronic device 104.[0035]). Ballinger does not teach calculating a score for each of the background inputs based on the corresponding weight and confidence level; However, Mirkovic teaches receiving background inputs from which the speech input was received (features from multiple sources of evidence regarding the speech recognition[0141]); calculating a score for each of the background inputs based on the corresponding weight and confidence level(the weighted confidence score for each feature of the input utterance is determined[0141]); ranking the plurality of textual candidates using a function that takes into account the scores and the confidence levels of the background inputs(The weighted confidence scores are then combined to rate the possible dialogue move candidates as the interpretation of the input utterance, 1404. Based on the highest confidence score, the optimum dialogue move is selected[0141]); It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention with the teachings of Ballinger to incorporate the weighted confidence score and subsequent rating of Mirkovic because it would allow for multiple sources of information to choose the highest scoring hypothesis(Abstract). Claims 8 and 15 recite similar limitations to Claim 1 and are rejected similarly. Regarding Claim 3, Ballinger teaches the background inputs further describe an input box from which the speech input was received (metadata schemes can define or describe the type of input field, such as an address field, a free form text field, a search field, or a social status field.[0091]). Claims 10 and 17 recite similar limitations to Claim 3 and are rejected similarly. Regarding Claim 6, Mirkovic teaches calculating a score for a background input comprises multiplying the weight of the background input by the confidence level of the background input(the weighted confidence score for each feature[0141]) Claim 13 recites similar limitations to Claim 6 and is rejected similarly. Regarding Claim 7, Ballinger teaches returning, to the user, the plurality of textual candidates and their rankings(The text 113, which can be a selected best candidate or can be a list of n-best candidates that correspond to the speech utterance, is provided back to the electronic device 104 over the network 112. The text 113 can be displayed to the user on a display 122 of the electronic device 104.[0035]). Claims 14 and 20 recite similar limitations and are rejected similarly Claim(s) 2, 9, 16 are rejected under 35 U.S.C. 103 as being unpatentable over Ballinger(US PGPub 20220405046) in view of Mirkovic(US PGPub 20060271364) as applied to claim 1 above, and further in view of Shires(Web Speech API Specification). Regarding Claim 2, neither Ballinger nor Mirkovic teach determining the language comprises retrieving a language code from at least one of the webpage and the application. However, Shires teaches (lang attribute: This attribute will set the language of the recognition for the request, using a valid BCP 47 language tag.[5.1.1 SpeechRecognition Attributes]). It would have been obvious to a person having ordinary skill in the art before the effective filing date with the teachings of Ballinger and Mirkovic to incorporate the use of a language code from Shires because it would enable developers to use speech recognition as an input for forms, continuous dictation and control. Claims 9 and 16 recite similar limitations to Claim 2 and are rejected similarly. Claim(s) 4-5, 11-12, 18-19 are rejected under 35 U.S.C. 103 as being unpatentable over Ballinger(US PGPub 20220405046) in view of Mirkovic(US PGPub 20060271364) as applied to claim 1 above, and further in view of Casado(US Pat 8862467). Regarding Claim 4, neither Ballinger nor Mirkovic teach the background inputs further describe a time when the speech input was received. However Casado teaches the background inputs further describe a time when the speech input was received (The context module 216 can also determine context for a spoken input from additional sources including… an indication of the time at which the spoken input was received by the client device 202 [Col 11 Line 16-21]). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention with the teaching of Ballinger and Mirkovich to incorporate the time context of Casado because it would improve the accuracy of speech recognition(Col 1 Line 44-45). Claims 11 and 18 recite similar limitations to Claim 4 and are rejected similarly. Regarding Claim 5, neither Ballinger nor Mirkovic teach the background inputs further describe a location from which the speech input was received. However Casado teaches the background inputs further describe a location from which the speech input was received (The context module 216 can also determine context for a spoken input from additional sources including… a location of the client device 202 [Col 11 Line 16-19]). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention with the teaching of Ballinger and Mirkovich to incorporate the time context of Casado because it would improve the accuracy of speech recognition(Col 1 Line 44-45). Claims 12 and 19 recite similar limitations to Claim 5 and are rejected similarly. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to ARJUN R SWAMY whose telephone number is (571)272-9763. The examiner can normally be reached Mon-Fri 8-5. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Hai Phan can be reached at (571) 272-6338. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ARJUN SWAMY/Examiner, Art Unit 2654 /HAI PHAN/Supervisory Patent Examiner, Art Unit 2654
Read full office action

Prosecution Timeline

Aug 07, 2023
Application Filed
Nov 24, 2023
Response after Non-Final Action
Sep 10, 2026
Non-Final Rejection mailed — §101, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month