Prosecution Insights
Last updated: September 23, 2026
Application No. 18/999,337

METHODS AND SYSTEMS FOR CORRECTING, BASED ON SPEECH, INPUT GENERATED USING AUTOMATIC SPEECH RECOGNITION

Non-Final OA §101§103§DOUBLEPATENT
Filed
Dec 23, 2024
Priority
May 24, 2017 — nonprovisional of PCTUS2017034229 +2 more
Examiner
SHIN, SEONG-AH A
Art Unit
Tech Center
Assignee
Adeia Technologies Inc.
OA Round
1 (Non-Final)
78%
Grant Probability
Favorable
1-2
OA Rounds
10m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 78% — above average
78%
Career Allowance Rate
331 granted / 422 resolved
+18.4% vs TC avg
Strong +21% interview lift
Without
With
+21.4%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
22 currently pending
Career history
445
Total Applications
across all art units

Statute-Specific Performance

§101
23.6%
-16.4% vs TC avg
§103
46.4%
+6.4% vs TC avg
§102
14.0%
-26.0% vs TC avg
§112
6.7%
-33.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 422 resolved cases

Office Action

§101 §103 §DOUBLEPATENT
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Status of Claims Claims 103-122 are pending in this application. Claims 1-102 are canceled. Double Patenting The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the claims at issue are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); and In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on a nonstatutory double patenting ground provided the reference application or patent either is shown to be commonly owned with this application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b). The USPTO internet Web site contains terminal disclaimer forms which may be used. Please visit http://www.uspto.gov/forms/. The filing date of the application will determine what form should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to http://www.uspto.gov/patents/process/file/efs/guidance/eTD-info-I.jsp. Claims 103 - 122 are rejected on the ground of nonstatutory double patenting over claims 1-19 of U.S. Patent No. 11,521,608. Although the claims at issue are not identical, they are not patentably distinct from each other because deleting inherent and/or unnecessary limitations/step and rearranging the claims would be within the level of one of ordinary skill in the art. It is well settled that the omission of an element, e.g. “receiving a non-speech input at a second time” and its function is an obvious expedient if the remaining elements perform the same function as before. In re Karlson, 136 USPQ 184 (CCPA 1963). Also note Ex parte Rainu, 168 USPQ 375 (Bd. App. 1969). Insertion of a reference element or step whose function is not needed would be obvious to one of ordinary skill in the art. Instant Application No. 18/999,337 U.S. Patent No. 11,521,608 103. A computer-implemented method, comprising: receiving first speech; generating, using automatic speech recognition, a first text input based at least in part on the first speech, wherein a content search query is performed using the first text input; providing for display, at a first time, content search results from the content search query; receiving, at a second time, a non-speech input indicating that the first speech was incorrectly interpreted; based at least in part on receiving the non-speech input, determining that a period of time between the first time when the content search results were provided for display and the second time when the non-speech input was received is less than a threshold period of time; and based at least in part on determining that the period of time is less than the threshold period of time, generating a corrected input of the first text input. 104. The method of claim 103, wherein receiving the non-speech input indicating that the first speech was incorrectly interpreted comprises: capturing an image of a face of a user using a camera device; analyzing the image of the face of the user using facial recognition; and based at least in part on the analyzing, determining that the face of the user is indicative of a dissatisfied emotion. 105. The method of claim 104, wherein receiving, at the second time, the non-speech input indicating that the first speech was incorrectly interpreted comprises: determining, based at least in part on the image, a first relative size of the face of the user; determining, at the second time, a second relative size of the face of the user based at least in part on a second image; comparing a relative size difference between the first relative size and the second relative size to a threshold relative size; and determining that the relative size difference is greater than the threshold relative size. 106. The method of claim 103, wherein generating the corrected input of the first text input is further based at least in part on: measuring a baseline environmental noise level; measuring an environmental noise level while the first speech is being received; and determining that an environmental noise level difference between the environmental noise level while the first speech is being received and the baseline environmental noise level is greater than a threshold environmental noise level. 107. The method of claim 103, wherein receiving, at the second time, the non-speech input indicating that the first speech was incorrectly interpreted comprises: measuring, between the first time and the second time, a second acceleration of a second motion associated with a user input device, wherein the user input device is configured to receive speech inputs; determining a difference in acceleration between a first acceleration of a first motion associated with the user input device and the second acceleration of the user input device, wherein the first acceleration is measured at the first time; and determining that the difference in acceleration is greater than a threshold acceleration. 108. The method of claim 103, wherein generating the corrected input is further based at least in part on: determining that no input associated with the content search results was received via a user interface between the first time when the content search results were provided for display and the second time when the non-speech input was received. 109. The method of claim 108, wherein determining that no input associated with the content search results was received via the user interface between the first time and the second time further comprises determining that no input to scroll through the content search results or access the content search results was received via the user interface between the first time and the second time. 110. The method of claim 103, wherein generating the corrected input of the first text input comprises modifying the first text input based at least in part on an interpretation of the non-speech input. 111. The method of claim 103, further comprising adjusting the threshold period of time based at least in part on respective average times between a plurality of first text inputs associated with previous speech inputs and a plurality of second inputs respectively associated with each of the plurality of first text inputs. 112. The method of claim 103, further comprising determining the first time by detecting a time when signals transmitted to pixels of a display screen first change subsequent to the first time. Capturing a first image of a face of a user while the first speech is being received; determining a first relative size of the face of the user in the first image; capturing, a second image of the face of the user while the second speech is being received; determining a second relative size of the face of the user in the second image; comparing the first relative size with the second relative size to derive a relative size difference; and wherein generating the corrected input is further based on determining that the relative size difference is greater than a threshold value. 1. A method for correcting, based on speech, input generated using automatic speech recognition, in the absence of an explicit indication in the speech that a user intended to correct the input with the speech, the method comprising: receiving, via a user input device, first speech; determining, using control circuitry and automatic speech recognition, a first input based on the first speech; retrieving, from a database, browsing search results based on the first input; generating for display, using the control circuitry, the browsing search results; determining, using the control circuitry, a first time when the browsing search results were generated for display; receiving, via the user input device, subsequent to receiving the first speech, second speech; determining, using the control circuitry and automatic speech recognition, a second input based on the second speech; determining, using the control circuitry, a second time when the second speech was received; calculating, using the control circuitry, a time difference between the second time and the first time; determining, using the control circuitry, whether the time difference between the second time and the first time is less than a threshold time; determining, using the control circuitry, that no input associated with the browsing search results was received via the user input device between the first time and the second time; and in response to determining that the time difference is less than the threshold time and determining that no input associated with the browsing search results was received via the user input device between the first time and the second time, generating, using the control circuitry, a corrected input based on the first input by replacing a portion of the first input with a portion of the second input. 2. The method of claim 1, wherein determining that no input associated with the browsing search results was received via the user input device between the first time and the second time comprises determining that no input to scroll through the browsing search results, read descriptions of the browsing search results, open the browsing search results, or play the browsing search results was received via the user input device between the first time and the second time. 3. The method of claim 1, further comprising: capturing, via the user input device, between the first time and the second time, an image of a face of a user; and wherein generating the corrected input is further based on determining, using the control circuitry, that the face of the user in the image is associated with a dissatisfied emotion. 4. The method of claim 1, further comprising: capturing, via the user input device, while the first speech is being received, a first image of a face of a user; determining, using the control circuitry, a first relative size of the face of the user in the first image; capturing, via the user input device, while the second speech is being received, a second image of the face of the user; determining, using the control circuitry, a second relative size of the face of the user in the second image; comparing, using the control circuitry, a relative size difference between the first relative size of the face of the user and the second relative size of the face of the user to a threshold relative size; based on comparing the relative size difference between the first relative size of the face of the user and the second relative size of the face of the user to the threshold relative size, determining, using the control circuitry, that the relative size difference is greater than the threshold relative size; and wherein generating the corrected input is further based on determining, using the control circuitry, that the relative size difference is greater than the threshold relative size. 5. The method of claim 1, further comprising: comparing, using the control circuitry, the time difference between the second time and the first time to another threshold time; based on comparing the time difference between the second time and the first time to the other threshold time, determining, using the control circuitry, that the time difference between the second time and the first time is greater than the other threshold time; and wherein generating the corrected input is further based on determining, using the control circuitry, that the time difference between the second time and the first time is greater than the other threshold time. 6. The method of claim 1, further comprising adjusting the threshold time based on an average time between inputs associated with a user. 7. The method of claim 1, further comprising: measuring, via the user input device, a baseline environmental noise level; measuring, via the user input device, an environmental noise level while the first speech is being received; comparing, using the control circuitry, an environmental noise level difference between the environmental noise level while the first speech is being received and the baseline environmental noise level to a threshold environmental noise level; based on comparing the environmental noise level difference between the environmental noise level while the first speech is being received and the baseline environmental noise level to the threshold environmental noise level, determining, using the control circuitry, that the environmental noise level difference is greater than the threshold environmental noise level; and wherein generating the corrected input is further based on determining, using the control circuitry, that the environmental noise level difference is greater than the threshold environmental noise level. 8. The method of claim 1, wherein determining the second time when the second speech was received comprises measuring, via the user input device, a time when an earliest pronunciation subsequent to the first time was received. 9. The method of claim 1, wherein determining the first time when the browsing search results were generated for display comprises detecting, using the control circuitry, a time when signals transmitted to pixels of a display screen first changed subsequent to the first time. Claims 103 - 122 are rejected on the ground of nonstatutory double patenting over claims 1-20 of U.S. Patent No. 12,211,501. Although the claims at issue are not identical, they are not patentably distinct from each other because deleting inherent and/or unnecessary limitations/step and rearranging the claims would be within the level of one of ordinary skill in the art. It is well settled that the omission of an element, e.g. “receiving a non-speech input at a second time” and its function is an obvious expedient if the remaining elements perform the same function as before. In re Karlson, 136 USPQ 184 (CCPA 1963). Also note Ex parte Rainu, 168 USPQ 375 (Bd. App. 1969). Insertion of a reference element or step whose function is not needed would be obvious to one of ordinary skill in the art. Instant Application No. 18/999,337 U.S. Patent No. 12,211,501 103. A computer-implemented method, comprising: receiving first speech; generating, using automatic speech recognition, a first text input based at least in part on the first speech, wherein a content search query is performed using the first text input; providing for display, at a first time, content search results from the content search query; receiving, at a second time, a non-speech input indicating that the first speech was incorrectly interpreted; based at least in part on receiving the non-speech input, determining that a period of time between the first time when the content search results were provided for display and the second time when the non-speech input was received is less than a threshold period of time; and based at least in part on determining that the period of time is less than the threshold period of time, generating a corrected input of the first text input. 104. The method of claim 103, wherein receiving the non-speech input indicating that the first speech was incorrectly interpreted comprises: capturing an image of a face of a user using a camera device; analyzing the image of the face of the user using facial recognition; and based at least in part on the analyzing, determining that the face of the user is indicative of a dissatisfied emotion. 105. The method of claim 104, wherein receiving, at the second time, the non-speech input indicating that the first speech was incorrectly interpreted comprises: determining, based at least in part on the image, a first relative size of the face of the user; determining, at the second time, a second relative size of the face of the user based at least in part on a second image; comparing a relative size difference between the first relative size and the second relative size to a threshold relative size; and determining that the relative size difference is greater than the threshold relative size. 106. The method of claim 103, wherein generating the corrected input of the first text input is further based at least in part on: measuring a baseline environmental noise level; measuring an environmental noise level while the first speech is being received; and determining that an environmental noise level difference between the environmental noise level while the first speech is being received and the baseline environmental noise level is greater than a threshold environmental noise level. 107. The method of claim 103, wherein receiving, at the second time, the non-speech input indicating that the first speech was incorrectly interpreted comprises: measuring, between the first time and the second time, a second acceleration of a second motion associated with a user input device, wherein the user input device is configured to receive speech inputs; determining a difference in acceleration between a first acceleration of a first motion associated with the user input device and the second acceleration of the user input device, wherein the first acceleration is measured at the first time; and determining that the difference in acceleration is greater than a threshold acceleration. 108. The method of claim 103, wherein generating the corrected input is further based at least in part on: determining that no input associated with the content search results was received via a user interface between the first time when the content search results were provided for display and the second time when the non-speech input was received. 109. The method of claim 108, wherein determining that no input associated with the content search results was received via the user interface between the first time and the second time further comprises determining that no input to scroll through the content search results or access the content search results was received via the user interface between the first time and the second time. 110. The method of claim 103, wherein generating the corrected input of the first text input comprises modifying the first text input based at least in part on an interpretation of the non-speech input. 111. The method of claim 103, further comprising adjusting the threshold period of time based at least in part on respective average times between a plurality of first text inputs associated with previous speech inputs and a plurality of second inputs respectively associated with each of the plurality of first text inputs. 112. The method of claim 103, further comprising determining the first time by detecting a time when signals transmitted to pixels of a display screen first change subsequent to the first time. Capturing a first image of a face of a user while the first speech is being received; determining a first relative size of the face of the user in the first image; capturing, a second image of the face of the user while the second speech is being received; determining a second relative size of the face of the user in the second image; comparing the first relative size with the second relative size to derive a relative size difference; and wherein generating the corrected input is further based on determining that the relative size difference is greater than a threshold value. 1. A method for correcting an input generated using automatic speech recognition for searching content, the method comprising: receiving first speech; generating, using automatic speech recognition, a first text input based on the first speech, wherein a content search query is performed using the first text input; providing for display content search results from the content search query using the first text input; determining a first time when the content search results were provided for display; receiving second speech at a second time; determining whether a period between the first time when the content search results were provided for display and the receiving the second speech at the second time is less than a threshold period; determining whether no input associated with the content search results was received via a user interface between the first time when the content search results were provided for display and the receiving the second speech at the second time; and in response to determining that the period is less than the threshold period and no input associated with the content search results was received via the user interface between the first time when the content search results were provided for display and the receiving the second speech at the second time, generating a corrected input of the first text input based on the second speech. 2. The method of claim 1, wherein providing for display the content search results from the content search query using the first text input further comprises retrieving, from a database, the content search results based on the first text input. 3. The method of claim 1, further comprising: capturing, via a user input device, an image of a face of a user between the first time and the second time; and wherein generating the corrected input is further based on determining that the face of the user in the image is associated with a dissatisfied emotion. 4. The method of claim 1, further comprising: capturing a first image of a face of a user while the first speech is being received; determining a first relative size of the face of the user in the first image; capturing, a second image of the face of the user while the second speech is being received; determining a second relative size of the face of the user in the second image; comparing the first relative size with the second relative size to derive a relative size difference; and wherein generating the corrected input is further based on determining that the relative size difference is greater than a threshold value. 5. The method of claim 1, further comprising: comparing the period between the second time and the first time to another threshold period; and wherein generating the corrected input is further based on determining that the period between the second time and the first time is greater than the other threshold period. 6. The method of claim 1, further comprising adjusting the threshold period based on an average period between inputs associated with a user. 7. The method of claim 1, further comprising: measuring a baseline environmental noise level; measuring an environmental noise level while the first speech is being received; comparing an environmental noise level difference between the environmental noise level while the first speech is being received and the baseline environmental noise level with a threshold environmental noise level; and wherein generating the corrected input is further based on determining that the environmental noise level difference is greater than the threshold environmental noise level. 8. The method of claim 1, further comprising: determining the second time comprises measuring a time when an earliest pronunciation subsequent to the first time is received. 9. The method of claim 1, further comprising: determining the first time comprises detecting a time when signals transmitted to pixels of a display screen first change subsequent to the first time. 10. The method of claim 1, wherein determining whether no input associated with the content search results was received via the user interface between the first time when the content search results were provided for display and the receiving the second speech at the second time comprises determining that no input to scroll through the browsing search results, read descriptions of the browsing search results, open the browsing search results, or play the browsing search results was received via the user interface between the first time and the second time. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 103-122 and are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 2A, Prong One: The independent claim 103 recites “receiving first speech; generating, using automatic speech recognition, a first text input based at least in part on the first speech, wherein a content search query is performed using the first text input; providing for display, at a first time, content search results from the content search query; receiving, at a second time, a non-speech input indicating that the first speech was incorrectly interpreted; based at least in part on receiving the non-speech input, determining that a period of time between the first time when the content search results were provided for display and the second time when the non-speech input was received is less than a threshold period of time; and based at least in part on determining that the period of time is less than the threshold period of time, generating a corrected input of the first text input.”. Claims 103 and 113 recite correcting intent from non-speech cues after receiving the first speech input when the cues indicate a likely correction. [Abstract idea indicators] Transcribing speech into text is the conversion of verbal content to written form— a task that a humans routinely performs mentally or with conventional tools. Determining to correct the first query based on user’s non-speech feedback, i.e., a decision-making process. Correcting the query is a cognitive step that are mental processes. Accordingly, the claims are directed to the judicial exception of a mental process. Step 2A, Prong Two: This judicial exception is not integrated into a practical application. The computer is recited at a high-level of generality (i.e., as performing a generic computer function and being used as an applying) such that it amounts no more than mere instructions to apply the exception using a generic computer. Accordingly, there additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea. Step 2B — Claims Do Not Recite an Inventive Concept That Transforms the Mental Process into Patent-Eligible Subject Matter The claims add generic, well-understood computer components (memory and control circuitry) and do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract idea into a practical application, the additional element of using a computer amounts to no more than mere instructions to apply an exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. The claim is not patent eligible. Applying Alice step two and relevant Federal Circuit precedent: The recitation of conventional computer components (memory and processor) performing routine functions does not supply an inventive concept. The claims recite high-level, result-oriented steps (e.g., “receiving,” “generating,” “providing”, “determining”) that describe mental processes rather than specific technical means for performing those processes. Because the claims lack limitations that tie the mental-process steps to a particular way of achieving a technological improvement (for example, a novel model architecture, specialized data representation, unique training regimen that yields demonstrable technical performance gains, a specialized streaming/decoding pipeline that reduces latency by a quantifiable amount, or hardware/software co-design), the additional elements do not transform the mental processes into significantly more. Therefore, claims 103 and 113 fail to recite an inventive concept sufficient to transform the judicial exception into patent-eligible subject matter. With respect to dependent claims 104 and 114, the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. With respect to dependent claims 105 and 115, the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. With respect to dependent claims 106 and 116, the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. With respect to dependent claims 107 and 117, the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. With respect to dependent claims 108 and 118, the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. With respect to dependent claims 109 and 119, the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. With respect to dependent claims 110 and 120, the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. With respect to dependent claims 111 and 121, the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. With respect to dependent claims 112 and 122, the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. Conclusion — Rejection Claims 103 -122 are rejected under 35 U.S.C. § 101 as being directed to a judicial exception (mental processes) and failing to recite additional elements that amount to significantly more than the judicial exception. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action. The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 103, 104, 110, 113, 114, and 120 are rejected under pre-AIA 35 U.S.C. 103(a) as being unpatentable over Sanders et al., (US 2016/0103833 A1) in view of Taubman et al., (US 2020/0082829 A1). Regarding claim 103, Sanders discloses a computer-implemented method, comprising: receiving first speech (Figs. 1 and 2A, [0018] receiving an audio query); generating, using automatic speech recognition, a first text input based at least in part on the first speech, wherein a content search query is performed using the first text input (Figs. 1 and 2A, [0018][0019] transforming the audio query into one or more terms to provide search result); providing for display, at a first time, content search results from the content search query (Fig. 2A, [0028] displaying search result 204 for the search query 202); receiving, at a second time, a non-speech input indicating that the first speech was incorrectly interpreted (the limitation has been interpreted based on Applicant Specification, Sanders, [0088][0089] “If first speech 106 was incorrectly recognized, and search results 112 presented do not match what the user intended, the user may be dissatisfied, and therefore the face of the user may exhibit a dissatisfied expression”; Fig. 2A, [0029][0030] gathering biometric feedback form a user for the query 202 after displaying the result and identifying that the search result 204 indicates negative engagement); based at least in part on determining that the period of time is less than the threshold period of time, generating a corrected input of the first text input (Fig. 2A, [0032] providing an additional input with biometric parameters which are captured from the user and searching including the additional input). Sanders does not explicitly teach, however Taubman does explicitly teach: based at least in part on receiving the [non-speech] input, determining that a period of time between the first time when the content search results were provided for display and the second time when the non-speech input was received is less than a threshold period of time (Fig. 2A-2C, 10, [0096][0116][0148] “determining that the voice input is classified as feedback may include determining that a time difference between a time associated with providing the answer and a time associated with receiving the voice input is within a predetermined time… a user's facial expressions may be used in addition to, or in lieu of, words spoken by the user in order to determine user feedback”). Therefore, it would have been obvious to one of ordinary skill before the effective filing date of the claimed invention to incorporate the method of Speech recognition using repeated and/or corrected utterances as taught by Sanders with the method of determining that the voice input is classified as feedback within a predetermined time as taught by Taubman to dynamically improve answers provided to questions—thus improving users' experience (Taubman, [0007]). Regarding claim 104, Sanders in view of Taubman discloses the method of claim 103, and Sanders further discloses: wherein receiving the non-speech input indicating that the first speech was incorrectly interpreted comprises: capturing an image of a face of a user using a camera device ([0029] capturing the facial feature of a user using a camera); analyzing the image of the face of the user using facial recognition ([0029][0030] analyzing biometric parameter, e.g., facial feature); and based at least in part on the analyzing, determining that the face of the user is indicative of a dissatisfied emotion ([0030] which can indicate likely negative engagement). Regarding claim 110, Sanders in view of Taubman discloses the method of claim 103, and Sanders further discloses: wherein generating the corrected input of the first text input comprises modifying the first text input based at least in part on an interpretation of the non-speech input (Fig. 2A, [0032] providing an additional input with biometric parameters which are captured from the user and searching including the additional input). Regarding claims 113, 114, and 120, Claims 113, 114, and 120 are the corresponding system claims to method claims 103, 104 and 110. Therefore, claims 113, 114, and 120 are rejected using the same rationale as applied to claims 104, 104 and 110 above. Claims 108, 109, 118 and 119 are rejected under pre-AIA 35 U.S.C. 103(a) as being unpatentable over Sanders et al., (US 2016/0103833 A1) in view of Taubman et al., (US 2020/0082829 A1) and further in view of Shaw et al. (US Patent 9,123,339). Regarding claim 108, Sanders in view of Taubman discloses the method of claim 103. Sanders in view of Taubman does not explicitly teach however Shaw does explicitly teach: wherein generating the corrected input is further based at least in part on: determining that no input associated with the content search results was received via a user interface between the first time when the content search results were provided for display and the second time when the non-speech input was received (Shaw, Fig. 3, steps 306 and 308, Col. 13, lines 27-50, between input of first and second, any input associated with browsing search results was not received using user input device). Therefore, it would have been obvious to one of ordinary skill before the effective filing date of the claimed invention to incorporate the method of Speech recognition using repeated and/or corrected utterances as taught by Sanders in view of Taubman with the method of identifying additional input from a user as taught by Shaw to provide improved text conversion. Regarding claim 109, Sanders in view of Taubman and further in view of Shaw discloses the method of claim 108 and Shaw further discloses; wherein determining that no input associated with the content search results was received via the user interface between the first time and the second time further comprises determining that no input to scroll through the content search results or access the content search results was received via the user interface between the first time and the second time (Shaw, Fig. 3, steps 306 and 308, Col. 13, lines 27-50, between input of first and second, any input associated with browsing search results was not received using user input device). The previous motivation statement as in claim 108 is still applied. Regarding claims 118 and 119, Claims 118 and 119 are the corresponding system claims to method claims 108 and 109. Therefore, claims 118 and 109 are rejected using the same rationale as applied to claims 108 and 109 above. Claims 111 and 121 are rejected under pre-AIA 35 U.S.C. 103(a) as being unpatentable over Sanders et al., (US 2016/0103833 A1) in view of Taubman et al., (US 2020/0082829 A1) and further in view of Aleksic et al., (US 2017/0069309 A1). Regarding claim 111, Sanders in view of Taubman discloses the method of claim 103. Sanders in view of Taubman does not explicitly teach however Aleksic does explicitly teach: adjusting the threshold period of time based at least in part on respective average times between a plurality of first text inputs associated with previous speech inputs and a plurality of second inputs respectively associated with each of the plurality of first text inputs ([0025][0037] dynamically adjusting an end of speech timeout). Therefore, it would have been obvious to one of ordinary skill before the effective filing date of the claimed invention to incorporate the method of Speech recognition using repeated and/or corrected utterances as taught by Sanders in view of Taubman with the method of adjusting EOS timeout period as taught by Aleksic to improve speech end pointing, achieving decreased speech recognition latency and improved speech recognition accuracy (Aleksic, [0004]). Regarding claim 121, Claim 121 is the corresponding system claim to the method claim 111. Therefore, claim 121 is rejected using the same rationale as applied to claim 111 above. Claims 112 and 122 are rejected under pre-AIA 35 U.S.C. 103(a) as being unpatentable over Sanders et al., (US 2016/0103833 A1) in view of Taubman et al., (US 2020/0082829 A1) and further in view of Kawase et al. (JP2016180917A). Regarding claim 112, Sanders in view of Taubman discloses the method of claim 103. Sanders in view of Taubman does not explicitly teach however Kawase does explicitly teach: determining the first time comprises detecting a time when signals transmitted to pixels of a display screen first change subsequent to the first time (Kawase, page 4, 2nd -7th paragraph, determining the (m-1)th time for presentation time). Therefore, it would have been obvious to one of ordinary skill before the effective filing date of the claimed invention to incorporate the method of Speech recognition using repeated and/or corrected utterances as taught by Sanders in view of Taubman with the method of adapting reaction time within predetermined time period between the time of result presentation and the time of receiving second speech input as taught by Kawase to provide visual presentation that can determine whether or not there is a corrected utterance without using a change in acoustic feature amount for each utterance as a basis (Kawase, page 2). Regarding claim 122, Claim 122 is the corresponding system claim to the method claim 112. Therefore, claim 122 is rejected using the same rationale as applied to claim 112 above. Allowable Subject Matter Claims 105-107 and 115-117 are objected to as being dependent upon a rejected base claim but would be allowable if rewritten in independent form including all the limitations of the base claim and any intervening claims and if rewritten or amended to overcome the rejection(s) under 35 U.S.C. 101 and a nonstatutory double patenting. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Please see attached form PTO-892. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SEONG-AH A. SHIN whose telephone number is (571)272-5933. The examiner can normally be reached 9 AM-3PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre-Louis Desir can be reached at 571-272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. Seong-ah A. Shin Primary Examiner Art Unit 2659 /SEONG-AH A SHIN/Primary Examiner, Art Unit 2659
Read full office action

Prosecution Timeline

Dec 23, 2024
Application Filed
Aug 20, 2026
Non-Final Rejection mailed — §101, §103, §DOUBLEPATENT (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12725610
VOICE RECOGNITION SYSTEM, SERVER, DISPLAY APPARATUS AND CONTROL METHODS THEREOF
3y 11m to grant Granted Sep 01, 2026
Patent 12725609
Hotwording by Degree
2y 2m to grant Granted Sep 01, 2026
Patent 12694220
METHOD AND SYSTEM FOR PERSONALIZED EMBEDDING SEARCH ENGINE
3y 5m to grant Granted Jul 28, 2026
Patent 12682898
KEY PHRASE SPOTTING
2y 3m to grant Granted Jul 14, 2026
Patent 12670904
SELECTING AN AUTOMATED ASSISTANT AS THE PRIMARY AUTOMATED ASSISTANT FOR A DEVICE BASED ON DETERMINED AFFINITY SCORES FOR CANDIDATE AUTOMATED ASSISTANTS
3y 6m to grant Granted Jun 30, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
78%
Grant Probability
99%
With Interview (+21.4%)
2y 7m (~10m remaining)
Median Time to Grant
Low
PTA Risk
Based on 422 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month