Prosecution Insights
Last updated: August 06, 2026
Application No. 18/094,280

TASK-ORIENTED DIALOG MODELING AND ACTION DETERMINATION

Non-Final OA §101§103
Filed
Jan 06, 2023
Examiner
LEE, MICHAEL CHRISTOPHER
Art Unit
2128
Tech Center
2100 — Computer Architecture & Software
Assignee
Toyota Connected North America Inc.
OA Round
3 (Non-Final)
62%
Grant Probability
Moderate
3-4
OA Rounds
0m
Est. Remaining
88%
With Interview

Examiner Intelligence

Grants 62% of resolved cases
62%
Career Allowance Rate
95 granted / 153 resolved
+7.1% vs TC avg
Strong +26% interview lift
Without
With
+26.4%
Interview Lift
resolved cases with interview
Typical timeline
3y 3m
Avg Prosecution
47 currently pending
Career history
197
Total Applications
across all art units

Statute-Specific Performance

§101
30.0%
-10.0% vs TC avg
§103
45.2%
+5.2% vs TC avg
§102
10.7%
-29.3% vs TC avg
§112
12.7%
-27.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 153 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 6/22/2026 has been entered. Response to Amendment Applicant’s Amendment and remarks submitted on 6/22/2026 have been considered. Claims 1-20 are pending. Claim Objections. The objections to claims 2-8, 10, and 18 are withdrawn in view of Applicant’s amendments to such claims. Response to Arguments On page 8 of Applicant’s 6/22/2026 Amendment and remarks, Applicant asserts that at least paras. 0056-0058, 0072, and 0078-79 of the instant specification provide written description support for the claim amendments. The examiner agrees that the portions of the disclosure identified by Applicant provide sufficient written description support for the claim amendments. On page 9 of Applicant’s 6/22/2026 Amendment and remarks, with respect to the rejections under 35 U.S.C. 101, Applicant argues that the claim amendments overcome the rejections. The examiner respectfully disagrees. As explained in the detailed rejections below, padding a data structure is a mental process. Moreover, having one ML model call another ML model merely pertains to one generic computer component calling another, and is not sufficient to integrate the abstract idea into a practical application. On pages 9-10 of Applicant’s 6/22/2026 Amendment and remarks, with respect to the rejections under 35 U.S.C. 103, Applicant argues that the amendments to the independent claims overcome the previous rejections. In view of the amendments to the independent claims, new grounds of rejection under 35 U.S.C. 103 are provided below. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Regarding Step 1 of the Alice/Mayo framework, Claims 1-8 are directed to an apparatus (a machine), Claims 9-16 are directed to a method (a process), and Claims 17-20 are directed to a “non-transitory computer-readable medium” (an article of manufacture) which each fall within one of the four statutory categories of inventions. Regarding Claim 1 Step 2A, prong 1 (Is the claim directed to a law of nature, a natural phenomenon or an abstract idea). Claim 1 recites the following mental processes, that in each case under the broadest reasonable interpretation, covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components (e.g., “audio sensor”, “processor”, “machine learning model”, and “virtual assistant”). determine, ... that the utterance lacks sufficient information to determine a task to be performed by the vehicle (under the broadest reasonable interpretation, a human can perform this limitation mentally, e.g., a human such as a human assistant can listen/read an utterance mentally and then mentally determine that there is not enough information available to determine a task to be performed by the vehicle, such as the human hearing an utterance of “umm umm umm” and knowing there is not enough information to determine a task) determine, ... that sufficient information is available to determine the action based on the additional utterance (under the broadest reasonable interpretation, a human can perform this limitation mentally, e.g., a human such as a human assistant can listen/read additional utterances mentally and then mentally determine that there is now enough information available to determine an action, e.g., the human assistant mentally determines that the human assistant has enough information in order to understand and carry-out the user’s request) store each word of the utterance in a corresponding slot of a data structure to create a stored utterance (under the broadest reasonable interpretation, a human can perform this limitation mentally or using pencil and paper, such as by writing down each word of the utterance in an array on a piece of paper) normalize a size of the stored utterance by adding padding to the data structure (under the broadest reasonable interpretation, a human can perform this limitation mentally or using pencil and paper, such as by writing down each word of the utterance in an array on a piece of paper, and normalizing the array by adding padding so that each array is the same size) in response to the utterance lacking the sufficient information, ... (under the broadest reasonable interpretation, a human can perform this limitation mentally, e.g., a human such as a human assistant can listen/read additional utterances mentally and then mentally determine that there is now enough information available to determine an action) generate a converted data structure with the padding removed (under the broadest reasonable interpretation, a human can perform this limitation mentally or using pencil and paper, such as by writing down each word of the utterance in an array on a piece of paper, and normalizing the array by adding padding so that each array is the same size, and then deleting such padding at a later point in time by erasing them) determine, ... the task based on the utterance and the additional utterances (under the broadest reasonable interpretation, a human can perform this limitation mentally, e.g., a human can mentally determine the task to be performed based on the utterances, e.g., can predict that the user wants the virtual assistant of the vehicle to answer a question) Step 2A, prong 2 (Does the claim recite additional elements that integrate the judicial exception into a practical application?). The judicial exception is not integrated into a practical application. In particular, the claim recites the additional elements (e.g., “receiver”, “processor”, “machine learning model”, and “virtual assistant”) which are recited at a high-level of generality such that they amount to no more than mere instructions to apply the exception using a generic computer component (See MPEP 2106.05(f)). Regarding the “A vehicle” limitation, such limitation amounts to no more than generally linking the use of a judicial exception to a particular technological environment or field of use (e.g., the mental processes are now limited to being performed within a vehicle). As explained by the Supreme Court, a claim directed to a judicial exception cannot be made eligible "simply by having the applicant acquiesce to limiting the reach of the patent for the formula to a particular technological use." Diamond v. Diehr, 450 U.S. 175, 192 n.14, 209 USPQ 1, 10 n. 14 (1981). Thus, limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception do not integrate a judicial exception into a practical application. Regarding the “a processor is configured to” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception. In particular, the claim only recites the additional element of a processor. This additional element is recited at a high-level of generality and amounts to no more than mere instructions to apply the exception using a generic computer component (a processor). Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)). Regarding the “call a first machine learning (ML) model in response to an audio sensor receiving an utterance from a user” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception. In particular, the claim only recites the additional element of calling a machine learning model. This additional element is recited at a high-level of generality and amounts to no more than mere instructions to apply the exception using a generic computer component (calling a machine learning model). Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)). Regarding the “via the first ML model” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception. In particular, the claim only recites the additional element of a machine learning model. This additional element is recited at a high-level of generality and amounts to no more than mere instructions to apply the exception using a generic computer component (a machine learning model). Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)). Regarding the “call, by the first ML model, a second ML model” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception. In particular, the claim only recites the additional element of calling a machine learning model. This additional element is recited at a high-level of generality and amounts to no more than mere instructions to apply the exception using a generic computer component (calling a machine learning model). Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)). Regarding the “cause, by the second ML model, a prompt to be provided to the user based on the converted data structure” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, because the limitation attempts to cover a solution to an identified problem with no restriction on how the result is accomplished, or provides no description of the mechanism for accomplishing the result. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)). Regarding the “receive, via the audio sensor, an additional utterance from the user in response to the prompt” limitation, such additional element of a data gathering step is recited at a high level of generality and amounts to extra-solution activity of receiving data, i.e. pre-solution activity of gathering data for use in the claimed process (see MPEP 2106.05(g)). Regarding the “cause the vehicle to perform the task” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, because the limitation attempts to cover a solution to an identified problem with no restriction on how the result is accomplished, or provides no description of the mechanism for accomplishing the result. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)). Accordingly, at Step 2A, prong two, after considering all claim elements individually and as an ordered combination, it is determined that the claims do not integrate the judicial exception into a practical application. Step 2B (Does the claim recite additional elements that amount to significantly more than the judicial exception?) In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception. As discussed above, the additional elements (e.g., “receiver”, “processor”, “machine learning model”, and “virtual assistant”) are recited at a high-level of generality such that they amount to no more than mere instructions to apply the exception using a generic computer component (See MPEP 2106.05(f)). Regarding the “A vehicle” limitation, such limitation amounts to no more than generally linking the use of a judicial exception to a particular technological environment or field of use as explained above, which does not amount to significantly more than the judicial exception. MPEP 2106.05(h). Regarding the “a processor is configured to” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, because the limitation merely provides instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not add significantly more than the judicial exception. (See MPEP 2106.05(f)). Regarding the “call a first machine learning (ML) model in response to an audio sensor receiving an utterance from a user” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, because the limitation merely provides instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not add significantly more than the judicial exception. (See MPEP 2106.05(f)). Regarding the “via the first ML model” limitations, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, because the limitation merely provides instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not add significantly more than the judicial exception. (See MPEP 2106.05(f)). Regarding the “call, by the first ML model, a second ML model” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, because the limitation merely provides instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not add significantly more than the judicial exception. (See MPEP 2106.05(f)). Regarding the “cause, by the second ML model, a prompt to be provided to the user based on the converted data structure” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, because the limitation attempts to cover a solution to an identified problem with no restriction on how the result is accomplished, or provides no description of the mechanism for accomplishing the result. Accordingly, this additional element does not add significantly more than the judicial exception. (See MPEP 2106.05(f)). Regarding the “receive, via the audio sensor, an additional utterance from the user in response to the prompt” limitation, as discussed above, the additional element of a data gathering step is recited at a high level of generality and amounts to extra-solution activity of receiving data, i.e. pre-solution activity of gathering data for use in the claimed process. The courts have found limitations directed to obtaining information electronically, recited at a high level of generality, to be well-understood, routine, and conventional (see MPEP 2106.05(d)(II), “receiving or transmitting data over a network”, "electronic record keeping," and "storing and retrieving information in memory"). Regarding the “cause the vehicle to perform the task” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, because the limitation attempts to cover a solution to an identified problem with no restriction on how the result is accomplished, or provides no description of the mechanism for accomplishing the result. Accordingly, this additional element does not add significantly more than the judicial exception. (See MPEP 2106.05(f)). Accordingly, at Step 2B after considering all claim elements individually and as an ordered combination, it is determined that the claims do not integrate the judicial exception into a practical application. Regarding Claim 2 Step 2A, Prong 1 prompt the user for the additional utterance based on content included in a most recently received utterance (under the broadest reasonable interpretation, a human can perform this limitation mentally, e.g., a human such as a human assistant can listen/read additional utterances mentally and then mentally determine that additional information is needed, and therefore make a mental decision to prompt the user for an additional utterance) Regarding Step 2A, Prong 2, the claim does not include any additional elements that integrate the judicial exception into a practical application and regarding Step 2B, there are no additional elements recited that amount to significantly more than the judicial exception. Regarding Claim 3 Step 2A, Prong 2 Regarding the “wherein the prompt is provided by a virtual assistant of the vehicle, and the user is a passenger within the vehicle” limitation, such limitation amounts to no more than generally linking the use of a judicial exception to a particular technological environment or field of use (automotive assistants). As explained by the Supreme Court, a claim directed to a judicial exception cannot be made eligible "simply by having the applicant acquiesce to limiting the reach of the patent for the formula to a particular technological use." Diamond v. Diehr, 450 U.S. 175, 192 n.14, 209 USPQ 1, 10 n. 14 (1981). Thus, limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception do not integrate a judicial exception into a practical application. Step 2B Regarding the “wherein the prompt is provided by a virtual assistant of the vehicle, and the user is a passenger within the vehicle” limitation, such limitation amounts to no more than generally linking the use of a judicial exception to a particular technological environment or field of use as explained above, which does not amount to significantly more than the judicial exception. MPEP 2106.05(h). Regarding Claim 4 Step 2A, Prong 1 request the user to confirm content included in a previously received. (under the broadest reasonable interpretation, this limitation merely relates to the social activity of having a conversation with another person, and requesting the user to confirm content in a previously received utterance) Regarding Step 2A, Prong 2, the claim does not include any additional elements that integrate the judicial exception into a practical application and regarding Step 2B, there are no additional elements recited that amount to significantly more than the judicial exception. Regarding Claim 5 Step 2A, Prong 1 predict an intent of the user after each utterance based on an aggregation of utterances during a conversation with the user (under the broadest reasonable interpretation, a human can perform this limitation mentally, e.g., a human such as a human assistant can review the aggregation of utterances and then mentally predict an intent of the user, e.g., the intent of the user is to receive information) when ... determines that the sufficient information is available ... determine that the sufficient information is available based on the aggregation of utterances (under the broadest reasonable interpretation, a human can perform this limitation mentally, e.g., a human such as a human assistant can mentally determine that there is enough information available in order to understand and carry-out the user’s request) Step 2A, Prong 2 Regarding the “the processor” and “the processor is configured to” limitations, such limitations are recited at a high-level of generality and amount to no more than adding the words “apply it” (or an equivalent) with the judicial exception. In particular, the claim only recites the additional elements of a processor. These additional elements are recited at a high-level of generality and amount to no more than mere instructions to apply the exception using generic computer components (a processor). Accordingly, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)). Step 2B Regarding the “the processor” and “the processor is configured to” limitations, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, because the limitation merely provides instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not add significantly more than the judicial exception. (See MPEP 2106.05(f)). Regarding Claim 6 Step 2A, Prong 1 determine, ... that a sequence of utterances is required to make a prediction based on the utterance. (under the broadest reasonable interpretation, a human can perform this limitation mentally, e.g., a human such as a human assistant can mentally determine that multiple utterances will be needed to collect sufficient information from a user) Step 2A, Prong 2 Regarding the “via the first ML model” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception. In particular, the claim only recites the additional element of a machine learning model. This additional element is recited at a high-level of generality and amounts to no more than mere instructions to apply the exception using a generic computer component (a machine learning model). Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)). Step 2B Regarding the “via the first ML model” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, because the limitation merely provides instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not add significantly more than the judicial exception. (See MPEP 2106.05(f)). Regarding Claim 7 Step 2A, Prong 1 predict, ... a next action to be taken ... based on the utterance and a previously received utterance associated with a same task (under the broadest reasonable interpretation, a human can perform this limitation mentally, e.g., a human such as a human assistant can mentally predict a next action to be taken based on the received utterances, e.g., can mentally predict that the next action will be to do research to find an answer to the user’s question) Step 2A, Prong 2 Regarding the “wherein the prompt is provided by a virtual assistant” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception. In particular, the claim only recites the additional element of a virtual assistant. This additional element is recited at a high-level of generality and amounts to no more than mere instructions to apply the exception using a generic computer component (a virtual assistant). Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)). Regarding the “via the second ML model” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception. In particular, the claim only recites the additional element of a machine learning model. This additional element is recited at a high-level of generality and amounts to no more than mere instructions to apply the exception using a generic computer component (a machine learning model). Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)). Regarding the “by the virtual assistant” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception. In particular, the claim only recites the additional element of a virtual assistant. This additional element is recited at a high-level of generality and amounts to no more than mere instructions to apply the exception using a generic computer component (a virtual assistant). Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)). Step 2B Regarding the “wherein the prompt is provided by a virtual assistant” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, because the limitation merely provides instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not add significantly more than the judicial exception. (See MPEP 2106.05(f)). Regarding the “via the second ML model” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, because the limitation merely provides instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not add significantly more than the judicial exception. (See MPEP 2106.05(f)). Regarding the “by the virtual assistant” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, because the limitation merely provides instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not add significantly more than the judicial exception. (See MPEP 2106.05(f)). Regarding Claim 8 Step 2A, Prong 1 identify, ... a word of interest within the additional utterance and an additional question ... to ask the user based on the word of interest (under the broadest reasonable interpretation, a human can perform this limitation mentally, e.g., a human such as a human assistant can identify a word of interest within an utterance and then based on such word of interest, form an additional question mentally to ask the user) Step 2A, Prong 2 Regarding the “wherein the prompt is provided by a virtual assistant” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception. In particular, the claim only recites the additional element of a virtual assistant. This additional element is recited at a high-level of generality and amounts to no more than mere instructions to apply the exception using a generic computer component (a virtual assistant). Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)). Regarding the “via the second ML model” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception. In particular, the claim only recites the additional element of a machine learning model. This additional element is recited at a high-level of generality and amounts to no more than mere instructions to apply the exception using a generic computer component (a machine learning model). Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)). Regarding the “for the virtual assistant” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception. In particular, the claim only recites the additional element of a virtual assistant. This additional element is recited at a high-level of generality and amounts to no more than mere instructions to apply the exception using a generic computer component (a virtual assistant). Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)). Step 2B Regarding the “wherein the prompt is provided by a virtual assistant” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, because the limitation merely provides instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not add significantly more than the judicial exception. (See MPEP 2106.05(f)). Regarding the “via the second ML model” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, because the limitation merely provides instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not add significantly more than the judicial exception. (See MPEP 2106.05(f)). Regarding the “for the virtual assistant” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, because the limitation merely provides instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not add significantly more than the judicial exception. (See MPEP 2106.05(f)). Regarding Claim 9 Step 2A, Prong 1 Claim 9 recites a method that corresponds to the apparatus of claim 1, and therefore the analysis under Step 2A, Prong 1 with respect to claim 1 also applies to this claim 9. While claim 9 recites additional generic computing components (e.g., “machine learning model”, and “virtual assistant”, and “audio sensor”), such additional generic computing components do not change the analysis under Step 2A, Prong 1. Step 2A, Prong 2 Claim 9 recites a method that corresponds to the apparatus of claim 1, and therefore the analysis under Step 2A, Prong 2 with respect to claim 1 also applies to this claim 9. While claim 9 recites additional generic computing components (e.g., “machine learning model”, and “virtual assistant”, and “audio sensor”), such additional generic computing components do not change the analysis under Step 2A, Prong 2. Step 2B Claim 9 recites a method that corresponds to the apparatus of claim 1, and therefore the analysis under Step 2B with respect to claim 1 also applies to this claim 9. While claim 9 recites additional generic computing components (e.g., “machine learning model”, and “virtual assistant”, and “audio sensor”), such additional generic computing components do not change the analysis under Step 2B Claims 10-16 depend from claim 9, and correspond to the apparatuses of claims 2-8, and are therefore rejected for the same reasons explained above with respect to claim 9 and claims 2-8, respectively. Regarding Claim 17 Step 2A, Prong 1 Claim 17 recites a non-transitory computer-readable medium that corresponds to the apparatus of claim 1, and therefore the analysis under Step 2A, Prong 1 with respect to claim 1 also applies to this claim 17. While claim 17 recites additional generic computing components (e.g., “non-transitory computer-readable medium”, “processor”, “machine learning model”, “virtual assistant”, and “audio sensor”), such additional generic computing components do not change the analysis under Step 2A, Prong 1. Step 2A, Prong 2 Claim 17 recites a non-transitory computer-readable medium that corresponds to the apparatus of claim 1, and therefore the analysis under Step 2A, Prong 2 with respect to claim 1 also applies to this claim 17. While claim 17 recites additional generic computing components (e.g., “non-transitory computer-readable medium”, “processor”, “machine learning model”, “virtual assistant”, and “audio sensor”), such additional generic computing components do not change the analysis under Step 2A, Prong 2. Step 2B Claim 17 recites a non-transitory computer-readable medium that corresponds to the apparatus of claim 1, and therefore the analysis under Step 2B with respect to claim 1 also applies to this claim 17. While claim 17 recites additional generic computing components (e.g., “non-transitory computer-readable medium”, “processor”, “machine learning model”, “virtual assistant”, and “audio sensor”), such additional generic computing components do not change the analysis under Step 2B. Claims 18-20 depend from claim 17, and correspond to the apparatuses of claims 2-4, and are therefore rejected for the same reasons explained above with respect to claim 17 and claims 2-4, respectively. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1-5, 7, 9-13, 15, and 17-20 are rejected under 35 U.S.C. 103 as being unpatentable over US 10997963 B1, hereinafter referenced as BALIGAR in view of US 12204866 B1, hereinafter referenced as ATLURI, and further in view of US 20190236155 A1, hereinafter referenced as BACHRACH, and further in view of US 20230259717 A1, hereinafter referenced as DANG. Regarding Claim 1 BALIGAR teaches: a processor configured to (BALIGAR, col. 5, lines 48-57: “FIG. 2 is a block diagram of an illustrative computing architecture 200 of a voice assistant service, such as the voice assistant service 102. The computing architecture 200 may be implemented in a distributed or non-distributed computing environment. The computing architecture 200 may include one or more processors 202 and one or more computer-readable media 204 that stores various modules, applications, programs, or other data. The computer-readable media 204 may include instructions that, when executed by the one or more processors 202, cause the processors to perform the operations described herein.”) determine, ... that the utterance lacks sufficient information to determine a task to be performed, (BALIGAR, col. 8, line 60 – col. 9, line 5: “At 408, the voice assistant service 102 may determine whether to request more information to determine a context of an audio request. For example, when the voice assistant service 102 includes enough information to understand and respond to an audio request that is supplemented by context information derived from the context queue, then the voice assistant service 102 may not request additional information from the user. When the voice assistant service 102 does not include enough information to understand and respond to the audio request that is supplemented by context information derived from the context queue (following the “yes” route from the decision operation 408), then the process 400 may advance to an operation 410.”; Examiner’s Note: the voice assistant determines if enough information is available to understand and respond to the user’s audio request in order for the assistant to “respond to an audio request” (where such response corresponds to the recited “performance of a task by the system”)) in response to the utterance lacking the sufficient information: (BALIGAR, col. 9, lines 6-13: “At 410, the voice assistant service 102 may request additional information from the user by sending an audio request to the user via the voice assistant device or voice assistant application. The request may be a choice, such as “did you mean “buy a toaster or diapers”. In some instances, the response may be a request for more specific information, such as “Please provide more information so I can fulfill your request”.”; BALIGAR, col. 9, lines 14-23: “When the voice assistant service 102 does include enough information to understand and respond to the audio request that is supplemented by context information derived from the context queue (following the “no” route from the decision operation 408), then the process may advance to an operation 412. At 412, the voice assistant service 102 may implement supplemented audio request. For example, the voice assistant service 102 may provide an audio response to the user via the voice assistant device or voice assistant application.”; Examiner’s Note: BALIGAR discloses requesting additional information from the user, which is received via user speech input, when additional information is required to fulfill the user’s request) cause, ... a prompt to be provided to the user ... (BALIGAR, col. 9, lines 6-13: “) At 410, the voice assistant service 102 may request additional information from the user by sending an audio request to the user via the voice assistant device or voice assistant application. The request may be a choice, such as “did you mean “buy a toaster or diapers”. In some instances, the response may be a request for more specific information, such as “Please provide more information so I can fulfill your request”) receive, via the audio sensor, an additional utterance from the user in response to the prompt, (BALIGAR, col. 4, lines 1-5: “Instead, the voice assistant device 110 may receive input from users by receiving spoken commands, which are converted to signals by the voice assistant device 110 and/or by a cloud service, and then processed, such as by an exchange of data with voice assistant service 102.”; BALIGAR, col. 2, lines 11-13: “The voice assistant system may include a user device that typically includes at least a network interface, a microphone, and a speaker.” BALIGAR, col. 5, lines 10-13: “Meanwhile, the voice assistant service 102 may receive a message 122 of the audible request of “buy this”, which was received via a microphone of the voice assistance device 110.”; BALIGAR, col. 8, line 60 – col. 9, line 5: “At 408, the voice assistant service 102 may determine whether to request more information to determine a context of an audio request. For example, when the voice assistant service 102 includes enough information to understand and respond to an audio request that is supplemented by context information derived from the context queue, then the voice assistant service 102 may not request additional information from the user. When the voice assistant service 102 does not include enough information to understand and respond to the audio request that is supplemented by context information derived from the context queue (following the “yes” route from the decision operation 408), then the process 400 may advance to an operation 410.”; Examiner’s Note: voice assistant device 110 has a microphone (corresponding to recited “audio sensor”) that receives additional speech inputs from users in response to the request for more information in operation 410) determine, ... the task based on the utterance and the additional utterance, and (BALIGAR, col. 9, lines 19-23: “At 412, the voice assistant service 102 may implement supplemented audio request. For example, the voice assistant service 102 may provide an audio response to the user via the voice assistant device or voice assistant application.”; Examiner’s Note: for example, a task may be to provide the user with information in the form of an audio response) ... perform the task. (BALIGAR, col. 12, lines 13-18: “At 618, the voice assistant service 102 may transmit a request to the content provider to cause a refresh of information served to a device associated with the user who made the audio request, to cause output of more books, such as by graphically outputting recommended books similar to the book referenced via the contextual information.”; Examiner’s Note: BALIGAR provides several examples of a voice assistant transmitting requests to other systems (such as a content provider) to perform a task requested by the user, such as causing the system to output an audio book) However, BALIGAR fails to explicitly teach: A vehicle comprising: call a first machine learning (ML) model in response to an audio sensor receiving an utterance from a user ... via the first machine learning (ML) model ... by the vehicle ... store each word of the utterance in a corresponding slot of a data structure to create a stored utterance normalize a size of the stored utterance by adding padding to the data structure generate a converted data structure with the padding removed call, by the first ML model, a second ML model by the second ML model ... based on the converted data structure, ... controlled by a second ML model, different than the first ML model ... ... via the first ML model ... cause the vehicle ... However, in a related field of endeavor (automated speech recognition and natural language understanding, see col. 2, lines 7-18), ATLURI teaches and makes obvious: A vehicle comprising: (ATLURI, col. 45, lines 7-13: “For example, a speech-controllable device 110a, a smart phone 110b, a smart watch 110c, a tablet computer 110d, a vehicle 110e, a speech-controllable display device 110f, a smart television 110g, a washer/dryer 110h, a refrigerator 110i, and/or a microwave 110j may be connected to the network(s) 199 through a wireless service provider, over a Wi-Fi or cellular network connection, or the like.”; Examiner’s Note: the BALIGAR-ATLURI combination now implements the voice assistant of BALIGAR into the vehicle of ATLURI as depicted in Fig. 11 of ATLURI) call a first machine learning (ML) model in response to an audio sensor receiving an utterance from a user (ATLURI, col. 5, lines 11-32: “The search skill component 190a may include at least an action handler component 330 and a suggested user input component 340, which are described in further detail below in relation to FIGS. 3-5. The search skill component 190a may process the ASR data and determine an action to be performed in response to the user input. The action handler component 330 may determine to search one or more data sources for information relating to an entity included in the user input. For example, for the user input “show me gifts for mom,” the action handler component 330 may determine to search a retail catalog for products. As another example, for the user input “what is happening in California,” the action handler component 330 may determine to search news articles. As described below in relation to FIG. 3, the action handler component 330 may invoke one or more components to retrieve item results corresponding to the information requested in the user input. The action handler component 330 may determine which of the retrieved item results are to be present to the user 105. Such determination may be based on user profile data associated with the user 105.” ATLURI, col. 8, lines 32-38: “The input data 202 may be generated by a user via a keyboard, touchscreen, microphone, camera, or other such input device associated with the device 110. In other embodiments, the input data 202 is generated by the ASR component 150, as described herein, from audio data corresponding to a spoken input received from the user 105.”; ATLURI, col. 21, lines 26-32: “In some embodiments, the action handler component 330 may use rules-based processing and/or ML models to determine which component(s) to invoke based on the action data 322. The action handler component 330 may process with respect to each of the actions included in the N-best list in the action data 322. In some embodiments, the rules and/or ML models may be domain-specific.” ATLURI, col. 42, lines 48-54: “One or more of the components described herein may employ a machine learning (ML) model(s). Generally, ML models may be trained and operated according to various machine learning techniques. Such techniques may include, for example, neural networks (such as deep neural networks and/or recurrent neural networks), inference engines, trained classifiers, etc.”; Examiner’s Note: ATLURI discloses that components, including the action handler component 330 (corresponding to the recited “first ML model”) can be implemented using machine learning models and techniques; the BALIGAR-ATLURI combination now implements the voice assistant of BALIGAR into the vehicle of ATLURI, such that the voice assistant of BALIGAR utilizes ASR as in ATLURI to determine tasks of the vehicle (such as asking the vehicle, including the assistant of the vehicle, to perform a task such as looking up information)) determine, via the first machine learning (ML) model, that the utterance lacks sufficient information to determine a task to be performed, by the vehicle (ATLURI, col. 5, lines 11-32: “The search skill component 190a may include at least an action handler component 330 and a suggested user input component 340, which are described in further detail below in relation to FIGS. 3-5. The search skill component 190a may process the ASR data and determine an action to be performed in response to the user input. The action handler component 330 may determine to search one or more data sources for information relating to an entity included in the user input. For example, for the user input “show me gifts for mom,” the action handler component 330 may determine to search a retail catalog for products. As another example, for the user input “what is happening in California,” the action handler component 330 may determine to search news articles. As described below in relation to FIG. 3, the action handler component 330 may invoke one or more components to retrieve item results corresponding to the information requested in the user input. The action handler component 330 may determine which of the retrieved item results are to be present to the user 105. Such determination may be based on user profile data associated with the user 105.” ATLURI, col. 8, lines 32-38: “The input data 202 may be generated by a user via a keyboard, touchscreen, microphone, camera, or other such input device associated with the device 110. In other embodiments, the input data 202 is generated by the ASR component 150, as described herein, from audio data corresponding to a spoken input received from the user 105.”; ATLURI, col. 42, lines 48-54: “One or more of the components described herein may employ a machine learning (ML) model(s). Generally, ML models may be trained and operated according to various machine learning techniques. Such techniques may include, for example, neural networks (such as deep neural networks and/or recurrent neural networks), inference engines, trained classifiers, etc.”; Examiner’s Note: ATLURI discloses that components, including the action handler component 330 (corresponding to the recited “first ML model”) can be implemented using machine learning models and techniques; the BALIGAR-ATLURI combination now implements the voice assistant of BALIGAR into the vehicle of ATLURI, such that the voice assistant of BALIGAR utilizes the action handler component 330 as in ATLURI to determine tasks of the vehicle (such as asking the vehicle, including the assistant of the vehicle, to perform a task such as looking up information)) call, by the first ML model, a second ML model (ATLURI, col. 21, lines 26-32: “In some embodiments, the action handler component 330 may use rules-based processing and/or ML models to determine which component(s) to invoke based on the action data 322. The action handler component 330 may process with respect to each of the actions included in the N-best list in the action data 322. In some embodiments, the rules and/or ML models may be domain-specific.”; Examiner’s Note: the BALIGAR-ATLURI combination now has the action handler component 330 of ATLURI call a second model that can be implemented using machine learning, such as Q&A module 334 as depicted in Fig. 3) cause, by the second ML model, a prompt to be provided to the user (ATLURI, col. 25, lines 37-42: “For example, the action handler component 330 may select the system response to be “AnswerQuestion” and may invoke the Q&A component 334 to retrieve item results, and may use the received item results to determine the system response data 348 (without invoking the item retrieval component 332 or the navigation component 336).”; Examiner’s Note: the BALIGAR-ATLURI combination now has the action handler component 330 of ATLURI invoke Q&A module 334 as depicted in Fig. 3 to perform Q&A tasks, where such tasks can iteratively request additional information as taught by BALIGAR) determine, via the first ML model, the task based on the utterance and the additional utterance (ATLURI, col. 25, lines 37-42: “For example, the action handler component 330 may select the system response to be “AnswerQuestion” and may invoke the Q&A component 334 to retrieve item results, and may use the received item results to determine the system response data 348 (without invoking the item retrieval component 332 or the navigation component 336).”; Examiner’s Note: the BALIGAR-ATLURI combination now has the action handler component 330 of ATLURI invoke Q&A module 334 as depicted in Fig. 3 to perform Q&A tasks, where such tasks can iteratively request additional information as taught by BALIGAR, where the original request corresponds to the recited “utterance” and the follow-up requests of BALIGAR correspond to the recited “additional utterance”) cause the vehicle to perform the task (ATLURI, col. 45, lines 7-13: “For example, a speech-controllable device 110a, a smart phone 110b, a smart watch 110c, a tablet computer 110d, a vehicle 110e, a speech-controllable display device 110f, a smart television 110g, a washer/dryer 110h, a refrigerator 110i, and/or a microwave 110j may be connected to the network(s) 199 through a wireless service provider, over a Wi-Fi or cellular network connection, or the like.”; Examiner’s Note: the BALIGAR-ATLURI combination now implements the voice assistant of BALIGAR into the vehicle of ATLURI as depicted in Fig. 11 of ATLURI in order to perform tasks, such as Q&A tasks as disclosed by ATLURI) Before the effective filing date of the present application, it would have been obvious to one of ordinary skill in the art to combine the teachings of BALIGAR with the teachings of ATLURI as explained above. As disclosed by ATLURI, one of ordinary skill would have been motivated to do so in order to provide “techniques for enabling a conversational search and browsing experience for a user, where the user may search and explore items or topics of interest using a multi-modal interface and a multi-turn dialog exchange. A system of the present disclosure can receive a spoken input from a user, respond to the spoken input by displaying search results, and receive a further spoken input interacting with the displayed search results.” (col. 2, lines 56-63). However, BALIGAR and ATLURI fail to explicitly teach: store each word of the utterance in a corresponding slot of a data structure to create a stored utterance normalize a size of the stored utterance by adding padding to the data structure generate a converted data structure with the padding removed ... based on the converted data structure However, in a related field of endeavor (text processing by conversational agents, see para. 0001), BACHRACH teaches and makes obvious: store each word of the utterance in a corresponding slot of a data structure to create a stored utterance (BACHRACH, para. 0033: “In certain examples, each text string 215 may be pre-processed. One method of pre-processing is text tokenization. Text tokenization splits a continuous sequence of characters into one or more discrete sets of characters, e.g. where each character is represented by a character encoding. ... FIG. 2 shows an example result of text tokenization performed on the sequence of strings 210 in the form of character set arrays 220. For example, character set array 225—[‘how’, ‘can’, ‘i’, ‘help’, ‘?’]— is generated by tokenizing string 215—“How can I help?”. Each character set array may be of a different length. In certain cases, a maximum array length may be set, e.g. as 50 or 100 tokens. In these cases, entries in the array that follow the end of a message may be padded with a special token (e.g. <PAD>).” Examiner’s Note: BACHRACH discloses storing text data in character arrays (corresponding to recited “data structure”); the BALIGAR-ATLURI-BACHRACH combination now uses the ASR of ATLURI to convert speech to text, and then stores such text utterances in character arrays as disclosed by BACHRACH) normalize a size of the stored utterance by adding padding to the data structure (BACHRACH, para. 0033: “In certain examples, each text string 215 may be pre-processed. One method of pre-processing is text tokenization. Text tokenization splits a continuous sequence of characters into one or more discrete sets of characters, e.g. where each character is represented by a character encoding. ... FIG. 2 shows an example result of text tokenization performed on the sequence of strings 210 in the form of character set arrays 220. For example, character set array 225—[‘how’, ‘can’, ‘i’, ‘help’, ‘?’]— is generated by tokenizing string 215—“How can I help?”. Each character set array may be of a different length. In certain cases, a maximum array length may be set, e.g. as 50 or 100 tokens. In these cases, entries in the array that follow the end of a message may be padded with a special token (e.g. <PAD>).” Examiner’s Note: BACHRACH discloses that each character array can be set as a specific size, and using special <PAD> tokens for any blanks; the BALIGAR-ATLURI-BACHRACH combination now uses the ASR of ATLURI to convert speech to text, and then stores such text utterances in character arrays as disclosed by BACHRACH, where the character arrays are in a specific size and filled with <PAD> tokens) cause, by the second ML model, a prompt to be provided to the user based on the converted data structure (BACHRACH, para. 0033: “In certain examples, each text string 215 may be pre-processed. One method of pre-processing is text tokenization. Text tokenization splits a continuous sequence of characters into one or more discrete sets of characters, e.g. where each character is represented by a character encoding. ... FIG. 2 shows an example result of text tokenization performed on the sequence of strings 210 in the form of character set arrays 220. For example, character set array 225—[‘how’, ‘can’, ‘i’, ‘help’, ‘?’]— is generated by tokenizing string 215—“How can I help?”. Each character set array may be of a different length. In certain cases, a maximum array length may be set, e.g. as 50 or 100 tokens. In these cases, entries in the array that follow the end of a message may be padded with a special token (e.g. <PAD>).” Examiner’s Note: BACHRACH discloses storing text data in character arrays (corresponding to recited “data structure”); the BALIGAR-ATLURI-BACHRACH combination now uses the ASR of ATLURI to convert speech to text, and then stores such text utterances in character arrays as disclosed by BACHRACH, where such character array is utilized by the Q&A module 334 of ATLURI when determining Q&A task information) Before the effective filing date of the present application, it would have been obvious to one of ordinary skill in the art to combine the teachings of BALIGAR with the teachings of ATLURI and BACHRACH as explained above. As disclosed by BACHRACH, one of ordinary skill would have been motivated to do so in order to “efficiently provide responses to a large number of user queries.” (para. 0005). One of ordinary skill would understand that using a character array for storage is a straightforward and efficient structure that allows for direct access to looking up data. However, BALIGAR, ATLURI, and BACHRACH fail to explicitly teach: generate a converted data structure with the padding removed However, in a related field of endeavor (natural language processing, see para. 0004), DANG teaches and makes obvious: generate a converted data structure with the padding removed (DANG, para. 0112: “This is because, even if unwanted characters and pads to be removed or reduced by regularization are included in mini-batches, they would be given very small attention weights under optimized parameter values and, thus, contribute little to model accuracy.”; Examiner’s Note: BACHRACH discloses storing text data in character arrays (corresponding to recited “data structure”); the BALIGAR-ATLURI-BACHRACH-DANG combination now removes padding as disclosed by DANG) Before the effective filing date of the present application, it would have been obvious to one of ordinary skill in the art to combine the teachings of BALIGAR with the teachings of ATLURI, BACHRACH, and DANG as explained above. One of ordinary skill would be motivated to do so in order to reduce unwanted tokens from being processed by a machine learning model, which would require additional computational resources. Regarding Claim 2 BALIGAR, ATLURI, BACHRACH, and DANG disclose the vehicle of claim 1 as explained above. BALIGAR further teaches: prompt the user for the additional utterance based on content included in a most recently received utterance. (BALIGAR, col. 8, lines 60-67: “At 408, the voice assistant service 102 may determine whether to request more information to determine a context of an audio request. For example, when the voice assistant service 102 includes enough information to understand and respond to an audio request that is supplemented by context information derived from the context queue, then the voice assistant service 102 may not request additional information from the user”; Examiner’s Note: As shown in Fig. 4, in the first iteration the request for information is based on the initial utterance (corresponding to recited “most recently received utterance”) and can request more information at step 408 for an additional audio input from the user) Regarding Claim 3 BALIGAR, ATLURI, BACHRACH, and DANG disclose the vehicle of claim 1 as explained above. BALIGAR further teaches: wherein the prompt is provided by a virtual assistant ... (BALIGAR, col. 2, lines 1-21: “The voice assistant system may include any system and/or device that receives audio commands from a user, processes the audio, possibly using speech to text algorithms and/or natural language processing (NLP) algorithms, to determine text, returns a reply based on the text, converts the reply to an audio output using text to speech algorithms, and causes a speaker to output the audio output. Examples of voice assistant systems include Alexa® provided by Amazon.com® of Seattle, Wash., Siri® provided by Apple Corp.® of Cupertino, Calif., and Cortana® provided by Microsoft Corp.® of Redmond, Wash. The voice assistant system may include a user device that typically includes at least a network interface, a microphone, and a speaker. The user device may be a smart phone, a dedicated device, and/or other devices controlled by users and located proximate to the users. The voice assistant system may include a service engine, which may be stored in a remote location (e.g., via remote computing devices such as in a cloud computing configuration, etc.), stored in a local device (e.g., a smartphone, a dedicated voice assistant device, etc.) and/or a combination of both.”) However, BALIGAR fails to explicitly teach: ... of the vehicle, and the user is a passenger within vehicle However, in a related field of endeavor (“establishing hands-free communications between an electronic device and a user via a software assistant” see para. 0002), NOY teaches: ... of the vehicle, and the user is a passenger within vehicle (ATLURI, col. 45, lines 7-13: “For example, a speech-controllable device 110a, a smart phone 110b, a smart watch 110c, a tablet computer 110d, a vehicle 110e, a speech-controllable display device 110f, a smart television 110g, a washer/dryer 110h, a refrigerator 110i, and/or a microwave 110j may be connected to the network(s) 199 through a wireless service provider, over a Wi-Fi or cellular network connection, or the like.”; Examiner’s Note: the BALIGAR-ATLURI-BACHRACH-DANG combination now implements the voice assistant of BALIGAR into the vehicle of ATLURI as depicted in Fig. 11 of ATLURI) Before the effective filing date of the present application, it would have been obvious to one of ordinary skill in the art to combine the teachings of BALIGAR with the teachings of ATLURI, BACHRACH, and DANG as explained above. As disclosed by ATLURI, one of ordinary skill would have been motivated to do so in order to provide “techniques for enabling a conversational search and browsing experience for a user, where the user may search and explore items or topics of interest using a multi-modal interface and a multi-turn dialog exchange. A system of the present disclosure can receive a spoken input from a user, respond to the spoken input by displaying search results, and receive a further spoken input interacting with the displayed search results.” (col. 2, lines 56-63). Regarding Claim 4 BALIGAR, ATLURI, BACHRACH, and DANG disclose the vehicle of claim 1 as explained above. BALIGAR further teaches: request the user to confirm content included in a previously received utterance. (BALIGAR, col. 4, lines 16-21: “As discussed herein, the voice assistant service 102 may be configured to engage in a dialog to receive an order of one or more items from the user, including facilitating selection and confirmation of items, and cause those items to be fulfilled and delivered to the use, or for other tasks or fulfillment of audible requests.”; BALIGAR, col. 11, lines 14-16: “A confirmation page or reply (possibly via audio) may confirm completion of this action.”; Examiner’s Note: BALIGAR discloses asking the user to confirm ordered items via dialog) Regarding Claim 5 BALIGAR, ATLURI, BACHRACH, and DANG disclose the vehicle of claim 1 as explained above. BALIGAR further teaches: wherein the processor is configured to predict an intent of the user after each utterance based on an aggregation of utterances during a conversation with the user, and (BALIGAR, col. 9, lines 19-23: “At 412, the voice assistant service 102 may implement supplemented audio request. For example, the voice assistant service 102 may provide an audio response to the user via the voice assistant device or voice assistant application.”; Examiner’s Note: for example, the voice assistant can predict that the user has asked a question that needs to be responded to (e.g., the “intent” is the request for information), and as shown by Fig. 4, this can be after a number of iterations for additional information, where the cumulative user speech inputs corresponds to the recited “aggregation of utterances during a conversation”) when the processor determines that the sufficient information is available, the processor is configured to determine that sufficient information is available based on the aggregation of utterances. (BALIGAR, col. 8, lines 60-67: “At 408, the voice assistant service 102 may determine whether to request more information to determine a context of an audio request. For example, when the voice assistant service 102 includes enough information to understand and respond to an audio request that is supplemented by context information derived from the context queue, then the voice assistant service 102 may not request additional information from the user”; Examiner’s Note: As shown in Fig. 4, after step 410 (request for more information via audio in instances where not enough information is available), the flow returns to step 406 and then step 408 to analyze whether the additional information provided by the user is sufficient to understand and respond to the user’s request, where such repeated iterations resulting in cumulative user speech inputs corresponds to the recited “aggregation of utterances”) Regarding Claim 7 BALIGAR, ATLURI, BACHRACH, and DANG disclose the vehicle of claim 1 as explained above. BALIGAR further teaches: predict, ... a next action to be taken by the virtual assistant based on the utterance and a previously received utterance associated with a same task. (BALIGAR, col. 9, lines 19-23: “At 412, the voice assistant service 102 may implement supplemented audio request. For example, the voice assistant service 102 may provide an audio response to the user via the voice assistant device or voice assistant application.”; Examiner’s Note: for example, the voice assistant can predict that the user has asked a question that needs to be responded to and then predict the response, and as shown by Fig. 4, this can be after a number of iterations for additional information, where the cumulative user speech inputs corresponds to the recited “previously received utterance associated with a same task”) However, BALIGAR fails to explicitly teach: via the second ML model, ... However, in a related field of endeavor (automated speech recognition and natural language understanding, see col. 2, lines 7-18), ATLURI teaches and makes obvious: predict, via the second ML model, a next action to be taken by the virtual assistant based on the utterance and a previously received utterance associated with a same task. (ATLURI, col. 25, lines 37-42: “For example, the action handler component 330 may select the system response to be “AnswerQuestion” and may invoke the Q&A component 334 to retrieve item results, and may use the received item results to determine the system response data 348 (without invoking the item retrieval component 332 or the navigation component 336).”; Examiner’s Note: the BALIGAR-ATLURI-BACHRACH-DANG combination now has the action handler component 330 of ATLURI invoke Q&A module 334 as depicted in Fig. 3 to perform Q&A tasks, where the next action is to provide the answer based on the utterances having sufficient information) Before the effective filing date of the present application, it would have been obvious to one of ordinary skill in the art to combine the teachings of BALIGAR with the teachings of ATLURI, BACHRACH, and DANG as explained above. As disclosed by ATLURI, one of ordinary skill would have been motivated to do so in order to provide “techniques for enabling a conversational search and browsing experience for a user, where the user may search and explore items or topics of interest using a multi-modal interface and a multi-turn dialog exchange. A system of the present disclosure can receive a spoken input from a user, respond to the spoken input by displaying search results, and receive a further spoken input interacting with the displayed search results.” (col. 2, lines 56-63). Claim 9 recites a method that corresponds to the vehicle of claim 1 and is therefore rejected for the same reasons explained above with respect to claim 1. Claim 10 depends from claim 9 and claims a method that corresponds to the apparatus of claim 2, and is therefore rejected for the same reasons explained above with respect to claims 2 and 9. Claim 11 depends from claim 9 and claims a method that corresponds to the apparatus of claim 3, and is therefore rejected for the same reasons explained above with respect to claims 3 and 9. Claim 12 depends from claim 9 and claims a method that corresponds to the apparatus of claim 4, and is therefore rejected for the same reasons explained above with respect to claims 4 and 9. Claim 13 depends from claim 9 and claims a method that corresponds to the apparatus of claim 5, and is therefore rejected for the same reasons explained above with respect to claims 5 and 9. Claim 15 depends from claim 9 and claims a method that corresponds to the apparatus of claim 7, and is therefore rejected for the same reasons explained above with respect to claims 7 and 9. Claim 17 recites a non-transitory computer-readable medium that corresponds to the vehicle of claim 1 and is therefore rejected for the same reasons explained above with respect to claim 1. Claim 18 depends from claim 17 and claims a non-transitory computer-readable medium that corresponds to the apparatus of claim 2, and is therefore rejected for the same reasons explained above with respect to claims 2 and 17. Claim 19 depends from claim 17 and claims a non-transitory computer-readable medium that corresponds to the apparatus of claim 3, and is therefore rejected for the same reasons explained above with respect to claims 3 and 17. Claim 20 depends from claim 17 and claims a non-transitory computer-readable medium that corresponds to the apparatus of claim 4, and is therefore rejected for the same reasons explained above with respect to claims 4 and 17. Claims 6 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over BALIGAR in view of PANDEY and NOY and further in view of US 20170132199 A1, hereinafter referenced as VESCOVI. Regarding Claim 6 BALIGAR, ATLURI, BACHRACH, and DANG disclose the vehicle of claim 1 as explained above. However, BALIGAR, ATLURI, BACHRACH, and DANG fail to explicitly teach: wherein the processor is configured to determine, via the first ML model, that a sequence of utterances is required to make a prediction based on the utterance. However, in a related field of endeavor, (virtual assistants, see para. 0003), VESCOVI teaches: wherein the processor is configured to determine, via the first ML model, that a sequence of utterances is required to make a prediction based on the utterance. (VESCOVI, para. 0277: “As illustrated in the example of FIG. 8A, the user requests that the digital assistant let a visitor (in this example, Tomas) into his apartment when that visitor arrives. The digital assistant determines whether the user request corresponds to at least one of a plurality of plan templates 802, as described below in greater detail relative to FIGS. 9A-9F. A plan template 802 includes a set of instructions 804 and corresponding inputs/outputs 806. As illustrated in the example of FIG. 8B, a generic plan template 802 includes a set of ordered instructions 804, beginning with one or more instructions 804 to gather information.”; VESCOVI, para. 0283: “Referring to FIGS. 8E-8H, the inputs 806 associated with the “time expected” and “date expected” instructions in this particular example are not optional. Because those inputs 806 are not optional, the digital assistant has insufficient information to generate a plan with this plan template 802 if no input 806 is received in association with either of the “time expected” or “date expected” instructions 804. “Sufficient information” is the minimum information with which the digital assistant can generate a plan. Because the digital assistant cannot generate a plan based on the plan template 802 if it does not receive inputs associated with both the “time expected” and “date expected” instructions 804, the digital assistant initiates communication with the user to request sufficient information to generate a plan based on the plan template. As shown in FIG. 8E, the digital assistant requests 810 from the user: “on what date is he coming?” As shown in FIG. 8F, the user replies 812 with “tonight.” The digital assistant recognizes that the word “tonight” is associated with the same date on which the user spoke the reply 812, and as a result obtains today's date from the calendar module 248 or other suitable source. The time at which the visitor is to arrive is still required, so as shown in FIG. 8G, the digital assistant requests 814 “what time is he coming?” As shown in FIG. 8H, the user replies with “about 8:00 p.m.” Having received information associated with both the “time expected” and “date expected” instructions 804, the digital assistant now has sufficient information to generate a plan based on the plan template 802.”; Examiner’s Note: VESCOVI teaches a virtual assistant that uses plan templates 802, including plan templates that require multiple required inputs, in order to perform a task, where the multiple required inputs can be gathered based on a sequence of utterances (see Fig. 8H, requiring input utterances 812 and 816); the BALIGAR-ATLURI-BACRHACH-DANG-VESCOVI combination now modifies the voice assistant of BALIGAR to use the plan templates of VESCOVI to solicit multiple input utterances from users in order to carry out a prediction (e.g., answering a query as in BALIGAR) using the plan templates of VESCOVI. Before the effective filing date of the present application, it would have been obvious to one of ordinary skill in the art to combine the teachings of BALIGAR with the teachings of ATLURI, BACHRACH, DANG and VESCOVI as explained above. As disclosed by VESCOVI, one of ordinary skill would have been motivated to do so because such plan templates of VESCOVI enable “more complex actions, and actions that rely upon contingent inputs.” (para. 0302). Claim 14 depends from claim 9 and claims a method that corresponds to the apparatus of claim 6, and is therefore rejected for the same reasons explained above with respect to claims 6 and 9. Claims 8 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over BALIGAR in view of ATLURI, BACHRACH, and DANG and further in view of US 20230353516 A1, hereinafter referenced as BANERJEE. Regarding Claim 8 BALIGAR, ATLURI, BACHRACH, and DANG disclose the vehicle of claim 1 as explained above. BALIGAR further teaches: wherein the prompt is provided by a virtual assistant (BALIGAR, col. 9, lines 6-13: “) At 410, the voice assistant service 102 may request additional information from the user by sending an audio request to the user via the voice assistant device or voice assistant application. The request may be a choice, such as “did you mean “buy a toaster or diapers”. In some instances, the response may be a request for more specific information, such as “Please provide more information so I can fulfill your request”) However, BALIGAR, ATLURI, BACHRACH, and DANG fail to explicitly teach: wherein the processor is configured to: identify, via the second ML model, a word of interest within the additional utterance and an additional question for the virtual assistant to ask the user based on the word of interest. However, in a related field of endeavor (AI voice systems for interacting with users, see para. 0001), BANERJEE teaches: identify, via the second ML model, a word of interest within the additional utterance and an additional question for the virtual assistant to ask the user based on the word of interest. (BANERJEE, para. 0092: “In either case, the machine learning algorithm determines the next question or action based on the user response. In the case of a first question in the form of a menu selection, the next question or action from the system is likely predetermined based on the selection. In the case of a first question asking for a free-form spoken or text input by the user, the system may analyze the response for keywords or keyword patterns to determine the next question or action.”); Examiner’s Note: BANERJEE teaches using a machine learning algorithm to analyze a user speech input for keywords or keyword patterns (corresponding to recited “word of interest within an additional utterance”) and using such keywords to determine a next question or action (corresponding to recited “additional question for the virtual assistant to ask the user based on the identified word of interest”); the BALIGAR-ATLURI-BACHRACH-DANG-BANERJEE combination now modifies the voice assistant of BALIGAR to use the machine learning algorithms of BANERJEE to analyze user speech for keywords and then to determine the next question for the voice assistant of BALIGAR to ask based on such identified keyword of BANERJEE). Before the effective filing date of the present application, it would have been obvious to one of ordinary skill in the art to combine the teachings of BALIGAR with the teachings of ATLURI, BACHRACH, DANG, and BANERJEE as explained above. As disclosed by BANERJEE, one of ordinary skill would have been motivated to do so in order to use an “AI algorithm [that] adaptively guides the dialog to achieve the most favorable outcome based on the current status of the dialog.” (para. 0005). Claim 16 depends from claim 9 and claims a method that corresponds to the apparatus of claim 8, and is therefore rejected for the same reasons explained above with respect to claims 8 and 9. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. US 20210103700 A1 (Toplyn). “In an exemplary embodiment, the NUMPY library for PYTHON is used to create an array from the list of pre-padded sequences and each sequence is split into input and output elements, where the output element is the last element in the sequence and is converted to a binary class matrix marking the corresponding word in the vocabulary.” (para. 0254). Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHAEL C LEE whose telephone number is (571)272-4933. The examiner can normally be reached M-F 12:00 pm - 8:00 pm ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Omar Fernandez Rivas can be reached at 571-272-2589. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MICHAEL C. LEE/Examiner, Art Unit 2128
Read full office action

Prosecution Timeline

Show 1 earlier event
Jan 08, 2026
Non-Final Rejection mailed — §101, §103
Mar 05, 2026
Response Filed
May 06, 2026
Final Rejection mailed — §101, §103
May 29, 2026
Examiner Interview Summary
May 29, 2026
Applicant Interview (Telephonic)
Jun 22, 2026
Request for Continued Examination
Jun 24, 2026
Response after Non-Final Action
Jul 16, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12645972
Performing Property Estimation Using Quantum Gradient Operation on Quantum Computing System
3y 7m to grant Granted Jun 02, 2026
Patent 12603081
METHOD AND SERVER FOR A TEXT-TO-SPEECH PROCESSING
4y 7m to grant Granted Apr 14, 2026
Patent 12602605
QUANTUM COMPUTER ARCHITECTURE BASED ON MULTI-QUBIT GATES
3y 11m to grant Granted Apr 14, 2026
Patent 12591915
METHODS AND SYSTEMS FOR DETERMINING RECOMMENDATIONS BASED ON REAL-TIME OPTIMIZATION OF MACHINE LEARNING MODELS
5y 0m to grant Granted Mar 31, 2026
Patent 12585743
INTERFACE ACCESS PROCESSING METHOD, COMPUTER DEVICE AND STORAGE MEDIUM
1y 6m to grant Granted Mar 24, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
62%
Grant Probability
88%
With Interview (+26.4%)
3y 3m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 153 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month