Prosecution Insights
Last updated: October 02, 2026
Application No. 18/793,044

METHOD AND DEVICE FOR CLASSIFYING UTTERANCE INTENT CONSIDERING CONTEXT SURROUNDING VEHICLE AND DRIVER

Final Rejection §103
Filed
Aug 02, 2024
Priority
Sep 11, 2023 — RE 10-2023-0120613 +1 more
Examiner
DUGDA, MULUGETA TUJI
Art Unit
2653
Tech Center
2600 — Communications
Assignee
Kia Corporation
OA Round
2 (Final)
83%
Grant Probability
Favorable
3-4
OA Rounds
9m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 83% — above average
83%
Career Allowance Rate
48 granted / 58 resolved
+20.8% vs TC avg
Strong +21% interview lift
Without
With
+21.2%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
15 currently pending
Career history
78
Total Applications
across all art units

Statute-Specific Performance

§101
18.6%
-21.4% vs TC avg
§103
59.9%
+19.9% vs TC avg
§102
18.6%
-21.4% vs TC avg
§112
2.6%
-37.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 58 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claims 1, 4-7, 10-13, and 16-18 are pending and claims 1, 7 and 13 are independent claims. Response to Arguments Applicant’s arguments, see Arguments page 6, filed on 06/22/2026, with respect to 35 USC § 101 claim rejections have been fully considered and are persuasive. The 35 USC § 101 claim rejections of claims 1, 4-7, 10-13, and 16-18 have been withdrawn. The reason for this withdrawal is the fact that the Applicant amended claim 1 and corresponding independent claims, and the Applicant particularly added an additional claim limitation “causing, via a vehicle voice recognition system, execution of at least one in-vehicle function corresponding to the determined intent” on claim 1 with a specific intention of overcoming the 101 and that limitation shows a practical application of the presently claimed invention (Arguments, page 6). Applicant's arguments, see Arguments pages 7-13, filed on 06/22/2026, with respect to 35 USC § 102 and 35 USC § 103 claim rejections have been fully considered but they are not persuasive. The Applicant argues that “Ullrich fails to disclose providing the utterance data to an intent classification model to attempt to determine an intent of the utterance and obtaining context information related to the utterance in response to failing to determine the intent of the utterance using the utterance data, the context information including status information of the vehicle…. Last fails to account for these deficiencies of Ullrich.” (Arguments, pages 7-8). The Examiner respectfully disagrees. The Examiner used a combination of references, Ullrich and Last, to reject the two limitations that were originally part of claim 2. In Ullrich, the context subsystem 300 may form a context prompt 138 by tokenizing received context data. The context prompt 138 may indicate an inference of the user's conversation intent based on data such as historical speech patterns and known device identities; OR, the method includes generating a context prompt at block 1310. For example, the context subsystem 300 illustrated in FIG. 3 may generate a context prompt. The context prompt may be a string of tokens representing an inference of the user's intent with regard to their multimedia artifacts. The second portion of the claim limitation argued by the Applicant has been mapped using Last, since Last discloses obtaining the context information related to the utterance in response to failing to determine the intent of the utterance. In Last, there is provided a computer-implemented method for determining a user intent from a speech input to effect a user intended action, the method comprising:… attempting to interpret, by each processing node, the same speech input based on the subset of words directly relevant to its associated context, to extract therefrom an output indicative of user intent, whereby each processing node is unable to interpret any portion of the same speech input containing a word outside of its subset of words, whereby a portion of the same speech input relating to the particular context of a first of the nodes is interpretable to the first processing node but is not interpretable to a second of the processing nodes; According to a fourth aspect, there is provided herein a processing node of a computing system for determining a user intent from a speech input to effect a user intended action, wherein the processing node is capable of understanding only a subset of words directly relevant to a particular context associated with the processing node, wherein the processing node is configured to attempt to interpret the speech input based on the subset of words directly relevant to the associated context, to extract therefrom an output indicative of user intent, whereby each processing node is unable to interpret any portion of the same speech input containing a word outside of its subset of words (Ullrich, para 0173 and Ullrich, para 0183; Last, para 0030, OR Last, para 0031). The Applicant argues that “Ullrich fails to disclose; obtaining a context-aware sentence from an output of a generative large language model by providing the prompt to the generative large language model; and providing the context-aware sentence to an intent classification model to determine the intent of the utterance, as is presently claimed… Last fails to account for these deficiencies of Ullrich” (Arguments, pages 9-10). The Examiner respectfully disagrees. The Examiner noted that the limitations presented here and argued were originally part of claims 7 and 13. These limitations were taught by Ullrich because Ullrich uses all contextual information and prior conversation history to modulate your responses. After each input, review the prior inputs and modify your subsequent predictions based on the context of the thread. Taking into account the current context, with spartan language, return a JSON string called ‘suggestions’ with three different and unique phrases without quotes. They should be complete sentences longer than two words. Do not include explanations. The phrases you respond with will be spoken by my speech generating device. The GenAI 118 may take in the prompt 144 from the prompt composer 116 and use this to generate a multimodal output 146. The GenAI 118 may consist of a pre-trained machine learning model, such as GPT. The GenAI 118 may generate a multimodal output 146 in the form of a token sequence that may be converted back into plaintext, or which may be consumed by a user agency process directly as a token sequence. Moreover, in Ullrich, the context prompt 138 may indicate an inference of the user's conversation intent based on data such as historical speech patterns and known device identities which is basically providing the context-aware sentence to an intent classification model to determine the intent of the utterance (Ullrich, para 0173, Ullrich, para 0086-0087). The Applicant argues that “Finally, Ullrich does not disclose the claimed technical architecture in which the output of the generative large language model is used to obtain a context-aware sentence, and that context-aware sentence is then supplied to an intent classification model to determine the intent of the utterance… Ullrich merely discloses generating contextual output, feedback, or multimodal output using a generative AI model,..” (Arguments, pages 11-13). The Examiner respectfully disagrees. Ullrich discloses technical architecture in which the output of the generative large language model is used to obtain a context-aware sentence, and that context-aware sentence is then supplied to an intent classification model to determine the intent of the utterance… Ullrich merely discloses generating contextual output, feedback, or multimodal output using a generative AI model. Please look at Figure 1, 10 and 14. Ullrich discloses context prompt 138, and user input prompt 140 may be sent to a prompt composer 116. The prompt composer 116 may consume the data including the biosignals prompt 136, context prompt 138, and user input prompt 140 tokens, and may construct a single token, a set of tokens, or a series of conditional or unconditional commands suitable to use as a prompt 144 for a GenAI 118 such as a Large Language Model (LLM), a Generative Pre-trained Transformer (GPT) like GPT-4, or a generalist agent such as Gato. Besides, the method includes utilizing, by a pre-trained Generative Artificial Intelligence (GenAI) model, the prompt to generate a multimodal output at block 410. For example, the GenAI 118 illustrated in FIG. 1 may utilize, by a pre-trained GenAI model, the prompt to generate a multimodal output. In one embodiment, the GenAI model may be at least one of large language models (LLMs), Generative Pre-trained Transformer (GPT) models, text-to-image creators, visual art creators, and generalist agent models. the method includes providing the user with multimodal output including personalized feedback helping them adjust their communication style to better connect with their conversation partners at block 1412. For example, the GenAI 118 illustrated in FIG. 1 may provide the user with multimodal output including personalized feedback helping them adjust their communication style to better connect with their conversation partners. The large multimodal output stage 120 may include a set of suggestions that are rendered into private audio for the user 102 (Ullrich, para 0068, 0111, 0195). Thus, the 35 USC § 103 claim rejections will be maintained. Claim Rejections - 35 USC § 103 5. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 4-7, 10-13 and 16-18 1- are rejected under 35 U.S.C. 103 as being unpatentable over Ullrich et al. Pat. App. No. US 20240419246 A1 (Ullrich) in view of Last et al. Pat App No. US 20260011328 A1 (Last) (EFD: 2023-08-08 & Foreign Priority Date: 2022-08-09). Regarding Claim 1, Ullrich discloses a computer-implemented method for determining an intent of a user’s utterance (Ullrich, para 0183, the context prompt may be a string of tokens representing an inference of the user's intent with regard to their multimedia artifacts), the method comprising: obtaining, by a processor, utterance data representing an utterance occurred within a vehicle (Ullrich, para 0207-0209, the method includes receiving context data from the vehicle's surroundings and performance at block 1606. For example, the context subsystem 300 illustrated in FIG. 3 may receive context data from the vehicle's surroundings and performance. Such data may be in the form of sensor data 110 from cameras and lidar, in addition to other device data 112 from the vehicle's computerized vehicle management unit… the method includes receiving at least one of the biosignals prompt, the context prompt, and an optional user input prompt at block 1610. For example, the prompt composer 116 illustrated in FIG. 1 may receive at least one of the biosignals prompt, the context prompt, and an optional user input prompt. The user input prompt 140 may include data from vehicle passengers; [i.e., “The user input prompt” can be an “utterance” from a driver or passenger in the vehicle]); providing, by the processor, the utterance data to an intent classification model to attempt to determine an intent of the utterance (Ullrich, para 0173, The context subsystem 300 may form a context prompt 138 by tokenizing received context data. The context prompt 138 may indicate an inference of the user's conversation intent based on data such as historical speech patterns and known device identities; OR, Ullrich, para 0183, the method includes generating a context prompt at block 1310. For example, the context subsystem 300 illustrated in FIG. 3 may generate a context prompt. The context prompt may be a string of tokens representing an inference of the user's intent with regard to their multimedia artifacts); the context information including status information of the vehicle (Ullrich, para 0065, This embodiment may provide the user 102 with capability augmentation or agency support by utilizing inference of the user's environment, physical state, history, and current desired capabilities as a user context, to be gathered at a context subsystem 300, described in greater detail with respect to FIG. 3.; Ullrich, para 207 - 208, the method includes receiving context data from the vehicle's surroundings and performance at block 1606. For example, the context subsystem 300 illustrated in FIG. 3 may receive context data from the vehicle's surroundings and performance.… the method includes generating a context prompt based on vehicle surroundings and performance at block 1608. For example, the context subsystem 300 illustrated in FIG. 3 may generate a context prompt based on vehicle surroundings and performance; [i.e., the received context data is based on the vehicle's surroundings and performance or based on the status of the vehicle]); generating, by the processor, a prompt based on the utterance data and the context information, the prompt including a task description, a function inventory, guided learning examples, the context information, and the utterance data (Ullrich, para 0208-0213, Figure 16, According to some examples, the method includes generating a context prompt based on vehicle surroundings and performance at block 1608. For example, the context subsystem 300 illustrated in FIG. 3 may generate a context prompt based on vehicle surroundings and performance…According to some examples, the method includes generating a string of tokens based on at least one of the biosignals prompt, the context prompt, and the optional user input prompt at block 1612. For example, the prompt composer 116 illustrated in FIG. 1 may generate a string of tokens based on at least one of the biosignals prompt, the context prompt, and the optional user input prompt…According to some examples, the method includes providing multimodal output of the real-time feedback at block 1616. For example, the GenAI 118 illustrated in FIG. 1 may provide multimodal output of the real-time feedback. In one embodiment, the real-time feedback may instruct the vehicle's autonomous control system to adjust speed, route, and other factors to improve passenger safety and comfort. The output of the GenAI 118 may be converted into multimodal sensations through vehicle instruments such as the steering wheel, driver heads up display and/or over the in-vehicle audio system. In one embodiment the system may utilize estimates of the user's anxiety or comfort, derived from one or more biosensors, in order to adapt the navigation or driving style of an autonomous vehicle…Personalizing E-Commerce Experience: FIG. 17 illustrates an example routine 1700 for personalizing e-commerce experience. Although the example routine 1700 depicts a particular sequence of operations, the sequence may be altered without departing from the scope of the present disclosure. For example, some of the operations depicted may be performed in parallel or in a different sequence that does not materially affect the function of the routine 1700. In other examples, different components of an example device or system that implements the routine 1700 may perform functions at substantially the same time or in a specific sequence; [i.e., Various task descriptions, function inventory, guided learning examples, context information, and the audio/utterance data are described in these paragraphs and some more sample information with related paragraphs/citations, which make use of the same components described above, i.e., GenAIs 118, Figures 1, 3, 16 and 17, which are also used to create the various prompts described above are given below]; “task descriptions” sample: Ullrich, para 0230, The GenAI 118, in one embodiment an LLM, may utilize additional cloud-based resources to facilitate this translation function, including connecting to travel related booking services on the user's behalf in order to provide concrete option planning for the user 102; “function inventory” sample: Ullrich, para 0204-0212, FIG. 16 illustrates an example routine 1600 for enhancing autonomous vehicle safety and comfort. Although the example routine 1600 depicts a particular sequence of operations, the sequence may be altered without departing from the scope of the present disclosure. For example, some of the operations depicted may be performed in parallel or in a different sequence that does not materially affect the function of the routine 1600. In other examples, different components of an example device or system that implements the routine 1600 may perform functions at substantially the same time or in a specific sequence; “guided learning examples” and “context information” sample: Ullrich, para 0128, the user 102 may explicitly select or direct components of the system. For example, the user 102 may be able to choose between GenAIs 118 that have been trained on a different corpus or training set if they prefer to have a specific type of interaction. In one example, the user 102 may select between a GenAI 118 trained on clinical background data or a GenAI 118 trained on legal background data. These models may provide distinct output tokens that are potentially more appropriate for a specific user-intended task or context; “audio/utterance data” sample: Ullrich, para 0225, the method includes receiving context data related to the user's surroundings at block 1806. For example, the context subsystem 300 illustrated in FIG. 3 may receive context data related to the user's surroundings. Context data may include sensor data 110 such as audio and video data capturing body language and spoken words from a user's conversation partner); obtaining, by the processor, a context-aware sentence from an output of a generative large language model by providing the prompt to the generative large language model (Ullrich, para 0086-0087, Use all contextual information and prior conversation history to modulate your responses. After each input, review the prior inputs and modify your subsequent predictions based on the context of the thread. Taking into account the current context, with spartan language, return a JSON string called ‘suggestions’ with three different and unique phrases without quotes. They should be complete sentences longer than two words. Do not include explanations. The phrases you respond with will be spoken by my speech generating device. The GenAI 118 may take in the prompt 144 from the prompt composer 116 and use this to generate a multimodal output 146. The GenAI 118 may consist of a pre-trained machine learning model, such as GPT. The GenAI 118 may generate a multimodal output 146 in the form of a token sequence that may be converted back into plaintext, or which may be consumed by a user agency process directly as a token sequence); and providing, by the processor, the context-aware sentence to an intent classification model to determine the intent of the utterance (Ullrich, para 0173, The context prompt 138 may indicate an inference of the user's conversation intent based on data such as historical speech patterns and known device identities). causing, by the processor, via a vehicle voice recognition system, execution of at least one in-vehicle function corresponding to the determined intent (Ullrich, para 0212, According to some examples, the method includes providing multimodal output of the real-time feedback at block 1616. For example, the GenAI 118 illustrated in FIG. 1 may provide multimodal output of the real-time feedback. In one embodiment, the real-time feedback may instruct the vehicle's autonomous control system to adjust speed, route, and other factors to improve passenger safety and comfort. The output of the GenAI 118 may be converted into multimodal sensations through vehicle instruments such as the steering wheel, driver heads up display and/or over the in-vehicle audio system. In one embodiment the system may utilize estimates of the user's anxiety or comfort, derived from one or more biosensors, in order to adapt the navigation or driving style of an autonomous vehicle; [i.e., “the real-time feedback may instruct the vehicle's autonomous control system to adjust speed, route, and other factors to improve passenger safety and comfort” as “execution of at least one in-vehicle function corresponding to the determined intent”]). Ulrich does not specifically disclose obtaining the context information related to the utterance in response to failing to determine the intent of the utterance. However, Last, in the same field of endeavor, discloses obtaining, by the processor, context information related to the utterance in response to failing to determine the intent of the utterance using the utterance data (Last para 0030, According to a third aspect, there is provided a computer-implemented method for determining a user intent from a speech input to effect a user intended action, the method comprising:… attempting to interpret, by each processing node, the same speech input based on the subset of words directly relevant to its associated context, to extract therefrom an output indicative of user intent, whereby each processing node is unable to interpret any portion of the same speech input containing a word outside of its subset of words, whereby a portion of the same speech input relating to the particular context of a first of the nodes is interpretable to the first processing node but is not interpretable to a second of the processing nodes; OR, Last, para 0031, According to a fourth aspect, there is provided herein a processing node of a computing system for determining a user intent from a speech input to effect a user intended action, wherein the processing node is capable of understanding only a subset of words directly relevant to a particular context associated with the processing node, wherein the processing node is configured to attempt to interpret the speech input based on the subset of words directly relevant to the associated context, to extract therefrom an output indicative of user intent, whereby each processing node is unable to interpret any portion of the same speech input containing a word outside of its subset of words … ; Last, para 0007-0008, If the node receives a speech input that does not relate to its narrow context, the node's intent recognition will fail. To build a system that is capable of handling a wider range of contexts, rather than attempting to expand the capabilities of an individual node, multiple such nodes, targeted on different narrow contexts, are deployed in parallel. When a speech input is received, the speech input is provided to each node for processing in parallel with each other node. Each node attempts to interpret the voice input and extract a user intent therefrom; in a typical scenario, only one node will be able to do so; [i.e., “context information related to the utterance in response to failing to determine the intent of the utterance using the utterance data” is determined using specific node intent recognition system]). Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Last in the method of Ullrich because this would enable implementation of different forms of intent recognition with various contexts, one such context being the currently more prevalent online ‘chatbots’ which typically have user interfaces with a chatbot through text input typed on a physical or virtual keyboard or provided via some other text input mechanism at their device, unlike conventional chatbots which are based on text rather than voice input (Last, para 0002). Regarding Claim 4, Ullrich discloses the method of claim 1, wherein the function inventory includes at least one in-vehicle function accessible through a vehicle voice recognition system (Ullrich, para 0070, Figure 3, the context subsystem 300 may generate a context prompt 138 token… Such a context prompt 138 may be generated by utilizing sensors… Such sensors may include … microphones configured to feed audio to a speech to text (STT) device; Ullrich, para 212, The GenAI 118 … may provide multimodal output of the real-time feedback. In one embodiment, the real-time feedback may instruct the vehicle's autonomous control system to adjust speed, route, and other factors to improve passenger safety and comfort. The output of the GenAI 118 may be converted into multimodal sensations through vehicle instruments such as the steering wheel, driver heads up display and/or over the in-vehicle audio system). Regarding Claim 5, Ullrich discloses the method of claim 1, wherein the guided learning examples include example utterance data (Ullrich, para 0225, the method includes receiving context data related to the user's surroundings at block 1806. For example, the context subsystem 300 illustrated in FIG. 3 may receive context data related to the user's surroundings. Context data may include sensor data 110 such as audio and video data capturing body language and spoken words from a user's conversation partner; OR, Ullrich, para 0091-0092, The user 102 may respond to the multimodal output 146 in a manner detectable through biosignals 106, and thus a channel may be provided to train the GenAI 118 based on user 102 response to multimodal output 146. In general, the user agency and capability augmentation system 100 may be viewed as a kind of application framework that uses the biosignals prompt 136, context prompt 138, and user input prompt 140 sequences to facilitate interaction with an application, much as a user 102 would use their finger to interact with a mobile phone application running on a mobile phone operating system. Unlike a touchscreen or mouse/keyboard interface, this system incorporates real time user inputs along with an articulated description of their physical context and historical context to facilitate extremely efficient interactions to enable user agency. FIG. 1 shows the pathways signals take from input, by sensing devices, stored data, or the user 102, to output in the form of text-to-speech utterances 124, written text 126; [“written text” as “a sentence”]), an example context-aware sentence, and an example process of reasoning the example context-aware sentence from the example utterance data (Ullrich, para 0086-0087, Use all contextual information and prior conversation history to modulate your responses. After each input, review the prior inputs and modify your subsequent predictions based on the context of the thread. Taking into account the current context, with spartan language, return a JSON string called ‘suggestions’ with three different and unique phrases without quotes. They should be complete sentences longer than two words. Do not include explanations. The phrases you respond with will be spoken by my speech generating device. The GenAI 118 may take in the prompt 144 from the prompt composer 116 and use this to generate a multimodal output 146. The GenAI 118 may consist of a pre-trained machine learning model, such as GPT. The GenAI 118 may generate a multimodal output 146 in the form of a token sequence that may be converted back into plaintext, or which may be consumed by a user agency process directly as a token sequence). Regarding Claim 6, Ullrich discloses the method of claim 1, wherein the guided learning examples include example utterance data and an example context-aware sentence (Ullrich, para 0063, FIG. 1 illustrates a user agency and capability augmentation system 100 in accordance with one embodiment. The user agency and capability augmentation system 100 comprises a user 102, a wearable computing and biosignal sensing device 104, biosignals 106, background material 108, sensor data 110, other device data 112, application context 114 a prompt composer 116, a GenAI 118, a multimodal output stage 120, an encoder/parser 132, output modalities 122 such as an utterance 124, a written text 126, a multimodal artifact 128, an other user agency 130, and a non-language user agency device 134, a biosignals subsystem 200, and a context subsystem 300; [“written text” as “context-aware sentence”]; OR, Ullrich, para 0091-0092, The user 102 may respond to the multimodal output 146 in a manner detectable through biosignals 106, and thus a channel may be provided to train the GenAI 118 based on user 102 response to multimodal output 146. In general, the user agency and capability augmentation system 100 may be viewed as a kind of application framework that uses the biosignals prompt 136, context prompt 138, and user input prompt 140 sequences to facilitate interaction with an application, much as a user 102 would use their finger to interact with a mobile phone application running on a mobile phone operating system. Unlike a touchscreen or mouse/keyboard interface, this system incorporates real time user inputs along with an articulated description of their physical context and historical context to facilitate extremely efficient interactions to enable user agency. FIG. 1 shows the pathways signals take from input, by sensing devices, stored data, or the user 102, to output in the form of text-to-speech utterances 124, written text 126; [“written text” as “a sentence”]). Regarding Claim 7, Ullrich discloses a computing apparatus comprising: at least one processor (Ullrich, para 0447, The components of computing device 5000 may include, but are not limited to, one or more processors or processing units 5004); and a memory operably coupled to the at least one processor (Ullrich, para 0447, The components of computing device 5000 may include, but are not limited to, one or more processors or processing units 5004, a system memory 5002, and a bus 5024 that couples various system components including system memory 5002 to processor processing units 5004.), wherein the memory stores instructions for causing the at least one processor to perform operations in response to instructions executed by the at least one processor (Ullrich, para 1063, A “credit distribution circuit configured to distribute credits to a plurality of processor cores” is intended to cover, for example, an integrated circuit that has Circuitry that performs this function during operation, even if the integrated circuit in question is not currently being used (e.g., a power supply is not connected to it). Thus, an entity described or recited as “configured to” perform some task refers to something physical, such as a device, circuit, memory storing program instructions executable to implement the task, etc.), the operations including: obtaining utterance data representing an utterance occurred within a vehicle (Ullrich, para 0207-0209, the method includes receiving context data from the vehicle's surroundings and performance at block 1606. For example, the context subsystem 300 illustrated in FIG. 3 may receive context data from the vehicle's surroundings and performance. Such data may be in the form of sensor data 110 from cameras and lidar, in addition to other device data 112 from the vehicle's computerized vehicle management unit… the method includes receiving at least one of the biosignals prompt, the context prompt, and an optional user input prompt at block 1610. For example, the prompt composer 116 illustrated in FIG. 1 may receive at least one of the biosignals prompt, the context prompt, and an optional user input prompt. The user input prompt 140 may include data from vehicle passengers; [i.e., “The user input prompt” can be an “utterance” from a driver or passenger in the vehicle]); providing, by the processor, the utterance data to an intent classification model to attempt to determine an intent of the utterance (Ullrich, para 0173, The context subsystem 300 may form a context prompt 138 by tokenizing received context data. The context prompt 138 may indicate an inference of the user's conversation intent based on data such as historical speech patterns and known device identities; OR, Ullrich, para 0183, the method includes generating a context prompt at block 1310. For example, the context subsystem 300 illustrated in FIG. 3 may generate a context prompt. The context prompt may be a string of tokens representing an inference of the user's intent with regard to their multimedia artifacts); the context information including status information of the vehicle (Ullrich, para 0065, This embodiment may provide the user 102 with capability augmentation or agency support by utilizing inference of the user's environment, physical state, history, and current desired capabilities as a user context, to be gathered at a context subsystem 300, described in greater detail with respect to FIG. 3.; Ullrich, para 207 - 208, the method includes receiving context data from the vehicle's surroundings and performance at block 1606. For example, the context subsystem 300 illustrated in FIG. 3 may receive context data from the vehicle's surroundings and performance.… the method includes generating a context prompt based on vehicle surroundings and performance at block 1608. For example, the context subsystem 300 illustrated in FIG. 3 may generate a context prompt based on vehicle surroundings and performance; [i.e., the received context data is based on the vehicle's surroundings and performance or based on the status of the vehicle]); generating a prompt based on the utterance data and the context information, the prompt including a task description, a function inventory, guided learning examples, the context information, and the utterance data (Ullrich, para 0208-0213, According to some examples, the method includes generating a context prompt based on vehicle surroundings and performance at block 1608. For example, the context subsystem 300 illustrated in FIG. 3 may generate a context prompt based on vehicle surroundings and performance…According to some examples, the method includes generating a string of tokens based on at least one of the biosignals prompt, the context prompt, and the optional user input prompt at block 1612. For example, the prompt composer 116 illustrated in FIG. 1 may generate a string of tokens based on at least one of the biosignals prompt, the context prompt, and the optional user input prompt…According to some examples, the method includes providing multimodal output of the real-time feedback at block 1616. For example, the GenAI 118 illustrated in FIG. 1 may provide multimodal output of the real-time feedback. In one embodiment, the real-time feedback may instruct the vehicle's autonomous control system to adjust speed, route, and other factors to improve passenger safety and comfort. The output of the GenAI 118 may be converted into multimodal sensations through vehicle instruments such as the steering wheel, driver heads up display and/or over the in-vehicle audio system. In one embodiment the system may utilize estimates of the user's anxiety or comfort, derived from one or more biosensors, in order to adapt the navigation or driving style of an autonomous vehicle…Personalizing E-Commerce Experience: FIG. 17 illustrates an example routine 1700 for personalizing e-commerce experience. Although the example routine 1700 depicts a particular sequence of operations, the sequence may be altered without departing from the scope of the present disclosure. For example, some of the operations depicted may be performed in parallel or in a different sequence that does not materially affect the function of the routine 1700. In other examples, different components of an example device or system that implements the routine 1700 may perform functions at substantially the same time or in a specific sequence; [i.e., Various task descriptions, function inventory, guided learning examples, context information, and the audio/utterance data are described in these paragraphs and some more sample information with related paragraphs/citations, which make use of the same components described above, i.e., GenAIs 118, Figures 1, 3, 16 and 17, which are also used to create the various prompts described above are given below]; “task descriptions” sample: Ullrich, para 0230, The GenAI 118, in one embodiment an LLM, may utilize additional cloud-based resources to facilitate this translation function, including connecting to travel related booking services on the user's behalf in order to provide concrete option planning for the user 102; “function inventory” sample: Ullrich, para 0204-0212, FIG. 16 illustrates an example routine 1600 for enhancing autonomous vehicle safety and comfort. Although the example routine 1600 depicts a particular sequence of operations, the sequence may be altered without departing from the scope of the present disclosure. For example, some of the operations depicted may be performed in parallel or in a different sequence that does not materially affect the function of the routine 1600. In other examples, different components of an example device or system that implements the routine 1600 may perform function at substantially the same time or in a specific sequence; “guided learning examples” and “context information” sample: Ullrich, para 0128, the user 102 may explicitly select or direct components of the system. For example, the user 102 may be able to choose between GenAIs 118 that have been trained on a different corpus or training set if they prefer to have a specific type of interaction. In one example, the user 102 may select between a GenAI 118 trained on clinical background data or a GenAI 118 trained on legal background data. These models may provide distinct output tokens that are potentially more appropriate for a specific user-intended task or context; “audio/utterance data” sample: Ullrich, para 0225, the method includes receiving context data related to the user's surroundings at block 1806. For example, the context subsystem 300 illustrated in FIG. 3 may receive context data related to the user's surroundings. Context data may include sensor data 110 such as audio and video data capturing body language and spoken words from a user's conversation partner); obtaining a context-aware sentence from an output of a generative large language model by providing the prompt to the generative large language model (Ullrich, para 0086-0087, Use all contextual information and prior conversation history to modulate your responses. After each input, review the prior inputs and modify your subsequent predictions based on the context of the thread. Taking into account the current context, with spartan language, return a JSON string called ‘suggestions’ with three different and unique phrases without quotes. They should be complete sentences longer than two words. Do not include explanations. The phrases you respond with will be spoken by my speech generating device. The GenAI 118 may take in the prompt 144 from the prompt composer 116 and use this to generate a multimodal output 146. The GenAI 118 may consist of a pre-trained machine learning model, such as GPT. The GenAI 118 may generate a multimodal output 146 in the form of a token sequence that may be converted back into plaintext, or which may be consumed by a user agency process directly as a token sequence); and providing the context-aware sentence to the intent classification model to determine the intent of the utterance (Ullrich, para 0173, The context prompt 138 may indicate an inference of the user's conversation intent based on data such as historical speech patterns and known device identities). causing, by the processor, via a vehicle voice recognition system, execution of at least one in-vehicle function corresponding to the determined intent (Ullrich, para 0212, According to some examples, the method includes providing multimodal output of the real-time feedback at block 1616. For example, the GenAI 118 illustrated in FIG. 1 may provide multimodal output of the real-time feedback. In one embodiment, the real-time feedback may instruct the vehicle's autonomous control system to adjust speed, route, and other factors to improve passenger safety and comfort. The output of the GenAI 118 may be converted into multimodal sensations through vehicle instruments such as the steering wheel, driver heads up display and/or over the in-vehicle audio system. In one embodiment the system may utilize estimates of the user's anxiety or comfort, derived from one or more biosensors, in order to adapt the navigation or driving style of an autonomous vehicle; [i.e., “the real-time feedback may instruct the vehicle's autonomous control system to adjust speed, route, and other factors to improve passenger safety and comfort” as “execution of at least one in-vehicle function corresponding to the determined intent”]). Ulrich does not specifically disclose obtaining the context information related to the utterance in response to failing to determine the intent of the utterance. However, Last, in the same field of endeavor, discloses obtaining, by the processor, context information related to the utterance in response to failing to determine the intent of the utterance using the utterance data (Last para 0030, According to a third aspect, there is provided a computer-implemented method for determining a user intent from a speech input to effect a user intended action, the method comprising:… attempting to interpret, by each processing node, the same speech input based on the subset of words directly relevant to its associated context, to extract therefrom an output indicative of user intent, whereby each processing node is unable to interpret any portion of the same speech input containing a word outside of its subset of words, whereby a portion of the same speech input relating to the particular context of a first of the nodes is interpretable to the first processing node but is not interpretable to a second of the processing nodes; OR, Last, para 0031, According to a fourth aspect, there is provided herein a processing node of a computing system for determining a user intent from a speech input to effect a user intended action, wherein the processing node is capable of understanding only a subset of words directly relevant to a particular context associated with the processing node, wherein the processing node is configured to attempt to interpret the speech input based on the subset of words directly relevant to the associated context, to extract therefrom an output indicative of user intent, whereby each processing node is unable to interpret any portion of the same speech input containing a word outside of its subset of words … ; Last, para 0007-0008, If the node receives a speech input that does not relate to its narrow context, the node's intent recognition will fail. To build a system that is capable of handling a wider range of contexts, rather than attempting to expand the capabilities of an individual node, multiple such nodes, targeted on different narrow contexts, are deployed in parallel. When a speech input is received, the speech input is provided to each node for processing in parallel with each other node. Each node attempts to interpret the voice input and extract a user intent therefrom; in a typical scenario, only one node will be able to do so; [i.e., “context information related to the utterance in response to failing to determine the intent of the utterance using the utterance data” is determined using specific node intent recognition system]). Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Last in the method of Ullrich because this would enable implementation of different forms of intent recognition with various contexts, one such context being the currently more prevalent online ‘chatbots’ which typically have user interfaces with a chatbot through text input typed on a physical or virtual keyboard or provided via some other text input mechanism at their device, unlike conventional chatbots which are based on text rather than voice input (Last, para 0002). Regarding Claim 10, Ullrich discloses the computing apparatus of claim 7, wherein the function inventory includes at least one in-vehicle function accessible through a vehicle voice recognition system (Ullrich, para 0070, Figure 3, the context subsystem 300 may generate a context prompt 138 token… Such a context prompt 138 may be generated by utilizing sensors… Such sensors may include … microphones configured to feed audio to a speech to text (STT) device; Ullrich, para 212, The GenAI 118 … may provide multimodal output of the real-time feedback. In one embodiment, the real-time feedback may instruct the vehicle's autonomous control system to adjust speed, route, and other factors to improve passenger safety and comfort. The output of the GenAI 118 may be converted into multimodal sensations through vehicle instruments such as the steering wheel, driver heads up display and/or over the in-vehicle audio system). Regarding Claim 11, Ullrich discloses the computing apparatus of claim 7, wherein the guided learning examples include example utterance data (Ullrich, para 0225, the method includes receiving context data related to the user's surroundings at block 1806. For example, the context subsystem 300 illustrated in FIG. 3 may receive context data related to the user's surroundings. Context data may include sensor data 110 such as audio and video data capturing body language and spoken words from a user's conversation partner; OR, Ullrich, para 0091-0092, The user 102 may respond to the multimodal output 146 in a manner detectable through biosignals 106, and thus a channel may be provided to train the GenAI 118 based on user 102 response to multimodal output 146. In general, the user agency and capability augmentation system 100 may be viewed as a kind of application framework that uses the biosignals prompt 136, context prompt 138, and user input prompt 140 sequences to facilitate interaction with an application, much as a user 102 would use their finger to interact with a mobile phone application running on a mobile phone operating system. Unlike a touchscreen or mouse/keyboard interface, this system incorporates real time user inputs along with an articulated description of their physical context and historical context to facilitate extremely efficient interactions to enable user agency. FIG. 1 shows the pathways signals take from input, by sensing devices, stored data, or the user 102, to output in the form of text-to-speech utterances 124, written text 126; [“written text ” as “a sentence”]), an example context-aware sentence, and an example process of reasoning the example context-aware sentence from the example utterance data (Ullrich, para 0086-0087, Use all contextual information and prior conversation history to modulate your responses. After each input, review the prior inputs and modify your subsequent predictions based on the context of the thread. Taking into account the current context, with spartan language, return a JSON string called ‘suggestions’ with three different and unique phrases without quotes. They should be complete sentences longer than two words. Do not include explanations. The phrases you respond with will be spoken by my speech generating device. The GenAI 118 may take in the prompt 144 from the prompt composer 116 and use this to generate a multimodal output 146. The GenAI 118 may consist of a pre-trained machine learning model, such as GPT. The GenAI 118 may generate a multimodal output 146 in the form of a token sequence that may be converted back into plaintext, or which may be consumed by a user agency process directly as a token sequence). Regarding Claim 12, Ullrich discloses the computing apparatus of claim 7, wherein the guided learning examples include example utterance data and an example context-aware sentence (Ullrich, para 0063, FIG. 1 illustrates a user agency and capability augmentation system 100 in accordance with one embodiment. The user agency and capability augmentation system 100 comprises a user 102, a wearable computing and biosignal sensing device 104, biosignals 106, background material 108, sensor data 110, other device data 112, application context 114 a prompt composer 116, a GenAI 118, a multimodal output stage 120, an encoder/parser 132, output modalities 122 such as an utterance 124, a written text 126, a multimodal artifact 128, an other user agency 130, and a non-language user agency device 134, a biosignals subsystem 200, and a context subsystem 300; [“written text” as “context-aware sentence”]; OR, Ullrich, para 0091-0092, The user 102 may respond to the multimodal output 146 in a manner detectable through biosignals 106, and thus a channel may be provided to train the GenAI 118 based on user 102 response to multimodal output 146. In general, the user agency and capability augmentation system 100 may be viewed as a kind of application framework that uses the biosignals prompt 136, context prompt 138, and user input prompt 140 sequences to facilitate interaction with an application, much as a user 102 would use their finger to interact with a mobile phone application running on a mobile phone operating system. Unlike a touchscreen or mouse/keyboard interface, this system incorporates real time user inputs along with an articulated description of their physical context and historical context to facilitate extremely efficient interactions to enable user agency. FIG. 1 shows the pathways signals take from input, by sensing devices, stored data, or the user 102, to output in the form of text- to-speech utterances 124, written text 126; [“written text ” as “a sentence”]). Regarding Claim 13,s Ullrich discloses a non-transitory computer-readable recording medium in which instructions are stored, the instructions causing a computer including a processor to perform, when executed by the computer (Ullrich, para 450, System memory 5002 may include computer system readable media in the form of volatile memory, such as Random access memory (RAM) 5006 and/or cache memory 5010. Computing device 5000 may further include other removable/non-removable, volatile/non-volatile computer system storage media…): obtaining utterance data representing an utterance occurred within a vehicle (Ullrich, para 0207-0209, the method includes receiving context data from the vehicle's surroundings and performance at block 1606. For example, the context subsystem 300 illustrated in FIG. 3 may receive context data from the vehicle's surroundings and performance. Such data may be in the form of sensor data 110 from cameras and lidar, in addition to other device data 112 from the vehicle's computerized vehicle management unit… the method includes receiving at least one of the biosignals prompt, the context prompt, and an optional user input prompt at block 1610. For example, the prompt composer 116 illustrated in FIG. 1 may receive at least one of the biosignals prompt, the context prompt, and an optional user input prompt. The user input prompt 140 may include data from vehicle passengers; [i.e., “The user input prompt” can be an “utterance” from a driver or passenger in the vehicle]); providing, by the processor, the utterance data to an intent classification model to attempt to determine an intent of the utterance (Ullrich, para 0173, The context subsystem 300 may form a context prompt 138 by tokenizing received context data. The context prompt 138 may indicate an inference of the user's conversation intent based on data such as historical speech patterns and known device identities; OR, Ullrich, para 0183, the method includes generating a context prompt at block 1310. For example, the context subsystem 300 illustrated in FIG. 3 may generate a context prompt. The context prompt may be a string of tokens representing an inference of the user's intent with regard to their multimedia artifacts); the context information including status information of the vehicle (Ullrich, para 0065, This embodiment may provide the user 102 with capability augmentation or agency support by utilizing inference of the user's environment, physical state, history, and current desired capabilities as a user context, to be gathered at a context subsystem 300, described in greater detail with respect to FIG. 3.; Ullrich, para 207 - 208, the method includes receiving context data from the vehicle's surroundings and performance at block 1606. For example, the context subsystem 300 illustrated in FIG. 3 may receive context data from the vehicle's surroundings and performance.… the method includes generating a context prompt based on vehicle surroundings and performance at block 1608. For example, the context subsystem 300 illustrated in FIG. 3 may generate a context prompt based on vehicle surroundings and performance; [i.e., the received context data is based on the vehicle's surroundings and performance or based on the status of the vehicle]); generating a prompt based on the utterance data and the context information, the prompt including a task description, a function inventory, guided learning examples, the context information, and the utterance data (Ullrich, para 0208-0213, According to some examples, the method includes generating a context prompt based on vehicle surroundings and performance at block 1608. For example, the context subsystem 300 illustrated in FIG. 3 may generate a context prompt based on vehicle surroundings and performance…According to some examples, the method includes generating a string of tokens based on at least one of the biosignals prompt, the context prompt, and the optional user input prompt at block 1612. For example, the prompt composer 116 illustrated in FIG. 1 may generate a string of tokens based on at least one of the biosignals prompt, the context prompt, and the optional user input prompt…According to some examples, the method includes providing multimodal output of the real-time feedback at block 1616. For example, the GenAI 118 illustrated in FIG. 1 may provide multimodal output of the real-time feedback. In one embodiment, the real-time feedback may instruct the vehicle's autonomous control system to adjust speed, route, and other factors to improve passenger safety and comfort. The output of the GenAI 118 may be converted into multimodal sensations through vehicle instruments such as the steering wheel, driver heads up display and/or over the in-vehicle audio system. In one embodiment the system may utilize estimates of the user's anxiety or comfort, derived from one or more biosensors, in order to adapt the navigation or driving style of an autonomous vehicle…Personalizing E-Commerce Experience: FIG. 17 illustrates an example routine 1700 for personalizing e-commerce experience. Although the example routine 1700 depicts a particular sequence of operations, the sequence may be altered without departing from the scope of the present disclosure. For example, some of the operations depicted may be performed in parallel or in a different sequence that does not materially affect the function of the routine 1700. In other examples, different components of an example device or system that implements the routine 1700 may perform functions at substantially the same time or in a specific sequence; [i.e., Various task descriptions, function inventory, guided learning examples, context information, and the audio/utterance data are described in these paragraphs and some more sample information with related paragraphs/citations, which make use of the same components described above, i.e., GenAIs 118, Figures 1, 3, 16 and 17, which are also used to create the various prompts described above are given below]; “task descriptions” sample: Ullrich, para 0230, The GenAI 118, in one embodiment an LLM, may utilize additional cloud-based resources to facilitate this translation function, including connecting to travel related booking services on the user's behalf in order to provide concrete option planning for the user 102; “function inventory” sample: Ullrich, para 0204-0212, FIG. 16 illustrates an example routine 1600 for enhancing autonomous vehicle safety and comfort. Although the example routine 1600 depicts a particular sequence of operations, the sequence may be altered without departing from the scope of the present disclosure. For example, some of the operations depicted may be performed in parallel or in a different sequence that does not materially affect the function of the routine 1600. In other examples, different components of an example device or system that implements the routine 1600 may perform function at substantially the same time or in a specific sequence; “guided learning examples” and “context information” sample: Ullrich, para 0128, the user 102 may explicitly select or direct components of the system. For example, the user 102 may be able to choose between GenAIs 118 that have been trained on a different corpus or training set if they prefer to have a specific type of interaction. In one example, the user 102 may select between a GenAI 118 trained on clinical background data or a GenAI 118 trained on legal background data. These models may provide distinct output tokens that are potentially more appropriate for a specific user-intended task or context; “audio/utterance data” sample: Ullrich, para 0225, the method includes receiving context data related to the user's surroundings at block 1806. For example, the context subsystem 300 illustrated in FIG. 3 may receive context data related to the user's surroundings. Context data may include sensor data 110 such as audio and video data capturing body language and spoken words from a user's conversation partner); obtaining a context-aware sentence from an output of a generative large language model by providing the prompt to the generative large language model (Ullrich, para 0086-0087, Use all contextual information and prior conversation history to modulate your responses. After each input, review the prior inputs and modify your subsequent predictions based on the context of the thread. Taking into account the current context, with spartan language, return a JSON string called ‘suggestions’ with three different and unique phrases without quotes. They should be complete sentences longer than two words. Do not include explanations. The phrases you respond with will be spoken by my speech generating device. The GenAI 118 may take in the prompt 144 from the prompt composer 116 and use this to generate a multimodal output 146. The GenAI 118 may consist of a pre-trained machine learning model, such as GPT. The GenAI 118 may generate a multimodal output 146 in the form of a token sequence that may be converted back into plaintext, or which may be consumed by a user agency process directly as a token sequence); and providing the context-aware sentence to the intent classification model to determine the intent of the utterance (Ullrich, para 0173, The context prompt 138 may indicate an inference of the user's conversation intent based on data such as historical speech patterns and known device identities). causing, by the processor, via a vehicle voice recognition system, execution of at least one in-vehicle function corresponding to the determined intent (Ullrich, para 0212, According to some examples, the method includes providing multimodal output of the real-time feedback at block 1616. For example, the GenAI 118 illustrated in FIG. 1 may provide multimodal output of the real-time feedback. In one embodiment, the real-time feedback may instruct the vehicle's autonomous control system to adjust speed, route, and other factors to improve passenger safety and comfort. The output of the GenAI 118 may be converted into multimodal sensations through vehicle instruments such as the steering wheel, driver heads up display and/or over the in-vehicle audio system. In one embodiment the system may utilize estimates of the user's anxiety or comfort, derived from one or more biosensors, in order to adapt the navigation or driving style of an autonomous vehicle; [i.e., “the real-time feedback may instruct the vehicle's autonomous control system to adjust speed, route, and other factors to improve passenger safety and comfort” as “execution of at least one in-vehicle function corresponding to the determined intent”]). Ulrich does not specifically disclose obtaining, by the processor, context information related to the utterance in response to failing to determine the intent of the utterance using the utterance data. However, Last, in the same field of endeavor, discloses obtaining, by the processor, context information related to the utterance in response to failing to determine the intent of the utterance using the utterance data (Last para 0030, According to a third aspect, there is provided a computer-implemented method for determining a user intent from a speech input to effect a user intended action, the method comprising:… attempting to interpret, by each processing node, the same speech input based on the subset of words directly relevant to its associated context, to extract therefrom an output indicative of user intent, whereby each processing node is unable to interpret any portion of the same speech input containing a word outside of its subset of words, whereby a portion of the same speech input relating to the particular context of a first of the nodes is interpretable to the first processing node but is not interpretable to a second of the processing nodes; OR, Last, para 0031, According to a fourth aspect, there is provided herein a processing node of a computing system for determining a user intent from a speech input to effect a user intended action, wherein the processing node is capable of understanding only a subset of words directly relevant to a particular context associated with the processing node, wherein the processing node is configured to attempt to interpret the speech input based on the subset of words directly relevant to the associated context, to extract therefrom an output indicative of user intent, whereby each processing node is unable to interpret any portion of the same speech input containing a word outside of its subset of words … ; Last, para 0007-0008, If the node receives a speech input that does not relate to its narrow context, the node's intent recognition will fail. To build a system that is capable of handling a wider range of contexts, rather than attempting to expand the capabilities of an individual node, multiple such nodes, targeted on different narrow contexts, are deployed in parallel. When a speech input is received, the speech input is provided to each node for processing in parallel with each other node. Each node attempts to interpret the voice input and extract a user intent therefrom; in a typical scenario, only one node will be able to do so; [i.e., “context information related to the utterance in response to failing to determine the intent of the utterance using the utterance data” is determined using specific node intent recognition system]). Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Last in the method of Ullrich because this would enable implementation of different forms of intent recognition with various contexts, one such context being the currently more prevalent online ‘chatbots’ which typically have user interfaces with a chatbot through text input typed on a physical or virtual keyboard or provided via some other text input mechanism at their device, unlike conventional chatbots which are based on text rather than voice input (Last, para 0002). Regarding Claim 16, Ullrich discloses the non-transitory computer-readable recording medium of claim 13, wherein the function inventory includes at least one in-vehicle function accessible through a vehicle voice recognition system (Ullrich, para 0070, Figure 3, the context subsystem 300 may generate a context prompt 138 token… Such a context prompt 138 may be generated by utilizing sensors… Such sensors may include … microphones configured to feed audio to a speech to text (STT) device; Ullrich, para 212, The GenAI 118 … may provide multimodal output of the real-time feedback. In one embodiment, the real-time feedback may instruct the vehicle's autonomous control system to adjust speed, route, and other factors to improve passenger safety and comfort. The output of the GenAI 118 may be converted into multimodal sensations through vehicle instruments such as the steering wheel, driver heads up display and/or over the in-vehicle audio system). Regarding Claim 17, Ullrich discloses the non-transitory computer-readable recording medium of claim 13, wherein the guided learning examples include example utterance data (Ullrich, para 0225, the method includes receiving context data related to the user's surroundings at block 1806. For example, the context subsystem 300 illustrated in FIG. 3 may receive context data related to the user's surroundings. Context data may include sensor data 110 such as audio and video data capturing body language and spoken words from a user's conversation partner; OR, Ullrich, para 0091-0092, The user 102 may respond to the multimodal output 146 in a manner detectable through biosignals 106, and thus a channel may be provided to train the GenAI 118 based on user 102 response to multimodal output 146. In general, the user agency and capability augmentation system 100 may be viewed as a kind of application framework that uses the biosignals prompt 136, context prompt 138, and user input prompt 140 sequences to facilitate interaction with an application, much as a user 102 would use their finger to interact with a mobile phone application running on a mobile phone operating system. Unlike a touchscreen or mouse/keyboard interface, this system incorporates real time user inputs along with an articulated description of their physical context and historical context to facilitate extremely efficient interactions to enable user agency. FIG. 1 shows the pathways signals take from input, by sensing devices, stored data, or the user 102, to output in the form of text-to-speech utterances 124, written text 126; [“written text ” as “a sentence”]), an example context-aware sentence, and an example process of reasoning the example context-aware sentence from the example utterance data (Ullrich, para 0086-0087, Use all contextual information and prior conversation history to modulate your responses. After each input, review the prior inputs and modify your subsequent predictions based on the context of the thread. Taking into account the current context, with spartan language, return a JSON string called ‘suggestions’ with three different and unique phrases without quotes. They should be complete sentences longer than two words. Do not include explanations. The phrases you respond with will be spoken by my speech generating device. The GenAI 118 may take in the prompt 144 from the prompt composer 116 and use this to generate a multimodal output 146. The GenAI 118 may consist of a pre-trained machine learning model, such as GPT. The GenAI 118 may generate a multimodal output 146 in the form of a token sequence that may be converted back into plaintext, or which may be consumed by a user agency process directly as a token sequence). Regarding Claim 18, Ullrich discloses the non-transitory computer-readable recording medium of claim 13, wherein the guided learning examples include example utterance data and an example context-aware sentence (Ullrich, para 0063, FIG. 1 illustrates a user agency and capability augmentation system 100 in accordance with one embodiment. The user agency and capability augmentation system 100 comprises a user 102, a wearable computing and biosignal sensing device 104, biosignals 106, background material 108, sensor data 110, other device data 112, application context 114 a prompt composer 116, a GenAI 118, a multimodal output stage 120, an encoder/parser 132, output modalities 122 such as an utterance 124, a written text 126, a multimodal artifact 128, an other user agency 130, and a non-language user agency device 134, a biosignals subsystem 200, and a context subsystem 300; [“written text” as “context-aware sentence”]; OR, Ullrich, para 0091-0092, The user 102 may respond to the multimodal output 146 in a manner detectable through biosignals 106, and thus a channel may be provided to train the GenAI 118 based on user 102 response to multimodal output 146. In general, the user agency and capability augmentation system 100 may be viewed as a kind of application framework that uses the biosignals prompt 136, context prompt 138, and user input prompt 140 sequences to facilitate interaction with an application, much as a user 102 would use their finger to interact with a mobile phone application running on a mobile phone operating system. Unlike a touchscreen or mouse/keyboard interface, this system incorporates real time user inputs along with an articulated description of their physical context and historical context to facilitate extremely efficient interactions to enable user agency. FIG. 1 shows the pathways signals take from input, by sensing devices, stored data, or the user 102, to output in the form of text- to-speech utterances 124, written text 126; [“written text ” as “a sentence”]). Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to MULUGETA T. DUGDA whose telephone number is (703)756-1106. The examiner can normally be reached Mon - Fri, 4:30am - 7:00pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Paras D. Shah can be reached at 571-270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MULUGETA TUJI DUGDA/Examiner, Art Unit 2653 /Paras D Shah/Supervisory Patent Examiner, Art Unit 2653 09/19/2026
Read full office action

Prosecution Timeline

Aug 02, 2024
Application Filed
Mar 20, 2026
Non-Final Rejection mailed — §103
Jun 22, 2026
Response Filed
Sep 22, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12725624
SYSTEMS AND METHODS FOR NOISE SUPPRESSION
2y 9m to grant Granted Sep 01, 2026
Patent 12717830
Compressing Information Provided to a Machine-Trained Generative Model
3y 1m to grant Granted Aug 25, 2026
Patent 12700419
SOUND SOURCE SEPARATION METHOD, SOUND SOURCE SEPARATION APPARATUS, AND PROGARM
3y 0m to grant Granted Aug 04, 2026
Patent 12694338
TECHNIQUES FOR TRAINING AND DEPLOYING A NAMED ENTITY RECOGNITION MODEL
3y 2m to grant Granted Jul 28, 2026
Patent 12670918
VOICE MODIFICATION
2y 5m to grant Granted Jun 30, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
83%
Grant Probability
99%
With Interview (+21.2%)
2y 11m (~9m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 58 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month