DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
In response to the Advisory Action from 6/11/2026, Applicant has filed a Request for Continued Examination (RCE) on 6/25/2026. In this reply, Applicant has amended independent claims 1, 10, and 19 to further recite that the LLM instruction requests the LLM to infer whether the spoken words are directed towards the automated agent or constitute ambient conversation not intended for the automated agent and that the structured output includes an inference result indicating whether the spoken words are directed towards the automated agent or constitute ambient conversation and when the inference result indicates that the spoken words are directed towards the automated agent, identifying the intent of the user. Claims 3 and 12 have been cancelled while new claims 21-22 have been added.
Applicant also argues that the prior art of record fails to teach the limitations added via the instant amendment (Remarks, Pages 10-12).
These arguments have been fully considered, however, are moot with respect to the new grounds of rejection further in view of Wagner, et al. ("Multimodal Data and Resource Efficient Device-directed Speech Detection with Large Foundation Models," 2023).
With the entry of the amendments of claims 5, 7-9, 14, and 16-18, Applicant argues that the rejections under 35 U.S.C. 112(b) should be withdrawn (Remarks, Pages 9-10).
In response to these amendments also included in the RCE that correct antecedent basis and remove indefinite "such as" claim language, the 35 U.S.C. 112(b) rejections are now moot and have been withdrawn.
Response to Arguments
Applicant arguments directed towards the teachings of Beauml, et al (U.S. PG Publication: 2023/0074406 A1) (see Remarks, Pages 10-12) are moot in favor of the new grounds of rejection in view of Wagner, et al. ("Multimodal Data and Resource Efficient Device-directed Speech Detection with Large Foundation Models," 2023) that more closely addressed Applicant concerns by teaching LLM instructions to "generate a decision on whether an unseen utterance is directed towards a device or not" (see Wagner, Section 3, Page 2).
Applicant has also commented on the remarks provided in 6/11/2026 Advisory Action by submitting more narrow claim construction that the decision of whether the spoken words are directed towards to automated agent or constitute ambient conversation not intended for the ambient agent. Particularly, Applicant contends that the "or" in the limitation is functional because it "lives inside" the configuration of a single instruction that asks the LLM to perform a binary classification decision (Remarks, Pages 12-13).
The examiner respectfully disagrees with applicants claim construction under the broadest reasonable interpretation (BRI). The prompt as claimed includes an instruction to request the LLM to make an interference that is distributed to either the inference of spoken words being directed towards the automated agent or ambient conversation. Prior art, such as Baeuml (as discussed in the Advisory Action), teaching the request of the LLM to infer whether the spoken words are directed towards the automated agent addresses the claim limitation set forth in the alternative. That analysis being taken into consideration, Applicant arguments/statements do effectively create an estoppel on the record narrowing the claim scope that the claim should be construed in the manner indicated (i.e., that the LLM is instructed to analyze both alternatives to make a binary classification decision) and ultimately this narrowed claim scope is addressed via the teachings of Wagner rendering Applicant arguments moot.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 6, 9-10, 15, 18-19, and 21-22 are rejected under 35 U.S.C. 103 as being unpatentable over Beauml, et al (U.S. PG Publication: 2023/0074406 A1) in view of Wagner, et al. ("Multimodal Data and Resource Efficient Device-directed Speech Detection with Large Foundation Models," 2023).
With respect to Claim 1, Baeuml discloses:
A system for processing spoken language to determine user intent for interaction with an automated agent, the system comprising:
at least one processor (“one or more processors,” Paragraph 0040);
at least one memory component storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising (“one or more memories,” Paragraph 0040):
continuously monitoring ambient audio via a microphone integrated with a device ("monitor a stream of audio data generated by one or more microphones of the client device," Paragraph 0068; see also continuous monitoring at the loop in Element 362 of Fig. 3);
converting captured spoken words from the ambient audio into text using a speech-to-text conversion process ("system can process the stream of audio data (e.g., the stream of audio data 201 of FIG. 2) using the ASR engine" to "produce recognized text that corresponds to the spoken utterance" and "monitor for one or more particular words or phrases included in the stream of audio data" Paragraphs 0045 and 0068-0069);
generating a structured prompt that includes at least the converted text and an instruction for a Large Language Model (LLM), wherein the instruction is configured to request the LLM to infer whether the spoken words are directed towards the automated agent (recognized text is processed by an LLM along with additional data (e.g., the context of the dialog) to constitute a “structured prompt” to determine whether the transcription relates to an “automated assistant”, Paragraphs 0053, 0056 (describing LLM prompting in the form of "textual input corresponding to the assistant query"), 0059, 0060, 0063, and 0089);
transmitting the structured prompt to the LLM (transmission of the LLM input to an LLM remotely located from the client device, Paragraphs 0022, 0030, and 0059);
receiving, from the LLM, a structured output in a standardized format, wherein the structured output includes an inference result indicating whether the spoken words are directed towards the automated agent ("LLMs can determine an intent associated with the given assistant query ( e.g., based on the stream of NLU output 204 generated using the NLU engine)" wherein intents detected pertain to voice assistant queries, LLM inference occurs after the target words that are being monitored are detected, Paragraphs 0053, 0059 (discussing example structuring of the LLM output), 0060-0061, 0073, and 0100); and
executing an action by the automated agent based on the identified intent of the user when the inference result indicates that the spoken words are directed towards for the automated agent (“automated assistant” is caused to take an action such as generating an output based upon the LLM processing including intent identification, Paragraphs 0001, 0035, 0060, 0073, and 0088).
Baeuml does not teach an LLM instruction that requests the LLM to infer whether the spoken words are directed towards the automated agent or constitute ambient conversation not intended for the automated agent as "binary classification" including the consideration of both classifications in rendering a structured output (see 6/25/2025 Remarks, Pages 12-13). Wagner, however, discloses an LLM that receives instructions to process a prompt (e.g., "directed decision" prompt along with data included to be assessed) "to generate decision about device directedness" that is binary and determines "whether an unseen utterance is directed towards a device or not" (Abstract; Section 1, Pages 1-2; Section 2, Pages 2-3; Fig. 1).
Baeuml and Wagner are analogous art because they are from a similar field of endeavor in voice assistants utilizing LLMs. Thus, it would have been obvious to one of ordinary skill before the effective filing date to utilize the LLM device directedness decision taught by Wagner in the virtual assistant interaction taught by Baeuml to provide a predictable result in the form of making virtual assistant interactions more natural by obviating the need for a trigger phrase (Wagner, Abstract).
With respect to Claim 6, Baeuml further discloses:
The system of claim 1, wherein the automated agent is integrated into an augmented reality (AR) device (client device in the form of a "augmented reality computing device," Paragraph 0031), and the action executed by the automated agent includes displaying relevant information via an application executing within an AR environment of the AR device ("assistant outputs 207...to be visually rendered by a display of the client device," Paragraph 0051).
With respect to Claim 9, Baeuml further discloses:
The system of claim 1, wherein the automated agent is integrated into an automobile’s infotainment system, and the action executed by the automated agent includes receiving spoken commands related to vehicle control functions, wherein the LLM is fine-tuned to recognize and process commands specific to automotive operations (client device is part of an in-vehicle system such that inputs causing actions are related to operation of an "in-vehicle navigation system" or "in-vehicle entertainment system," Paragraphs 0031, 0035, and 0066).
Claim 10 recites a corresponding to the functionality performed by the system of claim 1, and thus, is rejected under similar rationale.
Claim 15 contains subject matter similar to Claim 6, and thus, is rejected under similar rationale.
Claim 18 contains subject matter similar to Claim 9, and thus, is rejected under similar rationale.
Claim 19 is directed towards an embodiment of the functionality performed by the system of claim 1 realized as a non-transitory computer-readable storage medium storing processor executable instructions for carrying out that functionality, and thus is rejected under similar rationale. Moreover, claim 19 includes some narrower subject matter related to "domain-specific" automated agents and the LLM being fine-tuned for prompts related to a domain of the domain-specific automated agent. These limitations are addressed by Baeuml (see- automated assistant with trained/re-trained LLMs pertaining to different domains such as restaurant suggestion or weather, Paragraphs 0060, 0066, and 0100). Lastly, Baeuml teaches method implementation as a non-transitory computer-readable storage medium storing program instruction (Paragraph 0136).
With respect to Claim 21, the combination of Baeuml and Wagner disclose:
The system of claim 1, wherein the instruction is a first instruction, the structured prompt is a first structured prompt, the inference result is a first inference result, and the operations (Wagner- initial operation prior to a voice assistant acting on a user utterance or command where an LLM that receives instructions to process a prompt (e.g., "directed decision" prompt along with data included to be assessed) "to generate decision about device directedness" that is binary and determines "whether an unseen utterance is directed towards a device or not," Abstract; Section 1, Pages 1-2; Section 2, Pages 2-3; Fig. 1, prior to triggering the voice assistant to act upon the device microphone input) further comprise:
in response to the first inference result indicating that the spoken words are directed towards the automated agent, generating a second structured prompt that includes at least the converted text and a second instruction for the LLM (see the triggering of the virtual agent without a trigger phrase as discussed above in regards to Wagner), the second instruction being configured to request the LLM to identify a specific intent of the user from the converted text (Baeuml teaches the "automated" assistant (see Abstract) prompting an LLM using different sets of parameters for a domain (e.g., restaurants) and the NLU result for "contextual scenarios" pulling from different data structures to identify related intents (e.g., user profiles/preferences), Paragraphs 0060-0063, 0068-0069, 0073, and 0100);
transmitting the second structured prompt to the LLM (transmission of the LLM input to an LLM remotely located from the client device, Paragraphs 0022, 0030, and 0059); and
receiving, from the LLM, a second structured output that identifies the specific intent of the user, wherein the executing of the action is further based on the specific intent of the user identified in the second structured output (Baeuml discloses that the LLM performs intent determination and that the “automated assistant” is caused to take an action such as generating an output based upon the LLM processing including intent identification, Paragraphs 0001, 0035, 0060-0061, 0073, and 0088).
Thus, the combination of Baeuml in view of Wagner teaches yields multi-decision processing for LLMs- the first of which is the binary device-directedness decision taught by Wagner that triggers the intent determination sequence of processing by Baeuml that yields more natural interactions with a voice assistant by obviating the need for a trigger phrase (Wagner, Abstract).
Claim 22 contains subject matter similar to Claim 21, and thus, is rejected under similar rationale.
Claims 2, 11, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Baeuml, et al. in view of Wagner, et al. and further in view of Bose, et al. (U.S. PG Publication: 2024/0379096 A1).
With respect to Claim 2, Baeuml in view of Wagner discloses the system for monitoring audio for specific words for prompting an LLM to generate a voice assistant reply as applied to Claim 1. Baeuml also discloses that the automated agent is a domain-specific automated agent (automated assistant pertaining to different domains such as restaurant suggestion or weather, Paragraphs 0060 and 0100). Baeuml in view of Wagner, however, does not teach LLM prompting related to multi/few shot training as set forth in claim 2. Bose, however, discloses:
providing a system prompt distinct from the structured prompt as input to the LLM, the system prompt providing multi-shot fine-tuning examples to the LLM for a domain of the domain-specific automated agent (LLM is for a voice user assistant, Paragraph 0059; where domain specific utterance examples are provided by Baeuml as noted above), each example comprising a sample of spoken words and a corresponding structured output that indicates either a specific intent or absence of intent ("few-shot examples are pairs of an utterance with its corresponding intent derived from known data" to fine tune the LLM, Paragraphs 0015-0016, 0020, 0022, 0027-0028, and 0035 (discussing a task domain)).
Baeuml, Wagner, and Bose are analogous art because they are from a similar field of endeavor in voice assistants using LLMs. Thus, it would have been obvious to one of ordinary skill in the art, before the effective filing date, to utilize the few-shot intent examples taught by Bose in the prompting of Baeuml in view of Wagner to provide a predictable result of dynamically guiding the model at inference to generalize on unseen data using a few labeled examples (Bose, Paragraph 0016).
Claims 11 and 20 contain subject matter similar to Claim 2, and thus, are rejected under similar rationale.
Claims 4 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Baeuml, et al. in view of Wagner, et al. and further in view of Walters, et al. (U.S. PG Publication: 2025/0232766 A1).
With respect to Claim 4, Baeuml in view of Wagner discloses the system for monitoring audio for specific words for prompting an LLM to generate a voice assistant reply as applied to Claim 1.
Baeuml in view of Wagner fails to teach that the structured output is in JSON format, and the structured output includes a field for the identified intent of the user that is populated when the inference result is positive. Walters, however, discloses the structured output is in JSON format, and the structured output includes a field for the identified intent of the user that is populated when the inference result is positive (LLM generation in a "format...such as JSON” that includes an indication of intent that is positively detected, Paragraphs 0029-0030).
Baeuml, Wagner, and Walters are analogous art because they are from a similar field of endeavor in voice assistants utilizing LLMs. Thus, it would have been obvious to one of ordinary skill before the effective filing date to use the JSON format for LLM outputs as taught by Walters in the LLM output generation including intent taught by Baeuml in view of Wagner to provide a predictable result of implementing an output format that can readily be ingested by other processing modules (Walters, Paragraph 0017).
Claim 13 contain subject matter respectively similar to Claim 4, and thus, are rejected under similar rationale.
Claims 5, 8, 14, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Baeuml, et al. in view of Wagner, et al. and further in view of Pandita, et al. (U.S. PG Publication: 2024/0143289 A1).
With respect to Claim 5, Baeuml in view of Wagner discloses the system for monitoring audio for specific words for prompting an LLM to generate an automated agent reply as applied to Claim 1. Baeuml in view of Wagner does not teach that contextual information is relied upon from previous interactions is used in an instruction for the LLM to correct errors and receive a corrected text from the LLM before executing an action. Pandita, however, disclose (previous context information included with a prompt to an LLM to correct speech-to-text errors where the LM returns the corrected transcription that is corrected prior to executing an action (Paragraphs 0073-0080).
Baeuml, Wagner, and Pandita are analogous art because they are from a similar field of endeavor in speech interfaces using LLMs. Thus, it would have been obvious to one of ordinary skill in the art, before the effective filing date, to utilize the prompting related to speech-to-text correction based upon context information taught by Pandita with the LLM prompts for voice assistant prompting taught by Baeuml in view of Wagner to provide a predictable result of making better predictions by an LLM by resolving speech-to-text errors (Pandita, Paragraph 0079).
With respect to Claim 8, Baeuml (in combination with Pandita along with Wagner to teach the directedness decision as per claim 1) further discloses:
The system of claim 1, wherein the LLM is configured to utilize a function calling capability that ensures the structured output is provided in a standardized format, the function calling capability enabling the LLM to execute predefined functions within the structured prompt that correspond to specific tasks including correction of transcription errors resulting from the conversion of the captured spoken words from the ambient audio into text using the speech-to-text conversion process, and the inference of user intent from the text corresponding with the captured spoken words (the LLM calls various functions to perform specific tasks including the generation of intent based upon transcribed speech, the consideration of context information and other data such as probability information, Paragraphs 0045 and 0059-0060 wherein the added functionality of transcription error correction is taught by Pandita as applied to Claim 5).
Claim 14 contains subject matter similar to Claim 5, and thus, is rejected under similar rationale.
Claim 17 contains subject matter similar to Claim 8, and thus, is rejected under similar rationale.
Claims 7 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Baeuml, et al. in view of Wagner, et al. and further in view of Bonar, et al. (U.S. PG Publication: 2024/0169974 A1).
With respect to Claim 7, Baeuml in view of Wagner discloses the system for monitoring audio for specific words for prompting an LLM to generate a voice assistant reply as applied to Claim 1. Baeuml in view of Wagner does not teach pause detection that identifies the end of a spoken sentence and triggers transmission of the structured prompt to the LLM. Bonar, however, discloses that speech from a user is continually gathered until a "user pauses," which triggers transcription and LLM prompting (Paragraph 0020; Fig. 1, Element 112 (showing audio input transcription in the form of a sentence prior to prompting 116 that is triggered by the pause detection).
Baeuml, Wagner, and Bonar are analogous art because they are from a similar field of endeavor in voice assistants using LLMs. Thus, it would have been obvious to one of ordinary skill in the art, before the effective filing date, to utilize the pause detection taught by Bonar in the LLM prompting taught by Baeuml in view of Wagner to provide a predictable result of better ensuring that all user audio that completes a thought is gathered prior to prompting the LLM.
Claim 16 contains subject matter similar to Claim 7, and thus, is rejected under similar rationale.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
Huang, et al. ("A Study for Improving Device-Directed Speech Detection toward Frictionless Human-Machine Interaction," 2019)- teaches an approach to "device-directed utterance detection" that aims to distinguish voice queries intended for a smart-home device from background speech using a machine learning model (Abstract).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JAMES S WOZNIAK whose telephone number is (571)272-7632. The examiner can normally be reached 7-3, off alternate Fridays.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant may use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Flanders can be reached at (571)272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
JAMES S. WOZNIAK
Primary Examiner
Art Unit 2655
/JAMES S WOZNIAK/Primary Examiner, Art Unit 2655