Prosecution Insights
Last updated: August 17, 2026
Application No. 18/981,291

SPEECH CONTROL METHOD, APPARATUS, AND DEVICE

Non-Final OA §101§103
Filed
Dec 13, 2024
Priority
Jun 13, 2022 — CN 202210667673.7 +1 more
Examiner
BOGGS JR., JAMES
Art Unit
Tech Center
Assignee
Huawei Technologies Co., Ltd.
OA Round
1 (Non-Final)
63%
Grant Probability
Moderate
1-2
OA Rounds
1y 6m
Est. Remaining
97%
With Interview

Examiner Intelligence

Grants 63% of resolved cases
63%
Career Allowance Rate
75 granted / 119 resolved
+3.0% vs TC avg
Strong +34% interview lift
Without
With
+34.0%
Interview Lift
resolved cases with interview
Typical timeline
3y 2m
Avg Prosecution
28 currently pending
Career history
142
Total Applications
across all art units

Statute-Specific Performance

§101
11.8%
-28.2% vs TC avg
§103
50.4%
+10.4% vs TC avg
§102
15.7%
-24.3% vs TC avg
§112
18.5%
-21.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 119 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1, 10 and 16 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Claim 1 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim recites a speech control method, comprising: obtaining a speech stream comprising at least two intents; determining, based on the speech stream, an instruction corresponding to each intent included in a first intent, wherein the first intent comprises a part of each of the at least two intents; and executing, in a first sequence, the instruction corresponding to each intent in the first intent and determining an instruction corresponding to each intent in a second intent, wherein the first sequence is a delivery sequence of the at least two intents in the speech stream, and the second intent comprises at least one intent that includes at least two intents and that is arranged after the first intent. The claim 1 limitations, under their broadest reasonable interpretation, cover performance of the limitations in the mind but for the recitation of generic computer components. For example, “obtaining a speech stream” in the context of this claim encompasses a person listening to audio of someone speaking, “determining, based on the speech stream, an instruction corresponding to each intent” in the context of this claim encompasses a person determining requests of the speaker, “executing, in a first sequence, the instruction corresponding to each intent” in the context of this claim encompasses a person responding to requests of the speaker, and “determining an instruction corresponding to each intent in a second intent” in the context of this claim encompasses a person determining additional requests of the speaker. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea. This judicial exception is not integrated into a practical application. The claim does not include any additional elements that integrate the abstract idea into a practical application, and therefore does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claim is not patent eligible. Claim 10 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim recites a speech control method, comprising: obtaining a speech stream comprising at least two intents; sending the speech stream to a server; receiving, from the server, an instruction corresponding to each intent in a first intent, wherein the first intent comprises a part of each of the at least two intents; executing, in a first sequence, the instruction corresponding to each intent in the first intent, wherein the first sequence is a delivery sequence of the at least two intents in the speech stream; and when determining that instructions corresponding to all of the at least two intents are not completely obtained, obtaining, from the server, an instruction corresponding to each intent in a second intent, wherein the second intent comprises at least one intent that includes at least two intents and that is arranged after the first intent; and executing, in the first sequence, the instruction corresponding to each intent in the second intent, until an instruction corresponding to each intent in the at least two intents is executed. The claim 10 limitations, under their broadest reasonable interpretation, cover performance of the limitations in the mind but for the recitation of generic computer components. That is, other than reciting “a server”, nothing in the claim elements preclude the actions from practically being performed in the mind. For example, “obtaining a speech stream” in the context of this claim encompasses a person listening to audio of someone speaking, “receiving an instruction corresponding to each intent” in the context of this claim encompasses a person determining requests of the speaker, “executing, in a first sequence, the instruction corresponding to each intent in the first intent” in the context of this claim encompasses a person responding to requests of the speaker, “obtaining an instruction corresponding to each intent in a second intent” in the context of this claim encompasses a person determining additional requests of the speaker, and “executing, in the first sequence, the instruction corresponding to each intent in the second intent” in the context of this claim encompasses a person responding to the additional requests until all of the requests have been answered. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea. This judicial exception is not integrated into a practical application. In particular, the claim only recites the additional element “a server”. The additional element amounts to no more than mere instructions to apply the exception using generic computer components. Examples of generic computer components can be found in paragraph 0049 of the specification, “According to an eighth aspect, a server includes a memory and a processor. The memory is configured to store program instructions, and the processor is configured to invoke the program instructions in the memory, so that the server performs the speech control method in any one of the third aspect and the possible designs of the third aspect.”. Accordingly, the additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional element amounts to no more than mere instructions to apply the exception using generic computer components. Mere instructions to apply an exception using generic computer components cannot provide an inventive concept. The claim is not patent eligible. Claim 16 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim recites a speech control method, comprising: receiving, from a terminal device, a speech stream comprising at least two intents; determining an instruction corresponding to each intent in a first intent, wherein the first intent comprises a part of each of the at least two intents; sending the instruction corresponding to each intent in the first intent to the terminal device; determining an instruction corresponding to each intent in a second intent, wherein the second intent comprises at least one intent arranged after the first intent in the at least two intents in a first sequence, the first sequence being a delivery sequence of the at least two intents in the speech stream; and sending the instruction corresponding to each intent in the second intent to the terminal device. The claim 16 limitations, under their broadest reasonable interpretation, cover performance of the limitations in the mind but for the recitation of generic computer components. That is, other than reciting “a terminal device”, nothing in the claim elements preclude the actions from practically being performed in the mind. For example, “receiving a speech stream” in the context of this claim encompasses a person listening to audio of someone speaking, “determining an instruction corresponding to each intent” in the context of this claim encompasses a person determining requests of the speaker, and “determining an instruction corresponding to each intent in a second intent” in the context of this claim encompasses a person determining additional requests of the speaker. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea. This judicial exception is not integrated into a practical application. In particular, the claim only recites the additional element “a terminal device”. The additional element amounts to no more than mere instructions to apply the exception using generic computer components. Examples of generic computer components can be found in paragraph 0046 of the specification, “According to a fifth aspect, a terminal device includes a memory and a processor. The memory is configured to store program instructions, and the processor is configured to invoke the program instructions in the memory, so that the terminal device performs the speech control method in any one of the first aspect and the possible designs of the first aspect.”. Accordingly, the additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional element amounts to no more than mere instructions to apply the exception using generic computer components. Mere instructions to apply an exception using generic computer components cannot provide an inventive concept. The claim is not patent eligible. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1 – 6, 10 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Gruber et al. (US Patent No. 9,966,065), hereinafter Gruber, in view of Teserra et al. (US Patent No. 11,195,532), hereinafter Teserra. Regarding claim 1, Gruber discloses a speech control method, comprising: obtaining a speech stream comprising at least two intents (Column 1, line 58 - Column 2, line 4, "Systems and processes are disclosed for processing a multi-part voice command. In one example, speech input can be received from a user that includes a single utterance having one or more actionable commands. A text string can be generated based on the speech input using a speech transcription process. The text string can be parsed into multiple candidate substrings. Probabilities can be determined for each of the candidate substrings indicating whether they are likely to correspond to actionable commands. In response to the probabilities exceeding a threshold, user intents can be determined for each of the candidate substrings. Processes associated with the user intents can then be executed. An acknowledgment can also be provided to the user associated with the various user intents."; A multi-part voice command that includes a single utterance having one or more actionable commands reads on a speech stream comprising at least two intents.); determining, based on the speech stream, an instruction corresponding to each intent included in a first intent, wherein the first intent comprises a part of each of the at least two intents (Column 1, line 58 - Column 2, line 4, "Systems and processes are disclosed for processing a multi-part voice command. In one example, speech input can be received from a user that includes a single utterance having one or more actionable commands. A text string can be generated based on the speech input using a speech transcription process. The text string can be parsed into multiple candidate substrings. Probabilities can be determined for each of the candidate substrings indicating whether they are likely to correspond to actionable commands. In response to the probabilities exceeding a threshold, user intents can be determined for each of the candidate substrings. Processes associated with the user intents can then be executed. An acknowledgment can also be provided to the user associated with the various user intents."; Executing processes associated with the user intents reads on determining an instruction corresponding to each intent included in a first intent.); and executing, in a first sequence, the instruction corresponding to each intent in the first intent (Column 50, lines 36-54, "Referring again to process 800 of FIG. 8, at block 814, a first process associated with the first intent and a second process associated with the second intent can be executed. With user intents determined for each candidate substring (or some candidate substrings), the processes associated with the user intents can be executed. For example, messages can be composed and sent, emails can be deleted, notifications can be dismissed, or the like. In some examples, multiple tasks or processes can be associated with individual user intents, and the various tasks or processes can be executed at block 814. In other examples, the virtual assistant can engage the user in a dialogue to acquire additional information as necessary for completing task flows. As mentioned above, in some examples, if only a subset of the substrings of a multi-part command can be interpreted into user intents, the processes associated with those user intents can be executed, and the virtual assistant can handle the remaining substrings in a variety of ways (e.g., request more information, return an error, etc.)."; Executing a first process associated with the first intent and a second process associated with the second intent reads on executing, in a first sequence, the instruction corresponding to each intent in the first intent.); and determining an instruction corresponding to each intent in a second intent, [wherein the first sequence is a delivery sequence of the at least two intents in the speech stream,] and the second intent comprises at least one intent that includes at least two intents and that is arranged after the first intent (Column 1, line 58 - Column 2, line 4, "Systems and processes are disclosed for processing a multi-part voice command. In one example, speech input can be received from a user that includes a single utterance having one or more actionable commands. A text string can be generated based on the speech input using a speech transcription process. The text string can be parsed into multiple candidate substrings. Probabilities can be determined for each of the candidate substrings indicating whether they are likely to correspond to actionable commands. In response to the probabilities exceeding a threshold, user intents can be determined for each of the candidate substrings. Processes associated with the user intents can then be executed. An acknowledgment can also be provided to the user associated with the various user intents."; Executing processes associated with the user intents reads on determining an instruction corresponding to each intent in a second intent.). Gruber does not specifically disclose: wherein the first sequence is a delivery sequence of the at least two intents in the speech stream. Teserra teaches: wherein the first sequence is a delivery sequence of the at least two intents in the speech stream (Column 21, lines 45-61, "Precedence analyzer 430 analyzes each utterance constructed by the splitter 420 to determine an order in which to handle the utterances, e.g., an order in which the utterance are to be provided to an intent classifier. The order can be determined based on the structure of the utterances and the presence of certain words that signal precedence. For instance, as indicated in the example list of conjunctions above, the conjunctive words “before”, “after”, “and then”, and “only if” are some examples of words that can indicate order. The precedence analyzer 430 can take into consideration the extracted information 404 when deciding the order. For instance, the extracted information 404 can indicate important parts of the utterance 402 such as verbs and predicates (e.g. noun phrases) associated with the verbs. The precedence analyzer 430 places the utterances constructed by the splitter 420 into the queue 440 in the determined order."; Determining an order in which to handle the utterances reads on a delivery sequence of the at least two intents in the speech stream.). Teserra is considered to be analogous to the claimed invention because it is in the same field of voice recognition systems. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Gruber to incorporate the teachings of Teserra to determine an order in which to handle utterances. Doing so would allow for detecting that there are multiple intents represented in an utterance and matching each detected intent to an intent associated with a chatbot in a chatbot system (Teserra; Column 2, lines 3-7). Regarding claim 2, Gruber in view of Teserra discloses the method as claimed in claim 1. Gruber further discloses: wherein a terminal device comprises a first application module and a third application module (Column 5, lines 8-12, "One or more processing modules 114 can utilize data and models 116 to process speech input and determine the user's intent based on natural language input. Further, one or more processing modules 114 perform task execution based on inferred user intent."); and the method comprises: obtaining, by the first application module, the speech stream (Column 15, lines 13-16, "For example, digital assistant client module 229 can be capable of accepting voice input (e.g., speech input), text input, touch input, and/or gestural input through various user interfaces"; Column 33, lines 1-5, "User interface module 722 can receive commands and/or inputs from a user via I/O interface 706 (e.g., from a keyboard, touch screen, pointing device, controller, and/or microphone), and generate user interface objects on a display."; A digital assistant client module accepting voice input reads on the first application module obtaining the speech stream.); determining, by the third application module based on the at least two intents corresponding to the speech stream, the instruction corresponding to each intent in the first intent, and continuing to determine the instruction corresponding to each intent in the second intent (Column 5, lines 8-10, "One or more processing modules 114 can utilize data and models 116 to process speech input and determine the user's intent based on natural language input."; Column 35, lines 42-57, "Natural language processing module 732 (“natural language processor”) of the digital assistant can take the sequence of words or tokens (“token sequence”) generated by STT processing module 730, and attempt to associate the token sequence with one or more “actionable intents” recognized by the digital assistant. An “actionable intent” can represent a task that can be performed by the digital assistant, and can have an associated task flow implemented in task flow models 754. The associated task flow can be a series of programmed actions and steps that the digital assistant takes in order to perform the task. The scope of a digital assistant's capabilities can be dependent on the number and variety of task flows that have been implemented and stored in task flow models 754, or in other words, on the number and variety of “actionable intents” that the digital assistant recognizes."; A processing module determining actionable intents that have associated task flows reads on the third application module determining the instruction corresponding to each intent in the first intent and each intent in the second intent.); and obtaining, by the first application module from the third application module, the instruction corresponding to each intent in the first intent and executing, in the first sequence, the instruction corresponding to each intent in the first intent (Column 35, lines 42-57, "Natural language processing module 732 (“natural language processor”) of the digital assistant can take the sequence of words or tokens (“token sequence”) generated by STT processing module 730, and attempt to associate the token sequence with one or more “actionable intents” recognized by the digital assistant. An “actionable intent” can represent a task that can be performed by the digital assistant, and can have an associated task flow implemented in task flow models 754. The associated task flow can be a series of programmed actions and steps that the digital assistant takes in order to perform the task. The scope of a digital assistant's capabilities can be dependent on the number and variety of task flows that have been implemented and stored in task flow models 754, or in other words, on the number and variety of “actionable intents” that the digital assistant recognizes."; Column 38, line 64 - Column 39, line 9, "In some examples, natural language processing module 732 can pass the generated structured query (including any completed parameters) to task flow processing module 736 (“task flow processor”). Task flow processing module 736 can be configured to receive the structured query from natural language processing module 732, complete the structured query, if necessary, and perform the actions required to “complete” the user's ultimate request. In some examples, the various procedures necessary to complete these tasks can be provided in task flow models 754. In some examples, task flow models 754 can include procedures for obtaining additional information from the user and task flows for performing actions associated with the actionable intent."; The natural language processing module passing the generated structured query to the task flow processing module reads on obtaining, by the first application module from the third application module, the instruction corresponding to each intent in the first intent, and the task flow processing module performing the actions required to complete the user's request reads on executing the instruction corresponding to each intent in the first intent.), and [when determining that instructions corresponding to all of the at least two intents are not completely obtained,] obtaining, from the third application module, the instruction corresponding to each intent in the second intent until the instructions corresponding to all of the at least two intents are completely obtained (Column 35, lines 42-57, "Natural language processing module 732 (“natural language processor”) of the digital assistant can take the sequence of words or tokens (“token sequence”) generated by STT processing module 730, and attempt to associate the token sequence with one or more “actionable intents” recognized by the digital assistant. An “actionable intent” can represent a task that can be performed by the digital assistant, and can have an associated task flow implemented in task flow models 754. The associated task flow can be a series of programmed actions and steps that the digital assistant takes in order to perform the task. The scope of a digital assistant's capabilities can be dependent on the number and variety of task flows that have been implemented and stored in task flow models 754, or in other words, on the number and variety of “actionable intents” that the digital assistant recognizes."; Column 38, line 64 - Column 39, line 9, "In some examples, natural language processing module 732 can pass the generated structured query (including any completed parameters) to task flow processing module 736 (“task flow processor”). Task flow processing module 736 can be configured to receive the structured query from natural language processing module 732, complete the structured query, if necessary, and perform the actions required to “complete” the user's ultimate request. In some examples, the various procedures necessary to complete these tasks can be provided in task flow models 754. In some examples, task flow models 754 can include procedures for obtaining additional information from the user and task flows for performing actions associated with the actionable intent."; The natural language processing module passing the generated structured query to the task flow processing module reads on obtaining, by the first application module from the third application module, the instruction corresponding to each intent in the second intent until the instructions corresponding to all of the at least two intents are completely obtained.). Teserra further teaches: and when determining that instructions corresponding to all of the at least two intents are not completely obtained, obtaining, from the third application module, the instruction corresponding to each intent in the second intent until the instructions corresponding to all of the at least two intents are completely obtained (Column 30, lines 22-35, “At 712, the utterance and the extracted information are analyzed to determine that the utterance includes two or more parts (e.g., conjuncts), where each part represents a separate intent. The analysis in 712 involves applying one or more detection rules (e.g., whichever ones of Detection Rules 1 to 7 described above are applicable) to match the utterance to a sentence pattern indicative of the presence of multiple intents. As explained above, one feature that a detection rule can look for is a coordinating conjunction. However, the presence of a coordinating conjunction alone may not be sufficient to conclusively determine that there are multiple intents. Accordingly, the detection rules described above are designed to analyze overall sentence structure and capture the most prevalent cases of multiple intents.”; Column 30, line 63 - Column 31, line 5, “In 716, a determination is made as to whether there exists an order of precedence for the two or more parts identified in 712. For example, as indicated above, precedence can be determined through detection of independent or dependent markers. If there is an order of precedence, then the utterances constructed in 714 are processed in the same order as their corresponding parts. Otherwise, the utterances constructed in 714 can be processed in the order in which the parts appear in the original utterance received in 710, e.g., from left to right.”; Column 31, lines 14-20, “At 718, the utterances constructed in 714 are sent for further processing one at a time and according to the order determined in 716. For example, the utterances can be placed into a queue based on the order determined in 716 and then de-queued for explicit invocation analysis, where de-queuing occurs upon returning from a conversation associated with a particular bot intent.”; Separating an utterance into parts with different intents, placing the utterances in a queue, and de-queuing the utterance parts for explicit invocation analysis reads on obtaining the instruction corresponding to each intent in the second intent until the instructions corresponding to all of the at least two intents are completely obtained when determining that instructions corresponding to all of the at least two intents are not completely obtained, where determining that utterances are in the queue reads on determining that instructions corresponding to all of the at least two intents are not completely obtained.). Teserra is considered to be analogous to the claimed invention because it is in the same field of voice recognition systems. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Gruber in view of Teserra to further incorporate the teachings of Teserra to separate an utterance into parts with different intents, place the utterances in a queue, and de-queue the utterance parts for explicit invocation analysis. Doing so would allow for detecting that there are multiple intents represented in an utterance and matching each detected intent to an intent associated with a chatbot in a chatbot system (Teserra; Column 2, lines 3-7). Regarding claim 3, Gruber in view of Teserra discloses the method as claimed in claim 2. Gruber further discloses: further comprising: obtaining, by the first application module, the at least two intents based on the speech stream (Column 1, line 58 - Column 2, line 4, "Systems and processes are disclosed for processing a multi-part voice command. In one example, speech input can be received from a user that includes a single utterance having one or more actionable commands. A text string can be generated based on the speech input using a speech transcription process. The text string can be parsed into multiple candidate substrings. Probabilities can be determined for each of the candidate substrings indicating whether they are likely to correspond to actionable commands. In response to the probabilities exceeding a threshold, user intents can be determined for each of the candidate substrings. Processes associated with the user intents can then be executed. An acknowledgment can also be provided to the user associated with the various user intents."; Column 15, lines 13-16, "For example, digital assistant client module 229 can be capable of accepting voice input (e.g., speech input), text input, touch input, and/or gestural input through various user interfaces"; The digital assistant client module accepting voice input that includes a single utterance having one or more actionable commands reads on obtaining, by the first application module, the at least two intents based on the speech stream.); and sending, by the first application module, the at least two intents to the third application module (Column 35, lines 42-57, "Natural language processing module 732 (“natural language processor”) of the digital assistant can take the sequence of words or tokens (“token sequence”) generated by STT processing module 730, and attempt to associate the token sequence with one or more “actionable intents” recognized by the digital assistant. An “actionable intent” can represent a task that can be performed by the digital assistant, and can have an associated task flow implemented in task flow models 754. The associated task flow can be a series of programmed actions and steps that the digital assistant takes in order to perform the task. The scope of a digital assistant's capabilities can be dependent on the number and variety of task flows that have been implemented and stored in task flow models 754, or in other words, on the number and variety of “actionable intents” that the digital assistant recognizes."; The natural language processing module associating the token sequence with one or more actionable intents recognized by the digital assistant reads on sending the at least two intents to the third application module.). Regarding claim 4, Gruber in view of Teserra discloses the method as claimed in claim 3. Gruber further discloses: wherein: the first application module is a voice assistant application and the third application module is a dialogue management module; or the first application module and the third application module are modules in a voice assistant application (Column 4, lines 15-22, "FIG. 1 illustrates a block diagram of system 100 according to various examples. In some examples, system 100 can implement a digital assistant. The terms “digital assistant,” “virtual assistant,” “intelligent automated assistant,” or “automatic digital assistant” can refer to any information processing system that interprets natural language input in spoken and/or textual form to infer user intent, and performs actions based on the inferred user intent."; Column 5, lines 8-12, "One or more processing modules 114 can utilize data and models 116 to process speech input and determine the user's intent based on natural language input. Further, one or more processing modules 114 perform task execution based on inferred user intent."; One or more processing modules in a digital assistant system to process speech input, determine the user's intent based on natural language input, and perform task execution based on inferred user intent reads on the first application module and the third application module being modules in a voice assistant application.). Regarding claim 5, Gruber in view of Teserra discloses the method as claimed in claim 2. Gruber further discloses: wherein the terminal device further comprises a second application module, the method further comprising: sending, by the first application module, the speech stream to the second application module; obtaining, by the second application module, the at least two intents based on the speech stream; sending, by the second application module, the at least two intents to the third application module, wherein the obtaining, by the first application module from the third application module, an instruction corresponding to any intent comprises: obtaining, by the first application module from the third application module through the second application module, the instruction corresponding to any intent (Column 5, lines 8-12, "One or more processing modules 114 can utilize data and models 116 to process speech input and determine the user's intent based on natural language input. Further, one or more processing modules 114 perform task execution based on inferred user intent."; Column 38, lines 30-34, "In some examples, once natural language processing module 732 identifies an actionable intent (or domain) based on the user request, natural language processing module 732 can generate a structured query to represent the identified actionable intent."; Column 38, line 64 - Column 39, line 6, "In some examples, natural language processing module 732 can pass the generated structured query (including any completed parameters) to task flow processing module 736 (“task flow processor”). Task flow processing module 736 can be configured to receive the structured query from natural language processing module 732, complete the structured query, if necessary, and perform the actions required to “complete” the user's ultimate request. In some examples, the various procedures necessary to complete these tasks can be provided in task flow models 754."; One or more processing modules to process speech input, determine the user's intent based on natural language input, and perform task execution based on inferred user intent reads on the first application module, the second application module, and the third application module.). Regarding claim 6, Gruber in view of Teserra discloses the method as claimed in claim 5. Gruber further discloses: wherein: the first application module is a voice assistant application, the second application module is a central control unit, and the third application module is a dialogue management module (Column 5, lines 8-12, "One or more processing modules 114 can utilize data and models 116 to process speech input and determine the user's intent based on natural language input. Further, one or more processing modules 114 perform task execution based on inferred user intent."; A module to process speech input reads on a voice assistant application module, a module to determine the user's intent based on natural language input reads on a dialogue management module, and a module to perform task execution based on inferred user intent reads on a central control unit.). Regarding claim 10, Gruber discloses a speech control method, comprising: obtaining a speech stream comprising at least two intents (Column 1, line 58 - Column 2, line 4, "Systems and processes are disclosed for processing a multi-part voice command. In one example, speech input can be received from a user that includes a single utterance having one or more actionable commands. A text string can be generated based on the speech input using a speech transcription process. The text string can be parsed into multiple candidate substrings. Probabilities can be determined for each of the candidate substrings indicating whether they are likely to correspond to actionable commands. In response to the probabilities exceeding a threshold, user intents can be determined for each of the candidate substrings. Processes associated with the user intents can then be executed. An acknowledgment can also be provided to the user associated with the various user intents."; A multi-part voice command that includes a single utterance having one or more actionable commands reads on a speech stream comprising at least two intents.); sending the speech stream to a server (Column 6, lines 3-6, “For example, DA client 102 of user device 104 can be configured to transmit information (e.g., a user request received at user device 104) to DA server 106 via second user device 122.”; Column 6, lines 29-40, “Although the digital assistant shown in FIG. 1 can include both a client-side portion (e.g., DA client 102) and a server-side portion (e.g., DA server 106), in some examples, the functions of a digital assistant can be implemented as a standalone application installed on a user device. In addition, the divisions of functionalities between the client and server portions of the digital assistant can vary in different implementations. For instance, in some examples, the DA client can be a thin-client that provides only user-facing input and output processing functions, and delegates all other functionalities of the digital assistant to a backend server.”; Transmitting a user request received at user device to a server reads on sending the speech stream to a server.); receiving, from the server, an instruction corresponding to each intent in a first intent, wherein the first intent comprises a part of each of the at least two intents (Column 1, line 58 - Column 2, line 4, "Systems and processes are disclosed for processing a multi-part voice command. In one example, speech input can be received from a user that includes a single utterance having one or more actionable commands. A text string can be generated based on the speech input using a speech transcription process. The text string can be parsed into multiple candidate substrings. Probabilities can be determined for each of the candidate substrings indicating whether they are likely to correspond to actionable commands. In response to the probabilities exceeding a threshold, user intents can be determined for each of the candidate substrings. Processes associated with the user intents can then be executed. An acknowledgment can also be provided to the user associated with the various user intents."; Column 6, lines 29-40, “Although the digital assistant shown in FIG. 1 can include both a client-side portion (e.g., DA client 102) and a server-side portion (e.g., DA server 106), in some examples, the functions of a digital assistant can be implemented as a standalone application installed on a user device. In addition, the divisions of functionalities between the client and server portions of the digital assistant can vary in different implementations. For instance, in some examples, the DA client can be a thin-client that provides only user-facing input and output processing functions, and delegates all other functionalities of the digital assistant to a backend server.”; Executing processes associated with the user intents reads on receiving an instruction corresponding to each intent included in a first intent.); executing, in a first sequence, the instruction corresponding to each intent in the first intent (Column 50, lines 36-54, "Referring again to process 800 of FIG. 8, at block 814, a first process associated with the first intent and a second process associated with the second intent can be executed. With user intents determined for each candidate substring (or some candidate substrings), the processes associated with the user intents can be executed. For example, messages can be composed and sent, emails can be deleted, notifications can be dismissed, or the like. In some examples, multiple tasks or processes can be associated with individual user intents, and the various tasks or processes can be executed at block 814. In other examples, the virtual assistant can engage the user in a dialogue to acquire additional information as necessary for completing task flows. As mentioned above, in some examples, if only a subset of the substrings of a multi-part command can be interpreted into user intents, the processes associated with those user intents can be executed, and the virtual assistant can handle the remaining substrings in a variety of ways (e.g., request more information, return an error, etc.)."; Executing a first process associated with the first intent and a second process associated with the second intent reads on executing, in a first sequence, the instruction corresponding to each intent in the first intent.); and [when determining that instructions corresponding to all of the at least two intents are not completely obtained,] obtaining, from the server, an instruction corresponding to each intent in a second intent, wherein the second intent comprises at least one intent that includes at least two intents and that is arranged after the first intent (Column 35, lines 42-57, "Natural language processing module 732 (“natural language processor”) of the digital assistant can take the sequence of words or tokens (“token sequence”) generated by STT processing module 730, and attempt to associate the token sequence with one or more “actionable intents” recognized by the digital assistant. An “actionable intent” can represent a task that can be performed by the digital assistant, and can have an associated task flow implemented in task flow models 754. The associated task flow can be a series of programmed actions and steps that the digital assistant takes in order to perform the task. The scope of a digital assistant's capabilities can be dependent on the number and variety of task flows that have been implemented and stored in task flow models 754, or in other words, on the number and variety of “actionable intents” that the digital assistant recognizes."; Column 38, line 64 - Column 39, line 9, "In some examples, natural language processing module 732 can pass the generated structured query (including any completed parameters) to task flow processing module 736 (“task flow processor”). Task flow processing module 736 can be configured to receive the structured query from natural language processing module 732, complete the structured query, if necessary, and perform the actions required to “complete” the user's ultimate request. In some examples, the various procedures necessary to complete these tasks can be provided in task flow models 754. In some examples, task flow models 754 can include procedures for obtaining additional information from the user and task flows for performing actions associated with the actionable intent."; The natural language processing module passing the generated structured query to the task flow processing module reads on obtaining an instruction corresponding to each intent in a second intent.); and executing, in the first sequence, the instruction corresponding to each intent in the second intent, until an instruction corresponding to each intent in the at least two intents is executed (Column 50, lines 36-54, "Referring again to process 800 of FIG. 8, at block 814, a first process associated with the first intent and a second process associated with the second intent can be executed. With user intents determined for each candidate substring (or some candidate substrings), the processes associated with the user intents can be executed. For example, messages can be composed and sent, emails can be deleted, notifications can be dismissed, or the like. In some examples, multiple tasks or processes can be associated with individual user intents, and the various tasks or processes can be executed at block 814. In other examples, the virtual assistant can engage the user in a dialogue to acquire additional information as necessary for completing task flows. As mentioned above, in some examples, if only a subset of the substrings of a multi-part command can be interpreted into user intents, the processes associated with those user intents can be executed, and the virtual assistant can handle the remaining substrings in a variety of ways (e.g., request more information, return an error, etc.)."; Executing a first process associated with the first intent and a second process associated with the second intent reads on and executing the instruction corresponding to each intent in the second intent until an instruction corresponding to each intent in the at least two intents is executed.). Gruber does not specifically disclose: wherein the first sequence is a delivery sequence of the at least two intents in the speech stream; and when determining that instructions corresponding to all of the at least two intents are not completely obtained, obtaining, from the server, an instruction corresponding to each intent in a second intent, wherein the second intent comprises at least one intent that includes at least two intents and that is arranged after the first intent. Teserra teaches: wherein the first sequence is a delivery sequence of the at least two intents in the speech stream (Column 21, lines 45-61, "Precedence analyzer 430 analyzes each utterance constructed by the splitter 420 to determine an order in which to handle the utterances, e.g., an order in which the utterance are to be provided to an intent classifier. The order can be determined based on the structure of the utterances and the presence of certain words that signal precedence. For instance, as indicated in the example list of conjunctions above, the conjunctive words “before”, “after”, “and then”, and “only if” are some examples of words that can indicate order. The precedence analyzer 430 can take into consideration the extracted information 404 when deciding the order. For instance, the extracted information 404 can indicate important parts of the utterance 402 such as verbs and predicates (e.g. noun phrases) associated with the verbs. The precedence analyzer 430 places the utterances constructed by the splitter 420 into the queue 440 in the determined order."; Determining an order in which to handle the utterances reads on a delivery sequence of the at least two intents in the speech stream.); and when determining that instructions corresponding to all of the at least two intents are not completely obtained, obtaining, from the server, an instruction corresponding to each intent in a second intent, wherein the second intent comprises at least one intent that includes at least two intents and that is arranged after the first intent (Column 30, lines 22-35, “At 712, the utterance and the extracted information are analyzed to determine that the utterance includes two or more parts (e.g., conjuncts), where each part represents a separate intent. The analysis in 712 involves applying one or more detection rules (e.g., whichever ones of Detection Rules 1 to 7 described above are applicable) to match the utterance to a sentence pattern indicative of the presence of multiple intents. As explained above, one feature that a detection rule can look for is a coordinating conjunction. However, the presence of a coordinating conjunction alone may not be sufficient to conclusively determine that there are multiple intents. Accordingly, the detection rules described above are designed to analyze overall sentence structure and capture the most prevalent cases of multiple intents.”; Column 30, line 63 - Column 31, line 5, “In 716, a determination is made as to whether there exists an order of precedence for the two or more parts identified in 712. For example, as indicated above, precedence can be determined through detection of independent or dependent markers. If there is an order of precedence, then the utterances constructed in 714 are processed in the same order as their corresponding parts. Otherwise, the utterances constructed in 714 can be processed in the order in which the parts appear in the original utterance received in 710, e.g., from left to right.”; Column 31, lines 14-20, “At 718, the utterances constructed in 714 are sent for further processing one at a time and according to the order determined in 716. For example, the utterances can be placed into a queue based on the order determined in 716 and then de-queued for explicit invocation analysis, where de-queuing occurs upon returning from a conversation associated with a particular bot intent.”; Separating an utterance into parts with different intents, placing the utterances in a queue, and de-queuing the utterance parts for explicit invocation analysis reads on obtaining the instruction corresponding to each intent in the second intent when determining that instructions corresponding to all of the at least two intents are not completely obtained, where determining that utterances are in the queue reads on determining that instructions corresponding to all of the at least two intents are not completely obtained.). Teserra is considered to be analogous to the claimed invention because it is in the same field of voice recognition systems. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Gruber to incorporate the teachings of Teserra to determine an order in which to handle utterances and to separate an utterance into parts with different intents, place the utterances in a queue, and de-queue the utterance parts for explicit invocation analysis. Doing so would allow for detecting that there are multiple intents represented in an utterance and matching each detected intent to an intent associated with a chatbot in a chatbot system (Teserra; Column 2, lines 3-7). Regarding claim 16, Gruber discloses a speech control method, comprising: receiving, from a terminal device, a speech stream comprising at least two intents (Column 1, line 58 - Column 2, line 4, "Systems and processes are disclosed for processing a multi-part voice command. In one example, speech input can be received from a user that includes a single utterance having one or more actionable commands. A text string can be generated based on the speech input using a speech transcription process. The text string can be parsed into multiple candidate substrings. Probabilities can be determined for each of the candidate substrings indicating whether they are likely to correspond to actionable commands. In response to the probabilities exceeding a threshold, user intents can be determined for each of the candidate substrings. Processes associated with the user intents can then be executed. An acknowledgment can also be provided to the user associated with the various user intents."; Column 4, line 58 - Column 5, line 2, “As shown in FIG. 1, in some examples, a digital assistant can be implemented according to a client-server model. The digital assistant can include client-side portion 102 (hereafter “DA client 102”) executed on user device 104 and server-side portion 106 (hereafter “DA server 106”) executed on server system 108. DA client 102 can communicate with DA server 106 through one or more networks 110. DA client 102 can provide client-side functionalities such as user-facing input and output processing and communication with DA server 106. DA server 106 can provide server-side functionalities for any number of DA clients 102 each residing on a respective user device 104.”; A multi-part voice command that includes a single utterance having one or more actionable commands reads on a speech stream comprising at least two intents, and a user device reads on a terminal device.); determining an instruction corresponding to each intent in a first intent, wherein the first intent comprises a part of each of the at least two intents (Column 1, line 58 - Column 2, line 4, "Systems and processes are disclosed for processing a multi-part voice command. In one example, speech input can be received from a user that includes a single utterance having one or more actionable commands. A text string can be generated based on the speech input using a speech transcription process. The text string can be parsed into multiple candidate substrings. Probabilities can be determined for each of the candidate substrings indicating whether they are likely to correspond to actionable commands. In response to the probabilities exceeding a threshold, user intents can be determined for each of the candidate substrings. Processes associated with the user intents can then be executed. An acknowledgment can also be provided to the user associated with the various user intents."; Executing processes associated with the user intents reads on determining an instruction corresponding to each intent included in a first intent.); sending the instruction corresponding to each intent in the first intent to the terminal device (Column 4, line 58 - Column 5, line 2, “As shown in FIG. 1, in some examples, a digital assistant can be implemented according to a client-server model. The digital assistant can include client-side portion 102 (hereafter “DA client 102”) executed on user device 104 and server-side portion 106 (hereafter “DA server 106”) executed on server system 108. DA client 102 can communicate with DA server 106 through one or more networks 110. DA client 102 can provide client-side functionalities such as user-facing input and output processing and communication with DA server 106. DA server 106 can provide server-side functionalities for any number of DA clients 102 each residing on a respective user device 104.”; Column 35, lines 42-57, "Natural language processing module 732 (“natural language processor”) of the digital assistant can take the sequence of words or tokens (“token sequence”) generated by STT processing module 730, and attempt to associate the token sequence with one or more “actionable intents” recognized by the digital assistant. An “actionable intent” can represent a task that can be performed by the digital assistant, and can have an associated task flow implemented in task flow models 754. The associated task flow can be a series of programmed actions and steps that the digital assistant takes in order to perform the task. The scope of a digital assistant's capabilities can be dependent on the number and variety of task flows that have been implemented and stored in task flow models 754, or in other words, on the number and variety of “actionable intents” that the digital assistant recognizes."; The natural language processing module associating the token sequence with one or more actionable intents recognized by the digital assistant reads on sending the instruction corresponding to each intent in the first intent to the terminal device, where a user device reads on a terminal device.); determining an instruction corresponding to each intent in a second intent, wherein the second intent comprises at least one intent arranged after the first intent in the at least two intents in a first sequence (Column 1, line 58 - Column 2, line 4, "Systems and processes are disclosed for processing a multi-part voice command. In one example, speech input can be received from a user that includes a single utterance having one or more actionable commands. A text string can be generated based on the speech input using a speech transcription process. The text string can be parsed into multiple candidate substrings. Probabilities can be determined for each of the candidate substrings indicating whether they are likely to correspond to actionable commands. In response to the probabilities exceeding a threshold, user intents can be determined for each of the candidate substrings. Processes associated with the user intents can then be executed. An acknowledgment can also be provided to the user associated with the various user intents."; Executing processes associated with the user intents reads on determining an instruction corresponding to each intent included in a second intent.). and sending the instruction corresponding to each intent in the second intent to the terminal device (Column 4, line 58 - Column 5, line 2, “As shown in FIG. 1, in some examples, a digital assistant can be implemented according to a client-server model. The digital assistant can include client-side portion 102 (hereafter “DA client 102”) executed on user device 104 and server-side portion 106 (hereafter “DA server 106”) executed on server system 108. DA client 102 can communicate with DA server 106 through one or more networks 110. DA client 102 can provide client-side functionalities such as user-facing input and output processing and communication with DA server 106. DA server 106 can provide server-side functionalities for any number of DA clients 102 each residing on a respective user device 104.”; Column 35, lines 42-57, "Natural language processing module 732 (“natural language processor”) of the digital assistant can take the sequence of words or tokens (“token sequence”) generated by STT processing module 730, and attempt to associate the token sequence with one or more “actionable intents” recognized by the digital assistant. An “actionable intent” can represent a task that can be performed by the digital assistant, and can have an associated task flow implemented in task flow models 754. The associated task flow can be a series of programmed actions and steps that the digital assistant takes in order to perform the task. The scope of a digital assistant's capabilities can be dependent on the number and variety of task flows that have been implemented and stored in task flow models 754, or in other words, on the number and variety of “actionable intents” that the digital assistant recognizes."; The natural language processing module associating the token sequence with one or more actionable intents recognized by the digital assistant reads on sending the instruction corresponding to each intent in the second intent to the terminal device, where a user device reads on a terminal device.). Gruber does not specifically disclose: the first sequence being a delivery sequence of the at least two intents in the speech stream. Teserra teaches: the first sequence being a delivery sequence of the at least two intents in the speech stream (Column 21, lines 45-61, "Precedence analyzer 430 analyzes each utterance constructed by the splitter 420 to determine an order in which to handle the utterances, e.g., an order in which the utterance are to be provided to an intent classifier. The order can be determined based on the structure of the utterances and the presence of certain words that signal precedence. For instance, as indicated in the example list of conjunctions above, the conjunctive words “before”, “after”, “and then”, and “only if” are some examples of words that can indicate order. The precedence analyzer 430 can take into consideration the extracted information 404 when deciding the order. For instance, the extracted information 404 can indicate important parts of the utterance 402 such as verbs and predicates (e.g. noun phrases) associated with the verbs. The precedence analyzer 430 places the utterances constructed by the splitter 420 into the queue 440 in the determined order."; Determining an order in which to handle the utterances reads on a delivery sequence of the at least two intents in the speech stream.). Teserra is considered to be analogous to the claimed invention because it is in the same field of voice recognition systems. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Gruber to incorporate the teachings of Teserra to determine an order in which to handle utterances. Doing so would allow for detecting that there are multiple intents represented in an utterance and matching each detected intent to an intent associated with a chatbot in a chatbot system (Teserra; Column 2, lines 3-7). Allowable Subject Matter Claims 7 – 9, 11 – 15 and 17 – 20 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The following is a statement of reasons for the indication of allowable subject matter: The primary reason claims 7 – 9 would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims is the Inclusion, in all the claims, of the limitations “when determining that the instructions corresponding to all of the at least two intents are not completely obtained, sending a first request to the third application module, wherein the first request is used to request the instruction corresponding to each intent in the second intent”, “updating, by the third application module, the first intent to a next intent arranged after the first intent in the first sequence, updating the historical dialogue status to slot information corresponding to the at least two intents after the instruction corresponding to the first intent is executed, and determining an instruction corresponding to an updated first intent”, and “continuing, by the first application module, to obtain from the third application module, the instruction corresponding to the updated first intent and terminating sending the first request to the third application module when determining that the instructions corresponding to all of the at least two intents are obtained” in combination with the limitations to obtain a speech stream comprising at least two intents, determine an instruction corresponding to each intent included in a first intent based on the speech stream, where the first intent comprises a part of each of the at least two intents, execute in a first sequence, the instruction corresponding to each intent in the first intent, and determine an instruction corresponding to each intent in a second intent, where the first sequence is a delivery sequence of the at least two intents in the speech stream, and the second intent comprises at least one intent that includes at least two intents and is arranged after the first intent. Gruber in view of Teserra discloses the method as claimed in claim 6. However, Gruber and Teserra, individually or in combination, do not disclose the limitations “when determining that the instructions corresponding to all of the at least two intents are not completely obtained, sending a first request to the third application module, wherein the first request is used to request the instruction corresponding to each intent in the second intent”, “updating, by the third application module, the first intent to a next intent arranged after the first intent in the first sequence, updating the historical dialogue status to slot information corresponding to the at least two intents after the instruction corresponding to the first intent is executed, and determining an instruction corresponding to an updated first intent”, and “continuing, by the first application module, to obtain from the third application module, the instruction corresponding to the updated first intent and terminating sending the first request to the third application module when determining that the instructions corresponding to all of the at least two intents are obtained”. The primary reason claims 11 – 13 would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims is the Inclusion, in all the claims, of the limitations “executing the instruction corresponding to the first intent, and when determining that the instructions corresponding to all of the at least two intents are not completely obtained, sending a first request to the server, wherein the instruction corresponding to each intent in the second intent is requested based on the first request so that after receiving the first request, the server updates the first intent to a next intent arranged after the first intent in the first sequence, updates the historical dialogue status to slot information corresponding to the at least two intents after the instruction corresponding to the first intent is executed, and determines an instruction corresponding to an updated first intent” and “continuing to obtain, from the server, the instruction corresponding to the updated first intent and stopping sending the first request to the server when determining that the instructions corresponding to all of the at least two intents are obtained”, in combination with the limitations to obtain a speech stream comprising at least two intents, send the speech stream to a server, receive, from the server, an instruction corresponding to each intent in a first intent, where the first intent comprises a part of each of the at least two intents, execute the instruction corresponding to each intent in the first intent, where the first sequence is a delivery sequence of the at least two intents in the speech stream, when determining that instructions corresponding to all of the at least two intents are not completely obtained, obtain, from the server, an instruction corresponding to each intent in a second intent, where the second intent comprises at least one intent that includes at least two intents and that is arranged after the first intent, and execute the instruction corresponding to each intent in the second intent until an instruction corresponding to each intent in the at least two intents is executed. Gruber in view of Teserra discloses the method as claimed in claim 10. However, Gruber and Teserra, individually or in combination, do not disclose the limitations “executing the instruction corresponding to the first intent, and when determining that the instructions corresponding to all of the at least two intents are not completely obtained, sending a first request to the server, wherein the instruction corresponding to each intent in the second intent is requested based on the first request so that after receiving the first request, the server updates the first intent to a next intent arranged after the first intent in the first sequence, updates the historical dialogue status to slot information corresponding to the at least two intents after the instruction corresponding to the first intent is executed, and determines an instruction corresponding to an updated first intent” and “continuing to obtain, from the server, the instruction corresponding to the updated first intent and stopping sending the first request to the server when determining that the instructions corresponding to all of the at least two intents are obtained”. The primary reason claims 14 – 15 would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims is the Inclusion, in all the claims, of the limitations “executing the instruction corresponding to the 1st intent and sending the first identifier to the server, wherein the instruction corresponding to each intent in the second intent is requested based on the first identifier”, “receiving, from the server, a written queue element arranged in sequence in the first queue, wherein an initial queue element of the written queue element comprises at least an instruction corresponding to a next intent that is in the remaining intent and that is arranged after the 1st intent”, and “executing, in the first sequence, an instruction corresponding to each intent corresponding to the written queue element, and when determining that the instructions corresponding to all of the at least two intents are not completely obtained, continuing to send the first identifier to the server to extracts from the first queue based on the first identifier, a newly written queue element arranged after the written queue element until it is determined that the instructions corresponding to all of the at least two intents are completely obtained”, in combination with the limitations to obtain a speech stream comprising at least two intents, send the speech stream to a server, receive, from the server, an instruction corresponding to each intent in a first intent, where the first intent comprises a part of each of the at least two intents, execute the instruction corresponding to each intent in the first intent, where the first sequence is a delivery sequence of the at least two intents in the speech stream, when determining that instructions corresponding to all of the at least two intents are not completely obtained, obtain, from the server, an instruction corresponding to each intent in a second intent, where the second intent comprises at least one intent that includes at least two intents and that is arranged after the first intent, and execute the instruction corresponding to each intent in the second intent until an instruction corresponding to each intent in the at least two intents is executed. Gruber in view of Teserra discloses the method as claimed in claim 10. However, Gruber and Teserra, individually or in combination, do not disclose the limitations “executing the instruction corresponding to the 1st intent and sending the first identifier to the server, wherein the instruction corresponding to each intent in the second intent is requested based on the first identifier”, “receiving, from the server, a written queue element arranged in sequence in the first queue, wherein an initial queue element of the written queue element comprises at least an instruction corresponding to a next intent that is in the remaining intent and that is arranged after the 1st intent”, and “executing, in the first sequence, an instruction corresponding to each intent corresponding to the written queue element, and when determining that the instructions corresponding to all of the at least two intents are not completely obtained, continuing to send the first identifier to the server to extracts from the first queue based on the first identifier, a newly written queue element arranged after the written queue element until it is determined that the instructions corresponding to all of the at least two intents are completely obtained”. The primary reason claims 17 – 19 would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims is the Inclusion, in all the claims, of the limitations “sending the instruction corresponding to the first intent to the terminal device”, “receiving a first request from the terminal device, wherein the instruction corresponding to each intent in the second intent is requested based on the first request”, and “updating the first intent to a next intent arranged after the first intent in the first sequence, updating the historical dialogue status to slot information corresponding to the at least two intents after the instruction corresponding to the first intent is executed, determining an instruction corresponding to an updated first intent, and sending the instruction corresponding to the updated first intent to the terminal device, until the first request is not received from the terminal device after first duration”, in combination with the limitations to receive, from a terminal device, a speech stream comprising at least two intents, determine an instruction corresponding to each intent in a first intent, where the first intent comprises a part of each of the at least two intents, send the instruction corresponding to each intent in the first intent to the terminal device, determine an instruction corresponding to each intent in a second intent, where the second intent comprises at least one intent arranged after the first intent in the at least two intents in a first sequence, the first sequence being a delivery sequence of the at least two intents in the speech stream, and send the instruction corresponding to each intent in the second intent to the terminal device. Gruber in view of Teserra discloses the method as claimed in claim 16. However, Gruber and Teserra, individually or in combination, do not disclose the limitations “sending the instruction corresponding to the first intent to the terminal device”, “receiving a first request from the terminal device, wherein the instruction corresponding to each intent in the second intent is requested based on the first request”, and “updating the first intent to a next intent arranged after the first intent in the first sequence, updating the historical dialogue status to slot information corresponding to the at least two intents after the instruction corresponding to the first intent is executed, determining an instruction corresponding to an updated first intent, and sending the instruction corresponding to the updated first intent to the terminal device, until the first request is not received from the terminal device after first duration”. The primary reason claim 20 would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims is the Inclusion of the limitations “writing, into a first queue corresponding to the first identifier, an instruction corresponding to each intent in a remaining intent, wherein the remaining intent comprises another intent other than the 1st intent in the at least two intents, and an arrangement sequence of queue elements in the first queue is consistent with the first sequence”, “sending an instruction corresponding to the 1st intent and the first identifier to the terminal device”, “receiving the first identifier from the terminal device, wherein the first identifier is used to request the instruction corresponding to each intent in the second intent”, and “extracting, from the first queue based on the first identifier, a written queue element arranged in sequence, wherein an initial queue element of the written queue element comprises at least an instruction corresponding to a next intent that is in the remaining intent and that is arranged after the 1st intent; and sending the written queue element to the terminal device until the first identifier is not received from the terminal device after second duration”, in combination with the limitations to receive, from a terminal device, a speech stream comprising at least two intents, determine an instruction corresponding to each intent in a first intent, where the first intent comprises a part of each of the at least two intents, send the instruction corresponding to each intent in the first intent to the terminal device, determine an instruction corresponding to each intent in a second intent, where the second intent comprises at least one intent arranged after the first intent in the at least two intents in a first sequence, the first sequence being a delivery sequence of the at least two intents in the speech stream, and send the instruction corresponding to each intent in the second intent to the terminal device. Gruber in view of Teserra discloses the method as claimed in claim 16. However, Gruber and Teserra, individually or in combination, do not disclose the limitations “writing, into a first queue corresponding to the first identifier, an instruction corresponding to each intent in a remaining intent, wherein the remaining intent comprises another intent other than the 1st intent in the at least two intents, and an arrangement sequence of queue elements in the first queue is consistent with the first sequence”, “sending an instruction corresponding to the 1st intent and the first identifier to the terminal device”, “receiving the first identifier from the terminal device, wherein the first identifier is used to request the instruction corresponding to each intent in the second intent”, and “extracting, from the first queue based on the first identifier, a written queue element arranged in sequence, wherein an initial queue element of the written queue element comprises at least an instruction corresponding to a next intent that is in the remaining intent and that is arranged after the 1st intent; and sending the written queue element to the terminal device until the first identifier is not received from the terminal device after second duration”. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: Marzinzik (US Patent No. 11,705,114) teaches a system for parsing separate intents in natural language speech. Solomon et al. (US Patent No. 11,017,765) teaches a method for receiving natural language user input from a user, parsing the user input to determine an intent template with slots, populating the slots in the intent template with information from user input, and performing resolution on the intent template to partially resolve unresolved information. Evermann et al. (US Patent No. 10,769,385) teaches a method for determining a plurality of user intents from a selected subset of text strings. Fujii et al. (US Patent No. 9,530,405) teaches a method for extracting one or more intention estimations from a text inputted in a natural language. Stewart (US Patent No. 9,454,960) teaches a method for disambiguating a user utterance containing at least two user intents. Park (US Patent Application Publication No. 2019/0164540) teaches a voice recognition system for analyzing an uttered command having multiple intents. Any inquiry concerning this communication or earlier communications from the examiner should be directed to James Boggs whose telephone number is (571)272-2968. The examiner can normally be reached M-F 8:00 AM - 5:00 PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at (571)272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JAMES BOGGS/Examiner, Art Unit 2657
Read full office action

Prosecution Timeline

Dec 13, 2024
Application Filed
Jul 28, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12694880
METHOD AND DEVICE FOR AUDIO BAND-WIDTH DETECTION AND AUDIO BAND-WIDTH SWITCHING IN AN AUDIO CODEC
3y 3m to grant Granted Jul 28, 2026
Patent 12682911
AUTOMATIC DETECTION AND ATTENUATION OF SPEECH-ARTICULATION NOISE EVENTS
3y 5m to grant Granted Jul 14, 2026
Patent 12682181
MULTIMODAL DIALOGS USING LARGE LANGUAGE MODEL(S) AND VISUAL LANGUAGE MODEL(S)
3y 0m to grant Granted Jul 14, 2026
Patent 12670922
AUDIO PROCESSING METHOD AND APPARATUS
2y 3m to grant Granted Jun 30, 2026
Patent 12651112
INTELLIGENTLY IDENTIFYING FRESHNESS OF TERMS IN DOCUMENTATION
3y 9m to grant Granted Jun 09, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
63%
Grant Probability
97%
With Interview (+34.0%)
3y 2m (~1y 6m remaining)
Median Time to Grant
Low
PTA Risk
Based on 119 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month