Prosecution Insights
Last updated: October 04, 2026
Application No. 18/323,641

HUMAN-MACHINE DIALOGUE SYSTEM AND METHOD

Final Rejection §103
Filed
May 25, 2023
Priority
Jun 01, 2022 — CN 202210615940.6
Examiner
ZHU, RICHARD Z
Art Unit
2654
Tech Center
2600 — Communications
Assignee
Alibaba Damo (Hangzhou) Technology Co., Ltd.
OA Round
4 (Final)
69%
Grant Probability
Favorable
5-6
OA Rounds
0m
Est. Remaining
85%
With Interview

Examiner Intelligence

Grants 69% — above average
69%
Career Allowance Rate
509 granted / 734 resolved
+7.3% vs TC avg
Strong +16% interview lift
Without
With
+15.7%
Interview Lift
resolved cases with interview
Typical timeline
3y 3m
Avg Prosecution
25 currently pending
Career history
765
Total Applications
across all art units

Statute-Specific Performance

§101
13.1%
-26.9% vs TC avg
§103
59.7%
+19.7% vs TC avg
§102
20.5%
-19.5% vs TC avg
§112
4.4%
-35.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 734 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Acknowledgement Acknowledgement is made of applicant’s amendment made on 08/04/2026. Applicant’s submission filed has been entered and made of record. Status of the Claims Claims 1-20 are pending. Claims 5-13 are withdrawn. Response to Applicant’s Arguments In response to “Applicant respectfully submits that, in the cited paragraphs, Xing does not teach or suggest using a fusion feature to determine "whether the inserted voice reflects an interrupt intention," as recited by claims 1 and 14. Xing, at best, merely discusses using audio-textual representation to generate a "semantic prediction" that can be mapped to a "command action" representing a speaker's general intent. See Xing, paragraphs [0046], [0047], [0060]. Xing is silent about detecting an interrupt intention of an inserted voice during playback of a dialogue response. Xing's system is used to extract a general-purpose command and does not evaluate whether a user's voice, inserted while a dialogue response is being played, reflects an intention to interrupt that response. Thus, the fusion feature disclosed in Xing serves a fundamentally different purpose from the fusion feature recited by claims 1 and 14” and “Shen's rule-based and word-matching approach to interrupt detection is fundamentally different from the claimed joint modeling of speech and text using a fusion feature. Neither Shen nor Xing, alone or in combination, suggests or teaches determining an interrupt intention from a fused text-and-voice feature, as required by pending claims 1 and 14”. Shen teaches user outputting a voice stream during voice robot’s voice output process comprises two purposes / intents (p. 4, ¶12): (1) interrupt the voice robot directly (e.g., “Please stop”) or indirectly (e.g., “XX can explain it in more detail?”) or (2) not interrupting the voice robot with expressions like “oh” or “O”. Shen performs judgment no matter what kind of voice the user sends (p. 4, ¶12) by performing speech recognition on the speech signals to obtain speech recognition results to determine whether the number of characters in the user’s incoming voice stream exceeds a preset text threshold (p. 4, ¶14, step S104). If the number of text in the coming voice stream exceeds the text threshold, it is judged that the user sends out interrupted voice stream (p. 4, ¶18) and detection device interrupts the voice robot (p. 4, ¶19, step S105) by immediately stopping the voice robot currently outputting a sentence (p. 5, ¶20). If the number of text in the coming voice stream does not exceed the text threshold, determine if the recognition result matches a mood word in a preset mood bag database (p. 5, ¶21) comprising common mood words like “ah”, “oh”, “no”, “?”, “say again” indicating the user has an intention to interrupt (p. 5, ¶26). If the recognition result matches a mood word, then executes an action of interrupting the voice robot (p. 5, ¶25). Therefore, Shen’s determination of whether the inserted voice reflects an interrupt intention was entirely based on text feature. However, according to Xing, it is challenging for computers to accurately interpret intent or emotion associated with an utterance using linguistic content alone comprising ASR transcription of speech to text (Xing ¶3). When applied to the scenario of Shen where matching a mood word in the mood word database (Shen, p. 5, ¶25) or when user sends a tone alone to assume that user fails to understand clearly or does not want to continue (Shen, p. 5, ¶26), relying entirely on speech recognition result makes prediction of user intent inaccurate because speech recognition text transcription of speech signal cannot capture intent and emotion communicated through timing, intensity, intonation, and pitch (Xing ¶33). Accordingly, Xing teaches a multimodal language understanding system generating semantic predictions of a command action representing a speaker’s intent based on encoded and fused audio-textual representation (Xing ¶6). Specifically, as new speech chunks and corresponding text predictions are obtained, update the semantic prediction by processing the audio-textual representation (Xing ¶6) comprising: (1) encode ASR text transcript corresponding to a speech chunk to generate encoded representation of text prediction (Xing ¶75) (2) encode the speech chunk to generate an encoded representation of the speech chunk (Xing ¶74), and (3) generate the audio-textual representation by concatenating the encoded representation of the text prediction and encoded speech (Xing ¶78) corresponding to fused multimodal feature integrating speech features and text features enabling a model to learn a joint representation of the modalities that enables additional semantic information to be extracted from the speech modality to help capture important semantic cues that are not present in the text transcript (Xing ¶79). Applying such joint modeling of speech and text comprising fused multimodal feature integrating speech features with text features to situations of user inserted voices comprising mood words or sending a tone alone (Shen, p. 5, ¶26), inaccuracy in intent determination based on general understanding can be avoided by considering important semantic cues such as timing, intensity, intonation, and pitch. In response to “The Office merely stated a generic motivation (to "improve performance of spoken language understanding" and make "semantic predictions representing a speaker's intent"), without explaining why a person of ordinary skill in the art would have applied the command- extraction fusion of Xing specifically to the interrupt-intention decision of Shen. Xing's fusion feature is designed to extract general semantic commands from speech, and a person of ordinary skill in the art would not have been motivated to use the teachings of Xing to modify Shen to determine whether an inserted voice during playback of a dialogue response reflects an interrupt intention”. It would’ve been obvious to one ordinarily skilled in the art before the effective filing date of the invention to determine whether the inserted voice reflects an interrupt intention by a joint modeling of speech and text, according to a fusion feature obtained by combining a text feature and a voice feature of the inserted voice when user inserted voice comprises mood words or tones alone (Shen, p. 5, ¶26) in order to provide fusion feature combining text feature and voice feature to improve performance of spoken language understanding (Xing, ¶44) making semantic predictions representing a speaker’s intent (Xing, Abstract) by considering important semantic cues such as timing, intensity, intonation, and pitch (Xing, ¶33). In response to “In addition, Applicant further respectfully submits that the Office's proposed modification would change the principle of operation of the cited art. In particular, the proposed modification would change the principle of operation taught in Shen of using a discrete, rule- based, staged mechanism for interrupt determination” and “Modifying Shen so that the interrupt intention is instead determined "by a joint modeling of speech and text, according to a fusion feature obtained by combining a text feature and a voice feature of the inserted voice," as recited by claims 1 and 14, would dispense with and replace Shen's character-count/preset-text-threshold gating and its separate dictionary/tone matching, which is the core mechanism operated by Shen. Such a modification would render Shen unsatisfactory for its intended purpose of providing a lightweight, rule-based interrupt filter”. In Shen, user intent prediction in cases where user inserted voice comprises mood words or tones alone is based on general understanding that the user fails to understand clearly. According to Xing, such general understanding makes intent prediction inaccurate. Specifically, Xing suggests such reliance on text transcription alone to determine user intent is unsatisfactory because speech recognition text transcription of speech signal cannot capture intent and emotion communicated through timing, intensity, intonation, and pitch (Xing ¶33). Therefore, Xing provides a suggestion to modify the operations of Shen to capture the intent and emotion of mood words or tones alone based on timing, intensity, intonation, and pitch captured in the fused multimodal features of speech and text in a single joint representation (Xing ¶79) to improve performance of spoken language understanding (Xing, ¶44) at least regarding whether user’s intent is to really interrupt the voice robot when using mood words or sending tones alone. Claim Rejections - 35 USC § 103 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 103 that form the basis for the rejections under this section made in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 14, and 20 are rejected under 35 USC 103(a) as being unpatentable over Chatterjee et al. (US 11087094 B2) in view of Shen et al. (CN 110853638 A, see attached EPO translation) and Xing et al. (US 2023/0223018 A1). Regarding Claims 1 and 14, Chatterjee discloses a human-machine dialogue system (Fig. 1), comprising: a non-transitory computer-readable storage medium storing a set of instructions that are executable by one or more processors of a device to cause the device to perform a human-machine dialogue method and one or more processors configured to execute instructions to cause the human-machine dialogue system to perform operations of the human-machine dialogue method (Col 2, Rows 50-54 and Col 3, Rows 19-24) comprising: performing intention clustering of a dialogue data sample based on a semantic representation of the dialogue data sample (Col 15, Rows 52-63, receiving a set of conversations and related meta data (e.g., Col 17, Rows 6-7, conversation summary), converting each of the set of conversations and related meta data into a set of feature representations in a multi-dimensional vector space, and performing density based spatial clustering of applications with noise to identify clusters among the set of feature representations; per Col 17, Rows 28-34, once the cluster is correctly formed with a similar set of conversations, add a label to the cluster to create an intent model to infer user goals during conversations); constructing, based on a clustering result, a dialogue procedure corresponding to the dialogue data sample (Col 17, Rows 53-57, create nodes and assign word sequences to a node in a conversation graph and form corresponding transitional paths to represent the plurality of conversations that have occurred and were directed to a particular task, intent, goal, etc.; per Col 25, Rows 24-28, the conversation graph embeds utterance patterns for agents as a plurality of nodes and user utterances as a plurality of transitional paths or edges such that the conversation graph encodes an agent’s behavior based on user responses); receiving a voice dialogue from a user (Col 5, Rows 54-58, request 102 / caller voice request; Fig. 10, receiving user utterance 1010); converting the voice dialogue into a dialogue text (Col 5, Rows 65-67, automatic speech recognition system 110 processes request 102); obtaining a semantic representation corresponding to the dialogue text of the voice dialogue of the user (Col 6, Rows 5-8, pass the sequence of words to natural language understanding system 112; Col 25, Rows 36-52, Fig. 10, passing a first input / current user utterances 1010 into an embedding layer 1012 and LSTM layer 1014 where embedding layer 1012 generates embedding vectors of the current user utterances 1010); performing intention analysis on the semantic representation corresponding to the voice dialog to obtain an intention analysis result (Col 6, Rows 10-20, natural language understanding system outputs dialogue act category, the intent of the user, and slot names and values; Col 25, Rows 36-41, Fig. 10, passing a first input / current user utterances 1010 into an embedding layer 1012 and LSTM layer 1014; in view of Col 9, Rows 4-18 and Col 11, Rows 28-33, the LSTM layer 1014 corresponds to a spoken language understanding system including a bidirectional RNN comprising LSTM to determine slot, intent, and dialogue act classification (per Col 8, Rows 44-46, determine if the dialog is a question, query, greeting, command, or request, and information)); determining, according to the intention analysis result and the dialogue procedure constructed in advance, a dialogue response (Col 25, Rows 55-60, neural network dense layer 1040 output for predicting the next best node and sample an utterance randomly from the node to automatically generate a conversation flow to allow a system to respond in real-time to different customer utterances and conversation paths); and performing voice interaction of the dialogue response with the user (Col 6, Rows 38-44, track current state of the dialog between virtual agent 100 and customer and to respond to the request in a conversational manner) by converting the dialogue response into voice so as to interact with the user by voice (Col 6, Rows 52-53, convert words into audible response 104 by text to speech synthesis unit 118), wherein the dialogue response is an answer response to the voice dialogue (Col 7, Rows 28-37, look up answer to a particular question), or a clarification response to clarify a dialogue intention of the voice dialogue (Col 6, Rows 45-53, “when would you like to leave?”). Chatterjee does not disclose in a process of performing voice dialogue interaction with the user, performing: determining whether an inserted voice by the user is detected before the dialogue response being played is finished; in response to determining that the inserted voice is detected before the dialogue being played is finished, determining whether the inserted voice reflects an interrupt intention and in response to the determination that the inserted voice does not reflect the interrupt intention, continuing playing the dialogue response. Shen teaches a human machine dialogue system (p. 1, “Summary of the Invention”, “a method and a device for interrupting a voice robot in real time during a voice interaction process, so that a user can interrupt a voice robot’s voice output at any time during the voice interaction process with the voice robot”) where in a process of performing voice dialogue interaction with the user (p. 3, ¶6, step S101, the voice robot performs voice interaction with the user’s voice), performing: determining whether an inserted voice by the user is detected before the dialogue response being played is finished (p. 3, ¶8 and p. 3, ¶10, step S102, determine whether the voice robot outputs a voice, Step S103, when the voice robot outputs a voice, start a detection device to detect whether the user sends an incoming voice stream); in response to determining that the inserted voice is detected before the dialogue being played is finished, determining whether the inserted voice reflects an interrupt intention (p. 4, ¶12, user’s incoming voice stream is the voice stream output by the user during the voice robot’s voice output process where purpose of user’s voice is to interrupt the voice robot or not interrupt; p. 4, ¶18, judging whether user sends out interrupted voice stream when the number of text in the incoming voice stream does not exceed a preset text threshold) and in response to the determination that the inserted voice does not reflect the interrupt intention, continuing playing the dialogue response (p. 5, ¶21 and ¶26, when the number of texts in the incoming voice stream does not exceed the preset text threshold, execute step S106 to match the recognition result of the incoming voice stream with a preset mood bag database corresponding to user having an intention to interrupt; p. 5, ¶27, when the recognition result of the incoming voice stream does not match the preset mood bag database, the detection device does not perform the action of interrupting the voice robot in response to the incoming voice stream). It would’ve been obvious to one ordinarily skilled in the art before the effective filing date of the invention to determine whether an inserted voice by the user is detected before the dialogue response being played is finished, determine whether the inserted voice reflects an interrupt intention, and continue playing the dialogue response when it is determined that the inserted voice does not reflect the interrupt intention in order to enable the user to interrupt the voice (i.e., dialogue response) being output by the voice robot in real-time during the voice interaction process between the user and the voice robot, and to filter the non-interrupted user’s input voice (Shen, p. 3, ¶3). The combination of Chatterjee and Shen does not disclose determining whether the inserted voice reflects an interrupt intention by a joint modeling of speech and text, according to a fusion feature obtained by combining a text feature and a voice feature of the inserted voice. Xing discloses a multimodal language understanding system making a semantic prediction representing speaker’s intent based on an audio-textual representation (¶30 and Figs. 2-4) by determining speaker’s intent by a joint modeling of speech and text, according to a fusion feature obtained by combining a text feature and a voice feature of the speaker’s inserted voice (¶62, generate a sequence of speech chunks 230 and corresponding text transcripts 240 from speech signal 210; ¶74, encode the speech chunks into encoded speech embeddings 414; ¶75, encode the text transcript 240 to generate encoded word embeddings 416; ¶¶75-77, generate a uniform representation 420 comprising temporally aligned speech embedding and word embedding; ¶¶78-79, concatenate the uniform representation to generate an audio-textual representation 424 constituting fused multimodal feature integration of speech and text; ¶80, transform the audio-textual embeddings 424 into speaker’s intent). In Shen, the determination of user’s intention to interrupt is based on recognition of user’s tone and user’s words expressing certain mood (Shen, p. 5, ¶26) that Xing’s fusion feature can capture (Xing, ¶79, audio-textual embeddings 424 being a joint representation of both speech and text to help to capture important semantic cues not present in the text transcript; per ¶33, intent and emotion may be communicated through semantic cues such as timing, intensity, intonation, pitch not captured within a text transcript of speaker speech). It would’ve been obvious to one ordinarily skilled in the art before the effective filing date of the invention to determine whether the inserted voice reflects an interrupt intention by a joint modeling of speech and text, according to a fusion feature obtained by combining a text feature and a voice feature of the inserted voice when user inserted voice comprises mood words or tones alone (Shen, p. 5, ¶26) in order to provide fusion feature combining text feature and voice feature to improve performance of spoken language understanding (Xing, ¶44) making semantic predictions representing a speaker’s intent (Xing, Abstract) by considering important semantic cues such as timing, intensity, intonation, and pitch (Xing, ¶33). Regarding Claim 20, Chatterjee discloses wherein one part of the dialogue data sample is tagged data, and the other part of the dialogue data sample is untagged data (Col 17, Rows 28-46 and Col 18, Rows 38-44, 90%-97% of the conversation data can be annotated with a classifier for user intent; i.e., 3% to 10% of the conversation data are not annotated, which may require a human expert to add the label (per Col 17, Rows 21-27) in semi-supervised labeling algorithm). Claims 2-3 and 18-19 are rejected under 35 USC 103(a) as being unpatentable over Chatterjee et al. (US 11087094 B2) in view of Shen et al. (CN 110853638 A) and Xing et al. (US 2023/0223018 A1) as applied to claims 1 and 14, in view of Dubey et al. (US 2019/0124202 A1). Regarding claims 2 and 18, Chatterjee discloses wherein the operations further comprise: performing dialogue semantic cluster segmentation on the dialogue data sample in advance based on the semantic representation of the dialogue data sample (Col 15, Row 57 – Col 16, Row 4 and Col 16, Rows 48-55, convert each of the set of conversations and related meta data into feature representations in multi-dimensional vector space and using density based spatial clustering of applications with noise (DBSCAN) to determine a first set of clusters and a second set of conversations not mapped to the first set of clusters); performing density clustering according to a semantic cluster obtained by segmentation and a dialogue representation vector corresponding to the dialogue data sample (Col 15, Rows 3-7 and Col 16, Rows 8-12, clustering model runs multiple iterations over the data to adapt the problem space to find large clusters first and then increasingly smaller clusters in subsequent iterations); obtaining, according to the clustering result, at least one start intention and dialogue data corresponding to each start intention (Col 17, Rows 28-34, once the cluster is correctly formed with a similar set of conversations, add label to the cluster to train a classification model to create a intent model to infer user goals during conversations; Col 17, Rows 34-41, in an embodiment training is done on actual conversations where training data is collected by extracting first few utterances of the customer and an intent label is used as an end class to build a mapping function from conversation to intent); for each start intention, performing dialogue path mining based on dialogue data corresponding to the start intention (Col 17, Rows 55-60, create nodes and assign word sequences into a node in a conversation graph as well as forming corresponding transitional paths, the conversation graph represents a plurality of conversations that have occurred and directed to particular task, intent, goal, and objective; see e.g., Fig. 4 and Col 18, Rows 15-54, for each collection of conversations determined to have the same intent (per Col 18, Rows 15-17 and Rows 52-54, intent label), perform steps 410-440 to extract agent utterances comprising agent questions and corresponding question classification, agent resolutions and corresponding resolution classification); and constructing, according to a mining result, a dialogue procedure corresponding to the dialogue data sample (Col 18, Rows 54-56, convert extracted question-resolution sequence from each conversation to a graph; see e.g., Fig. 5, Rows 57-67, a conversation graph with set V of nodes and set E of edges (transitional paths)). Chatterjee does not disclose that the density clustering is hierarchical density clustering. Dubey teaches using natural language processing techniques to identify topics associated with conversations and perform hierarchical density-based cluster analysis (H-DBSCAN) to group conversations under one or more topics (¶31). It would’ve been obvious to one ordinarily skilled in the art before the effective filing date of the invention to perform hierarchical density clustering according to the semantic cluster in order to group conversations under one or more topics (Dubey, ¶31) such as dialogue act classifications. Regarding Claims 3 and 19, Chatterjee discloses wherein the dialogue procedure corresponding to the dialogue data sample is constructed according to the mining result by: obtaining, according to the mining result, dialogue semantic clusters respectively corresponding to a user and a robot customer service corresponding to the dialogue data (Col 16, Rows 34-41, clustering process obtains set of conversations comprising user’s utterances and responsive robot’s utterance); constructing a key dialogue transfer matrix according to the dialogue semantic clusters respectively corresponding to the user and the robot customer service (Col 16, Rows 40-55, use word2vec model to convert the set of conversations to feature representation, run the feature representation through density based spatial clustering of applications with noise to generate a first set of clusters and a second set of conversations not mapped to the first set of clusters, and generate a conversation feature matrix); generating, according to the key dialogue transfer matrix, a dialogue path used to indicate a dialogue procedure (Col 17, Rows 53-57, use the obtained data (i.e., conversation feature matrix) to create nodes and assign words sequences to a node in a conversation graph as well as form corresponding transitional paths); and mounting the generated dialogue path to the start intention to construct the dialogue procedure corresponding to the dialogue data sample (Col 18, Row 54 – Col 19, Row 22, convert to graph such as the conversation graph 500 of Fig. 5 comprise set V of nodes and set E of edges (transitional paths) with different sets of nodes representing different intents such as information seeking or action node to provide an action or resolution; per Col 17, Rows 57-60, a conversation graph represents a plurality of conversations that have occurred and were directed to a particular task, intent, goal, and/or objective). Claims 4 and 15-17 are rejected under 35 USC 103(a) as being unpatentable over Chatterjee et al. (US 11087094 B2) in view of Shen et al. (CN 110853638 A) and Xing et al. (US 2023/0223018 A1) as applied to claims 1 and 14, in view of Walters et al. (US 2014/0244712 A1). Regarding Claims 4 and 15, Chatterjee does not disclose in a process of performing dialogue interaction with the user, detecting whether an insertion timing for a set speech exists, and inserting the set speech when the insertion timing is detected. Walters discloses interaction environment comprising a dialog interface to receive input from a user and an interaction engine core determines a user’s intent to provide feedback or responses back to the user (¶86) where in a process of performing dialogue interaction with the user, detecting whether an insertion timing for a set speech exists (¶142, when determining to resume a task after an interruption of a shorter time period or an interruption of a longer time period), and inserting the set speech when the insertion timing is detected (¶142, with an interruption of a shorter time period, virtual assistant 710 utters “So, what date did you want to fly” and with an interruption of a longer time period, virtual assistant 710 utters “Should we continue with the flight booking to New York?”; i.e., insert “So, what date did you want to fly” or “Should we continue with the flight booking to New York?” based on length of time of interruption and chance of user forgetting where the dialog left off). It would’ve been obvious to one ordinarily skilled in the art before the effective filing date of the invention to insert set speech when insertion timing is detected in order to implement task resumption strategies based on length of interruption timing and based on chance of user forgetting where the dialog left off (Walters, ¶142). Regarding Claims 4 and 16, Chatterjee does not disclose in a process of performing dialogue interaction with the user, detecting the inserted voice by the user, and in response to determining that an intention corresponding to the inserted voice is to interrupt a dialogue voice, processing the inserted voice. Walters discloses in a process of performing dialogue interaction with the user, detecting an inserted voice by the user, and in response to determining that an intention corresponding to the inserted voice is to interrupt a dialogue voice, processing the inserted voice (¶¶108-110, when user requested “Can you check flights to Paris for Friday” and user interrupts the dialog “We will have to continue later I am home now”, virtual assistant executes the interruption, place a time stamp on the dialog, and store it in task repository). It would’ve been obvious to one ordinarily skilled in the art before the effective filing date of the invention to process inserted voice by a user in response to a determination that an intention corresponding to the inserted voice is to interrupt a dialogue voice in order to maintain a persistence of knowledge by preserving dialog sessions and techniques for dialog resumption (Walters, ¶105). Regarding Claims 4 and 17, Chatterjee does not disclose detecting a pause of the user in a dialogue interaction process, and in response to a detection result indicating that a dialogue corresponding to the pause is incomplete, inserting a guide language to guide the user to complete the dialogue. Walters discloses in a process of performing dialogue interaction with the user, detecting a pause of the user in a dialogue interaction process (¶¶108-110, as virtual assistant responds “I have found ten flights for Friday. When do you need to be there?” to user request “Can you check flights to Paris for Friday?”, determine that user has paused the conversation based on user interruption “We will have to continue later I am home now”), and in response to a detection result indicating that a dialogue corresponding to the pause is incomplete, inserting a guide language to guide the user to complete the dialogue (¶142, with an interruption (i.e., pause) of a shorter time period, virtual assistant 710 utters “So, what date did you want to fly”). It would’ve been obvious to one ordinarily skilled in the art before the effective filing date of the invention to inserting a guide language to guide the user to complete the dialogue in response to a detection result indicating that a dialogue corresponding to the pause is incomplete in order to implement task resumption strategies based on chance of user forgetting where the dialog left off (Walters, ¶142). Conclusion Prior art made of record and not relied upon is considered pertinent to applicant's disclosure: CN113488024A discloses an intelligent communication robot in a voice dialogue interaction with a user detected an user inserted voice before logistics prompt voice being played is finished (¶55, step S201) and detect an interrupt intention by a joint modeling of speech and text, according to a fusion feature obtained by combining a text feature and a voice feature of the inserted voice (¶57 and step S203, using a CNN-LSTM model to judge whether user’s current voice is likely to be real semantic breaking according to voice feature, text feature, and combined system voice for semantic prediction; ¶76 and Fig. 4, CNN-LSTM model splicing semantic feature of the audio feature and the semantic feature of the text feature as the final semantic representation indicating breaking voice is the real breaking voice). THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to examiner Richard Z. Zhu whose telephone number is 571-270-1587 or examiner’s supervisor Hai Phan whose telephone number is 571-272-6338. Examiner Richard Zhu can normally be reached on M-Th, 0730:1700. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /RICHARD Z ZHU/Primary Examiner, Art Unit 2654 09/02/2026
Read full office action

Prosecution Timeline

Show 2 earlier events
Jan 07, 2026
Response Filed
Jan 28, 2026
Final Rejection mailed — §103
Mar 30, 2026
Response after Non-Final Action
Apr 23, 2026
Request for Continued Examination
Apr 24, 2026
Response after Non-Final Action
May 06, 2026
Non-Final Rejection mailed — §103
Aug 04, 2026
Response Filed
Sep 04, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12738263
Method and System for Constructing Speech Recognition Model and Speech Processing
2y 10m to grant Granted Sep 15, 2026
Patent 12699692
TECHNIQUES FOR EFFICIENT ENCODING IN NEURAL SEMANTIC PARSING SYSTEMS
2y 6m to grant Granted Aug 04, 2026
Patent 12694875
METHOD, DEVICE AND SYSTEM OF CONTEXTUAL AND USER LOCATION BASED OPERATIONAL EXECUTION VIA A GENERATIVE ARTIFICIAL INTELLIGENCE (AI) COMPUTING PLATFORM IN RESPONSE TO USER INTERACTION THEREWITH
2y 7m to grant Granted Jul 28, 2026
Patent 12657538
AUDIO SIGNAL PROCESSING AND DYNAMIC NATURAL LANGUAGE UNDERSTANDING
3y 5m to grant Granted Jun 16, 2026
Patent 12633287
DOMAIN MODEL DRIVEN PROCESSING OF DIALOG INCLUDING AMBIGUOUS INTENTS
2y 5m to grant Granted May 19, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
69%
Grant Probability
85%
With Interview (+15.7%)
3y 3m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 734 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month