DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 6-10-2025 was filed after the mailing date of the claims on 9-12-2024. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1-4, 6-14 and 16-20 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Rossi et al. (U.S. 2024/0347058).
With regard to claim 1, Rossi teaches an information processing [abstract] method, comprising:
providing a first control associated with a speech input interface (Fig. 1, conversation system 110; [abstract] a user's spoken input is received), wherein the speech input interface is displayed with a first text ([abstract] a user's spoken input is received and converted to text, which forms a prompt for the Large Language Models (LLMs); [0005] The system receives a first speech input from a client device of a user, and converts the first speech input into first text), and the first text is obtained by converting a first speech input ([0005] The system receives a first speech input from a client device of a user, and converts the first speech input into first text); and
displaying a second text on the speech input interface in response to an operation associated with the first control ([abstract] the LLM generates a second text response based on the interruption, which is converted to speech and played back to the user), wherein the second text is obtained by processing the first text based on a processing process corresponding to the first control ([abstract] Upon generating a text response by the LLM, the text response is then converted back into speech and played to the user. If the user interrupts while the response is being played, the playback stops, and the interruption is captured as a new spoken input. This interruption is used to generate a new prompt for the LLM. Subsequently, the LLM generates a second text response based on the interruption, which is converted to speech and played back to the user; [0006] Upon detecting the vocal interruption during the play of the first speech response, the system causes the client device to stop playing the first speech response. The system converts the second speech input into a second text input, and generates a second prompt to the LLM based on the vocal interruption).
With regard to claim 2, the limitations are addressed above and Rossi teaches wherein the processing process corresponding to the first control is a processing process based on an artificial intelligence technology ([0003] A large language model (LLM) like GPT (Generative Pre-trained Transformer) is an advanced artificial intelligence system designed to understand and generate human-like text based on the input it receives; [0015] FIG. 5 illustrates an example process of training or retraining a machine learning model 520; [0037] a machine learning model may be trained to generate or modify responses based on a current state when a vocal interruption is received).
With regard to claim 3, the limitations are addressed above and Rossi teaches wherein the first control comprises one or more of:
a control for invoking a digital assistant interactive interface ([0060] The client device 130 may include (but is not limited to) a smartphone, a tablet, a laptop, a desktop, a smartwatch, an e-reader, a gaming console, a smart home device, a wearable fitness tracker, a virtual reality (VR) and/or augmented reality (AR) headset, a digital camera, a portable media player, a drone, among others. The client device 130 may have a user agent application 300 installed); and
a shortcut command control for preset processing ([0053] if the vocal interruption occurs beyond a second threshold time In some embodiments, the response modification module 270 may also be configured to shorten a response based on a type of interruption).
With regard to claim 4, the limitations are addressed above and Rossi teaches wherein the first control comprises a shortcut command control for preset processing ([0053] if the vocal interruption occurs beyond a second threshold time In some embodiments, the response modification module 270 may also be configured to shorten a response based on a type of interruption), and displaying the second text on the speech input interface in response to the operation associated with the first control ([abstract] the LLM generates a second text response based on the interruption, which is converted to speech and played back to the user) comprises:
displaying the second text on the speech input interface in response to a trigger operation for the shortcut command control ([0072] the conversation system 110 also determines whether to roll back the conversation state based in part on a pause 442D between user's two speeches 412D and 416D, a pause 444D between the starting of the audio response 414D and the start of the user speech 416D, and/or a pause 446D between user's two speeches 416D and 420D. For example, if the pause 444D is less than a threshold time, the conversation system 110 determines that the pause is a short pause), wherein the second text is obtained by processing the first text based on a preset processing process corresponding to the shortcut command control ([0072] the conversation system 110 also determines whether to roll back the conversation state based in part on a pause 442D between user's two speeches 412D and 416D, a pause 444D between the starting of the audio response 414D and the start of the user speech 416D, and/or a pause 446D between user's two speeches 416D and 420D. For example, if the pause 444D is less than a threshold time, the conversation system 110 determines that the pause is a short pause).
With regard to claim 6, the limitations are addressed above and Rossi teaches wherein the first control comprises a control for invoking a digital assistant interactive interface ([0060] The client device 130 may include (but is not limited to) a smartphone, a tablet, a laptop, a desktop, a smartwatch, an e-reader, a gaming console, a smart home device, a wearable fitness tracker, a virtual reality (VR) and/or augmented reality (AR) headset, a digital camera, a portable media player, a drone, among others. The client device 130 may have a user agent application 300 installed), and displaying the second text on the speech input interface in response to the operation associated with the first control ([abstract] This interruption is used to generate a new prompt for the LLM. Subsequently, the LLM generates a second text response based on the interruption, which is converted to speech and played back to the user; [0006] Upon detecting the vocal interruption during the play of the first speech response, the system causes the client device to stop playing the first speech response. The system converts the second speech input into a second text input, and generates a second prompt to the LLM based on the vocal interruption) comprises:
displaying the digital assistant interactive interface in response to a trigger operation for the first control ([0072] the conversation system 110 also determines whether to roll back the conversation state based in part on a pause 442D between user's two speeches 412D and 416D, a pause 444D between the starting of the audio response 414D and the start of the user speech 416D, and/or a pause 446D between user's two speeches 416D and 420D. For example, if the pause 444D is less than a threshold time, the conversation system 110 determines that the pause is a short pause), wherein the digital assistant interactive interface provides a plurality of candidate command controls (Fig. 7; [0082] The text record 700 includes text of a user's speech requests 710, 730, and text responses 720, 740 generated by the LLM system 120. The text record 700 also includes metadata associated with the user's speech and responses, including the timing and duration of each speech request or speech response); and
displaying the second text on the speech input interface in response to a trigger operation for a target command control from the plurality of candidate command controls ([abstract] the LLM generates a second text response based on the interruption, which is converted to speech and played back to the user), wherein the second text is obtained by processing the first text based on a processing process corresponding to the target command control ([abstract] Upon generating a text response by the LLM, the text response is then converted back into speech and played to the user. If the user interrupts while the response is being played, the playback stops, and the interruption is captured as a new spoken input. This interruption is used to generate a new prompt for the LLM. Subsequently, the LLM generates a second text response based on the interruption, which is converted to speech and played back to the user; [0006] Upon detecting the vocal interruption during the play of the first speech response, the system causes the client device to stop playing the first speech response. The system converts the second speech input into a second text input, and generates a second prompt to the LLM based on the vocal interruption).
With regard to claim 7, the limitations are addressed above and Rossi teaches wherein displaying the digital assistant interactive interface in response to the trigger operation for the first control ([0072] the conversation system 110 also determines whether to roll back the conversation state based in part on a pause 442D between user's two speeches 412D and 416D, a pause 444D between the starting of the audio response 414D and the start of the user speech 416D, and/or a pause 446D between user's two speeches 416D and 420D. For example, if the pause 444D is less than a threshold time, the conversation system 110 determines that the pause is a short pause) comprises:
obtaining, in response to the trigger operation for the first control, information of a business scenario to which the speech input interface belongs ([0004] LLM can also be integrated into chatbots and virtual assistants; businesses can automate customer service, providing quick and accurate responses to inquiries, which enhances customer experience and operational efficiency; [0005] The system receives a first speech input from a client device of a user, and converts the first speech input into first text); and
displaying the digital assistant interactive interface ([0052] The interface module 260 is configured to transmit text prompt generated by the prompt generation module 220 to the LLM system 120 and receive text response generated by the LLM system 120; [0053] if an interruption is detected shortly after a speech response starts to play, the response modification module 270 may be configured to generate a continuer response, such as “Uhm, go on.” In some embodiments, the response is generated from a small, instant language model that is coupled to the conversation system 110), wherein the digital assistant interactive interface provides a plurality of candidate command controls corresponding to the business scenario to which the speech input interface belongs ([0032] Upon receiving the speech input, the conversation system 110 is configured to convert the speech input into a text input. The conversation system 110 then generates a text prompt for the LLM system 120 based on the text input and sends the text prompt to the LLM system 120; [0053] if an interruption is detected shortly after a speech response starts to play, the response modification module 270 may be configured to generate a continuer response, such as “Uhm, go on.” In some embodiments, the response is generated from a small, instant language model that is coupled to the conversation system 110).
With regard to claim 8, the limitations are addressed above and Rossi teaches wherein the first control comprises a control for invoking a digital assistant interactive interface ([0052] The interface module 260 is configured to transmit text prompt generated by the prompt generation module 220 to the LLM system 120 and receive text response generated by the LLM system 120; [0053] if an interruption is detected shortly after a speech response starts to play, the response modification module 270 may be configured to generate a continuer response, such as “Uhm, go on.” In some embodiments, the response is generated from a small, instant language model that is coupled to the conversation system 110), and displaying the second text on the speech input interface in response to the operation associated with the first control ([abstract] the LLM generates a second text response based on the interruption, which is converted to speech and played back to the user) comprises:
displaying the digital assistant interactive interface in response to a trigger operation for the first control ([0072] the conversation system 110 also determines whether to roll back the conversation state based in part on a pause 442D between user's two speeches 412D and 416D, a pause 444D between the starting of the audio response 414D and the start of the user speech 416D, and/or a pause 446D between user's two speeches 416D and 420D. For example, if the pause 444D is less than a threshold time, the conversation system 110 determines that the pause is a short pause), wherein the digital assistant interactive interface is used to receive a content input by a user ([0006] A vocal interruption comprises a second speech input from the user. Upon detecting the vocal interruption during the play of the first speech response, the system causes the client device to stop playing the first speech response. The system converts the second speech input into a second text input, and generates a second prompt to the LLM based on the vocal interruption…The system transmits the second prompt to the LLM, causing the LLM to generate a second text response); and
displaying the second text on the speech input interface in response to an input operation on the digital assistant interactive interface ([0006] A vocal interruption comprises a second speech input from the user. Upon detecting the vocal interruption during the play of the first speech response, the system causes the client device to stop playing the first speech response. The system converts the second speech input into a second text input, and generates a second prompt to the LLM based on the vocal interruption…The system transmits the second prompt to the LLM, causing the LLM to generate a second text response), wherein the second text is obtained by processing the first text based on a processing process indicated by an input content on the digital assistant interactive interface ([abstract] Upon generating a text response by the LLM, the text response is then converted back into speech and played to the user. If the user interrupts while the response is being played, the playback stops, and the interruption is captured as a new spoken input. This interruption is used to generate a new prompt for the LLM. Subsequently, the LLM generates a second text response based on the interruption, which is converted to speech and played back to the user; [0006] Upon detecting the vocal interruption during the play of the first speech response, the system causes the client device to stop playing the first speech response. The system converts the second speech input into a second text input, and generates a second prompt to the LLM based on the vocal interruption), and the input content is described in natural language ([abstract] managing interruptions during oral interactions between users and Large Language Models (LLMs); [0073] FIG. 5 illustrates an example process 500 of training or retraining a machine learning model 520 in accordance with one or more embodiments. In some embodiments, the machine-learning model may be a language model configured to generate a continuer response (e.g., “Uhm, go on.”) upon detecting a particular type of interruption. In some embodiments, the machine-learning model is a large language model used by the LLM system 120).
With regard to claim 9, the limitations are addressed above and Rossi teaches wherein the first text is displayed in a first area of the speech input interface ([abstract] a user's spoken input is received and converted to text, which forms a prompt for the Large Language Models (LLMs); [0005] The system receives a first speech input from a client device of a user, and converts the first speech input into first text), and the method further comprises:
displaying the second text in the first area of the speech input interface in response to a replacement operation ([abstract] Upon generating a text response by the LLM, the text response is then converted back into speech and played to the user. If the user interrupts while the response is being played, the playback stops, and the interruption is captured as a new spoken input. This interruption is used to generate a new prompt for the LLM. Subsequently, the LLM generates a second text response based on the interruption, which is converted to speech and played back to the user; [0006] Upon detecting the vocal interruption during the play of the first speech response, the system causes the client device to stop playing the first speech response. The system converts the second speech input into a second text input, and generates a second prompt to the LLM based on the vocal interruption); or
displaying the first text and the second text in the first area of the speech input interface in response to an insertion operation.
With regard to claim 10, the limitations are addressed above and Rossi teaches further comprising:
displaying a selected first text on the speech input interface in a set display manner in response to a selection operation for the first text (Fig. 7; [abstract] a user's spoken input is received and converted to text, which forms a prompt for the Large Language Models (LLMs); [0005] The system receives a first speech input from a client device of a user, and converts the first speech input into first text); and
wherein displaying the second text on the speech input interface in response to the operation associated with the first control ([abstract] Upon generating a text response by the LLM, the text response is then converted back into speech and played to the user. If the user interrupts while the response is being played, the playback stops, and the interruption is captured as a new spoken input. This interruption is used to generate a new prompt for the LLM. Subsequently, the LLM generates a second text response based on the interruption, which is converted to speech and played back to the user; [0006] Upon detecting the vocal interruption during the play of the first speech response, the system causes the client device to stop playing the first speech response. The system converts the second speech input into a second text input, and generates a second prompt to the LLM based on the vocal interruption) comprises:
displaying the second text on the speech input interface in response to the operation associated with the first control ([abstract] Upon generating a text response by the LLM, the text response is then converted back into speech and played to the user. If the user interrupts while the response is being played, the playback stops, and the interruption is captured as a new spoken input. This interruption is used to generate a new prompt for the LLM. Subsequently, the LLM generates a second text response based on the interruption, which is converted to speech and played back to the user; [0006] Upon detecting the vocal interruption during the play of the first speech response, the system causes the client device to stop playing the first speech response. The system converts the second speech input into a second text input, and generates a second prompt to the LLM based on the vocal interruption), wherein the second text is obtained by processing the selected first text based on the processing process corresponding to the first control ([0094] the LLM generates and sends a second text response to the conversation system 110. The conversation system 110 receives 960 the second text response from the LLM and converts 965 the second text response into a second speech response. The conversation system 110 transmits 970 the second text response back to the client device 130, causing the client device 130 to play the second speech response to the user).
With regard to claim 11, the device claim corresponds to the method claim 1, respectively, and therefore is rejected with the same rationale.
With regard to claim 12, the device claim corresponds to the method claim 2, respectively, and therefore is rejected with the same rationale.
With regard to claim 13, the device claim corresponds to the method claim 3, respectively, and therefore is rejected with the same rationale.
With regard to claim 14, the device claim corresponds to the method claim 4, respectively, and therefore is rejected with the same rationale.
With regard to claim 16, the device claim corresponds to the method claim 6, respectively, and therefore is rejected with the same rationale.
With regard to claim 17, the device claim corresponds to the method claim 7, respectively, and therefore is rejected with the same rationale.
With regard to claim 19, the device claim corresponds to the method claim 9, respectively, and therefore is rejected with the same rationale.
With regard to claim 20, the medium claim corresponds to the method claim 1, respectively, and therefore is rejected with the same rationale.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 5 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Rossi et al. (U.S. 2024/0347058) in view of Naik et al. (U.S. 2017/0178619).
With regard to claim 5, the limitations are addressed above and Rossi teaches displaying the second text on the speech input interface in response to a trigger operation ([0072] the conversation system 110 also determines whether to roll back the conversation state based in part on a pause 442D between user's two speeches 412D and 416D, a pause 444D between the starting of the audio response 414D and the start of the user speech 416D, and/or a pause 446D between user's two speeches 416D and 420D. For example, if the pause 444D is less than a threshold time, the conversation system 110 determines that the pause is a short pause), wherein the second text is obtained by processing the first text based on a processing process ([0072] the conversation system 110 also determines whether to roll back the conversation state based in part on a pause 442D between user's two speeches 412D and 416D, a pause 444D between the starting of the audio response 414D and the start of the user speech 416D, and/or a pause 446D between user's two speeches 416D and 420D. For example, if the pause 444D is less than a threshold time, the conversation system 110 determines that the pause is a short pause); and wherein the method further comprises:
displaying an updated second text on the speech input interface in response to a switching operation ([0061] The listening module 310 is configured to listen to a user's speech input… When the listening module 310 detects the wake word, it switches to an active listening mode, ready to record a next spoken speech. The recording may be an audio file of the user's speech), wherein the updated second text is obtained by processing the first text based on a processing process ([abstract] This interruption is used to generate a new prompt for the LLM. Subsequently, the LLM generates a second text response based on the interruption, which is converted to speech and played back to the user; [0006] Upon detecting the vocal interruption during the play of the first speech response, the system causes the client device to stop playing the first speech response. The system converts the second speech input into a second text input, and generates a second prompt to the LLM based on the vocal interruption). However, Rossi does not specifically teach:
- a first sub-control and a second sub-control
- wherein the second text is obtained by processing the first text based on a processing process corresponding to the first sub-control;
Naik teaches a system and method of digital assistants that make use of user-specified pronunciations of words for speech synthesis and recognition [0002]. Naik teaches a first sub-control ([0053] the digital assistant client module 264 provides the context information or a subset thereof with the user input to the digital assistant server to help infer the user's intent; [0067] the digital assistant module 326 includes the following sub-modules, or a subset or superset thereof: an input/output processing module 328, a speech-to-text (STT) processing module 330, a phonetic alphabet conversion module 331, a natural language processing module 332, a dialogue flow processing module 334, a task flow processing module 336, a service processing module 338, and a speech interaction error detection module 339) and a second sub-control ([0067] Each of these modules has access to one or more of the following data and models of the digital assistant 326, or a subset or superset thereof: ontology 360, vocabulary index 344, user data 348, task flow models 354, and service models 356; [0084] Property nodes “restaurant,” “date/time” (for the reservation), and “party size” are each directly linked to the actionable intent node (i.e., the “restaurant reservation” node). In addition, property nodes “cuisine,” “price range,” “phone number,” and “location” are sub-nodes of the property node “restaurant,” and are each linked to the “restaurant reservation” node (i.e., the actionable intent node) through the intermediate property node “restaurant.”; [0085] sub-property nodes “cuisine,” “price range,” “phone number,” and “location.”). Naik goes on to say that wherein the second text is obtained by processing the first text based on a processing process corresponding to the first sub-control ([0053] the digital assistant client module 264 provides the context information or a subset thereof with the user input to the digital assistant server to help infer the user's intent; [0067] the digital assistant module 326 includes the following sub-modules, or a subset or superset thereof: an input/output processing module 328, a speech-to-text (STT) processing module 330, a phonetic alphabet conversion module 331, a natural language processing module 332, a dialogue flow processing module 334, a task flow processing module 336, a service processing module 338, and a speech interaction error detection module 339). Therefore, it would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which said subject matter pertains to have modified the system which converts first speech input into first text taught by Rossi, with the sub-controls for words for speech synthesis and recognition as taught by Naik, to have achieved a system and method for managing interruptions during verbal conversations between users and large language models (LLMs).
With regard to claim 15, the device claim corresponds to the method claim 5, respectively, and therefore is rejected with the same rationale.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Lee (US Patent No. 12,586,587) teaches method for analyzing a user utterance and processing the user utterance based on an utterance cache and an electronic device performing the method.
Yeri et al. (US 2026/0065034) teaches a method may include generating a test set of prompts; executing a generative artificial intelligence (GenAI) machine learning model using the test set of prompts.
Cohen (US 2025/0080480) teaches a message edit and reply management in a quote-reply messaging system.
Cohen (US 2023/0359812) teaches a system and method for allowing users of electronic devices to populate fields of form displayed using voice input and providing text to user device to populate first element with text.
Choi et al. (US 2022/0310096) teaches a device for recognizing speech input of user and operating in which the speech input is to be converted by recognizing a speech input using an automatic speech recognition (ASR) model.
Li (US Patent No. 11,238,860) teaches a system and method for implementing speech control between a first keyword text and a second keyword text.
Hu et al. (US 2021/0124805) teaches a natural language processor (120) which is configured to process the text-based representation.
Vendrow et al. (US 2015/0222572) teaches a conversation timeline for heterogeneous messaging system.
Yu et al. (US 2010/0125795) teaches a system for concatenating audio and video clips.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ANDREA C. LEGGETT whose telephone number is (571)270-7700. The examiner can normally be reached M-F 9am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kieu Vu can be reached at 571-272-4057. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ANDREA C LEGGETT/Primary Examiner, Art Unit 2171