Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Drawings
The drawings were received on 11/20/2024. These drawings are accepted.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-7 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea in the form of mental process without significantly more. The claim(s) recite(s) a conversation interaction between a user and AI assistant, determining an interruption, tracking the point of interruption and storing such information to use for resuming the conversation. Such limitations are directed towards actions performed by a human mentally during a conversation. The claim also recites generic devices performing the abstract idea such as “one or more computer processors”, “one or more computer memories”, “a set of instructions …” and “AI assistant”. This judicial exception is not integrated into a practical application because the claimed language is merely directed towards the abstract idea without positively recited limitations integrating the abstract idea into practical application. The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the claimed language is merely directed towards the judicial exception without positively recited language indicating significantly more than the judicial exception.
Claims 2-6 recites language adding to the abstract idea, but does not include positively recited language indicating significantly more and/or integrating the abstract idea into practical application.
Claim 7 recites language adding to the abstract idea with addition of generic device, such as “using machine learning”, performing the abstract idea, but does not include positively recited language indicating significantly more and/or integrating the abstract idea into practical application.
Claims 8-14 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea in the form of mental process without significantly more. The claim(s) recite(s) a conversation interaction between a user and AI assistant, determining an interruption, tracking the point of interruption and storing such information to use for resuming the conversation. Such limitations are directed towards actions performed by a human mentally during a conversation. This judicial exception is not integrated into a practical application because the claimed language is merely directed towards the abstract idea without positively recited limitations integrating the abstract idea into practical application. The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the claimed language is merely directed towards the judicial exception without positively recited language indicating significantly more than the judicial exception.
Claims 8-13 recites language adding to the abstract idea, but does not include positively recited language indicating significantly more and/or integrating the abstract idea into practical application.
Claim 14 recites language adding to the abstract idea with addition of generic device, such as “using machine learning”, performing the abstract idea, but does not include positively recited language indicating significantly more and/or integrating the abstract idea into practical application.
Claims 15-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea in the form of mental process without significantly more. The claim(s) recite(s) a conversation interaction between a user and AI assistant, determining an interruption, tracking the point of interruption and storing such information to use for resuming the conversation. Such limitations are directed towards actions performed by a human mentally during a conversation. The claim also recites generic devices performing the abstract idea such as “non-transitory computer readable storage medium …”. This judicial exception is not integrated into a practical application because the claimed language is merely directed towards the abstract idea without positively recited limitations integrating the abstract idea into practical application. The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the claimed language is merely directed towards the judicial exception without positively recited language indicating significantly more than the judicial exception.
Claims 16-20 recites language adding to the abstract idea, but does not include positively recited language indicating significantly more and/or integrating the abstract idea into practical application.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 1 recites the limitation “an intended interruption” in “in response to determining the audio input represents …”, "the interrupted response" in “maintaining context awareness …”, “the point of interruption” in in “maintaining context awareness …” of claim 1. There is insufficient antecedent basis for this limitation in the claim.
Claim 4 recites the limitation "the converted text" in claim 1. There is insufficient antecedent basis for this limitation in the claim.
Claim 7 recites the limitation "the user profile" in claim 1,6. There is insufficient antecedent basis for this limitation in the claim.
Claim 8 recites the limitation “an intended interruption” in “in response to determining the audio input represents …”, "the interrupted response" in “maintaining context awareness …”, “the point of interruption” in in “maintaining context awareness …” of claim 8. There is insufficient antecedent basis for this limitation in the claim.
Claim 11 recites the limitation "the converted text" in claim 8. There is insufficient antecedent basis for this limitation in the claim.
Claim 14 recites the limitation "the user profile" in claim 8,13. There is insufficient antecedent basis for this limitation in the claim.
Claim 15 recites the limitation “an intended interruption” in “in response to determining the audio input represents …”, "the interrupted response" in “maintaining context awareness …”, “the point of interruption” in in “maintaining context awareness …” of claim 15. There is insufficient antecedent basis for this limitation in the claim.
Claim 18 recites the limitation "the converted text" in claim 15. There is insufficient antecedent basis for this limitation in the claim.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim(s) 1,3-5,8,10-12,15,17-19 is/are rejected under 35 U.S.C. 102a1 as being anticipated by Rossi et al (US Publication No.: 20240347058).
Claim 1, Rossi et al discloses
one or more computer processors (paragraph 100);
one or more computer memories (paragraph 100);
a set of instructions stored in the one or more computer memories (paragraph 102), the set of instructions configuring the one or more computer processors to perform operations (paragraph 102), the operations comprising:
receiving audio input (Fig. 4d, label 410d,412d,414d,416d as conversation between user and bot. Paragraph 5 discloses verbal conversations between users and AI bot (LLMs).) during an artificial intelligence (AI) voice interaction between a user and an AI assistant (Paragraph 5 discloses verbal conversations between users and AI bot (LLMs).);
analyzing the audio input in real-time to determine whether the audio input represents an intended interruption of the AI assistant's speech (Fig. 4d shows the timeline in which the audio input is analyzed (real time). Label 430D indicates a detection of interruption of the AI assistant’s speech by the human (label 416D).);
in response to determining the audio input represents an intended interruption (Fig. 4d, label 430D indicates a detection of interruption with possible responses.), stopping the AI assistant's speech and tracking what portion of a response was actually spoken (Fig. 7,8 and Paragraph 82 discloses an example where the user interrupts the AI’ response, and the AI stops the response. A text record including metadata associated with the user’s speech and response, including timing and duration of each speech request or speech response. Paragraph 51 discloses “the conversation state module 250 may also log metadata related to conversation states or interruptions, and cause the prompt generation module 220 to incorporate the metadata into prompts.”); and
maintaining context awareness by storing information about the interrupted response to allow resuming from the point of interruption (Fig. 7 and paragraph 82 discloses cancellation the response or answer and a new response with the new information from the interruption is generated. The response is generated from the point of interruption. Fig. 8, label 820 indicates a detected vocal interruption, 803 response is stopped, 840 a new response is generated based on vocal interruption and 850 the new response is played.).
Claim 3, Rossi et al discloses wherein the analyzing the audio input (Fig. 4d, label 430d) comprises:
detecting whether the audio input represents an affirmative acknowledgment rather than an intended interruption (Fig. 4d, label 430d indicates detected interrupt, label short pause, 452d as the response to interrupt: roll back state, wherein an affirmative acknowledge is interpreted as a short response.); and
continuing the AI assistant's speech without interruption in response to detecting an affirmative acknowledgment (Fig. 4d, label 452D, rollback state where label 424d as a bot represents periods during which an audio response generated by the conversation system is being played. This indicates a previous state for roll back.).
Claim 4, Rossi et al discloses wherein the analyzing the audio input (Fig. 4d, label 430d) comprises:
processing the audio input using an on-premise processor to perform initial voice-to-text conversion locally (Fig. 2, label 210);
performing preprocessing of the converted text before communicating with a language model (Fig. 2, label 220 preprocesses the text output by label 220. Paragraph 45 discloses 220 cleans the original text generated by the TTS such as removing irrelevant words from the text.); and
determining interrupt intent based on the preprocessing (Paragraph 45 discloses “prompt generation module 220 is configured to add context or additional instructions that help the LLM system understand the prompt’s intent, such as metadata associated with a current state of a conversation or a vocal interruption.” By determining the current state of a conversation or vocal interruption, a determination of the intent occurs based on preprocessed text from the TTS.).
Claim 5, Rossi et al discloses wherein the analyzing the audio input comprises: analyzing sentiment through voice characteristics including volume, tone, or speaking rate (Paragraph 57 discloses machine learning model may be trained to understand the flow of a conversation, recognize when an interruption occurs and determine the state of the conversation at the moment. … In some embodiments, the machine learning model 296 is trained to dynamically adjust the conversation state based on a determined type of interruption and generate a response appropriate to the type of interruption and the new state of the conversation”. Paragraph 56 discloses 296 is trained using collected data such as examples of interruptions and their corresponding metadata about their states and timing, wherein such examples includes audio features such as speech rate, volume and pitch.); and
adjusting interruption sensitivity based on the analyzed sentiment (Fig. 4d, label 434d indicates the recognition of emotions or sentiments of the user during an interruption. Label 436d indicates the use of the emotions or sentiments as training data to train label 296. As per paragraph 57, model 296 is is trained to recognize when an interruption occurs (interruption sensitivity), determine type of interruption and response to the type of interruption. This indicates the emotions or sentiments, determined at label 434d and used for training of 296 as per label 436d and paragraph 71, trains the model to detect interruption, hence adjusting the interruption sensitivity.).
Claim 8 recites similar limitations in claim 1 and is rejected on the same grounds as claim 1.
Claim 10 recites similar limitations in claim 3 and is rejected on the same grounds as claim 3.
Claim 11 recites similar limitations in claim 4 and is rejected on the same grounds as claim 4.
Claim 12 recites similar limitations in claim 5 and is rejected on the same grounds as claim 5.
Claim 15 recites similar limitations in claim 1 and is rejected on the same grounds as claim 1.
Claim 17 recites similar limitations in claim 3 and is rejected on the same grounds as claim 3.
Claim 18 recites similar limitations in claim 4 and is rejected on the same grounds as claim 4.
Claim 19 recites similar limitations in claim 5 and is rejected on the same grounds as claim 5.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 2,9,16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Rossi et al (US Publication No.: 20240347058) in view of Al Muntasir et al (JP 2025065586).
Claim 2, Rossi et al discloses wherein the analyzing the audio input (Fig. 4d), but fails to disclose accessing a user profile containing historical interaction patterns for the user; and determining whether the audio input matches known vocal patterns associated with intended interruptions for the user based on the historical interaction patterns.
Al Muntasir et al discloses
accessing a user profile containing historical interaction patterns for the user (paragraph 90 discloses “ASR and NLU models related to registered user profiles …”.); and
determining whether the audio input matches known vocal patterns associated with intended interruptions for the user based on the historical interaction patterns (Paragraph 90 discloses “adding information related to user interaction sessions involves updating and training ASR and NLU models related to registered user profiles using voice data collected from corresponding user interaction sessions and utilizing features of voice data collected from user speech sessions. Updating and training models related to determining speech segments … and interrupt speech segments associated with the registered user profile.”). It would be obvious to one skilled in the art before the effective filing date of the application to modify Rossi et al’s interruption detection by incorporating audio features of a user profile to detect interruption in a conversation as disclosed by Al Muntasir et al so to improve the system’s ability to detect interruptions by the user, and improving the conversation between AI and user by accurately detecting interruption, hence improving the user’s experience with conversation with virtual characters.
Claim 9 recites similar limitations in claim 2 and is rejected on the same grounds as claim 2.
Claim 16 recites similar limitations in claim 2 and is rejected on the same grounds as claim 2.
Claim(s) 6,7,13-14,20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Rossi et al (US Publication No.: 20240347058) in view of Kennedy et al (US Publication No.: 20250225986).
Claim 6, Rossi et al discloses detecting an interruption (Fig. 4d) including categorizing sounds in the audio input (paragraph 56-57 discloses training a model 296 with examples interruptions and metadata include acoustic features, where the model detects interruptions and type of interruptions.), but fails to disclose wherein the analyzing the audio input comprises: categorizing sounds in the audio input as either meaningful interruptions or non-interruptive vocal ticks; and continuing the AI assistant's speech without interruption in response to detecting a non-interruptive vocal tick.
Kennedy et al discloses
wherein the analyzing the audio input comprises: categorizing sounds in the audio input as either meaningful interruptions or non-interruptive vocal ticks (Paragraph 13 discloses classifying speech into interruption (meaningful interruptions) vs non-interruption inputs (non-interruptive vocal ticks).); and
continuing the AI assistant's speech without interruption in response to detecting a non-interruptive vocal tick (Paragraph 16 discloses re-rendering the remainder of the speech output by the AI character when the conversation turn is to be retained. Paragraph 14 discloses an interruption response such as re-render portions of the speech output by the AI.).
It would be obvious to one skilled in the art before the effective filing date of the application to modify Rossi et al’s conversation interruption response by incorporating a response continuing the AI”s response when the interruption occurred as disclosed by Kennedy et al so to improve the conversation flow between the user and AI, hence improving the user’s experience with a virtual character.
Claim 7, Rossi et al discloses wherein the categorizing of the sounds (paragraph 56-57 discloses training a model 296 with examples interruptions and metadata include acoustic features, where the model detects interruptions and type of interruptions.) comprises:
updating the user profile with specific vocal patterns specific to the user over time by tracking speaking habits and common vocal expressions (Fig. 4d, label 434d,436d indicates updating of training data such as examples of varying changes in a user’s tone.);
identifying whether detected sounds match known vocal tick patterns in the user profile (Paragraph 56 discloses training the model 296 using examples of interruptions and corresponding metadata, wherein the metadata includes audio features such as any change in the user’s tone (vocal tick patterns).); and
using machine learning to detect whether sounds indicate acknowledgment or interruption intent (paragraph 57 discloses training the model 296 to detect interruptions and type of interruption.).
Claim 13 recites similar limitations in claim 6 and is rejected on the same grounds as claim 6.
Claim 14 recites similar limitations in claim 7 and is rejected on the same grounds as claim 7.
Claim 20 recites similar limitations in claim 6 and is rejected on the same grounds as claim 6.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to LINDA WONG whose telephone number is (571)272-6044. The examiner can normally be reached 9-5.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew C Flanders can be reached at 571-272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/LINDA WONG/Primary Examiner, Art Unit 2655