Prosecution Insights
Last updated: August 15, 2026
Application No. 18/511,796

METHOD AND SYSTEM FOR INTEGRATED MULTIMODAL INPUT PROCESSING FOR VIRTUAL AGENTS

Non-Final OA §101§102§103
Filed
Nov 16, 2023
Examiner
THOMPSON, ALMA BENNETT
Art Unit
Tech Center
Assignee
Quantiphi Inc.
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Office Action

§101 §102 §103
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-19 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. The claimed invention is directed to an abstract idea without significantly more. Regarding claim 1 and analogous claims 10 and 19: For step 2A prong 1, the following limitations recite a mental process: identifying a plurality of principal entities within the multimodal input (observation) extracting information about each entity of the plurality of principle entities (observation) generating a response based on the extracted information (evaluation) For step 2A prong 2 and step 2B, the following limitations fail to integrate into a practical application or recite significantly more than the judicial exception. A computer-implemented method for multimodal input processing for a virtual agent (well-understood, routine, and conventional generic computer, see MPEP 2106.05(f)) A computer system multimodal input processing for a virtual agent comprising, the computer system comprising: one or more computer processors, one or more computer readable memories, one or more computer readable storage devices, and program instructions stored on the one or more computer readable storage devices for execution by the one or more computer processors via the one or more computer readable memories (well-understood, routine, and conventional generic computer, see MPEP 2106.05(f)) A non-transitory computer-readable storage medium having stored thereon computer executable instruction which when executed by one or more processors, cause the one or more processors to carry out operations for multimodal input processing for a virtual agent (well-understood, routine, and conventional generic computer, see MPEP 2106.05(f)) obtaining a multimodal input by the virtual agent from a user (insignificant extra-solution activity of mere data gathering, see MPEP 2106.05(g)) wherein the virtual agent employs an Artificial Intelligence model (well-understood, routine, and conventional generic computer, see MPEP 2106.05(f)) Regarding claim 2 and analogous claim 11: For step 2A prong 1, the following limitations recite a mental process: The computer-implemented method of claim 1 For step 2A prong 2 and step 2B, the following limitations fail to integrate into a practical application or recite significantly more than the judicial exception: wherein the AI model is a generative AI model (well-understood, routine, and conventional generic computer, see MPEP 2106.05(f)) Regarding claim 3 and analogous claim 12: For step 2A prong 1, the following limitations recite a mental process: The computer-implemented method of claim 1 For step 2A prong 2 and step 2B, the following limitations fail to integrate into a practical application or recite significantly more than the judicial exception: further comprising storing the extracted information within an associated database in each cycle of input processing (well-understood, routine, and conventional generic computer, see MPEP 2106.05(f)) Regarding claim 4 and analogous claim 13: For step 2A prong 1, the following limitations recite a mental process: The computer-implemented method of claim 1 For step 2A prong 2 and step 2B, the following limitations fail to integrate into a practical application or recite significantly more than the judicial exception: wherein the multimodal data comprises data from modalities comprising sensors, ensembled data, speech, text, vision (well-understood, routine, and conventional generic computer, see MPEP 2106.05(f)) Regarding claim 5 and analogous claim 14: For step 2A prong 1, the following limitations recite a mental process: The computer-implemented method of claim 1 further comprising dynamically adapting the users accustomed communication style based on historical interactions of the virtual agent with the user (evaluation) Claims 5 and 14 do not recite significantly more or integrate into a practical application. Regarding claim 6 and analogous claim 15: For step 2A prong 1, the following limitations recite a mental process: The computer-implemented method of claim 1 wherein the AI model employs a role-based approach following user-provided instructions and prompts (evaluation) Claims 6 and 15 do not recite significantly more or integrate into a practical application. Regarding claim 7 and analogous claim 16: For step 2A prong 1, the following limitations recite a mental process: The computer-implemented method of claim 1 For step 2A prong 2 and step 2B, the following limitations fail to integrate into a practical application or recite significantly more than the judicial exception: wherein the AI model is trained to understand and respond to user emotions conveyed through the multimodal input (well-understood, routine, and conventional generic computer, see MPEP 2106.05(f)) Regarding claim 8 and analogous claim 17: For step 2A prong 1, the following limitations recite a mental process: The computer-implemented method of claim 1 further comprising continuously monitoring user engagement and satisfaction during interactions (observation) Claims 8 and 17 do not recite significantly more or integrate into a practical application. Regarding claim 9 and analogous claim 18: For step 2A prong 1, the following limitations recite a mental process: The computer implemented method of claim 1 wherein the principal entities are selected from a group of a name, a date, a time, a numeric value, an address, a location, a sentiment, an emotional cue, a facial feature, a visual cue, a gesture, a body language, a parameter, an object, a command, and a keyword (evaluation) Claims 9 and 17 do not recite significantly more or integrate into a practical application. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claim(s) 1, 2, 4, 5, 7-11, 13, 14, and 16-19 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Maitra et al. (US Patent Publication No. US 2022/0230632 A1, hereafter referred to as Maitra). Regarding claim 1 and analogous claims 10 and 19: Maitra teaches A computer-implemented method for multimodal input processing for a virtual agent (In paragraph [0067] of Maitra, “FIG. 3 is a diagram of an example environment 300 in which systems and/or methods described herein may be implemented. As shown in FIG. 3, environment 300 may include a wellness system 301, which may include one or more elements of and/or may execute within a cloud computing system [computer-implemented] 302.” In paragraph [0019] of Maitra, “As shown in FIG. 1A, and by reference number 105, the wellness system may receive, from the user device, text data identifying text input by the user to the user device, audio data identifying audio associated with the user, and video data identifying a video associated with the user. The text data may include text input by the user via an input component (e.g., a keyboard) to the user device, text that is spoken by the user and provided to the user device via another input component (e.g., a microphone or a camera with a microphone), and/or the like. In some implementations, the wellness system performs natural language processing on the text that is spoken by the user in order to convert voice data of the user to textual data. The audio data may include audio captured by the other input component (e.g., the microphone or the camera with the microphone) of the user device. For example, the audio data may include data identifying a prosody, an intonation, a rhythm, a pitch, an intensity, a loudness, an energy, jitter, and/or the like associated with a voice of the user, background noise associated with the user device, and/or the like. The video data may include video captured by the other input component (e.g., the camera) of the user device. For example, the video data may include video and/or images of the user, visual features of the user (e.g., a yaw, a pitch, and/or roll angles associated with the user's head, an eye gaze of the user, an intensity of a contraction of a facial muscle of the user, and/or the like), video and/or images of a background associated with the user, and/or the like.” The system of Maitra gathers and processes a mixture of text input, audio input, and camera input; this falls under the broadest reasonable interpretation of multimodal input processing. In paragraph [0043] of Maitra, “For example, the wellness system may provide an empathetic, context-aware, and multimodal conversational agent [a virtual agent] to assist mental wellness of the user.”) comprising: obtaining a multimodal input by the virtual agent from a user (In paragraph [0019] of Maitra, “As shown in FIG. 1A, and by reference number 105, the wellness system [the virtual agent] may receive, from the user device [from a user], text data identifying text input by the user to the user device, audio data identifying audio associated with the user, and video data identifying a video associated with the user.” The input is multimodal, as discussed above.) wherein the virtual agent employs an Artificial Intelligence (AI) model; (In paragraph [0033] of Maitra, “As shown in FIG. 1D, and by reference number 130, the wellness system may process the text data, the audio data, and the video data, with a generative pretrained transformer (GPT2) language model [an Artificial Intelligence (AI model)], to determine a response to the user.”) identifying a plurality of principal entities within the multimodal input; (in paragraph [0030] of Maitra, “The wellness system may pre-process the multi-modal data [within the multimodal input], as described above, to extract faces and audio signals [identifying a plurality of principal entities]. The wellness system may utilize the faces for further visual feature extraction and may utilize the audio signals for further audio feature extraction.”) extracting information about each entity of the plurality of principal entities (in paragraph [0031] of Maitra,” The wellness system may extract the audio features and the visual features as described above. The wellness system may concatenate the audio features to form a feature vector and may combine frame-level audio features to form sequences for model training. With respect to visual feature extraction, the wellness system may determine face pose features, such as a head pose, an eye gaze, action unit intensities, and/or the like [extracting information about each entity of the plurality of principal entities]. The wellness system may combine the face pose features to form sequence level features.”) and generating a response based on the extracted information (in paragraph [0033] of Maitra, As shown in FIG. 1D, and by reference number 130, the wellness system may process the text data, the audio data, and the video data, [based on the extracted information] with a generative pretrained transformer (GPT2) language model, to determine a response [generating a response] to the user.”) Regarding claim 2 and analogous claim 11: Maitra teaches The computer-implemented method of claim 1 wherein the AI model is a generative AI model (in paragraph [0033] of Maitra, “As shown in FIG. 1D, and by reference number 130, the wellness system may process the text data, the audio data, and the video data, with a generative pretrained transformer (GPT2) language model [a generative AI model], to determine a response to the user.”) Regarding claim 4 and analogous claim 13: Maitra teaches The computer-implemented method of claim 1 wherein the multimodal input comprises data from modalities comprising sensors, ensembled data, speech, text and vision (in paragraph [0019] of Maitra, “As shown in FIG. 1A, and by reference number 105, the wellness system may receive, from the user device, text data identifying text input by the user to the user device, audio data identifying audio associated with the user, and video data identifying a video associated with the user. The text data may include text input [text] by the user via an input component (e.g., a keyboard) to the user device, text that is spoken by the user [speech] and provided to the user device via another input component (e.g., a microphone or a camera with a microphone), and/or the like. In some implementations, the wellness system performs natural language processing on the text that is spoken by the user in order to convert voice data of the user to textual data. The audio data may include audio captured by the other input component (e.g., the microphone or the camera with the microphone) of the user device. For example, the audio data may include data identifying a prosody, an intonation, a rhythm, a pitch, an intensity, a loudness, an energy, jitter, and/or the like associated with a voice of the user, background noise associated with the user device, and/or the like. The video data [vision] may include video captured by the other input component (e.g., the camera) of the user device. For example, the video data may include video and/or images of the user, visual features of the user (e.g., a yaw, a pitch, and/or roll angles associated with the user's head, an eye gaze of the user, an intensity of a contraction of a facial muscle of the user, and/or the like), video and/or images of a background associated with the user, and/or the like.” in paragraph [0078] of Maitra, “Input component 450 enables device 400 to receive input, such as user input and/or sensed inputs [the multimodal input]. For example, input component 450 may include a touch screen, a keyboard, a keypad, a mouse, a button, a microphone, a switch, a sensor [sensors], a global positioning system component, an accelerometer, a gyroscope, an actuator, and/or the like.” Ensembled data is not defined in the specification, but it can reasonably be interpreted to mean several different sources of data being combined, which is taught in paragraph [0019] of Maitra.) Regarding claim 5 and analogous claim 14: Maitra teaches The computer-implemented method of claim 1 further comprising dynamically adapting user’s accustomed communication style (in paragraph [0048] of Maitra, “In some implementations, the one or more actions include the wellness system retraining [dynamically adapting user’s accustomed communications style] one or more of the support vector machine model, the different regression models, the deep learning convolutional neural network model, the classifier model, the generative pretrained transformer language model, the plug and play language model, or the dialog manager models based on the contextual conversation data.” In paragraph [0049] of Maitra, “In some implementations, the one or more actions include the wellness system conversing with a user [dynamically adapting user’s accustomed communications style]. For example, the wellness system may act as a friend or a companion, who tries to keep the user in good spirits, tries to guide the user to good practices” based on historical interactions of the virtual agent with the user (in paragraph [0043] of Maitra, “As shown in FIG. 1F, and by reference number 145, the wellness system may perform one or more actions based on the contextual conversation data [historical interactions of the virtual agent with the user.” Regarding claim 7 and analogous claim 16: Maitra teaches The computer-implemented method of claim 1 wherein the AI model is trained to understand and respond to user emotions conveyed through the multimodal input (in paragraph [0033] of Maitra, “The wellness system may generate an empathetic response via a sentiment head by passing the hidden state of the last token through a linear layer, by applying a softmax model to the hidden state of the last token to obtain an emotion class, and by applying a cross-entropy loss to train the sentiment head to classify emotion correctly. The wellness system may further generate a more empathetic response [respond to user emotions] by adding an emotion token in every sentence into the GPT2 language model and by causing the GPT2 language model to learn an association [understand user emotions] between the emotion token and relevant emotionally colored words.” Regarding claim 8 and analogous claim 17: Maitra teaches The computer-implemented method of claim 1 further comprising continuously monitoring user engagement and satisfaction during interactions (in paragraph [0017] of Maitra, “In some implementations, the wellness system may be utilized to identify an affect in general on the user [continuously monitoring during interactions]. An affect may include pain, haste, an engagement level [user engagement], despair, longing, fondness [satisfaction], and/or the like. In such implementations, the wellness system may remain unaltered, but the models described herein may need different sets of training data to be applicable to a corresponding end application.”) Regarding claim 9 and analogous claim 18: Maitra teaches The computer-implemented method of claim 1 wherein the principal entities are selected from a group of a name, a date, a time, a numeric value, an address, a location, a sentiment, an emotional cue, a facial feature, a visual cue, a gesture, a body language, a parameter, an object, a command, and a keyword (in paragraph [0031] of Maitra,” The wellness system may extract the audio features and the visual features as described above. The wellness system may concatenate the audio features to form a feature vector and may combine frame-level audio features to form sequences for model training. With respect to visual feature extraction, the wellness system may determine face pose features [a body language], such as a head pose, an eye gaze, action unit intensities, and/or the like. The wellness system may combine the face pose features to form sequence level features”). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 3 is rejected under 35 U.S.C. 103 as being unpatentable over Maitra in view of Fan et al. (Fan Y, Bowden KK, Cui W, Chen W, Harrison V, Ramirez A, Agashe S, Liu XG, Pullabhotla N, Bheemanpally NQ, Garg S. Athena 3.0: Personalized multimodal chatbot with neuro-symbolic dialogue generators. Alexa Prize Soc Bot Grand Challenge 5. September 2023. Hereafter referred to as Fan). Maitra teaches: The computer-implemented method of claim 1 Maitra fails to teach further comprising storing the extracted information within an associated database in each cycle of input processing Fan teaches further comprising storing the extracted information within an associated database in each cycle of input processing (On page 3 of Fan, “This year’s focus on the discourse model requires changes to the Response Generators (RGs), which are responsible for creating the system response detailed in Section 5. When the RGs introduce an entity [the extracted information] in the system utterance [each cycle of input processing], the entity and its knowledge graph ID are recorded in the Discourse model [storing]”. On page 2 of Fan, “The inputs to Athena are the ASR hypotheses for a user’s turn, as well as a conversation ID that is used to retrieve the conversation history and state information from a back-end database [an associated database]”) Fan and Maitra are both related to the same field of endeavor (i.e. interactive chatbots). In view of the teachings of Fan, it would have been obvious for a person of ordinary skill in the art before the effective filing date of the claimed invention to apply the teachings of Fan to Maitra so that the virtual agent could keep conversations “on-topic” over long periods of time (on page 3 of Fan, “Further, we build a novel history-aware topic classifier to reduce the chance of outputting off-topic responses”). Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Maitra in view of Xu et al. (Xu B, Yang A, Lin J, Wang Q, Zhou C, Zhang Y, Mao Z. Expertprompting: Instructing large language models to be distinguished experts. arXiv preprint arXiv:2305.14688. 2023 May 24. Hereafter referred to as Xu.) Maitra teaches: The computer-implemented method of claim 1 Maitra fails to teach wherein the AI model employs a role-based approach following user-provided instructions and prompts Xu teaches wherein the AI model employs a role-based approach following user-provided instructions and prompts (on page 1 of Xu, "For each specific instruction, ExpertPrompting [the AI model] first envisions a distinguished expert agent that is best suited for the instruction [following user-provided instructions and prompts], and then asks the LLMs to answer the instruction conditioned on such expert identity [role-based approach].") Maitra and Xu are both related to the same field of endeavor (i.e. interactive chatbots). In view of the teachings of Xu, it would have been obvious for a person of ordinary skill in the art before the effective filing date of the claimed invention to apply the teachings of Xu to Maitra in order to generate higher responses from the virtual agent (on page 4 of Xu, “We then randomly sample 500 instructions, and compare these answers using GPT-4 based evaluation. Results in Figure 3 show that ExpertPrompting answers are preferred at 48.5% by the reviewer model, compare to 23% of the vanilla answer, which demonstrates clear superiority.”) Conclusion Any inquiry concerning this communication or earlier communication from the examiner should be directed to Alma Thompson whose telephone number is +1 (571) 270-1810. The examiner can normally be reached Monday-Friday, 9:00 am – 5:00 pm EST. If attempts to reach the examiner are unsuccessful, the examiner’s supervisor, Michael Huntley can be reached at +1 (303) 297-4307. Should applicant desire to communicate via email or schedule an Applicant-initiated interview, the Applicant may use the Automated Interview Request (AIR) form online at <http://www.uspto.gov/patent/uspto-automated-interview-request-air-form> In particular, the AIR form allows the Applicant to include authorization to communicate via email with the examiner for interview-related communications. Information regarding the status of an application may be obtained from the Patent Application Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pairdirect.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at (866) 217-9197 (toll-free). If you would like assistance from a USPTO Customer Service representative or access to the automated information system, call (800) 786-9199 (IN USA OR CANADA) or (571) 272-1000. /MICHAEL J HUNTLEY/Supervisory Patent Examiner, Art Unit 2129
Read full office action

Prosecution Timeline

Nov 16, 2023
Application Filed
Aug 05, 2026
Non-Final Rejection mailed — §101, §102, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month