Prosecution Insights
Last updated: August 06, 2026
Application No. 18/438,677

System and Method for Automated Digital Twin Behavior Modeling for Multimodal Conversations

Final Rejection §101§103
Filed
Feb 12, 2024
Priority
Sep 24, 2021 — divisional of 11/942,075
Examiner
MASTERS, KRISTEN MICHELLE
Art Unit
2659
Tech Center
2600 — Communications
Assignee
Openstream Inc.
OA Round
2 (Final)
65%
Grant Probability
Favorable
3-4
OA Rounds
6m
Est. Remaining
86%
With Interview

Examiner Intelligence

Grants 65% — above average
65%
Career Allowance Rate
32 granted / 49 resolved
+3.3% vs TC avg
Strong +21% interview lift
Without
With
+21.2%
Interview Lift
resolved cases with interview
Typical timeline
3y 0m
Avg Prosecution
23 currently pending
Career history
84
Total Applications
across all art units

Statute-Specific Performance

§101
37.6%
-2.4% vs TC avg
§103
48.3%
+8.3% vs TC avg
§102
7.8%
-32.2% vs TC avg
§112
3.5%
-36.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 49 resolved cases

Office Action

§101 §103
Detailed Action This communication is in response to the Arguments and Amendments filed on 4/14/2026. Claims 1-20 are pending and have been examined. Claims 1-20 are rejected. Hence, this action has been made Final. Independent Claims 1, 11 are device and method claims, respectively. Apparent priority: 9/24/2021. Any previous objection/rejection not mentioned in this Office Action has been withdrawn by the Examiner. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendments and Arguments The Applicant has amended the claims to include “audio data and image data representing” “audio data and the image data” “extracting audio features from the audio data and visual features from the image data; generate a multimodal feature representation based on the audio features and the visual features;” “generate, via the digital twin platform based on the multimodal feature representation, a learned digital twin behavior model by transferring” “ Rejections under 35 U.S.C. §101 Applicant notes the Office asserts that claim 1 is directed to a mental process at Step 2A, Prong One. (Office Action, pp. 2-3). MPEP § 2106.04(a)(2), subsection III(A) is titled "A Claim With Limitation(s) That Cannot Practically be Performed in the Human Mind Does Not Recite a Mental Process." Claim 1 as amended recites elements that cannot practically be performed in the human mind. Specifically, claim 1 as amended recites "extracting audio features from the audio data and visual features from the image data" and generating "a multimodal feature representation based on the audio features and the visual features." Applicant notes the human mind, even with the aid of a pen and paper, is incapable of extracting audio features and visual features and using the audio features and visual features to generate a multimodal feature representation. Examiner notes the features of claim 1 are directed to an abstract idea. A human can naturally extract audio and visual features. In the case of visual feature extraction, the retina performs preprocessing, converting light into signals, the visual cortex detects features, objects become recognizable shapes, a human can draw those shapes or convey them through speech in an organized way, and then categorize and conceptualize those shapes into an index. Applicant notes Moreover, claim 1 as amended recites that the processor is further configured to execute instructions to "generate, via the digital twin platform based on the multimodal feature representation, a learned digital twin behavior model by transferring knowledge through the social simulations to the virtual humans, wherein social and functional behavior of humans is encoded in the learned digital twin behavior model." The human mind is incapable of generating a learned digital twin model by transferring knowledge from social simulations to virtual humans. Examiner notes humans naturally transfer knowledge they acquire from social simulations to other humans. Step 2A, Prong Two When evaluating claim 1 as a whole at Step 2A, Prong Two, any judicial exception as may be identified for claim 1 under Step 2A, Prong One of the Alice/Mayo test (noting that Applicant respectfully does not concede that such a judicial exception exists) is integrated into a practical application at least in that claim 1 describes an improvement to the functioning of a computer. Specifically, Applicant explains in paragraph [0020] that "digital twins of human beings have been fraught with many problems, such as lip-syncing, syncing of facial expressions, hand- eye coordination, smart body movements and motion, and the like, as there is no bidirectional communication between the avatar and the human." To address these problems, Applicant teaches a system that "analyzes the multimodal interactions, determines the need of a human in the loop, and replaces the virtual human seamlessly during the conversation without the user noticing the substitution." (1[0022]). Accordingly, claim 1 as a whole integrates any judicial exception into a practical application, i.e., an improved multimodal conversational system that "enables seamless transitions between a human agent and a conversational agent, where the term conversational agent can be interchanged with or refer to a virtual agent, a trained virtual human agent, a trained virtual human clone, or a virtual twin." (1[0017]). Examiner notes on the present claim wording, the limitations are largely functional and outcome oriented (receive data, recognize content, extract features , train , process social situations, generate model, generate responses) without concrete computational detail or a recitation of how the arrangements materially improve the functioning of the computer system itself (e.g., speed/latency reductions, memory or computational efficiency, novel data representations that reduce error by a measurable metric, or specific unconventional network architectures constrained in a way that produces the improvement). Step 2B Further, when evaluating claim 1 at Step 2B, claim 1 as a whole amounts to significantly more than any such judicial exception based at least upon the inventive concept recited therein. MPEP 2106.05 explains that consideration "of the elements in combination is particularly important, because even if an additional element does not amount to significantly more on its own, it can still amount to significantly more when considered in combination with the other elements of the claim." Here, the combination of receiving audio data and image data, parsing the audio data and the image data, extracting audio features from the audio data and visual features from the image data, generating a multimodal feature representation based on the audio features and the visual features, process social simulations which model behavior of humans and virtual humans, and generating a learned digital twin behavior model by transferring knowledge through the social simulations to the virtual humans provides a technological improvement to a multimodal conversational system. Accordingly, claim 1 amounts to significantly more based at least upon the inventive concept recited therein. Examiner notes the additional elements must supply an “inventive concept.” The claim recites known functional components (storage, processor) and high-level data transformations. Without claim specificity tying those components to particular unconventional architectures, constrained parameterizations, training/regimen steps, or demonstrable improvements, the recited elements appear to be routine, conventional uses of generic software components and models and therefore fail to supply an inventive concept. Rejections under 35 U.S.C. §103 Applicant notes the Office asserts that Leeds teaches that "each pair of a human and a virtual human form digital twins," as recited in claim 1. (Office Action, pp. 8-9). Leeds discloses a system of modular conversational agents ("subminds") that participate in collaborative dialogue, but these are task-oriented agents that are not tied to any specific human counterpart and do not form digital twins. The Office has not identified, within Leeds or elsewhere, any disclosure of a one- to-one twin pairing. Examiner notes Leeds establishes (6:2-10) (42) “Theory of Mind” or ToM is the modeling of one mind by another. It is used to describe the ability to attribute mental states to others, in the context of the present invention, for social interaction and collaboration in particular. ToM can be a model of another entity (which could be human or AI) held by a submind that captures that submind's understanding of the other entity. Examiner notes it is known in the prior art that we have been modeling virtual humans for interacting with humans. In Applicants invention the virtual human is interacting with a virtual assistant. This is also taught by Leeds. Examiner further notes a Digital twin is known in the field of art of Gaming. In Games, every person has an avatar that mimics the person. So both user and agent are Digital Twins. Examiner further notes you can generate a duplicate/twin/avatar/clone of a human by using a database of speech and facial gestures. Applicants’ invention adds a step, a method of data augmentation. Normally: we take the corpus of data of human behavior and train the automated agent. Applicant takes the same data; first generates a twin; then trains the agent. Applicant notes The Office asserts that Lok teaches that social and functional behavior of humans is encoded in the learned digital twin behavior model. (Office Action, p. 10). Lok discloses after-action review, logging, and visualization of human-virtual human interactions, including analysis of social, temporal, and spatial signals for feedback purposes. However, Lok is directed to educational evaluation, not to generation of a learned behavioral model. While Lok captures and displays interaction data to facilitate review, it does not disclose or suggest transferring behavioral characteristics of a specific human to a virtual human, nor does it generate a model that encodes such behavior. That is, Lok observes and visualizes behavior but does not learn, transfer, or embed that behavior into a virtual human. Accordingly, Lok fails to teach or suggest generating a "learned digital twin behavior model by transferring knowledge through the social simulations to the virtual humans," as recited in claim 1. Examiner notes Lok discloses the portions of the limitations of claim 1 as written. Examiner notes applicant describes a specific human. The claim wording is humans. The learned digital twin behavior model is based on encoded information of humans. Examiner further notes there are two mentions of humans. Examiner suggests clarifying the claim wording to a specific human. Lok observes and visualizes behavior, and learns, transfers, and embeds that behavior into virtual humans, as presented in independent claim 1. [0028] The virtual person is a full-body virtual character with facial animation and gestural capabilities including lipsynched speech, eye blinking, breathing, pointing, idle behaviors (e.g., looking around), the ability to maintain eye contact with the user, as well as scenario specific animations. [0030] … virtual human responses. A conversational domain expert manages the dataset of inputs and responses (e.g., an educator). [0031] Anticipating the utterances that the user will say to a VH through student interactions, and generating the responses of the VH through educator feedback, are the two elements used in order to accurately model human-VH conversations. Accurate modeling is important to expand the application of VH experiences to the training of communications skills. An asynchronous acquisition of knowledge for modeling human-VH conversations includes educators receiving new utterances from students and as a result, creating new responses. In turn, students speak to the VH (utterances) and receive responses from the VH. [0033] …VHs are being developed to play the role of virtual patients (VPs) Applicant notes The Office acknowledges that "Leeds in view of Lok does not specifically teach" these elements, but asserts that Hayashida teaches these elements. (Office Action, p. 12). Hayashida describes monitoring non-verbal behavior, detecting specific gestures, and selecting an avatar response. ( [0073]). However, Hayashida does not describe receiving "audio data and image data," parsing the audio data and the image data, and "extracting audio features from the audio data and visual features from the image data," as recited in amended claim 1. Examiner notes Hayashida contains [0092] 160. The other sensor includes, for example, a moving image imaging device, an audio obtaining device, extracts features of the user Fig. 4, 8, 11-12 Additionally, Hayashida contains [0097] Description will next be made of an image of the virtual reality space which image includes images of the user avatars of the user 160 and the user 170 and an image of the machine avatar. FIG. 2 is a diagram illustrating an example of an image of a virtual reality space. Applicant notes Additionally, Applicant amends claim 1 to recite that the processor is configured to execute the instructions to cause the system to "generate a multimodal feature representation based on the audio features and the visual features." Leeds, Lok, and Hayashida, alone and in combination, do not teach this element of amended claim 1. Examiner notes Leeds teaches this limitation (13:32-36) “(43) (43) Similar to participation in any collaborative forum, the subminds may be AI, human, or any hybrid thereof with human and AI components in any proportion. FIG. 8 shows the same four subminds as in FIG. 7, but here one of the subminds is a human, two are AI (although there is no requirement that they be of the same design or implementation), and one is a hybrid containing some human and AI function combined (such as a human operating a separate console program for lookups and analysis, or an AI that can call upon human for assistance in cases it has difficulty with). As disclosed in the '516 patent, all subminds have a two-way data connection to the forum where they can transmit data to the forum and read data from the forum. Though this preferred embodiment will be described as text-only for the sake of teaching the invention, data described here may include any combination of text, audio, images, binary or other computer readable form, video, ambient UX, haptics, olfaction, or any non-human or internet of things (IoT, e.g., device) perception. Moreover, the Office has not provided a sufficient motivation to combine Leeds, Lok, and Hayashida. The motivations to combine provided by the Office are conclusory and unsupported by the references. The Supreme Court, in KSR International Co. v. Teleflex Inc., provided that "rejections on obviousness cannot be sustained by mere conclusory statements; instead, there must be some articulated reasoning with some rational underpinning to support the legal conclusion of obviousness." KSR, 550 U.S. 398, 418 (2007) (quoting In re Kahn, 441 F.3d 977, 988 (Fed. Cir. 2006)). The Office relies on restatements of Applicant's claim language to assert that it would have been obvious to combine Leeds and Lok because they are "in the same field of endeavor of signal processing" and that it would allow "users to model virtual conversations as recognized by Lok [0038]." (Office Action, pp. 11-12). Similarly, the Office asserts that it would have been obvious to combine Hayashida with Leeds and Lok because they are "in the same field of endeavor of signal processing" and that it would allow "improved determination accuracy as recognized by Hayashida [0407]." (Office Action, pp. 12-13). The motivations asserted by the Office rely on broad, generalized statements that the references are in the same "field of endeavor" and would yield vague improvements, without identifying how the proposed improvements would fit into the teachings of the other references. The fact that references may relate generally to "signal processing" does not provide a meaningful rationale for combining them, particularly where the claimed invention is directed to digital twin behavioral modeling. Thus, the Office has not articulated how or why a person of ordinary skill would be motivated to modify the cited systems to arrive at the claimed invention. Examiner notes Leeds relates to Artificial life, i.e. computing arrangements simulating life based on simulated virtual individual or collective life forms, and, Semantic analysis Examiner notes Lok relates to image of a human Lok claim 1 “…virtual human the output of the simulation system is based on both the sensed physical interaction with the model and the communication representing physical and vocal interaction with a human represented by the virtual human” Examiner notes Hayashida relates to animation of characters, e.g. humans, animals or virtual beings and their interaction with the human body Examiner notes the systems of Leeds, Lok and Hayashida are closely related. Applicant notes, the cited passage merely references "quantum computing," "probabilistic calculations," and "particle entanglement" at a high-level or conceptual level in the context of conversational interpretation, and does not provide any technical disclosure of quantum-based mechanisms. Leeds describes probabilistic handling of conversational ambiguity, but does not describe any implementation of quantum state modeling, quantum information processing, or transfer of system states. Accordingly, Leeds fails to disclose or suggest any concrete quantum framework, such as state encoding, entangled system representations, state transitions, or measurable quantum-space relationships. Its references to quantum concepts are abstract and non-operational, and do not describe any specific mechanism by which such concepts are implemented in a system. As such, the cited portion of Leeds is insufficient to support the rejection of claims 3-10. Examiner notes Applicant is taking standard methods of the industry and applying words like "quantum computing", "particle entanglement" etc. Examiner considers this a label without technical effect and does not apply significant patentable weight to these terms. If applicant can show specific quantum algorithms, or specific quantum hardware that improves the modeling in a novel, non-obvious way, the analysis may change. Applicant's arguments have been fully considered but they are not persuasive. Updated mappings to reflect the amendments to the independent claims have been provided Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Independent claim 1 recites, “1. A computing device comprising: a computer-readable medium storing instructions for digital twin behavioral modeling; and a processor configured to execute the instructions to cause a system including at least a multimodal dialog manager and a digital twin platform to: receive audio data and image data representing one or more multimodal queries or conversations; (This relates to a human using auditory processing to receive conversations) parse the audio data and the image data for content; (This relates to a human using natural language processing to parse conversations for content in the human mind.) recognize and sense one or more multimodal content by extracting audio features from the audio data and visual features from the image data; (This relates to a human using natural language understanding to recognize and sense content in the human mind. This relates to a human extracting audio features from audio in the human mind and extracting visual features from image data in the human mind.) generate a multimodal feature representation based on the audio features and the visual features; [this relates to a human generating a feature representation using pen and paper] train the multimodal dialog manager and a virtual human, wherein to train, the processor is further configured to execute instructions to: (This relates to a human using pen and paper to execute instructions.) process social simulations which model behavior of humans and virtual humans, wherein each pair of a human and a virtual human form digital twins; (This relates to a human using natural language understanding and awareness to process social situations.) generate, via the digital twin platform based on the multimodal feature representation, a learned digital twin behavior model by transferring knowledge through the social simulations to the virtual humans, wherein social and functional behavior of humans is encoded in the learned digital twin behavior model; (This relates to a human using natural learning to transfer knowledge through speech or pen and paper.) and generate responses to the multimodal queries or conversations based on the learned digital twin behavior model. (This relates to a human using pen and paper to generate a response) The Dependent Claims do not include additional limitations that could incorporate the abstract idea into a practical application or cause the Claim as a whole to amount to significantly more than the underlying abstract idea. This judicial exception is not integrated into a practical application. In particular, claim 1 recites the additional elements of “processor”. For example, in [0035] of the as filed specification, there is description of using The computing device 100 includes at least one processor 102. Accordingly, these additional elements do not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claims are directed to an abstract idea. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract idea into a practical application, the additional element of using a computer noted as a general computer. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. Further, the additional limitation in the claims noted above are directed towards insignificant solution activity. The claims are not patent eligible. Regarding Independent Claim 11, claim 11 is a method claim with limitations similar to that of Claim 1 and is rejected under the same rational. no additional elements. Dependent claim 2 recites, “2. The computing device of claim 1, wherein the social and functional behavior include verbal and non-verbal behavioral patterns. (This relates to a human using contextual awareness and learning to include verbal and non-verbal patterns.) no additional elements. Dependent claim 3 recites, “3. The computing device of claim 1, wherein quantum learning and quantum transfer is used to transfer the knowledge among the digital twins. (This relates to a human using speech or pen and paper to transfer knowledge.) no additional elements. Dependent claim 4 recites, “4. The computing device of claim 3, wherein quantum teleportation and quantum entanglement is used to transfer conversational control states among the digital twins. (This relates to a human using speech or pen and paper to transfer states.) no additional elements. Dependent claim 5 recites, “5. The computing device of claim 4, wherein the quantum teleportation transfers a conversational control state to one of the human or the virtual human without communicating a control transfer to the human or the virtual human currently having the conversational control state. (This relates to a human using speech or pen and paper to transfer control state.) no additional elements. Dependent claim 6 recites, “6. The computing device of claim 5, wherein the quantum entanglement describes quantum states of the human and the virtual human with reference to each other. (This relates to a human having awareness in the human mind.) no additional elements. Dependent claim 7 recites, “7. The computing device of claim 6, wherein the human and the virtual human exist in a superposition quantum state. (This relates to a human existing.) no additional elements. Dependent claim 8 recites, “8. The computing device of claim 7, wherein the human and the virtual human can switch between quantum states to capture digital twin behavioral patterns represented in the learned digital twin behavior model. (this relates to a human applying logic and reasoning to capture behavior patterns.) no additional elements. Dependent claim 9 recites, “9. The computing device of claim 4, wherein the digital twin platform is a multi-layer quantum framework for performing the quantum teleportation and the quantum entanglement as between the digital twins. (This relates to a human existing) no additional elements. Dependent claim 10 recites, “10. The computing device of claim 3, wherein quantum information difference is minimized between data points for the human in quantum space and data points for the virtual human in quantum space so that a behavior of the virtual human is substantially equivalent to a behavior of the human. (This relates to a human behaving the same as another) no additional elements. As to claim 12, claim 12 is a parallel method claim with limitations similar to that of claim 2 and is rejected under the same rationale. As to claim 13, claim 13 is a parallel method claim with limitations similar to that of claim 3 and is rejected under the same rationale. As to claim 14, claim 14 is a parallel method claim with limitations similar to that of claim 4 and is rejected under the same rationale. As to claim 15, claim 15 is a parallel method claim with limitations similar to that of claim 5 and is rejected under the same rationale. As to claim 16, claim 16 is a parallel method claim with limitations similar to that of claim 6 and is rejected under the same rationale. As to claim 17, claim 17 is a parallel method claim with limitations similar to that of claim 7 and is rejected under the same rationale. As to claim 18, claim 18 is a parallel method claim with limitations similar to that of claim 8 and is rejected under the same rationale. As to claim 19, claim 19 is a parallel method claim with limitations similar to that of claim 9 and is rejected under the same rationale. As to claim 20, claim 20 is a parallel method claim with limitations similar to that of claim 10 and is rejected under the same rationale. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Leeds (U.S. Patent Number US 11431660 B1), in view of Lok (U.S. Patent Number US 20120139828 A1), and further in view of Hayashida (U.S. Patent Number US 20180120928 A1). Regarding independent Claim 1, Leeds teaches process social simulations which model behavior of humans and virtual humans, wherein each pair of a human and a virtual human form digital twins; (see Leeds (6:2-10) (42) “Theory of Mind” or ToM is the modeling of one mind by another. It is used to describe the ability to attribute mental states to others, in the context of the present invention, for social interaction and collaboration in particular. ToM can be a model of another entity (which could be human or AI) held by a submind that captures that submind's understanding of the other entity.”) and generate, via the digital twin platform based on the multimodal feature representation, a learned digital twin behavior model by transferring knowledge through the social simulations to the virtual humans, (see Leeds “(55:60-56:67) “(241… The present invention can be used to coordinate goal-oriented processes managed by facilitators and implemented by collaborative subminds. Quantum level challenges of NLU implementation, submind coordination and human language comprehensibility may be addressed by the present invention, as such computing resources become more readily available, and intelligent biological systems mechanisms are further elucidated. Quantum computing and effects utilized by subminds may include probabilistic calculations, particle entanglement and calculation with matrices. For example, all possible interpretations of a conversation segment in context may exist as probabilities (for example metaphor, allusion, sarcasm, off subject, in error and falsehood) until the time of chosen interpretation by the recipient, or qualified conversationally, or confirmed by action by the speaker. Results may be nondeterministic except as statistical norms, based on probabilistic states of quantum devices. Cooperative behavior for security and defense/offense are possible applications. Layers of security can be implemented within the architecture, e.g., the layering of forums. The open platform aspects may fit well with zero-trust architectures. Animal defense and swarm cooperation to avoid, confuse or distract other entities via cooperation, for example using cooperating submind applications on mobile devices (such as mobile phones and digital assistants) to emit the sounds of moving and barking animals, or high and low frequency tones with beat frequencies, for cumulative dispersed effects, or to respond with other organizing directions for humans and devices. Team marketing, e.g., feature and price testing (by buyer and seller) may be addressable by the present invention, particularly in the enlistment and teaming of existing marketing AI chatbots. Subminds with specific product or service knowledge may be added to a conversation, without the initial submind team member leaving, enabling improved sales opportunities and customer service engagement. Collaborative confusion where the purpose of the collaboration is to mislead, distract, or confuse, possibly through purposeful ambiguity. Entertainment animation, including interactive immersive AR/VR, recorded, modeled, or historical, or combined, will find the present invention of use. Lip synchronization during animated speech is a challenging area that requires balance among the various factors and ways to evaluate them, which the present invention could excel at. Real time flexible immersive simulation will be possible using subminds for efficient component simulations.”) generate a multimodal feature representation based on the audio features and the visual features; (See Leeds (13:32-36) “(43) Similar to participation in any collaborative forum, the subminds may be AI, human, or any hybrid thereof with human and AI components in any proportion. FIG. 8 shows the same four subminds as in FIG. 7, but here one of the subminds is a human, two are AI (although there is no requirement that they be of the same design or implementation), and one is a hybrid containing some human and AI function combined (such as a human operating a separate console program for lookups and analysis, or an AI that can call upon human for assistance in cases it has difficulty with). As disclosed in the '516 patent, all subminds have a two-way data connection to the forum where they can transmit data to the forum and read data from the forum. Though this preferred embodiment will be described as text-only for the sake of teaching the invention, data described here may include any combination of text, audio, images, binary or other computer readable form, video, ambient UX, haptics, olfaction, or any non-human or internet of things (IoT, e.g., device) perception. Leeds does not specifically teach 1. A computing device comprising: a computer-readable medium storing instructions for digital twin behavioral modeling; and a processor configured to execute the instructions to cause a system including at least a multimodal dialog manager and a digital twin platform to: However, Lok does teach this limitation (see Lok [0074] As described above, the embodiments of the invention may be embodied in the form of hardware, software, firmware, or any processes and/or apparatuses for practicing the embodiments. Embodiments of the invention may also be embodied in the form of computer program code containing instructions embodied in tangible media, such as floppy diskettes, CD-ROMs, hard drives, or any other computer-readable storage medium, wherein, when the computer program code is loaded into and executed by a computer, the computer becomes an apparatus for practicing the invention. The present invention can also be embodied in the form of computer program code, for example, whether stored in a storage medium, loaded into and/or executed by a computer, wherein, when the computer program code is loaded into and executed by a computer, the computer becomes an apparatus for practicing the invention. When implemented on a general-purpose microprocessor, the computer program code segments configure the microprocessor to create specific logic circuits.”) train the multimodal dialog manager and a virtual human, wherein to train, the processor is further configured to execute instructions to: (see Lok [0032] “The creation of VHs for practicing interview skills is logistically difficult and time consuming. The logistical hurdles involve the efficient acquisition of knowledge for the conversational model; specifically, the portion of the model that enables a VH to respond to user speech. Acquiring this knowledge has been a problem because it required extensive VH developer time to program the conversational model. Embodiments of the invention include a method for implementing a Virtual People Factory.”) wherein social and functional behavior of humans is ncoded in the learned digital twin behavior model; (see Lok [0047] To improve skills education, H-VH interactions with AAR are augmented. AAR enables students to review their H-VH interaction to evaluate their actions, and receive feedback on how to improve future real-world experiences. AAR for H-VH interactions incorporates three design principles: 1. An H-VH interaction is composed of social, temporal, and spatial characteristics. These characteristics will be explored by students via AAR visualizations. 2. An H-VH interaction is a set of signals. Interaction signals are captured, logged, and processed to produce visualizations. 3. An H-VH interaction is complex. Students gain insight into this complexity by reviewing multiple visualizations, such as audio, video, text, and graphs to enable AAR, IPSViz processes the signals characterizing an HVH interaction to provide an array of visualizations. The visualizations are used to facilitate interpersonal skills education. Novel visualizations can be produced by leveraging the many signals that are captured in an H-VH interaction. Given an H-VH interaction, AAR is facilitated through the following visualization types: the H-VH interaction can be 3D rendered from any perspective, including that of the conversation partner (the virtual camera is located at the VH's eyes). These are called "spatial visualizations." Students are able to perceive "what it was like to talk to themselves"; events in the H-VH interaction are visualized with respect to an interaction timeline. These are called "temporal visualizations." Students are able to discern the relationship between conversation events; Verbal and nonverbal behaviors are presented in log, graph, and 3D formats. These are called "social visualizations." Students are able to understand how their behavior affects the conversation.”) (see Lok [0028] The virtual person is a full-body virtual character with facial animation and gestural capabilities including lipsynched speech, eye blinking, breathing, pointing, idle behaviors (e.g., looking around), the ability to maintain eye contact with the user, as well as scenario specific animations.”)(see Lok [0030] … virtual human responses. A conversational domain expert manages the dataset of inputs and responses (e.g., an educator).”) (see Lok [0033] …VHs are being developed to play the role of virtual patients (VPs)”) and generate responses to the multimodal queries or conversations based on the learned digital twin behavior model. (see Lok [0021] “Users are able to interact with the MRH through a combination of verbal, gestural, and haptic communication techniques. The user communicates verbally with the MRH patient using natural speech. Wireless microphone 104 transmits the user's speech to the simulation system 112, which performs speech recognition. Recognized speech is matched to a database of question-answer pairs using a keyword based approach. The database for a scenario consists of 100-300 question responses paired with 1000-3000 questions. The many syntactical ways of expressing a question are handled by the keyword-based approach and a list of common synonyms. The MRH responds to matched user speech with speech pre-recorded by a human patient through the HMD 102.”) (see Lok [0031] Anticipating the utterances that the user will say to a VH through student interactions, and generating the responses of the VH through educator feedback, are the two elements used in order to accurately model human-VH conversations. Accurate modeling is important to expand the application of VH experiences to the training of communications skills. An asynchronous acquisition of knowledge for modeling human-VH conversations includes educators receiving new utterances from students and as a result, creating new responses. In turn, students speak to the VH (utterances) and receive responses from the VH.”) Leeds and Lok are in the same field of endeavor of signal processing, therefore, it would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the processing of social simulations which model behavior of humans and virtual humans, wherein each pair of a human and a virtual human form digital twins; and transfer knowledge through the social simulations to the virtual humans of Leeds to incorporate the teachings of Lok to include the device and the training the multimodal dialog manager and a virtual human, wherein to train, the processor is further configured to execute instructions to: wherein social and functional behavior of humans is ncoded in the learned digital twin behavior model; and generate responses to the multimodal queries or conversations based on the learned digital twin behavior model. This allows users to model virtual conversations as recognized by Lok [0038]. Leeds in view of Lok does not specifically teach receive audio data and image data representing one or more multimodal queries or conversations; parse the audio data and the image data for content; recognize and sense one or more multimodal content by extracting audio features from the audio data and visual features from the image data; However, Hayashida does teach this limitation (see Hayashida [0073] “The action instructing unit 125 monitors the non-verbal behavior of the communication partner user using the log data used in generating the image of the user avatar. In addition, the action instructing unit 125 determines whether or not the communication partner user has performed a particular non-verbal behavior based on a result of the monitoring. Further, when the action instructing unit 125 determines that the communication partner user has performed a particular non-verbal behavior, the action instructing unit 125 determines an appropriate image of the machine avatar for bringing the user into a post desirable-change user state, and gives an instruction to the machine avatar information display processing unit 121.”) (see Hayashida [0097] Description will next be made of an image of the virtual reality space which image includes images of the user avatars of the user 160 and the user 170 and an image of the machine avatar. FIG. 2 is a diagram illustrating an example of an image of a virtual reality space.”) Leeds in view of Lok and Hayashida are in the same field of endeavor of signal processing, therefore, it would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the device of combination Leeds and Lok to incorporate receive audio data and image data representing one or more multimodal queries or conversations; parse the audio data and the image data for content; recognize and sense one or more multimodal content by extracting audio features from the audio data and visual features from the image data; generate a multimodal feature representation based on the audio features and the visual features of Hayashida. This allows improved determination accuracy as recognized by Hayashida [0407]. As to Independent Claim 11, Claim 11 is a parallel method claim with limitations similar to that of claim 1 and is rejected under the same rationale. As to Claim 2, Leeds in view of Lok and further in view of Hayashida teaches 2. The computing device of claim 1, Furthermore, Hayashida teaches wherein the social and functional behavior include verbal and non-verbal behavioral patterns. (see Hayashida [0073] “The action instructing unit 125 monitors the non-verbal behavior of the communication partner user using the log data used in generating the image of the user avatar. In addition, the action instructing unit 125 determines whether or not the communication partner user has performed a particular non-verbal behavior based on a result of the monitoring. Further, when the action instructing unit 125 determines that the communication partner user has performed a particular non-verbal behavior, the action instructing unit 125 determines an appropriate image of the machine avatar for bringing the user into a post desirable-change user state, and gives an instruction to the machine avatar information display processing unit 121.”) Leeds in view of Lok and further in view of Hayashida are in the same field of endeavor of signal processing, therefore, it would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the device of combination of Leeds and Lok and Hayashida to incorporate wherein the social and functional behavior include verbal and non-verbal behavioral patterns of Hayashida. This allows improved determination accuracy as recognized by Hayashida [0407]. As to Claim 3, Leeds in view of Lok and further in view of Hayashida teaches 3. The computing device of claim 1, Furthermore, Leeds teaches wherein quantum learning and quantum transfer is used to transfer the knowledge among the digital twins. (see Leeds “(55:60-56:67) “(… The present invention can be used to coordinate goal-oriented processes managed by facilitators and implemented by collaborative subminds. Quantum level challenges of NLU implementation, submind coordination and human language comprehensibility may be addressed by the present invention, as such computing resources become more readily available, and intelligent biological systems mechanisms are further elucidated. Quantum computing and effects utilized by subminds may include probabilistic calculations, particle entanglement and calculation with matrices. For example, all possible interpretations of a conversation segment in context may exist as probabilities (for example metaphor, allusion, sarcasm, off subject, in error and falsehood) until the time of chosen interpretation by the recipient, or qualified conversationally, or confirmed by action by the speaker. … Subminds with specific product or service knowledge may be added to a conversation, without the initial submind team member leaving, enabling improved sales opportunities and customer service engagement. Collaborative confusion where the purpose of the collaboration is to mislead, distract, or confuse, possibly through purposeful ambiguity. Entertainment animation, including interactive immersive AR/VR, recorded, modeled, or historical, or combined, will find the present invention of use. Lip synchronization during animated speech is a challenging area that requires balance among the various factors and ways to evaluate them, which the present invention could excel at. Real time flexible immersive simulation will be possible using subminds for efficient component simulations.”) As to Claim 4, Leeds in view of Lok and further in view of Hayashida teaches 4. The computing device of claim 3, Furthermore, Leeds teaches wherein quantum teleportation and quantum entanglement is used to transfer conversational control states among the digital twins. (see Leeds “(55:60-56:67) “(… The present invention can be used to coordinate goal-oriented processes managed by facilitators and implemented by collaborative subminds. Quantum level challenges of NLU implementation, submind coordination and human language comprehensibility may be addressed by the present invention, as such computing resources become more readily available, and intelligent biological systems mechanisms are further elucidated. Quantum computing and effects utilized by subminds may include probabilistic calculations, particle entanglement and calculation with matrices. For example, all possible interpretations of a conversation segment in context may exist as probabilities (for example metaphor, allusion, sarcasm, off subject, in error and falsehood) until the time of chosen interpretation by the recipient, or qualified conversationally, or confirmed by action by the speaker. … Subminds with specific product or service knowledge may be added to a conversation, without the initial submind team member leaving, enabling improved sales opportunities and customer service engagement. Collaborative confusion where the purpose of the collaboration is to mislead, distract, or confuse, possibly through purposeful ambiguity. Entertainment animation, including interactive immersive AR/VR, recorded, modeled, or historical, or combined, will find the present invention of use. Lip synchronization during animated speech is a challenging area that requires balance among the various factors and ways to evaluate them, which the present invention could excel at. Real time flexible immersive simulation will be possible using subminds for efficient component simulations.”) As to Claim 5, Leeds in view of Lok and further in view of Hayashida teaches 5. The computing device of claim 4, Furthermore, Leeds teaches wherein the quantum teleportation transfers a conversational control state to one of the human or the virtual human without communicating a control transfer to the human or the virtual human currently having the conversational control state. (see Leeds “(55:60-56:67) “(… The present invention can be used to coordinate goal-oriented processes managed by facilitators and implemented by collaborative subminds. Quantum level challenges of NLU implementation, submind coordination and human language comprehensibility may be addressed by the present invention, as such computing resources become more readily available, and intelligent biological systems mechanisms are further elucidated. Quantum computing and effects utilized by subminds may include probabilistic calculations, particle entanglement and calculation with matrices. For example, all possible interpretations of a conversation segment in context may exist as probabilities (for example metaphor, allusion, sarcasm, off subject, in error and falsehood) until the time of chosen interpretation by the recipient, or qualified conversationally, or confirmed by action by the speaker. … Subminds with specific product or service knowledge may be added to a conversation, without the initial submind team member leaving, enabling improved sales opportunities and customer service engagement. Collaborative confusion where the purpose of the collaboration is to mislead, distract, or confuse, possibly through purposeful ambiguity. Entertainment animation, including interactive immersive AR/VR, recorded, modeled, or historical, or combined, will find the present invention of use. Lip synchronization during animated speech is a challenging area that requires balance among the various factors and ways to evaluate them, which the present invention could excel at. Real time flexible immersive simulation will be possible using subminds for efficient component simulations.”) As to Claim 6, Leeds in view of Lok and further in view of Hayashida teaches 6. The computing device of claim 5, Furthermore, Leeds teaches wherein the quantum entanglement describes quantum states of the human and the virtual human with reference to each other. (see Leeds “(55:60-56:67) “(… The present invention can be used to coordinate goal-oriented processes managed by facilitators and implemented by collaborative subminds. Quantum level challenges of NLU implementation, submind coordination and human language comprehensibility may be addressed by the present invention, as such computing resources become more readily available, and intelligent biological systems mechanisms are further elucidated. Quantum computing and effects utilized by subminds may include probabilistic calculations, particle entanglement and calculation with matrices. For example, all possible interpretations of a conversation segment in context may exist as probabilities (for example metaphor, allusion, sarcasm, off subject, in error and falsehood) until the time of chosen interpretation by the recipient, or qualified conversationally, or confirmed by action by the speaker. … Subminds with specific product or service knowledge may be added to a conversation, without the initial submind team member leaving, enabling improved sales opportunities and customer service engagement. Collaborative confusion where the purpose of the collaboration is to mislead, distract, or confuse, possibly through purposeful ambiguity. Entertainment animation, including interactive immersive AR/VR, recorded, modeled, or historical, or combined, will find the present invention of use. Lip synchronization during animated speech is a challenging area that requires balance among the various factors and ways to evaluate them, which the present invention could excel at. Real time flexible immersive simulation will be possible using subminds for efficient component simulations.”) As to Claim 7, Leeds in view of Lok and further in view of Hayashida teaches 7. The computing device of claim 6, Furthermore, Leeds teaches, wherein the human and the virtual human exist in a superposition quantum state. (see Leeds “(55:60-56:67) “(… The present invention can be used to coordinate goal-oriented processes managed by facilitators and implemented by collaborative subminds. Quantum level challenges of NLU implementation, submind coordination and human language comprehensibility may be addressed by the present invention, as such computing resources become more readily available, and intelligent biological systems mechanisms are further elucidated. Quantum computing and effects utilized by subminds may include probabilistic calculations, particle entanglement and calculation with matrices. For example, all possible interpretations of a conversation segment in context may exist as probabilities (for example metaphor, allusion, sarcasm, off subject, in error and falsehood) until the time of chosen interpretation by the recipient, or qualified conversationally, or confirmed by action by the speaker. … Subminds with specific product or service knowledge may be added to a conversation, without the initial submind team member leaving, enabling improved sales opportunities and customer service engagement. Collaborative confusion where the purpose of the collaboration is to mislead, distract, or confuse, possibly through purposeful ambiguity. Entertainment animation, including interactive immersive AR/VR, recorded, modeled, or historical, or combined, will find the present invention of use. Lip synchronization during animated speech is a challenging area that requires balance among the various factors and ways to evaluate them, which the present invention could excel at. Real time flexible immersive simulation will be possible using subminds for efficient component simulations.”) As to Claim 8, Leeds in view of Lok and further in view of Hayashida teaches 8. The computing device of claim 7, Furthermore, Leeds teaches, wherein the human and the virtual human can switch between quantum states to capture digital twin behavioral patterns represented in the learned digital twin behavior model. (see Leeds “(55:60-56:67) “(… The present invention can be used to coordinate goal-oriented processes managed by facilitators and implemented by collaborative subminds. Quantum level challenges of NLU implementation, submind coordination and human language comprehensibility may be addressed by the present invention, as such computing resources become more readily available, and intelligent biological systems mechanisms are further elucidated. Quantum computing and effects utilized by subminds may include probabilistic calculations, particle entanglement and calculation with matrices. For example, all possible interpretations of a conversation segment in context may exist as probabilities (for example metaphor, allusion, sarcasm, off subject, in error and falsehood) until the time of chosen interpretation by the recipient, or qualified conversationally, or confirmed by action by the speaker. … Subminds with specific product or service knowledge may be added to a conversation, without the initial submind team member leaving, enabling improved sales opportunities and customer service engagement. Collaborative confusion where the purpose of the collaboration is to mislead, distract, or confuse, possibly through purposeful ambiguity. Entertainment animation, including interactive immersive AR/VR, recorded, modeled, or historical, or combined, will find the present invention of use. Lip synchronization during animated speech is a challenging area that requires balance among the various factors and ways to evaluate them, which the present invention could excel at. Real time flexible immersive simulation will be possible using subminds for efficient component simulations.”) As to Claim 9, Leeds in view of Lok and further in view of Hayashida teaches 9. The computing device of claim 4, Furthermore, Leeds teaches, wherein the digital twin platform is a multi-layer quantum framework for performing the quantum teleportation and the quantum entanglement as between the digital twins. (see Leeds “(55:60-56:67) “(… The present invention can be used to coordinate goal-oriented processes managed by facilitators and implemented by collaborative subminds. Quantum level challenges of NLU implementation, submind coordination and human language comprehensibility may be addressed by the present invention, as such computing resources become more readily available, and intelligent biological systems mechanisms are further elucidated. Quantum computing and effects utilized by subminds may include probabilistic calculations, particle entanglement and calculation with matrices. For example, all possible interpretations of a conversation segment in context may exist as probabilities (for example metaphor, allusion, sarcasm, off subject, in error and falsehood) until the time of chosen interpretation by the recipient, or qualified conversationally, or confirmed by action by the speaker. … Subminds with specific product or service knowledge may be added to a conversation, without the initial submind team member leaving, enabling improved sales opportunities and customer service engagement. Collaborative confusion where the purpose of the collaboration is to mislead, distract, or confuse, possibly through purposeful ambiguity. Entertainment animation, including interactive immersive AR/VR, recorded, modeled, or historical, or combined, will find the present invention of use. Lip synchronization during animated speech is a challenging area that requires balance among the various factors and ways to evaluate them, which the present invention could excel at. Real time flexible immersive simulation will be possible using subminds for efficient component simulations.”) As to Claim 10, Leeds in view of Lok and further in view of Hayashida teaches 10. The computing device of claim 3, Furthermore, Leeds teaches wherein quantum information difference is minimized between data points for the human and data points for the virtual human in quantum space so that a behavior of the virtual human is substantially equivalent to a behavior of the human. (see Leeds “(55:60-56:67) “(… The present invention can be used to coordinate goal-oriented processes managed by facilitators and implemented by collaborative subminds. Quantum level challenges of NLU implementation, submind coordination and human language comprehensibility may be addressed by the present invention, as such computing resources become more readily available, and intelligent biological systems mechanisms are further elucidated. Quantum computing and effects utilized by subminds may include probabilistic calculations, particle entanglement and calculation with matrices. For example, all possible interpretations of a conversation segment in context may exist as probabilities (for example metaphor, allusion, sarcasm, off subject, in error and falsehood) until the time of chosen interpretation by the recipient, or qualified conversationally, or confirmed by action by the speaker. … Subminds with specific product or service knowledge may be added to a conversation, without the initial submind team member leaving, enabling improved sales opportunities and customer service engagement. Collaborative confusion where the purpose of the collaboration is to mislead, distract, or confuse, possibly through purposeful ambiguity. Entertainment animation, including interactive immersive AR/VR, recorded, modeled, or historical, or combined, will find the present invention of use. Lip synchronization during animated speech is a challenging area that requires balance among the various factors and ways to evaluate them, which the present invention could excel at. Real time flexible immersive simulation will be possible using subminds for efficient component simulations.”) As to claim 12, claim 12 is a parallel method claim with limitations similar to that of claim 2 and is rejected under the same rationale. As to claim 13, claim 13 is a parallel method claim with limitations similar to that of claim 3 and is rejected under the same rationale. As to claim 14, claim 14 is a parallel method claim with limitations similar to that of claim 4 and is rejected under the same rationale. As to claim 15, claim 15 is a parallel method claim with limitations similar to that of claim 5 and is rejected under the same rationale. As to claim 16, claim 16 is a parallel method claim with limitations similar to that of claim 6 and is rejected under the same rationale. As to claim 17, claim 17 is a parallel method claim with limitations similar to that of claim 7 and is rejected under the same rationale. As to claim 18, claim 18 is a parallel method claim with limitations similar to that of claim 8 and is rejected under the same rationale. As to claim 19, claim 19 is a parallel method claim with limitations similar to that of claim 9 and is rejected under the same rationale. As to claim 20, claim 20 is a parallel method claim with limitations similar to that of claim 10 and is rejected under the same rationale. Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to KRISTEN MICHELLE MASTERS whose telephone number is (703)756-1274. The examiner can normally be reached M-F 8:30 AM - 5:00 PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre Louis Desir can be reached at 571-272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /KRISTEN MICHELLE MASTERS/Examiner, Art Unit 2659 /PIERRE LOUIS DESIR/Supervisory Patent Examiner, Art Unit 2659
Read full office action

Prosecution Timeline

Feb 12, 2024
Application Filed
Jan 14, 2026
Non-Final Rejection mailed — §101, §103
Apr 14, 2026
Response Filed
Jul 16, 2026
Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12694889
PROFANITY FILTER FOR COLLABORATION SESSIONS IN HETEROGENOUS COMPUTING PLATFORMS
3y 10m to grant Granted Jul 28, 2026
Patent 12664366
CROSS-DOMAIN LABEL-ADAPTIVE STANCE DETECTION
3y 11m to grant Granted Jun 23, 2026
Patent 12592219
Hearing Device User Communicating With a Wireless Communication Device
4y 5m to grant Granted Mar 31, 2026
Patent 12548569
METHOD AND SYSTEM OF DETECTING AND IMPROVING REAL-TIME MISPRONUNCIATION OF WORDS
3y 2m to grant Granted Feb 10, 2026
Patent 12548564
SYSTEM AND METHOD FOR CONTROLLING A PLURALITY OF DEVICES
3y 7m to grant Granted Feb 10, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
65%
Grant Probability
86%
With Interview (+21.2%)
3y 0m (~6m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 49 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month