DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1; 2, 3, 4, 6, 7, 8 and 9 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hoque et al. U.S. Pub. No. 2014/0356822 in view of Kita U.S. Pub. No. 2019/0087736 and Whitney et al. U.S. Pub. No. 2021/0104100.
Re: claim 1, Hoque teaches
1. (Currently Amended) A system comprising: one or more processors to: initiate a first virtual agent corresponding to an instance of a first application that is hosted using one or more first computing devices; (“For example, a user could use the virtual coach for a mock job interview as follows: (In this example... the human user is named Mike and the virtual coach is named John). Mike chooses to interact with an animated coach to improve his interview skills. Mike chooses one of two virtual counselors, John, who appears on the screen, greets him, and starts the interview.”; Hoque, [0077])
For example, for a mock job interview, the user, Mike, requests a mock job interview and selects a virtual coach named John (initiate a first virtual agent), who appears on the screen, greets Mike and starts the interview.
(“Based on the analysis of the sensor data and social media data 101, the computer then determines the behavior of a virtual coach, using a behavior generation software module 131. The computer causes animations of the virtual coach’s behavior to be displayed visually on the screen 135 and audibly by speakers.”; Hoque, [0064])
Based on the analysis of the sensor data, the computer determines the behavior of the virtual coach, using the behavior generation software module 131 (corresponding to an instance of a first application that is hosted using one or more first computing devices). The behavior generation software module is considered to include an instance of the application for the virtual coach.
generate, remotely from the one or more first computing devices hosting the instance of the first application, first data corresponding to a first graphical representation of the first virtual agent, (“Through the network 409, the user’s computer 401 may interface with servers (e.g., 411, 413) that control the virtual coach and UI. The servers (e.g., 411, 413) may store data in, and retrieve data from, memory devices (e.g., one or more hard drives) 412, 414.”; Hoque, [0071], Fig. 4)
Fig. 4 illustrates that the user’s computer communicates with servers, that control the virtual coach, over the network (generate, remotely from the one or more first computing devices hosting the instance of the first application, first data corresponding to a first graphical representation of the first virtual agent).
(“... virtual coaching is provided remotely over the Internet. The processors that control the virtual coach and UI may be located on one or more servers (or a computer connected to one or more servers) that are remote from the user. A display screen and speakers at the user’s location may display the virtual coach and UI.”; Hoque, [0079])
The virtual coach is provided remotely over the internet (generate, remotely from one or more first computing devices hosting the instance of the first application, first data corresponding to a first graphical representation of the first virtual agent).
send the first data to the one or more first computing devices using one or more wireless networks, the first graphical representation of the first virtual agent being included in a first stream of data of the instance of the first application for presentation on one or more first client devices that are communicating with the one or more first computing devices; (“... virtual coaching is provided remotely over the Internet. The processors that control the virtual coach and the UI may be located on one or more servers... that are remote from the user. A display screen and speakers at the user’s location may display the virtual coach and UI.”; Hoque, [0079])
Virtual coaching is provided remotely over the internet. Fig. 4 illustrates network 409 and servers 411, 413 that control the virtual coach and communicate with the computer 401. The servers control the virtual coach and displays the virtual coach at the user’s device.
(“Depending on the particular implementation of this invention, feedback to a human user can be provided in many different ways, including by a virtual coach speaking or by a UI visual display.”; Hoque, [0079])
Feedback is provided (send, to the one or more first computing devices) to the user, for example, of the displaying the virtual coach speaking (first data corresponding to a first graphical representation of the first virtual agent).
(“... one or more electronic processors are specially adapted:... (3) to compute behaviors of a virtual coach, and to compute the audiovisual animations of that behavior; (4) to control an audio-visual display of a user interface;... The processor may be located in any position or positions within or outside of the automated conversation coach. For example:... (b) at least some of the processors may be remote from other components of the automatic conversation coach. The processors may be connected to each other or to other components in the automatic conversation coach either: (a) wirelessly... For example, one or more electronic processors may be housed in a computer (including computer 400 or servers 411, 413 in Fig. 4).”; Hoque, [0085])
The processors at the servers compute behaviors, audiovisual animations of the virtual coach and control (send, to the one or more first computing devices) an audiovisual (first stream of data) display of the virtual coach (the first graphical representation of the first virtual agent being included in a first stream of data) via network 409, at the user’s device (for presentation on one or more first client devices that are communicating with the one or more first computing devices).
Hoque is silent regarding initiate a second virtual agent corresponding to an instance of a second application that is hosted using one or more second computing devices, however, Kita teaches
initiate a second virtual agent corresponding to an instance of a second application that is hosted using one or more second computing devices; (“The AI selection processing is a set of processings in which when the user uses an artificial intelligence, the artificial intelligence to be used by the user (for example, an artificial intelligence agent) is selected, and the selected artificial intelligence is suggested to the user... for example, the artificial intelligence agent.. which communicates with the user, responds according to the request of the user, provides a service, or operates various electronic devices, is suggested, as the artificial intelligence to be suggested to the user.”; Kita, [0022], Fig. 2)
Fig. 2 illustrates AI selection processing (second computing device) where an AI agent is initiated (initiate a second virtual agent), when AI selection processing selects an AI agent to be used by the user (corresponding to an instance of a second application that is hosted using one or more second computing devices). The AI selection processing communicates with the user through the communication unit 21.
(“Thus, it is possible to select an optimal artificial intelligence for each user according to the evaluation point based on a sense or sensibility of each of the users, without selecting the artificial intelligence by each of the users according to a unified standard... it is possible to select the optimal artificial intelligence for each of the users, according to the evaluation point based on the sense or the sensibility of each of the users.”; Kita, [0039])
The optimal AI agent is selected for each user (which includes a second virtual agent), according to the evaluation point based on the sense or sensibility of each user. Therefore, it would have been obvious to one of ordinary skill in the art at the time of effective filing date to modify the method of Hoque by adding the feature of initiate a second virtual agent corresponding to an instance of a second application that is hosted using one or more second computing devices, in order to enable the AI selection processing to select an artificial intelligence agent, which communicates with the user in response to user requests, as taught by Kita. ([0022])
Hoque and Kita are silent regarding generate, remotely from the one or more second computing devices hosting the instance of the second application, second data corresponding to a second graphical representation of the second virtual agent, and send the second data to the one or more second computing devices using the one or more wireless networks, the second graphical representation of the second virtual agent being included in a second stream of data of the instance of the second application for presentation on one or more second client devices that are communicating with the one or more second computing devices, however, Whitney teaches
generate, remotely from the one or more second computing devices hosting the instance of the second application, second data corresponding to a second graphical representation of the second virtual agent, (“The wearable system 600 can comprise an avatar processing and rendering system 690. The avatar processing and rendering system 690 can be configured to generate, update, animate, and render an avatar based on contextual information. Some or all of the avatar processing and rendering system 600 can be implemented as part of the local processing and data module 260 or the remote processing module 270 alone or in combination...”; Whitney, [0103], Figs. 2, 6A and 9A)
Fig. 9A illustrates an avatar processing and rendering system that is implemented as the remote processing module 270. The avatar processing and rendering system generates and renders an avatar.
(“For example, a virtual avatar may represent a real person or may represent a non-user character, such as a virtual assistant that is configured to interface with users... multiple avatar processing and rendering systems 690 (e.g., as implemented on different wearable devices) can be used for rendering the virtual avatar 670. For example, a first user’s wearable device may be used to determine the first user’s intent, while a second user’s wearable device can determine an avatar’s characteristics and render the avatar of the first user based on the intent received from the first user’s wearable device. The first user’s wearable device and the second user’s wearable device... can communicate via a network...”; Whitney, [0033], [0103], Fig. 6A)
Fig. 6A illustrates multiple avatar processing and rendering systems 690, implemented on different wearable devices, that are used for rendering the virtual avatar (virtual assistant). For example, a first user’s wearable device is used to determine the first user’s intent, and the second user’s wearable device renders the first user’s avatar based on the first user’s intent. The first user’s wearable device and the second user’s wearable device communicate via a network.
(“Fig. 9A schematically illustrates an overall system view depicting multiple user devices interacting with each other. The computing environment 900 includes user devices 930a, 930b, 930c. The user devices 930a, 930b, 930c can communicate with each other through a network 990. The user devices 930a-930c can each include a network interface to communicate via the network 990 with a remote computing system 920... The computing environment 900 can also include one or more remote computing systems 920. The remote computing system 920 may include server computer systems that are clustered and located at different geographic locations. The user devices 930a, 930b, and 930c may communicate with the remote computing system 920 via the network 990.”; Whitney, [0118], Fig. 9A)
Fig. 9A illustrates multiple user devices 930a, 930b, 930c that communicate, via the network, with each other and with one or more remote computing systems 920, which include server computer systems.
(“The remote computing system 920 may include a remote data repository 980 which can maintain information about a specific user’s physical and/or virtual worlds. Data storage 980 can store information related to users, user’s environment... or configurations of avatars of the users... The remote processing module 970 may include one or more processors which can communicate with the user devices (930a, 930b, 930c) and the remote data repository 980.”; Whitney, [0119])
The remote computing system includes a remote data repository/data storage, which stores configurations of user’s avatars. And, the remote processing module includes processors which communicate with user devices and the remote repository.
(“A wearable device can use information acquired of a first user and the environment to animate a virtual avatar that will be rendered by a second user’s wearable device to create a tangible sense of presence of the first user in the second user’s environment. For example, the wearable devices 902 and 904, the remote computing system 920, alone or in combination, may process Alice’s images or movements for presentation by Bob’s wearable device 904 or may process Bob’s images or movements for presentation by Alice’s wearable device 902... the avatars can be rendered based on contextual information such as, e.g., a user’s intent, an environment of the user or an environment in which the avatar is rendered...”; Whitney, [0126])
A second user’s wearable device uses information obtained from a first user to animate a first virtual avatar that is rendered by the second user’s wearable device. For example, the second user’s device uses information from a first user to render a virtual avatar on the second user’s device and the first user’s device uses information from a second user to render a second virtual avatar that is rendered on the first users device.
and send the second data to the one or more second computing devices using the one or more wireless networks, the second graphical representation of the second virtual agent being included in a second stream of data of the instance of the second application for presentation on one or more second client devices that are communicating with the one or more second computing devices (“The local processing and data module 260 may be operatively coupled by communication links 262 or 264, such as via wired or wireless communication links, to the remote processing module 270 or remote data repository 280 such that these remote modules are available as resources to the local processing and data module 260. In addition, remote processing module 270 and remote data repository 280 may be operatively coupled to each other.”; Whitney, [0045], Fig. 2)
Fig. 2 illustrates that the wearable system 200 includes local processing and data module that is in wireless communication with the remote processing module and the remote repository. The remote processing module and the remote repository are also in wireless communication with each other.
(“The wearable system 600 can comprise an avatar processing and rendering system 690. The avatar processing and rendering system 690 can be configured to generate, update, animate, and render an avatar based on contextual information. Some or all of the avatar processing and rendering system 600 can be implemented as part of the local processing and data module 260 or the remote processing module 270 alone or in combination...”; Whitney, [0103], Figs. 2 and 6A)
Fig. 9A illustrates an avatar processing and rendering system that is implemented as the remote processing module 270. The avatar processing and rendering system generates and renders an avatar.
(“Fig. 9A schematically illustrates an overall system view depicting multiple user devices interacting with each other. The computing environment 900 includes user devices 930a, 930b, 930c. The user devices 930a, 930b, 930c can communicate with each other through a network 990. The user devices 930a-930c can each include a network interface to communicate via the network 990 with a remote computing system 920... The computing environment 900 can also include one or more remote computing systems 920. The remote computing system 920 may include server computer systems that are clustered and located at different geographic locations. The user devices 930a, 930b, and 930c may communicate with the remote computing system 920 via the network 990.”; Whitney, [0118], Fig. 9A)
Fig. 9A illustrates multiple user devices 930a, 930b, 930c that communicate, via the network, with each other and with one or more remote computing systems 920, which include server computer systems. The wireless communication between the user devices and the remote computing systems is via a network.
(“The remote computing system 920 may include a remote data repository 980 which can maintain information about a specific user’s physical and/or virtual worlds. Data storage 980 can store information related to users, user’s environment... or configurations of avatars of the users... The remote processing module 970 may include one or more processors which can communicate with the user devices (930a, 930b, 930c) and the remote data repository 980.”; Whitney, [0119])
The remote computing system includes a remote data repository/data storage, which stores configurations of user’s avatars. And, the remote processing module includes processors which communicate with user devices and the remote repository. The communication between the remote processing module, the user devices and the remote data repository is wireless and via a network.
(“A wearable device can use information acquired of a first user and the environment to animate a virtual avatar that will be rendered by a second user’s wearable device to create a tangible sense of presence of the first user in the second user’s environment. For example, the wearable devices 902 and 904, the remote computing system 920, alone or in combination, may process Alice’s images or movements for presentation by Bob’s wearable device 904 or may process Bob’s images or movements for presentation by Alice’s wearable device 902... the avatars can be rendered based on contextual information such as, e.g., a user’s intent, an environment of the user or an environment in which the avatar is rendered...”; Whitney, [0126])
A second user’s wearable device uses information obtained from a first user (send the second data to the one or more second computing devices using the one or more wireless networks) to animate a first virtual avatar that is rendered by the second user’s wearable device (the second graphical representation of the second virtual agent being included in a second stream of data of the instance of the second application for presentation on one or more second client devices that are communicating with the one or more second computing devices). For example, the second user’s device uses information from a first user to render a virtual avatar on the second user’s device and the first user’s device uses information from a second user to render a second virtual avatar that is rendered on the first users device. Whitney is combined with Hoque and Kita such that the virtual avatar of Whitney are the AI agent of Kita and the virtual coach of Hoque and computing environment 900 of Whitney is included in the computing environments of Kita and Hoque. Therefore, it would have been obvious to one of ordinary skill in the art at the time of effective filing date to modify the method of Hoque by adding the feature of generate, remotely from the one or more second computing devices hosting the instance of the second application, second data corresponding to a second graphical representation of the second virtual agent, and send the second data to the one or more second computing devices using the one or more wireless networks, the second graphical representation of the second virtual agent being included in a second stream of data of the instance of the second application for presentation on one or more second client devices that are communicating with the one or more second computing devices, in order to display a virtual assistant configured to assist the user with contextual objects and suggestions depending on what virtual content the user is interacting with, as taught by Whitney ([0033]).
Re: claims 2 and 11 (which are rejected under the same rationale), Hoque, Kita and Whitney teach
(Original) The system of claim 1, wherein the one or more processors are further to at least one of: receive, from the one or more first computing devices, a first request associated with the first virtual agent, wherein the first virtual agent is initiated based at least on the first request; or receive, from the one or more second computing devices, a second request associated with the second virtual agent, wherein the second virtual agent is initiated based at least on the second request. (“For example, a user could use the virtual coach for a mock job interview as follows: (In this example... the human user is named Mike and the virtual coach is named John). Mike chooses to interact with an animated coach to improve his interview skills. Mike chooses one of two virtual counselors, John, who appears on the screen, greets him, and starts the interview.”; Hoque, [0077])
For example, for a mock job interview, the user, Mike, selects a virtual coach named John (receive, from the one or more computing devices, a first request associated with the first virtual agent), who appears on the screen, greets Mike and starts the interview (receive from the one or more first computing devices, a first request associated with the first virtual agent, wherein the first virtual agent is initiated based at least on the first request).
Re: claim 3, Hoque, Kita and Whitney teach
(Original) The system of claim 1, wherein the one or more processors are further to at least one of: receive, from the one or more first computing devices, a first selection of the first virtual agent for the instance of the first application, wherein the first virtual agent is initiated based at least on the first selection; or receive, from the one or more second computing devices, a second selection of the second virtual agent for the instance of the second application, wherein the second virtual agent is initiated based at least on the second selection. (“... the human user may select a particular persona for a virtual coach, out of a set of personas. For example, the human user may select a virtual coach who appears to be female or a virtual coach who appears to be male.”; Hoque, [0012])
The user selects a persona, from a set of personas, for the virtual coach (receive, from the one or more computing devices, a selection of the first virtual agent).
(“For example, a user could use the virtual coach for a mock job interview as follows: (In this example... the human user is named Mike and the virtual coach is named John). Mike chooses to interact with an animated coach to improve his interview skills. Mike chooses one of two virtual counselors, John, who appears on the screen, greets him, and starts the interview.”; Hoque, [0077])
For example, for a mock job interview, the user, Mike, selects a virtual coach named John, who appears on the screen, greets Mike and starts the interview (the first virtual agent is initiated based on the first selection).
(“Based on the analysis of the sensor data and social media data 101, the computer then determines the behavior of a virtual coach, using a behavior generation software module 131. The computer causes animations of the virtual coach’s behavior to be displayed visually on the screen 135 and audibly by speakers.”; Hoque, [0064])
Based on the analysis of the sensor data, the computer determines the behavior of the virtual coach, using the behavior generation software module 131 (instance of the first application).
Re: claim 4, Hoque, Kita and Whitney teach
(Original) The system of claim 1, wherein the one or more processors are further to at least one of: send, to the one or more first computing devices, first audio data corresponding to first speech associated with the first virtual agent, wherein the first audio data is further included in the first stream of data; or send, to the one or more second computing devices, second audio data corresponding to second speech associated with the second virtual agent, wherein the second audio data is further included in the second stream of data. (“... virtual coaching is provided remotely over the Internet. The processors that control the virtual coach and the UI may be located on one or more servers... that are remote from the user. A display screen and speakers at the user’s location may display the virtual coach and UI.”; Hoque, [0079])
Virtual coaching is provided remotely over the internet. Fig. 4 illustrates network 409 and servers 411, 413 that control the virtual coach and communicate with the computer 401. The servers control the virtual coach and displays the virtual coach at the user’s device.
(“Depending on the particular implementation of this invention, feedback to a human user can be provided in many different ways, including by a virtual coach speaking or by a UI visual display.”; Hoque, [0079])
Feedback is provided (send, to the one or more computing devices) to the user, for example, of the virtual coach speaking (first audio data corresponding to first speech associated with the first virtual agent).
(“... one or more electronic processors are specially adapted:... (3) to compute behaviors of a virtual coach, and to compute the audiovisual animations of that behavior; (4) to control an audio-visual display of a user interface;”; Hoque, [0085])
The processors at the servers compute behaviors, audiovisual animations of the virtual coach and control (send, to the one or more first computing devices) an audiovisual (first stream of data) display of the virtual coach (first audio data is further included in the first stream of data) at the user’s device.
Re: claim 6, Hoque, Kita and Whitney teach
(Original) The system of claim 1, wherein the one or more processors are further to at least one of: receive, from the one or more first computing devices, third data representative of one or more first interactions associated with the first virtual agent; or receive, from the one or more second computing devices, fourth data representative of one or more second interactions associated with the second virtual agent. (“... the virtual coach: (a) appears cartoon-like (rather than humanlike); or (b) in the middle of a mock job interview or other mock social interaction, provides feedback on how well the user is doing (instead of waiting until after the mock interaction to provide that feedback). For example, if the interaction is going well, the virtual coach may continue to look interested. But if the user does something odd, (e.g., monologue, poor eye contact), then the coach may provide subtle cues of being disinterested.”; Hoque, [0083])
The virtual coach appears cartoon-like and provides feedback (third data) to the client device (receive, from the one or more first computing devices, third data) on how well the user is doing in the mock job interview. If the interaction is going well, the virtual coach continues to look interested (third data representative of one or more first interactions associated with the first virtual agent), but if the user does something odd, the virtual coach provides clues of being disinterested (third data representative of one or more first interactions associated with the first virtual agent).
Re: claim 7, Hoque, Kita and Whitney teach
(Currently Amended) The system of claim 1, wherein at least one of: the first graphical representation of the first virtual agent is generated based at least on third data representative of one or more first interactions associated with the first virtual agent; or the second graphical representation of the second virtual agent is generated based at least on fourth data representative of one or more second interactions associated with the second virtual agent. (“... the virtual coach: (a) appears cartoon-like (rather than humanlike); or (b) in the middle of a mock job interview or other mock social interaction, provides feedback on how well the user is doing (instead of waiting until after the mock interaction to provide that feedback). For example, if the interaction is going well, the virtual coach may continue to look interested. But if the user does something odd, (e.g., monologue, poor eye contact), then the coach may provide subtle cues of being disinterested.”; Hoque, [0083])
The virtual coach appears cartoon-like (the first graphical representation of the first virtual agent is generated) and provides feedback (third data representative of one or more first interactions associated with the first virtual agent) on how well the user is doing in the mock job interview. If the interaction is going well, the virtual coach continues to look interested (third data representative of one or more first interactions associated with the virtual agent), but if the user does something odd, the virtual coach provides clues of being disinterested (third data representative of one or more first interactions associated with the virtual agent).
Re: claim 8, Hoque, Kita and Whitney teach
8. (Currently Amended) The system of claim 1, wherein the one or more first computing devices include at least one separate device as the one or more second computing devices. (“Fig. 9A schematically illustrates an overall system view depicting multiple user devices interacting with each other. The computing environment 900 includes user devices 930a, 930b, 930c. The user devices 930a, 930b, 930c can communicate with each other through a network 990. The user devices 930a-930c can each include a network interface to communicate via the network 990 with a remote computing system 920... The computing environment 900 can also include one or more remote computing systems 920. The remote computing system 920 may include server computer systems that are clustered and located at different geographic locations. The user devices 930a, 930b, and 930c may communicate with the remote computing system 920 via the network 990.”; Whitney, [0118], Fig. 9A)
Fig. 9A illustrates multiple user devices 930a, 930b, 930c that communicate, via the network, with each other and with one or more remote computing systems (separate device), which include server computer systems. For example user device 930b is considered to be a first computing device and user device 930c is considered to be a second computing device. Therefore, it would have been obvious to one of ordinary skill in the art at the time of effective filing date to modify the method of Hoque by adding the feature of the one or more first computing devices include at least one
Re: claim 9, Hoque, Kita and Whitney teach
9. (Original) The system of claim 1, wherein the system is comprised in at least one of: a control system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing collaborative content creation for 3D assets; a system that provides one or more cloud gaming applications; a system for performing one or more deep learning operations; a system implemented using an edge device; a system for performing one or more generative AI operations; a system for performing one or more conversational AI operations; a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system incorporating one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources. (“For example, the image of the character in which the artificial intelligence (for example, the artificial intelligence agent) which is the communication partner... is displayed, and a lip synchronization operation in which the character performs sound conversation, expression, gesture, or the like is displayed by an animation method or the like of computer graphic synthesis, according to the conversation or the answer of the artificial intelligence.”; Kita, [0051])
An animated image of the character of the artificial intelligence agent is displayed and a lip synchronization operation is performed (conversational AI operation) such that the character performs a sound conversation, for example, to answer a question from the user (a system for performing one or more conversational AI operations). Therefore, it would have been obvious to one of ordinary skill in the art at the time of effective filing date to modify the method of Hoque by adding the feature of the system is comprised in at least one of: a control system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing collaborative content creation for 3D assets; a system that provides one or more cloud gaming applications; a system for performing one or more deep learning operations; a system implemented using an edge device; a system for performing one or more generative AI operations; a system for performing one or more conversational AI operations; a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system incorporating one or more virtual machines (VMs);a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources, in order to enable the AI selection processing to select an artificial intelligence agent, which communicates with the user in response to user requests, as taught by Kita. ([0022])
Claim(s) 5 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hoque in view of Kita and Whitney as applied to claim 1 above, and further in view of Nasar et al. U.S. Pub. No. 2020/0365160.
Re: claims 5 and 15 (which are rejected under the same rationale), Hoque and Kita are silent regarding the one or more processors are further to at least one of: generate the first graphical representation using at least one of a first virtual camera that captures the first virtual agent or a first virtual microphone that captures first audible output from the first virtual agent; or generate the second graphical representation using at least one of a second virtual camera that captures the second virtual agent or a second virtual microphone that captures second audible output from the second virtual agent, however, Nassar teaches
(Currently Amended) The system of claim 1, wherein the one or more processors are further to at least one of: the first graphical representation is generated using at least one of a first virtual camera that captures the first virtual agent or a first virtual microphone that captures first audible output from the first virtual agent; or the second graphical representation is generated using at least one of a second virtual camera that captures the second virtual agent or a second virtual microphone that captures second audible output from the second virtual agent. (“Interactive virtual meeting assistant 132 may also generate visual, sound-based, and/or text-based responses to the commands that are outputted over a virtual webcam, virtual microphone, and/or virtual keyboard.”; Nassar, [0021], Fig. 1)
Fig. 1 illustrates an interactive virtual meeting assistant that generates, for example, visual and sound-based responses (the first graphical representation is generated) to commands that are outputted over a virtual webcam (using at least one first virtual camera that captures the first virtual agent) and a virtual microphone (first virtual microphone that captures the first audible output from the first virtual agent). Therefore, it would have been obvious to one of ordinary skill in the art at the time of the effective filing date, to modify the method of Hoque, by adding the feature of the one or more processors are further to at least one of: performance of interactive virtual assistants, as taught by Nassar ([0006]).
Claim(s) 10; 11, 12, 13, 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hoque in view of Whitney.
Re: claim 10, Hoque teaches
10. (Currently Amended) A method comprising: generating, by a system that is separate from one or more first computing devices that are hosting an instance of a first application, a first graphical representation of a first virtual agent corresponding to an instance of a first application that is hosted using a first system; (“Through the network 409, the user’s computer 401 may interface with servers (e.g., 411, 413) that control the virtual coach and UI. The servers (e.g., 411, 413) may store data in, and retrieve data from, memory devices (e.g., one or more hard drives) 412, 414.”; Hoque, [0071], Fig. 4)
Fig. 4 illustrates that the user’s computer communicates with servers, that control the virtual coach, over the network (generating, by a system that is separate from one or more first computing devices that are hosting an instance of a first application, a first graphical representation of the first virtual agent corresponding to an instance of a first application that is hosed using a first system).
(“... virtual coaching is provided remotely over the Internet. The processors that control the virtual coach and UI may be located on one or more servers (or a computer connected to one or more servers) that are remote from the user. A display screen and speakers at the user’s location may display the virtual coach and UI.”; Hoque, [0079])
The virtual coach is provided remotely over the internet (generating, by a system that is separate from one or more first computing devices that are hosting an instance of a first application, a first graphical representation of the first virtual agent corresponding to an instance of a first application that is hosed using a first system).
sending, to the one or more first computing devices, first data corresponding to the first graphical representation of the first virtual agent, the first graphical representation of the first virtual agent being included in a first stream of data of the instance of the first application for presentation on one or more first client devices that are communicating with the one or more first computing devices; (“... virtual coaching is provided remotely over the Internet. The processors that control the virtual coach and the UI may be located on one or more servers... that are remote from the user. A display screen and speakers at the user’s location may display the virtual coach and UI.”; Hoque, [0079])
Virtual coaching is provided remotely over the internet. Fig. 4 illustrates network 409 and servers 411, 413 that control the virtual coach and communicate with the computer 401. The servers (instance of the first application) control the virtual coach and display the virtual coach at the user’s device.
(“Depending on the particular implementation of this invention, feedback to a human user can be provided in many different ways, including by a virtual coach speaking or by a UI visual display.”; Hoque, [0079])
Feedback is provided (sending, to the one or more first computing devices) to the user, for example, of the displaying the virtual coach speaking (first data corresponding to a first graphical representation of the first virtual agent). .
(“... one or more electronic processors are specially adapted:... (3) to compute behaviors of a virtual coach, and to compute the audiovisual animations of that behavior; (4) to control an audio-visual display of a user interface; ... The processor may be located in any position or positions within or outside of the automated conversation coach. For example:... (b) at least some of the processors may be remote from other components of the automatic conversation coach. The processors may be connected to each other or to other components in the automatic conversation coach either: (a) wirelessly... For example, one or more electronic processors may be housed in a computer (including computer 400 or servers 411, 413 in Fig. 4).”; Hoque, [0085])
The processors at the servers compute behaviors, audiovisual animations of the virtual coach and control (sending, to the one or more first computing devices) an audiovisual (first stream of data) display of the virtual coach (the first graphical representation of the first virtual agent being included in a first stream of data) via network 409, at the user’s device (for presentation on one or more first client devices that are communicating with the one or more first computing devices).
Hoque is silent regarding generating, by the system that is separate from one or more second computing devices that are hosting an instance of a second application, a second graphical representation of a second virtual agent corresponding to
generating, by the system that is separate from one or more second computing devices that are hosting an instance of a second application, a second graphical representation of a second virtual agent corresponding to the instance of the second application ; (“The wearable system 600 can comprise an avatar processing and rendering system 690. The avatar processing and rendering system 690 can be configured to generate, update, animate, and render an avatar based on contextual information. Some or all of the avatar processing and rendering system 600 can be implemented as part of the local processing and data module 260 or the remote processing module 270 alone or in combination...”; Whitney, [0103], Figs. 2, 6A and 9A)
Fig. 9A illustrates an avatar processing and rendering system that is implemented as the remote processing module 270. The avatar processing and rendering system generates and renders an avatar.
(“For example, a virtual avatar may represent a real person or may represent a non-user character, such as a virtual assistant that is configured to interface with users... multiple avatar processing and rendering systems 690 (e.g., as implemented on different wearable devices) can be used for rendering the virtual avatar 670. For example, a first user’s wearable device may be used to determine the first user’s intent, while a second user’s wearable device can determine an avatar’s characteristics and render the avatar of the first user based on the intent received from the first user’s wearable device. The first user’s wearable device and the second user’s wearable device... can communicate via a network...”; Whitney, [0033], [0103], Fig. 6A)
Fig. 6A illustrates multiple avatar processing and rendering systems 690, implemented on different wearable devices, that are used for rendering the virtual avatar (virtual assistant). For example, a first user’s wearable device is used to determine the first user’s intent, and the second user’s wearable device renders the first user’s avatar based on the first user’s intent. The first user’s wearable device and the second user’s wearable device communicate via a network.
(“Fig. 9A schematically illustrates an overall system view depicting multiple user devices interacting with each other. The computing environment 900 includes user devices 930a, 930b, 930c. The user devices 930a, 930b, 930c can communicate with each other through a network 990. The user devices 930a-930c can each include a network interface to communicate via the network 990 with a remote computing system 920... The computing environment 900 can also include one or more remote computing systems 920. The remote computing system 920 may include server computer systems that are clustered and located at different geographic locations. The user devices 930a, 930b, and 930c may communicate with the remote computing system 920 via the network 990.”; Whitney, [0118], Fig. 9A)
Fig. 9A illustrates multiple user devices 930a, 930b, 930c that communicate, via the network, with each other and with one or more remote computing systems 920, which include server computer systems.
(“The remote computing system 920 may include a remote data repository 980 which can maintain information about a specific user’s physical and/or virtual worlds. Data storage 980 can store information related to users, user’s environment... or configurations of avatars of the users... The remote processing module 970 may include one or more processors which can communicate with the user devices (930a, 930b, 930c) and the remote data repository 980.”; Whitney, [0119])
The remote computing system (system that is separate from one or more second computing devices that are hosting an instance of a second application) includes a remote data repository/data storage, which stores configurations of user’s avatars (hosting an instance of a second application). And, the remote processing module includes processors which communicate with user devices and the remote repository.
(“A wearable device can use information acquired of a first user and the environment to animate a virtual avatar that will be rendered by a second user’s wearable device to create a tangible sense of presence of the first user in the second user’s environment. For example, the wearable devices 902 and 904, the remote computing system 920, alone or in combination, may process Alice’s images or movements for presentation by Bob’s wearable device 904 or may process Bob’s images or movements for presentation by Alice’s wearable device 902... the avatars can be rendered based on contextual information such as, e.g., a user’s intent, an environment of the user or an environment in which the avatar is rendered...”; Whitney, [0126])
A second user’s wearable device uses information obtained from a first user to animate a first virtual avatar that is rendered by the second user’s wearable device. For example, the second user’s device uses information from a first user to render a virtual avatar on the second user’s device and the first user’s device uses information from a second user to render a second virtual avatar that is rendered on the first users device (generating... a second graphical representation of a second virtual agent corresponding to the instance of the second application).
and sending, to the one or more second computing devices, second data corresponding to the second graphical representation of the second virtual agent, the second graphical representation of the second virtual agent being included in a second stream of data of the instance of the second application for presentation on one or more second client devices that are communicating with the one or more second computing devices. (“The local processing and data module 260 may be operatively coupled by communication links 262 or 264, such as via wired or wireless communication links, to the remote processing module 270 or remote data repository 280 such that these remote modules are available as resources to the local processing and data module 260. In addition, remote processing module 270 and remote data repository 280 may be operatively coupled to each other.”; Whitney, [0045], Fig. 2)
Fig. 2 illustrates that the wearable system 200 includes local processing and data module that is in wireless communication with the remote processing module and the remote repository. The remote processing module and the remote repository are also in wireless communication with each other.
(“The wearable system 600 can comprise an avatar processing and rendering system 690. The avatar processing and rendering system 690 can be configured to generate, update, animate, and render an avatar based on contextual information. Some or all of the avatar processing and rendering system 600 can be implemented as part of the local processing and data module 260 or the remote processing module 270 alone or in combination...”; Whitney, [0103], Figs. 2 and 6A)
Fig. 9A illustrates an avatar processing and rendering system that is implemented as the remote processing module 270. The avatar processing and rendering system generates and renders an avatar.
(“Fig. 9A schematically illustrates an overall system view depicting multiple user devices interacting with each other. The computing environment 900 includes user devices 930a, 930b, 930c. The user devices 930a, 930b, 930c can communicate with each other through a network 990. The user devices 930a-930c can each include a network interface to communicate via the network 990 with a remote computing system 920... The computing environment 900 can also include one or more remote computing systems 920. The remote computing system 920 may include server computer systems that are clustered and located at different geographic locations. The user devices 930a, 930b, and 930c may communicate with the one or more remote computing system 920 via the network 990.”; Whitney, [0118], Fig. 9A)
Fig. 9A illustrates multiple user devices 930a, 930b, 930c that communicate, via the network, with each other and with one or more remote computing systems 920, which include server computer systems. The wireless communication between the user devices and the remote computing systems is via a network.
(“The remote computing system 920 may include a remote data repository 980 which can maintain information about a specific user’s physical and/or virtual worlds. Data storage 980 can store information related to users, user’s environment... or configurations of avatars of the users... The remote processing module 970 may include one or more processors which can communicate with the user devices (930a, 930b, 930c) and the remote data repository 980.”; Whitney, [0119])
The remote computing system includes a remote data repository/data storage, which stores configurations of user’s avatars. And, the remote processing module includes processors which communicate with user devices and the remote repository. The communication between the remote processing module, the user devices and the remote data repository is wireless and via a network.
(“A wearable device can use information acquired of a first user and the environment to animate a virtual avatar that will be rendered by a second user’s wearable device to create a tangible sense of presence of the first user in the second user’s environment. For example, the wearable devices 902 and 904, the remote computing system 920, alone or in combination, may process Alice’s images or movements for presentation by Bob’s wearable device 904 or may process Bob’s images or movements for presentation by Alice’s wearable device 902... the avatars can be rendered based on contextual information such as, e.g., a user’s intent, an environment of the user or an environment in which the avatar is rendered...”; Whitney, [0126])
A second user’s wearable device uses information obtained from a first user (sending to the one or more second computing devices) to animate a first virtual avatar that is rendered by the second user’s wearable device (second data corresponding to the second graphical representation of the second virtual agent being included in a second stream of data of the instance of the second application for presentation on one or more second client devices that are communicating with the one or more second computing devices). For example, the second user’s device uses information from a first user to render a virtual avatar on the second user’s device and the first user’s device uses information from a second user to render a second virtual avatar that is rendered on the first users device. Whitney is combined with Hoque and Kita such that the virtual avatar of Whitney are the AI agent of Kita and the virtual coach of Hoque and computing environment 900 of Whitney is included in the computing environments of Kita and Hoque. Therefore, it would have been obvious to one of ordinary skill in the art at the time of effective filing date to modify the method of Hoque by adding the feature of generating, by the system that is separate from one or more second computing devices that are hosting an instance of a second application, a second graphical representation of a second virtual agent corresponding to the instance of the second application, and sending, to the one or more second computing devices, second data corresponding to the second graphical representation of the second virtual agent, the second graphical representation of the second virtual agent being included in a second stream of data of the instance of the second application for presentation on one or more second client devices that are communicating with the one or more second computing devices, in order to display a virtual assistant configured to assist the user with contextual objects and suggestions depending on what virtual content the user is interacting with, as taught by Whitney ([0033]).
Claim 11 is a method analogous to the system of claim 2, is similar in scope and is rejected under the same rationale.
Re: claim 12, Hoque and Whitney teach
12. (Currently Amended) The method of claim 10, further comprising at least one of: receiving, from the one or more first computing device, a first selection of the first virtual agent for the instance of the first application, wherein the generating the first graphical representation of the first virtual agent is based at least on the first selection; or receiving, from the second system, a second selection of the second virtual agent for the instance of the second application, wherein the generating the second graphical representation of the second virtual agent is based at least on the second selection. (“... the human user may select a particular persona for a virtual coach, out of a set of personas. For example, the human user may select a virtual coach who appears to be female or a virtual coach who appears to be male.”; Hoque, [0012])
The user selects a persona, from a set of personas, for the virtual coach (receiving, from the one or more first computing device, a selection of the first virtual agent).
(“For example, a user could use the virtual coach for a mock job interview as follows: (In this example... the human user is named Mike and the virtual coach is named John). Mike chooses to interact with an animated coach to improve his interview skills. Mike chooses one of two virtual counselors, John, who appears on the screen, greets him, and starts the interview.”; Hoque, [0077])
For example, for a mock job interview, the user, Mike, selects a virtual coach named John, who appears on the screen, greets Mike and starts the interview (generating the first graphical representation of the first virtual agent is based at least on the first selection).
(“Based on the analysis of the sensor data and social media data 101, the computer then determines the behavior of a virtual coach, using a behavior generation software module 131. The computer causes animations of the virtual coach’s behavior to be displayed visually on the screen 135 and audibly by speakers.”; Hoque, [0064])
Based on the analysis of the sensor data, the computer determines the behavior of the virtual coach, using the behavior generation software module 131 (instance of the first application).
Re: claim 13, Hoque and Whitney teach
13. (Currently Amended) The method of claim 10, further comprising at least one of: sending, to the one or more first computing device, first audio data corresponding to first speech associated with the first virtual agent, wherein the first audio data is further included in the first stream of data; or sending, to the one or more second computing devices, second audio data corresponding to second speech associated with the second virtual agent, wherein the second audio data is further included in the second stream of data. (“... virtual coaching is provided remotely over the Internet. The processors that control the virtual coach and the UI may be located on one or more servers... that are remote from the user. A display screen and speakers at the user’s location may display the virtual coach and UI.”; Hoque, [0079])
Virtual coaching is provided remotely over the internet. Fig. 4 illustrates network 409 and servers 411, 413 that control the virtual coach and communicate with the computer 401. The servers control the virtual coach and displays the virtual coach at the user’s device.
(“Depending on the particular implementation of this invention, feedback to a human user can be provided in many different ways, including by a virtual coach speaking or by a UI visual display.”; Hoque, [0079])
Feedback is provided (sending, to the one or more first computing devices) to the user, for example, of the virtual coach speaking (first audio data corresponding to first speech associated with the first virtual agent).
(“... one or more electronic processors are specially adapted:... (3) to compute behaviors of a virtual coach, and to compute the audiovisual animations of that behavior; (4) to control an audio-visual display of a user interface;”; Hoque, [0085])
The processors at the servers compute behaviors, audiovisual animations of the virtual coach and control (sending, to the one or more first computing devices) an audiovisual (first stream of data) display of the virtual coach (first audio data is further included in the first stream of data) at the user’s device.
Re: claim 14, Hoque and Whitney teach
14. (Currently Amended) The method of claim 10, further comprising at least one of: receiving, from the one or more first computing device, third data representative of one or more first interactions associated with the first virtual agent, wherein the generating the first graphical representation of the first virtual agent is based at least on the third data; or receiving, from the first system, fourth data representative of one or more second interactions associated with the second virtual agent, wherein the generating the second graphical representation of the second virtual agent is based at least on the fourth data. (“... the virtual coach: (a) appears cartoon-like (rather than humanlike); or (b) in the middle of a mock job interview or other mock social interaction, provides feedback on how well the user is doing (instead of waiting until after the mock interaction to provide that feedback). For example, if the interaction is going well, the virtual coach may continue to look interested. But if the user does something odd, (e.g., monologue, poor eye contact), then the coach may provide subtle cues of being disinterested.”; Hoque, [0083])
The virtual coach appears cartoon-like and provides feedback (third data) to the user’s device (receiving, from the one or more first computing device, third data) on how well the user is doing in the mock job interview. If the interaction is going well, the virtual coach continues to look interested (third data representative of one or more first interactions associated with the virtual agent), but if the user does something odd, the virtual coach provides clues of being disinterested (the generating the first graphical representation of the first virtual agent is based at least on the third data).
Claim(s) 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hoque in view of Whitney as applied to claim 10 above, and further in view of Nasar.
Claim 15 is a method analogous to the system of claim 5, is similar in scope and is rejected under the same rationale.
Claim(s) 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hoque in view of Whitney as applied to claim 10 above, and further in view of Bar-Zeev et al. U.S. Pub. No. 2020/0098188.
Re: claim 16, Hoque and Kita are silent regarding at least one of: determining at least one of a first background or one or more first objects associated with the first virtual agent, wherein the generating the first graphical representation is based on the at least one of the first background or the one or more first objects; or determining at least one of a second background or one or more second objects associated with the second virtual agent, however, Bar-Zeev teaches
16. (Currently Amended) The method of claim 10, further comprising at least one of: determining at least one of a first background or one or more first objects associated with the first virtual agent, wherein the generating the first graphical representation is based on the at least one of the first background or the one or more first objects; or determining at least one of a second background or one or more second objects associated with the second virtual agent, wherein the generating the second graphical representation is based at least on the at least one of the second background or the one or more second objects. (“In FIG. 4A, the image data characterizes a view of a street including a restaurant 401 and a street light on a sidewalk. The user 405 looks at the store front of the restaurant 401 and also points at the restaurant sign on the restaurant 401... As used herein, a contextual trigger for a contextual CGR digital assistant includes one or more objects, information... and/or location associated with an available action performed by the contextual CGR digital assistant. For example, in Fig. 4A, multiple factors can trigger the contextual CGR digital assistant that can scout and return with information associated with the restaurant. Such factors can include the user 405 looking at the restaurant 401, voice commands from the user 405 to find out more information about the restaurant 401, the user 405 pointing at the restaurant 401 as shown in Fig. 4A...”; Bar-Zeev, [0065], Figs. 4A-4C)
Fig. 4 illustrates a restaurant, where the user points at the restaurant sign (first object). A contextual trigger for the CGR digital assistant (first virtual agent) includes, for example the user pointing at objects, such as the restaurant and the restaurant sign (one or more first objects associated with the first virtual agent).
(“... a computer-generated reality dog is selected as the visual representation of the contextual CGR digital assistant based at least in part on the context associated with the CGR environment 400, including a contextual meaning corresponding to a dog... Once the dog is selected as the visual representation, as shown in FIG. 4B, a highlight 402 is displayed around the restaurant sign indicating the proximate location of the contextual trigger and a computer-generated dog 404 appears in the CGR environment 400 to assist the information retrieval. The visual representation of the computer-generated dog 404 is composited into the CGR environment 400 as shown in FIG. 4C. ”; Bar-Zeev, [0066])
Fig. 4B illustrates that a computer-generated dog appears as the CGR representation of the digital assistant (generating the first graphical representation) based on the context associated with the CGR environment, which includes the restaurant and the restaurant sign (based on the one or more first objects). Therefore, it would have been obvious to one of ordinary skill in the art at the time of the effective filing date, to modify the method of Hoque by adding the feature of further comprising at least one of: determining at least one of a first background or one or more first objects associated with the first virtual agent, wherein the generating the first graphical representation is based on the at least one of the first background or the one or more first objects; or determining at least one of a second background or one or more second objects associated with the second virtual agent, wherein the generating the second graphical representation is based on the at least one of the second background or the one or more second objects in order to improve the user’s experience by subtly drawing the user’s attention to relevant computer-generated media content, as taught by Bar-Zeev ([0026]).
Claim(s) 17; 18 and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hoque in view of Whitney, Kita and Deac U.S. Pub. No. 2019/0273707.
Re: claim 17, Hoque teaches
17. (Currently Amended) A data center comprising:... wherein one or more components of the data center are to: generate graphical representations of virtual agents for instances of applications that are hosted using computing devices the graphical representations being generated separately from the computing devices hosting the instances of the applications; (“... the human user may select a particular persona for a virtual coach, out of a set of personas. For example, the human user may select a virtual coach who appears to be female or a virtual coach who appears to be male.”; Hoque, [0012])
The user selects a persona, from a set of personas, for the virtual coach.
(“For example, a user could use the virtual coach for a mock job interview as follows: (In this example... the human user is named Mike and the virtual coach is named John). Mike chooses to interact with an animated coach to improve his interview skills. Mike chooses one of two virtual counselors, John, who appears on the screen, greets him, and starts the interview.”; Hoque, [0077])
For example, for a mock job interview, the user, Mike, requests a mock job interview and selects a virtual coach named John, who appears on the screen, greets Mike and starts the interview (generate graphical representations of virtual agents).
(“Based on the analysis of the sensor data and social media data 101, the computer then determines the behavior of a virtual coach, using a behavior generation software module 131. The computer causes animations of the virtual coach’s behavior to be displayed visually on the screen 135 and audibly by speakers.”; Hoque, [0064])
Based on the analysis of the sensor data, the computer determines the behavior of the virtual coach, using the behavior generation software module 131 (instances of the applications hosted using the computing devices). The behavior generation software module is considered to include instances of the applications for the set of virtual coaches.
(“Through the network 409, the user’s computer 401 may interface with servers (e.g., 411, 413) that control the virtual coach and UI. The servers (e.g., 411, 413) may store data in, and retrieve data from, memory devices (e.g., one or more hard drives) 412, 414.”; Hoque, [0071], Fig. 4)
Fig. 4 illustrates that the user’s computer communicates with servers, that control the virtual coach, over the network (generate graphical representations of virtual agents for instances of applications that are hosted using computing devices).
(“... virtual coaching is provided remotely over the Internet. The processors that control the virtual coach and UI may be located on one or more servers (or a computer connected to one or more servers) that are remote from the user. A display screen and speakers at the user’s location may display the virtual coach and UI.”; Hoque, [0079])
The virtual coach is provided remotely over the internet (generate graphical representation of the virtual agents for instances of applications that are hosted using computing devices). Hoque is silent regarding the graphical representations being generated separately from the computing devices hosting the instances of the applications, however, Whitney teaches
(“The wearable system 600 can comprise an avatar processing and rendering system 690. The avatar processing and rendering system 690 can be configured to generate, update, animate, and render an avatar based on contextual information. Some or all of the avatar processing and rendering system 600 can be implemented as part of the local processing and data module 260 or the remote processing module 270 alone or in combination...”; Whitney, [0103], Figs. 2, 6A and 9A)
Fig. 9A illustrates an avatar processing and rendering system that is implemented as the remote processing module 270. The avatar processing and rendering system generates and renders an avatar.
(“For example, a virtual avatar may represent a real person or may represent a non-user character, such as a virtual assistant that is configured to interface with users... multiple avatar processing and rendering systems 690 (e.g., as implemented on different wearable devices) can be used for rendering the virtual avatar 670. For example, a first user’s wearable device may be used to determine the first user’s intent, while a second user’s wearable device can determine an avatar’s characteristics and render the avatar of the first user based on the intent received from the first user’s wearable device. The first user’s wearable device and the second user’s wearable device... can communicate via a network...”; Whitney, [0033], [0103], Fig. 6A)
Fig. 6A illustrates multiple avatar processing and rendering systems 690, implemented on different wearable devices, that are used for rendering the virtual avatar (virtual assistant). For example, a first user’s wearable device is used to determine the first user’s intent, and the second user’s wearable device renders the first user’s avatar based on the first user’s intent.
(“Fig. 9A schematically illustrates an overall system view depicting multiple user devices interacting with each other. The computing environment 900 includes user devices 930a, 930b, 930c. The user devices 930a, 930b, 930c can communicate with each other through a network 990. The user devices 930a-930c can each include a network interface to communicate via the network 990 with a remote computing system 920... The computing environment 900 can also include one or more remote computing systems 920. The remote computing system 920 may include server computer systems that are clustered and located at different geographic locations. The user devices 930a, 930b, and 930c may communicate with the remote computing system 920 via the network 990.”; Whitney, [0118], Fig. 9A)
Fig. 9A illustrates multiple user devices 930a, 930b, 930c that communicate, via the network, with each other and with one or more remote computing systems 920, which include server computer systems.
(“The remote computing system 920 may include a remote data repository 980 which can maintain information about a specific user’s physical and/or virtual worlds. Data storage 980 can store information related to users, user’s environment... or configurations of avatars of the users... The remote processing module 970 may include one or more processors which can communicate with the user devices (930a, 930b, 930c) and the remote data repository 980.”; Whitney, [0119])
The remote computing system (separate from one or more computing devices hosting instances of applications) includes a remote data repository/data storage, which stores configurations of user’s avatars (hosting instances of applications). And, the remote processing module includes processors which communicate with user devices and the remote repository.
(“A wearable device can use information acquired of a first user and the environment to animate a virtual avatar that will be rendered by a second user’s wearable device to create a tangible sense of presence of the first user in the second user’s environment. For example, the wearable devices 902 and 904, the remote computing system 920, alone or in combination, may process Alice’s images or movements for presentation by Bob’s wearable device 904 or may process Bob’s images or movements for presentation by Alice’s wearable device 902... the avatars can be rendered based on contextual information such as, e.g., a user’s intent, an environment of the user or an environment in which the avatar is rendered...”; Whitney, [0126])
A second user’s wearable device uses information obtained from a first user to animate a first virtual avatar that is rendered by the second user’s wearable device. For example, the second user’s device uses information from a first user to render a virtual avatar on the second user’s device and the first user’s device uses information from a second user to render a second virtual avatar that is rendered on the first users device (graphical representations being generated separately from the computing devices hosting the instances of the applications). Therefore, it would have been obvious to one of ordinary skill in the art at the time of effective filing date to modify the method of Hoque by adding the feature of generate graphical representations of virtual agents for instances of applications that are hosted using computing devices the graphical representations being generated separately from the computing devices hosting the instances of the applications, in order to display a virtual assistant configured to assist the user with contextual objects and suggestions depending on what virtual content the user is interacting with, as taught by Whitney ([0033]).
Hoque is silent regarding one or more central processing units (CPUs)... send, to the computing devices, data corresponding to the graphical representations of the virtual agents, the graphical representations of the virtual agents being included in streams of data for presentation using client devices that are communicating with the computing devices, however, Kita teaches
... one or more central processing units (CPUs);... (“As illustrated in Fig. 1, the information processing apparatus 1 includes a central processing unit (CPU) 11... ”; Kita, [0013], Fig. 1)
Fig. 1 illustrates that the information processing apparatus includes a CPU.
and send, to the computing devices, data corresponding to the graphical representations of the virtual agents, the graphical representations of the virtual agents being included in streams of data associated with the instance of the applications for presentation using client devices that are communicating with the computing devices. (“Thus, it is possible to select an optimal artificial intelligence for each user according to the evaluation point based on a sense or sensibility of each of the users, without selecting the artificial intelligence by each of the users according to a unified standard... it is possible to select the optimal artificial intelligence for each of the users, according to the evaluation point based on the sense or the sensibility of each of the users.”; Kita, [0039])
The optimal AI agent is selected for each user (which includes other user’s device and other virtual agents), according to the evaluation point based on the sense or sensibility of each user.
(“... virtual coaching is provided remotely over the Internet. The processors that control the virtual coach and the UI may be located on one or more servers... that are remote from the user. A display screen and speakers at the user’s location may display the virtual coach and UI.”; Hoque, [0079])
Virtual coaching is provided remotely over the internet. Fig. 4 illustrates network 409 and servers 411, 413 that control the virtual coach and communicate with the computer 401. The servers control the virtual coach and displays the virtual coach at the user’s device.
(“Depending on the particular implementation of this invention, feedback to a human user can be provided in many different ways, including by a virtual coach speaking or by a UI visual display.”; Hoque, [0079])
Feedback is provided (send, to the computing devices) to the user, for example, of the displayed virtual coach speaking (data corresponding to the graphical representations of the virtual agents).
(“... one or more electronic processors are specially adapted:... (3) to compute behaviors of a virtual coach, and to compute the audiovisual animations of that behavior; (4) to control an audio-visual display of a user interface;”; Hoque, [0085])
The processors at the servers compute behaviors, audiovisual animations of the virtual coach and control (send, to the computing devices) an audiovisual streams of data associated with the instance of the applications) display of the virtual coach (the graphical representations of the virtual agents begin included in streams of data for presentation using client devices that are communicating with the computing devices.
Hoque is silent regarding, one or more graphics processing units (GPUs); one or more data processing units (DPUs), however, Deac teaches
... one or more graphics processing units (GPUs); one or more data processing units (DPUs); (“the message support definition generator 234 generates the intelligent agents that are loaded into each of the selected user computer devices...”; Deac, [0137])
Fig. 2 illustrates a support definition generator that generates intelligent agents.
(“Embodiments of the invention may be implemented using specifically designed hardware, configurable hardware, programmable data processors configured by the provision of software... capable of executing on the data processors... Examples of programmable data processors are: microprocessors... graphics processors... For example, one or more data processors in a control circuit for a device may implement methods described herein by executing software instructions in a program memory accessible to the processors.”; Deac, [0483])
The invention is implemented using one or more programmable data processors (DPUs), such as microprocessors (CPUs) and graphics processors (GPUs). Therefore, it would have been obvious to one of ordinary skill in the art at the time of the effective filing date, to modify the method of Hoque by adding the feature of one or more graphics processing units (GPUs); one or more data processing units (DPUs), in order to send and receive messages faster and become easier to use, as taught by Deac ([0013]).
Re: claim 18, Hoque, Whitney, Kita and Deac teach
18. (Original) The data center of claim 17, wherein the one or more components of the data center are further to: receive, from the computing devices, requests associated with the virtual agents; and initiate, based at least on the requests, the virtual agents for the instances of the applications hosted using the computing devices. (“For example, a user could use the virtual coach for a mock job interview as follows: (In this example... the human user is named Mike and the virtual coach is named John). Mike chooses to interact with an animated coach to improve his interview skills. Mike chooses one of two virtual counselors, John, who appears on the screen, greets him, and starts the interview.”; Hoque, [0077])
For example, for a mock job interview, the user, Mike, requests a mock job interview and selects (requests) a virtual coach named John (receive, from the computing devices, requests associated with the virtual agents), who appears on the screen, greets Mike and starts the interview (initiate, based at least on the first request).
(“Based on the analysis of the sensor data and social media data 101, the computer then determines the behavior of a virtual coach, using a behavior generation software module 131. The computer causes animations of the virtual coach’s behavior to be displayed visually on the screen 135 and audibly by speakers.”; Hoque, [0064])
Based on the analysis of the sensor data, the computer determines the behavior of the virtual coach, using the behavior generation software module 131 (instances of the applications hosted using the computing devices). The behavior generation software module is considered to include instances of the applications for the set of virtual coaches.
Claim 20 is a data center analogous to the system of claim 9, is similar in scope and is rejected under the same rationale.
Re: claim 19, Claim(s) 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hoque in view of Whitney, Kita and Deac as applied to claim 17 above, and further in view of Nasar.
Re: claim 19, Hoque, Kita and Deac are silent regarding the generation of the graphical representations of the virtual agents uses at least one of virtual cameras that capture the virtual agents or virtual microphones that capture audible outputs from the virtual agents, however, Nasar teaches
19. (Original) The data center of claim 17, wherein the generation of the graphical representations of the virtual agents uses at least one of virtual cameras that capture the virtual agents or virtual microphones that capture audible outputs from the virtual agents. (“Interactive virtual meeting assistant 132 may also generate visual, sound-based, and/or text-based responses to the commands that are outputted over a virtual webcam, virtual microphone, and/or virtual keyboard.”; Nassar, [0021], Fig. 1)
Fig. 1 illustrates an interactive virtual meeting assistant that generates, for example, visual and sound-based responses (generation of the graphical representations of the virtual agents) to commands that are outputted over a virtual webcam (uses at least one first virtual camera that capture the virtual agent) and a virtual microphone (virtual microphones that capture the audible outputs from the virtual agents). Therefore, it would have been obvious to one of ordinary skill in the art at the time of the effective filing date, to modify the method of Hoque, by adding the feature of the generation of the graphical representations of the virtual agents uses at least one of virtual cameras that capture the virtual agents or virtual microphones that capture audible outputs from the virtual agents, in order to provide technological improvements in the interactivity, functionality, and performance of interactive virtual assistants, as taught by Nassar ([0006]).
Response to Arguments
Applicant’s arguments, see Amendment/Request for Reconsideration-After Non-Final Rejection, filed 6/12/2026, with respect to the Claim Objection to claim 14 have been fully considered and are persuasive. The Claim Objection of the previous Office Action has been withdrawn.
Applicant’s arguments, see Amendment/Request for Reconsideration-After Non-Final Rejection, filed 6/12/2026, with respect to the Objection to the Specification have been fully considered and are persuasive. The Objection to the Specification of the previous Office Action has been withdrawn.
Applicant’s arguments, see Amendment/Request for Reconsideration-After Non-Final Rejection, filed 6/12/2026, with respect to the rejection(s) of claim(s) 1 under 35 U.S.C § 103 have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of Hoque, Kita and Whitney.
“Applicant respectfully submits that Hoque and Kita, whether taken alone or in combination, do not teach or suggest, at least, "initiat[ing] a first virtual agent corresponding to an instance of a first application that is hosted using one or more first computing devices; generat[ing], remotely from the one or more first computing devices hosting the instance of the first application, first data corresponding to a first graphical representation of the first virtual agent; [and] send[ing] the first data to the one or more first computing devices using one or more wireless networks, the first graphical representation of the first virtual agent being included in a first stream of data of the instance of the first application for presentation on one or more first client devices that are communicating with the one or more first computing devices," as amended claim 1 recites... As shown, Hoque initially describes that a computer determines the behavior of a virtual coach using a behavior generation software module. Id. Hoque then describes that the computer causes animation of the virtual coach's behavior to be displayed on a screen. Id. The Office appears to recite the behavior generation software module of Hoque as allegedly teaching the "instance of the first application" of independent claim 1. However, even if the behavior generation software module of Hoque does teach the "instance of the first application" of amended claim 1, which Applicant does not concede, Hoque still does not teach or suggest that a separate "system" initiates the virtual coach and also generate the animation of the virtual coach separate from the behavior generation software module and/or a computer executing the behavior generation software module... As shown, Hoque further describes that the virtual coaching may be provided remotely over the internet, such as on one or more servers. Id. However, Hoque again does not teach or suggest that, in this implementation, the servers are separate from "one or more ... computing devices" that are hosting an "application" that displays the virtual coach. Rather, in this implementation, the servers would be hosting the behavior generation software module. Therefore, Hoque again does not teach or suggest that a separate "system" initiates the virtual coach and also generate the animation of the virtual coach separate from the behavior generation software module and/or the servers executing the behavior generation software module....”
Examiner agrees. A new rejection is set forth necessitated by the amendment. Whitney teaches this amended limitation. Hoque and Kita are combined with Whitney, which teaches separate systems that initiate the virtual avatar. Whitney illustrates in Fig. 9A, an avatar processing and rendering system that is implemented as the remote processing module 270 (separate system that initiates the virtual avatar). The avatar processing and rendering system generates and renders an avatar. (Whitney, [0103], Figs. 2, 6A and 9A). Whitney illustrates in Fig. 6A, multiple avatar processing and rendering systems 690, implemented on different wearable devices (separate system that initiates the virtual avatar), that are used for rendering the virtual avatar (virtual assistant). For example, a first user’s wearable device is used to determine the first user’s intent, and the second user’s wearable device renders the first user’s avatar based on the first user’s intent. The first user’s wearable device and the second user’s wearable device communicate via a network. (Whitney, [0033], [0103], Fig. 6A). Fig. 6A illustrates multiple avatar processing and rendering systems 690, implemented on different wearable devices (separate system that initiates the virtual avatar), that are used for rendering the virtual avatar (virtual assistant). For example, a first user’s wearable device is used to determine the first user’s intent, and the second user’s wearable device renders the first user’s avatar based on the first user’s intent. The first user’s wearable device and the second user’s wearable device communicate via a network. (Whitney, [0033], [0103], Fig. 6A).
Applicant argues:
“Kita does not remedy the deficiencies of Hoque. In the rejection of previously presented independent claim 1, the Office cites Kita as allegedly teaching, "initiate a second virtual agent corresponding to an instance of a second application that is hosted using one or more second computing devices; and send, to the one or more second computing devices and using the one or more wireless networks, second data corresponding to a second graphical representation of the second virtual agent, the second graphical representation of the second virtual agent being included in a second stream of data for presentation on one or more second client devices that are communicating with the one or more second computing devices... As shown, Kita describes that a CPU 11 may output an image of a character to a user. Id., paras. [0022], [0050], and [0051]. However, Kita does not teach or suggest that the CPU 11 that performs the AI processing is separate from "one or more ... computing devices" that are hosting an "application" that uses the character. Additionally, Kita does not teach or suggest that a separate "system" initiates the character independently from the CPU 11. Consequently, the combination of Hoque and Kita does not teach or suggest "initiat[ing] a first virtual agent corresponding to an instance of a first application that is hosted using one or more first computing devices; generat[ing], remotely from the one or more first computing devices hosting the instance of the first application, first data corresponding to a first graphical representation of the first virtual agent; [and] send[ing] the first data to the one or more first computing devices using one or more wireless networks, the first graphical representation of the first virtual agent being included in a first stream of data of the instance of the first application for presentation on one or more first client devices that are communicating with the one or more first computing devices," as amended claim 1 recites.”
Examiner disagrees. As discussed immediately above, Hoque and Kita are combined with Whitney to teach the amended limitation, which includes a separate systems for initiating the virtual avatars.
Applicant's arguments filed 6/12/2026 have been fully considered but they are not persuasive. Applicant argues regarding claims 5 and 15:
“As described above, the Hoque and Kita references fail to teach or suggest each and every feature of independent claims 1 and 10. Further, the Nasar reference fails to overcome the deficiencies described above with respect to independent claims 1 and 10, nor was it cited to for doing so. As such, the combination of the references used under this section to teach or suggest each and every feature of dependent claims 5 and 15. Accordingly, Applicant respectfully requests withdrawal of the 35 U.S.C. § 103 rejections of claims 5 and 15.”
Examiner disagrees. Claims 1, 10 and claims 5 and 15 have been rejected. Please see the corresponding rejections.
Applicant's arguments filed 6/12/2026 have been fully considered but they are not persuasive. Applicant argues regarding claim 8:
“As described above, the Hoque and Kita references fail to teach or suggest each and every feature of independent claim 1. Further, the Honda reference fails to overcome the deficiencies described above with respect to independent claim 1, nor was it cited to for doing so. As such, the combination of the references used under this section to teach or suggest each and every feature of dependent claim 8. Accordingly, Applicant respectfully requests withdrawal of the 35 U.S.C. § 103 rejection of claim 8.”
Examiner disagrees. Claims 1 and 17 have been rejected. Please see the corresponding rejections.
Applicant's arguments filed 6/12/2026 have been fully considered but they are not persuasive. Applicant argues regarding claim 16:
“As described above, the Hoque and Kita references fail to teach or suggest each and every feature of independent claim 10. Further, the Honda reference fails to overcome the deficiencies described above with respect to independent claim 10, nor was it cited to for doing so. As such, the combination of the references used under this section to teach or suggest each and every feature of dependent claim 16. Accordingly, Applicant respectfully requests withdrawal of the 35 U.S.C. § 103 rejection of claim 16.”
Examiner disagrees. Claims 1 and 10 are rejected. Please see the corresponding rejections.
Applicant's arguments filed 6/12/2026 have been fully considered but they are not persuasive. Applicant argues regarding claims 17, 18 and 20:
“Applicant respectfully submits that the combination of Hoque and Kita does not teach or suggest the features of amended claim 17 for at least the reasons above with respect to amended independent claim 1. Additionally, Deac does not remedy the deficiencies of Hoque and Kita. As such, amended independent claim 17 is patentably distinguishable over the cited references and withdrawal of the rejection is respectfully requested. Thus, amended independent claim 17, along with each claim depending therefrom rejected under this section, is patentably distinguishable over the cited reference and withdrawal of the rejections is respectfully requested.”
Examiner disagrees. Claims 1 and 17 have been rejected. Please see the corresponding rejections.
Applicant's arguments filed 6/12/2026 have been fully considered but they are not persuasive. Applicant argues regarding claim 19:
“As described above, the Hoque, Kita, and Deac references fail to teach or suggest each and every feature of independent claim 17. Further, the Nasar reference fails to overcome the deficiencies described above with respect to independent claim 17, nor was it cited to for doing so. As such, the combination of the references used under this section to teach or suggest each and every feature of dependent claim 19. Accordingly, Applicant respectfully requests withdrawal of the 35 U.S.C. § 103 rejection of claim 19.”
Examiner disagrees. Claims 17 and 19 have been rejected. Please see the corresponding rejections.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DONNA J RICKS whose telephone number is (571)270-7532. The examiner can normally be reached on M-F 7:30am-5pm EST (alternate Fridays off).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Devona Faulk can be reached on 571-272-7776. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see https://ppair-my.uspto.gov/pair/PrivatePair. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Donna J. Ricks/Examiner, Art Unit 2618
/DEVONA E FAULK/Supervisory Patent Examiner, Art Unit 2618