DETAILED ACTION
Notice of AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant’s amendments and arguments filed 04/22/2026, with respect to claim(s) 1-20 have been fully considered.
35 U.S.C 101 rejections of Claims 1-20 have been withdrawn in view of the amended claims filed on 04/22/2026.
Applicant’s arguments filed 04/22/2026, with respect to claim(s) 1-20, under 35 U.S.C. 102/103 have been fully considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 2, 4, 5, 8-11, 13, 17, 18, 20 and 21 are rejected under 35 U.S.C. 103 as being unpatentable over Babbar et al. ( 20250371289 A1), hereinafter referenced as Babbar, in view of Shubhanshu et al. ( 20250086685 A1), hereinafter referenced as Shubhanshu , further in view of Gelfenbeyn et al. ( 20230351120 A1), hereinafter referenced as Gelfenbeyn.
Regarding Claim 1, Babbar teaches a method comprising:
obtaining, during a session associated with a gaming application, one or more first embeddings associated with the information corresponding one or more images depicting content associated with a gaming application( Babbar: Para.[0039], [0042],Fig. 1, media device 106 can generate one or more embeddings based on one or more signals, such as visual signal ( e.g. image data) in one or more frames of the segment of the media content in the media content server 120, which may store content 122 and metadata 124. Content 122 may include gaming applications) ;
generating one or more second embeddings corresponding to a textual input associated with a first character of the gaming application ( Babbar: Para.[0039], [0042],Fig. 1, media device 106 can generate one or more embeddings based on one or more signals, such as closed caption signal ( e.g. text data) in one or more frames of the segment of the media content in the media content server 120, which may store content 122 and metadata 124. Content 122 may include gaming applications. Para.[0082], audio related text data is related to a character) ;
determining, based at least on comparing the one or more second embeddings with respect to the one or more first embeddings, at least an image of the one or more images that is associated with a current state of the gaming application ( Babbar: Para.[0042],Fig. 1, media device 106 can use the one or more embeddings to determine a category for the segment of the media content that describes, represents, summarizes, classifies, and/or identifies the segment of the media content, the content of the segment of the media content, a context(s) of the content of the segment of the media content, and/or one or more characteristics of the segment of the media content and/or the content of the segment of the media content);
Babbar, while teaching the method of claim 1, fails to explicitly teach the claimed, determining, based at least on one or more language models processing input data associated with the textual input and the image, a textual output [[for ]]that is associated with the textual input and the current state of the gaming application; and causing a second character of the gaming application to output speech associated with the textual output.
However, Shubhanshu does teach the claimed, determining, based at least on one or more language models processing input data associated with the textual input and the image, a textual output [[for ]]that is associated with the textual input and the current state of the [gaming] application ( Shubhanshu: Para.[0058], [ 0081],[0082], Fig. 4, The image 410 captured by the user's device is sent to the online concierge system for detection of relevant items and to answer questions from the customer in the interface 400. In this example, the user enters the question 430 “Which ice cream is which?” in natural language. The online concierge system receives the image and the question from the user, detects items, and uses the large language model to evaluate the user's request and output a natural language response 440 provided for display to the user. Para.[0075], the text-based output is provided in an interface as an interactive chat, such that the user may enter further questions related to the image and the detected item(s) within. As such, multiple inputs and related outputs may be generated as a part of an interaction with the large language model.
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Shubhanshu’s teaching of an online concierge system which assists users in identifying additional information about items in an image, into the system and method for using AI/ML models to generate context-aware metadata for a media content item, taught by Babbar, because, by effective processing of the item information by the language model, would improve user interactions with the online concierge system. (Shubhanshu, Para.[0003],[0066]).
Babbar in view of Shubhanshu, while teaching the method of claim 1, fails to explicitly teach the claimed, determining, based at least on one or more language models processing input data associated with the textual input and the image], a textual output [[for ]]that is associated with the textual input and the current state of the gaming application; and causing a second character of the gaming application to output speech associated with the textual output.
However, Gelfenbeyn does teach the claimed, determining, based at least on one or more language models processing input data associated with the textual input and the image], a textual output [[for ]]that is associated with the textual input and the current state of the gaming application ( Gelfenbeyn: Para.[0029], [0069], [0073], Fig. 6B, The dialogue prompts 636 ( input data) may be provided to a LLM 646 to generate dialogue output 654. The dialogue output 654, the client-side narrative triggers 656, and the animation controls 658 provide the dialogue, the events, the client-side triggers, and the animations that need to be enacted on the client side);
and causing a second character of the gaming application to output speech associated with the textual output ( Gelfenbeyn: Para.[0069], [0074], Fig. 6B, The dialogue output 654, the client-side narrative triggers 656, the animation controls 658, and the voice parameters 644 may be processed using text to speech conversion 660. The output data obtained upon applying the text to speech conversion 660 are sent as a stream to the client 662. The game engine animates the AI character ( second object) based on the received data by instructing the AI character on what to say, how to move, what to enact, and the like).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Gelfenbeyn’s teaching of system and method for observation-based training of an Artificial Intelligence (AI) character model , into the system and method, taught by Babbar in view of Shubhanshu, because, by allowing virtual character models to train to change their interactions with users based on observing interactions between users and virtual characters, would improve user’s experience. (Gelfenbeyn, Para.[0003],[0004]).
Regarding Claim 2, Babbar in view of Shubhanshu, further in view of Gelfenbeyn teach the method of claim 1. Gelfenbeyn further teaches, further comprising: determining an identifier associated with the second character ( Gelfenbeyn: Para. [0065], Fig. 6A illustrates different machine learning models such as goals model 622, which identify which goals need to be activated for the AI character, safety model 624, which identify which unsafe responses need to be filtered out) ,
wherein the determining the textual output is further based at least on the identifier ( Gelfenbeyn: Para.[0062], [0065], Fig. 6A, all different machine learning models, such as goals model 622, safety model 624, which identify different characteristics of the AI character, are configured to process the embeddings 618 ( preprocessed data stream in the form of text and/or embeddings) to recognize what needs to be activated).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Gelfenbeyn’s teaching of system and method for observation-based training of an Artificial Intelligence (AI) character model , into the system and method for using AI/ML models to generate context-aware metadata for a media content item, taught by Babbar in view of Shubhanshu, because, by allowing virtual character models to train to change their interactions with users based on observing interactions between users and virtual characters, would improve user’s experience. (Gelfenbeyn, Para.[0003],[0004]).
Regarding Claim 4, Babbar in view of Shubhanshu, further in view of Gelfenbeyn teach the method of claim 1. Gelfenbeyn further teaches, further comprising obtaining information associated with the gaming application, the information including one or more of: [first information indicating one or more settings associated with the gaming application];
second information indicating one or more locations associated with the gaming application ( Gelfenbeyn: Para.[0095], Fig.9, A scene 902 may be driven by a plurality of parameters. The parameters may include scene and location knowledge 904); [third information indicating one or more tasks associated with the gaming application;]
fourth information associated with the second character ( Gelfenbeyn: Para.[0076], [0079], [0084],[0085], Fig.7A, AI character personality and background description 706, An identity profile 718 specify elements of an AI character ( e.g., role, interests). Voice configuration 738 ( information) may be used to determine the configuration of voice in real-time, which can allow AI characters to show different expressions when pursuing a goal. Dialogue style controls 742 may be used to control a dialogue style of an AI character); [fifth information associated a user of thegaming application];
sixth information indicating one or more actions that occurred with respect to the gaming application ( Gelfenbeyn: Para.[0086], Fig.7A, Goals and actions 746 received from the user may be processed to specify the goals that an AI character has per scene, and then set up the actions that the AI character has available to pursue the goal);
seventh information associated with a context for [[a]] the current state associated with the gaming application( Gelfenbeyn: Para.[0083], Fig.7A, The contextual knowledge 734 may be processed to include information about an environment or context to contextualize pursuit of the goal);[or one or more images corresponding to the interactive application.]
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Gelfenbeyn’s teaching of system and method for observation-based training of an Artificial Intelligence (AI) character model , into the system and method for using AI/ML models to generate context-aware metadata for a media content item, taught by Babbar in view of Shubhanshu, because, by allowing virtual character models to train to change their interactions with users based on observing interactions between users and virtual characters, would improve user’s experience. (Gelfenbeyn, Para.[0003],[0004]).
Regarding Claim 5, Babbar in view of Shubhanshu, further in view of Gelfenbeyn teach the method of claim 1. Babbar further teaches, one or more images depicting the content associated with the gaming application ( Babbar: Para.[0039], [0042],Fig. 1, media device 106 can generate one or more embeddings based on one or more signals, such as visual signal ( e.g. image data) in one or more frames of the segment of the media content in the media content server 120, which may store content 122 and metadata 124. Content 122 may include gaming applications) ;
Regarding Claim 8, Babbar in view of Shubhanshu, further in view of Gelfenbeyn teach the method of claim 1. Gelfenbeyn further teaches, further comprising: determining one or more filters associated with at least one of the textual input, the second character, or thegaming application ( Gelfenbeyn: Para.[0028], filters are determined to be used, such as, prior to sending a request to the LLM, the platform may classify and filter the user questions and messages to change words based on the personalities of AI characters, emotional states of AI characters );
and determining, based at least on the one or more filters, at least a first embeddings ( Gelfenbeyn: Para. [0065],[0066], Fig. 6A, filtering of unsafe response for the AI model can be configured based on the embeddings 618 and safety model 624, event model 630 (second portion));
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Gelfenbeyn’s teaching of system and method for observation-based training of an Artificial Intelligence (AI) character model , into the system and method, taught by Babbar in view of Shubhanshu, because, by allowing virtual character models to train to change their interactions with users based on observing interactions between users and virtual characters, would improve user’s experience. (Gelfenbeyn, Para.[0003],[0004]).
Babbar further teaches, wherein the determining the image that is associated with the current state of the application is based at least on comparing the one or more second embeddings with the at least the first embeddings ( Babbar: Para.[0042],Fig. 1, media device 106 can use the one or more embeddings to determine a category for the segment of the media content that describes, represents, summarizes, classifies, and/or identifies the segment of the media content, the content of the segment of the media content, a context(s) of the content of the segment of the media content, and/or one or more characteristics of the segment of the media content and/or the content of the segment of the media content).
Regarding Claim 9, Babbar in view of Shubhanshu, further in view of Gelfenbeyn teach the method of claim 1. Gelfenbeyn further teaches, wherein the causing the second character of the gaming application to output the speech corresponding to the textual output comprises: generating audio data representative of the speech associated with the textual output ( Gelfenbeyn: Para.[0090], Fig.7B, at block 762, the text to speech conversion model determines how the AI character speaks his lines (audio) to pursue the goal),
and sending, to a client device, the audio data along with image data representative of one or more images corresponding to at least the second character ( Gelfenbeyn: Para.[0090]-[0092], Fig.7B, the outputs obtained in blocks, dialogue output (audio or text) 766, the client side narrative triggers 768, and the animation 770 may be provided to a client 772 (e.g., a client engine, a game engine, a web application, and the like)).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Gelfenbeyn’s teaching of system and method for observation-based training of an Artificial Intelligence (AI) character model , into the system and method, taught by Babbar in view of Shubhanshu, because, by allowing virtual character models to train to change their interactions with users based on observing interactions between users and virtual characters, would improve user’s experience. (Gelfenbeyn, Para.[0003],[0004]).
Regarding Claim 10, Babbar teaches a system comprising: one or more processors to: obtain one or more first embeddings corresponding to one or more sources of information associated with one or more states of an interactive application, the one or more states including at least a current state of the interactive application ( Babbar: Para.[0039], [0042],Fig. 1, media device 106 can generate one or more embeddings based on one or more signals, such as visual signal ( e.g. image data/contextual signal) in one or more frames of the segment of the media content in the media content server 120, which may store content 122 and metadata 124. Content 122 may include gaming applications) ;
generate one or more second embeddings corresponding to a textual input associated with a first object of the interactive application ( Babbar: Para.[0039], [0042],Fig. 1, media device 106 can generate one or more embeddings based on one or more signals, such as closed caption signal ( e.g. text data) in one or more frames of the segment of the media content in the media content server 120, which may store content 122 and metadata 124. Content 122 may include gaming applications) ;
determine, based at least on comparing the one or more second embeddings with respect to the one or more first embeddings, at least a source of information from the one or more the source of information being associated with the current state of the interactive application ( Babbar: Para.[0042],Fig. 1, media device 106 can use the one or more embeddings to determine a category for the segment of the media content that describes, represents, summarizes, classifies, and/or identifies the segment of the media content, the content of the segment of the media content, a context(s) of the content of the segment of the media content, and/or one or more characteristics of the segment of the media content and/or the content of the segment of the media content) ;
Babbar, while teaching the system of claim 10, fails to explicitly teach the claimed, generate input data based at least on the textual input and the the textual input; and cause a second object of the interactive application to output speech associated with the textual output.
However, Shubhanshu does teach the claimed, generate input data based at least on the textual input and the ( Shubhanshu: Para.[0058], [ 0081],[0082], Fig. 4, The image 410 captured by the user's device is sent to the online concierge system for detection of relevant items and to answer questions from the customer in the interface 400. In this example, the user enters the question 430 “Which ice cream is which?” in natural language. The online concierge system receives the image and the question from the user, detects items, and uses the large language model to evaluate the user's request and output a natural language response 440 provided for display to the user. Para.[0075], the text-based output is provided in an interface as an interactive chat, such that the user may enter further questions related to the image and the detected item(s) within. As such, multiple inputs and related outputs may be generated as a part of an interaction with the large language model.
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Shubhanshu’s teaching of an online concierge system which assists users in identifying additional information about items in an image, into the system and method for using AI/ML models to generate context-aware metadata for a media content item, taught by Babbar, because, by effective processing of the item information by the language model, would improve user interactions with the online concierge system. (Shubhanshu, Para.[0003],[0066]).
Babbar in view of Shubhanshu, while teaching the system of claim 10, fails to explicitly teach the claimed, and cause a second object of the interactive application to output speech associated with the textual output.
However, Gelfenbeyn does teach the claimed, and cause a second object of the interactive application to output speech associated with the textual output ( Gelfenbeyn: Para.[0069], [0074], Fig. 6B, The dialogue output 654, the client-side narrative triggers 656, the animation controls 658, and the voice parameters 644 may be processed using text to speech conversion 660. The output data obtained upon applying the text to speech conversion 660 are sent as a stream to the client 662. The game engine animates the AI character ( second object) based on the received data by instructing the AI character on what to say, how to move, what to enact, and the like).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Gelfenbeyn’s teaching of system and method for observation-based training of an Artificial Intelligence (AI) character model , into the system and method for using AI/ML models to generate context-aware metadata for a media content item, taught by Babbar, because, by allowing virtual character models to train to change their interactions with users based on observing interactions between users and virtual characters, would improve user’s experience. (Gelfenbeyn, Para.[0003],[0004]).
Regarding Claim 11, Babbar in view of Shubhanshu, further in view of Gelfenbeyn teach the system of claim 10. Gelfenbeyn further teaches, wherein the one or more processors are further to: determine an identifier associated with the second object( Gelfenbeyn: Para. [0065], Fig. 6A illustrates different machine learning models such as goals model 622, which identify which goals need to be activated for the AI character, safety model 624, which identify which unsafe responses need to be filtered out) ,
wherein the determination of the ( Gelfenbeyn: Para.[0083], Fig. 7A, The contextual knowledge 734 may be processed to include information about an environment or context to contextualize pursuit of the goal).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Gelfenbeyn’s teaching of system and method for observation-based training of an Artificial Intelligence (AI) character model , into the system and method, taught by Babbar in view of Shubhanshu, because, by allowing virtual character models to train to change their interactions with users based on observing interactions between users and virtual characters, would improve user’s experience. (Gelfenbeyn, Para.[0003],[0004]).
Regarding Claim 13, Babbar in view of Shubhanshu, further in view of Gelfenbeyn teach the system of claim 10. Gelfenbeyn further teaches, wherein the one or more sources of information include one or more images depicting at least one of one or more objects associated with the interactive application or one or more areas of an environment associated with the interactive application ( Gelfenbeyn: Para.[0061],[0062], the information can be received from an environment of AI characters with which the user interacts in a game).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Gelfenbeyn’s teaching of system and method for observation-based training of an Artificial Intelligence (AI) character model , into the system and method, taught by Babbar in view of Shubhanshu, because, by allowing virtual character models to train to change their interactions with users based on observing interactions between users and virtual characters, would improve user’s experience. (Gelfenbeyn, Para.[0003],[0004]).
Regarding Claim 17, Babbar in view of Shubhanshu, further in view of Gelfenbeyn teach the system of claim 10. Gelfenbeyn further teaches, wherein the system is comprised in at least one of: [a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing one or more deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for performing operations using one or more vision language models (VLMs);a system for generating synthetic data; a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system incorporating one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.]
a system that provides one or more cloud gaming applications ( Gelfenbeyn: Para.[0029], the platform can be used in game applications);
a system for performing one or more generative Al operations ( Gelfenbeyn: Para.[0030], the platform can be used in generative AI operations, such as generating AI character models);
a system for performing operations using one or more large language models (LLMs) ( Gelfenbeyn: Para.[0028], the system may utilize a LLM in conversations with the users.);
a system for performing one or more conversational Al operations ( Gelfenbeyn: Para.[0030], the AI character model can engage in conversation with the user).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Gelfenbeyn’s teaching of system and method for observation-based training of an Artificial Intelligence (AI) character model , into the system and method, taught by Babbar in view of Shubhanshu, because, by allowing virtual character models to train to change their interactions with users based on observing interactions between users and virtual characters, would improve user’s experience. (Gelfenbeyn, Para.[0003],[0004]).
Regarding Claim 18, Babbar teaches one or more processors comprising: processing circuitry to: obtain one or more first embeddings corresponding to contextual information associated with an interactive application ( Babbar: Para.[0039], [0042],Fig. 1, media device 106 can generate one or more embeddings based on one or more signals, such as visual signal ( e.g. image data/contextual signal) in one or more frames of the segment of the media content in the media content server 120, which may store content 122 and metadata 124. Content 122 may include gaming applications) ;
generate one or more second embeddings corresponding to a textual input associated with a first object of the interactive application( Babbar: Para.[0039], [0042],Fig. 1, media device 106 can generate one or more embeddings based on one or more signals, such as closed caption signal ( e.g. text data) in one or more frames of the segment of the media content in the media content server 120, which may store content 122 and metadata 124. Content 122 may include gaming applications) ;
determine, based at least on comparing the one or more second embeddings with respect to the one or more first embeddings, at least a portion of the contextual information that is associated with a state of the interactive application( Babbar: Para.[0042],Fig. 1, media device 106 can use the one or more embeddings to determine a category for the segment of the media content that describes, represents, summarizes, classifies, and/or identifies the segment of the media content, the content of the segment of the media content, a context(s) of the content of the segment of the media content, and/or one or more characteristics of the segment of the media content and/or the content of the segment of the media content) ;
Babbar, while teaching the processor of claim 18, fails to explicitly teach the claimed, determine, based at least on one or more language models processing input data associated with the textual input and the at least the portion of the contextual information, a textual output for the textual input; and cause a second object of the interactive application to output speech associated with the textual output.
However, Shubhanshu does teach the claimed, determine, based at least on one or more language models processing input data associated with the textual input and the at least the portion of the contextual information, a textual output for the textual input ( Shubhanshu: Para.[0058], [ 0081],[0082], Fig. 4, The image 410 captured by the user's device is sent to the online concierge system for detection of relevant items and to answer questions from the customer in the interface 400. In this example, the user enters the question 430 “Which ice cream is which?” in natural language. The online concierge system receives the image and the question from the user, detects items, and uses the large language model to evaluate the user's request and output a natural language response 440 provided for display to the user. Para.[0075], the text-based output is provided in an interface as an interactive chat, such that the user may enter further questions related to the image and the detected item(s) within. As such, multiple inputs and related outputs may be generated as a part of an interaction with the large language model.
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Shubhanshu’s teaching of an online concierge system which assists users in identifying additional information about items in an image, into the system and method for using AI/ML models to generate context-aware metadata for a media content item, taught by Babbar, because, by effective processing of the item information by the language model, would improve user interactions with the online concierge system. (Shubhanshu, Para.[0003],[0066]).
Babbar in view of Shubhanshu, while teaching the method of claim 1, fails to explicitly teach the claimed, and cause a second object of the interactive application to output speech associated with the textual output.
However, Gelfenbeyn does teach the claimed, and cause a second object of the interactive application to output speech associated with the textual output ( Gelfenbeyn: Para.[0069], [0074], Fig. 6B, The dialogue output 654, the client-side narrative triggers 656, the animation controls 658, and the voice parameters 644 may be processed using text to speech conversion 660. The output data obtained upon applying the text to speech conversion 660 are sent as a stream to the client 662. The game engine animates the AI character ( second object) based on the received data by instructing the AI character on what to say, how to move, what to enact, and the like).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Gelfenbeyn’s teaching of system and method for observation-based training of an Artificial Intelligence (AI) character model , into the system and method, taught by Babbar in view of Shubhanshu , because, by allowing virtual character models to train to change their interactions with users based on observing interactions between users and virtual characters, would improve user’s experience. (Gelfenbeyn, Para.[0003],[0004]).
Claim 20 is a processor claim performing the steps in system claim 17 above and as such, claim 20 is similar in scope and content to claim 17 and therefore, claim 20 is rejected under similar rationale as presented against claim 17 above.
Regarding Claim 21, Babbar in view of Shubhanshu, further in view of Gelfenbeyn teach the one or more processors of claim 18. Gelfenbeyn further teaches, wherein the processing circuitry is further to: determine, based at least on the at least the portion of the contextual information, text describing the state of the interactive application; and generate the input data to represent at least the textual input and the text describing the state of the interactive application ( Gelfenbeyn: Para.[0051], The language model 304 may form a request for the LLM, receive a response from the LLM, and process the response from the LLM to form a response to voice messages of the user. The request for the LLM can include classification and adjustment of the text requests from the integration interface 206, according to the current scene, environmental parameters, an emotional state of the AI character, an emotional state of the user, and current context of the conversation with the user).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Gelfenbeyn’s teaching of system and method for observation-based training of an Artificial Intelligence (AI) character model , into the system and method, taught by Babbar in view of Shubhanshu, because, by allowing virtual character models to train to change their interactions with users based on observing interactions between users and virtual characters, would improve user’s experience. (Gelfenbeyn, Para.[0003],[0004]).
Claims 3, 6, 7, 12, 14, 15 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Babbar et al. ( 20250371289 A1), hereinafter referenced as Babbar, in view of Shubhanshu et al. ( 20250086685 A1), hereinafter referenced as Shubhanshu , further in view of Gelfenbeyn et al. ( 20230351120 A1), hereinafter referenced as Gelfenbeyn, further in view of Ahafonov et al. ( US 20240355010 A1), hereinafter referenced as Ahafonov.
Regarding Claim 3, Babbar in view of Shubhanshu, further in view of Gelfenbeyn teach the method of claim 1. Babbar further teaches, wherein the obtaining the one or more first embeddings corresponding to the one or more images comprises generating, based at least on the image data, the one or more first embeddings corresponding to the one or more images depicting the content associated with the gaming application ( Babbar: Para.[0039], [0042],Fig. 1, media device 106 can generate one or more embeddings based on one or more signals, such as visual signal ( e.g. image data) in one or more frames of the segment of the media content in the media content server 120, which may store content 122 and metadata 124. Content 122 may include gaming applications).
Babbar in view of Shubhanshu, further in view of Gelfenbeyn while teaching the method of claim 3, fail to explicitly teach the claimed, further comprising: receiving second input data representative of one or more inputs; and generating, based at least on the second input data, image data representative of the one or more images depicting the content associated with the
However, Ahafonov does teach the claimed, further comprising: receiving second input data representative of one or more inputs ( Ahafonov: Para.[0158], Fig. 5, the personal AI agent 502 determines that the user is using the AR/VR device 526 and receives input via the AR/VR device 526 including a voice prompt “What can I cook with these ingredients?");
and generating, based at least on the second input data, image data representative of the one or more images depicting the content associated with the ( Ahafonov: Para.[0158], The personal AI agent 502 accesses a camera feed on the AR/VR device 526 and identifies objects within the camera feed, which include potatoes and carrots. Para.[0259], Fig. 11, can be a gaming application ),
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Ahafonov’s teaching of system and method for generating an extended reality (XR) try-on experience , into the system and method, taught by Babbar in view of Shubhanshu, further in view of Gelfenbeyn, because, this would improve the efficiency of using an electronic device by intelligently and automatically generating images that depict real-world objects in a real-world scene in a simple and intuitive manner. (Ahafonov, Para.[0018]).
Regarding Claim 6, Babbar in view of Shubhanshu, further in view of Gelfenbeyn teach the method of claim 1. Babbar in view of Shubhanshu, further in view of Gelfenbeyn fail to explicitly teach the claimed, further comprising: ; generating a second textual input describing a least in the image; and generating a prompt based at least the textual input and the second textual input; wherein the input data represents at least the prompt.
However, Ahafonov does teach the claimed, generating a second textual input describing a least in the image; and generating a prompt based at least the textual input and the second textual input( Ahafonov: Para.[0167], [0168], Fig.6,a prompt is generated based on the intent 622, where the intent is generated based on the information ( could be textual) extracted/obtained from the multimodal memory);
wherein the input data represents at least the prompt ( Ahafonov: Para.[0168], Fig. 6, The user 638 interacts with a device, such as by entering a prompt as user input 602).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Ahafonov’s teaching of system and method for generating an extended reality (XR) try-on experience , into the system and method, taught by Babbar in view of Shubhanshu, further in view of Gelfenbeyn, because, this would improve the efficiency of using an electronic device by intelligently and automatically generating images that depict real-world objects in a real-world scene in a simple and intuitive manner. (Ahafonov, Para.[0018]).
Regarding Claim 7, Babbar in view of Shubhanshu, further in view of Gelfenbeyn teach the method of claim 1. Babbar in view of Shubhanshu, further in view of Gelfenbeyn fail to explicitly teach the claimed, wherein: first embeddings include[[s]] at least one or more textual embeddings corresponding to the image;
However, Ahafonov does teach the claimed, wherein: first embeddings include[[s]] at least one or more textual embeddings corresponding to the image( Ahafonov: Para.[0147], Fig.5, the embeddings used by the personal AI agent can be image embeddings);
and the input data is associated with the textual input and the one or more textual embeddings ( Ahafonov: Para.[0152], Fig.5, the generative machine learning models are trained to receive a prompt as input (which can include any combination of text, images) and which can be derived from multimodal memory 508, where embeddings are stored);
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Ahafonov’s teaching of system and method for generating an extended reality (XR) try-on experience , into the system and method, taught by Babbar in view of Shubhanshu, further in view of Gelfenbeyn, because, this would improve the efficiency of using an electronic device by intelligently and automatically generating images that depict real-world objects in a real-world scene in a simple and intuitive manner. (Ahafonov, Para.[0018]).
Regarding Claim 12, Babbar in view of Shubhanshu, further in view of Gelfenbeyn teach the system of claim 10. Babbar further teaches, wherein the one or more sources of information includes the one or more images, and wherein the one or more first embeddings corresponding to the one or more images are obtained by at least generating the one or more first embeddings corresponding to the one or more images. ( Babbar: Para.[0039], [0042],Fig. 1, media device 106 can generate one or more embeddings based on one or more signals, such as visual signal ( e.g. image data) in one or more frames of the segment of the media content in the media content server 120, which may store content 122 and metadata 124. Content 122 may include gaming applications).
Babbar in view of Shubhanshu, further in view of Gelfenbeyn while teaching the system of claim 12, fail to explicitly teach the claimed, wherein the one or more processors are further to: receive second input data representative of one or more inputs; and generate, based at least on the second input data, image data representative of one or more images associated with [[a]] the current state of the interactive application.
However, Ahafonov does teach the claimed, wherein the one or more processors are further to:
receive second input data representative of one or more inputs( Ahafonov: Para.[0158], Fig. 5, the personal AI agent 502 determines that the user is using the AR/VR device 526 and receives input via the AR/VR device 526 including a voice prompt “What can I cook with these ingredients?");
and generate, based at least on the second input data, image data representative of one or more images associated with [[a]] the current state of the interactive application ( Ahafonov: Para.[0158], The personal AI agent 502 accesses a camera feed on the AR/VR device 526 and identifies objects within the camera feed, which include potatoes and carrots),
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Ahafonov’s teaching of system and method for generating an extended reality (XR) try-on experience , into the system and method, taught by Babbar in view of Shubhanshu, further in view of Gelfenbeyn, because, this would improve the efficiency of using an electronic device by intelligently and automatically generating images that depict real-world objects in a real-world scene in a simple and intuitive manner. (Ahafonov, Para.[0018]).
Regarding Claim 14, Babbar in view of Shubhanshu, further in view of Gelfenbeyn teach the system of claim 10. Babbar in view of Shubhanshu, further in view of Gelfenbeyn fail to explicitly teach the claimed, wherein the one or more processors are further to: retrieve text corresponding to the
However, Ahafonov does teach the claimed, wherein the one or more processors are further to: retrieve text corresponding to the ( Ahafonov: Para.[0167], [0168], Fig.6, the information communicated are extracted from multimodal memory 626, the data can be text or image data, video data, audio data, electronic documents, links to data stored on the Internet or the client system 606);
and generate a prompt based at least the textual input and the text, wherein the input data represents at least the prompt ( Ahafonov: Para.[0167], [0168], Fig.6,a prompt is generated based on the intent 622, where the intent is generated based on the information ( could be textual) extracted/obtained from the multimodal memory. The user 638 interacts with a device, such as by entering a prompt as user input 602).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Ahafonov’s teaching of system and method for generating an extended reality (XR) try-on experience , into the system and method, taught by Babbar in view of Shubhanshu, further in view of Gelfenbeyn, because, this would improve the efficiency of using an electronic device by intelligently and automatically generating images that depict real-world objects in a real-world scene in a simple and intuitive manner. (Ahafonov, Para.[0018]).
Regarding Claim 15, Babbar in view of Shubhanshu, further in view of Gelfenbeyn teach the system of claim 10. Babbar in view of Shubhanshu, further in view of Gelfenbeyn fail to explicitly teach the claimed, wherein: the current state of the interactive application; the one or more processors are further to determine text based at least on the one or more images; and the input data is associated with the textual input and the text.
However, Ahafonov does teach the claimed, wherein: the current state of the interactive application ( Ahafonov: Para.[0147], Fig.5, the embeddings ( source of information) used by the personal AI agent can be image embeddings);
the one or more processors are further to determine text based at least on the one or more images ( Ahafonov: Para.[0147], Fig.5, embeddings of an image could have corresponding text description);
and the input data is associated with the textual input and the text ( Ahafonov: Para.[0152], Fig.5, the generative machine learning models are trained to receive a prompt as input (which can include any combination of text, images) and which can be derived from multimodal memory 508, where embeddings are stored);
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Ahafonov’s teaching of system and method for generating an extended reality (XR) try-on experience , into the system and method, taught by Babbar in view of Shubhanshu, further in view of Gelfenbeyn, because, this would improve the efficiency of using an electronic device by intelligently and automatically generating images that depict real-world objects in a real-world scene in a simple and intuitive manner. (Ahafonov, Para.[0018]).
Regarding Claim 19, Babbar in view of Shubhanshu, further in view of Gelfenbeyn teach the one or more processors of claim 18. Babbar in view of Shubhanshu, further in view of Gelfenbeyn fail to explicitly teach the claimed, wherein the processing circuitry is further to: generate one or more images associated with a context of the interactive application, wherein the one or more first embeddings areobtained based at least on generating the one or more first embeddings using the one or more images.
However, Ahafonov does teach the claimed, wherein the processing circuitry is further to: generate one or more images associated with a context of the interactive application ( Ahafonov: Para.[0065], the artificial intelligence and machine learning system can implement one or more machine learning models that generate artificial images of a person or object wearing an artificially generated fashion item corresponding to a textual description or prompt and later generate a new image in which the artificially generated fashion item is replaced with an object or XR object that resembles (looks like) a real-world fashion item or product),
wherein the one or more first embeddings areobtained based at least on generating the one or more first embeddings using the one or more images ( Ahafonov: Para.[0145], Fig. 5, personal AI agent 502 can use the user database 504 to determine that a user posts a picture captured from a mobile phone with the user's dog and the caption or comments refer to the dog as Jake. The personal AI agent 502 can update the multimodal memory 508 ( as embeddings) for the user to store a link that associates the user with a dog named Jake, such as a link between an entity representing a dog ( image) and another entity representing the name Jake).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Ahafonov’s teaching of system and method for generating an extended reality (XR) try-on experience , into the system and method, taught by Babbar in view of Shubhanshu, further in view of Gelfenbeyn, because, this would improve the efficiency of using an electronic device by intelligently and automatically generating images that depict real-world objects in a real-world scene in a simple and intuitive manner. (Ahafonov, Para.[0018]).
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to NADIRA SULTANA whose telephone number is (571)272-4048. The examiner can normally be reached M-F,7:30 am-5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Paras D. Shah can be reached on (571) 270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/NADIRA SULTANA/Examiner, Art Unit 2653
/Paras D Shah/Supervisory Patent Examiner, Art Unit 2653
06/30/2026