DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 1-34 are presented for examination.
Claim Rejections - 35 USC § 103
Claims 1, 5-7, 12-13, 17-19, 24-25, and 30 are rejected under 35 U.S.C. 103 as being unpatentable over Fadel Argerich et al. (US 20210150417) (“Fadel Argerich”) in view of Osman et al. (US 20210394073) (“Osman”) and further in view of Akhoundi et al. (US 20210312689) (“Akhoundi”).
Regarding claim 1, Fadel Argerich discloses “[o]ne or more processors (Fadel Argerich Fig. 6, processor(s) 602 [circuits]), comprising:
circuitry (see mapping of limitation above) to:
identify, in one or more gameplay images of a video game, one or more objects (in a scenario where a reinforcement learning system has as its environment the ATARI game Breakout, the tutor takes a frame [gameplay image] from the video game as an input and outputs a suggested direction that the bar should be moved – Fadel Argerich, paragraph 163; one of the actions in Breakout is to throw the ball [object] – id. at paragraph 172);
determine, using one or more neural networks, one or more suggested interactions of one or more allowed interactions between a user-controlled game element and the one or more objects, … [wherein the] allowed interactions [are] associated with each of the one or more objects (in a scenario where a reinforcement learning system has as its environment the ATARI game Breakout, the tutor takes a frame from the video game as an input and outputs a suggested direction that the bar [user-controlled game element; user = agent] should be moved – Fadel Argerich, paragraph 163; four actions are available in this environment: (1) no operation; (2) fire (starts the game by “throwing the ball”) [ball = object]; (3) right, and (4) left [these four collectively comprising the allowed interactions associated with the ball] – id. at paragraph 172; see also paragraph 102 (disclosing that the agent learns from its experience using a neural network)) …;
generate, using the one or more neural networks, one or more video frames depicting the user-controlled game element performing the one or more suggested interactions with the one or more objects and further depicting a resulting behavior of the one or more objects responsive to the interaction (reinforcement learning system may have as its environment the video game ATARI Breakout; the tutor takes a frame from the video game as an input and outputs a suggested direction [interaction] that the bar should be moved; thus, for every timestep, the tutor interacts with the agent [user] and gives advice to the agent for making better decisions – Fadel Argerich, paragraph 163 [note that, when the agent takes the tutor’s advice, the result is a sequence of video frames depicting the suggested interaction being performed]; four actions are available in this environment: (1) no operation; (2) fire (starts the game by “throwing the ball”) [ball = object]; (3) right, and (4) left; each action is repeatedly performed for a duration of k = 4 frames [i.e., the action of moving the bar being performed, and the resulting behavior of the ball responsive to moving the bar, are displayed on-screen] – id. at paragraph 172; see also paragraph 102 (discussing the use of neural networks in the system)) …; and
cause a presentation, via a display device, of at least one video frame of the one or more video frames (processing system can include one or more user interfaces, which may include a display – Fadel Argerich, paragraphs 177 and 184) ….”
Fadel Argerich appears not to disclose explicitly the further limitations of the claim. However, Osman discloses “caus[ing] a presentation … of at least one video frame of the one or more video frames to the user (pop-up dashboard [video] may be provided [presented] to keep the player [user] informed on various metrics – Osman, paragraph 83; suggested next move tab of the dashboard may identify a specific move or a specific sequence of moves that the player can select to perform – id. at paragraph 86) ….”
Osman and the instant application both relate to machine learning systems for video games and are analogous. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Fadel Argerich to generate videos suggesting the interactions for display to a user, as disclosed by Osman, and an ordinary artisan could reasonably expect to have done so successfully. Doing so would improve the engagement level of the player and spectators watching the player. See Osman, paragraph 1.
Neither Fadel Argerich nor Osman appears to disclose explicitly the further limitations of the claim. However, Akhoundi discloses that “the … interactions are determined by an encoder, the encoder having stored mappings of … interactions associated with each of the one or more objects (in a system for enhanced pose generation based on conditional modeling of inverse kinematics, a variational autoencoder that generates a latent feature space based on distributions of latent variables may be used – Akhoundi, paragraphs 58-60; the encoder of the autoencoder may learn to map input features of poses to the latent feature space [note that these mappings would need to be stored somewhere to be used in downstream operations], and a decoder may map the latent feature space to an output defining features of poses; thus, the autoencoder may be trained to generate an output pose that reproduces an input pose – id. at paragraph 40 [object = character being posed; interaction = generation of a pose, which is determined partly by the encoder]; see also paragraph 71 (disclosing that a JSON data structure may be used to store the locations for each generated pose)) ….”
Akhoundi and the instant application both relate to the use of neural networks on visual input and are analogous. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination of Fadel Argerich and Osman to employ an encoder to encode mappings related to the objects, as disclosed by Akhoundi, and an ordinary artisan could reasonably expect to have done so successfully. Doing so would save memory space by allowing the mappings to be encoded in a lower-dimensional space. See Akhoundi, paragraph 12.
Claims 7, 13, and 19 are system, method, and non-transitory computer-readable medium claims, respectively, corresponding to processor claim 1 and are rejected for the same reasons as given in the rejection of that claim. Similarly, claim 25 is a player training system claim corresponding to processor claim 1 and is rejected for the same reasons as given in the rejection of that claim, except insofar as claim 25 also contains the following limitation, taught by Fadel Argerich: “memory to store network parameters for the one or more neural networks (processors can perform operations embodying a function, method, or operation by executing code stored on memory – Fadel Argerich, paragraph 179; see also paragraph 164 (disclosing that a threshold parameter may be defined for the agent to control when it will take suggested actions from the tutor instead of using its own decision; this parameter must be stored in memory))”.
Regarding claim 5, Fadel Argerich, as modified by Osman/Akhoundi, discloses that “the one or more video frames depicting the suggested interactions are presented in response to an identification of the one or more objects (suggestions are provided to the players to assist the players in improving the engagement level of the spectators in the video game; the suggestions to improve the engagement level of the spectators may include requests to the players to perform certain types of actions or a certain sequence of actions in the game play, wherein the actions may be identified based on the preference of the spectators [spectator preference = object] – Osman, paragraph 5; see also paragraph 83 (disclosing that the suggestions are presented on a dashboard [video])).” It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Fadel Argerich/Akhoundi to depict the suggested interactions in response to identification of objects, as disclosed by Osman, and an ordinary artisan could reasonably expect to have done so successfully. Doing so would improve the engagement level of the player and spectators watching the player. See Osman, paragraph 1.
Regarding claim 6, Fadel Argerich/Osman/Akhoundi discloses “the one or more video frames comprise one or more frames of video content (tutor takes a frame from the video game as an input and outputs a suggested direction that the bar should be moved [i.e., the system suggests that a frame should be modified so that the bar is in the suggested position] – Fadel Argerich, paragraph 163), and … the video content includes one or more segments representing the suggested interactions, the one or more segments further representing resulting behaviors for the suggested interactions (in the Breakout environment, the observation is an RGB image of the screen, which is an array, and four actions are available: no operation; fire (“throwing the ball”); right; and left; the guide function takes as input the pre-processed frame, locates the position of the ball and the bar in the X-axis [ball, bar = segments representing the interaction between the ball and the bar], and returns “fire” if no ball is found or the action to move in the direction of the ball if it is not above the bar [behaviors = fire, left, and right, which are represented/determined by the relative position of the ball and the bar] – Fadel Argerich, paragraphs 172-73).”
Claims 12, 18, 24, and 30 are system, method, non-transitory computer-readable medium, and player training system claims, respectively, corresponding to processor claim 6 and are rejected for the same reasons as given in the rejection of that claim.
Regarding claim 17, Fadel Argerich, as modified by Osman/Akhoundi, discloses that “the one or more videos depicting the suggested interactions is caused to be presented in response to one or more changes to a state of the video game (suggestions are provided to the players to assist the players in improving the engagement level of the spectators in the video game; the suggestions to improve the engagement level of the spectators may include requests to the players to perform certain types of actions or a certain sequence of actions in the game play, wherein the actions may be identified based on the preference of the spectators [announcement of spectator preference = change in state] – Osman, paragraph 5; see also paragraph 83 (disclosing that the suggestions are presented on a dashboard [video])).” It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Fadel Argerich/Akhoundi to depict the suggested interactions in response to changes to a state of the game, as disclosed by Osman, and an ordinary artisan could reasonably expect to have done so successfully. Doing so would improve the engagement level of the player and spectators watching the player. See Osman, paragraph 1.
Claims 2-3, 8-9, 11, 14-15, 20-21, 23, and 26-28 are rejected under 35 U.S.C. 103 as being unpatentable over Fadel Argerich in view of Osman and Akhoundi and further in view of Bae et al. (US 20210352307) (“Bae”).
Regarding claim 2, Fadel Argerich/Osman/Akhoundi appears not to disclose explicitly the further limitations of the claim. However, Bae discloses that “the circuitry is further to perform instance segmentation to identify features for the one or more objects in one or more input images (if an instance segmentation technique is used, each pixel of the image may be further associated with a label of an instance of objects of the same class; for example, for a class of “individuals,” the instance segmentation technique can differentiate and associate each pixel in the class with labels of “person 1,” “person 2,” and so on [object = individual; feature = number assigned to each individual, e.g., 1, 2, etc.] – Bae, paragraph 97; see also paragraph 4 (indicating that the method is performed with a processor [circuit])).”
Bae and the instant application both relate to the use of neural networks in image processing and are analogous. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Fadel Argerich/Osman/Akhoundi to identify features of images with instance segmentation, as disclosed by Bae, and an ordinary artisan could reasonably expect to have done so successfully. Doing so would allow for the determination of more granular features of the image than would be possible with semantic segmentation, thereby enhancing the system’s understanding of the image. See Bae, paragraph 97.
Claims 8, 14, 20, and 26 are system, method, non-transitory computer-readable medium, and player training system claims, respectively, corresponding to processor claim 2 and are rejected for the same reasons as given in the rejection of that claim.
Regarding claim 3, neither Fadel Argerich, Osman, nor Bae appears to disclose explicitly the further limitations of the claim. However, Akhoundi discloses that “the one or more neural networks include a variational autoencoder (VAE) to encode the features of the one or more objects into a latent space, the VAE further maintaining one or more mappings between the suggested interactions and the one or more objects (in a system for enhanced pose generation based on conditional modeling of inverse kinematics, a variational autoencoder that generates a latent feature space based on distributions of latent variables [features] may be used – Akhoundi, paragraphs 58-60; the encoder of the autoencoder may learn to map input features of poses to the latent feature space, and a decoder may map the latent feature space to an output defining features of poses; thus, the autoencoder may be trained to generate an output pose that reproduces an input pose – id. at paragraph 40 [object = character being posed; interaction = generation of a pose (which entails the system suggesting that the pose should be generated), so the mapping of the pose features to a latent feature space and vice versa is a mapping between the character/object and the generation of its poses]).” It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination of Fadel Argerich, Osman, and Bae to employ a variational autoencoder to encode features of objects into a latent space, as disclosed by Akhoundi, and an ordinary artisan could reasonably expect to have done so successfully. Doing so would allow the system to generate new data once trained, thereby allowing the system to adapt to previously unknown situations. See Akhoundi, paragraph 60.
Claims 9, 15, 21, and 27 are system, method, non-transitory computer-readable medium, and player training system claims, respectively, corresponding to processor claim 3 and are rejected for the same reasons as given in the rejection of that claim.
Regarding claim 11, Fadel Argerich, as modified by Osman, Bae, and Akhoundi, discloses that “the VAE is trained using unsupervised learning to determine the suggested interactions for one or more potential states (autoencoder is an unsupervised machine learning technique capable of learning efficient representations of input data – Akhoundi, paragraph 56; autoencoder may learn to map input features of poses to a latent feature space; a decoder may learn to map the latent feature space to an output defining features of poses; thus; the autoencoder may be trained to generate an output pose that reproduces an input pose [state = pose; interaction = generation of the pose] – id. at paragraph 40).” It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination of Fadel Argerich, Osman, and Bae to train the VAE using unsupervised learning to determine an interaction, as disclosed by Akhoundi, and an ordinary artisan could reasonably expect to have done so successfully. Doing so would allow the system to generate new data once trained, thereby allowing the system to adapt to previously unknown situations. See Akhoundi, paragraph 60.
Claims 23 and 28 are non-transitory computer-readable medium and player training system claims, respectively, corresponding to processor claim 11 and are rejected for the same reasons as given in the rejection of that claim.
Claims 4, 10, 16, 22, and 29 are rejected under 35 U.S.C. 103 as being unpatentable over Fadel Argerich in view of Osman and Bae and further in view of Akhoundi and Robinson et al. (US 20220227379) (“Robinson”).
Regarding claim 4, the rejection of claim 3 is incorporated. Osman further discloses “one or more video frames depicting [the] suggested interactions”, as shown above in the rejection of claim 1. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Fadel Argerich to show videos depicting the suggested interactions, as disclosed by Osman, for substantially the same reasons as given in the rejection of claim 1.
Neither Fadel Argerich, Bae, Osman, nor Akhoundi appears to disclose explicitly the further limitations of the claim. However, Robinson discloses that “the one or more networks include a generative network for generating the … suggested interactions, the generative network accepting as input at least the latent space and the mappings (in a system for the detection of edge cases through application of a neural network to predict future vehicle environment data, sensor data are encoded by mapping the sensor data onto a corresponding latent space; encoded data are sent to a fusion node in a level above the node, and the fusion node combines encoded sensor data from multiple nodes; the encoded combination is decoded into the latent space of the node [i.e., the mapped latent space is input to the decoder] – Robinson, paragraphs 155-56; neural network processes the environment data to determine predicted environment data for a second time and, in response to determining that predicted environment data indicate an environmental state associated with danger, an alert is issued to a vehicle control system to recommend [suggest] taking remedial action [interaction] to adapt to the predicted environmental state – id. at paragraph 27; the system may form a generative adversarial network – id. at paragraph 102).”
Robinson and the instant application both relate to generative networks used to determine suggested actions and are analogous. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination of Fadel Argerich, Osman, Bae, and Akhoundi to use a generative network that accepts a latent representation of the data as input to generate the recommendations, as disclosed by Robinson, and an ordinary artisan could reasonably expect to have done so successfully. Doing so would allow the system more accurately to produce output by training the generator to produce better outputs using the feedback of the discriminator. See Robinson, paragraph 102.
Claims 10, 16, and 22 are system, method, and non-transitory computer-readable medium claims, respectively, corresponding to processor claim 4 and are rejected for the same reasons as given in the rejection of that claim.
Regarding claim 29, Fadel Argerich, as modified by Bae, Osman, Akhoundi, and Robinson, discloses that “the one or more neural networks include a generative adversarial network (GAN) to accept the latent space as input and generate the one or more recommendations based at least in part upon the one or more cumulative changes of state determined from the latent space (in a system for the detection of edge cases through application of a neural network to predict future vehicle environment data, sensor data are encoded by mapping the sensor data onto a corresponding latent space; encoded data are sent to a fusion node in a level above the node, and the fusion node combines encoded sensor data from multiple nodes; the encoded combination is decoded into the latent space of the node [i.e., the latent space is input to the decoder] – Robinson, paragraphs 155-56; neural network processes the environment data to determine predicted environment data for a second time and, in response to determining that predicted environment data indicate an environmental state associated with danger, an alert is issued to a vehicle control system to recommend taking remedial action to adapt to the predicted environmental state [change of state] – id. at paragraph 27; the system may form a generative adversarial network – id. at paragraph 102).” It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination of Fadel Argerich, Bae, Osman, and Akhoundi to use a GAN accepting a latent space as input to provide recommendations based on state changes, as disclosed by Robinson, and an ordinary artisan could reasonably expect to have done so successfully. Doing so would allow the system more accurately to produce output by training the generator to produce better outputs using the feedback of the discriminator. See Robinson, paragraph 102.
Claims 31-34 are rejected under 35 U.S.C. 103 as being unpatentable over Fadel Argerich in view of Osman and Akhoundi and further in view of Bennett (US 20200134447) (“Bennett”).
Regarding claim 31, Fadel Argerich, as modified by Osman, Akhoundi, and Bennett, discloses that “the user is one or more human users (in a system for providing synchronized input feedback in a video game, an input event [from a human] that precedes an action of an avatar in the video stream is placed at a time in the audio stream of the video game before the actions of the avatar occur; encoded input embedded in the output stream during reproduction is undetectable to a user who is a human being with average vision and hearing faculties – Bennett, paragraph 23).”
Bennett and the instant application both relate to the use of machine learning in video games and are analogous. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Fadel Argerich/Osman/Akhoundi such that the system receives input and interacts with a human user, as disclosed by Bennett, and an ordinary artisan could reasonably expect to have done so successfully. Doing so would allow the system to tailor its output such that it is integrated seamlessly with the user experience. See Bennett, paragraph 23.
Claims 32-34 are system, method, and non-transitory computer-readable medium claims, respectively, corresponding to processor claim 31 and are rejected for the same reasons as given in the rejection of that claim.
Response to Arguments
Applicant's arguments filed July 13, 2026 (“Remarks”) have been fully considered but they are, except insofar as rendered moot by the introduction of a new ground of rejection, not persuasive.
Applicant’s argument that the Fadel Argerich/Osman combination does not teach that the interactions are determined by an encoder storing mappings of interactions associated with objects, Remarks at 11-12, is moot by virtue of the use of Akhoundi to teach this limitation. Applicant further argues that Fadel Argerich allegedly fails to disclose displaying video frames of game elements performing suggested interactions with objects and further depicting a resulting behavior of the objects responsive to the interaction being performed, Remarks at 12-13, paragraph 172 of Fadel Argerich discloses that each action suggested by the tutor is performed for four frames. To the extent that the action performed is throwing the ball, the display (see paragraph 184 of Fadel Argerich) must show, over the course of those four frames, both (a) the ball being thrown (i.e., the performance of the suggested interaction), and (b) the movement of the ball over those four frames as a result of being thrown (i.e., the behavior of the object responsive to the action). Therefore, Fadel Argerich discloses this limitation.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to RYAN C VAUGHN whose telephone number is (571)272-4849. The examiner can normally be reached M-R 7:00a-5:00p ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kamran Afshar, can be reached at 571-272-7796. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/RYAN C VAUGHN/Primary Examiner, Art Unit 2125