DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Objections
Claims 1, 14, and 18 are objected to because of the following informalities:
“playing the videogame data to a discriminator” recited in claim 1, ln. 6-7, claim 14, ln. 7-8, and claim 18, ln. 8-9 should likely read “playing the videogame
“by the one or more processors” recited in claim 14, ln. 2 should likely read “by .
Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 4-5, 17, and 21 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 4 recites “The method according to claim 1, wherein the videogame data comprises image data of the videogame and action data of the videogame.” It is indefinite as to whether “the videogame data” is referring to the videogame data generated by the human, the videogame data of the autonomous agent, or both.
Claims 17 and 21 are rejected for the same reasoning.
Claim 5 is rejected by virtue of its dependency on claim 4.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1, 4, 6, 9-11, 14, 17-18, and 21 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Chang (U.S. Pub. 2022/0176248 A1).
Regarding claim 1, Chang discloses a computer-implemented method (Figs. 1-3; [0002]; [0006-0007]), comprising:
providing videogame data generated by a human playing a videogame as input data to a training network for training an autonomous agent for playing the videogame ([0043]; [0049-0050], where game data of a real game user (human) is collected and provided as input to train a machine learning model so as to obtain a game AI model that imitates a user game action to make a game decision);
generating videogame data of the autonomous agent playing the videogame ([0049-0050], where the game AI model makes a game decision (generates videogame data));
providing the videogame data of the autonomous agent playing the videogame data to a discriminator of a generative adversarial network (GAN), wherein the discriminator of the GAN is trained to distinguish between human gameplay and non-human gameplay ([0050-0051]; [0109-0116]; [0122-0123]; [0127], where training samples/input data are input into a discriminator model of a generative adversarial network (GAN) such that the discriminator is trained to classify the training samples as model game samples/model game data (non-human gameplay) or user game samples/user game data (human gameplay));
generating a classification, by the discriminator of the GAN, of the videogame data of the autonomous agent playing the videogame as human or non-human (Fig. 3; [0049-0051]; [0109-0116]; [0122-0123]; [0127], where the discriminator classifies the data inputted into the model as user game data (human gameplay) or model game data (non-human gameplay)); and
updating at least one of the training network and the discriminator of the GAN based on the classification generated by the discriminator of the GAN ([0051]; [0097]; [0116-0128], updating the action model and/or model parameters of the discriminator model accordingly).
Regarding claim 4, Chang further discloses wherein the videogame data comprises image data of the videogame and action data of the videogame ([0049-0050]; [0057]; [0067], game scenario (environment information/image data) and game action data).
Regarding claim 6, Chang further discloses wherein the training network comprises at least one of an inverse reinforcement learning network and an imitation learning network (Figs. 2-3; [0043]; [0049-0051]; [0129]; [0140-0142]; [0157], imitation learning network).
Regarding claim 9, Chang further discloses wherein, when the discriminator of the GAN generates a correct classification, the method comprises: performing at least one of (1) adjusting one or more parameters of the training network by providing the training network with a low reward or (2) adjusting one or more parameters of the discriminator of the GAN by providing the discriminator of the GAN with a high reward ([0123-0129]; [0140-0141], where the parameters of the action model and discriminator model are updated in an adversarial game manner, wherein the discriminator model improves its own discrimination capability as much as possible so as to improve accuracy of the discrimination information by updating and optimizing the model parameters).
Regarding claim 10, Chang further discloses wherein, when the discriminator generates an incorrect classification, the method comprises: performing at least one of (1) adjusting one or more parameters of the discriminator of the GAN by providing the discriminator of the GAN with a low reward or (2) adjusting one or more parameters of the training network by providing the training network with a high reward ([0123-0129]; [0140-0141], where the parameters of the action model and discriminator model are updated in an adversarial game manner, wherein the action model improves its own imitation capability as much as possible and updates and optimizes the model parameters to output a model game sample whose probability distribution is close to that of a user game sample, thus making it more difficult for the discriminator model to accurately distinguish the user game sample/data from the model game sample/data).
Regarding claim 11, Chang further discloses wherein the low reward comprises a penalty ([0123-0129]; [0140-0141], where the parameters of the action model and discriminator model are updated in an adversarial game manner, wherein the discriminator model improves its own discrimination capability as much as possible so as to improve accuracy of the discrimination information by updating and optimizing the model parameters, while the action model’s objective is to deceive the discriminator model so that it is difficult to discriminate the data).
Regarding claim 14, claim 14 is one or more non-transitory computer readable media that store instructions which, when executed by the one or more processors, cause the one or more processors to perform the operations of claim 1, and is therefore rejected for similar reasoning as claim 1 (see further Chang, [0006]; [0009]; [0022]; [0217-0218]).
Regarding claim 17, claim 17 is one or more non-transitory computer readable media that store instructions which, when executed by the one or more processors, cause the one or more processors to perform the operations of claim 4, and is therefore rejected for similar reasoning as claim 4.
Regarding claim 18, claim 18 is a system for performing the operations of claim 1, and is therefore rejected for similar reasoning as claim 1 (see further Chang, Fig. 1; [0047-0048]; [0217-0218]).
Regarding claim 21, claim 21 is a system for performing the operations of claim 4, and is therefore rejected for similar reasoning as claim 4.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 2-3, 7-8, 15-16, and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Chang in view of Perry et al. (U.S. Pub. 2022/0309364 A1) (hereinafter “Perry”).
Regarding claim 2, Chang further discloses repeating the steps of the method ([0050-0051]; [0096-0097]; [0123-0129]; [0158], continuously collecting user game data of a real user and acquiring model game data of the action model and continuously updating the model parameters of the action model and the discriminator model) so that the model action data is close to an action decision habit of a real game user or meets an action decision expectation of a real game user). However, Chang may not explicitly disclose repeating the steps of the method until the classification generated by the discriminator of the GAN satisfies a predetermined condition. Nevertheless, Perry, directed to the creation of human-like non-player character behavior ([0015]), teaches repeating training until the accuracy of the trained network reaches a satisfactory level ([0034]). It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to repeat the training of at least the discriminator of the GAN until the classification accuracy of the discriminator reaches a satisfactory level (satisfies a predetermined condition), as taught by Perry, in order to ensure that the model action data meets an action decision expectation of a real game user, wherein improvement of the discriminator improves the training network.
Regarding claim 3, Chang may not further explicitly disclose wherein the predetermined condition is that the discriminator of the GAN generates a classification with an accuracy of at least a predetermined threshold. However, Perry teaches repeating training until the accuracy of the trained network reaches a satisfactory level (predetermined threshold) ([0034]). It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to repeat the training, as taught by Perry, of the discriminator of the GAN in Chang in order to ensure that the discriminator generates a classification with a satisfactory level of accuracy (Perry, [0034]; Chang, [0123]; [0158]).
Regarding claim 7, Chang further discloses wherein ML and deep learning generally include technologies including reinforcement learning ([0042]; [0051]; [0146-0154], further noting that a learning algorithm of the action model may be optimized by using a policy gradient algorithm in deep reinforcement learning). However, Chang does not further disclose wherein, when the training network comprises an inverse reinforcement learning network, the method comprises: iteratively training the autonomous agent using the inverse reinforcement learning network on the videogame data generated by the human playing the videogame; and generating the videogame data of the autonomous agent playing the videogame after each training iteration. Nevertheless, Perry teaches these limitations ([0015-0018]; [0023]; [0031]; [0033]; [0064-0067]; [0071]; [0075]; [0078-0079]). It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to use inverse reinforcement learning, as taught by Perry, in the invention of Chang as a substitute method to perform the training of the autonomous agent and corresponding generation of the videogame data of the autonomous agent.
Regarding claim 8, Chang further discloses wherein, when the training network comprises an imitation learning network, the method comprises: training the autonomous agent using the imitation learning network on the videogame data generated by the human playing the videogame; and generating the videogame data of the autonomous agent playing the videogame (Figs. 2-3; [0043]; [0049-0051]; [0129]; [0140-0142]; [0157]). However, Chang may not explicitly disclose training for a predetermined number of epochs. Nevertheless, Perry teaches training an autonomous agent using an imitation learning network for a predetermined number of epochs ([0023]; [0027]; [0031]; [0034]; [0071], wherein training epochs are repeated until accuracy reaches a satisfactory level (where the predetermined number is a number that is reached when the satisfactory level is achieved)). It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to perform the training of the autonomous agent in Chang for a predetermined number of epochs, as taught by Perry, in order to perform the training of the autonomous agent and corresponding generation of the videogame data of the autonomous agent such the autonomous agent is trained to make effective decisions that conform to human action logic with sufficient accuracy (Perry, [0015]; [0034]; Chang, [0043]; [0158]).
Regarding claim 15, claim 15 is one or more non-transitory computer readable media that store instructions which, when executed by the one or more processors, cause the one or more processors to perform the operations of claim 2, and is therefore rejected for similar reasoning as claim 2.
Regarding claim 16, claim 16 is one or more non-transitory computer readable media that store instructions which, when executed by the one or more processors, cause the one or more processors to perform the operations of claim 3, and is therefore rejected for similar reasoning as claim 3.
Regarding claim 19, claim 19 is a system for performing the operations of claim 2, and is therefore rejected for similar reasoning as claim 2.
Regarding claim 20, claim 20 is a system for performing the operations of claim 3, and is therefore rejected for similar reasoning as claim 3.
Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Chang in view of Somers et al. (U.S. Pub. 2021/0001229 A1) (hereinafter “Somers”).
Regarding claim 5, Chang does not further disclose wherein the image data of the videogame comprises at least one of first-person video and third person video of the videogame. However, Somers, directed to training a machine learning model to control an in-game character or other entity in a video game in a manner that aims to imitate how a particular player would control the character or entity (Abstract), teaches where the videogame may be played from a first person or third person point of view ([0031]). It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention for the videogame of Chang to be played in a first or third person point of view such that the videogame data includes image data comprising a first-person video or a third person video, as taught by Somers, in order to achieve the claimed invention and/or to allow for an autonomous agent to be trained for different types or modes of video games.
Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Chang in view of Beltran et al. (CN113365706A) (hereinafter “Beltran”).
Regarding claim 12, Chang does not further disclose wherein the videogame data generated by the human playing the videogame comprises data from a plurality of human players of the videogame. Nevertheless, Beltran, directed to training an AI model associated with a game play process of a game application (e.g., training the player, providing opponents to the player, etc.) ([0006]), teaches training an AI model using training state data collected during a plurality of game play processes of the game application, wherein a plurality of game play processes are controlled by a plurality of players via a plurality of client devices ([0103]). It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to utilize videogame data generated from a plurality of human players of the videogame as input data to the training network, as taught by Beltran, in order to provide a wider range of data to generate a more accurate and/or efficient training network.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
U.S. Pub. 2024/0350909 A1 – This reference teaches where training data of the gameplay of a number of players is used to train a model to recognize how players respond in a range of game contexts.
U.S. Pub. 2020/0387739 A1 – This reference teaches where a generative adversarial network comprises two neural networks: a generative network which learns to output data with target features, and a discriminative network which learns to distinguish candidates produced by the generative network from true target data based on such features.
U.S. Pub. 2014/0292803 A1 – This reference teaches where first-person shooter games may include a representation of an object or a part of the player’s character, while a third-person point of view allows a person to view a representation of the player’s character from a third-person perspective.
DE102022204369A1 – This reference teaches where GANs offer a more cost-effective solution for creating your own training, validating, and testing artificial neural networks.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALYSSA N BIANCAMANO whose telephone number is (571)272-4280. The examiner can normally be reached M-F: 8:30am-5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Dmitry Suhol, can be reached at (571)272-4430. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ALYSSA N BIANCAMANO/Examiner, Art Unit 3715