Prosecution Insights
Last updated: October 02, 2026
Application No. 19/008,339

METHOD, COMPUTER PROGRAM AND APPARATUS FOR TRAINING AN AUTONOMOUS AGENT

Non-Final OA §102§103§112
Filed
Jan 02, 2025
Priority
Jan 18, 2024 — GB 2400678.5
Examiner
BIANCAMANO, ALYSSA N
Art Unit
Tech Center
Assignee
Sony Group Corporation
OA Round
1 (Non-Final)
56%
Grant Probability
Moderate
1-2
OA Rounds
1y 5m
Est. Remaining
92%
With Interview

Examiner Intelligence

Grants 56% of resolved cases
56%
Career Allowance Rate
100 granted / 179 resolved
-4.1% vs TC avg
Strong +36% interview lift
Without
With
+36.4%
Interview Lift
resolved cases with interview
Typical timeline
3y 2m
Avg Prosecution
41 currently pending
Career history
222
Total Applications
across all art units

Statute-Specific Performance

§101
17.1%
-22.9% vs TC avg
§103
33.9%
-6.1% vs TC avg
§102
14.3%
-25.7% vs TC avg
§112
32.0%
-8.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 179 resolved cases

Office Action

§102 §103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Objections Claims 1, 14, and 18 are objected to because of the following informalities: “playing the videogame data to a discriminator” recited in claim 1, ln. 6-7, claim 14, ln. 7-8, and claim 18, ln. 8-9 should likely read “playing the videogame “by the one or more processors” recited in claim 14, ln. 2 should likely read “by . Appropriate correction is required. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 4-5, 17, and 21 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 4 recites “The method according to claim 1, wherein the videogame data comprises image data of the videogame and action data of the videogame.” It is indefinite as to whether “the videogame data” is referring to the videogame data generated by the human, the videogame data of the autonomous agent, or both. Claims 17 and 21 are rejected for the same reasoning. Claim 5 is rejected by virtue of its dependency on claim 4. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1, 4, 6, 9-11, 14, 17-18, and 21 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Chang (U.S. Pub. 2022/0176248 A1). Regarding claim 1, Chang discloses a computer-implemented method (Figs. 1-3; [0002]; [0006-0007]), comprising: providing videogame data generated by a human playing a videogame as input data to a training network for training an autonomous agent for playing the videogame ([0043]; [0049-0050], where game data of a real game user (human) is collected and provided as input to train a machine learning model so as to obtain a game AI model that imitates a user game action to make a game decision); generating videogame data of the autonomous agent playing the videogame ([0049-0050], where the game AI model makes a game decision (generates videogame data)); providing the videogame data of the autonomous agent playing the videogame data to a discriminator of a generative adversarial network (GAN), wherein the discriminator of the GAN is trained to distinguish between human gameplay and non-human gameplay ([0050-0051]; [0109-0116]; [0122-0123]; [0127], where training samples/input data are input into a discriminator model of a generative adversarial network (GAN) such that the discriminator is trained to classify the training samples as model game samples/model game data (non-human gameplay) or user game samples/user game data (human gameplay)); generating a classification, by the discriminator of the GAN, of the videogame data of the autonomous agent playing the videogame as human or non-human (Fig. 3; [0049-0051]; [0109-0116]; [0122-0123]; [0127], where the discriminator classifies the data inputted into the model as user game data (human gameplay) or model game data (non-human gameplay)); and updating at least one of the training network and the discriminator of the GAN based on the classification generated by the discriminator of the GAN ([0051]; [0097]; [0116-0128], updating the action model and/or model parameters of the discriminator model accordingly). Regarding claim 4, Chang further discloses wherein the videogame data comprises image data of the videogame and action data of the videogame ([0049-0050]; [0057]; [0067], game scenario (environment information/image data) and game action data). Regarding claim 6, Chang further discloses wherein the training network comprises at least one of an inverse reinforcement learning network and an imitation learning network (Figs. 2-3; [0043]; [0049-0051]; [0129]; [0140-0142]; [0157], imitation learning network). Regarding claim 9, Chang further discloses wherein, when the discriminator of the GAN generates a correct classification, the method comprises: performing at least one of (1) adjusting one or more parameters of the training network by providing the training network with a low reward or (2) adjusting one or more parameters of the discriminator of the GAN by providing the discriminator of the GAN with a high reward ([0123-0129]; [0140-0141], where the parameters of the action model and discriminator model are updated in an adversarial game manner, wherein the discriminator model improves its own discrimination capability as much as possible so as to improve accuracy of the discrimination information by updating and optimizing the model parameters). Regarding claim 10, Chang further discloses wherein, when the discriminator generates an incorrect classification, the method comprises: performing at least one of (1) adjusting one or more parameters of the discriminator of the GAN by providing the discriminator of the GAN with a low reward or (2) adjusting one or more parameters of the training network by providing the training network with a high reward ([0123-0129]; [0140-0141], where the parameters of the action model and discriminator model are updated in an adversarial game manner, wherein the action model improves its own imitation capability as much as possible and updates and optimizes the model parameters to output a model game sample whose probability distribution is close to that of a user game sample, thus making it more difficult for the discriminator model to accurately distinguish the user game sample/data from the model game sample/data). Regarding claim 11, Chang further discloses wherein the low reward comprises a penalty ([0123-0129]; [0140-0141], where the parameters of the action model and discriminator model are updated in an adversarial game manner, wherein the discriminator model improves its own discrimination capability as much as possible so as to improve accuracy of the discrimination information by updating and optimizing the model parameters, while the action model’s objective is to deceive the discriminator model so that it is difficult to discriminate the data). Regarding claim 14, claim 14 is one or more non-transitory computer readable media that store instructions which, when executed by the one or more processors, cause the one or more processors to perform the operations of claim 1, and is therefore rejected for similar reasoning as claim 1 (see further Chang, [0006]; [0009]; [0022]; [0217-0218]). Regarding claim 17, claim 17 is one or more non-transitory computer readable media that store instructions which, when executed by the one or more processors, cause the one or more processors to perform the operations of claim 4, and is therefore rejected for similar reasoning as claim 4. Regarding claim 18, claim 18 is a system for performing the operations of claim 1, and is therefore rejected for similar reasoning as claim 1 (see further Chang, Fig. 1; [0047-0048]; [0217-0218]). Regarding claim 21, claim 21 is a system for performing the operations of claim 4, and is therefore rejected for similar reasoning as claim 4. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 2-3, 7-8, 15-16, and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Chang in view of Perry et al. (U.S. Pub. 2022/0309364 A1) (hereinafter “Perry”). Regarding claim 2, Chang further discloses repeating the steps of the method ([0050-0051]; [0096-0097]; [0123-0129]; [0158], continuously collecting user game data of a real user and acquiring model game data of the action model and continuously updating the model parameters of the action model and the discriminator model) so that the model action data is close to an action decision habit of a real game user or meets an action decision expectation of a real game user). However, Chang may not explicitly disclose repeating the steps of the method until the classification generated by the discriminator of the GAN satisfies a predetermined condition. Nevertheless, Perry, directed to the creation of human-like non-player character behavior ([0015]), teaches repeating training until the accuracy of the trained network reaches a satisfactory level ([0034]). It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to repeat the training of at least the discriminator of the GAN until the classification accuracy of the discriminator reaches a satisfactory level (satisfies a predetermined condition), as taught by Perry, in order to ensure that the model action data meets an action decision expectation of a real game user, wherein improvement of the discriminator improves the training network. Regarding claim 3, Chang may not further explicitly disclose wherein the predetermined condition is that the discriminator of the GAN generates a classification with an accuracy of at least a predetermined threshold. However, Perry teaches repeating training until the accuracy of the trained network reaches a satisfactory level (predetermined threshold) ([0034]). It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to repeat the training, as taught by Perry, of the discriminator of the GAN in Chang in order to ensure that the discriminator generates a classification with a satisfactory level of accuracy (Perry, [0034]; Chang, [0123]; [0158]). Regarding claim 7, Chang further discloses wherein ML and deep learning generally include technologies including reinforcement learning ([0042]; [0051]; [0146-0154], further noting that a learning algorithm of the action model may be optimized by using a policy gradient algorithm in deep reinforcement learning). However, Chang does not further disclose wherein, when the training network comprises an inverse reinforcement learning network, the method comprises: iteratively training the autonomous agent using the inverse reinforcement learning network on the videogame data generated by the human playing the videogame; and generating the videogame data of the autonomous agent playing the videogame after each training iteration. Nevertheless, Perry teaches these limitations ([0015-0018]; [0023]; [0031]; [0033]; [0064-0067]; [0071]; [0075]; [0078-0079]). It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to use inverse reinforcement learning, as taught by Perry, in the invention of Chang as a substitute method to perform the training of the autonomous agent and corresponding generation of the videogame data of the autonomous agent. Regarding claim 8, Chang further discloses wherein, when the training network comprises an imitation learning network, the method comprises: training the autonomous agent using the imitation learning network on the videogame data generated by the human playing the videogame; and generating the videogame data of the autonomous agent playing the videogame (Figs. 2-3; [0043]; [0049-0051]; [0129]; [0140-0142]; [0157]). However, Chang may not explicitly disclose training for a predetermined number of epochs. Nevertheless, Perry teaches training an autonomous agent using an imitation learning network for a predetermined number of epochs ([0023]; [0027]; [0031]; [0034]; [0071], wherein training epochs are repeated until accuracy reaches a satisfactory level (where the predetermined number is a number that is reached when the satisfactory level is achieved)). It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to perform the training of the autonomous agent in Chang for a predetermined number of epochs, as taught by Perry, in order to perform the training of the autonomous agent and corresponding generation of the videogame data of the autonomous agent such the autonomous agent is trained to make effective decisions that conform to human action logic with sufficient accuracy (Perry, [0015]; [0034]; Chang, [0043]; [0158]). Regarding claim 15, claim 15 is one or more non-transitory computer readable media that store instructions which, when executed by the one or more processors, cause the one or more processors to perform the operations of claim 2, and is therefore rejected for similar reasoning as claim 2. Regarding claim 16, claim 16 is one or more non-transitory computer readable media that store instructions which, when executed by the one or more processors, cause the one or more processors to perform the operations of claim 3, and is therefore rejected for similar reasoning as claim 3. Regarding claim 19, claim 19 is a system for performing the operations of claim 2, and is therefore rejected for similar reasoning as claim 2. Regarding claim 20, claim 20 is a system for performing the operations of claim 3, and is therefore rejected for similar reasoning as claim 3. Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Chang in view of Somers et al. (U.S. Pub. 2021/0001229 A1) (hereinafter “Somers”). Regarding claim 5, Chang does not further disclose wherein the image data of the videogame comprises at least one of first-person video and third person video of the videogame. However, Somers, directed to training a machine learning model to control an in-game character or other entity in a video game in a manner that aims to imitate how a particular player would control the character or entity (Abstract), teaches where the videogame may be played from a first person or third person point of view ([0031]). It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention for the videogame of Chang to be played in a first or third person point of view such that the videogame data includes image data comprising a first-person video or a third person video, as taught by Somers, in order to achieve the claimed invention and/or to allow for an autonomous agent to be trained for different types or modes of video games. Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Chang in view of Beltran et al. (CN113365706A) (hereinafter “Beltran”). Regarding claim 12, Chang does not further disclose wherein the videogame data generated by the human playing the videogame comprises data from a plurality of human players of the videogame. Nevertheless, Beltran, directed to training an AI model associated with a game play process of a game application (e.g., training the player, providing opponents to the player, etc.) ([0006]), teaches training an AI model using training state data collected during a plurality of game play processes of the game application, wherein a plurality of game play processes are controlled by a plurality of players via a plurality of client devices ([0103]). It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to utilize videogame data generated from a plurality of human players of the videogame as input data to the training network, as taught by Beltran, in order to provide a wider range of data to generate a more accurate and/or efficient training network. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. U.S. Pub. 2024/0350909 A1 – This reference teaches where training data of the gameplay of a number of players is used to train a model to recognize how players respond in a range of game contexts. U.S. Pub. 2020/0387739 A1 – This reference teaches where a generative adversarial network comprises two neural networks: a generative network which learns to output data with target features, and a discriminative network which learns to distinguish candidates produced by the generative network from true target data based on such features. U.S. Pub. 2014/0292803 A1 – This reference teaches where first-person shooter games may include a representation of an object or a part of the player’s character, while a third-person point of view allows a person to view a representation of the player’s character from a third-person perspective. DE102022204369A1 – This reference teaches where GANs offer a more cost-effective solution for creating your own training, validating, and testing artificial neural networks. Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALYSSA N BIANCAMANO whose telephone number is (571)272-4280. The examiner can normally be reached M-F: 8:30am-5:00pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Dmitry Suhol, can be reached at (571)272-4430. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ALYSSA N BIANCAMANO/Examiner, Art Unit 3715
Read full office action

Prosecution Timeline

Jan 02, 2025
Application Filed
Mar 26, 2026
Response after Non-Final Action
Aug 21, 2026
Non-Final Rejection mailed — §102, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12708290
ASYMMETRICAL RHYTHMIC AUDITORY CUEING BASED GAIT MODIFICATION
4y 9m to grant Granted Aug 18, 2026
Patent 12705998
Systems and Methods for an Educational Generative Artificial Intelligence Model
2y 7m to grant Granted Aug 11, 2026
Patent 12697674
WELD TRACKING SYSTEMS
4y 6m to grant Granted Aug 04, 2026
Patent 12678169
PRESSURE LIMITED TRAINING TOURNIQUET
2y 9m to grant Granted Jul 14, 2026
Patent 12682781
ROBOT SYSTEM
2y 1m to grant Granted Jul 14, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
56%
Grant Probability
92%
With Interview (+36.4%)
3y 2m (~1y 5m remaining)
Median Time to Grant
Low
PTA Risk
Based on 179 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month