Prosecution Insights
Last updated: October 01, 2026
Application No. 18/296,795

EVENT TABLES FOR EFFICIENT EXPERIENCE REPLAY

Final Rejection §102§103§112
Filed
Apr 06, 2023
Priority
May 13, 2022 — provisional 63/364,665
Examiner
DRAPEAU, SIMEON PAUL
Art Unit
Tech Center
Assignee
Sony Group Corporation
OA Round
2 (Final)
19%
Grant Probability
At Risk
3-4
OA Rounds
9m
Est. Remaining
89%
With Interview

Examiner Intelligence

Grants only 19% of cases
19%
Career Allowance Rate
3 granted / 16 resolved
-41.2% vs TC avg
Strong +70% interview lift
Without
With
+70.3%
Interview Lift
resolved cases with interview
Typical timeline
4y 3m
Avg Prosecution
33 currently pending
Career history
49
Total Applications
across all art units

Statute-Specific Performance

§101
33.3%
-6.7% vs TC avg
§103
31.0%
-9.0% vs TC avg
§102
17.1%
-22.9% vs TC avg
§112
16.0%
-24.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 16 resolved cases

Office Action

§102 §103 §112
DETAILED ACTION The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claims 1-19 are presented for examination based on the amended claims in the application filed on July 29, 2026. Claims 17-18 are rejected under 35 U.S.C. § 112(b) or 35 U.S.C. § 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. § 112, the applicant), regards as the invention. Claims 1-7, 10-12, and 14-19 are rejected under 35 U.S.C. § 102(a)(1) as being anticipated by Sharma, Anil, Mayank K. Pal, Saket Anand, and Sanjit K. Kaul. "Stratified sampling based experience replay for efficient camera selection decisions." In 2020 IEEE Sixth International Conference on Multimedia Big Data (BigMM), pp. 144-151. IEEE, 2020 [herein “Sharma”]. Claim 8 is rejected under 35 U.S.C. § 103 as being unpatentable over Sharma in view of Wurman, Peter R. et al. "Outracing champion Gran Turismo drivers with deep reinforcement learning." Nature 602, no. 7896 (2022): 223-228. Claim 9 is rejected under 35 U.S.C. § 103 as being unpatentable over Sharma in view of Luo, Jieliang, and Hui Li. "Dynamic experience replay." In Conference on robot learning, pp. 1191-1200. PMLR, 2020. Claim 13 is rejected under 35 U.S.C. § 103 as being unpatentable over Sharma in view of Daley, Brett, Cameron Hickert, and Christopher Amato. "Stratified experience replay: Correcting multiplicity bias in off-policy reinforcement learning." arXiv preprint arXiv:2102.11319 (2021). This action is made Final. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment The amendment filed July 29, 2026 has been entered. Claims 1-19 remain pending in the application. Applicant’s amendments to the Specification, Drawings, and Claims have overcome each and every objection and 112(b) rejections previously set forth in the Non-Final Office Action mailed June 30, 2026, with the exception of the objections to the drawings and claims as provided below. Drawings The drawings are objected to because of the following: FIGS. 2A-2E (Para. 0035), FIGS. 6A-6E (Para. 0039), FIGS. 7A-7F (Para. 0040), FIG. 9 (Para. 0042), FIG. 10 (Para. 0043), and FIGS. 11A-11C (Para. 0044) fail to show “shaded region”. FIG. 17A (Para. 00114) fails to show “blue labeled lines”. FIG. 3A, 4A, and 6A recite “Get to the green square”, but the figures are not in color. Any structural detail that is essential for a proper understanding of the disclosed invention should be shown in the drawing (see MPEP § 608.02(d) and 37 CFR 1.83(a)). Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. The figure or figure number of a n amended drawing should not be labeled as “amended.” If a drawing figure is to be canceled, the appropriate figure must be removed from the replacement sheet, and where necessary, the remaining figures must be renumbered and appropriate changes made to the brief description of the several views of the drawings for consistency. Additional replacement sheets may be necessary to show the renumbering of the remaining figures. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin a s either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance. Claim Objections Claim 5 is objected to because of the following informality: Claim 5, “each table” in Ln. 1-2, should be “each event table”. Applicant is advised that should claim 16 be found allowable, claim 19 will be objected to under 37 CFR 1.75 as being a substantial duplicate thereof. When two claims in an application are duplicates or else are so close in content that they both cover the same thing, despite a slight difference in wording, it is proper after allowing one claim to object to the other as being a substantial duplicate of the allowed claim. See MPEP § 608.01(m). Claim Rejections - 35 U.S.C. § 112 The following is a quotation of 35 U.S.C. § 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. § 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 17-18 are rejected under 35 U.S.C. § 112(b) or 35 U.S.C. § 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. § 112, the applicant), regards as the invention. Claim 17 recites “for utilizing the event tables with a reinforcement learning system” in Ln 2-13. This phrase renders the claim indefinite, because it merely recites a use without any steps delimiting the use (See MPEP § 2173.05(q), “Attempts to claim a process without setting forth any steps involved in the process generally raises an issue of indefiniteness under 35 U.S.C. § 112(b) or pre-AIA 35 U.S.C. § 112, second paragraph”). Claim 18, being dependent on claim 17, is also rejected. Claim Rejections - 35 U.S.C. § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. § 102 and 103 (or as subject to pre-AIA 35 U.S.C. § 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. § 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claims 1-7, 10-12, and 14-19 are rejected under 35 U.S.C. § 102(a)(1) as being anticipated by Sharma, Anil, Mayank K. Pal, Saket Anand, and Sanjit K. Kaul. "Stratified sampling based experience replay for efficient camera selection decisions." In 2020 IEEE Sixth International Conference on Multimedia Big Data (BigMM), pp. 144-151. IEEE, 2020 [herein “Sharma”]. As per claim 1, Sharma teaches “A computer-implemented method for improving convergence in training a policy for a reinforcement learning agent”. (Pg. 145 Sect. II, “Experience replay has been incorporated in many deep reinforcement learning methods to memorize and replay past experiences of the agent-environment interaction. It has been observed that ER stabilizes the training process by breaking the temporal correlations of the sequential online transitions [1]. Uniform sampling is the most common approach for sampling the transitions from the replay memory. However, researchers have demonstrated that uniform sampling fails to create a diverse minibatch for many applications [7] and hence the RL algorithms fails in such scenarios” and Pg. 145 Sect. I, “This is an important observation with respect to training a deep RL model, which requires appropriately handling of imbalanced state transitions during experience replay. Our proposed approach, referred to as Stratified Experience Replay (SER) resolves this challenge by sampling in the imbalanced replay memory created by the different episodic runs of the agent-environment interaction. Our specific contributions are the following: 1) We propose a novel experience replay method to segregate transitions into multiple replay memories” [i.e., a method for improving convergence in training a policy for a reinforcement learning agent]. Pg. 148 Sect. V, “We implemented the DQN algorithm using PyTorch framework and utilized a server with 128 GBs of RAM and a 11-GB Nvidia RTX 2080 Ti GPU for training” [A computer-implemented method]. Further see Sect. I-II and IV-V. The examiner has interpreted that using PyTorch framework and a server for training a deep reinforcement learning (RL) model to handle state transitions using a diverse minibatch during experience replay to stabilize the training process as a computer-implemented method for improving convergence in training a policy for a reinforcement learning agent.) Sharma also teaches “partitioning an experience replay buffer for the reinforcement learning agent into event tables based on event conditions and history lengths, wherein each event table contains data where a corresponding event condition was true; and each event table includes a history length that reaches back to a previous occurrence of the corresponding event condition or to an initial state, the event table storing historic data of the reinforcement learning agent for the history length”. (Pg. 145 Sect. I, “This is an important observation with respect to training a deep RL model, which requires appropriately handling of imbalanced state transitions during experience replay. Our proposed approach, referred to as Stratified Experience Replay (SER) resolves this challenge by sampling in the imbalanced replay memory created by the different episodic runs of the agent-environment interaction. Our specific contributions are the following: 1) We propose a novel experience replay method to segregate transitions into multiple replay memories” [partitioning an experience replay buffer for a reinforcement learning agent into event tables]. Pg. 146 Sect. III, “during transitions or occlusions, the state at time t captures three elements to handle this partially observable state of the target: 1) xt: the last observed location of the target…, 2) ht: the action history that maintains a list of previously selected actions by the policy. The history is stored as a list of cameras encoded as a one-hot vector. 3) τ: This is the time-elapsed vector that captures the time since the target’s most recent observation by the agent in any camera” [event conditions and history lengths, and each event table includes a history length that reaches back to a previous occurrence of the corresponding event condition and to an initial state, the event table storing historic data of the reinforcement learning agent for the history length]. Furthermore, Pg. 147 Sect. IV, “Therefore, we have segregated transitions into multiple replay memories to enable sampling of all kinds of transitions to create the minibatch. Also as observed in supervised learning, we need present the rare transitions more often to the network. To ensure efficient sampling of rare and other transitions, we segregate transitions in different replay memories to compensate for searching in significantly large replay memory” [partitioning an experience replay buffer for a reinforcement learning agent into event tables based on event conditions and history lengths]. Pg. 148 Sect. IV, “The multiple replays are; Rf, which stores the transitions which pertain frequently occurring action C×; R−, which stores the transitions that receive a negative reward, and R+ which stores the transitions that receive a positive reward” [e.g., wherein each table contains data where a corresponding event condition was true]. Further see Sect. I-IV. The examiner has interpreted that training a deep reinforcement learning (RL) model to handle state transition during experience replay by segregating transitions into multiple relay memories to create a minibatch to compensate for searching in large replay memory where transitions have states that capture locations of targets, a history of previous selected actions from the policy, a time of last captured location that was captured by the agent, and store transitions which pertain frequently occurring action as partitioning an experience replay buffer for the reinforcement learning agent into event tables based on event conditions and history lengths, wherein each event table contains data where a corresponding event condition was true; and each event table includes a history length that reaches back to a previous occurrence of the corresponding event condition or to an initial state, the event table storing historic data of the reinforcement learning agent for the history length.) Sharma also teaches “sampling individual steps from each event table to build training samples for off-policy reinforcement learning; and performing off-policy training of the reinforcement learning agent”. (Pg. 145 Sect. I, “This is an important observation with respect to training a deep RL model, which requires appropriately handling of imbalanced state transitions during experience replay. Our proposed approach, referred to as Stratified Experience Replay (SER) resolves this challenge by sampling in the imbalanced replay memory created by the different episodic runs of the agent-environment interaction” [reinforcement learning system]. Pg. 148 Sect. IV, “For learning, a minibatch is prepared from the stored transitions in the multiple replays. The transitions are sampled uniformly from Rf, and R− replay. R+ stores both rare and other positive reward transitions, it follows a prioritized sampling. For priority sampling, a higher weight is assigned to the rare transitions and a lower weight to other transitions. A probability value is assigned to ith transition as w i / ∑ w i ” [sampling individual steps from each event table to build training samples for off-policy reinforcement learning] Pg. 148 Sect. IV, “The minibatch is used to learn the policy” [performing training of the reinforcement learning agent]. Pg. 148 Sect. V, “we will be comparing performance of different ER methods applied to off-policy DQN” [off-policy reinforcement learning/off-policy training]. Further see Sect. I-II and IV-V. The examiner has interpreted that sampling from the stored transitions to create a minibatch to learn the policy and train a deep reinforcement model applied off-policy as sampling individual steps from each event table to build training samples for off-policy reinforcement learning; and performing off-policy training of the reinforcement learning agent.) As per claim 2, Sharma teaches “where the experience replay buffer includes a default table that holds all incoming data.” (Pg. 147 Sect. IV, “replay memory of size R is used, which stores the last |R| transitions” [where the experience replay buffer includes a default table that holds all incoming data]. Further see Sect. IV.) As per claim 3, Sharma teaches “where the event conditions are based on histories and not just a single state.” (Pg. 146 Sect. III, “during transitions or occlusions, the state at time t captures three elements to handle this partially observable state of the target: 1) xt: the last observed location of the target…, 2) ht: the action history that maintains a list of previously selected actions by the policy. The history is stored as a list of cameras encoded as a one-hot vector. 3) τ: This is the time-elapsed vector that captures the time since the target’s most recent observation by the agent in any camera” [where the event conditions are based on histories and not just a single state]. Further see Sect. III-V. The examiner has interpreted that transitions previous selected actions in an action history as where the event conditions are based on histories and not just a single state.) As per claim 4, Sharma teaches “where each event table has a specified capacity in the experience replay buffer proportional to an overall size of the experience replay buffer.” (Pg. 147 Sect. IV, “Please note that state of-the-art replay methods use single replay memory with recommended size of 106” [an overall size of the experience replay buffer]. Pg. 148 Sect. V, “For the proposed method, we used three replays to separate the rare transitions, where the size of each replay was set to 103 acquired via hyperparameter tuning” [e.g., where each event table has a specified capacity in the experience replay buffer proportional to an overall size of the experience replay buffer]. Further see Sect. IV-V. The examiner has interpreted that separating a single replay memory with recommended size of 106 into three separate replays to store transitions of size 103 as where each event table has a specified capacity in the experience replay buffer proportional to an overall size of the experience replay buffer.) As per claim 5, Sharma teaches “where each table has a specified sampling probability.” (Pg. 148 Sect. IV, “The multiple replays are; Rf, which stores the transitions which pertain frequently occurring action C×; R−, which stores the transitions that receive a negative reward, and R+ which stores the transitions that receive a positive reward” [e.g., wherein each table contains certain transitions]. Pg. 148 Sect. IV, “The transitions are sampled uniformly from Rf, and R− replay. R+ stores both rare and other positive reward transitions, it follows a prioritized sampling. For priority sampling, a higher weight is assigned to the rare transitions and a lower weight to other transitions. A probability value is assigned to ith transition as w i / ∑ w i ” [transitions are sampled having uniform probability and others have a probability value, e.g. where each table has a specified sampling probability]. Further see Sect. IV. The examiner has interpreted that having multiple replays having a uniform and a priority sampling with a probability have that are sampled into three replays as where each table has a specified sampling probability.) As per claim 6, Sharma teaches “wherein a mapping from the event conditions to the event tables is surjective, with multiple ones of the event conditions funneling data into the same one of the event tables.” (Pg. 148 Sect. IV “The multiple replays are; Rf, which stores the transitions which pertain frequently occurring action C×; R−, which stores the transitions that receive a negative reward, and R+ which stores the transitions that receive a positive reward…This transition is stored in a replay which satisfies the above criteria for segregation” [wherein a mapping from the event conditions to the event tables is surjective]. The examiner would like to refer the applicant to Figure 3 from Sharma, shown below as Figure 1, that depicts that transitions are stored in one of three replay memories. As shown, if rt+1 = rcx, then the transition is stored in Rf. If rt+1 < rcx, then the transition is stored in R_. If rt+1 > rcx, then the transition is stored in R+. Additionally, this is also seen in Algorithm 1, shown below as Figure 2, annotated lines 15-21 shown below that shows that transitions are binned into one of three replay memories based on the reward received. This shows that transitions can get binned into the same replay memories, e.g., with multiple ones of the event conditions funneling data into the same one of the event tables. Further see Sect. IV. The examiner has interpreted that storing transitions into one of three replay memories based on the received rewards as wherein a mapping from the event conditions to the event tables is surjective, with multiple ones of the event conditions funneling data into the same one of the event tables.) PNG media_image1.png 508 724 media_image1.png Greyscale Figure 1: Figure 3 - Overview of the proposed experience replay method [AltContent: textbox (Transitions are stored in one of the three replay memories)][AltContent: ][AltContent: textbox (1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35)] PNG media_image2.png 930 565 media_image2.png Greyscale Figure 2: Algorithm 1 DQN with Stratified Experience Replay As per claim 7, Sharma teaches “where the event conditions represent a handling or manipulation of objects by an artificial agent.” (Pg. 146 Sect. IV, “We model camera selection decision as a finite horizon discounted sum reward problem, where an RL agent is responsible for deciding the presence of a target given a camera frame at a discrete time step t. At each time step, the agent receives the location of the target to be tracked contained in st. The agent needs to select one of the cameras represented using the action space A = {0,1,...,N,C×}. The agent interacts with the environment Ɛ by selecting an action” [where the event conditions represent a handling of objects by an artificial agent]. Further see Sect. IV.) As per claim 10, Sharma teaches “wherein at least some of the event conditions are designed to help the artificial agent retain memory of, and avoid, adverse outcomes.” (Pg. 148 Sect. IV “The multiple replays are; Rf, which stores the transitions which pertain frequently occurring action C×; R−, which stores the transitions that receive a negative reward, and R+ which stores the transitions that receive a positive reward…This transition is stored in a replay which satisfies the above criteria for segregation” [e.g., wherein at least some of the event conditions are designed to help the artificial agent retain memory of, and avoid, adverse outcomes]. Further see Sect. IV.) Re Claim 11, it is a method claim, having similar limitations of claim 1. Thus, claim 11 is also rejected under the similar rationale as cited in the rejection of claim 1. Furthermore, regarding claim 11, Sharma teaches “permitting stratified sampling from the event tables (SSET) for utilizing the event tables with a reinforcement learning system.” (Pg. 145 Sect. I, “This is an important observation with respect to training a deep RL model, which requires appropriately handling of imbalanced state transitions during experience replay” [A method comprising an experience replay buffer for a reinforcement learning agent]. “Our proposed approach, referred to as Stratified Experience Replay (SER) resolves this challenge by sampling in the imbalanced replay memory created by the different episodic runs of the agent-environment interaction. Our specific contributions are the following: 1) We propose a novel experience replay method to segregate transitions into multiple replay memories. Our investigations show that stratified sampling helps learning a better policy for camera selection in a camera network” [permitting stratified sampling from the event tables (SSET) for utilizing the event tables with a reinforcement learning system]. Further see Sect. I. The examiner has interpreted that using stratified sampling to help train a deep reinforcement learning model using the segregate transitions into multiple memories as permitting stratified sampling from the event tables (SSET) for utilizing the event tables with a reinforcement learning system.) As per claim 12, Sharma teaches “blocking sampling from one of the event tables until it has a predetermined minimal data requirement.” (Pg. 148 Sect. V, “For the proposed method, we used three replays to separate the rare transitions, where the size of each replay was set to 103 acquired via hyperparameter tuning” [event tables size or amount of data]. Pg. 147 Sect. IV, “To ensure efficient sampling of rare and other transitions, we segregate transitions in different replay memories to compensate for searching in significantly large replay memory. We create three replay memories named Rf, R+, and R−. Given the three replay memories, we sample a minibatch” [create the tables first then sample, e.g., blocking sampling from one of the event tables until it has a predetermined minimal data requirement]. Furthermore, Pg. 148 Sect. IV, “This transition is stored in a replay which satisfies the above criteria for segregation. For learning, a minibatch is prepared from the stored transitions in the multiple replays. The transitions are sampled uniformly from Rf and R− replay. R+ stores both rare and other positive reward transitions, it follows a prioritized sampling” [storing then sampling, e.g., blocking sampling from one of the event tables until it has a predetermined minimal data requirement]. Further see Sect. IV-V. The examiner has interpreted that segregating transitions into three separate replays of size 103 and then sampling the transitions of the multiple replays as blocking sampling from one of the event tables until it has a predetermined minimal data requirement.) As per claim 14, Sharma teaches “applying a prioritization scheme inside each of the event tables while sampling.” (Pg. 148 Sect. IV, “The transitions are sampled uniformly from Rf, and R− replay. R+ stores both rare and other positive reward transitions, it follows a prioritized sampling. For priority sampling, a higher weight is assigned to the rare transitions and a lower weight to other transitions. A probability value is assigned to ith transition as w i / ∑ w i ” [applying a prioritization scheme inside each of the event tables while sampling]. Further see Sect. IV. The examiner has interpreted that having multiple replays having a uniform and a priority sampling with a probability have that are sampled into three replays as applying a prioritization scheme inside each of the event tables while sampling.) As per claim 15, Sharma teaches “multi-task training to balance gradient updates to respect data from each of the event tables.” (As shown above Figure 2, line 32 performs gradient descent updates. Pg. 150 Sect. VI, “Our experiments showed that SER based DQN resulted in high recall even in camera networks that have long transition times, thus showing that SER successfully balanced rare and frequent actions while learning a policy for camera selection” [multi-task training to balance gradient updates to respect data from each of the event tables]. Further see Sect. IV and VI. The examiner has interpreted that performing gradient descent updates to successfully balance rare and frequent actions while learning a policy for camera selection as multi-task training to balance gradient updates to respect data from each of the event tables.) Re Claim 16, it is a method claim, having similar limitations of claim 1. Thus, claim 16 is also rejected under the similar rationale as cited in the rejection of claim 1. Furthermore, regarding claim 16, Sharma teaches “storing data from all time steps of the reinforcement learning agent into a default buffer; determining at least one event condition; determining at least one preceding state indicating a state preceding the event condition within a history length”. (Pg. 147 Sect. IV, “replay memory of size R is used, which stores the last |R| transitions” [storing data from all time steps of the reinforcement learning agent into a default buffer]. Pg. 146 Sect. III, “during transitions or occlusions, the state at time t captures three elements to handle this partially observable state of the target: 1) xt: the last observed location of the target…, 2) ht: the action history that maintains a list of previously selected actions by the policy. The history is stored as a list of cameras encoded as a one-hot vector. 3) τ: This is the time-elapsed vector that captures the time since the target’s most recent observation by the agent in any camera” [determining at least one event condition; determining at least one preceding state indicating a state preceding the event condition within a history length]. Further see Sect. III-IV. The examiner has interpreted using a replay memory that stores the last transitions where transitions that have states that capture locations of targets, a history of previous selected actions, time of last captured location that was captured by the agent, and store transitions which pertain frequently occurring action as storing data from all time steps of the reinforcement learning agent into a default buffer; determining at least one event condition; determining at least one preceding state indicating a state preceding the event condition within a history length.) Re Claim 17, it is a method claim, having similar limitations of claim 11. Thus, claim 17 is also rejected under the similar rationale as cited in the rejection of claim 11. Re Claim 18, it is a method claim, having similar limitations of claim 12. Thus, claim 18 is also rejected under the similar rationale as cited in the rejection of claim 12. Re Claim 19, it is a method claim, having similar limitations of claim 2. Thus, claim 19 is also rejected under the similar rationale as cited in the rejection of claim 2. Claim Rejections - 35 U.S.C. § 103 The following is a quotation of 35 U.S.C. § 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. § 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. § 102(b)(2)(C) for any potential 35 U.S.C. § 102(a)(2) prior art against the later invention. Claim 8 is rejected under 35 U.S.C. § 103 as being unpatentable over Sharma in view of Wurman, Peter R. et al. "Outracing champion Gran Turismo drivers with deep reinforcement learning." Nature 602, no. 7896 (2022): 223-228 [herein “Wurman”]. As per claim 8, Sharma does not specifically teach “wherein, in a car racing simulator, the event conditions are based on drafting in a car’s slipstream, winning a race, recovering from going off-course, and racing incidents.” However, in the same field of endeavor namely training a reinforcement learning algorithm from an experience replay buffer, Wurman teaches “wherein, in a car racing simulator, the event conditions are based on drafting in a car’s slipstream, winning a race, recovering from going off-course, and racing incidents.” (Pg. 224, “we describe how we used model-free, off-policy deep RL to build a champion-level racing agent, which we call Gran Turismo Sophy (GT Sophy). GT Sophy was developed to compete with the world’s best players of the highly realistic PlayStation 4 (PS4) game Gran Turismo (GT) Sport” [© 2019 Sony Interactive Entertainment Inc] [in a car racing simulator]. Pg. 226, “the opportunities to learn certain skills are rare. We call this the exposure problem; certain states of the world are not accessible to the agent without the ‘cooperation’ of its opponents. For example, to execute a slingshot pass, a car must be in the slipstream of an opponent on a long straightaway, a condition that may occur naturally a few times or not at all in an entire race. If that opponent always drives only on the right, the agent will learn to pass only on the left and would be easily foiled by a human who chose to drive on the left. To address this issue, we developed a process that we called mixed-scenario training. We worked with a retired competitive GT driver to identify a small number of race situations that were probably pivotal on each track” [wherein, in a car racing simulator, the event conditions are based on drafting in a car’s slipstream]. Pg. 229, “The reward function was a hand-tuned linear combination of reward components computed on the transition between the previous state s and current state s′. The reward components were: course progress (Rcp), off-course penalty (Rsoc or Rloc), wall penalty (Rw), tyre-slip penalty (Rts), passing bonus (Rps), any-collision penalty (Rc), rear-end penalty (Rr) and unsporting-collision penalty (Ruc)” [recovering from going off-course, and racing incidents]. Pg. 227, “One of the advantages of using deep RL to develop a racing agent is that it eliminates the need for engineers to program how and when to execute the skills needed to win the race” [winning a race]. Further see Pgs. 224, 226, 227, and 229. The examiner has interpreted that developing a reinforcement learning racing agent in a car racing game that learns states such as slighshotting, off-course penalty, and collision penalty, to win a race as wherein, in a car racing simulator, the event conditions are based on drafting in a car’s slipstream, winning a race, recovering from going off-course, and racing incidents.) Therefore, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to add “wherein, in a car racing simulator, the event conditions are based on drafting in a car’s slipstream, winning a race, recovering from going off-course, and racing incidents” as conceptually seen from the teaching of Wurman, into that of Sharma because this modification of including car racing states for the advantageous purpose of considering advanced sports tactics in teaching computers competitive tasks (Wurman, Pg. 227). Further motivation to combine be that Sharma and Wurman are analogous art to the current claim directed to training a reinforcement learning algorithm from an experience replay buffer. Claim 9 is rejected under 35 U.S.C. § 103 as being unpatentable over Sharma in view of Luo, Jieliang, and Hui Li. "Dynamic experience replay." In Conference on robot learning, pp. 1191-1200. PMLR, 2020 [herein “Luo”]. As per claim 9, Sharma does not specifically teach “wherein, in continuous control problems that already have rich reward signals, event conditions are based on exceeding thresholds in immediate rewards from a state.” However, in the same field of endeavor namely creating custom experience replay buffers for training reinforcement agents, Luo teaches “wherein, in continuous control problems that already have rich reward signals, event conditions are based on exceeding thresholds in immediate rewards from a state”. (Pg. 2 Sect. 2, “the equation can be solved by model free RL algorithms to avoid using dynamics. DDPG is a model-free off-policy RL algorithm for continuous action spaces. In DDPG, an actor policy π : S → A is created to explore the space and store the collected transition (sj, aj ,sj+1, rj) in a replay buffer R” [in a continuous control problem]. Pg. 4, Sect. 3, “During training, all successful transitions that are generated by workers are saved in a pool, which is sampled periodically by each replay buffer and stored in the demonstration zone” [e.g., event conditions are based on rewards from a state]. Pg. 4 Sect. 3, “We use a simple linear reward function based on the distance between the goal pose and the current pose of the timber piece attached to the robot arm for both tasks. Additionally we use a large positive reward (+1000 for the peg-in-hole and +100 for the lap-joint) if the object is within a small distance of the goal pose where x is the current pose of the object, g is the goal pose, ε is a distance threshold, and R is the large positive reward. We use negative distance as our reward function to discourage the behavior of loitering around the goal because the negative distance also contains time penalty” [wherein, in continuous control problems that already have rich reward signals]. Further see Sect. 2-3. The examiner has interpreted that saving successful transitions into a replay buffer for peg-in-hole and lap-joint tasks that contain large positive rewards in a continuous action space as wherein, in continuous control problems that already have rich reward signals, event conditions are based on exceeding thresholds in immediate rewards from a state.) Therefore, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to add “wherein, in continuous control problems that already have rich reward signals, event conditions are based on exceeding thresholds in immediate rewards from a state” as conceptually seen from the teaching of Luo, into that of Sharma because this modification of filtering events for a continuous control problem for the advantageous purpose of efficiently training RL agents on successful samples (Luo, Pg. 1 & 3). Further motivation to combine be that Sharma and Luo are analogous art to the current claim directed creating custom experience replay buffers for training reinforcement agents. Claim 13 is rejected under 35 U.S.C. § 103 as being unpatentable over Sharma in view of Daley, Brett, Cameron Hickert, and Christopher Amato. "Stratified experience replay: Correcting multiplicity bias in off-policy reinforcement learning." arXiv preprint arXiv:2102.11319 (2021) [herein “Daley”]. As per claim 13, Sharma does not specifically teach “determining a bias correction term for stochastic environments.” However, in the same field of endeavor namely training reinforcement learning using experience relay, Daley teaches “determining a bias correction term for stochastic environments.” (Sect. 3, “Dividing this by Pr(𝑠, 𝑎) to eliminate the multiplicity bias, and then normalizing to make the probabilities sum to 1 over the set S×A×S, we arrive at the ideal sampling distribution: Pr(s’ | 𝑠, 𝑎) / |S × A|. Remarkably, this indicates that we can sample from two uniform distributions in succession to counter multiplicity b” and Sect 4, “SER offers a theoretically well-motivated alternative to the uniform distribution for off-policy deep RL methods. By correcting for multiplicity bias, SER helps agents learn significantly faster in small MDPs” [determining a bias correction term for stochastic environments]. Further see Sect. 3-4. The examiner has interpreted that creating an ideal sampling distribution by dividing by an additional factor in a Markov Decision Process as determining a bias correction term for stochastic environments.) Therefore, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to add “determining a bias correction term for stochastic environments” as conceptually seen from the teaching of Daley, into that of Sharma because this modification of adding a bias correction for the advantageous purpose of correcting bias and training the reinforcement learning faster (Daley, Sect. 4). Further motivation to combine be that Sharma and Daley are analogous art to the current claim directed to training reinforcement learning using experience relay. Response to Arguments Applicant’s arguments, see Pg. 18-25, filed July 29, 2026, with respect to the rejection(s) of claims 1-19 under 35 U.S.C. § 101 have been fully considered and are persuasive with regards to the amended independent claims that integrate the claimed invention into a practical application. Therefore, the rejection has been withdrawn. Applicant's arguments filed on July 29, 2026, with respect to the rejection(s) of claim under 35 U.S.C. § 102(a)(1) have been fully considered but they are not persuasive. Applicant argues that reference does not teach each and every limitation in the amended independent claims because cited reference fails to teach “event tables based on history lengths” (See Applicant’s response, Pg. 25-27). MPEP § 2143.03 states “All words in a claim must be considered in judging the patentability of that claim against the prior art” and “Examiners must consider all claim limitations when determining patentability of an invention over the prior art.” As original mapped in the previous Office Action in claim 1, Sharma discloses “event tables based on history lengths” as transitions have states that capture locations of targets, a history of previous selected actions from the policy, a time of last captured location that was captured by the agent, and store transitions which pertain frequently occurring action. Specifically, the history length is the action history that maintains a list of previously selected actions by the policy and the time history observed by the agent. Sharma segregates the transitions based on this stored information in the transitions. Therefore, the claimed limitation is taught. Additional emphasis has been added to this mapping in the rejection above to the amended limitation. Therefore, all of the limitations of the amended independent claims are disclosed in Sharma. Therefore, applicant’s arguments are not persuasive and the rejection of amended independent claims as anticipated by Sharma is maintained. Applicant argues that reference does not teach each and every limitation in the amended independent claims because cited reference fails to teach “sampling individual steps from each event table to build training samples for off-policy reinforcement learning; and performing off-policy training of the reinforcement learning agent” (See Applicant’s response, Pg. 27). MPEP § 2143.03 states “All words in a claim must be considered in judging the patentability of that claim against the prior art” and “Examiners must consider all claim limitations when determining patentability of an invention over the prior art.” As provided for the new limitations in the rejection above, Sharma discloses “sampling individual steps from each event table to build training samples for off-policy reinforcement learning; and performing off-policy training of the reinforcement learning agent” as sampling from the stored transitions to create a minibatch to learn the policy and train a deep reinforcement model applied off-policy. By creating a minibatch that is sampled from the transitions that were stored to learn the policy of a reinforcement learning model that is applied as an off-policy, the claimed limitation is taught. Therefore, all of the limitations of the amended independent claims are disclosed in Sharma. Therefore, applicant’s arguments are not persuasive and the rejection of amended independent claims as anticipated by Sharma is maintained. Applicant argues that the rejection of claim 8 relies upon a disclosure that does not qualify as prior art under the U.S.C. § 102(b)(1)(A) exception (See Applicant’s response, Pg. 28). MPEP § 2153.01(a) states “A disclosure made within the grace period is not prior art under AIA 35 U.S.C. 102(a)(1) if it is apparent from the disclosure itself that it is an inventor-originated disclosure. Specifically, Office personnel may not apply a disclosure as prior art under AIA 35 U.S.C. 102(a)(1) if the disclosure: (1) was made one year or less before the effective filing date of the claimed invention; (2) names the inventor or a joint inventor as an author or an inventor; and (3) does not name additional persons as authors on a printed publication or joint inventors on a patent. This means that in circumstances where an application names additional persons as joint inventors relative to the persons named as authors in the publication (e.g., the application names as joint inventors A, B, and C, and the publication names as authors A and B), and the publication is one year or less before the effective filing date, it is apparent that the disclosure is a grace period inventor disclosure, and the publication is not prior art under AIA 35 U.S.C. 102(a)(1). If, however, the application names fewer joint inventors than a publication (e.g., the application names as joint inventors A and B, and the publication names as authors A, B and C), it would not be readily apparent from the publication that it is an inventor-originated disclosure and the publication would be treated as prior art under AIA 35 U.S.C. 102(a)(1) unless there is evidence of record that an exception under AIA 35 U.S.C. 102(b)(1) applies.” MPEP § 2153.01(a) states “The Office has provided a mechanism for filing an affidavit or declaration (under 37 CFR 1.130) to establish that a disclosure is not prior art under AIA 35 U.S.C. 102(a) due to an exception in AIA 35 U.S.C. 102(b). See MPEP § 717. In the situations in which it is not apparent from the grace period disclosure itself or the patent application specification that the disclosure is an inventor-originated disclosure, the applicant may establish that the AIA 35 U.S.C. 102(b)(1)(A) exception applies by way of an affidavit or declaration under 37 CFR 1.130(a). MPEP § 2155.01 discusses the use of affidavits or declarations to show that a disclosure was an inventor-originated disclosure made during the grace period” (emphasis added). Claim 8 is rejected by a combination of Sharma and Wurman. In addition to the inventors listed in the application (i.e., Varun Kompella, Thomas Walsh, Samuel Barrett, Peter Wurman, and Peter Stone), the Wurman reference also cites the additional inventors: Kenta Kawamoto, James MacGlashan, Kaushik Subramanian, Roberto Capobianco, Alisa Devlic, Franziska Eckert, Florian Fuchs, Leilani Gilpin, Piyush Khandelwal, HaoChih Lin, Patrick MacAlpine, Declan Oller, Takuma Seno, Craig Sherstan, Michael D. Thomure, Houmehr Aghabozorgi, Leon Barrett, Rory Douglas, Dion Whitehead, Peter Dürr, Michael Spranger, and Hiroaki Kitano. As noted in the above MPEP citations, since the additional inventors have not been included in the application, it is not apparent from the publication that the disclosure (i.e. Wurman reference) is an inventor-originated disclosure, and thus it qualifies as prior art under U.S.C. § 102(a)(1). As noted in the above MPEP citations, the applicant can overcome this rejection by filing a signed affidavit or declaration under 37 CFR 1.130(a) that provides details as to what role and contributions were of additional inventors listed in the patent application (i.e. Kenta Kawamoto, James MacGlashan, Kaushik Subramanian, Roberto Capobianco, Alisa Devlic, Franziska Eckert, Florian Fuchs, Leilani Gilpin, Piyush Khandelwal, HaoChih Lin, Patrick MacAlpine, Declan Oller, Takuma Seno, Craig Sherstan, Michael D. Thomure, Houmehr Aghabozorgi, Leon Barrett, Rory Douglas, Dion Whitehead, Peter Dürr, Michael Spranger, and Hiroaki Kitano) to establish that the disclosure is inventor-originated and does not constitute as prior art. Therefore, since Wurman qualifies as prior art, then all of the limitations of claim 8 are disclosed in Sharma or Wurman. Therefore, applicant’s arguments are not persuasive and the rejection of claim 8 over Sharma in view of Wurman is maintained. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Chen, Shi-Yong, Yang Yu, Qing Da, Jun Tan, Hai-Kuan Huang, and Hai-Hong Tang. "Stabilizing reinforcement learning in dynamic environment with application to online recommendation." In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pp. 1187-1196 2018 teaches a method for stabilizing results for reinforcement learning using a sampled buffer using stratified sampling replay. Liu, Yang, Yunan Luo, Yuanyi Zhong, Xi Chen, Qiang Liu, and Jian Peng. "Sequence modeling of temporal credit assignment for episodic reinforcement learning." arXiv preprint arXiv:1905.13420 (2019) teaches evaluating the time step for each trajectory in a replay buffer using stratified sampling. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Examiner’s Note: The examiner has cited particular columns and line numbers in the reference that applied to the claims above for the convenience of the applicant. Although the specified citations are representative of the art and are applied to specific limitations within the individual claim, other passages and figures may apply as well. It is respectfully requested from the applicant, to fully consider the references in their entirety as potentially teaching all or part of the claimed invention, as well as the context of the passage as taught by the prior art or disclosed by the examiner. In the case of amending the claimed invention, the applicant is respectfully requested to indicate the portion(s) of the specification which dictate(s) the structure relied on for the proper interpretation and also to verify and ascertain the metes and bound of the claimed invention. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Simeon P Drapeau whose telephone number is (571)-272-1173. The examiner can normally be reached Monday - Friday, 8 a.m. - 5 p.m. ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Ryan Pitaro can be reached on (571) 272-4071. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SIMEON P DRAPEAU/Examiner, Art Unit 2188 /RYAN F PITARO/Supervisory Patent Examiner, Art Unit 2188
Read full office action

Prosecution Timeline

Apr 06, 2023
Application Filed
Jun 30, 2026
Non-Final Rejection mailed — §102, §103, §112
Jul 29, 2026
Response Filed
Sep 09, 2026
Final Rejection mailed — §102, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12618324
PREDICTING FORMATION PORE PRESSURE IN REAL TIME BASED ON MUD GAS DATA
4y 4m to grant Granted May 05, 2026
Study what changed to get past this examiner. Based on 1 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
19%
Grant Probability
89%
With Interview (+70.3%)
4y 3m (~9m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 16 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month