Prosecution Insights
Last updated: August 18, 2026
Application No. 17/964,704

DEEP REINFORCEMENT LEARNING BASED WIRELESS NETWORK SIMULATOR

Final Rejection §103
Filed
Oct 12, 2022
Priority
Oct 28, 2021 — FI 20216115
Examiner
DIEP, DUY T
Art Unit
2123
Tech Center
2100 — Computer Architecture & Software
Assignee
Nokia Corporation
OA Round
2 (Final)
32%
Grant Probability
At Risk
3-4
OA Rounds
5m
Est. Remaining
52%
With Interview

Examiner Intelligence

Grants only 32% of cases
32%
Career Allowance Rate
10 granted / 31 resolved
-22.7% vs TC avg
Strong +19% interview lift
Without
With
+19.3%
Interview Lift
resolved cases with interview
Typical timeline
4y 4m
Avg Prosecution
21 currently pending
Career history
64
Total Applications
across all art units

Statute-Specific Performance

§101
31.0%
-9.0% vs TC avg
§103
57.8%
+17.8% vs TC avg
§102
3.1%
-36.9% vs TC avg
§112
8.2%
-31.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 31 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment The amendments filed 04/06/2026 have been entered. Claim 1 remain pending in the application. Applicant’s amendments, with respect to the claim objection of claim 10 filed 01/14/2026 have been fully considered and they are persuasive. Therefore, the previous objections as set forth in the previous office action has been removed. Applicant’s amendments, with respect to the claim rejections of claim 5 under 35 U.S.C 112b filed 01/14/2026 have been fully considered and they are persuasive. Therefore, the previous rejections as set forth in the previous office action has been removed. Applicant’s arguments and amendments, with respect to the claim rejection(s) of claim(s) 1-19 under 35 U.S.C 103 filed 01/14/2026 have been fully considered but they are not persuasive. Applicant argues that amended independent claim 1 overcomes the prior §103 rejection. Applicant identifies the amended limitations directed to configuring DRL agents to emulate components of a wireless/mobile network, formulating component behavior as an MDP, augmenting states using a VAE, estimating rewards using distributional regression and GPR, interconnecting DRL agents to emulate real network connections, simulating a real wireless network in a geographic area, using a user agent/mobile device to generate and receive traffic/performance data, using an RNN architecture, obtaining a massive number of states, and estimating rewards for the massive number of states based on similarity of the states. Applicant generally asserts that the cited and applied art does not teach or suggest at least the emphasized amended features. However, Applicant does not provide a specific explanation identifying which particular teaching is missing from the cited references or why the combination fails to render the amended claim obvious. Applicant also notes that claims 2-19 have been cancelled, rendering the rejections of those claims moot. Examiner respectfully disagrees. Applicant’s remarks are not persuasive because Applicant has not identified a specific deficiency in the applied references or provided a particular explanation as to why the prior combination fails to teach or suggest the amended limitations. Rather, Applicant generally asserts that the emphasized amended features are not taught or suggested. As discussed in the rejection, the teachings of Bouton, Andersen, and Levine address the main aspects of the amended claim. For example, Bouton teaches or at least suggests simulating a mobile/wireless communication network using reinforcement learning agents. Bouton models network optimization as a multi-agent Markov decision process, where an agent represents each base station. Bouton further teaches that the agents observe network performance indicators and local/network information as state inputs for the reinforcement learning process. Therefore, Bouton teaches or at least suggests configuring DRL agents to emulate operations of wireless-network components, using states representing wireless-network and component information, and formulating behavior emulation of the respective network component as an MDP. Bouton also teaches or at least suggests interconnection between network devices/agents through a coordination graph or similar network representation. Bouton further teaches communication between network devices, wireless connectivity, and forwarding of traffic between instances. These teachings correspond to interconnecting DRL agents to emulate real connections between components in the wireless network, including base stations, switches, and processing/coordination units. Bouton also teaches using an existing network deployment and geographic locations of base stations, which corresponds to simulating a real wireless network implemented in a certain geographical area. With respect to traffic and performance information, Bouton teaches user equipment, network traffic, throughput, SINR, number of UEs, and other performance indicators used in the reinforcement learning process. Bouton also teaches forwarding traffic between devices or instances. Therefore, Bouton teaches or at least suggests generating and receiving data traffic and user performance information within the wireless network, including information associated with user/mobile devices. With respect to state augmentation, Bouton teaches that in multi-agent reinforcement learning, the possible state/action space may grow exponentially with the number of agents, resulting in a large or massive number of possible states. Andersen further teaches a Dreaming Variational Autoencoder that generates probable future states from state-action pairs for reinforcement learning. Thus, Andersen teaches or at least suggests using a variational autoencoder to augment states and obtain additional states for the DRL agents. With respect to reward estimation, Bouton teaches that reinforcement learning agents receive rewards based on performance indicators after taking actions. Levine further teaches Gaussian Process Inverse Reinforcement Learning, in which rewards are learned as nonlinear functions of features using Gaussian process regression and a distribution over Gaussian process outputs. Levine also teaches that states having similar values for highly weighted features take on similar rewards. Therefore, Levine teaches or at least suggests estimating rewards using distributional regression and Gaussian process regression, including reward estimation based on similarity of the states. Accordingly, Applicant’s general assertion that the cited and applied art does not teach or suggest the amended limitations is not persuasive. The prior teachings of Bouton, Andersen, and Levine already address the main amended features relating to wireless-network DRL simulation, state augmentation using a VAE, and reward estimation using distributional/Gaussian-process regression based on state similarity. Therefore, the rejection of claim 1 under 35 U.S.C. § 103 is maintained. However, upon further consideration, new ground(s) of rejections have been raised (See Below.) Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim 1 is rejected under 35 U.S.C. 103 as being unpatentable over Bouton et.al (US 20240022950 A1) in view of Andersen et.al (NPL: The Dreaming Variational Autoencoder for Reinforcement Learning Environments), further in view of Dunning et.al (US 20210097373 A1), further in view of Levine et.al (NPL: Nonlinear Inverse Reinforcement Learning with Gaussian Processes) Regarding claim 1, Bouton teaches or at least suggest the limitation “A device for simulating a wireless network, comprising: at least one processor; and at least one memory including instructions stored thereon, which when executed by the at least one processor, to cause the device to” (paragraph 81 “a simulation of the target mobile communication network”, and paragraph 5 “The network device includes a non-transitory machine-readable storage medium having stored therein a optimization coordinator, and a processor coupled to the non-transitory machine-readable storage medium. The processor executes the optimization coordinator.” Bouton discloses a simulation of mobile communication network among network devices with wireless connection according to fig.7A, which corresponds to simulating the wireless network, as claimed. The simulation may be carried out by the network device includes a non-transitory machine-readable storage medium and a processor coupled to the non-transitory machine-readable storage medium to perform coordinator optimization of the radio access network (RAN) of the mobile communication network, which include the simulation of the mobile communication network to thereby perform the optimization. Bouton teaches or at least suggest the limitation “configure deep reinforced learning (DRL) agents, wherein each DRL agent is configured to emulate an operation of a component of the wireless network, and each DRL agent is configured to states representing information of the wireless network and information of the component, wherein the states comprise an inner state representing technical inner information of the component and wherein each DRL agent is configured to receive the inner state as an input, wherein the wireless network comprises a mobile network, wherein each DRL agent is configured to formulate behavior emulation of the respective component as a Markovian decision process (MDP)” (paragraph 43 “The embodiments address the problem of optimizing network performance by modeling the problem as a multiagent Markov decision process (MDP) where each base station is an agent”, paragraph 44 “Each base station can observe various performance indicators from the network such as the Signal to Interference and Noise Ratio (SINR), Reference Signal Received Power (RSRP), and/or Channel Quality Indicator (CQI) of each user equipment (UE) connected to it. This information can be processed and used as a state input to the reinforcement learning process”, and paragraph 91 “In cooperative multi agent reinforcement learning, as used herein, a group of agents is trying to maximize a common objective. Each agent controls its action using a policy which is a mapping from its local observation of the environment to an action”. Bouton discloses optimizing network performance by modeling the problem as a multiagent Markov decision process (MDP) where each base station is an agent, which corresponds to DRL agents that emulate an operation of a component of the wireless network as a Markovian decision process (MDP), as claimed. Bouton also discloses each agent (base station) observe various performance indicators from the network, then process and used as a state input to the Reinforcement learning process, which teaches or at least suggests states representing information of the wireless network and information of the component, wherein the states comprise an inner state representing technical inner information of the component, as claimed. Bouton also discloses the mobile communication network as recited above, which corresponds to the mobile network, as claimed. Bouton teaches or at least suggest part of the limitation “wherein the DRL agents are configured to receive and execute training data so that the states are augmented and reward estimated, ...” (paragraph 77 “a reinforcement learning loop with a custom update rule that only uses local information. The first action of the reinforcement loop comprises taking an action and collecting data. Once an optimal joint action is computed through message passing, each agent can decide to take this action ... After the agent takes this action, it receives a reward from the environment ... In this interaction each agent gathers an experience tuple, (si,ai,ri,si′), where si′ is the state observed after applying configuration ai”. Bouton discloses the Reinforcement learning loop at each agent (base station) that only uses local information, taking an action based on the Reinforcement learning result, receives a reward, and observe the state after applying the action, which teaches or at least suggests the execution of training data so that states are augmented and reward estimated by the DRL agents as claimed, because during reinforcement learning, the agent select an action, get rewarded, and obtaining new states after perform the selected action, thereby changing the current state.) Bouton teaches or at least suggest the limitation “inter-connect the DRL agents to emulate real connections between the components in the wireless network, wherein the DRL agent is configured to emulate a base station, a switch, and a data processor unit of the wireless network” (paragraph 33 “In the examples provided herein, each base station includes multiple antennas which are nodes in a network. The embodiments leverage communication between the nodes of the network, through a coordination graph or similar structure, to find a globally optimal joint antenna configuration (as opposed to optimizing each antenna individually) while learning locally”, paragraph 102 “FIG. 7A illustrates connectivity between network devices (NDs) within an exemplary network ... These NDs are physical devices, and the connectivity between these NDs can be wireless”, and paragraph 110 “In certain embodiments, the virtualization layer 754 includes a virtual switch that provides similar forwarding services as a physical Ethernet switch. Specifically, this virtual switch forwards traffic between instances”. Bouton discloses each base station includes multiple antennas which are nodes in a network and the embodiment leverage communication between the nodes of the network, through a coordination graph or similar structure, to represent the network communication, which teaches or at least suggests emulating real connections between the components in the wireless network, wherein the DRL agent is configured to emulate a base station, as claimed. Furthermore, Bouton discloses at fig 7.A communication between network devices, in which a person of ordinary skill in the art would have understood that the network devices are provided as exemplary physical implementation structures for carrying out node/base-station functionality discussed above. Since Bouton discloses the base stations as network nodes/agents that communicate through a coordination graph, a person of ordinary skill in the art would have understood that the base stations may be implemented by, or corresponds to, network devices executing the relevant communication and coordination functions, wherein each network device comprises of a virtual switch that provides similar forwarding services as a physical Ethernet switch switch and processor unit.) Bouton teaches or at least suggest the limitation “execute the DRL agents based on the states as inputs to simulate the wireless network online, wherein each DRL agent is configured to emulate an individual component in a real wireless network, wherein the component comprises the individual component and the wireless network comprises the real wireless network implemented in a certain geographical area” (paragraph 42 “In addition, the coordination graph or similar representation provides an excellent support for encoding prior knowledge and enabling knowledge transfer between simulation and real world base stations, as well as between different networks”, and paragraph 62 “In one example, an automatic construction of the coordination graph can use an existing network deployment. A coordination graph can be built using geographical locations of each base station. Each base station includes three antennas which are connected to each other”. Bouton discloses the execution of Reinforcement learning loop of each agent (base station) with states as input, and represents the connection of agent (base station) using a graph representation, thereby enabling knowledge transfer between simulation and real-world base stations, which corresponds to or at least suggest the simulation of the wireless network online, wherein each DRL agent is configured to emulate an individual component in a real wireless network, as claimed. Furthermore, the communication between base stations (agents/network devices) may be represented by a coordination graph that use an existing network deployment based on geographical locations of each base station, which corresponds to or at least suggest the real wireless network for the communication between agents implemented in a certain geographical area, as claimed. Bouton teaches or at least suggest the limitation “the user agent configured to generate data traffic of the wireless network and performances of the user within the wireless network” (paragraph 77 “The reward signal can include any type of performance indicator measurable by a base station such as the average SINR, the number of UEs with SINR greater than a threshold, the average throughput, the 10th percentile throughput or SINR for example, or similar performance indicators”, and paragraph 110 “In certain embodiments, the virtualization layer 754 includes a virtual switch that provides similar forwarding services as a physical Ethernet switch. Specifically, this virtual switch forwards traffic between instance”. Bouton discloses the communication between each base station (agent/network device) via a graph representation as disclosed above, which corresponds to the user agent configured to emulate operations of a user device of the wireless network, as claimed. Bouton further discloses that each network device may comprises of virtualization layer to forward traffic between devices, thus teach or at least suggesting the generating of data traffic of the network at each device. Bouton also discloses the reinforcement learning includes performance indicator, which corresponds to the performances of the user within the wireless network, as claimed.) Bouton teaches or at least suggest the limitation “wherein the DRL agents are configured to receive the data traffic and the performances of the user within the wireless network,...” (paragraph 77 “The reward signal can include any type of performance indicator measurable by a base station such as the average SINR, the number of UEs with SINR greater than a threshold, the average throughput, the 10th percentile throughput or SINR for example, or similar performance indicators”, paragraph 87 “To perform the updates, the following information needs to be shared: ... a performance indicator which can often be deduced from the knowledge”, and paragraph 110 “In certain embodiments, the virtualization layer 754 includes a virtual switch that provides similar forwarding services as a physical Ethernet switch. Specifically, this virtual switch forwards traffic between instance”. Bouton discloses that each network device (base station/agent) may comprises of virtualization layer to forward traffic between devices, thus teach or at least suggesting the generating of data traffic and communicating the data traffic to other device. Bouton also discloses the reinforcement learning includes performance indicator, which can be shared to other base station (network device/agent), thereby teach or at least suggest the receiving of performance of the user within the wireless network, as claimed.) Bouton teaches or at least suggest the limitation “wherein the user device comprises a mobile device” (paragraph 114 “The NDs of FIG. 7A, for example, may form part of the Internet or a private network; and other electronic devices (not shown; such as end user devices including workstations, laptops, netbooks, tablets, palm tops, mobile phones”. Bouton discloses the network devices (base stations/agents) may form part of the internet network, which further comprise communication with end user devices such as mobile phones, thereby corresponding to, or suggest the user device comprises a mobile device, as claimed.) Bouton teaches or at least suggest the limitation “wherein the device is configured to augment the states so that a massive number of states is obtained for the DRL agent” (paragraph 93 “In multi-agent reinforcement learning, the space of possible state and action grows exponentially with the number of agents making Q very difficult to represent and finding the optimal joint action is also challenging. To address this curse of dimensionality, one can rely on function approximation and represent Q by a neural network for example” Bouton discloses multi-agent reinforcement learning, wherein the agent continuing learning the state with new action being selected and reward, therefore, the state grows exponentially, which teach or at least suggest the augmenting of the state so that a massive number of states is obtained, because during reinforcement learning, the agent continuously select an action, get rewarded, and obtaining new states after perform the selected action, thereby changing the current state, which cause the state to grow exponentially, thereby corresponds to or at least suggest the augmenting step for a massive number of states, as claimed.) Bouton does not teach a part of the limitation “... wherein for augmenting, the device is further configured to use an autoencoder to augment the states, wherein the autoencoder comprises a variational autoencoder (VAE),...” However, Andersen teaches or at least suggest this (page 3 section 4 “The Dreaming Variational Autoencoder (DVAE) is an end-to-end solution for generating probable future states ... from an arbitrary state-space S using state-action pairs explored prior ...”, and page 4 figure 1 “Illustration of the DVAE model. The model consumes state and action pairs, yielding the input encoded in latent-space. Latent-space can then be decoded to a probable future state.” Andersen discloses the Dreaming Variational Autoencoder (DVAE), a neural network based generative modeling architecture for exploration in environments with sparse feedback. The DVAE can generate probable future states from an arbitrary state-space S using state-action pairs, which corresponds to use an autoencoder to augment the states, as claimed.) Before the effective filing date, it would have been obvious to one of ordinary skill in the art to combine the teaching of a method and system for optimizing radio access networks using reinforcement learning by Bouton with the teaching of the Dreaming Variational Autoencoder by Andersen. The motivation to do so is referred to in Andersen’s disclosure (page 2 section 1 “By combining the ideas of variational autoencoders with deep RL agents, we find that it is possible for agents to learn optimal policies using only generated training data samples. The approach is presented as the dreaming variational autoencoder”, and page 13 section 7 “This paper introduces the Dreaming Variational Autoencoder (DVAE) as a neural network based generative modeling architecture to enable exploration in environments with sparse feedback. The DVAE shows promising results in modeling simple non-continuous environments” Andersen discloses the motivation to combining the ideas of variational autoencoders with deep RL agents, thus help the RL agents to learn optimal policies using only generated training data samples in environments with sparse feedback. Therefore, one of ordinary skilled in the art would have been motivated to combine the teaching by Bouton with the teaching of Andersen to apply the DVAE to generate probable future states that does not depend too much on the environment, thereby improve the Reinforcement learning framework by Bouton.) Bouton/Andersen does not teach a part of the limitation “...wherein the DRL agents are configured to a recurrent neural network (RNN) architecture” However, Dunning teaches or at least suggest this (paragraph 41 “the system 100 can select the action to be performed by the agent in accordance with an exploration policy ...”, paragraph 42 “The policy neural network 110 includes a temporally hierarchical recurrent neural network 112 ... includes two recurrent neural networks: a fast updating recurrent neural network (RNN) that updates its hidden state at every time step and a slow updating recurrent neural network that updates its hidden state at less than all of the time steps.” Dunning discloses a Reinforcement learning method and system, wherein the agent can select an action to be performed in accordance with an exploration policy, wherein the agent employs a policy neural network that includes two recurrent neural networks to update the states in view of the action selected and reward, thereby teach or at least suggests the DRL agents are configured to a recurrent neural network (RNN) architecture, as claimed.) Bouton/Andersen does not teach the limitation “wherein the user agent is configured to the RNN architecture” However, Dunning teaches or at least suggest this (paragraph 41 “the system 100 can select the action to be performed by the agent in accordance with an exploration policy ...”, paragraph 42 “The policy neural network 110 includes a temporally hierarchical recurrent neural network 112 ... includes two recurrent neural networks: a fast updating recurrent neural network (RNN) that updates its hidden state at every time step and a slow updating recurrent neural network that updates its hidden state at less than all of the time steps.” Dunning discloses a Reinforcement learning method and system, wherein the agent can select an action to be performed in accordance with an exploration policy, wherein the agent employs a policy neural network that includes two recurrent neural networks to update the states in view of the action selected and reward, thereby teach or at least suggests the DRL agents are configured to a recurrent neural network (RNN) architecture, as claimed.) Before the effective filing date, it would have been obvious to one of ordinary skill in the art to combine the teaching of a method and system for optimizing radio access networks using reinforcement learning by Bouton, and the teaching of the Dreaming Variational Autoencoder by Andersen with the teaching of the user agent perform action selection based on a policy recurrent neural network by Dunning. The motivation to do so is referred to in ...’s disclosure (paragraph 46 “The policy neural network 110 then uses this updated fast updating hidden state to generate the action selection output 122. By making use of the temporally hierarchical RNN 112, the system 100 can select actions at each time step that are consistent with long-term plans for the agent and the system 100 can effectively control the agent even on tasks that can require that the action that is selected at any given time step be dependent on data received in observations at time steps that are a large number of time steps before the given time step.” Dunning discloses the benefit of employing a temporally hierarchical RNN in configuring the action to be selected by the agent in view of the state and reward. The recurrent neural network architecture can help the agent select actions at each time step that are consistent with long-term plans and effectively control the agent in selecting an action at any given time step. Therefore, one of ordinary skill in the art would have been motivated to further employ the temporally hierarchical RNN architecture for the neural network of the agent within Reinforcement learning, thereby improve the Reinforcement learning framework by Bouton/Andersen.) Bouton/Andersen/Dunning does not teach a part of the limitation “... wherein for the reward estimating the device is further configured to use distributional regression and gaussian process regression (GPR)” However, Levine teaches or at least suggest this limitation (page 1 section 1 “Previous IRL algorithms generally learn the reward as a linear combination of features, either by finding a reward under which the expert’s policy has a higher value than all other policies ... or else by maximizing the probability of the reward under a model of near-optimal expert behavior ... GPIRL is the first method to combine probabilistic reasoning about stochastic expert behavior with the ability to learn the reward as a nonlinear function of features, allowing it to outperform prior methods on tasks with inherently nonlinear rewards and suboptimal examples”, and page 2 section 3 “GPIRL represents the reward as a nonlinear function of feature values. This function is modeled as a Gaussian process, and its structure is determined by its kernel function. The Bayesian GP framework provides a principled method for learning the hyperparameters of this kernel, thereby learning the structure of the unknown reward ... we use Equation 1 to specify a distribution over GP outputs ... In GP regression, we use noisy observations y of the true underlying outputs u. GPIRL directly learns the true outputs u, which represent the rewards” Levine discloses using Gaussian Process Inverse Reinforcement Learning (GPIRL) method to combine probabilistic reasoning about stochastic expert behavior with the ability to learn the reward as a nonlinear function of features. Essentially, Levine discloses using Gaussian processes to learn the reward as a nonlinear function, wherein the Gaussian processes comprise GP regression equation to specify a distribution of rewards, which corresponds to the distributional regression and gaussian process regression for reward estimating, as claimed.) Bouton/Andersen/Dunning does not teach the limitation “wherein the device is configured to reward estimate the massive number of states by distributional regression based on similarity of the states” However, Levine teaches or at least suggest this (page 2 section 3 “GPIRL represents the reward as a nonlinear function of feature values. This function is modeled as a Gaussian process, and its structure is determined by its kernel function. The Bayesian GP framework provides a principled method for learning the hyperparameters of this kernel, thereby learning the structure of the unknown reward ... we use Equation 1 to specify a distribution over GP outputs ... In GP regression, we use noisy observations y of the true underlying outputs u. GPIRL directly learns the true outputs u, which represent the rewards ... In GP regression, we use noisy observations y of the true underlying outputs u. GPIRL directly learns the true outputs u, which represent the rewards”, and page 3 section 3 “States distinguished by highly-weighted features can take on different reward values, while those that have similar values for all highly-weighted features take on similar rewards” Levine disclose GPIRL as the function to calculate unknown reward using Gaussian Process and a distribution over Gaussian Process outputs, wherein the rewards are generated for many states such that the states with similar values for all highly-weighted features take on similar rewards, which corresponds to or at least suggest the claimed process of reward estimate the massive number of states by distributional regression based on similarity of the states, as claimed.) Before the effective filing date, it would have been obvious to one of ordinary skill in the art to combine the teaching of a method and system for optimizing radio access networks using reinforcement learning by Bouton, the teaching of the Dreaming Variational Autoencoder by Andersen, and the teaching of the user agent perform action selection based on a policy recurrent neural network by Dunning, with the teaching of Gaussian Process Inverse Reinforcement Learning (GPIRL) method by Levine. The motivation to do so is referred to in Levine’s disclosure (page 1 section 1 “Previous IRL algorithms generally learn the reward as a linear combination of features, either by finding a reward under which the expert’s policy has a higher value than all other policies ... or else by maximizing the probability of the reward under a model of near-optimal expert behavior ... GPIRL is the first method to combine probabilistic reasoning about stochastic expert behavior with the ability to learn the reward as a nonlinear function of features, allowing it to outperform prior methods on tasks with inherently nonlinear rewards and suboptimal examples” Levine discloses the GPIRL is the first method to combine probabilistic reasoning about stochastic expert behavior with the ability to learn the reward as a nonlinear function of features, allowing it to outperform prior methods on tasks with inherently nonlinear rewards and suboptimal examples. Therefore, one ordinary skilled in the art would have been motivated to incorporate the GPIRL method to perform learning of reward for the reinforcement learning to further improve the reward function within the reinforcement learning framework by Bouton/Andersen/Dunning.) Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to DUY TU DIEP whose telephone number is (703)756-1738. The examiner can normally be reached M-F 8-4:30. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey Shmatov can be reached at (571) 270-3428. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /DUY T DIEP/Examiner, Art Unit 2123 /ALEXEY SHMATOV/Supervisory Patent Examiner, Art Unit 2123
Read full office action

Prosecution Timeline

Oct 12, 2022
Application Filed
Jan 14, 2026
Non-Final Rejection mailed — §103
Apr 06, 2026
Response Filed
Jun 24, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12651158
NEURAL NETWORK TRAINING METHOD AND APPARATUS USING TREND
4y 1m to grant Granted Jun 09, 2026
Patent 12608642
MODEL PARAMETER LEARNING METHOD AND MOVEMENT MODE DETERMINATION METHOD
4y 7m to grant Granted Apr 21, 2026
Patent 12579428
METHOD FOR INJECTING HUMAN KNOWLEDGE INTO AI MODELS
4y 3m to grant Granted Mar 17, 2026
Patent 12488223
FEDERATED LEARNING FOR TRAINING MACHINE LEARNING MODELS
3y 11m to grant Granted Dec 02, 2025
Patent 12412129
DISTRIBUTED SUPPORT VECTOR MACHINE PRIVACY-PRESERVING METHOD, SYSTEM, STORAGE MEDIUM AND APPLICATION
4y 4m to grant Granted Sep 09, 2025
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
32%
Grant Probability
52%
With Interview (+19.3%)
4y 4m (~5m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 31 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month