DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This action is in response to a request for continued examination filed on May 12th, 2026. Claims 1-20 are pending in the current application, with claims 1, 8, and 15 being currently amended.
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on May 12th, 2026, has been entered.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-5, 8-12, and 15-19, are rejected under 35 U.S.C. 103 as being unpatentable over Di Wu et al. (Herein referred to as Wu) (Evading Machine Learning Botnet Detection Models via Deep Reinforcement Learning) in view of Ali Alizadeh et al. (Herein referred to as Alizadeh) (Automated Lane Change Decision Making using Deep Reinforcement Learning in Dynamic and Uncertain Highway Environment) and in further view of Borrajo et al. (Herein referred to as Borrajo) (Simulating and classifying behavior in adversarial environments based on action-state traces: an application to money laundering)
Regarding claim 1, Wu teaches a method comprising: training a reinforcement learning agent to learn an electronic policy data structure that maps states of a system (“The gym framework provides a standardized environment to produce benchmarks and trains the RL agent through some methods: reset, step, and render…”, pg. 3, right column, under “Fig. 1. Structure of the framework”) actions of the reinforcement learning agent with transition values, (“The agent follows a policy [Symbol font/0x70] (a|s) based on the estimated value determined by the s and r… An optimal policy can be built from the optimal Q-function by choosing, for a given state, the action with highest Q-value (i.e. Q-learning). However, because the space in the Arcade Learning Environment (ALE) is too large to tractably store a tabular representation of the Q-function, the Deep Q-Network (DQN) which uses a deep function (e.g. CNN) approximator to represent the state-action value function was proposed [21].”, pg. 3, left and right columns, under “A. Deep Reinforcement Learning”) (The parameters for the DQN correspond to transition values.) an electronic policy data structure that controls the reinforcement learning agent, when executed by one or more processors, to perform a task (“The agent follows a policy [Symbol font/0x70] (a|s) based on the estimated value determined by the s and r.”, pg. 3, left column, under “A. Deep Reinforcement Learning”) and evade one or more scenarios of a monitoring system that are configured to hinder the task; (“A reinforcement learning model consists of an agent and an environment. For each turn, the environment receives the action a chosen by the agent, and feeds back the observed state s[Symbol font/0xA2] (after executing a) and reward r. The agent follows a policy [Symbol font/0x70] (a|s) based on the estimated value determined by the s and r… In the context of botnet traffic evasion, we apply DQN in a reinforcement learning framework, as shown in Figure 1. “, pg. 3, left column under “A. Deep Reinforcement Learning “; pg. 3, right column, under “B. Framework Structure”) (The agent is configured to evade detection from a deep learning detector. The detector corresponds to a monitoring system) wherein the sampling includes accessing the transition values stored in the electronic policy data structure to select an action by the reinforcement learning agent (“An optimal policy can be built from the optimal Q-function by choosing, for a given state, the action with highest Q-value (i.e. Q-learning). However, because the space in the Arcade Learning Environment (ALE) is too large to tractably store a tabular representation of the Q-function, the Deep Q-Network (DQN) which uses a deep function (e.g. CNN) approximator to represent the state-action value function was proposed [21].”, pg. 3, left and right columns) (The parameters of the DQN correspond to transition values which determine Q-values, which are then used to select an action.) analyzing the steps taken in the episode to measure a strength of monitoring in the monitoring system; (“Through the score for each query, the attacker is able to directly measure the efficacy of any perturbation to the target black-box model.” Pg. 2, right column, under Score-based attack) (The more effective the attacker, the less effective the monitoring system is. For figures demonstrating this, see Wu’s TABLE III below) and presenting the strength of monitoring in an interface. (“…we use OpenAI gym as the environment interface, which is a toolkit for developing and comparing reinforcement learning algorithms. The gym framework provides a standardized environment to produce benchmarks”, pg. 3, right column, under “Fig. 1. Structure of the framework”) (Also see Wu’s TABLE III below for statistics involving Detection Model evasions, which teaches the strength of the monitoring system)
However, Wu does not explicitly teach to learn an electronic policy data structure that maps states of an electronic transaction system nor after training of the reinforcement learning agent is complete, sampling the policy to simulate an episode of steps taken by the reinforcement learning agent nor to write a step row into an episode data structure that records alert states of one or more scenarios resulting from executing the action;
Alizadeh teaches after training of the reinforcement learning agent is complete, sampling the policy to simulate an episode of steps taken by the reinforcement learning agent (“For this purpose, we define two periods, by which validation phase flag is activated, and the latest network weights are being recorded and exploited during validation. Depending on which period, the agent’s performance is evaluated for several episodes, and the achieved mean reward is compared to the latest maximum value. This process enables the agent to record the best trained model by validating on unseen scenarios.”, pg. 2, right column, second paragraph) and to write a step row into an episode data structure that records alert states of one or more scenarios resulting from executing the action; (“In Q-learning, a memory table Q[s,a] is built to store the Q-values for all the possible combinations of states and actions. By taking action on the current state, the reward Rand the new states are acquired to take the next action that has the maximum Q(s,a) in the memory table… However, if the combinations of state and actions are too large or states and actions are continuous, the memory and computation requirement for action-value function Q will be too high. To address this issue, Deep Q-Network (DQN) is utilized that approximates the action-value function Q(s,a)… Depending on which period, the agent’s performance is evaluated for several episodes, and the achieved mean reward is compared to the latest maximum value.”, pg. 2, left column, second to last paragraph; pg. 2, right column, above “III. SIMULATION SETUP”) (The Q-values, related to states, actions, and rewards in the memory table, correspond to step rows written into an episodic data structure.)
Therefore, it would have been considered obvious to one of ordinary skill in the art, prior to the current application’s filing date, to combine Wu’s Adversarial Reinforcement Learning Agent, with the policy sampling and episodic data structure of Alizadeh. One would be motivated to combine the teachings, prior to the filing date of the current application, as this allows for an optimized agent to be trained fast and be more generalized. (“This process enables the agent to record the best trained model by validating on unseen scenarios. Defining two various periods with a different number of episodes helps the training to be faster and record a more generalized model at the same time.”, pg. 2, right column, above “III. SIMULATION SETUP”)
However, Wu, as modified by Alizadeh, does not explicitly teach to learn an electronic policy data structure that maps states of an electronic transaction system.
Borrajo teaches an electronic transaction system electronic transaction system. (See Tables 1, 2 and 3) with states to learn a policy data structure (“The learning system takes as input traces of observable behavior. A trace 𝑡𝐶 is a sequence of states and actions executed by 𝐶 in those states: 𝑡𝐶 = (𝑠0, 𝑎1, 𝑠1, 𝑎2, 𝑠2, . . . , 𝑠𝑛−1, 𝑎𝑛, 𝑠𝑛), where 𝑠𝑖 is a state and 𝑎𝑖 is an action name and its parameters. States and actions correspond to the observable predicates and actions from the viewpoint of 𝐹 .”, pg. 3, right column, under “2.3 Traces of Behavior”) (While policies are not taught in this disclosure, the actions that are map to states of an electronic transaction system are taught, which in combination with the learning of a policy as disclosed in Wu, the limitation is fully taught.)
Therefore, it would have been considered obvious to one of ordinary skill in the art, prior to the current application’s filing date, to combine Wu’s Adversarial Reinforcement Learning Agent, modify it with Alizadeh, and combine it with the actions, predicates, and functions of Borrajo. One would be motivated to combine the teachings, prior to the filing date of the current application, as Borrajo’s method helps detect and classify good and bad behavior in a financial setting. (“The learning task can be defined as follows. Given: 𝑁 classes of behavior, ({good, bad} in our current application);3 and a set of labeled observed traces, 𝑇𝐶𝑖 , ∀𝐶𝑖 ∈ {g𝑜𝑜𝑑, 𝑏𝑎𝑑} Obtain: a classifier that takes as input a new (partial) trace 𝑡 (with unknown class) and outputs the predicted class… Good refers to standard customers’ behavior and bad corresponds to money laundering-related behavior.” Pg. 4, left column, under “3.1 Learning Task”)
Regarding claim 8, Wu teaches a computing system comprising: a processor; a memory operably connected to the processor; a non-transitory computer-readable medium operably connected to the processor and memory and storing computer-executable instructions (“In this paper, we propose a more general framework based on deep reinforcement learning (DRL), which effectively generates adversarial traffic flows to deceive the detection model by automatically adding perturbations to samples.”, pg. 1 Abstract) (This teaches the computer system, processor, and non-transitory computer readable medium operably connected to the processor and memory, as one would need a system comprising such features to run a method that generates adversarial traffic flows) to learn an electronic policy data structure that maps states of a system (“The gym framework provides a standardized environment to produce benchmarks and trains the RL agent through some methods: reset, step, and render.”, pg. 3, right column, under “Fig. 1. Structure of the framework”) actions of the reinforcement learning agent with transition values, (“The agent follows a policy [Symbol font/0x70] (a|s) based on the estimated value determined by the s and r… An optimal policy can be built from the optimal Q-function by choosing, for a given state, the action with highest Q-value (i.e. Q-learning). However, because the space in the Arcade Learning Environment (ALE) is too large to tractably store a tabular representation of the Q-function, the Deep Q-Network (DQN) which uses a deep function (e.g. CNN) approximator to represent the state-action value function was proposed [21].”, pg. 3, left and right columns, under “A. Deep Reinforcement Learning”) (The parameters for the DQN correspond to transition values.) an electronic policy data structure that controls the reinforcement learning agent, when executed by one or more processors, (“The agent follows a policy [Symbol font/0x70] (a|s) based on the estimated value determined by the s and r.”, pg. 3, left column, under “A. Deep Reinforcement Learning”) and evade one or more scenarios of a monitoring system that are configured to hinder the task; (“A reinforcement learning model consists of an agent and an environment. For each turn, the environment receives the action a chosen by the agent, and feeds back the observed state s[Symbol font/0xA2] (after executing a) and reward r. The agent follows a policy [Symbol font/0x70] (a|s) based on the estimated value determined by the s and r… In the context of botnet traffic evasion, we apply DQN in a reinforcement learning framework, as shown in Figure 1. “, pg. 3, left column under “A. Deep Reinforcement Learning “; pg. 3, right column, under “B. Framework Structure”) (The agent is configured to evade detection from a deep learning detector. The detector corresponds to a monitoring system) wherein the sampling includes accessing the transition values stored in the electronic policy data structure to select an action by the reinforcement learning agent (“An optimal policy can be built from the optimal Q-function by choosing, for a given state, the action with highest Q-value (i.e. Q-learning). However, because the space in the Arcade Learning Environment (ALE) is too large to tractably store a tabular representation of the Q-function, the Deep Q-Network (DQN) which uses a deep function (e.g. CNN) approximator to represent the state-action value function was proposed [21].”, pg. 3, left and right columns) (The parameters of the DQN correspond to transition values which determine Q-values, which are then used to select an action.) analyzing the steps taken in the episode to measure a strength of monitoring in the monitoring system; (“Through the score for each query, the attacker is able to directly measure the efficacy of any perturbation to the target black-box model.” Pg. 2, right column, under Score-based attack) (The more effective the attacker, the less effective the monitoring system is. For figures demonstrating this, see Wu’s TABLE III below) and presenting the strength of monitoring in an interface. (“…we use OpenAI gym as the environment interface, which is a toolkit for developing and comparing reinforcement learning algorithms. The gym framework provides a standardized environment to produce benchmarks”, pg. 3, right column, under “Fig. 1. Structure of the framework”) (Also see Wu’s TABLE III below for statistics involving Detection Model evasions, which teaches the strength of the monitoring system)
However, Wu does not explicitly teach to learn an electronic policy data structure that maps states of an electronic transaction system nor to transfer an amount from a source account to a destination account nor one or more scenarios of a monitoring system that are configure to hinder a transfer, nor once the policy has been learned, sample the policy to simulate multiple episodes of steps taken by the reinforcement learning agent without repeating the training process, nor to write a step row into an episode data structure that records alert states of one or more scenarios resulting from executing the action;
Alizadeh teaches once the policy has been learned, sample the policy to simulate multiple episodes of steps taken by the reinforcement learning agent without repeating the training process, (“For this purpose, we define two periods, by which validation phase flag is activated, and the latest network weights are being recorded and exploited during validation. Depending on which period, the agent’s performance is evaluated for several episodes, and the achieved mean reward is compared to the latest maximum value. This process enables the agent to record the best trained model by validating on unseen scenarios.”, pg. 2, right column, second paragraph) and to write a step row into an episode data structure that records alert states of one or more scenarios resulting from executing the action; (“In Q-learning, a memory table Q[s,a] is built to store the Q-values for all the possible combinations of states and actions. By taking action on the current state, the reward Rand the new states are acquired to take the next action that has the maximum Q(s,a) in the memory table… However, if the combinations of state and actions are too large or states and actions are continuous, the memory and computation requirement for action-value function Q will be too high. To address this issue, Deep Q-Network (DQN) is utilized that approximates the action-value function Q(s,a)… Depending on which period, the agent’s performance is evaluated for several episodes, and the achieved mean reward is compared to the latest maximum value.”, pg. 2, left column, second to last paragraph; pg. 2, right column, above “III. SIMULATION SETUP”) (The Q-values, related to states, actions, and rewards in the memory table, correspond to step rows written into an episodic data structure.)
Therefore, it would have been considered obvious to one of ordinary skill in the art, prior to the current application’s filing date, to combine Wu’s Adversarial Reinforcement Learning Agent, with the policy sampling and episodic data structure of Alizadeh. One would be motivated to combine the teachings, prior to the filing date of the current application, as this allows for an optimized agent to be trained fast and be more generalized. (“This process enables the agent to record the best trained model by validating on unseen scenarios. Defining two various periods with a different number of episodes helps the training to be faster and record a more generalized model at the same time.”, pg. 2, right column, above “III. SIMULATION SETUP”)
However, Wu, as modified by Alizadeh, does not explicitly teach to learn an electronic policy data structure that maps states of an electronic transaction system nor to transfer an amount from a source account to a destination account nor one or more scenarios of a monitoring system that are configured to hinder a transfer.
Borrajo teaches an electronic transaction system electronic transaction system. (See Tables 1, 2 and 3) with states to learn a policy data structure (“The learning system takes as input traces of observable behavior. A trace 𝑡𝐶 is a sequence of states and actions executed by 𝐶 in those states: 𝑡𝐶 = (𝑠0, 𝑎1, 𝑠1, 𝑎2, 𝑠2, . . . , 𝑠𝑛−1, 𝑎𝑛, 𝑠𝑛), where 𝑠𝑖 is a state and 𝑎𝑖 is an action name and its parameters. States and actions correspond to the observable predicates and actions from the viewpoint of 𝐹 .”, pg. 3, right column, under “2.3 Traces of Behavior”) (While policies are not taught in this disclosure, the actions that are map to states of an electronic transaction system are taught, which in combination with the learning of a policy as disclosed in Wu, the limitation is fully taught.) to transfer an amount from a source account to a destination account (“For each action-state pair, we created standard attributes used by other works for the two partial observability models (under the bank and full models, we could observe all these attributes). Examples are average, min and max values of the previous transactions of each type (e.g. wires, or deposits), balance of accounts or number of connected accounts”, pg. 7, left column, bottom paragraph) and one or more scenarios of a monitoring system that are configured to hinder a transfer. (“In this paper, we focus on AML. Over time, financial institutions have been mandated by law enforcement agencies to improve their processes to detect suspicious activity and raise the corresponding Suspicious Activity Reports (SARs). A typical prevalent AML model starts by observing transactions, public media, or a referral, and generates alerts [11]. Then, alerts are investigated by humans who decide whether they need to report a SAR to law enforcement for the alert.”, pg. 2, right column, second paragraph) (AMLs are designed to detect suspicious activity and prevent fraudulent transactions, hindering the transfer of an amount of money.)
Therefore, it would have been considered obvious to one of ordinary skill in the art, prior to the current application’s filing date, to combine Wu’s Adversarial Reinforcement Learning Agent, modify it with Alizadeh, and combine it with the actions, predicates, and functions of Borrajo. One would be motivated to combine the teachings, prior to the filing date of the current application, as Borrajo’s method helps detect and classify good and bad behavior in a financial setting. (“The learning task can be defined as follows. Given: 𝑁 classes of behavior, ({good, bad} in our current application);3 and a set of labeled observed traces, 𝑇𝐶𝑖 , ∀𝐶𝑖 ∈ {g𝑜𝑜𝑑, 𝑏𝑎𝑑} Obtain: a classifier that takes as input a new (partial) trace 𝑡 (with unknown class) and outputs the predicted class… Good refers to standard customers’ behavior and bad corresponds to money laundering-related behavior.” Pg. 4, left column, under “3.1 Learning Task”)
Regarding claim 15, Wu teaches a non-transitory computer-readable medium having stored thereon computer-executable instructions (“In this paper, we propose a more general framework based on deep reinforcement learning (DRL), which effectively generates adversarial traffic flows to deceive the detection model by automatically adding perturbations to samples.”, pg. 1 Abstract) (This teaches the non-transitory computer readable medium, as one would need a medium to run a method that generates adversarial traffic flows) train a reinforcement learning agent to learn an electronic policy data structure that maps states of a system (“The gym framework provides a standardized environment to produce benchmarks and trains the RL agent through some methods: reset, step, and render.”, pg. 3, right column, under “Fig. 1. Structure of the framework”) actions of the reinforcement learning agent with transition values, (“The agent follows a policy [Symbol font/0x70] (a|s) based on the estimated value determined by the s and r… An optimal policy can be built from the optimal Q-function by choosing, for a given state, the action with highest Q-value (i.e. Q-learning). However, because the space in the Arcade Learning Environment (ALE) is too large to tractably store a tabular representation of the Q-function, the Deep Q-Network (DQN) which uses a deep function (e.g. CNN) approximator to represent the state-action value function was proposed [21].”, pg. 3, left and right columns, under “A. Deep Reinforcement Learning”) (The parameters for the DQN correspond to transition values.) an electronic policy data structure that controls the reinforcement learning agent, when executed by one or more processors, to perform a task (“The agent follows a policy [Symbol font/0x70] (a|s) based on the estimated value determined by the s and r.”, pg. 3, left column, under “A. Deep Reinforcement Learning”) and evade one or more scenarios of a monitoring system that are configured to hinder the task; (“A reinforcement learning model consists of an agent and an environment. For each turn, the environment receives the action a chosen by the agent, and feeds back the observed state s[Symbol font/0xA2] (after executing a) and reward r. The agent follows a policy [Symbol font/0x70] (a|s) based on the estimated value determined by the s and r… In the context of botnet traffic evasion, we apply DQN in a reinforcement learning framework, as shown in Figure 1. “, pg. 3, left column under “A. Deep Reinforcement Learning “; pg. 3, right column, under “B. Framework Structure”) (The agent is configured to evade detection from a deep learning detector. The detector corresponds to a monitoring system) wherein the sampling includes accessing the transition values stored in the electronic policy data structure to select an action by the reinforcement learning agent (“An optimal policy can be built from the optimal Q-function by choosing, for a given state, the action with highest Q-value (i.e. Q-learning). However, because the space in the Arcade Learning Environment (ALE) is too large to tractably store a tabular representation of the Q-function, the Deep Q-Network (DQN) which uses a deep function (e.g. CNN) approximator to represent the state-action value function was proposed [21].”, pg. 3, left and right columns) (The parameters of the DQN correspond to transition values which determine Q-values, which are then used to select an action.) analyze the steps taken in the episode to measure a strength of monitoring in the monitoring system; (“Through the score for each query, the attacker is able to directly measure the efficacy of any perturbation to the target black-box model.” Pg. 2, right column, under Score-based attack) (The more effective the attacker, the less effective the monitoring system is. For figures demonstrating this, see Wu’s TABLE III below) and present the strength of monitoring in an interface. (“…we use OpenAI gym as the environment interface, which is a toolkit for developing and comparing reinforcement learning algorithms. The gym framework provides a standardized environment to produce benchmarks”, pg. 3, right column, under “Fig. 1. Structure of the framework”) (Also see Wu’s TABLE III below for statistics involving Detection Model evasions, which teaches the strength of the monitoring system)
However, Wu does not explicitly teach to learn an electronic policy data structure that maps states of an electronic transaction system nor after training of the reinforcement learning agent is complete, sample the policy to simulate an episode of steps taken by the reinforcement learning agent, nor to write a step row into an episode data structure that records alert states of one or more scenarios resulting from executing the action;
Alizadeh teaches once the policy has been learned, sample the policy to simulate multiple episodes of steps taken by the reinforcement learning agent without repeating the training process, (“For this purpose, we define two periods, by which validation phase flag is activated, and the latest network weights are being recorded and exploited during validation. Depending on which period, the agent’s performance is evaluated for several episodes, and the achieved mean reward is compared to the latest maximum value. This process enables the agent to record the best trained model by validating on unseen scenarios.”, pg. 2, right column, second paragraph) and to write a step row into an episode data structure that records alert states of one or more scenarios resulting from executing the action; (“In Q-learning, a memory table Q[s,a] is built to store the Q-values for all the possible combinations of states and actions. By taking action on the current state, the reward Rand the new states are acquired to take the next action that has the maximum Q(s,a) in the memory table… However, if the combinations of state and actions are too large or states and actions are continuous, the memory and computation requirement for action-value function Q will be too high. To address this issue, Deep Q-Network (DQN) is utilized that approximates the action-value function Q(s,a)… Depending on which period, the agent’s performance is evaluated for several episodes, and the achieved mean reward is compared to the latest maximum value.”, pg. 2, left column, second to last paragraph; pg. 2, right column, above “III. SIMULATION SETUP”) (The Q-values, related to states, actions, and rewards in the memory table, correspond to step rows written into an episodic data structure.)
Therefore, it would have been considered obvious to one of ordinary skill in the art, prior to the current application’s filing date, to combine Wu’s Adversarial Reinforcement Learning Agent, with the policy sampling and episodic data structure of Alizadeh. One would be motivated to combine the teachings, prior to the filing date of the current application, as this allows for an optimized agent to be trained fast and be more generalized. (“This process enables the agent to record the best trained model by validating on unseen scenarios. Defining two various periods with a different number of episodes helps the training to be faster and record a more generalized model at the same time.”, pg. 2, right column, above “III. SIMULATION SETUP”)
However, Wu, as modified by Alizadeh, does not explicitly teach to learn an electronic policy data structure that maps states of an electronic transaction system.
Borrajo teaches an electronic transaction system electronic transaction system. (See Tables 1, 2 and 3) with states to learn a policy data structure (“The learning system takes as input traces of observable behavior. A trace 𝑡𝐶 is a sequence of states and actions executed by 𝐶 in those states: 𝑡𝐶 = (𝑠0, 𝑎1, 𝑠1, 𝑎2, 𝑠2, . . . , 𝑠𝑛−1, 𝑎𝑛, 𝑠𝑛), where 𝑠𝑖 is a state and 𝑎𝑖 is an action name and its parameters. States and actions correspond to the observable predicates and actions from the viewpoint of 𝐹 .”, pg. 3, right column, under “2.3 Traces of Behavior”) (While policies are not taught in this disclosure, the actions that are map to states of an electronic transaction system are taught, which in combination with the learning of a policy as disclosed in Wu, the limitation is fully taught.)
Therefore, it would have been considered obvious to one of ordinary skill in the art, prior to the current application’s filing date, to combine Wu’s Adversarial Reinforcement Learning Agent, modify it with Alizadeh, and combine it with the actions, predicates, and functions of Borrajo. One would be motivated to combine the teachings, prior to the filing date of the current application, as Borrajo’s method helps detect and classify good and bad behavior in a financial setting. (“The learning task can be defined as follows. Given: 𝑁 classes of behavior, ({good, bad} in our current application);3 and a set of labeled observed traces, 𝑇𝐶𝑖 , ∀𝐶𝑖 ∈ {g𝑜𝑜𝑑, 𝑏𝑎𝑑} Obtain: a classifier that takes as input a new (partial) trace 𝑡 (with unknown class) and outputs the predicted class… Good refers to standard customers’ behavior and bad corresponds to money laundering-related behavior.” Pg. 4, left column, under “3.1 Learning Task”)
PNG
media_image1.png
144
325
media_image1.png
Greyscale
Wu’s TABLE III
Regarding claims 2, 9, and 16, Wu, as modified by Alizadeh and Borrajo teaches selecting the action from a current probability distribution of available actions for a current state of the reinforcement learning agent (“A MDP [Markov Decision Process] is defined as a tuple (S, A, P, R, [Symbol font/0x67]) where: S is the state space of the process; A is a finite set of actions; P is a Markovian transition model, where P(s, a, s[Symbol font/0xA2]) is the probability of making a transition to state s[Symbol font/0xA2] when taking action a in state s; R is a reward (or cost) function, such that R(s, a) is the expected reward for taking action a in state s; [Symbol font/0x67] [Symbol font/0xCE] [0, 1) is the discount factor for future rewards… For each turn, the environment receives the action a chosen by the agent...”, pg. 3, left column, under “Deep Reinforcement Learning”) wherein the current probability distribution favors a subset of the available actions that do not trigger an alert under the one or more scenarios (“Reward: a reward value [between 0 and R], where 0 denotes that the botnet flow was detected by the detection model and R is the reward for evading the model.”, pg. 3, right column, bottom bullets) (The less the attacker is detected, the greater the reward, so the agent is rewarded greater for actions that don’t alert the monitoring system) executing the action to move the reinforcement learning agent into a new state (“For each turn, the environment receives the action a chosen by the agent…", pg. 3, left column, under Deep Reinforcement Learning) and evaluating the new state with the one or more scenarios to determine alert states of the one or more scenarios resulting from the action (“For each turn, the environment receives the action a chosen by the agent, and feeds back the observed state s[Symbol font/0xA2] (after executing a) and reward r. The agent follows a policy [Symbol font/0x70] (a|s) based on the estimated value determined by the s and r.” pg. 3, left column, under “Deep Reinforcement Learning”)
However, as currently combined, the combination does not explicitly teach appending a record of the action and the alert states to the episode as the step row.
Alizadeh further teaches appending a record of the action and the alert states to the episode as the step row. (“In Q-learning, a memory table Q[s,a] is built to store the Q-values for all the possible combinations of states and actions. By taking action on the current state, the reward Rand the new states are acquired to take the next action that has the maximum Q(s,a) in the memory table… However, if the combinations of state and actions are too large or states and actions are continuous, the memory and computation requirement for action-value function Q will be too high. To address this issue, Deep Q-Network (DQN) is utilized that approximates the action-value function Q(s,a)… We also appended the real-time validation phase to the original DQN algorithm to record the best-trained model during the training.”, pg. 2, left column, bottom paragraph; pg. 2 right column, above “III. SIMULATION SETUP”)
Therefore, it would have been considered obvious to one of ordinary skill in the art,
prior to the current application’s filing date, to combine the RL agent of Wu with the appending of a record as disclosed by Alizadeh. One would be motivated to combine the teachings, prior to the filing date of the current application, as this allows for the recording of the best-trained model during the training, as disclosed by Alizadeh. (“We also appended the real-time validation phase to the original DQN algorithm to record the best-trained model during the training” pg. 2, right column, above “III. SIMULATION SETUP”)
Regarding claim 3, 10, and 17, Wu, as modified by Alizadeh and Borrajo, teaches the method, system, and non-transitory computer-readable medium of claims 2, 9, and 16 respectively, as well as, repeating the selecting the action, the executing the action, the evaluating the new state and the appending the record until the task/transfer is complete. (“The agent follows a policy [Symbol font/0x70] (a|s) based on the estimated value determined by the s and r. The process stops when a target state is reached through a series of exploration and exploitation.” Pg. 5, left column, under “C. Relevant Parameters” (Wu)) (The combination teaches that the task/transfer is complete when the agent reaches a target state.)
Regarding claim 4, 11, and 18 Wu, as modified by Alizadeh and Borrajo, teaches the method, system, and non-transitory computer-readable medium of claims 1, 8, and 15 respectively, as well as configuring probability distributions of available actions for states of the reinforcement learning agent to favor actions that do not trigger an alert under the one or more scenarios. (“P(s, a, s[Symbol font/0xA2]) is the probability of making a transition to state s[Symbol font/0xA2] when taking action a in state s; R is a reward (or cost) function, such that R(s, a) is the expected reward for taking action a in state s… Reward: a reward value [Symbol font/0xCE] {0, R}, where 0 denotes that the botnet flow was detected by the detection model and R is the reward for evading the model. In our experiments, we use R=10.”, pg. 3, left column, under Deep Reinforcement Learning; pg. 3 right column, the first bullet (Wu)) (The less the attacker is detected, the greater the reward, so the agent is rewarded greater for actions that don’t alert the monitoring system)
Regarding claim 5, 12, and 19, Wu, as modified by Alizadeh and Borrajo, teaches the method, system, and non-transitory computer-readable medium of claims 1, 8, and 15 respectively, as well as the instructions to analyze the steps taken in the episode to measure the strength of monitoring in the monitoring system further cause the computer to determine a number of steps in the episode. (“An episode is a finite sequence of states, actions, and rewards… During an episode, states are transitioned and reward values are emitted…”, pg. 43, under “4.2 Environment Designs”, “In our experiments, we adopt Boltzmann exploration to select next action and allow the agent to perform up to ten actions before declaring failure.”, Pg. 5, left column, under “C. Relevant Parameters” (Wu)) (This teaches it, as Wu needs to determine the number of actions performed before declaring failure at ten actions, and episodes comprise actions (state transitions and reward values emitted) performed by the agent. Wu’s Fig. 1 shows this, (as seen below) as the chart does not show average actions above 9 actions.)
PNG
media_image2.png
223
257
media_image2.png
Greyscale
Wu’s Fig. 1
Claims 6, 13, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Wu, in view of Alizadeh, in further view of Borrajo, and in further view of Khalid El-Awady (Adaptive Stress Testing for Adversarial Learning in a Financial Environment)
Regarding claims 6, 13, and 20, Wu, as modified by Alizadeh and Borrajo, teaches the method, system, and non-transitory computer-readable medium of claims 1, 8, and 15 respectively but does not teach determining a number of accounts used for transfer in the episode.
Khalid El-Awady teaches determining a number of accounts used for transfer in the episode. (“for each episode… Initialize the state, s0: agent randomly selects a new or old payment card account to use, daily transactions are set to 0, and the fraud indicator set to 0.”, pg. 6, Algorithm 1 AST Q-learning, (See Khalid El-Awady’s Algorithm below))
Therefore, it would have been considered obvious to one of ordinary skill in the art, prior to the current application’s filing date, to combine Wu’s adversarial reinforcement learning agent, as modified by Alizadeh and Borrajo, with the accounts disclosed in Khalid El-Awady. One would be motivated to combine the teachings, prior to the filing date of the current application, as one could use a machine learning AI in a financial environment, for Adaptive Stress Testing, as disclosed by Khalid El-Awady. (“Adaptive Stress Testing for Adversarial Learning in a Financial Environment”, pg. 1, Title)
PNG
media_image3.png
296
627
media_image3.png
Greyscale
Khalid El-Awady’s Algorithm 1
Claims 7 and 14 is rejected under 35 U.S.C. 103 as being unpatentable over Wu, in view of Alizadeh, in further view of Borrajo, and in further view of Yan et al. (Herein referred to as Yan) (U.S. Patent Application No. US 20180365696 A1)
Regarding claims 7 and 14, Wu, as modified by Alizadeh and Borrajo teaches the method and system of claims 1 and 8 respectively but does not teach determining a percentage of amount transferred to a destination account before a cutoff by one of (i) generation of an alert or (ii) reaching a cap on episode length.
Yan teaches determining a percentage of amount transferred to a destination account before a cutoff by reaching a cap on episode length. ("…the suspicious percentage detector 222 can be used to detect, e.g., suspicious remittances, however other transactions such as, e.g., cash transfers withdrawals, deposits, among others are contemplated. Therefore, the suspicious percentage detector 222 can receive remittance histories for each account holder in the account holder cluster 211, including, e.g., remittance percentages. Here, a remittance percentage is used to signify the remittance amount divided by an account balance for a given user…”, pg. 4, Paragraph 42)
Therefore, it would have been considered obvious to one of ordinary skill in the art, prior to the current application’s filing date, to combine Wu’s adversarial reinforcement learning agent, as modified by Alizadeh and Borrajo, with the fraud detectors of Yan. One would be motivated to combine the teachings, prior to the filing date of the current application, as one could combine the agent with the fraud detector systems to mitigate fraud in transactions, as disclosed by Yan. (“According to an aspect of the present principles, a method is provided for mitigating fraud in transactions.”, pg. 1, Paragraph 4)
Response to Arguments
Applicant's arguments filed on May 12th, 2026 have been fully considered but they are not fully persuasive. The applicant argues in substance:
Argument 1: The 103 rejections are rendered moot by the amendments
Applicant’s arguments with respect to claim(s) 1, 8, and 15 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Argument 2: The limitation “sampling the policy to simulate an episode of steps taken by the reinforcement learning agent” is not taught by Wu or Mhaisen.
Upon further consideration, it is determined that Wu does teach “sampling the policy” as evidenced in Wu’s disclosure on pg. 3. (“the Deep Q-Network (DQN) which uses a deep function (e.g. CNN) approximator to represent the state-action value function was proposed… The agent gets a reward (benign or botnet) given by the target model and an estimate of the environment state which [is] represented by a feature vector of the sample. Through such information, the Q-function and action policy determine which action to select next.”, right column, first paragraph; under “B. Framework Structure”; See also Fig. 1 on pg. 3) However, Wu does not explicitly teach that sampling a policy is used “to simulate an episode of steps taken by the reinforcement learning agent”. Mhaisen does teach this, as evidenced by Algorithm 2 (steps 8-12) The rejection was made properly on the last action, but has been modified in this action in light of the new amendment.
Argument 3: The term “sampling” is given an overly broad interpretation. The term is given the broadest possible interpretation, rather than the broadest reasonable interpretation, as it is improperly interpreted to be learning done by an RL agent.
The examiner respectfully disagrees. Mhaisen teaches sampled states and actions, and a policy is a mapping between states and actions. The broadest reasonable interpretation of “sampling a policy” is interpreted to be noting the states and actions taken as a part of a strategy used by an artificial intelligence, which Mhaisen teaches in pg. 5 of their disclosure. The rejection was made properly on the last action, but has been modified in this action in light of the new amendment.
Argument 4: The rationale to combine Mhaisen and Wu is improper, as the motivation to combine references must lead to a combination of the references in the way claimed, and not some arbitrary manner.
The examiner respectfully disagrees. As explained in the above rejection, It would have been considered obvious to combine Wu’s agent and Mhaisen’s transition values and writing of a step row, with the motivation to combine the teachings being that it would allow for real-time learning to develop an optimal policy. In the KSR decision, it was deemed that “If a person of ordinary skill can implement a predictable variation, § 103 likely bars its patentability. For the same reason, if a technique has been used to improve one device, and a person of ordinary skill in the art would recognize that it would improve similar devices in the same way, using the technique is obvious unless its actual application is beyond his or her skill.” If the technique of writing a step row would improve that reinforcement learning agent of Mhaisen, one of ordinary skill in the art would recognize that it would improve the agent of Wu, and implementing the technique would be obvious. Furthermore, upon further consideration, it is found that Wu does teach sampling the policy and supports monitor-strength evaluation frameworks, as evidenced by Wu’s pg. 5. In combination with the specific writing of a row step in an episode data structure of Mhaisan, this leads to a combination of the references in the way claimed.
Argument 5: The combination of Wu, Mhaisen and Borrajo is improper, as the combination would materially change the principles of operation of the cited systems
The examiner respectfully disagrees. As stated in the above rejection, the combination of Wu, Mhaisen, and Borrajo would have been considered obvious without compromising the principles of operation of the systems. Specifically, combining the reinforcement learning agent, training and learning of a policy, and episodes of Wu, with the specific episode data structure including the writing of a step row of Mhaisen, and applying that to the field of electronic transaction systems similar to Borrajo; none of the modifications would change how the system of Wu would fundamentally operate.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Tyler E Iles whose telephone number is (571)272-5442. The examiner can normally be reached 9:00am - 5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kakali Chaki can be reached at (571) 272-3719. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/T.E.I./Patent Examiner, Art Unit 2122
/KAKALI CHAKI/Supervisory Patent Examiner, Art Unit 2122