Prosecution Insights
Last updated: October 04, 2026
Application No. 18/999,274

METHOD AND APPARATUS FOR GENERATING CYBERATTACK SEQUENCE BASED ON REINFORCEMENT LEARNING

Final Rejection §102§103§112
Filed
Dec 23, 2024
Priority
Feb 23, 2024 — RE 10-2024-0026345
Examiner
SIMITOSKI, MICHAEL J
Art Unit
2493
Tech Center
2400 — Computer Networks
Assignee
Electronics and Telecommunications Research Institute
OA Round
2 (Final)
80%
Grant Probability
Favorable
3-4
OA Rounds
1y 4m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 80% — above average
80%
Career Allowance Rate
630 granted / 785 resolved
+22.3% vs TC avg
Strong +28% interview lift
Without
With
+28.4%
Interview Lift
resolved cases with interview
Typical timeline
3y 2m
Avg Prosecution
20 currently pending
Career history
802
Total Applications
across all art units

Statute-Specific Performance

§101
10.7%
-29.3% vs TC avg
§103
45.2%
+5.2% vs TC avg
§102
13.9%
-26.1% vs TC avg
§112
21.1%
-18.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 785 resolved cases

Office Action

§102 §103 §112
CTNF 18/999,274 CTNF 79943 Notice of Pre-AIA or AIA Status 07-03-aia AIA 15-10-aia The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA. DETAILED ACTION The IDS filed 12/23/2024 was received and considered. Claims 1-20 are pending. Eligibility under 35 U.S.C. §101 Claim 1 recites: (1) generating a cyberattack simulation environment (data gathering); (2) training a cyberattack agent model based on the cyberattack simulation environment (involves a broad array of techniques and/or activities that may involve or rely upon mathematical concepts, but does not set forth or describe any mathematical relationships, calculations, formulas, or equations using words or mathematical symbols) 1 ; (3) generating an attack sequence using the trained cyberattack agent model (producing output, the method uses reinforcement learning to arrive at a valid attack sequence, specification ¶8). Claim 11 recites an apparatus comprising a processor and memory, which perform the same method as recited in claim 1. For both claims, step 1 of the 2019 PEG 2 is satisfied (claim is directed to a process, machine, manufacture). For step 2A, prong one, the claims are not directed to an abstract idea. While (1) is considered an abstract idea and (3) results from observing output from the trained model (specification ¶113), (2) does not set forth or describe any mathematical relationships, calculations, formulas, or equations using words or mathematical symbols and is thus not directed to an abstract idea, as training a machine learning model does not fall within the established grouping of abstract ideas (does not recite mathematical concepts, cannot be reasonably performed mentally, not a method or organizing human activity) 3 . Therefore, the claims are considered eligible under 35 U.S.C. §101. Claim Rejections - 35 USC § 112 07-30-02 AIA The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. 07-34-01 Claims 9 and 19 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Regarding claim 9, the limitation “the state” (line 3) renders the claim indefinite, as it is unclear if “the state” references “a state for the generated cyberattack sequence” (claim 8) or “a state of a host” (claim 9). Regarding claim 19, the limitation “the state” (line 3) renders the claim indefinite, as it is unclear if “the state” references “a state for the generated cyberattack sequence” (claim 18) or “a state of a host” (claim 19). Claim Rejections - 35 USC § 102 07-06 AIA 15-10-15 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. 07-07-aia AIA 07-07 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – 07-08-aia AIA (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. 07-15 AIA Claim s 1-3, 5-8, 10-13, 15-18 and 20 are rejected under 35 U.S.C. 102( a)(1 ) as being anticipated by CN 117521070 A to Lu et al. ( Lu ) . Regarding claim 1, Lu discloses a method for generating a cyberattack sequence based on reinforcement learning (environmental feedback of the previous attack action is obtained from the new network environment information, and this feedback is delivered to the decision module to make a new decision, ¶95), comprising: generating a cyberattack simulation environment (obtain network environment sample information; collect information in a small range to identify the current network composition, including host PCs, servers, routers, and security devices, ¶¶75-76); training a cyberattack agent model based on the cyberattack simulation environment (make decisions based on the sample information of each network environment and obtain the corresponding reward signal, ¶¶76-77; later utilize model to attack from current tactical stage based on network environment information and decision-making model, ¶87); and generating an attack sequence using the trained cyberattack agent model (launch simulated attack, ¶104; including generating an attack payload, ¶¶89-92). Regarding claim 11, the claim is similar in scope to claim 1 and is therefore rejected using a similar rationale (Lu discloses a memory in which at least one program is recorded; and a processor for executing the program, wherein the program includes instructions, ¶¶129-133). Regarding claims 2 and 12, Lu discloses wherein the simulation environment is configured with a network model (environment information, ¶122), an action space (adapted algorithm to different network conditions, ¶117; ATT&CK framework comprising technical phases and actions, ¶66), a state space (current tactical stage, ¶123; based on current network environment, ¶73), and a reward function (feedback regarding actions in the form of reward value, ¶70; model obtains the corresponding reward value based on the tactical stage reached by the agent and the attack action executed, ¶73). Regarding claims 3 and 13, Lu discloses wherein generating the cyberattack simulation environment comprises generating the cyberattack simulation environment by receiving a predefined simulation scenario configuration file (current network composition, ¶76), and the simulation scenario configuration file includes network configuration information (system accesses current network composition, including state of the network environment, ¶76), host asset configuration information (identifying host PCs, servers, etc., ¶76), and information about an attack technique (action space includes elements from the Attack Tactics Standard (ATT&CK) framework, ¶66; attack actions include specific technical actions, such as Collection, Privilege Escalation, Defense Evasion, Discovery, Credential Access, Lateral Movement, Execution, Persistence, Impact, Command and Control, Exfiltration, Impact, Preparation, and Initial Access, ¶66). Regarding claims 5 and 15, Lu discloses wherein the action space is configured with a pair of a host (identifying host PCs, servers, etc., ¶76; ) and an attack technique (attack actions include specific technical actions, ¶66), and the attack technique includes pre-attack state information of the host (overall state perceived by the system, including system components, ¶76) and state information of the host in an event of a successful attack (based on test results (attack), systems is configured to obtain hidden state, ¶96; S1051. When the test result reaches the preset result, obtain the hidden state of the corresponding decision model at the previous moment, ¶99). Regarding claims 6 and 16, Lu discloses wherein the state space is configured with a state of a host constituting a network (overall state perceived by the system, including system components, ¶76, used to determine optimal attack, ¶¶87-88) and a result of execution (¶¶89-95) of an attack technique (when the test result is executed and reaches the preset result, obtain the hidden state of the corresponding decision model at the previous moment, ¶99). Regarding claims 7 and 17, Lu discloses wherein the reward function is calculated based on a value of a compromised host depending on a change in the state space (maximum reward (reward function) is set for successfully gaining control of the target system (compromised host), so as to encourage the model to continuously try towards the final goal, ¶71). Regarding claims 8 and 18, Lu discloses wherein training the cyberattack agent model comprises generating a cyberattack sequence for the cyberattack simulation environment (set reward for each attack action and tactic, ¶71) and performing training using a reward for a state for the generated cyberattack sequence (the model obtains the corresponding reward value based on the tactical stage reached by the agent and the attack action executed, ¶73; policy network Actor updates its policy based on the attack action reward value fed back by the value network Critic, ¶73). Regarding claims 10 and 20, Lu discloses wherein the attack technique corresponds to an attack technique of a MITRE ATT&CK framework (action space includes elements from the Attack Tactics Standard (ATT&CK) framework, ¶66) . Claim Rejections - 35 USC § 103 07-06 AIA 15-10-15 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. 07-20-aia AIA The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 07-21-aia AIA Claim s 4 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Lu , as applied to claims 2 and 12, in view of CN 115766113 A to Chen et al. ( Chen ) and “Applying reinforcement learning for enhanced cybersecurity against adversarial simulation” by Oh et al. ( Oh ) . Regarding claims 4 and 14, Lu discloses wherein the network model is configured with a host (system uses scanning behavior to collect information in a small range to identify the current network composition, including host PCs, servers, routers, and security devices; inputs from different network sensors and processes are fused and prepared to be input into the deep neural network of the training thread model, ¶76), but lacks wherein the network model is configured with a subnetwork, topology and a firewall, and an allocation value used to calculate a reward value is defined in the host. However, Chen, in an analogous art (network penetration testing, ¶58), teaches that it was known to utilize a target network environment model described using configurations of a host, specifically including subnets, network topology, firewalls, routers, etc. (¶58). Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify Lu such that the network model is configured with a subnetwork, topology and a firewall. One of ordinary skill in the art would have been motivated to perform such a modification to describe the state of a given target host based on a position within the target network, as taught by Chen. As modified, Lu lacks where an allocation value used to calculate a reward value is defined in the host. However, Oh, in an analogous art (reinforcement learning for cyberattacks), teaches that it was known to provide a reward for a successful target compromise, where the reward is based on the importance of the target host (p. 5, §2.3, ¶1). Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to further modify Lu such that an allocation value used to calculate a reward value is defined in the host. One of ordinary skill in the art would have been motivated to perform such a modification to utilize a reward based on the value of a target, as taught by Oh . 07-21-aia AIA Claim s 9 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Lu , as applied to claims 8 and 18, in view of “Security analysis of cyber-physical systems using reinforcement learning” by Ibrahim et al. ( Ibrahim ) . Regarding claims 9 and 19, Lu lacks wherein training the cyberattack agent model comprises analyzing the generated cyberattack sequence and changing, when a state of a host satisfies pre-attack state information, the state to state information of the host in an event of a successful attack. However, Ibrahim, in an analogous art (reinforcement learning for attack agents), teaches that it was known to change, when a state of a host satisfies pre-attack state information (when attacker’s initial state is node 1, p. 9, #3), the state to state information of the host in an event of a successful attack (when attack is successful, state moves to further node, p. 10, initial state 1 is provided as input; path is forecasted taking into account the highest value of action “a” for the state “s”; see also p. 11, Algorithm 1). Ibrahim teaches the sequence used to predict the optimal route based on determining a current state and attacking vulnerabilities successfully to access the next node. Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify Lu to include analyzing the generated cyberattack sequence (potential actions) and changing, when a state of a host satisfies pre-attack state (current state of the network environment, per Lu, ¶76), the state to state information of the host in an event of a successful attack (consider host attacked and move to the next node). One of ordinary skill in the art would have been motivated to perform such a modification to determine the optimal attack path through multiple hosts, as taught by Ibrahim . Conclusion 07-96 AIA The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. WO 2025051493 A1 (BEARD ALFIE et al.) teaches using reinforcement learning for cybersecurity (abstract), including for training a cybersecurity red agent to attack a network (p. 10). CN 113660241 A (ZHENG, Chao et al.) teaches penetration testing using DQN (see machine translation). US 20250168186 A1 (MCDONALD; Geoffrey Lyall Pearson et al.) teaches commanding an attacker machine based on an AI model (abstract), including performing attack simulation (¶¶21-29). US 20220210200 A1 (Crabtree; Jason et al.) teaches identifying attack trees using attack simulation, using tools such as MITRE ATT&CK framework (¶153; ¶¶160-161) and using reinforcement learning (¶¶187-192). US 20240143737 A1 (Zamir; Amos et al.) teaches simulating a cyberattack against a target network and training an AI model to identify the attack (Fig. 5). “Employing deep reinforcement learning to cyber-attack simulation for enhancing cybersecurity” (Oh, Sang Ho, et al.) teaches employing an agent to interact with a realistic cyberattack scenario provided by MITRE ATT&CK (p. 4, §3) using reinforcement learning and rewards (p. 10, §4). US 20210064762 A1 (Salji; Carl Joseph) teaches graphing a virtualized network instance (¶23, Fig. 2) and simulating a cyberattack on the instance (¶¶47-57). US 20200045069 A1 (Nanda; Soumendra et al.) teaches a cyberattack course of action (CoA) simulator (abstract). “Catch me if you can: Improving adversaries in cyber-security with Q-Learning algorithms” (Bandhana, Arti, et al.) teaches training an attack agent using Q-learning and based on the Mitre ATT&CK framework. “Cygil: A cyber gym for training autonomous agents over emulated network systems” (Li, Li, Raed Fayad, and Adrian Taylor) teaches using reinforcement learning with Deep-Q networks to train attack agents using action spaces such as the Mitre framework and using reward functions. Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHAEL J SIMITOSKI whose telephone number is (571)272-3841. The examiner can normally be reached Monday - Friday, 7:00-3:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Carl Colin can be reached at 571-272-3862. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Michael Simitoski/ Primary Examiner, Art Unit 2493 April 22, 2026 Application/Control Number: 18/999,274 Page 2 Art Unit: 2493 Application/Control Number: 18/999,274 Page 3 Art Unit: 2493 Application/Control Number: 18/999,274 Page 4 Art Unit: 2493 Application/Control Number: 18/999,274 Page 5 Art Unit: 2493 Application/Control Number: 18/999,274 Page 6 Art Unit: 2493 Application/Control Number: 18/999,274 Page 7 Art Unit: 2493 Application/Control Number: 18/999,274 Page 8 Art Unit: 2493 Application/Control Number: 18/999,274 Page 9 Art Unit: 2493 Application/Control Number: 18/999,274 Page 10 Art Unit: 2493 Application/Control Number: 18/999,274 Page 11 Art Unit: 2493 1 “While it is common for claims to AI inventions to involve abstract ideas, USPTO personnel must draw a distinction between a claim that “recites” an abstract idea (and thus requires further eligibility analysis) and one that merely involves, or is based on, an abstract idea.” (2024 Guidance Update on Patent Subject Matter Eligibility, Including on Artificial Intelligence) 2 2019 Revised Patent Subject Matter Eligibility Guidance 3 USPTO published “Subject Matter Eligibility Examples: Abstract Ideas”, “Example 39 - Method for Training a Neural Network for Facial Detection”
Read full office action

Prosecution Timeline

Dec 23, 2024
Application Filed
May 06, 2026
Non-Final Rejection mailed — §102, §103, §112
Aug 06, 2026
Response Filed
Oct 01, 2026
Final Rejection mailed — §102, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12732808
VEHICLE, IN-VEHICLE DEVICE, AND MANAGEMENT METHOD
5y 0m to grant Granted Sep 08, 2026
Patent 12712747
SECURE CHANNEL INITIATION BETWEEN CARD AND HOST
2y 0m to grant Granted Aug 18, 2026
Patent 12695631
INTERIM ROOT-OF-TRUST ENROLMENT AND DEVICE-BOUND PUBLIC KEY REGISTRATION
2y 10m to grant Granted Jul 28, 2026
Patent 12695722
Network Traffic Control Method and Related System
2y 4m to grant Granted Jul 28, 2026
Patent 12689646
Malicious application detection
3y 7m to grant Granted Jul 21, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
80%
Grant Probability
99%
With Interview (+28.4%)
3y 2m (~1y 4m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 785 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month