Prosecution Insights
Last updated: October 02, 2026
Application No. 19/207,601

Method for Training a Reinforcement Learning Agent for an Industrial Process System and System for Training a Reinforcement Learning Agent for an Industrial Process System

Non-Final OA §101§102
Filed
May 14, 2025
Priority
May 15, 2024 — EU 24176069
Examiner
ANDERSON, FOLASHADE
Art Unit
Tech Center
Assignee
ABB Schweiz AG
OA Round
1 (Non-Final)
35%
Grant Probability
At Risk
1-2
OA Rounds
2y 10m
Est. Remaining
73%
With Interview

Examiner Intelligence

Grants only 35% of cases
35%
Career Allowance Rate
191 granted / 543 resolved
-24.8% vs TC avg
Strong +37% interview lift
Without
With
+37.4%
Interview Lift
resolved cases with interview
Typical timeline
4y 3m
Avg Prosecution
18 currently pending
Career history
572
Total Applications
across all art units

Statute-Specific Performance

§101
36.8%
-3.2% vs TC avg
§103
36.5%
-3.5% vs TC avg
§102
14.2%
-25.8% vs TC avg
§112
11.5%
-28.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 543 resolved cases

Office Action

§101 §102
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Status of Claims Claims 1-15 are pending and examined herein per Applicant’s May 14, 2025 filing with the USPTO. Information Disclosure Statement The information disclosure statement (IDS) submitted on 05/14/2025 and 07/20/2026 were in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Drawings The drawings are objected to as failing to comply with 37 CFR 1.84(p)(5) because they include the following reference character(s) not mentioned in the description: 235 of Fig. 2 and Fig. 3. Corrected drawing sheets in compliance with 37 CFR 1.121(d), or amendment to the specification to add the reference character(s) in the description in compliance with 37 CFR 1.121(b) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Note to Applicant Claims 1-15 were examine under Bilski/Alice 35 U.S.C. 101 the claimed invention was found not to be directed to an abstract idea, when considered in light of USPTO examples 39 and 47; when the “reinforcement learning agent” is understood in light of the specification to solely occur within the industrial process system. With one exception, claims 9-15 fail to comply with step 1 of the Bilski/Alice analysis. Claims 9-15 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. The claim(s) does/do not fall within at least one of the four categories of patent eligible subject matter because the claims are to a device without a tangible form (hardware components). A device is interpreted as falling under the statutory class of a machine. Where a machine is a concrete thing, consisting of parts, or of certain devices and combination of devices. MPEP 2106.03(I). The courts' definitions of machines, manufactures and compositions of matter indicate a product must have a physical or tangible form in order to fall within one of these statutory categories. MPEP 2106.03(I). As currently claimed the claims are found to be software per se. Where products that do not have a physical or tangible form, such as information (often referred to as "data per se") or a computer program per se (often referred to as "software per se") when claimed as a product without any structural recitations. MPEP 2106.03(I). It is further noted that system only comprises a “data storage medium” and “low-fidelity simulator” Where the specification provides support for the claimed element in “data storage medium 210 is provided, the data storage medium is configured for storing plant historical data” (Instant Spec [23]) and “low-fidelity simulator may be constructed using plant historical data” (Instant Spec [24]) The Subject Matter Eligibility of Computer Readable Media memorandum (01/26/2010) states “The broadest reasonable interpretation of a claim drawn to a computer readable medium . . . typically covers forms of non-transitory tangible media and transitory propagating signal per se . . . particularly when the specification is silent”, Also see MPEP 2106.03(I). It is suggested that Applicant change the statutory class to a medium and add “non-transitory” to the claimed limitation to overcome the rejection; because the specification is silent on the components of for example the claimed “industrial process system”; thus it is also interrupted as unembodied software. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claim(s) 1-15 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Yu et al (US 2023/0419005 A1) Claim 1 Yu teaches a method for training a reinforcement learning (RL) agent for an industrial process system, (Yu [201] “method can include training an agent for motor directional drilling using deep reinforcement learning (DRL)”) comprising: training the RL agent with plant historical data of the industrial process system (Yu [127] “field management tool 520 may be configured with functionalities to store oilfield data (e.g., historical data” and [201] “training an agent for motor directional drilling using deep reinforcement learning (DRL)”); and retraining the RL agent using plant historical data and a low-fidelity simulator of the industrial process system (Yu [212] “increase fidelity, a method can include increasing one or more types of uncertainty (e.g., noise, etc.) during training and/or retraining of an agent or agents.”), the retraining the RL agent comprising: analyzing the plant historical data to identify white spots as regions of process states and dynamic behavior that have not been explored during the training the RL agent (Yu [237] “ agent with known dynamics . . . whose dynamics are unknown” and [367] “known/knowable or unknown/unknowable, various conditions may be estimated using one or more types of simulators”); and retraining the RL agent by prioritized exploration with information gained from the white spots and with simulated data provided by simulating the industrial process system with the low-fidelity simulator (Yu [412] “ if an agent demonstrates low fidelity when forward modeled with uncertainty, the agent may be subjected to retraining optionally with a different noise regime (e.g., more noise, different type of noise, etc.)”). Claim 2 Yu teaches all the limitations of the method of claim 1, wherein the analyzing the plant historical data to identify white spots comprises retrieving at least one bound selected from the group consisting of lower bounds for process state variables or upper bounds for process state variables (Yu [449] where the limitation is made in the alternative only one element needs to be present in the art); and identifying the white spots by variable space exploration using at least one of the bounds selected from the group consisting of the lower bounds or upper bounds (Yu [237] and [367] where the limitation is made in the alternative only one element needs to be present in the art). Claim 3 Yu teaches all the limitations of the method of claim 1, further comprising inferring dynamics of safety-related variables from at least one of plant historical data or by the prioritized exploration (Yu [210] and [236]); and leveraging the dynamics of the safety-related variables to construct a safety verifier configured to predict safety variables based on values of manipulated variables (Yu [236] and [240]). Claim 4 Yu teaches all the limitations of the method of claim 3, further comprising comparing predicted safety variables to pre-determined safety constraints (Yu [238-240]); and adjusting values of the manipulated variables to ensure compliance of the safety variables with the safety constraints (Yu [243]). Claim 5 Yu teaches all the limitations of the method of claim 3, further comprising manipulating the industrial process system to a predefined safe state by a safety guarantor when the safety verifier fails due to at least one incident selected from the group consisting of insufficient learning or non-compliance of the safety variable with the safety constraints (Yu [214]). Claim 6 Yu teaches all the limitations of the method of claim 1, further comprising fine tuning the RL agent by: using a high-fidelity simulator; (Yu [127-128] or [202], where the limitation is made in the alternative only one element needs to be present in the art) interacting the RL agent with the industrial process system; (Yu [127-128] or [202], where the limitation is made in the alternative only one element needs to be present in the art) using plant historical data; (Yu [127-128] or [202], where the limitation is made in the alternative only one element needs to be present in the art) or a combination thereof. (Yu [127-128] or [202], where the limitation is made in the alternative only one element needs to be present in the art) Claim 7 Yu teaches all the limitations of the method of claim 1, further comprising deploying the RL agent to the industrial process system (Yu [227]). Claim 8 Yu teaches all the limitations of the method of claim 7, further comprising fine tuning an RL agent policy by iteratively performing the steps of (Yu [99], [221], and [315]): monitoring a performance and behavior of the RL agent by collecting historical data and rewards (Yu [381] and [400]); analyzing the performance and behavior of the RL agent (Yu [228]); fine tuning the RL agent policy by adjusting policy parameters and exploring new actions to get an updated RL agent policy (Yu [99], [221], and [315]); subjecting the updated RL agent policy to offline validation using the collected historical data to evaluate its impact on the performance because of policy alterations (Yu [238]); and upon determining that the updated RL agent policy pass the validation, systematically rolling out the updated RL agent policy to the industrial process (Yu [107]). Claim 9 Yu a system for training a reinforcement learning (RL) agent for an industrial process system, (Yu [217]) comprising: a data storage medium configured for storing plant historical data of the industrial process system (Yu [121]); and a low-fidelity simulator of the industrial process system (Yu [379] and [412]). Claim 10 Yu teaches all the limitations of the system of claim 9, further comprising a safety verifier configured to predict safety variables based on values of manipulated variables (Yu [210]). Claim 11 Yu teaches all the limitations of the system of claim 9, further comprising a high-fidelity simulator of the industrial process system (Yu [419]). Claim 12 Yu teaches all the limitations of the system of claim 9, further comprising a training module configured to train the RL agent according to a method comprising (Yu [201]): training the RL agent with plant historical data of the industrial process system (Yu [127] and [201]); and retraining the RL agent using plant historical data and a low-fidelity simulator of the industrial process system (Yu [212]), the retraining the RL agent comprising: analyzing the plant historical data to identify white spots as regions of process states and dynamic behavior that have not been explored during the training the RL agent (Yu [237] and [367]); and retraining the RL agent by prioritized exploration with information gained from the white spots and with simulated data provided by simulating the industrial process system with the low-fidelity simulator (Yu [412]). Claim 13 Yu teaches all the limitations of the system of claim 10, further comprising a safety guarantor configured to manipulate the industrial process system to a predefined safe state in the event of a failure of the safety verifier (Yu [236] and [240]). Claim 14 Yu teaches all the limitations of the system of claim 13, further comprising a high-fidelity simulator of the industrial process system (Yu [419]). Claim 15 Yu teaches all the limitations of the system of claim 14, further comprising a training module configured to train an RL agent according to according to a method comprising (Yu [201]): training the RL agent with plant historical data of the industrial process system (Yu [127]); and retraining the RL agent using plant historical data and a low-fidelity simulator of the industrial process system, the retraining the RL agent comprising (Yu [212]): analyzing the plant historical data to identify white spots as regions of process states and dynamic behavior that have not been explored during the training the RL agent (Yu [237] and [367]); and retraining the RL agent by prioritized exploration with information gained from the white spots and with simulated data provided by simulating the industrial process system with the low-fidelity simulator (Yu [412]). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Traut et al (US 2021/0357692 A1) teaches training procedures for adjusting trainable parameters including reinforcement machine learning (e.g., deep Q learning based on feedback) may be trained independently of other components (e.g., offline training on historical data). Crabtree et al (US 2026/0147639 A1) teaches high-criticality scenarios identified by adaptive elastic funnel engine 230, the system allocates increased representational capacity by simultaneously increasing bond dimensions χ.sub.j in the relevant regions of the tensor network, deepening logical circuits in differentiable logic evaluation structure 310, and allocating additional computational resources through computational resource orchestrator 510. This coordinated precision management extends across all processing domains, creating a unified approach to resource allocation based on scenario importance. The dynamic precision mechanisms utilize real-time criticality signals, computational resource availability monitored by computational resource orchestrator 510, and feedback on decision confidence from decision engine 320. This enables the system to operate efficiently under varying computational constraints while maintaining high fidelity in critical scenario regions. Solowjow et al (US 2024/0296662 A1) teaches rendered images and the corresponding labels (e.g., generated as described above) may then be used to retrain the deep learning model. This closed loop approach can thus adjust the training data distribution to make sure that no blind spots exist in the dataset. Any inquiry concerning this communication or earlier communications from the examiner should be directed to FOLASHADE ANDERSON whose telephone number is (571)270-3331. The examiner can normally be reached Monday to Thursday 12:00 P.M. to 6:00 P.M. CST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Rutao Wu can be reached at (571) 272-6045. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /FOLASHADE ANDERSON/Primary Examiner, Art Unit 3623
Read full office action

Prosecution Timeline

May 14, 2025
Application Filed
Sep 08, 2026
Non-Final Rejection mailed — §101, §102 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743666
METHOD AND SYSTEM OF TRAINING OF CHAINED NEURAL NETWORKS FOR DELAY PREDICTION IN TRANSIT NETWORKS
2y 11m to grant Granted Sep 22, 2026
Patent 12718158
PREDICTING DOWNSTREAM SCHEDULE EFFECTS OF USER TASK ASSIGNMENTS
3y 0m to grant Granted Aug 25, 2026
Patent 12711449
Tagging Performance Evaluation and Improvement
5y 6m to grant Granted Aug 18, 2026
Patent 12682357
FRAMEWORK FOR CYBER-PHYSICAL INTERACTION AWARE TEST CASE GENERATION TO IDENTIFY OPERATIONAL CHANGES
3y 1m to grant Granted Jul 14, 2026
Patent 12646020
APPARATUS FOR CLASSROOM SCHEDULING AND METHOD OF USE
2y 8m to grant Granted Jun 02, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
35%
Grant Probability
73%
With Interview (+37.4%)
4y 3m (~2y 10m remaining)
Median Time to Grant
Low
PTA Risk
Based on 543 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month