Prosecution Insights
Last updated: October 04, 2026
Application No. 19/069,386

METHOD AND APPARATUS FOR OPTIMIZING SCHEDULING USING REINFORCEMENT LEARNING

Final Rejection §101§103
Filed
Mar 04, 2025
Priority
Jan 18, 2024 — RE 10-2024-0008196 +2 more
Examiner
BROWN, SARA GRACE
Art Unit
3625
Tech Center
3600 — Transportation & Electronic Commerce
Assignee
LG Management Development Institute Co. Ltd.
OA Round
2 (Final)
29%
Grant Probability
At Risk
3-4
OA Rounds
1y 10m
Est. Remaining
62%
With Interview

Examiner Intelligence

Grants only 29% of cases
29%
Career Allowance Rate
47 granted / 161 resolved
-22.8% vs TC avg
Strong +33% interview lift
Without
With
+33.2%
Interview Lift
resolved cases with interview
Typical timeline
3y 5m
Avg Prosecution
24 currently pending
Career history
196
Total Applications
across all art units

Statute-Specific Performance

§101
35.0%
-5.0% vs TC avg
§103
40.4%
+0.4% vs TC avg
§102
9.5%
-30.5% vs TC avg
§112
14.0%
-26.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 161 resolved cases

Office Action

§101 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Regarding the 35 USC 112(b) rejection, Examiner has fully considered Applicant’s arguments and amendments. Examiner has deemed Applicant’s amendments sufficient to overcome the 35 USC 112(b) rejection. Accordingly, the 35 USC 112(b) rejection is withdrawn. Regarding the 35 USC 101 rejection, Examiner has fully considered Applicant’s arguments and amendments. Regarding Applicant’s assertion of “Thus, the amended claims are not directed merely to generic reinforcement learning implemented on a generic computer, but instead recite a specific technical implementation for operation scheduling of a naphtha cracking center integrated into a practical industrial application.,” Examiner respectfully disagrees. The use of reinforcement learning, as drafted, is nothing more than mere use of a computer as a tool to generate the schedule information. The generation of an improved schedule, as drafted, would be an improvement to the abstract idea identified in the independent claims and not to the reinforcement learning model itself. MPEP 2106.05(a): “It is important to note, the judicial exception alone cannot provide the improvement. The improvement can be provided by one or more additional elements...” Additionally, as discussed in 2106.05(a)(II) improvements to technology or technical fields, “an improvement in the abstract idea itself … is not an improvement in technology” Regarding Applicant’s assertion of “Furthermore, the amended claims, which recite generating scheduling information for operation of a naphtha cracking center using multiple agents configured to be learned using the same reward in reinforcement learning, impose meaningful limitations on any alleged judicial exception and amount to significantly more than merely applying an abstract idea using generic computer components.,” Examiner respectfully disagrees. The present claims, as drafted, do not improve the functioning of the reinforcement learning model itself. The use of a “same reward” is nothing more than mere detailing the external evaluative feedback signal used to update the model, which is not an improvement to the functioning of the model itself. Rather, this merely improves the model’s accuracy with respect to the output information, namely, the scheduling information. The generation of an improved schedule, as drafted, would be an improvement to the abstract idea identified in the independent claims and not to the reinforcement learning model itself. MPEP 2106.05(a): “It is important to note, the judicial exception alone cannot provide the improvement. The improvement can be provided by one or more additional elements...” Additionally, as discussed in 2106.05(a)(II) improvements to technology or technical fields, “an improvement in the abstract idea itself … is not an improvement in technology” Accordingly, the present claims are rejected under 35 USC 101. Regarding the 35 USC 102 rejection, Examiner has fully considered Applicant’s arguments and amendments. Regarding Applicant’s assertion of “Thus, these amendments of claims 1, 14, and 18 encompass the features disclosed in the exemplary embodiments disclosed at paragraph [0057] of this application, but are nowhere disclosed in the references relied upon in the Office Action.,” Examiner has withdrawn the 35 USC 102 rejection; however, the present claims are rejected under 35 USC 103. Applicant’s arguments with respect to the previous prior art reference Copperthite of the record have been considered but are moot because the new grounds of rejection does not rely on any reference applied in the prior art rejection for any teachings or matter specifically challenged in the argument. The claims are rejected under a new grounds of rejection, which was necessitated by amendment. Examiner has introduced the Wu reference to cure the deficiencies of the prior art combination of the record. See the detailed rejection below. Accordingly, the present claims are rejected under 35 USC 103. Claim Objections Claim objected to because of the following informalities: Examiner suggests amending the claim to correct the minor antecedence issue of: “wherein the first agent, the second agent, and the third agent are configured to be learned using [[the]]a same reward in reinforcement learning.” Appropriate correction is required. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-3, 5-10, 12-14, 16, and 18 are rejected under 35 USC 101 because the claimed invention is directed to a judicial exception (i.e. abstract idea) without anything significantly more. Step 1: Claims 1-3, 5-10, and 12-13 are directed to a method, claims 14 and 16 are directed to a system, and claim 18 is directed to a non-transitory computer readable medium. Therefore, the claims are directed to patent eligible categories of invention. Step 2A, Prong 1: Claims 1, 14, and 18 recite determining operation information, constituting an abstract idea based on “Mental Processes” related to concepts performed in the human mind including observation, evaluation, judgment, and opinion. Claim 1 recites limitations, similarly recited in claims 14 and 18, including “obtaining input information; determining, incoming tank information based on the input information, determining, mixing tank combination information; determining, cracking furnace operation information; generating, one or more scheduling information for the naphtha cracking center based on the incoming tank information determined using the first agent, the mixing tank combination information determined using the second agent, and the cracking furnace operation information determined using the third agent.” These limitations, as drafted, but for the recitation of “by the at least one processor,” is a process that covers performance of the limitations in the mind but for the recitation of generic computer components. That is, but for the “by the at least one processor” language, nothing in the claim elements preclude the steps from practically being performed in the human mind. For example, with the exception of the “by the at least one processor” language, the claim steps in the context of the claim encompass a user mentally or manually performing the steps of the claim. Dependent claims 3, 6-9, and 12-13 further narrow the abstract idea identified in the independent claims and do not introduce further additional elements for consideration. Dependent claims 2 and 5 will be evaluated under Step 2A, Prong 2 below. Step 2A, Prong 2: Claims 1, 14, and 18 do not integrate the judicial exception into a practical application. Claim 1 is directed to a method performed “by the at least one processor.” Claim 14 is directed to a system comprising “at least one processor; and at least one memory having stored therein computer-readable instruction configured to cause the at least one processor to perform a method for scheduling a naphtha cracking center by at least one processor.” Claim 18 is directed to a “non-transitory computer-readable storage medium having computer-executable instructions stored thereon, which when executed by at least one processor, cause the at least one processor to perform a method for scheduling a naphtha cracking center by at least one processor.” Claims 1, 14, and 18 further recite performing the determinations of the claims using a “first agent,” “second agent,” and “third agent.” Claims 1, 14, and 18 further recite “wherein the first agent is a first artificial intelligence device configured to be learned by reinforcement learning,” “wherein the second agent is a second artificial intelligence device configured to be learned by reinforcement learning,” “wherein the third agent is a third artificial intelligence device configured to be learned by reinforcement learning,” and “wherein the first agent, the second agent, and the third agent are configured to be learned using the same reward in reinforcement learning.” These limitations of the agent being “learned by reinforcement learning” provide nothing more than mere instructions to implement an abstract idea on a generic computer. See MPEP 2106.05(f). MPEP 2106.05(f) provides the following considerations for determining whether a claim simply recites a judicial exception with the words “apply it” (or an equivalent), such as mere instructions to implement an abstract idea on a computer: (1) whether the claim recites only the idea of a solution or outcome i.e., the claim fails to recite details of how a solution to a problem is accomplished; (2) whether the claim invokes computers or other machinery merely as a tool to perform an existing process; and (3) the particularity or generality of the application of the judicial exception. These additional elements are mere instructions to implement an abstract idea using a computer in its ordinary capacity, or merely uses the computer as a tool to perform the identified abstract idea. Use of a computer or other machinery in its ordinary capacity for tasks (e.g., to receive, store, or transmit data) or simply adding a general purpose computer or computer components after the fact to an abstract idea (e.g., mental processes) does not integrate a judicial exception into a practical application. See MPEP 2106.05(f). Therefore, the additional elements of the independent claims, when considered both individually and in combination, are not sufficient to prove integration into a practical application. Dependent claims 3, 6-9, and 12-13 further narrow the abstract idea identified in the independent claims and do not introduce further additional elements for consideration, which does not integrate the judicial exception into a practical application. Dependent claim 2 introduces the additional element of “wherein: the first agent, the second agent, and the third agent are asynchronous multi-agents.” Dependent claims 10 and 16 introduce the additional element of “wherein: at least one of the first agent, the second agent, and the third agent comprises a plurality of agents.” These limitations related to reinforcement learning provide nothing more than mere instructions to implement an abstract idea on a generic computer. See MPEP 2106.05(f). MPEP 2106.05(f) provides the following considerations for determining whether a claim simply recites a judicial exception with the words “apply it” (or an equivalent), such as mere instructions to implement an abstract idea on a computer: (1) whether the claim recites only the idea of a solution or outcome i.e., the claim fails to recite details of how a solution to a problem is accomplished; (2) whether the claim invokes computers or other machinery merely as a tool to perform an existing process; and (3) the particularity or generality of the application of the judicial exception. These additional elements are mere instructions to implement an abstract idea using a computer in its ordinary capacity, or merely uses the computer as a tool to perform the identified abstract idea. Use of a computer or other machinery in its ordinary capacity for tasks (e.g., to receive, store, or transmit data) or simply adding a general purpose computer or computer components after the fact to an abstract idea (e.g., mental processes) does not integrate a judicial exception into a practical application. See MPEP 2106.05(f). Dependent claim 5 introduces the additional element of “wherein: the input information is obtained through a first user interface (UI), and the scheduling information is provided to a user through a second UI.” Use of a computer or other machinery in its ordinary capacity for tasks (e.g., to receive, store, or transmit data) or simply adding a general purpose computer or computer components after the fact to an abstract idea (e.g., mental processes) does not integrate a judicial exception into a practical application. See MPEP 2106.05(f). Therefore, the additional elements of the dependent claims, when considered both individually and in the context of the independent claims above, are not sufficient to prove integration into a practical application. Step 2B: Claims 1, 14, and 18 do not comprise anything significantly more than the judicial exception. Claim 1 is directed to a method performed “by the at least one processor.” Claim 14 is directed to a system comprising “at least one processor; and at least one memory having stored therein computer-readable instruction configured to cause the at least one processor to perform a method for scheduling a naphtha cracking center by at least one processor.” Claim 18 is directed to a “non-transitory computer-readable storage medium having computer-executable instructions stored thereon, which when executed by at least one processor, cause the at least one processor to perform a method for scheduling a naphtha cracking center by at least one processor.” Claims 1, 14, and 18 further recite performing the determinations of the claims using a “first agent,” “second agent,” and “third agent.” Claims 1, 14, and 18 further recite “wherein the first agent is a first artificial intelligence device configured to be learned by reinforcement learning,” “wherein the second agent is a second artificial intelligence device configured to be learned by reinforcement learning,” “wherein the third agent is a third artificial intelligence device configured to be learned by reinforcement learning,” and “wherein the first agent, the second agent, and the third agent are configured to be learned using the same reward in reinforcement learning.” These limitations of the agent being “learned by reinforcement learning” provide nothing more than mere instructions to implement an abstract idea on a generic computer. See MPEP 2106.05(f). MPEP 2106.05(f) provides the following considerations for determining whether a claim simply recites a judicial exception with the words “apply it” (or an equivalent), such as mere instructions to implement an abstract idea on a computer: (1) whether the claim recites only the idea of a solution or outcome i.e., the claim fails to recite details of how a solution to a problem is accomplished; (2) whether the claim invokes computers or other machinery merely as a tool to perform an existing process; and (3) the particularity or generality of the application of the judicial exception. These additional elements are mere instructions to implement an abstract idea using a computer in its ordinary capacity, or merely uses the computer as a tool to perform the identified abstract idea. Use of a computer or other machinery in its ordinary capacity for tasks (e.g., to receive, store, or transmit data) or simply adding a general purpose computer or computer components after the fact to an abstract idea (e.g., mental processes) is not anything significantly more than the judicial exception. See MPEP 2106.05(f). Therefore, the additional elements of the independent claims, when considered both individually and in combination, are not anything significantly more than the judicial exception. Dependent claims 3, 6-9, and 12-13 further narrow the abstract idea identified in the independent claims and do not introduce further additional elements for consideration, which is not anything significantly more than the judicial exception. Dependent claim 2 introduces the additional element of “wherein: the first agent, the second agent, and the third agent are asynchronous multi-agents.” Dependent claims 10 and 16 introduce the additional element of “wherein: at least one of the first agent, the second agent, and the third agent comprises a plurality of agents.” These limitations related to reinforcement learning provide nothing more than mere instructions to implement an abstract idea on a generic computer. See MPEP 2106.05(f). MPEP 2106.05(f) provides the following considerations for determining whether a claim simply recites a judicial exception with the words “apply it” (or an equivalent), such as mere instructions to implement an abstract idea on a computer: (1) whether the claim recites only the idea of a solution or outcome i.e., the claim fails to recite details of how a solution to a problem is accomplished; (2) whether the claim invokes computers or other machinery merely as a tool to perform an existing process; and (3) the particularity or generality of the application of the judicial exception. These additional elements are mere instructions to implement an abstract idea using a computer in its ordinary capacity, or merely uses the computer as a tool to perform the identified abstract idea. Use of a computer or other machinery in its ordinary capacity for tasks (e.g., to receive, store, or transmit data) or simply adding a general purpose computer or computer components after the fact to an abstract idea (e.g., mental processes) is not anything significantly more than the judicial exception. See MPEP 2106.05(f). Dependent claim 5 introduces the additional element of “wherein: the input information is obtained through a first user interface (UI), and the scheduling information is provided to a user through a second UI.” Use of a computer or other machinery in its ordinary capacity for tasks (e.g., to receive, store, or transmit data) or simply adding a general purpose computer or computer components after the fact to an abstract idea (e.g., mental processes) is not anything significantly more than the judicial exception. See MPEP 2106.05(f). Therefore, the additional elements of the dependent claims, when considered both individually and in the context of the independent claims above, are not anything significantly more than the judicial exception. Accordingly, claims 1-3, 5-10, 12-14, 16, and 18 are rejected under 35 USC 101. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claim(s) 1, 3, 5-9, 14, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Copperthite et al. (US 20240201639 A1) in view of Wu et al. (Hierarchical Hybrid Multi-Agent Deep Reinforcement Learning for Peer-to-Peer Energy Trading Among Multiple Heterogeneous Microgrids, November 2023). Regarding claim 1, Copperthite teaches a method for scheduling a naphtha cracking center by at least one processor ([0258-0260] teach a computer system including the software components of an IoT platform), comprising the steps of: obtaining input information ([0263] teaches the edge devices may monitor operations of a particular processing unit and may be connected to the network in order to send and receive information, wherein each edge device includes one or more controllers for monitoring and selectively controlling a respective edge device, wherein [0195] teaches the sensor input data includes particular data representing results of subsequent processing or derivation, wherein the sensor input data is associated with operation of one or more assets, e.g. processing units, that can be performed via the machine learning models, wherein [0268] teaches the IoT platform is a model-driven architecture that communicates with each layer to contextualize site data of the enterprise using an extensible object model, or asset model, and knowledge graphs where the equipment, e.g. edge devices, and processes of the enterprise are modeled, wherein [0270] teaches the models describe the assets, or nodes, of the enterprise at the edge devices and describe the relationship of the assets with other components or links, wherein the models are self-validating, wherein the models describe the types of sensors mounted on any given asset or edge device and the type of data being sensed by each sensor, wherein the IoT platform is extensible, model-driven end-to-end stack including two-way model sync and secure data exchange between the edge and the cloud, wherein [0319] teaches the industrial facility includes a number of individual processing units that may each embody an industrial asset that performs a particular function during operation of the industrial facility, wherein the oil refinery could include processing units including a storage tank, catalytic cracking unit, and blender, wherein [0094] teaches an asset refers to a machine, such as a storage tank or product tank, furnace, and blender, as well as in [0172] teaches the industrial control system includes a product tank, rundown blender or batch blender, and catalytic combustor; see also: [0033, 0276]); determining, by the at least one processor, incoming tank information using a first agent based on the input information ([0263] teaches the edge devices may monitor operations of a particular processing unit and may be connected to the network in order to send and receive information, wherein each edge device includes one or more controllers for monitoring and selectively controlling a respective edge device, wherein [0195] teaches the sensor input data includes particular data representing results of subsequent processing or derivation, wherein the sensor input data is associated with operation of one or more assets, e.g. processing units, that can be performed via the machine learning models, wherein [0268] teaches the IoT platform is a model-driven architecture that communicates with each layer to contextualize site data of the enterprise using an extensible object model, or asset model, and knowledge graphs where the equipment, e.g. edge devices, and processes of the enterprise are modeled, wherein [0270] teaches the models describe the assets, or nodes, of the enterprise at the edge devices and describe the relationship of the assets with other components or links, wherein the models are self-validating, wherein the models describe the types of sensors mounted on any given asset or edge device and the type of data being sensed by each sensor, wherein the IoT platform is extensible, model-driven end-to-end stack including two-way model sync and secure data exchange between the edge and the cloud, wherein [0276] teaches employing a variety of different process models that are linked with assets such that when an asset, or edge device, instance is created, any associated calculation instances and parameters are linked to the appropriate attributes of the asset and corresponding edge device, wherein [0319] teaches the industrial facility includes a number of individual processing units that may each embody an industrial asset that performs a particular function during operation of the industrial facility, wherein the oil refinery could include processing units including a storage tank, catalytic cracking unit, and blender, wherein [0094] teaches an asset refers to a machine, such as a storage tank or product tank, as well as in [0172] teaches the industrial control system includes a product tank; see also: [0033, 0160, 0219, 0271, 0374]), wherein the first agent is a first artificial intelligence device configured to be learned by reinforcement learning ([0263] teaches the edge devices may monitor operations of a particular processing unit and may be connected to the network in order to send and receive information, wherein each edge device includes one or more controllers for monitoring and selectively controlling a respective edge device, wherein Fig. 12 and [0266] teach the IoT platform is a platform for plantwide optimization that uses real-time accurate models and real-time data for real-time control for sustained peak performance of the enterprise, wherein the IoT platform is deployed in a cloud environment that supports end-to-end capability to execute emission optimization models for assets, wherein [0327] teaches the portions of the emission optimization pathway data can be provided to the plurality of production models associated with the industrial processes and industrial assets, wherein the blending information can be provided to the production model to optimize the site-wide operations and industrial processes, wherein [0268] teaches the IoT platform is a model-driven architecture that communicates with each layer to contextualize site data of the enterprise using an extensible object model, or asset model, and knowledge graphs where the equipment, e.g. edge devices, and processes of the enterprise are modeled, wherein [0276] teaches employing a variety of different process models that are linked with assets such that when an asset, or edge device, instance is created, any associated calculation instances and parameters are linked to the appropriate attributes of the asset and corresponding edge device, wherein [0270] teaches the models describe the assets, or nodes, of the enterprise at the edge devices and describe the relationship of the assets with other components or links, wherein the models are self-validating, wherein the models describe the types of sensors mounted on any given asset or edge device and the type of data being sensed by each sensor, wherein the IoT platform is extensible, model-driven end-to-end stack including two-way model sync and secure data exchange between the edge and the cloud, wherein [0163] teaches the optimization system includes a number of machine learning techniques to optimize the optimization model including reinforcement learning algorithms, as well as in [0219] teaches the optimization model is a reinforcement learning model, as well as in [0374] teaches the multi-optimization model can be a machine learning model, such as a reinforcement learning algorithm, wherein [0319] teaches the industrial facility includes a number of individual processing units that may each embody an industrial asset that performs a particular function during operation of the industrial facility, wherein the oil refinery could include processing units including a storage tank, catalytic cracking unit, and blender; see also: [0033, 0160, 0172, 0271, 0319]); determining, by the at least one processor, mixing tank combination information using a second agent ([0263] teaches the edge devices may monitor operations of a particular processing unit and may be connected to the network in order to send and receive information, wherein each edge device includes one or more controllers for monitoring and selectively controlling a respective edge device, wherein [0195] teaches the sensor input data includes particular data representing results of subsequent processing or derivation, wherein the sensor input data is associated with operation of one or more assets, e.g. processing units, that can be performed via the machine learning models, wherein [0268] teaches the IoT platform is a model-driven architecture that communicates with each layer to contextualize site data of the enterprise using an extensible object model, or asset model, and knowledge graphs where the equipment, e.g. edge devices, and processes of the enterprise are modeled, wherein [0270] teaches the models describe the assets, or nodes, of the enterprise at the edge devices and describe the relationship of the assets with other components or links, wherein the models are self-validating, wherein the models describe the types of sensors mounted on any given asset or edge device and the type of data being sensed by each sensor, wherein the IoT platform is extensible, model-driven end-to-end stack including two-way model sync and secure data exchange between the edge and the cloud, wherein [0276] teaches employing a variety of different process models that are linked with assets such that when an asset, or edge device, instance is created, any associated calculation instances and parameters are linked to the appropriate attributes of the asset and corresponding edge device, wherein [0319] teaches the industrial facility includes a number of individual processing units that may each embody an industrial asset that performs a particular function during operation of the industrial facility, wherein the oil refinery could include processing units including a blender, as well as in [0172] teaches the industrial control system includes rundown blenders and batch blenders, wherein [0094] teaches an asset refers to a machine, such as a blender; see also: [0033, 0160, 0219, 0271, 0374]), wherein the second agent is a second artificial intelligence device configured to be learned by reinforcement learning ([0263] teaches the edge devices may monitor operations of a particular processing unit and may be connected to the network in order to send and receive information, wherein each edge device includes one or more controllers for monitoring and selectively controlling a respective edge device, wherein Fig. 12 and [0266] teach the IoT platform is a platform for plantwide optimization that uses real-time accurate models and real-time data for real-time control for sustained peak performance of the enterprise, wherein the IoT platform is deployed in a cloud environment that supports end-to-end capability to execute emission optimization models for assets, wherein [0327] teaches the portions of the emission optimization pathway data can be provided to the plurality of production models associated with the industrial processes and industrial assets, wherein the blending information can be provided to the production model to optimize the site-wide operations and industrial processes, wherein [0268] teaches the IoT platform is a model-driven architecture that communicates with each layer to contextualize site data of the enterprise using an extensible object model, or asset model, and knowledge graphs where the equipment, e.g. edge devices, and processes of the enterprise are modeled, wherein [0276] teaches employing a variety of different process models that are linked with assets such that when an asset, or edge device, instance is created, any associated calculation instances and parameters are linked to the appropriate attributes of the asset and corresponding edge device, wherein [0270] teaches the models describe the assets, or nodes, of the enterprise at the edge devices and describe the relationship of the assets with other components or links, wherein the models are self-validating, wherein the models describe the types of sensors mounted on any given asset or edge device and the type of data being sensed by each sensor, wherein the IoT platform is extensible, model-driven end-to-end stack including two-way model sync and secure data exchange between the edge and the cloud, wherein [0163] teaches the optimization system includes a number of machine learning techniques to optimize the optimization model including reinforcement learning algorithms, as well as in [0219] teaches the optimization model is a reinforcement learning model, as well as in [0374] teaches the multi-optimization model can be a machine learning model, such as a reinforcement learning algorithm, wherein [0319] teaches the industrial facility includes a number of individual processing units that may each embody an industrial asset that performs a particular function during operation of the industrial facility, wherein the oil refinery could include processing units including a storage tank, catalytic cracking unit, and blender; see also: [0033, 0160, 0172, 0271, 0319]); determining, by the at least one processor, cracking furnace operation information using a third agent ([0263] teaches the edge devices may monitor operations of a particular processing unit and may be connected to the network in order to send and receive information, wherein each edge device includes one or more controllers for monitoring and selectively controlling a respective edge device, wherein [0195] teaches the sensor input data includes particular data representing results of subsequent processing or derivation, wherein the sensor input data is associated with operation of one or more assets, e.g. processing units, that can be performed via the machine learning models, wherein [0268] teaches the IoT platform is a model-driven architecture that communicates with each layer to contextualize site data of the enterprise using an extensible object model, or asset model, and knowledge graphs where the equipment, e.g. edge devices, and processes of the enterprise are modeled, wherein [0270] teaches the models describe the assets, or nodes, of the enterprise at the edge devices and describe the relationship of the assets with other components or links, wherein the models are self-validating, wherein the models describe the types of sensors mounted on any given asset or edge device and the type of data being sensed by each sensor, wherein the IoT platform is extensible, model-driven end-to-end stack including two-way model sync and secure data exchange between the edge and the cloud, wherein [0276] teaches employing a variety of different process models that are linked with assets such that when an asset, or edge device, instance is created, any associated calculation instances and parameters are linked to the appropriate attributes of the asset and corresponding edge device, wherein [0319] teaches the industrial facility includes a number of individual processing units that may each embody an industrial asset that performs a particular function during operation of the industrial facility, wherein the oil refinery could include processing units including a catalytic cracking unit, wherein [0094] teaches an asset refers to a machine, such as a furnace, as well as in [0172] teaches the industrial control system includes a catalytic combustor; see also: [0033, 0160, 0219, 0271, 0374]), wherein the third agent is a third artificial intelligence device configured to be learned by reinforcement learning ([0263] teaches the edge devices may monitor operations of a particular processing unit and may be connected to the network in order to send and receive information, wherein each edge device includes one or more controllers for monitoring and selectively controlling a respective edge device, wherein Fig. 12 and [0266] teach the IoT platform is a platform for plantwide optimization that uses real-time accurate models and real-time data for real-time control for sustained peak performance of the enterprise, wherein the IoT platform is deployed in a cloud environment that supports end-to-end capability to execute emission optimization models for assets, wherein [0327] teaches the portions of the emission optimization pathway data can be provided to the plurality of production models associated with the industrial processes and industrial assets, wherein the blending information can be provided to the production model to optimize the site-wide operations and industrial processes, wherein [0268] teaches the IoT platform is a model-driven architecture that communicates with each layer to contextualize site data of the enterprise using an extensible object model, or asset model, and knowledge graphs where the equipment, e.g. edge devices, and processes of the enterprise are modeled, wherein [0276] teaches employing a variety of different process models that are linked with assets such that when an asset, or edge device, instance is created, any associated calculation instances and parameters are linked to the appropriate attributes of the asset and corresponding edge device, wherein [0270] teaches the models describe the assets, or nodes, of the enterprise at the edge devices and describe the relationship of the assets with other components or links, wherein the models are self-validating, wherein the models describe the types of sensors mounted on any given asset or edge device and the type of data being sensed by each sensor, wherein the IoT platform is extensible, model-driven end-to-end stack including two-way model sync and secure data exchange between the edge and the cloud, wherein [0163] teaches the optimization system includes a number of machine learning techniques to optimize the optimization model including reinforcement learning algorithms, as well as in [0219] teaches the optimization model is a reinforcement learning model, as well as in [0374] teaches the multi-optimization model can be a machine learning model, such as a reinforcement learning algorithm, wherein [0319] teaches the industrial facility includes a number of individual processing units that may each embody an industrial asset that performs a particular function during operation of the industrial facility, wherein the oil refinery could include processing units including a storage tank, catalytic cracking unit, and blender; see also: [0033, 0160, 0172, 0271, 0319]); and generating, by the at least one processor, one or more scheduling information for the naphtha cracking center based on the incoming tank information determined using the first agent (Fig. 7 and [0236] teaches an optimization path task schedule interface can be provided wherein information can be displayed and inputs can be obtained, wherein [0199] teaches the optimization pathway generator receives input constraints from the user through a user interface, as well as in [0194] teaches the system can receive sensor input data through the user interface, wherein [0028] teaches generating a schedule for modifying at least one asset of a plurality of assets, wherein [0107] teaches the optimization pathway comprises a schedule of actions and proposed tasks, wherein [0356] teaches receiving information related to operating a plurality of assets of the processing plant and outputting a schedule for each transformation action within the optimized set of transformation actions, wherein [0316] teaches the industrial facility receives and/or processes ingredients as inputs to create a final product, such as a hydrocarbon processing plant, wherein [0319] teaches the industrial facility includes a number of individual processing units that may each embody an industrial asset that performs a particular function during operation of the industrial facility, wherein the oil refinery could include processing units including a storage tank, catalytic cracking unit, and blender; see also: [0102, 0315, 0330, 0366-0370]), the mixing tank combination information determined using the second agent (Fig. 7 and [0236] teaches an optimization path task schedule interface can be provided wherein information can be displayed and inputs can be obtained, wherein [0199] teaches the optimization pathway generator receives input constraints from the user through a user interface, as well as in [0194] teaches the system can receive sensor input data through the user interface, wherein [0028] teaches generating a schedule for modifying at least one asset of a plurality of assets, wherein [0107] teaches the optimization pathway comprises a schedule of actions and proposed tasks, wherein [0356] teaches receiving information related to operating a plurality of assets of the processing plant and outputting a schedule for each transformation action within the optimized set of transformation actions, wherein [0316] teaches the industrial facility receives and/or processes ingredients as inputs to create a final product, such as a hydrocarbon processing plant, wherein [0319] teaches the industrial facility includes a number of individual processing units that may each embody an industrial asset that performs a particular function during operation of the industrial facility, wherein the oil refinery could include processing units including a storage tank, catalytic cracking unit, and blender; see also: [0102, 0315, 0330, 0366-0370]), and the cracking furnace operation information determined using the third agent (Fig. 7 and [0236] teaches an optimization path task schedule interface can be provided wherein information can be displayed and inputs can be obtained, wherein [0199] teaches the optimization pathway generator receives input constraints from the user through a user interface, as well as in [0194] teaches the system can receive sensor input data through the user interface, wherein [0028] teaches generating a schedule for modifying at least one asset of a plurality of assets, wherein [0107] teaches the optimization pathway comprises a schedule of actions and proposed tasks, wherein [0356] teaches receiving information related to operating a plurality of assets of the processing plant and outputting a schedule for each transformation action within the optimized set of transformation actions, wherein [0316] teaches the industrial facility receives and/or processes ingredients as inputs to create a final product, such as a hydrocarbon processing plant, wherein [0319] teaches the industrial facility includes a number of individual processing units that may each embody an industrial asset that performs a particular function during operation of the industrial facility, wherein the oil refinery could include processing units including a storage tank, catalytic cracking unit, and blender; see also: [0102, 0315, 0330, 0366-0370]). However, Copperthite does not explicitly teach wherein the first agent, the second agent, and the third agent are configured to be learned using the same reward in reinforcement learning. From the same or similar field of endeavor, Wu teaches wherein the first agent, the second agent, and the third agent are configured to be learned using the same reward in reinforcement learning (Pg. 4652 teaches proposing a hierarchical hybrid multi-agent double deep Q network to capture the diversity advantage of resources in order to perform hierarchical scheduling, wherein the scheduling can be performed for industrial furnaces and other industrial resources, wherein the multi-agent network solves scheduling by splitting the workload into sub-tasks that share the same reward, as well as in Pgs. 4655-4656 teach formulating a scheduling problem based on a multi-agent algorithm, wherein the two controllers are cooperating to obtain an optimal scheduling policy in order to maximize a common reward, wherein Pg. 4657-4658 teach that the multi-agent problem has a same reward that is utilized in order for both of the sub-controllers to learn their optimal cooperative policies in order to maximize their common reward; see also: Pgs. 4650). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to modify Copperthite to incorporate the teachings of Wu to include wherein the first agent, the second agent, and the third agent are configured to be learned using the same reward in reinforcement learning. One would have been motivated to do so in order to split the optimal policy learning workload into two hierarchical sub-tasks that share the same reward, thus reducing operational complexity while maintaining operational flexibility (Wu, Pg. 4652). By incorporating the teachings of Wu, one would have been able to use the same reward to calculate the Q-learning value, thus motivating both sub-controllers to learn the optimal cooperative policies in order to maximize their common reward (Wu, Pg. 4658). Regarding claims 14 and 18, the claims recite limitations already addressed by the rejection of claim 1. Regarding claim 14, Copperthite teaches a system comprising (Fig. 2): at least one processor; and at least one memory having stored therein computer-readable instruction configured to cause the at least one processor to perform a method for scheduling a naphtha cracking center by at least one processor, comprising the steps of (Fig. 2 and [0180-0186] teach a computing system comprising a processor configured to control one or more functions through program instructions stored on a memory accessible to the processor). Regarding claim 18, Copperthite teaches a non-transitory computer-readable storage medium having computer-executable instructions stored thereon, which when executed by at least one processor, cause the at least one processor to perform a method for scheduling a naphtha cracking center by at least one processor ([0018] teaches a non-transitory computer readable medium having computer program code stored thereon that, in execution with at least one processor, configures the computer program product for performing the steps of the methods; see also: [0178-0180]), comprising the steps of. Accordingly, claims 14 and 18 are rejected as being unpatentable over Copperthite in view of Wu. Regarding claim 3, the combination of Copperthite and Wu teaches all the limitations of claim 1 above. Copperthite further teaches wherein: the input information comprises at least one of constraints, mixing tank operation information, cracking furnace operation plan information ([0263] teaches the edge devices may monitor operations of a particular processing unit and may be connected to the network in order to send and receive information, wherein each edge device includes one or more controllers for monitoring and selectively controlling a respective edge device, wherein [0245] teaches the optimization pathway generator may generate an updated optimization pathway based on observed values, wherein the observed value is a measured value captured by one or more sensing devices positioned to capture physical characteristics of a monitored asset of the plurality of assets, wherein the generator may utilize measurements obtained from sensing devices measuring data at the target assets in order to determine if the asset is operating at or near the level projected, wherein [0292] teaches the edge devices are associated with the industrial assets including furnaces and other equipment, wherein [0319] teaches the industrial facility includes a number of individual processing units that may each embody an industrial asset that performs a particular function during operation of the industrial facility, wherein the oil refinery could include processing units including a storage tank, catalytic cracking unit, and blender, wherein [0094] teaches an asset refers to a machine, such as a storage tank or product tank, furnace, and blender, as well as in [0172] teaches the industrial control system includes a product tank, rundown blender or batch blender, and catalytic combustor; see also: [0195, 0253]) Regarding claim 5, the combination of Copperthite and Wu teaches all the limitations of claim 1 above. Copperthite further teaches wherein: the input information is obtained through a first user interface (UI) ([0199] teaches the optimization pathway generator receives input constraints from the user through a user interface, as well as in [0194] teaches the system can receive sensor input data through the user interface, wherein [0028] teaches generating a schedule for modifying at least one asset of a plurality of assets, wherein [0107] teaches the optimization pathway comprises a schedule of actions and proposed tasks, wherein [0356] teaches receiving information related to operating a plurality of assets of the processing plant and outputting a schedule for each transformation action within the optimized set of transformation actions, wherein [0316] teaches the industrial facility receives and/or processes ingredients as inputs to create a final product, such as a hydrocarbon processing plant, wherein [0319] teaches the industrial facility includes a number of individual processing units that may each embody an industrial asset that performs a particular function during operation of the industrial facility, wherein the oil refinery could include processing units including a storage tank, catalytic cracking unit, and blender; see also: [0102, 0315, 0330, 0366-0370]), and the scheduling information is provided to a user through a second UI (Fig. 7 and [0236] teaches an optimization path task schedule interface can be provided wherein information can be displayed and inputs can be obtained, [0028] teaches generating a schedule for modifying at least one asset of a plurality of assets, wherein [0107] teaches the optimization pathway comprises a schedule of actions and proposed tasks, wherein [0356] teaches receiving information related to operating a plurality of assets of the processing plant and outputting a schedule for each transformation action within the optimized set of transformation actions, wherein [0316] teaches the industrial facility receives and/or processes ingredients as inputs to create a final product, such as a hydrocarbon processing plant, wherein [0319] teaches the industrial facility includes a number of individual processing units that may each embody an industrial asset that performs a particular function during operation of the industrial facility, wherein the oil refinery could include processing units including a storage tank, catalytic cracking unit, and blender; see also: [0102, 0315, 0330, 0366-0370]). Regarding claim 6, the combination of Copperthite and Wu teaches all the limitations of claim 1 above. Copperthite further teaches wherein: the scheduling information comprises at least one of incoming scheduling information (Fig. 7 and [0236] teaches an optimization path task schedule interface can be provided wherein information can be displayed and inputs can be obtained, wherein [0199] teaches the optimization pathway generator receives input constraints from the user through a user interface, as well as in [0194] teaches the system can receive sensor input data through the user interface, wherein [0028] teaches generating a schedule for modifying at least one asset of a plurality of assets, wherein [0107] teaches the optimization pathway comprises a schedule of actions and proposed tasks, wherein [0356] teaches receiving information related to operating a plurality of assets of the processing plant and outputting a schedule for each transformation action within the optimized set of transformation actions, wherein [0316] teaches the industrial facility receives and/or processes ingredients as inputs to create a final product, such as a hydrocarbon processing plant, wherein [0319] teaches the industrial facility includes a number of individual processing units that may each embody an industrial asset that performs a particular function during operation of the industrial facility, wherein the oil refinery could include processing units including a storage tank, catalytic cracking unit, and blender; see also: [0102, 0315, 0330, 0366-0370]), mixing scheduling information (Fig. 7 and [0236] teaches an optimization path task schedule interface can be provided wherein information can be displayed and inputs can be obtained, wherein [0199] teaches the optimization pathway generator receives input constraints from the user through a user interface, as well as in [0194] teaches the system can receive sensor input data through the user interface, wherein [0028] teaches generating a schedule for modifying at least one asset of a plurality of assets, wherein [0107] teaches the optimization pathway comprises a schedule of actions and proposed tasks, wherein [0356] teaches receiving information related to operating a plurality of assets of the processing plant and outputting a schedule for each transformation action within the optimized set of transformation actions, wherein [0316] teaches the industrial facility receives and/or processes ingredients as inputs to create a final product, such as a hydrocarbon processing plant, wherein [0319] teaches the industrial facility includes a number of individual processing units that may each embody an industrial asset that performs a particular function during operation of the industrial facility, wherein the oil refinery could include processing units including a storage tank, catalytic cracking unit, and blender; see also: [0102, 0315, 0330, 0366-0370]), cracking furnace scheduling information (Fig. 7 and [0236] teaches an optimization path task schedule interface can be provided wherein information can be displayed and inputs can be obtained, wherein [0199] teaches the optimization pathway generator receives input constraints from the user through a user interface, as well as in [0194] teaches the system can receive sensor input data through the user interface, wherein [0028] teaches generating a schedule for modifying at least one asset of a plurality of assets, wherein [0107] teaches the optimization pathway comprises a schedule of actions and proposed tasks, wherein [0356] teaches receiving information related to operating a plurality of assets of the processing plant and outputting a schedule for each transformation action within the optimized set of transformation actions, wherein [0316] teaches the industrial facility receives and/or processes ingredients as inputs to create a final product, such as a hydrocarbon processing plant, wherein [0319] teaches the industrial facility includes a number of individual processing units that may each embody an industrial asset that performs a particular function during operation of the industrial facility, wherein the oil refinery could include processing units including a storage tank, catalytic cracking unit, and blender; see also: [0102, 0315, 0330, 0366-0370]). Regarding claim 7, the combination of Copperthite and Wu teaches all the limitations of claim 1 above. Copperthite further teaches wherein: the incoming tank information comprises at least one of: naphtha incoming schedule information for each of the at least one incoming tank ([0102] teaches the optimization model is configured to process various inputs including a schedule of transformation actions, wherein [0308] teaches generating a user-interactive electronic interface that renders a visual representation of data associated with the modifications to the industrial processes, wherein the user can adjust a set-point or a schedule for the one or more industrial processes, as well as in [0298] teaches the optimization request is received in response to a user-initiated action initiated via the user interface of a computing device, wherein the request is received in response to a schedule for the one or more industrial processes satisfying a defined criterion, such as a schedule interval for the one or more industrial processes being above at threshold timer, wherein [0319] teaches the industrial facility includes a number of individual processing units that may each embody an industrial asset that performs a particular function during operation of the industrial facility, wherein the oil refinery could include processing units including a storage tank, catalytic cracking unit, and blender, wherein [0094] teaches an asset refers to a machine, such as a storage tank or product tank, as well as in [0172] teaches the industrial control system includes a product tank; see also: [0033, 0160, 0219, 0271, 0374]). Regarding claim 8, the combination of Copperthite and Wu teaches all the limitations of claim 1 above. Copperthite further teaches wherein: the mixing tank combination information comprises determining at least one of: information on the mixing schedule with the mixing tank for each of the at least one incoming tanks ([0102] teaches the optimization model is configured to process various inputs including a schedule of transformation actions, wherein [0308] teaches generating a user-interactive electronic interface that renders a visual representation of data associated with the modifications to the industrial processes, wherein the user can adjust a set-point or a schedule for the one or more industrial processes, as well as in [0298] teaches the optimization request is received in response to a user-initiated action initiated via the user interface of a computing device, wherein the request is received in response to a schedule for the one or more industrial processes satisfying a defined criterion, such as a schedule interval for the one or more industrial processes being above at threshold timer, wherein [0226] teaches the optimization model can be configured with pre-optimized inputs including outputting a production model with fuels blending to optimize site-wide operations, wherein mitigation actions can be performed, wherein [0319] teaches the industrial facility includes a number of individual processing units that may each embody an industrial asset that performs a particular function during operation of the industrial facility, wherein the oil refinery could include processing units including a storage tank, catalytic cracking unit, and blender, wherein [0094] teaches an asset refers to a machine, such as a storage tank or product tank, as well as in [0172] teaches the industrial control system includes a product tank and rundown blenders and batch blender, and wherein [0327] teaches the fuel blending information can include an optimization pathway to optimize the industrial process; see also: [0033, 0160, 0219, 0271, 0374]). Regarding claim 9, the combination of Copperthite and Wu teaches all the limitations of claim 1 above. Copperthite further teaches wherein: the cracking furnace operation information comprises at least one of: input rate ([0319] teaches the industrial facility includes a number of individual processing units that may each embody an industrial asset that performs a particular function during operation of the industrial facility, wherein the oil refinery could include processing units including a catalytic cracking unit, wherein [0094] teaches an asset refers to a machine, such as a furnace, as well as in [0172] teaches the industrial control system includes a catalytic combustor, wherein [0128] teaches gathering sensor data input indicating the physical condition of an asset including flow rate and other operating parameters; see also: [0178]), and cracking furnace operation schedule information ([0102] teaches the optimization model is configured to process various inputs including a schedule of transformation actions, wherein [0308] teaches generating a user-interactive electronic interface that renders a visual representation of data associated with the modifications to the industrial processes, wherein the user can adjust a set-point or a schedule for the one or more industrial processes, as well as in [0298] teaches the optimization request is received in response to a user-initiated action initiated via the user interface of a computing device, wherein the request is received in response to a schedule for the one or more industrial processes satisfying a defined criterion, such as a schedule interval for the one or more industrial processes being above at threshold timer, wherein [0319] teaches the industrial facility includes a number of individual processing units that may each embody an industrial asset that performs a particular function during operation of the industrial facility, wherein the oil refinery could include processing units including a catalytic cracking unit, wherein [0094] teaches an asset refers to a machine, such as a furnace, as well as in [0172] teaches the industrial control system includes a catalytic combustor; see also: [0033, 0160, 0219, 0271, 0374]). Claim(s) 2 is rejected under 35 U.S.C. 103 as being unpatentable over Copperthite et al. (US 20240201639 A1) in view of Wu et al. (Hierarchical Hybrid Multi-Agent Deep Reinforcement Learning for Peer-to-Peer Energy Trading Among Multiple Heterogeneous Microgrids, November 2023) in view of Bahrpeyma et al. (“A review of the applications of multi-agent reinforcement learning in smart factories,” December 01, 2022). Regarding claim 2, the combination of Copperthite and Wu teaches all the limitations of claim 1 above. However, Copperthite does not explicitly teach wherein: the first agent, the second agent, and the third agent are asynchronous multi-agents. From the same or similar field of endeavor, Bahrpeyma teaches wherein: the first agent, the second agent, and the third agent are asynchronous multi-agents (Pgs. 7-8 teach a global network that updates parameters based on the aggregate gradient of the exploring agents, and the exploring agents copy weights asynchronously from the global network, wherein the agents are deployed in separate and parallel environments, each exploring a different part of the problem space, so that they cannot affect each other, therefore cooperation between agents in the MARL setting is established through the global network, wherein Pg. 11 teaches the multi-agent problem can utilize reinforcement learning; see also: Pg. 11). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to modify the combination of Copperthite and Wu to incorporate the teachings of Bahrpeyma to include wherein: the first agent, the second agent, and the third agent are asynchronous multi-agents. One would have been motivated to do so in order to realize efficiency, agility, and automation all at once in the dynamic nature of manufacturing environments (Bahrpeyma, Pg. 1). By incorporating the teachings of Bahrpeyma, one would have been motivated to do so in order to allow agents to be deployed in separate and parallel environments, thus allowing them to cooperate under a multi-agent reinforcement learning setting through the global network (Bahrpeyma, Pgs. 7-8). Claim(s) 10 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Copperthite et al. (US 20240201639 A1) in view of Wu et al. (Hierarchical Hybrid Multi-Agent Deep Reinforcement Learning for Peer-to-Peer Energy Trading Among Multiple Heterogeneous Microgrids, November 2023) in view of Hinton (Hinton, David Corder. “A Hybrid Framework for Critical Infrastructures Interdependency Modeling, Simulation, and Analysis.” Missouri University of Science and Technology, 2023.). Regarding claims 10 and 16, the combination of Copperthite and Wu teaches all the limitations of claims 1 and 14 above. However, Copperthite does not explicitly teach wherein: at least one of the first agent, the second agent, and the third agent comprises a plurality of agents From the same or similar field of endeavor, Hinton teaches wherein: at least one of the first agent, the second agent, comprises a plurality of agents (Pgs. 43-44 teach a framework of software tools including a plurality of agents, wherein the agents may recursively encapsulate sub-agents, which behave similarly within the boundaries of their scope, wherein the agents can invoke updates on attached components with compartmentalized responsibility and scope for sub-agent containerization, wherein Pg. 5 teaches the framework can be applied to many systems including Pgs. 11-12 teach petroleum flow processing including naphtha; see also: 45-53). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to modify the combination of Copperthite and Wu to incorporate the teachings of Hinton to include wherein: at least one of the first agent, the second agent, and the third agent comprises a plurality of agents. One would have been motivated to do so in order to support functionally autonomous, run-time agents that recursively encapsulate sub-agents that behave similarly within the boundaries of their scope (Hinton, Pg. 43). By incorporating the teachings of Hinton, one would have been able to allow containerization of sub-agents that enables portability across frameworks (Hinton, Pg. 45). Claim(s) 12 is rejected under 35 U.S.C. 103 as being unpatentable over Copperthite et al. (US 20240201639 A1) in view of Wu et al. (Hierarchical Hybrid Multi-Agent Deep Reinforcement Learning for Peer-to-Peer Energy Trading Among Multiple Heterogeneous Microgrids, November 2023) in view of Lee et al. (“Data science and reinforcement learning for price forecasting and raw material procurement in petrochemical industry,” 2022). Regarding claim 12, the combination of Copperthite and Wu teaches all the limitations of claim 1 above. However, Copperthite does not explicitly teach wherein: the reward in the reinforcement learning is determined based on total earnings, facility operation costs, naphtha purchase costs, and costs associated with constraints. From the same or similar field of endeavor, Lee teaches wherein: the reward in the reinforcement learning is determined based on total earnings, facility operation costs, naphtha purchase costs, and costs associated with constraints (Pg. 1 teaches a two stage data science framework to predict the weekly price of butadiene and optimizing the procurement decision, wherein the first stage suggests several price prediction models with comprehensive information including contract price, supply rate, demand rate, and upstream and downstream information, wherein the second stage applies reinforcement learning to derive an optimal policy of procurement decision and reduce the total procurement cost, wherein Pg. 2 teaches the price forecasting is based on the capacity operating rate, which is the rate at which the factory utilizes its capacity, and downstream substitutes in the supply chain, wherein the BD price is based on the ethylene capacity operating rate from naphtha crackers, the demand of downstream products, and crude oil price, wherein Pgs. 5-6 teach collecting data including historical price, supply and demand, upstream/downstream material price and capacity, and other industrial data, wherein the forecasted price and historical data are utilized to build the reward function, wherein the reward can be described explicitly, wherein the objective of the MDP is to maximize the expected reward, wherein the framework aims to minimize the expected total procurement cost of the BD, wherein Pg. 8 teaches the price and supply of upstream materials significantly affect the BD price, wherein an increase of marginal profit increases the supply of upstream materials related to BD, and thus increasing demand of BD leads to the decrease of BD price, wherein there is a variance inflating factor that provides a metric that measures the variance of the estimated regression coefficient, wherein Pg. 10 teaches calculating the total procurement cost and average inventory and considering the procurement cost, which is based on the shipping of the product; see also: Pgs. 12-13). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to modify the combination of Copperthite and Wu to incorporate the teachings of Lee to include wherein: the reward in the reinforcement learning is determined based on total earnings, facility operation costs, naphtha purchase costs, and costs associated with constraints. One would have been able to improve the accuracy of price forecasts and reduce the procurement costs associated with raw materials through the use of reinforcement learning (Lee, Pg. 1). By incorporating the teachings of Lee, one would have been able to reduce the material cost in order to significantly improve profitability (Lee, Pgs. 12-13). Allowable Subject Matter Claim 13 overcome the prior art of record such that none of the cited prior art references can be applied to form the basis of a 35 USC 102 rejection nor can they be combined to fairly suggest in combination, the basis of a 35 USC 103 rejection when the limitations are read in the particular environment of the claims. However, Examiner notes that the claims remain rejected under 35 USC 101, as set forth above. With respect to claim 13, the prior art of the record does not teach or disclose: wherein: the reward in the reinforcement learning is determined by the following: R e w a r d = P r o f i t -   ∑ c ∈ C o n s t r a i n t s w c ∙ C o s t c , P r o f i t = R e v e n u e - E n e r g y   u s a g e - N a p h t h a   c o s t wherein the profit is determined based on subtracting facility operation costs and naphtha purchase costs from the total earnings, wc is the weight per each constraint and Costc is the cost per each constraint. The closest prior art of the record discloses: Lee et al. (“Data science and reinforcement learning for price forecasting and raw material procurement in petrochemical industry,” 2022) discloses assigning relative importance to a number of elements to extract weights based on the decision maker’s preference structure, wherein the weights correspond to each element, wherein this analytical hierarchy can be used to assess the rewards of the reinforcement learning model. However, Lee fails to explicitly teach or disclose the limitations above. Hubbs et al. (US 20220027817 A1) discloses the reinforcement learning agent can generate a reward, wherein the reward, or value function, can indicate a profit or other favorable economic value. Hubbs further discloses generating the reward, revenue, and inventory costs. However, Hubbs fails to explicitly teach or disclose the limitations above. Osawa Shohei (WO 2019207826 A1) discloses utilizing reinforcement learning with each agent in a multi-agent environment in order to improve the accuracy of the estimated virtual revenue, wherein the reward is calculated according to the profit. The reference further discloses the agent providing a reward by providing information generated by weighted information received from a plurality of information supply agents. However, Osawa Shohei fails to explicitly teach or disclose the limitations above. Radovic et al. (“Revealing Robust Oil and Gas Company Macro-Strategies using Deep Multi-Agent Reinforcement Learning,” 2022) discloses utilizing reinforcement learning to maximize payouts by utilizing a real-valued reward signal, which is a value predicated on the efficacy of its developed strategy towards achieving a goal. The reference further discloses an agent’s reward function focusing on maximizing shareholder value via payments through the use of an agented reinforcement learning model. However, Radovic fails to explicitly teach or disclose the limitations above. Claim 13 would be allowable if rewritten to overcome the rejection(s) under 35 U.S.C. 101, set forth in this Office action and to include all of the limitations of the base claim and any intervening claims. As allowable subject matter has been indicated, applicant's reply must either comply with all formal requirements or specifically traverse each requirement not complied with. See 37 CFR 1.111(b) and MPEP § 707.07(a). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: Selfridge (US 20160063992 A1) discloses using the same fixed reward for each agent in a hierarchical reinforcement learning arrangement Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Sara G Brown whose telephone number is (469)295-9145. The examiner can normally be reached M-F 8:00 am- 5:00 pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Brian Epstein can be reached at (571) 270-5389. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SARA GRACE BROWN/Primary Examiner, Art Unit 3625
Read full office action

Prosecution Timeline

Mar 04, 2025
Application Filed
Apr 07, 2026
Non-Final Rejection mailed — §101, §103
Jun 08, 2026
Response Filed
Sep 02, 2026
Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737701
APPARATUS AND METHOD FOR CROP YIELD PREDICTION
5y 3m to grant Granted Sep 15, 2026
Patent 12718306
COMPUTER-BASED METHOD AND SYSTEM FOR MANAGING A FOOD INVENTORY OF A FLIGHT
2y 8m to grant Granted Aug 25, 2026
Patent 12700047
ANALYZING AND ENHANCING PERFORMANCE OF OILFIELD ASSETS
2y 0m to grant Granted Aug 04, 2026
Patent 12619682
IDENTIFYING OPERATION ANOMALIES OF SUBTERRANEAN DRILLING EQUIPMENT
4y 6m to grant Granted May 05, 2026
Patent 12602620
APPARATUS AND A METHOD FOR THE IDENTIFICATION OF A BREAKAWAY POINT
2y 11m to grant Granted Apr 14, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
29%
Grant Probability
62%
With Interview (+33.2%)
3y 5m (~1y 10m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 161 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month