DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-18 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim(s) recite(s):
a method, non-transitory computer-readable medium storing instructions for implementation of a multi-agent reinforcement learning based decision system, and an apparatus, comprising:
initializing a simulation environment comprising decision points for dispatching materials and attributes, initializing reinforcement learning agents representative of the decision points, initializing domain expert heuristics for the decision points, and iteratively training the reinforcement learning based decision system with the initialized domain expert heuristics on the simulation environment.
The claims include an abstract idea directed to a mathematical concept, specifically mathematical calculations for determining and updating values associated with the reinforcement learning agents dispatch policies (decision points) in a simulator.
This judicial exception is not integrated into a practical application because the steps of: 1) initializing domain expert heuristics for the decision points (dispatch policies), and 2) iteratively training the reinforcement learning based decision system with the initialized domain expert heuristics, merely describe how the mathematical calculations are initialized within a generic simulator environment without imposing any meaningful limits that tie the abstract idea to a particular technological implementation or improvement. For example, the steps do not tie the math to any particular real-world warehouse hardware, or change how the simulator or any other tech (conveyor belt, robot, etc.) physically operates at a low level.
The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the additional elements - initializing domain expert heuristics, and iteratively training the reinforcement learning based decision system with the initialized domain expert heuristics - are just generic simulator and training operations that were well-understood, routine, and conventional in the art at the time of the invention. These elements do not add significantly more than the abstract idea. See MPEP §2106.04(d)
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Jenkins et al (US PG. Pub. 2026/0030514) in view of Meghani et al (US PG. Pub. 2023/0094381). Relative to claims 1-6, Jenkins discloses:
claim 1) A method for implementation of a multi-agent reinforcement learning based decision system for a system (system may include a facility or business, Para. 0029, for multi-agent reinforcement learning, see Para. 0032, and agent deployment system)(Fig. 1), the method comprising:
initializing an environment comprising decision points for dispatching materials and attributes (see device or equipment of any type, Para. 0029) in a system (Para. 0066),
the environment is configured to request a decision for dispatching materials and attributes of the system, at the decision points to the multi-agent reinforcement learning based decision system (Para. 0029, 0039);
initializing reinforcement learning agents representative of the decision points for the multi-agent reinforcement learning based decision system (Para. 0054, agents 26 are selected and deployed from the larger or total set of agents needed to perform an AI-based intervention based on a trigger);
claim 2) iteratively training the reinforcement learning based decision system on the environment (Para. 0055, the training may be iteratively done, see also iterative interaction with the data, Para. 0038) comprising:
for each time step in the environment:
receiving a current state and reward, the reward based on each successful material dispatch for the each of the reinforcement learning agents (Para. 0038, 0057); and
for a decision point from the decision points being reached, selecting a decision for the decision point from the domain expert heuristics or from a decision generated by the reinforcement learning based decision system (Para. 0038);
the each time step in the environment is iteratively executed until a convergence or a specified goal is reached (Para. 0066).
claim 3) the reinforcement learning based decision system is configured to select a decision from either the initialized domain expert heuristics or from the trained decision heuristic (Para. 0038).
claim 4) the environment is configurable via adjustable parameters at multiple levels of the system (see learned parameters, Para. 0033).
claim 5) the reinforcement learning agents (see different agents 26A-N, as well as other agents)(Fig. 1) are heterogenous for classes of decisions to be made, and the reinforcement learning agents are trained in parallel (Para. 0054).
claim 6) truncating data provided to the reinforcement learning based decision system at each iteration so that data received by the reinforcement learning based decision system is of a same size during the training (Para. 0055, see hyperparameters can include batch size).
Jenkins does not expressly disclose:
claim 1) the environment is a simulation environment, the multi-agent reinforcement learning based decision system is for a materials handling system;
the simulation environment comprises decision points for dispatching materials and attributes of a materials handling system,
initializing domain expert heuristics for the decision points for the multi-agent reinforcement learning based decision system;
iteratively training the reinforcement learning based decision system with the initialized domain expert heuristics on the simulation environment;
claim 2) the iterative training the reinforcement learning based decision system is initialized with domain expert heuristics on the simulation environment, and the current state and reward is from the simulator; or
claim 4) the configurable environment with the adjustable parameters is a simulation environment.
Meghani teaches:
claim 1) the environment is a simulation environment (Para. 0078), the multi-agent reinforcement learning based decision system is for a materials handling system (Para. 0096, see warehouse),
the simulation environment comprising the decision points is for dispatching materials and attributes of a materials handling system (Para. 0096, 0101, see RL begins the learning process by incorporating a computational agent to make decisions by trial and error),
initializing domain expert heuristics for the decision points for the multi-agent reinforcement learning based decision system (Para. 0170, see feedback from the domain expert/user),
claim 2) the iterative training is initialized with domain expert heuristics (Para. 0170, 0090), and the current state and reward is from the simulator (Para. 0079); and
claim 4) the configurable environment with the adjustable parameters is a simulation environment (Para. 0079).
Meghani teaches the: simulation environment, the multi-agent reinforcement learning based decision system is for a materials handling system, initializing domain expert heuristics, and iterative training the reinforcement learning based decision system with the domain expert heuristics as described above, for the purpose of providing an improved system and method for incorporating deep reinforcement learning trained artificial intelligence auto-schedulers to receive project database inputs and to automatically generate detailed, optimized schedules, while minimizing costs (Para. 0005).
It would have been obvious to one of ordinary skill in the art on or before the time of the filing to modify the system of Jenkins with the: simulation environment, the multi-agent reinforcement learning based decision system is for a materials handling system, initializing domain expert heuristics, and iterative training the reinforcement learning based decision system with the domain expert heuristics, as taught in Meghani for the purpose of providing an improved system and method for incorporating deep reinforcement learning trained artificial intelligence auto-schedulers to receive project database inputs and to automatically generate detailed, optimized schedules, while minimizing costs.
Relative to claims 7-12, the disclosure of Jenkins includes:
claim 7) A non-transitory computer readable medium, storing instructions for implementation of a multi-agent reinforcement learning based decision system (Para. 0073), the instructions comprising:
initializing an environment comprising decision points for dispatching materials and attributes (Para. 0066),
the environment is configured to request a decision for dispatching materials and attributes at the decision points to the multi-agent reinforcement learning based decision system (Para. 0029; 0039); and
initializing reinforcement learning agents representative of the decision points for the multi-agent reinforcement learning based decision system (Para. 0054).
claim 8) iteratively training the reinforcement learning based decision system on the simulation environment (Para. 0055, 0032) comprising:
for each time step in the environment:
receiving a current state and reward, the reward based on each successful material dispatch for the each of the reinforcement learning agents (Para. 0038, 0057); and
for a decision point from the decision points being reached, selecting a decision for the decision point from the domain expert heuristics or from a decision generated by the reinforcement learning based decision system (Para. 0038);
the each time step in the environment is iteratively executed until a convergence or a specified goal is reached (Para. 0066).
claim 9) the reinforcement learning based decision system is configured to select a decision from either the initialized domain expert heuristics or from the trained decision heuristic (Para. 0038).
claim 10) the environment is configurable via adjustable parameters at multiple levels of the materials handling system (Para. 0033).
claim 11) the reinforcement learning agents are heterogenous for classes of decisions to be made, and the reinforcement learning agents are trained in parallel (Para. 0054), and
claim 12) truncating data provided to the reinforcement learning based decision system at each iteration so that data received by the reinforcement learning based decision system is of a same size during the training (Para. 0055).
Jenkins does not expressly disclose:
claim 7) the environment is a simulation environment, the multi-agent reinforcement learning based decision system is for a materials handling system;
the simulation environment comprises decision points for dispatching materials and attributes of a materials handling system,
initializing domain expert heuristics for the decision points for the multi-agent reinforcement learning based decision system;
iteratively training the reinforcement learning based decision system with the initialized domain expert heuristics on the simulation environment;
claim 8) the iterative training the reinforcement learning based decision system is initialized with domain expert heuristics on the simulation environment, and the current state and reward is from the simulator; or
claim 10) the configurable environment with adjustable parameters is a simulation environment.
Meghani teaches:
claim 7) the environment is a simulation environment (Para. 0078), the multi-agent reinforcement learning based decision system is for a materials handling system (Para. 0096, see warehouse),
the simulation environment comprises decision points is for dispatching materials and attributes of a materials handling system (Para. 0096, 0101, see RL begins the learning process by incorporating a computational agent to make decisions by trial and error),
initializing domain expert heuristics for the decision points for the multi-agent reinforcement learning based decision system (Para. 0170, see feedback from the domain expert/user),
claim 8) the iterative training is initialized with domain expert heuristics (Para. 0170, 0090), the current state and reward is from the simulator (Para. 0079); and
claim 10) the configurable environment with adjustable parameters is a simulation environment (Para. 0079).
Meghani teaches the: simulation environment, the multi-agent reinforcement learning based decision system is for a materials handling system, initializing domain expert heuristics, and iterative training the reinforcement learning based decision system with the domain expert heuristics as described above, for the purpose of providing an improved system and method for incorporating deep reinforcement learning trained artificial intelligence auto-schedulers to receive project database inputs and to automatically generate detailed, optimized schedules, while minimizing costs (Para. 0005).
It would have been obvious to one of ordinary skill in the art on or before the time of the filing to modify the system of Jenkins with the simulation environment, the multi-agent reinforcement learning based decision system is for a materials handling system, initializing domain expert heuristics, and iterative training the reinforcement learning based decision system with the domain expert heuristics, as taught in Meghani for the purpose of providing an improved system and method for incorporating deep reinforcement learning trained artificial intelligence auto-schedulers to receive project database inputs and to automatically generate detailed, optimized schedules, while minimizing costs.
Relative to claims 13-18, the disclosure of Jenkins includes:
claim 13) An apparatus for implementation of a multi-agent reinforcement learning based decision system for a system (Para. 0032), the apparatus comprising:
a processor (Para. 0073), configured to:
initialize an environment (Para. 0029) comprising decision points for dispatching materials and attributes of the system (Para. 0066),
the environment configured to request a decision for materials dispatch at the decision points to the multi-agent reinforcement learning based decision system (Para. 0039); and
initialize reinforcement learning agents representative of the decision points for the multi-agent reinforcement learning based decision system (Para. 0054);
claim 14) the processor is configured to iteratively train the reinforcement learning based decision system on the environment (Para. 0055) by:
for each time step in the environment:
receiving a current state and reward, the reward based on each successful material dispatch for the each of the reinforcement learning agents (Para. 0038, 0057); and
for a decision point from the decision points being reached, selecting a decision for the decision point from the domain expert heuristics or from a decision generated by the reinforcement learning based decision system (Para. 0038);
the each time step in the environment is iteratively executed until a convergence or a specified goal is reached (Para. 0066).
claim 15) the reinforcement learning based decision system is configured to select a decision from either the initialized domain expert heuristics or from the trained decision heuristic (Para. 0038).
claim 16) the environment is configurable via adjustable parameters at multiple levels of the materials handling system (Para. 0033).
claim 17) the reinforcement learning agents are heterogenous for classes of decisions to be made, and the reinforcement learning agents are trained in parallel (Para. 0054), and
claim 18) truncating data provided to the reinforcement learning based decision system at each iteration so that data received by the reinforcement learning based decision system is of a same size during the training (Para. 0055).
Jenkins does not expressly disclose:
Jenkins does not expressly disclose:
claim 13) the environment is a simulation environment, the multi-agent reinforcement learning based decision system is for a materials handling system;
the simulation environment comprising the decision points for dispatching materials and attributes is for a materials handling system,
initializing domain expert heuristics for the decision points for the multi-agent reinforcement learning based decision system;
iteratively training the reinforcement learning based decision system with the initialized domain expert heuristics on the simulation environment;
claim 14) the iterative training the reinforcement learning based decision system is initialized with domain expert heuristics on the simulation environment, and the current state and reward is from the simulator; or
claim 16) the configurable environment with adjustable parameters is a simulation environment.
Meghani teaches:
claim 13) the environment is a simulation environment (Para. 0078), the multi-agent reinforcement learning based decision system is for a materials handling system (Para. 0096, see warehouse),
the simulation environment comprising the decision points is for dispatching materials and attributes for a materials handling system (Para. 0096, 0101, see RL begins the learning process by incorporating a computational agent to make decisions by trial and error),
initializing domain expert heuristics for the decision points for the multi-agent reinforcement learning based decision system (Para. 0170, see feedback from the domain expert/user),
claim 14) the iterative training is initialized with domain expert heuristics (Para. 0170, 0090), the current state and reward is from the simulator (Para. 0079); and
claim 16) the configurable environment with adjustable parameters is a simulation environment (Para. 0079).
Meghani teaches the: simulation environment, the multi-agent reinforcement learning based decision system is for a materials handling system, initializing domain expert heuristics, and iterative training the reinforcement learning based decision system with the domain expert heuristics as described above, for the purpose of providing an improved system and method for incorporating deep reinforcement learning trained artificial intelligence auto-schedulers to receive project database inputs and to automatically generate detailed, optimized schedules, while minimizing costs (Para. 0005).
It would have been obvious to one of ordinary skill in the art on or before the time of the filing to modify the system of Jenkins with the simulation environment, the multi-agent reinforcement learning based decision system is for a materials handling system, initializing domain expert heuristics, and iterative training the reinforcement learning based decision system with the domain expert heuristics, as taught in Meghani for the purpose of providing an improved system and method for incorporating deep reinforcement learning trained artificial intelligence auto-schedulers to receive project database inputs and to automatically generate detailed, optimized schedules, while minimizing costs.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Krnjaic US 2026/0054929: See Fig. 1A of warehouse environment, Para. 0003; 0008, 0030; 0032
Any inquiry concerning this communication or earlier communications from the examiner should be directed to YOLANDA RENEE CUMBESS whose telephone number is (571)270-5527. The examiner can normally be reached M-F 10-6.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Gene Crawford can be reached at 571-272-6911. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/YOLANDA R CUMBESS/Primary Examiner, Art Unit 3651