Prosecution Insights
Last updated: October 01, 2026
Application No. 18/888,905

MULTI-AGENT REINFORCEMENT LEARNING FRAMEWORK FOR DYNAMIC DISPATCHING IN MATERIAL HANDLING SYSTEMS

Non-Final OA §101§103
Filed
Sep 18, 2024
Examiner
CUMBESS, YOLANDA RENEE
Art Unit
3651
Tech Center
3600 — Transportation & Electronic Commerce
Assignee
Hitachi Ltd.
OA Round
1 (Non-Final)
87%
Grant Probability
Favorable
1-2
OA Rounds
3m
Est. Remaining
95%
With Interview

Examiner Intelligence

Grants 87% — above average
87%
Career Allowance Rate
981 granted / 1129 resolved
+34.9% vs TC avg
Moderate +8% lift
Without
With
+8.5%
Interview Lift
resolved cases with interview
Typical timeline
2y 3m
Avg Prosecution
30 currently pending
Career history
1151
Total Applications
across all art units

Statute-Specific Performance

§101
1.1%
-38.9% vs TC avg
§103
46.8%
+6.8% vs TC avg
§102
19.3%
-20.7% vs TC avg
§112
31.1%
-8.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1129 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-18 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim(s) recite(s): a method, non-transitory computer-readable medium storing instructions for implementation of a multi-agent reinforcement learning based decision system, and an apparatus, comprising: initializing a simulation environment comprising decision points for dispatching materials and attributes, initializing reinforcement learning agents representative of the decision points, initializing domain expert heuristics for the decision points, and iteratively training the reinforcement learning based decision system with the initialized domain expert heuristics on the simulation environment. The claims include an abstract idea directed to a mathematical concept, specifically mathematical calculations for determining and updating values associated with the reinforcement learning agents dispatch policies (decision points) in a simulator. This judicial exception is not integrated into a practical application because the steps of: 1) initializing domain expert heuristics for the decision points (dispatch policies), and 2) iteratively training the reinforcement learning based decision system with the initialized domain expert heuristics, merely describe how the mathematical calculations are initialized within a generic simulator environment without imposing any meaningful limits that tie the abstract idea to a particular technological implementation or improvement. For example, the steps do not tie the math to any particular real-world warehouse hardware, or change how the simulator or any other tech (conveyor belt, robot, etc.) physically operates at a low level. The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the additional elements - initializing domain expert heuristics, and iteratively training the reinforcement learning based decision system with the initialized domain expert heuristics - are just generic simulator and training operations that were well-understood, routine, and conventional in the art at the time of the invention. These elements do not add significantly more than the abstract idea. See MPEP §2106.04(d) Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Jenkins et al (US PG. Pub. 2026/0030514) in view of Meghani et al (US PG. Pub. 2023/0094381). Relative to claims 1-6, Jenkins discloses: claim 1) A method for implementation of a multi-agent reinforcement learning based decision system for a system (system may include a facility or business, Para. 0029, for multi-agent reinforcement learning, see Para. 0032, and agent deployment system)(Fig. 1), the method comprising: initializing an environment comprising decision points for dispatching materials and attributes (see device or equipment of any type, Para. 0029) in a system (Para. 0066), the environment is configured to request a decision for dispatching materials and attributes of the system, at the decision points to the multi-agent reinforcement learning based decision system (Para. 0029, 0039); initializing reinforcement learning agents representative of the decision points for the multi-agent reinforcement learning based decision system (Para. 0054, agents 26 are selected and deployed from the larger or total set of agents needed to perform an AI-based intervention based on a trigger); claim 2) iteratively training the reinforcement learning based decision system on the environment (Para. 0055, the training may be iteratively done, see also iterative interaction with the data, Para. 0038) comprising: for each time step in the environment: receiving a current state and reward, the reward based on each successful material dispatch for the each of the reinforcement learning agents (Para. 0038, 0057); and for a decision point from the decision points being reached, selecting a decision for the decision point from the domain expert heuristics or from a decision generated by the reinforcement learning based decision system (Para. 0038); the each time step in the environment is iteratively executed until a convergence or a specified goal is reached (Para. 0066). claim 3) the reinforcement learning based decision system is configured to select a decision from either the initialized domain expert heuristics or from the trained decision heuristic (Para. 0038). claim 4) the environment is configurable via adjustable parameters at multiple levels of the system (see learned parameters, Para. 0033). claim 5) the reinforcement learning agents (see different agents 26A-N, as well as other agents)(Fig. 1) are heterogenous for classes of decisions to be made, and the reinforcement learning agents are trained in parallel (Para. 0054). claim 6) truncating data provided to the reinforcement learning based decision system at each iteration so that data received by the reinforcement learning based decision system is of a same size during the training (Para. 0055, see hyperparameters can include batch size). Jenkins does not expressly disclose: claim 1) the environment is a simulation environment, the multi-agent reinforcement learning based decision system is for a materials handling system; the simulation environment comprises decision points for dispatching materials and attributes of a materials handling system, initializing domain expert heuristics for the decision points for the multi-agent reinforcement learning based decision system; iteratively training the reinforcement learning based decision system with the initialized domain expert heuristics on the simulation environment; claim 2) the iterative training the reinforcement learning based decision system is initialized with domain expert heuristics on the simulation environment, and the current state and reward is from the simulator; or claim 4) the configurable environment with the adjustable parameters is a simulation environment. Meghani teaches: claim 1) the environment is a simulation environment (Para. 0078), the multi-agent reinforcement learning based decision system is for a materials handling system (Para. 0096, see warehouse), the simulation environment comprising the decision points is for dispatching materials and attributes of a materials handling system (Para. 0096, 0101, see RL begins the learning process by incorporating a computational agent to make decisions by trial and error), initializing domain expert heuristics for the decision points for the multi-agent reinforcement learning based decision system (Para. 0170, see feedback from the domain expert/user), claim 2) the iterative training is initialized with domain expert heuristics (Para. 0170, 0090), and the current state and reward is from the simulator (Para. 0079); and claim 4) the configurable environment with the adjustable parameters is a simulation environment (Para. 0079). Meghani teaches the: simulation environment, the multi-agent reinforcement learning based decision system is for a materials handling system, initializing domain expert heuristics, and iterative training the reinforcement learning based decision system with the domain expert heuristics as described above, for the purpose of providing an improved system and method for incorporating deep reinforcement learning trained artificial intelligence auto-schedulers to receive project database inputs and to automatically generate detailed, optimized schedules, while minimizing costs (Para. 0005). It would have been obvious to one of ordinary skill in the art on or before the time of the filing to modify the system of Jenkins with the: simulation environment, the multi-agent reinforcement learning based decision system is for a materials handling system, initializing domain expert heuristics, and iterative training the reinforcement learning based decision system with the domain expert heuristics, as taught in Meghani for the purpose of providing an improved system and method for incorporating deep reinforcement learning trained artificial intelligence auto-schedulers to receive project database inputs and to automatically generate detailed, optimized schedules, while minimizing costs. Relative to claims 7-12, the disclosure of Jenkins includes: claim 7) A non-transitory computer readable medium, storing instructions for implementation of a multi-agent reinforcement learning based decision system (Para. 0073), the instructions comprising: initializing an environment comprising decision points for dispatching materials and attributes (Para. 0066), the environment is configured to request a decision for dispatching materials and attributes at the decision points to the multi-agent reinforcement learning based decision system (Para. 0029; 0039); and initializing reinforcement learning agents representative of the decision points for the multi-agent reinforcement learning based decision system (Para. 0054). claim 8) iteratively training the reinforcement learning based decision system on the simulation environment (Para. 0055, 0032) comprising: for each time step in the environment: receiving a current state and reward, the reward based on each successful material dispatch for the each of the reinforcement learning agents (Para. 0038, 0057); and for a decision point from the decision points being reached, selecting a decision for the decision point from the domain expert heuristics or from a decision generated by the reinforcement learning based decision system (Para. 0038); the each time step in the environment is iteratively executed until a convergence or a specified goal is reached (Para. 0066). claim 9) the reinforcement learning based decision system is configured to select a decision from either the initialized domain expert heuristics or from the trained decision heuristic (Para. 0038). claim 10) the environment is configurable via adjustable parameters at multiple levels of the materials handling system (Para. 0033). claim 11) the reinforcement learning agents are heterogenous for classes of decisions to be made, and the reinforcement learning agents are trained in parallel (Para. 0054), and claim 12) truncating data provided to the reinforcement learning based decision system at each iteration so that data received by the reinforcement learning based decision system is of a same size during the training (Para. 0055). Jenkins does not expressly disclose: claim 7) the environment is a simulation environment, the multi-agent reinforcement learning based decision system is for a materials handling system; the simulation environment comprises decision points for dispatching materials and attributes of a materials handling system, initializing domain expert heuristics for the decision points for the multi-agent reinforcement learning based decision system; iteratively training the reinforcement learning based decision system with the initialized domain expert heuristics on the simulation environment; claim 8) the iterative training the reinforcement learning based decision system is initialized with domain expert heuristics on the simulation environment, and the current state and reward is from the simulator; or claim 10) the configurable environment with adjustable parameters is a simulation environment. Meghani teaches: claim 7) the environment is a simulation environment (Para. 0078), the multi-agent reinforcement learning based decision system is for a materials handling system (Para. 0096, see warehouse), the simulation environment comprises decision points is for dispatching materials and attributes of a materials handling system (Para. 0096, 0101, see RL begins the learning process by incorporating a computational agent to make decisions by trial and error), initializing domain expert heuristics for the decision points for the multi-agent reinforcement learning based decision system (Para. 0170, see feedback from the domain expert/user), claim 8) the iterative training is initialized with domain expert heuristics (Para. 0170, 0090), the current state and reward is from the simulator (Para. 0079); and claim 10) the configurable environment with adjustable parameters is a simulation environment (Para. 0079). Meghani teaches the: simulation environment, the multi-agent reinforcement learning based decision system is for a materials handling system, initializing domain expert heuristics, and iterative training the reinforcement learning based decision system with the domain expert heuristics as described above, for the purpose of providing an improved system and method for incorporating deep reinforcement learning trained artificial intelligence auto-schedulers to receive project database inputs and to automatically generate detailed, optimized schedules, while minimizing costs (Para. 0005). It would have been obvious to one of ordinary skill in the art on or before the time of the filing to modify the system of Jenkins with the simulation environment, the multi-agent reinforcement learning based decision system is for a materials handling system, initializing domain expert heuristics, and iterative training the reinforcement learning based decision system with the domain expert heuristics, as taught in Meghani for the purpose of providing an improved system and method for incorporating deep reinforcement learning trained artificial intelligence auto-schedulers to receive project database inputs and to automatically generate detailed, optimized schedules, while minimizing costs. Relative to claims 13-18, the disclosure of Jenkins includes: claim 13) An apparatus for implementation of a multi-agent reinforcement learning based decision system for a system (Para. 0032), the apparatus comprising: a processor (Para. 0073), configured to: initialize an environment (Para. 0029) comprising decision points for dispatching materials and attributes of the system (Para. 0066), the environment configured to request a decision for materials dispatch at the decision points to the multi-agent reinforcement learning based decision system (Para. 0039); and initialize reinforcement learning agents representative of the decision points for the multi-agent reinforcement learning based decision system (Para. 0054); claim 14) the processor is configured to iteratively train the reinforcement learning based decision system on the environment (Para. 0055) by: for each time step in the environment: receiving a current state and reward, the reward based on each successful material dispatch for the each of the reinforcement learning agents (Para. 0038, 0057); and for a decision point from the decision points being reached, selecting a decision for the decision point from the domain expert heuristics or from a decision generated by the reinforcement learning based decision system (Para. 0038); the each time step in the environment is iteratively executed until a convergence or a specified goal is reached (Para. 0066). claim 15) the reinforcement learning based decision system is configured to select a decision from either the initialized domain expert heuristics or from the trained decision heuristic (Para. 0038). claim 16) the environment is configurable via adjustable parameters at multiple levels of the materials handling system (Para. 0033). claim 17) the reinforcement learning agents are heterogenous for classes of decisions to be made, and the reinforcement learning agents are trained in parallel (Para. 0054), and claim 18) truncating data provided to the reinforcement learning based decision system at each iteration so that data received by the reinforcement learning based decision system is of a same size during the training (Para. 0055). Jenkins does not expressly disclose: Jenkins does not expressly disclose: claim 13) the environment is a simulation environment, the multi-agent reinforcement learning based decision system is for a materials handling system; the simulation environment comprising the decision points for dispatching materials and attributes is for a materials handling system, initializing domain expert heuristics for the decision points for the multi-agent reinforcement learning based decision system; iteratively training the reinforcement learning based decision system with the initialized domain expert heuristics on the simulation environment; claim 14) the iterative training the reinforcement learning based decision system is initialized with domain expert heuristics on the simulation environment, and the current state and reward is from the simulator; or claim 16) the configurable environment with adjustable parameters is a simulation environment. Meghani teaches: claim 13) the environment is a simulation environment (Para. 0078), the multi-agent reinforcement learning based decision system is for a materials handling system (Para. 0096, see warehouse), the simulation environment comprising the decision points is for dispatching materials and attributes for a materials handling system (Para. 0096, 0101, see RL begins the learning process by incorporating a computational agent to make decisions by trial and error), initializing domain expert heuristics for the decision points for the multi-agent reinforcement learning based decision system (Para. 0170, see feedback from the domain expert/user), claim 14) the iterative training is initialized with domain expert heuristics (Para. 0170, 0090), the current state and reward is from the simulator (Para. 0079); and claim 16) the configurable environment with adjustable parameters is a simulation environment (Para. 0079). Meghani teaches the: simulation environment, the multi-agent reinforcement learning based decision system is for a materials handling system, initializing domain expert heuristics, and iterative training the reinforcement learning based decision system with the domain expert heuristics as described above, for the purpose of providing an improved system and method for incorporating deep reinforcement learning trained artificial intelligence auto-schedulers to receive project database inputs and to automatically generate detailed, optimized schedules, while minimizing costs (Para. 0005). It would have been obvious to one of ordinary skill in the art on or before the time of the filing to modify the system of Jenkins with the simulation environment, the multi-agent reinforcement learning based decision system is for a materials handling system, initializing domain expert heuristics, and iterative training the reinforcement learning based decision system with the domain expert heuristics, as taught in Meghani for the purpose of providing an improved system and method for incorporating deep reinforcement learning trained artificial intelligence auto-schedulers to receive project database inputs and to automatically generate detailed, optimized schedules, while minimizing costs. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Krnjaic US 2026/0054929: See Fig. 1A of warehouse environment, Para. 0003; 0008, 0030; 0032 Any inquiry concerning this communication or earlier communications from the examiner should be directed to YOLANDA RENEE CUMBESS whose telephone number is (571)270-5527. The examiner can normally be reached M-F 10-6. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Gene Crawford can be reached at 571-272-6911. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /YOLANDA R CUMBESS/Primary Examiner, Art Unit 3651
Read full office action

Prosecution Timeline

Sep 18, 2024
Application Filed
Aug 11, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12735268
CONVEYING APPARATUS AND CONVEYING METHOD
3y 1m to grant Granted Sep 15, 2026
Patent 12730446
TRANSFER ROBOT SYSTEM AND THE TRANSFER ROBOT SYSTEM DRIVING METHOD
3y 6m to grant Granted Sep 08, 2026
Patent 12715714
ROBOT INFORMED DYNAMIC PARCEL INFLOW GATING
3y 7m to grant Granted Aug 25, 2026
Patent 12709480
Apparatus and Method for Charging a Load Handling Device
4y 0m to grant Granted Aug 18, 2026
Patent 12709479
SYSTEM AND METHOD FOR EFFICIENTLY TRANSFERRING RECEPTACLES OUT OF A RECEPTACLE STORAGE AREA
2y 7m to grant Granted Aug 18, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
87%
Grant Probability
95%
With Interview (+8.5%)
2y 3m (~3m remaining)
Median Time to Grant
Low
PTA Risk
Based on 1129 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month