Prosecution Insights
Last updated: October 01, 2026
Application No. 17/546,022

INTEGRATED AI PLANNERS AND RL AGENTS THROUGH AI PLANNING ANNOTATION IN RL

Final Rejection §101§103
Filed
Dec 08, 2021
Examiner
ZHEN, LI B
Art Unit
2121
Tech Center
2100 — Computer Architecture & Software
Assignee
International Business Machines Corporation
OA Round
2 (Final)
54%
Grant Probability
Moderate
3-4
OA Rounds
4m
Est. Remaining
94%
With Interview

Examiner Intelligence

Grants 54% of resolved cases
54%
Career Allowance Rate
91 granted / 168 resolved
-0.8% vs TC avg
Strong +40% interview lift
Without
With
+39.9%
Interview Lift
resolved cases with interview
Typical timeline
5y 1m
Avg Prosecution
6 currently pending
Career history
174
Total Applications
across all art units

Statute-Specific Performance

§101
21.0%
-19.0% vs TC avg
§103
52.9%
+12.9% vs TC avg
§102
8.5%
-31.5% vs TC avg
§112
11.1%
-28.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 168 resolved cases

Office Action

§101 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Status of Claims Claims 1, 3-14, and 16-20 are pending and are examined herein. Claims 1, 3-14, and 16-20 are rejected under 35 USC 101 as being directed to an abstract idea without significantly more. Claims 1, 3-14, and 16-20 are rejected under 35 USC 103. Information Disclosure Statement The information disclosure statements (IDS) submitted on 8/22/2025, 3/23/2026, and 7/21/2026 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Response to Arguments/Amendments Amendment to claims 10 and 11 overcomes the 35 USC 112(b) rejection. On page 11 of the response, applicant argues that the claims, as currently presented, are directed toward patent-eligible subject matter, “the claim is directed to a computer-implemented method for integrating an AI planner with a reinforcement learning (RL) agent via a structured planning-annotated reinforcement learning (PaRL) framework executed by a particularly configured computing device comprising a hierarchical reinforcement learning module, a symbolic planning module, and a reinforcement learning module.” Response: Examiner respectfully disagrees. According to the specification, an example of a particularly configured computer hardware platform is described in Fig. 14, which describes generic computer components with blocks to represent the various modules that perform functions recited in the independent claims. Examiner interprets the “particularly configured computing device” as a generic computer the executes instructions to perform the abstract ideas recited in the independent claims. Neither the claims nor the specification includes additional details of the modules that perform the claimed functions. Therefore, the claims do not include additional elements that integrate the abstract ideas into a practical application and do not amount to significantly more. On page 12, applicant argues, “[t]he claim elements go beyond mere data processing or mathematical abstraction - they describe how the symbolic planning model is integrated into a reinforcement learning pipeline, producing symbolic options that are used to train an RL agent which solves a defined RL problem. This pipeline involves functional cooperation among discrete system components to achieve a real-world result.” Response: Examiner disagrees. Neither the claims nor the specification describes a reinforcement learning pipeline. The claims describe a process of mappings states from a Markov Decision process to an AI planning states of an AI planning model. According to Fig. 2 and [0047] of the published specification, the MDP states and AI plannings states are represented using descriptor languages (see also claim 9), which are human readable descriptors. Therefore, a human can generate the state mapping in their mind. On page 12, applicant argues, “[t]he claimed method is implemented by a particularly configured computing device that executes a structured pipeline to generate, annotate, and use symbolic representations for training an RL agent. This architecture, and the claimed interplay of modules, provides a technological solution to a technological problem - namely, improving the efficacy and structure of RL systems through the integration of symbolic AI planning models.” Response: Examiner disagrees, see response to argument 2) and 3) above regarding a particularly configured computing device and a structured pipeline. The additional elements in the claims are directed to mere instructions to apply the abstract idea on a generic computer. On page 14, applicant argues that Konidaris does not teach annotating an RL task with an AI planning task or generating a PaRL task as recited in amended claim 1. Response: Examiner disagrees. Applicant’s arguments do not explain why the cited sections of Konidaris fail to teach the recited claim language. As to applicant’s argument directed to “using an external symbolic planning model to construct symbolic options”, these features are not recited in the claims. It is also noted that the specification does not appear to describe the argued features. On page 15, applicant argues that Leonetti does not teach “mapping state spaces between MDP and AI planning models (as required by claim 1)” and “generating symbolic options from planning operators, nor using them to train an RL agent”. In addition, “Leonetti also does not annotate the RL task with planning tasks, nor describe a pipeline resembling the PaRL task conceptually or structurally.” Response: Examiner disagrees. Leonetti teaches mapping state spaces between MDP and AI planning models in at least Sections 2.3, 4, and 4.3 (see also prior art rejection below. Applicant alleges that Leonetti does not teach mapping state spaces between MDP and AI planning models without providing an explanation why the cited sections of Leonetti fail to teach the argued limitations. As to the argument regarding generating symbolic options and annotating the RL task, the primary reference Konidaris was relied up to teach this feature, see response to argument 5) above and the prior art rejection below. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1, 3-14, and 16-20 are rejected under 35 USC 101 because the claimed invention is directed to an abstract idea without significantly more. When considering subject matter eligibility under 35 U.S.C. 101, it must be determined whether the claim is directed to one of the four statutory categories of invention, i.e., process, machine, manufacture, or composition of matter (Step 1). If the claim does fall within one of the statutory categories, the second step in the analysis is to determine whether the claim is directed to a judicial exception (Step 2A). The Step 2A analysis is broken into two prongs. In the first prong (Step 2A, Prong 1), it is determined whether or not the claims recite a judicial exception (e.g., mathematical concepts, mental processes, certain methods of organizing human activity). If it is determined in Step 2A, Prong 1 that the claims recite a judicial exception, the analysis proceeds to the second prong (Step 2A, Prong 2), where it is determined whether or not the claims integrate the judicial exception into a practical application. If it is determined at step 2A, Prong 2 that the claims do not integrate the judicial exception into a practical application, the analysis proceeds to determining whether the claim is a patent-eligible application of the exception (Step 2B). If an abstract idea is present in the claim, any element or combination of elements in the claim must be sufficient to ensure that the claim integrates the judicial exception into a practical application, or else amounts to significantly more than the abstract idea itself. Applicant is advised to consult the 2019 PEG for more details of the analysis. Step 1 Analysis According to the first part of the analysis, in the instant case Claims 1-13 are directed to a method, Claims 14-19 are directed to a device, and Claim 20 is directed to storage medium; consequently, these claims fall within one of the four statutory categories (i.e. process, machine, manufacture, or composition of matter). Step 2 Analysis (Combined Step 2A Prong 1-2 and Step 2B Analysis) Claim 1 includes the following recitation of an abstract idea: identifying…an RL problem over the MDP (This is practical to perform in the human mind under its broadest reasonable interpretation aside from the recitation of generic computer components.) mapping…state spaces from the MDP states in the RL environment to Al planning states of the Al planning model (This is practical to perform in the human mind under its broadest reasonable interpretation aside from the recitation of generic computer components.) annotating the RL task with an Al planning task from the mapping to generate a PaRL task (This is practical to perform in the human mind under its broadest reasonable interpretation aside from the recitation of generic computer components.) Claim 1 recites the following additional elements which, considered individually and as an ordered combination, do not integrate the abstract idea into a practical application or amount to significantly more than the abstract idea: receiving, by a Hierarchical Reinforcement Learning (HRL) module of a particularly configured computing device, a Markov Decision Process (MDP) (This is insignificant extra-solution activity. See MPEP 2106.05(g). Moreover, sending or receiving data is well-understood, routine, conventional as evidenced by the court cases cited at MPEP 2106.05(d), example i. Receiving or transmitting data); receiving, by the computing device, a description of the Markov decision process (MDP) having a plurality of states in an RL environment to generate an RL task to solve the RL problem (This is insignificant extra-solution activity. See MPEP 2106.05(g). Moreover, sending or receiving data is well-understood, routine, conventional as evidenced by the court cases cited at MPEP 2106.05(d), example i. Receiving or transmitting data.) receiving, by an AI planning module of the computing device, an Al planning model described in a symbolic planning language (This is insignificant extra-solution activity. See MPEP 2106.05(g). Moreover, sending or receiving data is well-understood, routine, conventional as evidenced by the court cases cited at MPEP 2106.05(d), example i. Receiving or transmitting data.) training, by the reinforcement learning module, the RL agent using the PaRL task, wherein the PaRL task includes symbolic options derived from the AI planning task (The limitation recites training a reinforcement learning module using the generated PaRL task from the previous step. The limitation does not include details or implementation of training a reinforcement module; therefore, this is directed to mere instruction to apply the abstract idea on a generic computer. The additional limitation does not integrate the judicial exception into a practical application and does not amount to significantly more than the judicial exception. See MPEP 2106.05(f)); and solving the identified RL problem using the generated PaRL task (This is an attempt to limit the abstract idea to a particular field of use or technological environment, which does not integrate the abstract idea into a practical application or amount to significantly more than the abstract idea. See MPEP 2106.05(h).) In addition, the additional limitations of using the computing device to identify an RL problem and a reinforcement learning module of the computing device to map state spaces are interpreted as additional element of mere instructions to apply the abstract idea on a generic computer as the claims do not provide any specific implements of the identifying and mapping functions, see MPEP 2106.05(f). Claim 1 does not reflect an improvement to computer technology or any other technology. The claim as a whole, looking at the additional elements individually and in combination, does not integrate the judicial exception into a practical application and does not amount to significantly more than the identified judicial exception. Claim 3 recites at least the abstract idea identified above in the claim upon which it depends. Claim 3 includes the following recitation of an additional abstract idea: formulating by the PaRL task an options framework for the MDP (This is practical to perform in the human mind under its broadest reasonable interpretation aside from the recitation of generic computer components.) Claim 3 recites no further additional elements which, considered individually and as an ordered combination with the additional elements from the claim upon which it depends, that integrate the abstract idea into a practical application or amount to significantly more than the abstract ideas. Claim 3 does not reflect an improvement to computer technology or any other technology. Claim 4 recites at least the abstract idea identified above in the claim upon which it depends. Claim 4 includes the following recitation of an additional abstract idea: generating one or more sets of Al plans in the options framework (This is practical to perform in the human mind under its broadest reasonable interpretation aside from the recitation of generic computer components.) selecting options from the options framework for training the RL agent by ranking the options with scores (This is practical to perform in the human mind under its broadest reasonable interpretation aside from the recitation of generic computer components.) Claim 4 recites the following additional elements which, considered individually and as an ordered combination, do not integrate the abstract idea into a practical application or amount to significantly more than the abstract idea: sending the options to the RL agent (This is insignificant extra-solution activity. See MPEP 2106.05(g). Moreover, sending or receiving data is well-understood, routine, conventional as evidenced by the court cases cited at MPEP 2106.05(d), example i. Receiving or transmitting data.) Claim 4 does not reflect an improvement to computer technology or any other technology. The claim as a whole, looking at the additional elements individually and in combination, does not integrate the judicial exception into a practical application and does not amount to significantly more than the identified judicial exception. Claim 5 recites at least the abstract idea identified above in the claim upon which it depends. Claim 5 recites the following additional elements which, considered individually and as an ordered combination, do not integrate the abstract idea into a practical application or amount to significantly more than the abstract idea: the selecting options from the options framework is performed online or offline (This appears to be an attempt to limit data to a particular type or source, which amounts to no more than an attempt to limit the abstract idea to a particular field of use or technological environment. This does not provide the significantly more element required to overcome the abstract idea. See MPEP 2106.05(h).) Claim 5 does not reflect an improvement to computer technology or any other technology. The claim as a whole, looking at the additional elements individually and in combination, does not integrate the judicial exception into a practical application and does not amount to significantly more than the identified judicial exception. Claim 6 recites at least the abstract idea identified above in the claim upon which it depends. Claim 6 includes the following recitation of an additional abstract idea: generating a plan given trajectory (This is practical to perform in the human mind under its broadest reasonable interpretation aside from the recitation of generic computer components.) ranking options according to a scoring function (This is practical to perform in the human mind under its broadest reasonable interpretation aside from the recitation of generic computer components.) Claim 6 recites the following additional elements which, considered individually and as an ordered combination, do not integrate the abstract idea into a practical application or amount to significantly more than the abstract idea: sending the options with a highest score to the RL agent (This is insignificant extra-solution activity. See MPEP 2106.05(g). Moreover, sending or receiving data is well-understood, routine, conventional as evidenced by the court cases cited at MPEP 2106.05(d), example i. Receiving or transmitting data.) Claim 6 does not reflect an improvement to computer technology or any other technology. The claim as a whole, looking at the additional elements individually and in combination, does not integrate the judicial exception into a practical application and does not amount to significantly more than the identified judicial exception. Claim 7 recites at least the abstract idea identified above in the claim upon which it depends. Claim 7 includes the following recitation of an additional abstract idea: guiding a sampling process…to sample the options (This is practical to perform in the human mind under its broadest reasonable interpretation aside from the recitation of generic computer components.) Claim 7 recites the following additional elements which, considered individually and as an ordered combination, do not integrate the abstract idea into a practical application or amount to significantly more than the abstract idea: sending the options to the RL agent (This is insignificant extra-solution activity. See MPEP 2106.05(g). Moreover, sending or receiving data is well-understood, routine, conventional as evidenced by the court cases cited at MPEP 2106.05(d), example i. Receiving or transmitting data.) by a PaRL planner (This falls under mere instructions to apply a planner. See MPEP 2106.05(f).) Claim 7 does not reflect an improvement to computer technology or any other technology. The claim as a whole, looking at the additional elements individually and in combination, does not integrate the judicial exception into a practical application and does not amount to significantly more than the identified judicial exception. Claim 8 recites at least the abstract idea identified above in the claim upon which it depends. Claim 8 includes the following recitation of an additional abstract idea: at least one mapping selected from the group of: abstraction mapping in Al planning, heuristic mapping between state spaces, and rule-based mapping (This is practical to perform in the human mind under its broadest reasonable interpretation aside from the recitation of generic computer components.) Claim 8 recites no further additional elements which, considered individually and as an ordered combination with the additional elements from the claim upon which it depends, that integrate the abstract idea into a practical application or amount to significantly more than the abstract ideas. Claim 8 does not reflect an improvement to computer technology or any other technology. The claim as a whole, looking at the additional elements individually and in combination, does not integrate the judicial exception into a practical application and does not amount to significantly more than the identified judicial exception. Claim 9 recites at least the abstract idea identified above in the claim upon which it depends. Claim 9 includes the following recitation of an additional abstract idea: the planning language is selected from the group of a Planning Domain Definition Language (PDDL), a Stanford Research Institute Problem Solver (STRIPS), a Statistical Analysis Software (SAS+), and an Action Description Language (ADL) (This is practical to perform in the human mind under its broadest reasonable interpretation aside from the recitation of generic computer components.) Claim 9 recites the following additional elements which, considered individually and as an ordered combination, do not integrate the abstract idea into a practical application or amount to significantly more than the abstract idea: the receiving of the Al model described in the planning language (This is insignificant extra-solution activity. See MPEP 2106.05(g). Moreover, sending or receiving data is well-understood, routine, conventional as evidenced by the court cases cited at MPEP 2106.05(d), example i. Receiving or transmitting data.) Claim 9 does not reflect an improvement to computer technology or any other technology. The claim as a whole, looking at the additional elements individually and in combination, does not integrate the judicial exception into a practical application and does not amount to significantly more than the identified judicial exception. Claim 10 recites at least the abstract idea identified above in the claim upon which it depends. Claim 10 includes the following recitation of an additional abstract idea: producing a policy function and a probability distribution over RL environment actions per RL environment state (This is practical to perform in the human mind under its broadest reasonable interpretation aside from the recitation of generic computer components. The policy function and probability distribution could also fall under being a mathematical concept.) defining options for the RL environment based on the operators in the planning task (This is practical to perform in the human mind under its broadest reasonable interpretation aside from the recitation of generic computer components.) defining an initiation set of an option by a set of states of the RL environment that is mapped (L) to states satisfying the precondition of an action operator (This is practical to perform in the human mind under its broadest reasonable interpretation aside from the recitation of generic computer components.) defining the termination set of an option by the set of states of the RL environments that are mapped (L) to states satisfying the effects of the action operator (This is practical to perform in the human mind under its broadest reasonable interpretation aside from the recitation of generic computer components.) Claim 10 does not reflect an improvement to computer technology or any other technology. The claim as a whole, looking at the additional elements individually and in combination, does not integrate the judicial exception into a practical application and does not amount to significantly more than the identified judicial exception. Claim 11 recites at least the abstract idea identified above in the claim upon which it depends. Claim 11 includes the following recitation of an additional abstract idea: generating a sequence of options using an Al planner from a state of the RL environment (This is practical to perform in the human mind under its broadest reasonable interpretation aside from the recitation of generic computer components.) Claim 11 recites the following additional elements which, considered individually and as an ordered combination, do not integrate the abstract idea into a practical application or amount to significantly more than the abstract idea: obtaining the initial state of the planning task by mapping (L) from the RL environment state (This is insignificant extra-solution activity. See MPEP 2106.05(g). Moreover, sending or receiving data is well-understood, routine, conventional as evidenced by the court cases cited at MPEP 2106.05(d), example i. Receiving or transmitting data.) applying planning algorithms to generate a sequence of action operators that lead from an initial planning state to a planning goal ((This falls under mere instructions to apply the planning algorithms to get results. See MPEP 2106.05(f).) Claim 11 does not reflect an improvement to computer technology or any other technology. The claim as a whole, looking at the additional elements individually and in combination, does not integrate the judicial exception into a practical application and does not amount to significantly more than the identified judicial exception. Claim 12 recites at least the abstract idea identified above in the claim upon which it depends. Claim 12 recites the following additional elements which, considered individually and as an ordered combination, do not integrate the abstract idea into a practical application or amount to significantly more than the abstract idea: the producing of the policy function and the probability distribution over options per the RL environment state is performed by using a reinforcement learning algorithm (This is an attempt to limit the abstract idea to a particular field of use or technological environment, which does not integrate the abstract idea into a practical application or amount to significantly more than the abstract idea. See MPEP 2106.05(h).) Claim 12 does not reflect an improvement to computer technology or any other technology. The claim as a whole, looking at the additional elements individually and in combination, does not integrate the judicial exception into a practical application and does not amount to significantly more than the identified judicial exception. Claim 13 recites at least the abstract idea identified above in the claim upon which it depends. Claim 13 includes the following recitation of an additional abstract idea: the policy function comprises a set of option policy functions (This is practical to perform in the human mind under its broadest reasonable interpretation aside from the recitation of generic computer components.) Claim 13 does not reflect an improvement to computer technology or any other technology. Claim 13 recites no further additional elements which, considered individually and as an ordered combination with the additional elements from the claim upon which it depends, that integrate the abstract idea into a practical application or amount to significantly more than the abstract ideas. Claim 14 recites at least the abstract idea identified above in Claim 1. Claim 14 recites substantially similar subject matter to Claim 1 except it is a device performing the method instead of the method itself, respectively, and is rejected with the same rationale, mutatis mutandis. Claim 14 recites the following additional elements aside from those described which, considered individually and as an ordered combination with the additional elements from the claim upon which it depends, do not integrate the abstract idea into a practical application or amount to significantly more than the abstract idea: the device comprising: a processor; a memory coupled to the processor, the memory storing instructions to cause the processor to perform acts (This is a high-level recitation of generic computer components for performing the abstract idea. This does not integrate the abstract idea into a practical application or amount to significantly more than the abstract idea and falls under mere instructions to apply. See MPEP 2106.05(f).) Claim 14 does not reflect an improvement to computer technology or any other technology. The claim as a whole, looking at the additional elements individually and in combination, does not integrate the judicial exception into a practical application and does not amount to significantly more than the identified judicial exception. Claim 16 recites at least the abstract idea identified above in the claim upon which it depends. Claim 16 recites substantially similar subject matter to Claims 4 and 5, respectively, and are rejected with the same rationale, mutatis mutandis. Claim 16 does not reflect an improvement to computer technology or any other technology. The claim as a whole, looking at the additional elements individually and in combination, does not integrate the judicial exception into a practical application and does not amount to significantly more than the identified judicial exception. Claim 17 recites at least the abstract idea identified above in the claim upon which it depends. Claim 17 recites substantially similar subject matter to Claim 6, respectively, and are rejected with the same rationale, mutatis mutandis. Claim 17 does not reflect an improvement to computer technology or any other technology. The claim as a whole, looking at the additional elements individually and in combination, does not integrate the judicial exception into a practical application and does not amount to significantly more than the identified judicial exception. Claim 18 recites at least the abstract idea identified above in the claim upon which it depends. Claim 18 recites substantially similar subject matter to Claim 9, respectively, and are rejected with the same rationale, mutatis mutandis. Claim 18 does not reflect an improvement to computer technology or any other technology. The claim as a whole, looking at the additional elements individually and in combination, does not integrate the judicial exception into a practical application and does not amount to significantly more than the identified judicial exception. Claim 19 recites at least the abstract idea identified above in the claim upon which it depends. Claim 19 recites substantially similar subject matter to Claim 10, respectively, and are rejected with the same rationale, mutatis mutandis. Claim 19 does not reflect an improvement to computer technology or any other technology. The claim as a whole, looking at the additional elements individually and in combination, does not integrate the judicial exception into a practical application and does not amount to significantly more than the identified judicial exception. Claim 20 recites at least the abstract idea identified above in Claim 1. Claim 20 recites substantially similar subject matter to Claim 1 except it is a storage medium that stores the method instead of being the method itself, respectively, and is rejected with the same rationale, mutatis mutandis. Claim 20 recites the following additional elements aside from those described which, considered individually and as an ordered combination with the additional elements from the claim upon which it depends, do not integrate the abstract idea into a practical application or amount to significantly more than the abstract idea: a non-transitory computer readable storage medium tangibly embodying a computer readable program code having computer readable instructions that, when executed, causes a computer device to carry out a method (This is a high-level recitation of generic computer components for performing the abstract idea. This does not integrate the abstract idea into a practical application or amount to significantly more than the abstract idea and falls under mere instructions to apply. See MPEP 2106.05(f).) Claim 20 does not reflect an improvement to computer technology or any other technology. The claim as a whole, looking at the additional elements individually and in combination, does not integrate the judicial exception into a practical application and does not amount to significantly more than the identified judicial exception. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 1, 3, 8, 9, 14, 18, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over “Konidaris” (“From Skills to Symbols: Learning Symbolic Representations for Abstract High-Level Planning”, Published 2018) in view of “Leonetti” (“A Synthesis of Automated Planning and Reinforcement Learning for Efficient, Robust Decision-Making”, Published 2016). Regarding Claim 1, Konidaris teaches A computer-implemented method of integrating an Artificial Intelligence (AI) planner and a reinforcement learning (RL) agent through Al planning annotation in RL (PaRL) (Konidaris, Abstract recites “We consider the problem of constructing abstract representations for planning in high-dimensional, continuous environments. We assume an agent equipped with a collection of high-level actions, and construct representations provably capable of evaluating plans composed of sequences of those actions..” The examiner interprets constructing abstract representations as relating to AI Planning and the agent equipped with a collection of high-level actions as relating to reinforcement learning. In addition, this limitation is fulfilled when subsequent limitations are met as this is understood to be the preamble summarizing the following limitations.) receiving, by a Hierarchical Reinforcement Learning (HRL) module of a particularly configured computing device, a Markov Decision Process (MDP) (Konidaris Section 2 and Section 3, “Typically, the agent is concerned with maximizing R in a single MDP. However, we are interested in the multi-task reinforcement learning setting (Wilson, Fern, Ray, & Tadepalli, 2007; Brunskill & Li, 2013) where the agent must optimize expected reward over a distribution over tasks. Specifically, we are interested in changing goals over a fixed environment; we model the environment as a base MDP that specifies the environmental dynamics and a background reward function (typically representing action costs). The agent must solve multiple tasks in this environment, each of which is obtained by adding a set of task goal states, which are high-reward and in which execution terminates.”; see also Section 5.1 which discusses implementation of Konidaris’s process on a robot named Anathema Device where computing was performed partially onboard and partially on a workstation connected by a wireless network. These devices are computing devices.); identifying, by the computing device, an RL problem over the MDP (Konidaris, Abstract recites “We consider the problem of constructing abstract representations for planning in high dimensional, continuous environments. We assume an agent equipped with a collection of high-level actions…” and Section 2 and Section 3, “Typically, the agent is concerned with maximizing R in a single MDP. However, we are interested in the multi-task reinforcement learning setting (Wilson, Fern, Ray, & Tadepalli, 2007; Brunskill & Li, 2013) where the agent must optimize expected reward over a distribution over tasks.” Examiner interprets the agent equipped with high level actions as tackling a hierarchical reinforcement learning problem.) receiving, by the computing device a description of the Markov decision process (MDP) having a plurality of states in an RL environment to generate an RL task to solve the RL problem (Konidaris, 2 Background and Setting recites “We adopt the Markov-decision process (MDP) formalism of agent decision-making, which models the agent’s low-level state and action spaces with a tuple.” The agent receives a Markov decision process description of the environment in order to tackle the problem. See also Section 5.1 which discusses implementation of Konidaris’s process on a robot named Anathema Device where computing was performed partially onboard and partially on a workstation connected by a wireless network. These devices reads on the computing device.) receiving, by an AI planning module of the computing device, an Al planning model described in a symbolic planning language (Konidaris, 3.2 and 5.2 Constructing a STRIPS-Like PDDL Domain Description recites “We formulate our model as a set-theoretic high-level domain specification expressed using PDDL…which is the input format for most off-the-shelf general purpose planners.” Examiner interprets a STRIPS-Like PDDL Domain Description as a planning language. It is stated this language is the ‘input language’ which would mean the agent’s planner would receive this description. See also Section 5.1 which discusses implementation of Konidaris’s process on a robot named Anathema Device where computing was performed partially onboard and partially on a workstation connected by a wireless network. These devices reads on the computing device implementation instructions of AI planning.) annotating the RL task with an Al planning task from the mapping to generate a PaRL task (Konidaris, 2.3 and Section 3.2.1 The Semantics of Symbolic Representations states “A symbolic representation –or model– of a problem therefore consists of a vocabulary of symbols, plus a collection of operators that specific rules for manipulating those symbols to describe the effects of actions in the world.” Examiner interprets the use of a vocabulary and different operators as annotating the task to generate the PaRL task as identified by the instant claim.) training, by the reinforcement learning module, the RL agent using the PaRL task, wherein the PaRL task includes symbolic options (Konidaris Section 5.2 Learning a Symbolic Representation, see discussions on option initiation and option mask) derived from the AI planning task (Konidaris Section 5, “The preceding section has demonstrated that our framework enables an agent to learn a symbolic representation using only data obtained through interaction with its environment, and thereby compute plans very quickly using an off-the-shelf probabilistic high-level planner. That demonstration took place in a relatively simple computer game environment, but we are motivated primarily by achieving abstract planning on real robots. We now show that our framework can achieve that goal, by describing a robot manipulation system that acquires a symbolic representation directly from low-level sensor outputs-point clouds, map locations, and joint positions-and uses it to rapidly plan to achieve its goals.”); and solving the identified RL problem using the generated PaRL task (Konidaris, 1 Introduction states “We construct an agent that autonomously learns the correct abstract representation of a computer game domain and rapidly solves it.” Examiner interpretates the abstract representation as a generated PaRL task as claimed.) Konidaris does not appear to explicitly teach mapping state spaces from the MDP states in the RL environment to Al planning states of the Al planning model (specifically the exact details on the mapping) However, Leonetti, directed to analogous art (focusing on planning with reinforcement learning), teaches mapping, by a reinforcement learning module of the computing device, state spaces from the MDP states in the RL environment to Al planning states of the Al planning model (Leonetti, Section 4 and 4.3 Method definition recites “We assume that the domain can be modeled as an MDP D = (S, A, f, r), where S can be either discrete or continuous, A is finite, and f and r are stochastic and unknown…Furthermore, we expect the designer to create a model of D which we will denote as Dm = (Sm, A, fm)…The model Dm is the one on which a planning algorithm would compute a plan to execute in the real domain.” Each one of the variables that make up Dm are derived from the MDP. Sm has a specific mapping as recited by Leonetti, 4. Method definition, “We also require the designer to specify a mapping o : S ➔ Sm which for every state in S returns a corresponding state in Sm.”. ‘A’ being a direct mapping as it uses the same variable found din the MDP. ‘fm’ is derived from the Sm and A as seen in Leonetti, 2.3 Planning in answer set programming, “fm : Sm x A ➔ Sm is the transition function.” Examiner thus concludes that all elements of the Dm model that computes plans uses only elements derived from the MDP domain. See also Abstract, “Domain Approximation for Reinforcement LearnING (DARLING), a method that takes advantage of planning to constrain the behavior of the agent to reasonable choices, and of reinforcement learning to adapt to the environment”). It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Konidaris in view of Leonetti to teach specifics of how to map between the different steps of forming a plan. The motivation being directly from Leonetti. Leonetti, 1. Introduction states “The reason why plans are brittle can ultimately be attributed to imperfections in the models: relevant details overlooked, dynamics incorrectly represented, or assumptions violated. Nonetheless, making decisions in domains of any interest unavoidably involves abstraction and approximation, causing imperfect models to be widespread.” Said person after learning of Leonetti would try to integrate the art within Konidaris to strengthen the plans made by Konidaris. One main difference between the two arts is that Konidaris does not appear to directly map the MDP environment to the planning task. It would be obvious for said person that this potentially makes the plan weak and for said person to specifically use Leonetti’s way of mapping in order to obtain the art’s benefit of stronger plans. Regarding Claim 3, the rejection of Claim 1 is incorporated herein. Konidaris teaches formulating by the PaRL task an options framework for the MDP (Konidaris, 2.1 Hierarchical Reinforcement Learning recites “The options framework…is a hierarchical reinforcement learning framework that models high-level actions as options.” The ‘high-level actions’ as stated would be the various PaRL tasks as stated by the instant claim. Konidaris, 2.1 Hierarchical Reinforcement Learning continues “Replacing an agent’s low-level actions with higher-level options in this way results in a semi-Markov decision process (SMDP).” This options framework operates on and for the lower level MDP environment. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Konidaris in view of Leonetti as described above with respect to Claim 1. Regarding Claim 8, the rejection of Claim 1 is incorporated herein. Konidaris does not appear to explicitly teach the annotating of the RL task with the Al planning task from the mapping to generate the PaRL task comprises at least one mapping selected from the group of: abstraction mapping in Al planning, heuristic mapping between state spaces, and rule-based mapping However, Leonetti, directed to analogous art (teaching an options framework for an RL agent having an MDP environment), teaches the annotating of the RL task with the Al planning task from the mapping to generate the PaRL task comprises at least one mapping selected from the group of: abstraction mapping in Al planning, heuristic mapping between state spaces, and rule-based mapping (Leonetti, 4. Method definition recites “Through our method we will define a new MDP Dr (where the subscript stands for reduced), computed by planning over Dm, that models the same domain as D but is a reduced version of it.” Examiner interprets reducing the plan Dm to Dr as annotating the plan. This results in Dr which is the PaRL task as stated by the instant claim. Leonetti, 2.2 Answer set programing recites “A set of rules forms an ASP theory. A model of an ASP theory is an Answer Set, that is a set of atoms compatible with the theory.” It would be logical to assume that the mapping from the MDP to the planning language uses the rule set of the planning language while mapping. Using rule-based mapping is one of the options within the instant claim, thus satisfying the limitation.) It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Konidaris in view of Leonetti as described above with respect to Claim 1. Regarding Claim 9, the rejection of Claim 1 is incorporated herein. Konidaris teaches the receiving of the Al model described in the planning language is selected from the group of a Planning Domain Definition Language (PDDL), a Stanford Research Institute Problem Solver (STRIPS), a Statistical Analysis Software (SAS+), and an Action Description Language (ADL) (Konidaris, 3.2 Constructing a STRIPS-Like PDDL Domain Description recites “We formulate our model as a set-theoretic high-level domain specification expressed using PDDL, the Planning and Domain Definition Language (McDermott et al., 1998), which is the input format for most off-the-shelf general purpose planners.” The model’s input format is being expressed using PDDL which is one of the selections of the instant claim, thus satisfying the limitation.) It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Konidaris in view of Leonetti as described above with respect to Claim 1. Claim 14 recites substantially similar subject matter to Claim 1 except it is a device performing the method instead of the method itself, respectively, and is rejected with the same rationale, mutatis mutandis. Additionally, Konidaris teaches the device comprising: a processor; a memory coupled to the processor, the memory storing instructions to cause the processor to perform acts (Konidaris, Abstract recites “We assume an agent equipped with a collection of high-level actions, and construct representations provably capable of evaluating plans composed of sequences of those actions.” The broadest reasonable interpretation of an agent capable of evaluating plans would entail a device with appropriate hardware like a processor and memory to perform.) Claim 18 recites substantially similar subject matter to Claim 9, respectively, and are rejected with the same rationale, mutatis mutandis. Claim 20 recites substantially similar subject matter to Claim 1 except it is a storage medium that stores the method instead of being the method itself, respectively, and is rejected with the same rationale, mutatis mutandis. Additionally, Konidaris teaches a non-transitory computer readable storage medium tangibly embodying a computer readable program code having computer readable instructions that, when executed, causes a computer device to carry out a method (Konidaris, Abstract recites “We assume an agent equipped with a collection of high-level actions, and construct representations provably capable of evaluating plans composed of sequences of those actions.” The broadest reasonable interpretation of an agent capable of evaluating plans would entail a device with appropriate hardware such as computer readable memory storing the instructions to perform the disclosed steps.) Claims 4-7, 16, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over “Konidaris” (“From Skills to Symbols: Learning Symbolic Representations for Abstract High-Level Planning”, Published 2018) in view of “Leonetti” (“A Synthesis of Automated Planning and Reinforcement Learning for Efficient, Robust Decision-Making”, Published 2016) and “Sutton” (“Between MDPs and Semi-MDPs: A Framework for Temporal Abstraction in Reinforcement Learning”, Published 1999). Regarding Claim 4, the rejection of Claim 3 is incorporated herein. Konidaris teaches generating one or more sets of Al plans in the options framework (Konidaris, 5 Learning a Symbolic Representation of a Robot Manipulation Task recites “…thereby compute plans very quickly using an off-the-shelf probabilistic high-level planner.” Konidaris, Abstract recites “We assume an agent equipped with a collection of high-level actions, and construct representations provably capable of evaluating plans composed of sequences of those actions.” The agent uses a planner to generate a plurality of plans. A plan is later defined in Konidaris, Definition 2 as being “…a sequence of options to be executed from some state in Z.”) sending the options to the RL agent (Konidaris, 1 Introduction recites “An agent should therefore solve problems and generate behavior by identifying goals and then constructing plans composed of high-level skills in order to reach them.” A plan is later defined in Konidaris, Definition 2 as being “…a sequence of options to be executed from some state in Z.” The agent is constructing the options itself and therefore would obtain the options naturally.) Konidaris does not appear to explicitly teach selecting options from the options framework for training the RL agent by ranking the options with scores However, Sutton, directed to analogous art (teaching an options framework for an RL agent having an MDP environment), teaches selecting options from the options framework for training the RL agent by ranking the options with scores (Sutton, 1. The reinforcement learning (MDP) framework defines an action-value function, “This is known as the action-value function for policy.” Within the same section, it also defines “The optimal action-value function”. This function takes the maximum value. The action-value function for options is a type of scoring function. Due to every option being scored this way, a natural ranking occurs. Examiner interpretates the optimal value function as to requiring the options to be ranked in implementation. Examiner interpretates taking the maximum value only as a form of selecting an option.) It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Konidaris in view of Leonetti as described above with respect to Claim 1 and in view of Sutton to teach the implementation. The motivation to incorporate Sutton would be to recreate Konidaris. Konidaris mentions Sutton when first discussing options within 2.1 Hierarchical Reinforcement Learning, “The options framework (Sutton et al., 1999) is a hierarchical reinforcement learning framework that models high-level actions as options.” Konidaris continues to state “Our goal is to show how the addition of options to an MDP allows us to drastically reduce its state space in addition to its action space, resulting in an abstract representation that affords high-level planning.” Examiner interpretates Konidaris as building onto the work of Sutton. In order to replicate Konidaris’ work in its entirety, a person of ordinary skill in the art before the effective filing date of the claimed invention would be led to incorporate Sutton’s work. The modification in view of Leonetti would not remove or alter the options framework by Konidaris as Leonetti is incorporated to fill the gap how implementation would happen not change existing components. Regarding Claim 5, the rejection of Claim 4 is incorporated herein. Konidaris teaches the selecting options from the options framework is performed online or offline (Konidaris, 2.1 Hierarchical Reinforcement Learning recites “The options framework (Sutton et al., 1999) is a hierarchical reinforcement learning framework that models high-level actions as options.” Konidaris does not specific whether the options framework is online or offline. Konidaris, 1 Introduction recites “An agent should therefore solve problems and generate behavior by identifying goals and then constructing plans composed of high-level skills in order to reach them.” A plan is later defined in Konidaris, Definition 2 as being “…a sequence of options to be executed from some state in Z.” In context, Konidaris appears to be entirely offline or at least the agent has all the necessary components itself to tackle the problem. The options framework being offline satisfies the instant claim as being one of the choices.) It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Konidaris in view of Leonetti and Sutton as described above with respect to Claim 4. Regarding Claim 6, the rejection of Claim 5 is incorporated herein. Konidaris teaches generating a plan given trajectory (Konidaris, 2.1 Hierarchical RL recites “We will not assume that the agent has access to option models; instead, it must gather data about each option's performance and execution characteristics by actually executing it in the world.” The set of options that a plan is composed of, change based on the performance of past options. Thus, an overall plan is made using the trajectory of how the past plan performed.) sending the options with a highest score to the RL agent (Konidaris, 1 Introduction states “An agent should therefore solve problems and generate behavior by identifying goals and then constructing plans composed of high-level skills in order to reach them.” The agent generates all the options itself. The highest scored option is guaranteed to be with the agent.) Konidaris does not appear to explicitly teach ranking options according to a scoring function However, Sutton, directed to analogous art (teaching an options framework for an RL agent having an MDP environment), teaches ranking options according to a scoring function (Sutton, 1. The reinforcement learning (MDP) framework defines an action-value function, “This is known as the action-value function for policy.” Within the same section, it also defines “The optimal action-value function”. This function takes the maximum value. The action-value function for options is a type of scoring function. Due to every option being scored this way, a natural ranking occurs. Examiner interpretates the optimal value function as to requiring the options to be ranked in implementation.) It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Konidaris in view of Leonetti and Sutton as described above with respect to Claim 4. Regarding Claim 7, the rejection of Claim 6 is incorporated herein. Konidaris teaches sending the options to the RL agent further comprises guiding a sampling process by a PaRL planner to sample the options (Konidaris, 2.1 Hierarchical RL recites “One way to do so would be to use option execution experience to learn the relevant option models themselves. An agent possessing such models is capable of sample-based SMDP planning (Sutton et al., 1999; Kocsis & Szepesvari, 2006).”) It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Konidaris in view of Leonetti and Sutton as described above with respect to Claim 4. Claim 16 recites substantially similar subject matter to Claims 4 and 5, respectively, and are rejected with the same rationale, mutatis mutandis. Claim 17 recites substantially similar subject matter to Claim 6, respectively, and are rejected with the same rationale, mutatis mutandis. Claims 10-13 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over “Konidaris” (“From Skills to Symbols: Learning Symbolic Representations for Abstract High-Level Planning”, Published 2018) in view of “Leonetti” (“A Synthesis of Automated Planning and Reinforcement Learning for Efficient, Robust Decision-Making”, Published 2016) and “Moerland” (“Model-Based Reinforcement Learning: A Survey”, Published 2021). Regarding Claim 10, the rejection of Claim 1 is incorporated herein. Konidaris teaches producing a policy function and a probability distribution over RL environment actions per RL environment state (Konidaris, 4.1 Symbols for Probabilistic Planning recites “Distributional symbols are used to express a distribution over start states and the probabilistic image operator which, given a start state distribution and assuming successful option execution, returns the distribution over states in which the agent may find itself.” Examiner interprets the probabilistic image operator as a policy function and the output or the distribution over states as the probability distribution over the environment actions.) defining options for the RL environment based on the operators in the planning task (Konidaris, 3.1 Symbols for Deterministic Planning recites “We also define an operator expressing the consequences of executing an option from one of a set of states” Examiner interprets this to mean the operators are already defined at the time an option is created. Thus, an option is created with the operators involved.) defining an initiation set of an option by a set of states of the RL environment that is mapped (L) to states satisfying the precondition of an action operator (Konidaris, 2.1 Hierarchical Reinforcement Learning recites “an initiation set, which describes the low-level states in which the option may be executed.” Examiner interprets the ‘states in which the option may be executed’ as states that satisfy the precondition of an action operator. The low-level states are in relation to the RL environment that the agent is in.) Konidaris does not appear to explicitly teach defining the termination set of an option by the set of states of the RL environments that are mapped (L) to states satisfying the effects of the action operator However, Moerland, directed to analogous art (a survey on planning-based reinforcement learning models), teaches defining the termination set of an option by the set of states of the RL environments that are mapped by L to states satisfying the effects of the action operator (Moerland, 4.8 Temporal abstraction states “Options and goal-conditioned value functions are conceptually different. Most importantly, options have a separate sub-policy per option, while GCVFs attempt to generalize over goals/subpolicies. Moreover, options fix the initiation and termination set based on state information.” Examiner interprets the broadest reasonable interpretation in the art to be of the termination set to be like that of an initiation set yet that are mapped to states that are after the operators have been applied.) It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Konidaris in view of Leonetti as described above with respect to Claim 1 and in view of Moerland to assist with the implementation. The motivation to incorporate Moerland is for understanding. Konidaris already defines the ingredients for making a termination set in 2.1 Hierarchical Reinforcement Learning, “An option o consists of three components: an option policy, which is executed when the option is invoked, and maps low-level states to low-level actions; an initiation set, which describes the low-level states in which the option may be executed; and a termination condition, which describes the probability that an option will terminate upon reaching low-level state s.” A termination set would be the set of termination conditions for each state. Thus, said person would gain a better understanding of the consequences of an option having a full termination set of the initiation set. Konidaris mentions Sutton when first discussing options within 2.1 Hierarchical Reinforcement Learning, “The options framework (Sutton et al., 1999) is a hierarchical reinforcement learning framework that models high-level actions as options.” Konidaris continues to state “Our goal is to show how the addition of options to an MDP allows us to drastically reduce its state space in addition to its action space, resulting in an abstract representation that affords high-level planning.” Examiner interpretates Konidaris as building onto the work of Sutton. In order to replicate Konidaris’ work in its entirety, a person of ordinary skill in the art before the effective filing date of the claimed invention would be led to incorporate Sutton’s work. The modification in view of Leonetti would not remove or alter the options framework by Konidaris as Leonetti is incorporated to fill the gap how implementation would happen not change existing components. Regarding Claim 11, the rejection of Claim 10 is incorporated herein. Konidaris teaches generating a sequence of options using an Al planner from a state of the RL environment (Konidaris, 5 Learning a Symbolic Representation of a Robot Manipulation Task recites “…thereby compute plans very quickly using an off-the-shelf probabilistic high-level planner.” Konidaris, Definition 2 states “A plan p from a state set is a sequence of options to be executed from some state in Z.” The planner used by the agent generates a plan comprising of multiple options. Konidaris, 3.1.2 Abstract Subgoal Options recites “The key assumption of subgoal options is that they drive all aspects of the low-level state to the target subgoal.” These options from the low-level states, and thus the planner would logically use the states to generate a plan.) obtaining the initial state of the planning task by mapping (L) from the RL environment state (Konidaris, Definition 3 states “Both definitions use a set of start states, Z, rather than a single start state, s0. We do this for consistency, as will become clear shortly, although of course it is possible to plan from a single start state by setting Z={s0}.” Examiner interpretates this to mean that no matter how many start states there is within plan, the initial state of the task can be found by checking s0 or the first item in Z.) applying planning algorithms to generate a sequence of action operators that lead from an initial planning state to a planning goal (Konidaris, 3.1 Symbols for Deterministic Planning recites “We also define an operator expressing the consequences of executing an option from one of a set of states.” Konidaris, 3.2 Constructing a STRIPS-Like PDDL Domain Description recites “We formulate our model as a set-theoretic high-level domain specification expressed using PDDL…a set-theoretic specification consists of…a set of operators A={a1,…,am}.” and “Each operator describes the circumstances under which an option can be executed, and the resulting effect, though there may be multiple operators for each option. We assume that all options have abstract subgoals.” The model is formulated in a planning way which includes a set of operators that lead from an initial state to potential states.) It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Konidaris in view of Leonetti and Moerland as described above with respect to Claim 10. Regarding Claim 12, the rejection of Claim 11 is incorporated herein. Konidaris teaches the producing of the policy function and the probability distribution over options per the RL environment state is performed by using a reinforcement learning algorithm (Konidaris, 4.1 Symbols for Probabilistic Planning recites “Distributional symbols are used to express a distribution over start states and the probabilistic image operator which, given a start state distribution and assuming successful option execution, returns the distribution over states in which the agent may find itself.” Konidaris, Definition 23 states “The probabilistic image operator takes as input a start distribution and an option, and returns a distribution over states (i.e., a function of s).” Examiner interprets the probabilistic image operator as a policy function and the output or the distribution over states as the probability distribution over the environment actions. It is under the broadest reasonable interpretation to assume that the agent itself is performing these limitations.) It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Konidaris in view of Leonetti and Moerland as described above with respect to Claim 10. Regarding Claim 13, the rejection of Claim 12 is incorporated herein. Konidaris teaches the policy function comprises a set of option policy functions (Konidaris, 2.1 Hierarchical Reinforcement Learning recites “The options…that models high-level actions as options. An option o consists of three components: an option policy which is executed when the option is invoked.” The option which contains the policy is feed into the policy function in Konidaris, Definition 23, “The probabilistic image operator takes as input a start distribution and an option, and returns a distribution over states (i.e., a function of s).” Konidaris, Theorem 1 recites “Given an SMDP, the ability to represent the precondition of each option and to compute the image operator is sufficient for determining whether any plan tuple (Z, p) is feasible.” The policy is applied many times on each option within a plan, and thus many option policies would be included.) It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Konidaris in view of Leonetti and Moerland as described above with respect to Claim 10. Claim 19 recites substantially similar subject matter to Claim 10, respectively, and are rejected with the same rationale, mutatis mutandis. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Dicong Qiu et al, "Deep Reinforcement Learning with Abstract High-level Symbolic Planning" describes a method to combine abstracted high-level symbolic planning and low-level reinforcement learning and control. The method includes a breadth-first search planner to break down a complex task into easier-to-achieve sub-tasks, creating a CNN state monitor to detect changes in symbolic states, and implementing a scalable reinforcement learning agent executioner. THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to LI B ZHEN whose telephone number is (571)272-3768. The examiner can normally be reached M-F, 7:30a-4p. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Li B. Zhen/Supervisory Patent Examiner, Art Unit 2121
Read full office action

Prosecution Timeline

Dec 08, 2021
Application Filed
Mar 19, 2025
Non-Final Rejection mailed — §101, §103
May 23, 2025
Response Filed
Aug 26, 2026
Final Rejection mailed — §101, §103
Sep 16, 2026
Interview Requested

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 11928609
NODE SHARING FOR A RULE ENGINE CODED IN A COMPILED LANGUAGE
3y 3m to grant Granted Mar 12, 2024
Patent 10963807
SOCIAL COLLABORATION IN PROBABILISTIC PREDICTION
3y 6m to grant Granted Mar 30, 2021
Patent 10586147
NEUROMORPHIC COMPUTING DEVICE, MEMORY DEVICE, SYSTEM, AND METHOD TO MAINTAIN A SPIKE HISTORY FOR NEURONS IN A NEUROMORPHIC COMPUTING ENVIRONMENT
3y 5m to grant Granted Mar 10, 2020
Patent 9753713
Coordinated Upgrades In Distributed Systems
6y 10m to grant Granted Sep 05, 2017
Patent 8528003
COMMUNICATION AMONG BROWSER WINDOWS
9y 8m to grant Granted Sep 03, 2013
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
54%
Grant Probability
94%
With Interview (+39.9%)
5y 1m (~4m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 168 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month