Prosecution Insights
Last updated: October 02, 2026
Application No. 18/425,358

GUIDED EXPLORATION METHOD FOR REINFORCEMENT LEARNING TRAINING

Non-Final OA §101§103§112
Filed
Jan 29, 2024
Examiner
VONG, HAO THIEN
Art Unit
Tech Center
Assignee
Dell Products L.P.
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Office Action

§101 §103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after January 29, 2024, is being examined under the first inventor to file provisions of the AIA Status of Claims The present application is being examined under the claims filed on 01/29/2024. Claims 1-20 are rejected. Claims 1-20 are pending Specification The specification filed on January 29, 2024 is acceptable for examination purposes . Drawings The drawings filed on January 29, 2024 is acceptable for examination purposes. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claim 2, 6, 13, 16 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Claim 1 is method claim. Therefore, claims 1-20 are directed to either a process, machine, manufacture, or composition of matter Regarding claim 2, Step 2A Prong 1: wherein the performance quality is comparing a generalization of an optimizer to a generality of the reinforcement learning agent. Mental process: Defined as concepts that can practically be performed in the human mind, or by a human using pen and paper as a physical aid. Examples of mental processes include observation, evaluations, judgements, and opinions. Step 2A Prong 2: The judicial exceptions are not integrated into a practical application Additional element: wherein the performance quality is comparing a generalization of an optimizer to a generality of the reinforcement learning agent. The additional element merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception: Additional element: wherein the performance quality is comparing a generalization of an optimizer to a generality of the reinforcement learning agent. The additional element merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f). For the reasons above, claim 2 is rejected as being directed to an abstract idea without significantly more. Regarding claim 6, Step 2A Prong 1: wherein the metrics include: a first metric including a comparative evaluation of the reinforcement learning agent and a heuristic; a second metric including a processing time to complete the training epoch; and a third metric including resource utilization. Mental process: Defined as concepts that can practically be performed in the human mind, or by a human using pen and paper as a physical aid. Examples of mental processes include observation, evaluations, judgements, and opinions. Step 2A Prong 2: The judicial exceptions are not integrated into a practical application Additional element: wherein the metrics include: a first metric including a comparative evaluation of the reinforcement learning agent and a heuristic. Additional element adds more than activities that are incidental to the primary process or product or merely a nominal or tangential addition to the claim. See MPEP 2106.05(g). Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception: Additional element: wherein the metrics include: a first metric including a comparative evaluation of the reinforcement learning agent and a heuristic. Additional element adds more than activities that are incidental to the primary process or product or merely a nominal or tangential addition to the claim. See MPEP 2106.05(g). For the reasons above, claim 6 is rejected as being directed to an abstract idea without significantly more. Regarding claim 13, Step 2A Prong 1: wherein the performance quality is determined by comparing a generalization of an optimizer to a generality of the reinforcement learning agent, further comprising performing a predetermined number of training epochs to move toward the selection strategy, wherein the predetermined number of training epochs select actions like the heuristic. Mental process: Defined as concepts that can practically be performed in the human mind, or by a human using pen and paper as a physical aid. Examples of mental processes include observation, evaluations, judgements, and opinions. Step 2A Prong 2: The judicial exceptions are not integrated into a practical application Additional element: wherein the performance quality is comparing a generalization of an optimizer to a generality of the reinforcement learning agent. The additional element merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception: Additional element: wherein the performance quality is comparing a generalization of an optimizer to a generality of the reinforcement learning agent. The additional element merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f). For the reasons above, claim 13 is rejected as being directed to an abstract idea without significantly more. Regarding claim 16, Step 2A Prong 1: wherein the metrics include: a first metric including a comparative evaluation of the reinforcement learning agent and a heuristic; a second metric including a processing time to complete the training epoch; and a third metric including resource utilization. Mental process: Defined as concepts that can practically be performed in the human mind, or by a human using pen and paper as a physical aid. Examples of mental processes include observation, evaluations, judgements, and opinions. Step 2A Prong 2: The judicial exceptions are not integrated into a practical application Additional element: wherein the metrics include: a first metric including a comparative evaluation of the reinforcement learning agent and a heuristic. Additional element adds more than activities that are incidental to the primary process or product or merely a nominal or tangential addition to the claim. See MPEP 2106.05(g). Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception: Additional element: wherein the metrics include: a first metric including a comparative evaluation of the reinforcement learning agent and a heuristic. Additional element adds more than activities that are incidental to the primary process or product or merely a nominal or tangential addition to the claim. See MPEP 2106.05(g). For the reasons above, claim 16 is rejected as being directed to an abstract idea without significantly more. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claim 3, 13 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as failing to set forth the subject matter which the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the applicant regards as the invention. Regarding claims 3 and 13: The phrase “move toward the selection strategy” is indefinite because “the selection strategy” lacks of clear antecedent basis. Claim 1 and 12 recite both “a reference strategy” and “an action selection strategy”, but the phrase “selection strategy” did not recite which strategy it wanted to mention. Thus, it is unclear that the claim intend to mention “a reference strategy” or “an action selection strategy” Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 4, 5, 11, 12, 14, 15, 20 are rejected under 35 U.S.C. 103 as being unpatentable over Taylor et al. (US 20190236458 A1) in view of Tang et al. (CN 111144580 A) Regarding Claim 1, based on a performance quality of a heuristic: Taylor teaches: performing a preprocessing stage of a reinforcement learning agent being trained to perform a task: Taylor, 100 in Fig. 1A and paragraph [0084] , “The system 100 is a demonstrator data set bootstrapping engine that is configured for receiving data sets from one or more demonstrators for conducting improved pre-training of a machine learning model”. Examiner’s note: performing a preprocessing stage (i.e. pre-training) a reinforcement learning agent (i.e. of a machine learning model) trained to perform a task (i.e. is a demonstrator data set bootstrapping engine that is configured for receiving data sets from one or more demonstrators for conducting) based on a performance quality of a heuristic: Taylor, paragraph [0130], “The online confidence metric is measured via a temporal difference (TD) approach. For each action source, Applicants built a TD model to measure the confidence-based performance via experience… But if an RL agent follows the recommendation of an action from its prior knowledge (i.e., the demonstrator's action), the action source would then be the prior knowledge.” Examiner note: performance quality of a heuristic is determined (i.e. the quality of the prior knowledge determined by confidence-based performance via experience). performing an online stage of training the reinforcement learning agent”: Taylor, paragraph [0057], “DRoP (Dynamic Reuse of Prior) is an interactive method to boost Reinforcement Learning by addressing the above problems. DRoP uses temporal difference models to perform online confidence measurement on transferred knowledge." Taylor, paragraph [0017], “A combination of offline knowledge and online confidence-based performance analysis can be utilized to dynamically involve the demonstrator's knowledge”. Examiner note: performing an online stage of training the reinforcement learning agent(i.e. uses temporal difference models to perform online confidence measurement on transferred knowledge). Also it taught about combining offline and online stage that use the heuristic or demonstrator. dynamically adapting an action selection strategy for each training epoch: Taylor, 1510 in Fig 1.B and paragraph [0116], “The source selector 1510 is configured as an action selection mechanism (e.g., a switch) that selects between the actions posited by the trained classifiers corresponding to the demonstrator data…” . Taylor, paragraph [0109], “the action-decision models and their associated determinations … may shift in proportion as the machine learning model is improved over training epochs.” Examiner note: dynamically adapting…each training epoch(i.e. may shift in proportion as the machine learning model is improved over training) an action selection strategy for each training epoch (i.e. an action selection mechanism (e.g., a switch) that selects between the actions posited by the trained classifiers corresponding to the demonstrator data). Taylor does not explicitly teach: determine a reference strategy Tang teaches: determine a reference strategy: Tang, claim 1, “…using teaching data based on pre-training, simulating learning for determining the initial policy, then the training based on reinforcement learning based on the initial policy, determining the training model”. Examiner note: determine a reference strategy (i.e. determining the initial policy). It would have been obvious to a person having ordinary skill in the art at a time before the effective filing date of the application to have further modified the Taylor disclosures and teachings by performing pre-training and online confidence measurement with prior knowledge or demonstrator, which pre-training to determining the initial policy taught and suggested by Tang. Such a person would have been motivated to do so with a reasonable expectation of success to allow for pre-training to using teaching data based on pre-training, simulating learning for determining the initial policy, then the training based on reinforcement learning based on the initial policy, determining the training model in reinforcement learning agent training. (Tang, claim 1) Regarding Claim 4, The combination of Taylor and Tang teach: The method of claim 1, further comprising (preamble) adopting the reference strategy for a first step of the online stage.: Tang, claim 8, “wherein the training process comprises: performing initial policy according to the given initial state obtaining training data, and the teaching data and training data set as an empirical data pool, using empirical data pool data level reinforcement learning method based on determining the final training model. Examiner note: adopting the reference strategy for a first step of the online stage(i.e. the training process comprises: performing initial policy according to the given initial state obtaining training data…). That’s mean the first step for training process is initial policy. The reasons of obviousness have been noted in the rejection of Claim 1 above and applicable herein. Regarding Claim 5, The combination of Taylor and Tang teach: The method of claim 4, further comprising (preamble) performing an online training epoch: Taylor, paragraph [0056], “DRoP dynamically involves the demonstrator's knowledge, integrating it into the reinforcement learning agent's online learning loop to achieve efficient and robust learning.”. Examiner note: performing an online training epoch (i.e integrating it into the reinforcement learning agent's online learning loop). obtaining metrics.: Taylor, paragraph [0108], “After an action is executed…observes the outcome and associated rewards/states, and updates the machine learning model stored in model data storage”. Examiner note: obtaining metrics (i.e. observes the outcome and associated rewards/states). The reasons of obviousness have been noted in the rejection of Claim 1 above and applicable herein. Regarding claim 11, The combination of Taylor and Tang teach: The method of claim 1, further comprising (preamble) finishing the training when a budget is consumed or resources are unavailable.: Taylor, paragraph [0018], “Where it is determined that there is insufficient demonstration data to determine what to do, the system may pause and prompt for a demonstrator to provide more demonstration data (e.g., play the game).” Examiner note: finishing the training(i.e the system may pause). when a budget is consumed or resources are unavailable (i.e. there is insufficient demonstration data) The reasons of obviousness have been noted in the rejection of Claim 1 above and applicable herein. Regarding claim 12, it is rejected for similar reasons as claim 1. A non-transitory storage medium to perform the steps recited in claim 1 (Taylor’s claim 10) Regarding claim 14, it is rejected for similar reasons as claim 4. Regarding claim 15, it is rejected for similar reasons as claim 5. Regarding claim 20, it is rejected for similar reasons as claim 11. Claims 2, 13 are rejected under 35 U.S.C. 103 as being unpatentable over Taylor in view of Tang as applied in claim 1, and further in view of Paulo et al (Learning 2-Opt Heuristics for Routing Problems via Deep Reinforcement Learning) Regarding claim 2, The combination of Taylor and Tang teach: The method of claim 1 (preamble) performance quality is determined: Taylor, paragraph [0130], “The online confidence metric is measured via a temporal difference (TD) approach. For each action source, Applicants built a TD model to measure the confidence-based performance via experience… But if an RL agent follows the recommendation of an action from its prior knowledge (i.e., the demonstrator's action), the action source would then be the prior knowledge.” Examiner note: performance quality is determined (i.e. the quality of the prior knowledge determined by confidence-based performance via experience). The combination of Taylor and Tang did not comparing a generalization of an optimizer to a generality of the reinforcement learning agent: Paulo teaches: comparing a generalization of an optimizer to a generality of the reinforcement learning agent: Paulo, page 12, “Comparison to Exact and Heuristics Baselines The results for the set of 1000 instances are presented in Table 6. We observe that the learned policies are close to the performances of both Gurobi and LKH3 when solving instances with 20 nodes with 0.02%, 0.08% optimality gaps, respectively.” Paulo, page 12, “Moreover, as we increase the size of the instances the performance of Gurobi running for just 30 s decreases considerably taking significantly longer (8h) and yielding results far from LKH3.” Paulo, page 16, reference 9 ,“Gurobi optimizer reference manual” Examiner note: generalization of an optimizer (i.e. performance of Gurobi or optimizer) generality of the reinforcement learning agent (i.e. learning general policies), PNG media_image1.png 174 690 media_image1.png Greyscale It would have been obvious to a person having ordinary skill in the art at a time before the effective filing date of the application to have further modified the Taylor and Tang disclosures and teachings by determine the performance quality by comparing performance of Gurobi (or optimizer) to learning policies as taught and suggested by Paulo. Such a person would have been motivated to do so with a reasonable expectation of success to allow for pre-training Paulo performs a comparative performance among learned policies are close to the performances of both Gurobi (Paulo, page 12) Regarding claim 13, it is rejected for similar reasons as claim 2. Claims 3, 13 are rejected under 35 U.S.C. 103 as being unpatentable over Taylor in view of Tang as applied in claim 1, and further in view of Arnott et al (US 20220264331 A1) Regarding claim 3, The combination of Taylor and Tang teach: The method of claim 1, further comprising (preamble) move toward the selection strategy: Taylor, paragraph [0109], “As the machine learning model progresses, the action-decision models and their associated determinations as it relates to actor-source for actions …may shift in proportion as the machine learning model is improved over training epochs.” Examiner note: move toward the selection strategy (i.e. actor-source for actions…may shift in proportion as the machine learning model) select actions like the heuristic: Taylor, paragraph [0098], “In some embodiments, the confidence scores are utilized… which utilizes a selection function to determine an action for the machine learning model to take…” Examiner note: select actions like the heuristic (i.e. utilizes a selection function to determine an action for the machine learning model”. The combination of Taylor and Tang did not explicitly teach: predetermined number of training epochs Arnott teaches: predetermined number of training epochs: Arnott, paragraph [0092], “Training is performed in a series of ‘epochs’. An epoch consists of 390 iterations of 32 time steps each… In each iteration the following steps are performed.” Arnott, paragraph [0093], “At each time step the selected action and observed reward are stored in the experience replay memory, along with the neural network input data for the current state and observed next state”. Arnott, claim 13, “selecting at least one network optimisation action to be performed in at least one of the cellular regions is performed based on the probability ε, and wherein the probability ε gradually changes from an initial value to a final value over the plurality of learning iterations.” Examiner note: predetermined number of training epochs (i.e. Training is performed in a series of ‘epochs’) Paragraph [0093] and claim 13 explain how the epoch move to the selected action by probability ε gradually changes from an initial value to a final value over the plurality of learning iterations It would have been obvious to a person having ordinary skill in the art at a time before the effective filing date of the application to have further modified the Arnott disclosures and teachings by training is performed in a series of ‘epochs’ and move to actor-source for actions that may shift in proportion as the machine learning model is improved over training epochs, utilizes a selection function to determine an action for the machine learning model taught or discussed or suggested by Tang and Taylor. Such a person would have been motivated to do so with a reasonable expectation of success to allow training epoch move to the selected action (Arnott, par [0093] and claim 13), utilizes a selection function (Taylor, par [0098]) Regarding claim 13, it is rejected for similar reasons as claim 3. Claims 6, 16 are rejected under 35 U.S.C. 103 as being unpatentable over Taylor in view of Tang as applied in claim 5, and further in view of Calmon et al (US 20200241921 A1) Regarding claim 6, The combination of Taylor and Tang teach: The method of claim 5, wherein the metrics include (preamble) a first metric including a comparative evaluation of the reinforcement learning agent and a heuristic: Taylor, paragraph [0139], “The confidence Q knowledge model is denoted by CQ(s)”, Taylor, paragraph [0135], “The confidence prior knowledge model is denoted by CP(s).” Taylor, paragraph [0142]: “The hard decision model (HD) is greedy and attempts to maximize the current confidence expectation. Given current state s, action source AS is selected as: AS=arg max[{CQ(s),CP(s)}],” Examiner note: “comparative evaluation (i.e arg max[{CQ(s),CP(s)}]) reinforcement learning agent (i.e. The confidence Q knowledge model is denoted by CQ(s)) a heuristic (i.e. The confidence prior knowledge model is denoted by CP(s))” The combination of Taylor and Tang did not explicitly teach: second metric including a processing time to complete the training epoch a third metric including resource utilization. Calmon teaches: second metric including a processing time to complete the training epoch: Calmon, paragraph [0100], “Assuming that the SLA metric to be controlled is the execution time (et=1), the amount of time t it took to complete an epoch can be used as a feedback parameter and this time can be compared to the desired time per epoch, which is T/n.”. Examiner note: “second metric including a processing time to complete the training epoch (i.e. the amount of time t it took to complete an epoch) a third metric including resource utilization: Calmon, claim 1, “a domain model of the iterative workload that relates an amount of resources allocated in training data with one or more service metrics.”. Calmon, paragraph [0100], “If an epoch took longer than T/n to finish, more resources might me be needed. On the other hand, if the time t is significantly smaller than T/n, this indicates that the job does not need the amount of resources allocated to it and reducing the allocation can decrease cost and even make room for other jobs to run.” Examiner note: a third metric including resource utilization (i.e. iterative workload that relates an amount of resources allocated…), paragraph [0100] give an example on how to utilization resource. It would have been obvious to a person having ordinary skill in the art at a time before the effective filing date of the application to have further modified the Tang and Taylor disclosures and teachings by three types of action decision models in which first model is the hard decision model (HD) taught or suggested Tang and Taylor, and others are the amount of time t it took to complete an epoch and iterative workload that relates an amount of resources allocated, both taught and suggested by Calmon. Such a person would have been motivated to do so with a reasonable expectation of success to allow obtain metrics include: Taylor and Tang’s hard decision model, Calmon’s time to complete an epoch and an amount resource allocated. Regarding claim 16, it is rejected for similar reasons as claim 6 A non-transitory storage medium to perform the steps recited in claim 1 (Taylor’s claim 10) Claims 7, 8, 9, 17, 18 are rejected under 35 U.S.C. 103 as being unpatentable over Taylor in view of Tang, Calmon as applied in claim 6, and further in view of Lin et al ( US 20250217254 A1 ) Regarding claim 7, The combination of Taylor, Tang, Calmon teach: The method of claim 6 (preamble) the second metric and the third metric: Already rejected under 103 in claim 6. next exploitation aspect of a next action selection strategy used in a next training epoch: Taylor, paragraph [0109], “As the machine learning model progresses, the action-decision models…may shift in proportion as the machine learning model is improved over training epochs.”. Taylor, paragraph [0158], Algorithm 1 teaches Target Learning Bootstrap. Examiner note: next exploitation aspect (i.e. action from maximizes Q from Algorithm 1) next action selection strategy (i.e. the action-decision models) a next training epoch (i.e. machine learning model is improved over training epochs) PNG media_image2.png 249 461 media_image2.png Greyscale PNG media_image3.png 205 465 media_image3.png Greyscale The combination of Taylor, Tang, Calmon and did not explicitly teach: determining a step size base one step size relates to an increment Lin teaches: determining a step size base on: Lin, claim 1, “determine a step size based on the reward value, and determine a performance adjustment action to take based on the step size”. Lin, claim 2, “determines the step size through optimizing a learning rate based on the reward value and calculating the step size based on the learning rate, the target frame speed, and the actual frame speed.” Examiner note: determining a step size base on (i.e. determine a step size based on the reward value and optimizing a learning rate based on the reward value and calculating the step size based on the learning rate, the target frame speed, and the actual frame speed) step size relates to an increment: Lin, claim 7 “the agent module is further configured to calculate the new target frame speed through incrementing the target frame speed by the step size” Examiner note: step size relates to an increment (i.e. calculate the new target frame speed through incrementing the target frame speed by the step size) It would have been obvious to a person having ordinary skill in the art at a time before the effective filing date of the application to have further modified the Tang, Taylor and Clamon disclosures and teachings by obtain three metric which two of those are time to complete an epoch, resources utilization, action from maximizes Q, the action-decision models, machine learning model is improved over training epoch and how to determine a step size, step size relate to the increment of an the target frame speed (or an adjustment) taught or suggested by Lin. Such a person would have been motivated to do so with a reasonable expectation of success to allow determine step size base on time to complete an epoch (Lin’s claim 1 and 2 with Calmon’s par [0100]), step size relate to the increment of an adjustment ( Lin’s claim 7), an adjustment in here modified as action from maximizes Q of the action-decision models that is improved over training epoch (Tang’s [0109] and Algorithm 1) Regarding claim 8, The combination of Taylor, Tang, Calmon and Lin teach: The method of claim 7, further comprising (preamble) determining a next random aspect and a next heuristic aspect of the next training epoch: Taylor, Abstract, “dynamic action selection based on confidence levels associated with demonstrator data or portions thereof.”. – Taylor, paragraph [0109], “As the machine learning model progresses, the action-decision models and their associated determinations as it relates to actor-source for actions (e.g., decision to use demonstrator “knowledge” or the model's own “knowledge”) may shift in proportion as the machine learning model is improved over training epochs.”. Taylor, paragraph [0158], Algorithm 1 teaches Target Learning Bootstrap. Examiner note: a next random aspect (i.e. random action from Algorithm 1) a next heuristic aspect (i.e action front Prior Knowledge from Algorithm 1) the next training epoch(machine learning model is improved over training epochs). PNG media_image2.png 249 461 media_image2.png Greyscale PNG media_image3.png 205 465 media_image3.png Greyscale The reasons of obviousness have been noted in the rejection of Claim 7 above and applicable herein. Regarding claim 9, The combination of Taylor, Tang, Calmon and Lin teach: The method of claim 8, further comprising (preamble) changing the action strategy to a next action strategy based on the next exploitation aspect, the next random aspect, and the next heuristic aspect.: Taylor, paragraph [0109], “As the machine learning model progresses… may shift in proportion as the machine learning model is improved over training epochs”. Taylor, paragraph [0083], “FIG. 1A is a block schematic of an example system 100 for interactive reinforcement learning with dynamic reuse of prior knowledge.” Taylor, paragraph [0158], Algorithm 1. Examiner note: changing the action strategy to a next action strategy(i.e. actor-source for actions may shift in proportion as the machine learning model) The next exploitation aspect (i.e. action that maximizes Q from Algorithm 1), the next random aspect(i.e. action random Q from Algorithm 1), PNG media_image2.png 249 461 media_image2.png Greyscale next heuristic aspect(i.e. action front Prior Knowledge Q from Algorithm 1) PNG media_image3.png 205 465 media_image3.png Greyscale PNG media_image4.png 608 718 media_image4.png Greyscale The reasons of obviousness have been noted in the rejection of Claim 7 above and applicable herein. Regarding claim 17, it is rejected for similar reasons as claim 7. A non-transitory storage medium to perform the steps recited in claim 1 (Taylor’s claim 10) Regarding claim 18, it is rejected for similar reasons as claim 8 and claim 9. Claims 10, 19 are rejected under 35 U.S.C. 103 as being unpatentable over Taylor in view of Tang, Calmon and Lin as applied in claim 9, and further in view of Arnott et al (US 20220264331 A1) Regarding claim 10, The combination of Taylor, Tang, Calmon and Lin teach: The method of claim 9 (preamble) next action selection strategy (snext): Taylor, paragraph [0109], “As the machine learning model progresses, the action-decision models and their associated determinations as it relates to actor-source for actions (e.g., decision to use demonstrator “knowledge” or the model's own “knowledge”) may shift in proportion as the machine learning model is improved over training epochs.”. Examiner note: Paragraph [0109], explain the model’s confidence being updated by using the knowledge, the next action-decision models will be different from the previous. This recites the next action selection strategy v1 relates to a random aspect of the action-selection strategy Taylor, Algorithm 1 Examiner note: v1(random) (i.e. random action or exploration) v2 relates to an exploitation aspect of the action-selection strategy Taylor, Description [0130], “An action source is defined by where an agent gets its action from. That is to say, in the current state, if an RL agent chooses an action by arg max Q(s, a), the corresponding action source is its learned Q-value” Taylor, paragraph [0158], Algorithm 1 Examiner note: v2(exploitation) (i.e. action that maximizes Q) Q is Q-Value which explain action of model at state s, model choose action a, model expect the chosen action give be best reward. action maximizes Q = arg max Q(s,a) v3 relates to a heuristic aspect of the action-selection strategy. Taylor, paragraph [0158], Algorithm 1 Examiner note: v3(heuristic) (i.e. action from Prior Knoweldge) snext=(v1-az,v2-az,v3+2az), wherein a is a variation, z is a generalization score, and (v1,v2,v3) corresponds to a current strategy. This limitation recites the s(next) next action selection strategy (snext) determined by three-component transition to obtain the s(next), s(next) change when v1(random), v2(exploitation), v3(heuristic) change. Taylor, paragraph [0078], “As learning goes on, there will be a balance between the transferred knowledge and self-learned Q knowledge. That is, an action decision model will consider the confidence the agent has in all sources of knowledge and select the one most likely to yield high reward. Over time, if the transferred knowledge is sub-optimal, the self-learned Q knowledge will become selected more and more often.” Examiner note: Taylor’s paragraph [0078] explains the model over time will more and more make decision base on the self-learned Q (i.e. v2(exploitation)), that also mean that the decision less and less base on transfer knowledge (i.e. prior knowledge or v3(heuristic)) Taylor, Description paragraph [0078] teaches the action decision model depend when v3(heuristic) and v2(exploitation) change. The combination of Taylor, Tang, Calmon and Lin did not explicitly teach: v1(random) change Arnott teaches: v1(random) change: Arnott, paragraph [0099], “During training, the agent selects actions according to a modified ε-greedy policy, whereby with probability ε an action is selected uniformly at random and with probability 1-ε an action is selected based on Q(s.sub.t,a.sub.t,θ). The value of ε is linearly annealed from an initial value of 1 to a final value of 0.1 over the first 1500 training epochs. Rather than always select the action with the maximum Q(s.sub.t,a.sub.t,θ), we select action a with probability… custom-character(s.sub.t) is the set of actions allowed in state s.sub.t and α=1000. This is to encourage exploration in the case that there is more than one action with a Q-value close to the maximum.” Examiner note: Arnott’s paragraph [0099] explains that the agent selects based on three-component transition, where the value of ε is linearly annealed from an initial value of 1 to a final value of 0.1 (i.e. the agent exploration decreases) v1(random) (i.e. random action or exploration) from Taylor’s Algorithm Thus, the agent selects change when v1(random) changes. It would have been obvious to a person having ordinary skill in the art at a time before the effective filing date of the application to have further modified the Tang, Taylor, Calmon and Lin disclosures and teachings by determined the next selection decision model based on the three-component transition which two component are exploitation and heuristic(prior knowledge) change, and exploration change taught or suggested by Arnott. Such a person would have been motivated to do so with a reasonable expectation of success to allow determine next selection decision model based on the three-component transition, the components change during training to make the next selection decision give the best reward. Those components are v1(random) (Taylor’s Algorithm 1 with Arnott’s par [0099]), v2(exploitation) (Taylor’s Algorithm 1, par [0078] and par [0130]), v3(heuristic) (Taylor’s Algorithm 1, par [0078]) Regarding claim 19, it is rejected for similar reasons as claim 10 A non-transitory storage medium to perform the steps recited in claim 1 (Taylor’s claim 10) Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to HAO T VONG whose telephone number is (571)270-7701. The examiner can normally be reached Monday - Friday (8:00 AM - 6:00 AM). Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Viker A Lamardo can be reached at (571) 270-5871. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MICHAEL J HUNTLEY/Supervisory Patent Examiner, Art Unit 2129
Read full office action

Prosecution Timeline

Jan 29, 2024
Application Filed
Aug 27, 2026
Non-Final Rejection mailed — §101, §103, §112 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month