DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 01/26/2026 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-16 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea and does not integrate the judicial exception into a practical application or amount to significantly more than the judicial exception.
Regarding claim 1
Step 1: The claim recites a method; therefore, it falls into the statutory category of processes.
Step 2A Prong 1: The claim recites multiple mental processes, as explained below. The claim recites, inter alia:
“…each reward being associated with an action, the method comprising the steps of: setting a fixed length sliding window within which past actions are considered; receiving a plurality of rewards, each reward being associated with an action taken by a learning agent; defining a group of in-window rewards comprising rewards received while their associated actions are considered to be within the fixed length sliding window; and when an action performed by a learning agent is no longer considered to be within the fixed length sliding window, sending the learning agent reward information relating to all in-window rewards received for the action.”
This limitation is directed to the abstract idea of a mental process (concepts performed in the human mind, including observation and evaluation [see MPEP 2106.04(a)(2) III. C.]).
Step 2A Prong 2: This judicial exception is not integrated into a practical. In particular, the claim only recites additional elements that are mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. See MPEP 2106.05(f). The additional element of “A computer implemented method for processing rewards received during a reinforcement learning session in which one or more learning agents each perform one or more actions in a reinforcement-type machine learning environment,”, as drafted, is reciting generic computer components. The generic computer components in these steps are recited at a high-level of generality (i.e., as a generic computer component performing a generic computer function) such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into practical application, the additional element of using generic computer components to perform the abstract idea amounts to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. Thus, the claim is not patent eligible.
Regarding claim 2
Step 1: The claim recites a method; therefore, it falls into the statutory category of processes.
Step 2A Prong 1:
“wherein each in-window reward includes a value associated with a critical assessment of the action associated with the reward, and the reward information relating to all in-window rewards received for the action includes information associated to an aggregate of the values of the in-window rewards.”
This limitation is directed to the abstract idea of a mental process (concepts performed in the human mind, including observation and evaluation [see MPEP 2106.04(a)(2) III. C.]).
Step 2A Prong 2: This judicial exception is not integrated into a practical. In particular, the claim only recites additional elements that are mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. See MPEP 2106.05(f). The additional element of “computer implemented”, as drafted, is reciting generic computer components. The generic computer components in these steps are recited at a high-level of generality (i.e., as a generic computer component performing a generic computer function) such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of using generic computer components to perform the abstract idea amounts to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. Thus, the claim is not patent eligible.
Regarding claim 3
Step 1: The claim recites a method; therefore, it falls into the statutory category of processes.
Step 2A Prong 1:
“wherein the method further comprises the step of: sending each in-window reward, associated with an action performed by a learning agent, to the learning agent as soon as it is received.”
This limitation is directed to the abstract idea of a mental process (concepts performed in the human mind, including observation and evaluation [see MPEP 2106.04(a)(2) III. C.]).
Step 2A Prong 2: This judicial exception is not integrated into a practical. In particular, the claim only recites additional elements that are mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. See MPEP 2106.05(f). The additional element of “computer implemented”, as drafted, is reciting generic computer components. The generic computer components in these steps are recited at a high-level of generality (i.e., as a generic computer component performing a generic computer function) such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of using generic computer components to perform the abstract idea amounts to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. Thus, the claim is not patent eligible.
Regarding claim 4
Step 1: The claim recites a method; therefore, it falls into the statutory category of processes.
Step 2A Prong 1: The claim recites multiple mental processes, as explained below. The claim recites, inter alia:
“wherein the reward information relating to all in-window rewards received for the action includes an indication that all the rewards associated with the action have been received by the learning agent.”
This limitation is directed to the abstract idea of a mental process (concepts performed in the human mind, including observation and evaluation [see MPEP 2106.04(a)(2) III. C.]).
Step 2A Prong 2: This judicial exception is not integrated into a practical. In particular, the claim only recites additional elements that are mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. See MPEP 2106.05(f). The additional element of “computer implemented”, as drafted, is reciting generic computer components. The generic computer components in these steps are recited at a high-level of generality (i.e., as a generic computer component performing a generic computer function) such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into practical application, the additional element of using generic computer components to perform the abstract idea amounts to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. Thus, the claim is not patent eligible.
Regarding claim 5
Step 1: The claim recites a method; therefore, it falls into the statutory category of processes.
Step 2A Prong 1: The claim recites multiple mental processes, as explained below. The claim recites, inter alia:
“wherein the learning agent is configured to use less than all of the rewards received as a result of the sending step for reinforcement-type machine learning.”
This limitation is directed to the abstract idea of a mental process (concepts performed in the human mind, including observation and evaluation [see MPEP 2106.04(a)(2) III. C.]).
Step 2A Prong 2: This judicial exception is not integrated into a practical. In particular, the claim only recites additional elements that are mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. See MPEP 2106.05(f). The additional element of “computer implemented”, as drafted, is reciting generic computer components. The generic computer components in these steps are recited at a high-level of generality (i.e., as a generic computer component performing a generic computer function) such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into practical application, the additional element of using generic computer components to perform the abstract idea amounts to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. Thus, the claim is not patent eligible.
Regarding claim 6
Step 1: The claim recites a method; therefore, it falls into the statutory category of processes.
Step 2A Prong 1: The claim recites multiple mental processes, as explained below. The claim recites, inter alia:
“wherein the fixed length sliding window is implemented using a fixed length of time.”
This limitation under its broadest reasonable interpretation a human could perform matrix multiplication using pen and paper.
Step 2A Prong 2: This judicial exception is not integrated into a practical. In particular, the claim only recites additional elements that are mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. See MPEP 2106.05(f). The additional element of “computer implemented”, as drafted, is reciting generic computer components. The generic computer components in these steps are recited at a high-level of generality (i.e., as a generic computer component performing a generic computer function) such that it amounts no more than mere instructions to apply the exception using a generic computer component.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into practical application, the additional element of using generic computer components to perform the abstract idea amounts to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. Thus, the claim is not patent eligible.
Regarding claim 7
Step 1: The claim recites a method; therefore, it falls into the statutory category of processes.
Step 2A Prong 1: The claim recites multiple mental processes, as explained below. The claim recites, inter alia:
“wherein the fixed length of time comprises a linear period of time and the fixed length sliding window is implemented as a timer.”
This limitation is directed to the abstract idea of a mental process (concepts performed in the human mind, including observation and evaluation [see MPEP 2106.04(a)(2) III. C.]).
Step 2A Prong 2: This judicial exception is not integrated into a practical. In particular, the claim only recites additional elements that are mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. See MPEP 2106.05(f). The additional element of “computer implemented”, as drafted, is reciting generic computer components. The generic computer components in these steps are recited at a high-level of generality (i.e., as a generic computer component performing a generic computer function) such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into practical application, the additional element of using generic computer components to perform the abstract idea amounts to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. Thus, the claim is not patent eligible.
Regarding claim 8
Step 1: The claim recites a method; therefore, it falls into the statutory category of processes.
Step 2A Prong 1: The claim recites multiple mental processes, as explained below. The claim recites, inter alia:
“wherein the fixed length of time comprises a fixed number of discrete steps associated with the implementation of the reinforcement-type machine learning environment.”
This limitation is directed to the abstract idea of a mental process (concepts performed in the human mind, including observation and evaluation [see MPEP 2106.04(a)(2) III. C.]).
Step 2A Prong 2: This judicial exception is not integrated into a practical. In particular, the claim only recites additional elements that are mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. See MPEP 2106.05(f). The additional element of “computer implemented”, as drafted, is reciting generic computer components. The generic computer components in these steps are recited at a high-level of generality (i.e., as a generic computer component performing a generic computer function) such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into practical application, the additional element of using generic computer components to perform the abstract idea amounts to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. Thus, the claim is not patent eligible.
Regarding claim 9
Step 1: The claim recites a method; therefore, it falls into the statutory category of processes.
Step 2A Prong 1:
Step 2A Prong 2: This judicial exception is not integrated into a practical. In particular, the claim only recites additional elements that are mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. See MPEP 2106.05(f). The additional element of “wherein the fixed length sliding window is implemented using a circular memory buffer” as drafted, is reciting generic computer components. The generic computer components in these steps are recited at a high-level of generality (i.e., as a generic computer component performing a generic computer function) such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into practical application, the additional element of using generic computer components to perform the abstract idea amounts to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. Thus, the claim is not patent eligible.
Regarding claim 10
Step 1: The claim recites a method; therefore, it falls into the statutory category of processes.
Step 2A Prong 1: The claim recites multiple mental processes, as explained below. The claim recites, inter alia:
“wherein rewards comprise an assessment of the action associated with the reward.”
This limitation is directed to the abstract idea of a mental process (concepts performed in the human mind, including observation and evaluation [see MPEP 2106.04(a)(2) III. C.]).
Step 2A Prong 2: This judicial exception is not integrated into a practical. In particular, the claim only recites additional elements that are mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. See MPEP 2106.05(f). The additional element of “computer implemented”, as drafted, is reciting generic computer components. The generic computer components in these steps are recited at a high-level of generality (i.e., as a generic computer component performing a generic computer function) such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into practical application, the additional element of using generic computer components to perform the abstract idea amounts to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. Thus, the claim is not patent eligible.
Regarding claim 11
Step 1: The claim recites a method; therefore, it falls into the statutory category of processes.
Step 2A Prong 1: The claim recites multiple mental processes, as explained below. The claim recites, inter alia:
“wherein the assessment of the action associated with the reward can be a positive, negative or neutral assessment”
This limitation is directed to the abstract idea of a mental process (concepts performed in the human mind, including observation and evaluation [see MPEP 2106.04(a)(2) III. C.]).
Step 2A Prong 2: This judicial exception is not integrated into a practical. In particular, the claim only recites additional elements that are mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. See MPEP 2106.05(f). The additional element of “computer implemented”, as drafted, is reciting generic computer components. The generic computer components in these steps are recited at a high-level of generality (i.e., as a generic computer component performing a generic computer function) such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into practical application, the additional element of using generic computer components to perform the abstract idea amounts to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. Thus, the claim is not patent eligible.
Regarding claim 12
Step 1: The claim recites a method; therefore, it falls into the statutory category of processes.
Step 2A Prong 1: The claim recites multiple mental processes, as explained below. The claim recites, inter alia:
“wherein rewards comprise a numerical value representing the positive, negative or neutral assessment of the action associated with the reward.”
This limitation is directed to the abstract idea of a mental process (concepts performed in the human mind, including observation and evaluation [see MPEP 2106.04(a)(2) III. C.]).
Step 2A Prong 2: This judicial exception is not integrated into a practical. In particular, the claim only recites additional elements that are mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. See MPEP 2106.05(f). The additional element of “computer implemented”, as drafted, is reciting generic computer components. The generic computer components in these steps are recited at a high-level of generality (i.e., as a generic computer component performing a generic computer function) such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into practical application, the additional element of using generic computer components to perform the abstract idea amounts to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. Thus, the claim is not patent eligible.
Regarding claim 13
Step 1: The claim recites a method; therefore, it falls into the statutory category of processes.
Step 2A Prong 1: The claim recites multiple mental processes, as explained below. The claim recites, inter alia:
“wherein rewards are generated by one or more human evaluators.”
This limitation is directed to the abstract idea of a mental process (concepts performed in the human mind, including observation and evaluation [see MPEP 2106.04(a)(2) III. C.]).
Step 2A Prong 2: This judicial exception is not integrated into a practical. In particular, the claim only recites additional elements that are mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. See MPEP 2106.05(f). The additional element of “computer implemented”, as drafted, is reciting generic computer components. The generic computer components in these steps are recited at a high-level of generality (i.e., as a generic computer component performing a generic computer function) such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into practical application, the additional element of using generic computer components to perform the abstract idea amounts to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. Thus, the claim is not patent eligible.
Regarding claim 14
Step 1: The claim recites a method; therefore, it falls into the statutory category of processes.
Step 2A Prong 1: The claim recites multiple mental processes, as explained below. The claim recites, inter alia:
“wherein rewards are generated by one or more other learning agents.”
This limitation is directed to the abstract idea of a mental process (concepts performed in the human mind, including observation and evaluation [see MPEP 2106.04(a)(2) III. C.]).
Step 2A Prong 2: This judicial exception is not integrated into a practical. In particular, the claim only recites additional elements that are mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. See MPEP 2106.05(f). The additional element of “computer implemented”, as drafted, is reciting generic computer components. The generic computer components in these steps are recited at a high-level of generality (i.e., as a generic computer component performing a generic computer function) such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into practical application, the additional element of using generic computer components to perform the abstract idea amounts to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. Thus, the claim is not patent eligible.
Regarding claim 15
Step 1: The claim recites a method; therefore, it falls into the statutory category of processes.
Step 2A Prong 1: The claim recites multiple mental processes, as explained below. The claim recites, inter alia:
“wherein the values of rewards are weighted differently depending on which of the one or more human evaluators and one or more other learning agents generated the rewards.”
This limitation is directed to the abstract idea of a mental process (concepts performed in the human mind, including observation and evaluation [see MPEP 2106.04(a)(2) III. C.]).
Step 2A Prong 2: This judicial exception is not integrated into a practical. In particular, the claim only recites additional elements that are mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. See MPEP 2106.05(f). The additional element of “computer implemented”, as drafted, is reciting generic computer components. The generic computer components in these steps are recited at a high-level of generality (i.e., as a generic computer component performing a generic computer function) such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into practical application, the additional element of using generic computer components to perform the abstract idea amounts to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. Thus, the claim is not patent eligible.
Regarding claim 16
Step 1: The claim recites a method; therefore, it falls into the statutory category of processes.
Step 2A Prong 1:
Step 2A Prong 2: This judicial exception is not integrated into a practical. In particular, the claim only recites additional elements that are mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. See MPEP 2106.05(f). The additional element of “the system comprising: a processor; and at least one non-transitory memory containing instructions which when executed by the processor cause the system carry out the method of claim 1”, as drafted, is reciting generic computer components. The generic computer components in these steps are recited at a high-level of generality (i.e., as a generic computer component performing a generic computer function) such that it amounts no more than mere instructions to apply the exception using a generic computer component. In addition, the claim limitation “A system for processing rewards received during a reinforcement learning session in which one or more learning agents each perform one or more actions in a reinforcement-type machine learning environment, each reward being associated with an action…” as explained by the Supreme Court, the addition of insignificant extra-solution activity does not amount to an inventive concept, particularly when the activity is well-understood or conventional. See MPEP 2106.05(g). Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The additional element of using generic computer components to perform the abstract idea amounts to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. The only remaining limitation of the claim “A system for processing rewards received during a reinforcement learning session in which one or more learning agents each perform one or more actions in a reinforcement-type machine learning environment, each reward being associated with an action…” constitute storing and retrieving information in memory, which the courts have found to be well-understood, routine, and conventional. See MPEP 2106.05(d)(II); Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015). Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1-12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yang et al. (“Adaptive Reward Computation in Reinforcement Learning-Based Continuous Integration Testing”) in view of Deng et al. (“Value-based Algorithms Optimization with Discounted Multiple-step Learning Method in Deep Reinforcement Learning”).
Regarding claim 1 (Currently Amended)
Yang teaches a computer implemented method for processing rewards received during a reinforcement learning session in which one or more learning agents each perform one or more actions in a reinforcement-type machine learning environment, (section B “Reinforcement learning (RL) comprises four elements: agent, environment, action, and reward, which can solve decision-making with automatic and continuous decision-making mechanisms. RL’s goal is to maximize the cumulative reward in the process of agent interactions with the environment, which continually adjusts the behavior to optimize its adaptability to the environment by the feedback of the historical action results”)
each reward being associated with an action, the method comprising the steps of: setting a fixed length sliding window within which past actions are considered; (abstract “Firstly, the fixed-size sliding window is introduced to set a fixed length of recent historical information for each CItest.Then dynamic sliding window techniques are proposed, where the window size is continuously adaptive to each CI testing.”)
receiving a plurality of rewards, each reward being associated with an action taken by a learning agent; (section B “Reinforcement learning (RL) comprises four elements: agent, environment, action, and reward, which can solve decision-making with automatic and continuous decision-making mechanisms. RL’s goal is to maximize the cumulative reward in the process of agent interactions with the environment, which continually adjusts the behavior to optimize its adaptability to the environment by the feedback of the historical action results”)
defining a group of in-window rewards comprising rewards received while their associated actions are considered to be within the fixed length sliding window; (pg. 2 “From the aspect of the time dimension in CI testing, the latest execution can be more valuable for the reward computation of RL. Then the sliding window is usually started from the last CI testing cycle. The size of the sliding window determines the scope of the historical information used for reward computation.”)
Yang does not teach and when an action performed by a learning agent is no longer considered to be within the fixed length sliding window, sending the learning agent reward information relating to all in-window rewards received for the action.
Deng teaches and when an action performed by a learning agent is no longer considered to be within the fixed length sliding window, sending the learning agent reward information relating to all in-window rewards received for the action. (Abstract “In this method, regard the discounted truncated N-step return rather than accumulated discount reward as the important part of target network when computing the TD-error as a loss function of evaluate network” also see FIG. 1 “. N-step learning method. The red arrow depicts the policy learned by the agent, and square depicts the predicted value in this timestep. As shown in this picture, truncated n-step return R(n) t can be calculated by the rewards generated during this procedure”)
Yang and Deng are analogous art because they are all directed to machine learning.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combined adaptive reinforcement learning of Yang with optimization method in deep reinforcement learning of Deng.
One of ordinary skill in the art would have been motivated to make this modification in order to include N-step return method to produce “more accurate predict of value function, thereby outperform other optimal methods in terms of stability, overestimate and convergence” as disclosed by (Deng abstract “In this method, regard the discounted truncated N-step return rather than accumulated discount reward as the important part of target network when computing the TD-error as a loss function of evaluate network. In the experiment part, we perform experiments compare to value-based algorithms without this method, and prove this method can make more accurate predict of value function, thereby outperform other optimal methods in terms of stability, overestimate and convergence.”)
Regarding claim 2
Yang in view of Deng teaches the computer implemented method of claim 1.
Yang teaches wherein each in-window reward includes a value associated with a critical assessment of the action associated with the reward, and the reward information relating to all in-window rewards received for the action includes information associated to an aggregate of the values of the in-window rewards. (Section B “Table 1 lists the details of the fourteen datasets, including the number of test cases, the number of CI cycles, etc. ‘Results’ is the number of execution results, which refers to the total number of executions of all test cases during the entire CI process, ‘Failure Rate’ represents the failure ratio of total executions, and ‘Frequency’ is the probability of each test case occurring in each cycle.”)
Regarding claim 3
Yang in view of Deng teaches the computer implemented method of claim 1.
Yang further teaches wherein the method further comprises the step of: sending each in-window reward, associated with an action performed by a learning agent, to the learning agent as soon as it is received. (pg. 5 right col “According to the CI environment, the fixed-size sliding window is proposed to measure valid information with strong relevance to the test. Under the fixed-size sliding window, all test cases are extracted features with the same historical execution information scope. Then, the test cases that fail frequently can get more rewards. The size of the sliding window determines the selection of historical information for the test case and accordingly affects the reward computation.”)
Regarding claim 4
Yang in view of Deng teaches the computer implemented method of claim 3.
Yang further teaches wherein the reward information relating to all in-window rewards received for the action includes an indication that all the rewards associated with the action have been received by the learning agent. (pg. 5 right col “According to the CI environment, the fixed-size sliding window is proposed to measure valid information with strong relevance to the test. Under the fixed-size sliding window, all test cases are extracted features with the same historical execution information scope. Then, the test cases that fail frequently can get more rewards. The size of the sliding window determines the selection of historical information for the test case and accordingly affects the reward computation.”)
Regarding claim 5
Yang in view of Deng teaches the computer implemented method of claim 1.
Yang further teaches wherein the learning agent is configured to use less than all of the rewards received as a result of the sending step for reinforcement-type machine learning. (Pg. 2 “From the aspect of the time dimension in CI testing, the latest execution can be more valuable for the reward computation of RL. Then the sliding window is usually started from the last CI testing cycle. The size of the sliding window determines the scope of the historical information used for reward computation. Two dynamic methods are proposed based on the consideration of failure distribution with the time dimension and the spatial dimension of CI testing, respectively.”)
Regarding claim 6
Yang in view of Deng teaches the computer implemented method of claim 1.
Yang further teaches wherein the fixed length sliding window is implemented using a fixed length of time. (Section D “The fixed sliding window size is set to 5 based on the previous experimental results [16]. For the setting of dynamic sliding window size, to avoid the situations where the window is too small, which is close to the latest execution result, or the window is too large, which is close to the whole information,”)
Regarding claim 7
Yang in view of Deng teaches the computer implemented method of claim 1.
Yang further teaches wherein the fixed length of time comprises a linear period of time and the fixed length sliding window is implemented as a timer. (Section D “We keep the parameters consistent to RETECS [8] for the comparison of the experimental results. For the agent selection, a Network-based agent is selected, which is better for CI testing in RETECS [8]. However, the combination of RL and neural network may lead to instability of the algorithm [7], so the experiments are repeated 60 times. To simulate the rapid feedback mechanism of CI testing, the test cases are sorted in descending order of priority until reaching the specified time, which is half the execution time of all test cases in the cycle.”)
Regarding claim 10
Yang in view of Deng teaches the computer implemented method of claim 1.
Yang further teaches wherein rewards comprise an assessment of the action associated with the reward. (Abstract “wo methods are proposed, the test suite-based dynamic sliding window and the individual test case-based dynamic sliding window. The empirical studies are conducted on fourteen industrial-level programs, and the results reveal that under limited time, the sliding window-based reward function can effectively improve the TCP effect, where the NAPFD (Normalized Average Percentage of Faults Detected)”)
Regarding claim 16
Yang in view of Deng teaches the computer implemented method of claim 1.
Yang further teaches a system for processing rewards received during a reinforcement learning session in which one or more learning agents each perform one or more actions in a reinforcement-type machine learning environment, each reward being associated with an action, the system comprising: a processor; and at least one non-transitory memory containing instructions which when executed by the processor cause the system carry out the method of claim 1. (Section A “For complex optimization problems with multiple objects, a search-based approach is proposed by Li et al. [25], which was first used to solve multi-object TCP in software regression testing. There were a series of search algorithms for multi-object TCP problems, and the GPU parallel method was proposed [26] to further improve search efficiency. For the algorithm optimization problem of TCP, thee pistatic gene theory was proposed with the interdependence among test cases [27], [28].”)
Claim(s) 8-9 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yang et al. (“Adaptive Reward Computation in Reinforcement Learning-Based Continuous Integration Testing”) in view of Deng et al. (“Value-based Algorithms Optimization with Discounted Multiple-step Learning Method in Deep Reinforcement Learning”) and further in view of Heinrich et al. (“Deep Reinforcement Learning from Self-Play in Imperfect-Information Games”).
Regarding claim 8
Yang in view of Deng teaches the computer implemented method of claim 6.
Yang in view of Deng does not teach wherein the fixed length of time comprises a fixed number of discrete steps associated with the implementation of the reinforcement-type machine learning environment.
Heinrich teaches wherein the fixed length of time comprises a fixed number of discrete steps associated with the implementation of the reinforcement-type machine learning environment. (pg. 4 “NFSP uses ˆβi t+1 − ˆπi t ≈ d dtˆπi t as a discrete-time approximation of the derivative that is used in these anticipatory dynamics. Note that ∆ˆπi t ∝ ˆβi t+1 − ˆπi t is the normal-form update direction of common discrete-time fictitious play.” Also see pg. 7 “Each agent performed 2 stochastic gradient updates of mini-batch size 256 per network for every 256 steps in the game. The target network was refitted every 1000 updates. NFSP’s anticipatory parameter was set to η = 0.1. The-greedy policies’ exploration started at 0.08 and decayed to 0, more slowly than in Leduc Hold’em. In addition to NFSP’s main, average strategy profile we also evaluated the best response and greedy-average strategies, which deter”)
Yang, Deng and Heinrich are analogous art because they are all directed to machine learning.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combined adaptive reinforcement learning of Yang in view of Deng to include interactive reinforcement learning of Heinrich.
One of ordinary skill in the art would have been motivated to make this modification in order provide any deep learning agent learner to learn sequential decision-making policies quicker as disclosed by (Heinrich abstract “This paper proposes a novel approach combining IL with different types of RL methods, namely State–action–reward–state–action (SARSA) and Asynchronous Advantage Actor-Critic Agents (A3C), to overcome the problems of both stand-alone systems. It is addressed how to effectively leverage the teacher’s feedback– be it direct binary or indirect detailed– for the agent learner to learn sequential decision-making policies. The results of this study on various OpenAI-Gym environments show that this algorithmic method can be incorporated with different combinations, significantly decreases both human endeavor and tedious exploration process.
Regarding claim 9
Yang in view of Deng teaches the computer implemented method of claim 1.
Yang in view of Deng does not teach wherein the fixed length sliding window is implemented using a circular memory buffer.
Heinrich teaches wherein the fixed length sliding window is implemented using a circular memory buffer. (See Algorithm 1 “Initialize replay memories MRL (circular buffer) and MSL (reservoir”)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combined adaptive reinforcement learning of Yang in view of Deng to include interactive reinforcement learning of Heinrich.
One of ordinary skill in the art would have been motivated to make this modification in order provide any deep learning agent learner to learn sequential decision-making policies quicker as disclosed by (Heinrich abstract “This paper proposes a novel approach combining IL with different types of RL methods, namely State–action–reward–state–action (SARSA) and Asynchronous Advantage Actor-Critic Agents (A3C), to overcome the problems of both stand-alone systems. It is addressed how to effectively leverage the teacher’s feedback– be it direct binary or indirect detailed– for the agent learner to learn sequential decision-making policies. The results of this study on various OpenAI-Gym environments show that this algorithmic method can be incorporated with different combinations, significantly decreases both human endeavor and tedious exploration process.
Claim(s) 11-12 and 14-15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yang et al. (“Adaptive Reward Computation in Reinforcement Learning-Based Continuous Integration Testing”) in view of Deng et al. (“Value-based Algorithms Optimization with Discounted Multiple-step Learning Method in Deep Reinforcement Learning”) and further in view of Tampuu et al. (“Multiagent cooperation and competition with deep reinforcement learning”).
Regarding claim 11
Yang in view of Deng teaches the computer implemented method of claim 10.
Yang in view of Deng does not teach wherein the assessment of the action associated with the reward can be a positive, negative or neutral assessment.
Tampuu teaches wherein the assessment of the action associated with the reward can be a positive, negative or neutral assessment. (Pg. 4 “This makes it essentially a zero sum game, where a positive reward for the left player implies a negative reward of the same size for the right player and vice versa. Notice that ρ = 1 is the only case where the rewards of the two players sum up to zero, for all ρ < 1 the sum of rewards is negative.”)
Yang, Deng and Tampuu are analogous art because they are all directed to machine learning.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combined adaptive reinforcement learning of Yang in view of Deng to include multiagent cooperation using deep reinforcement learning of Tampuu.
One of ordinary skill in the art would have been motivated to make this modification in order show “how learning by playing against another adaptive agent, instead of against a hard-wired algorithm, results in more robust strategies” as disclosed by (Tampuu abstract “By manipulating the classical rewarding scheme of Pong we show how competitive and collaborative behaviors emerge. We also describe the progression from competitive to collaborative behavior when the incentive to cooperate is increased. Finally we show how learning by playing against another adaptive agent, instead of against a hard-wired algorithm, results in more robust strategies. The present work shows that Deep Q-Networks can become a useful tool for studying decentralized learning of multiagent systems coping with high-dimensional environments.”).
Regarding claim 12
Yang in view of Deng teaches the computer implemented method of claim 1.
Yang in view of Deng does not teach wherein rewards comprise a numerical value representing the positive, negative or neutral assessment of the action associated with the reward.
Tampuu teaches wherein rewards comprise a numerical value representing the positive, negative or neutral assessment of the action associated with the reward. (Pg. 4 “In the traditional rewarding scheme of Pong, also used in [7], the player who scores a point gets a positive reward of size 1 (ρ = 1). The player conceding a point is penalized with a reward of-1.”)
Yang, Deng and Tampuu are analogous art because they are all directed to machine learning.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combined adaptive reinforcement learning of Yang in view of Deng to include multiagent cooperation using deep reinforcement learning of Tampuu.
One of ordinary skill in the art would have been motivated to make this modification in order show “how learning by playing against another adaptive agent, instead of against a hard-wired algorithm, results in more robust strategies” as disclosed by (Tampuu abstract “By manipulating the classical rewarding scheme of Pong we show how competitive and collaborative behaviors emerge. We also describe the progression from competitive to collaborative behavior when the incentive to cooperate is increased. Finally we show how learning by playing against another adaptive agent, instead of against a hard-wired algorithm, results in more robust strategies. The present work shows that Deep Q-Networks can become a useful tool for studying decentralized learning of multiagent systems coping with high-dimensional environments.”).
Regarding claim 14
Yang in view of Deng teaches the computer implemented method of claim 1.
Yang in view of Deng does not teach wherein rewards are generated by one or more other learning agents.
Tampuu teaches wherein rewards are generated by one or more other learning agents. (Pg. 5 “In all of the experiments we let the agents learn for 50 epochs, 250000 time steps each. We limit the learning to 50 epochs because the Q-values predicted by the network have stabilized (see S1 Fig).”)
Yang, Deng and Tampuu are analogous art because they are all directed to machine learning.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combined adaptive reinforcement learning of Yang in view of Deng to include multiagent cooperation using deep reinforcement learning of Tampuu.
One of ordinary skill in the art would have been motivated to make this modification in order show “how learning by playing against another adaptive agent, instead of against a hard-wired algorithm, results in more robust strategies” as disclosed by (Tampuu abstract “By manipulating the classical rewarding scheme of Pong we show how competitive and collaborative behaviors emerge. We also describe the progression from competitive to collaborative behavior when the incentive to cooperate is increased. Finally we show how learning by playing against another adaptive agent, instead of against a hard-wired algorithm, results in more robust strategies. The present work shows that Deep Q-Networks can become a useful tool for studying decentralized learning of multiagent systems coping with high-dimensional environments.”).
Regarding claim 15
Yang in view of Deng teaches the computer implemented method of claim 14.
Yang in view of Deng does not teach wherein rewards are generated by one or more other learning agents.
Tampuu teaches wherein the values of rewards are weighted differently depending on which of the one or more human evaluators and one or more other learning agents generated the rewards. . (Pg. 5 “In all of the experiments we let the agents learn for 50 epochs, 250000 time steps each. We limit the learning to 50 epochs because the Q-values predicted by the network have stabilized (see S1 Fig).”)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combined adaptive reinforcement learning of Yang in view of Deng to include multiagent cooperation using deep reinforcement learning of Tampuu.
One of ordinary skill in the art would have been motivated to make this modification in order show “how learning by playing against another adaptive agent, instead of against a hard-wired algorithm, results in more robust strategies” as disclosed by (Tampuu abstract “By manipulating the classical rewarding scheme of Pong we show how competitive and collaborative behaviors emerge. We also describe the progression from competitive to collaborative behavior when the incentive to cooperate is increased. Finally we show how learning by playing against another adaptive agent, instead of against a hard-wired algorithm, results in more robust strategies. The present work shows that Deep Q-Networks can become a useful tool for studying decentralized learning of multiagent systems coping with high-dimensional environments.”).
Claim(s) 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yang et al. (“Adaptive Reward Computation in Reinforcement Learning-Based Continuous Integration Testing”) in view of Deng et al. (“Value-based Algorithms Optimization with Discounted Multiple-step Learning Method in Deep Reinforcement Learning”) and further in view of Neda et al. (“HUMAN / AI INTERACTION LOOP TRAINING: NEW APPROACH FOR INTERACTIVE REINFORCEMENT LEARNING”).
Regarding claim 13
Yang in view of Deng teaches the computer implemented method of claim 1.
Yang in view of Deng does not teach wherein rewards are generated by one or more human evaluators.
Neda teaches wherein rewards are generated by one or more human evaluators. (Pg. 2 “The teacher assistance considers both direct dual feedback, with positive and negative reward, and indirect detailed feedback, with access to action domain feedback using online policy IL process. Management of teacher’s feedback in the “feedback management” block is one of the features of the structure (see Fig. 2). Also this structure reflects the online teacher feedback as soon as the learner takes an action and deals with quantity of teacher’s feedback”)
Yang, Deng and Neda are analogous art because they are all directed to machine learning.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combined adaptive reinforcement learning of Yang in view of Deng to include multiagent cooperation using deep reinforcement learning of Neda.
One of ordinary skill in the art would have been motivated to make this modification in order show “how to effectively leverage the teacher’s feedback– be it direct binary or indirect detailed– for the agent learner to learn sequential decision-making policies” as disclosed by (Neda abstract “his paper proposes a novel approach combining IL with different types of RL methods, namely State–action–reward–state–action (SARSA) and Asynchronous Advantage Actor-Critic Agents (A3C), to overcome the problems of both stand-alone systems. It is addressed how to effectively leverage the teacher’s feedback– be it direct binary or indirect detailed– for the agent learner to learn sequential decision-making policies. The results of this study on various OpenAI-Gym environments show that this algorithmic method can be incorporated with different combinations, significantly decreases both human endeavor and tedious exploration process.”)
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to VAN C MANG whose telephone number is (571)270-7598. The examiner can normally be reached Mon - Fri 8:00-5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, David Yi can be reached at 5712707519. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/VAN C MANG/Primary Examiner, Art Unit 2126