DETAILED ACTION
This action is responsive to the amendment filed on 04/14/2026. Claims 1-7, 9-14, and 16-22 are pending in the case. Claims 1, 9-10, and 16-17 are currently amended in the case. Claims 1, 16, and 17 are independent claims.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Acknowledgement is made of applicant’s claim for domestic priority based on international application no. PCT/EP2021/074892 filed 09/10/2021, which claims priority to provisional application no. 63/076876 filed 09/10/2020.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 04/15/2026 is being considered by the examiner.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-7, 9-14, and 16-22 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding claim 1: Step 1 Statutory Category: Claim 1 is directed to a method, which falls under one of the four statutory categories.
Step 2A Prong 1 Judicial Exception: Claim 1 recites, in part, “generate a policy output that defines a control policy for controlling the agent”. This limitation, under the broadest reasonable interpretation, covers the recitation of mathematical concepts, see MPEP §2106.04(a)(2)(I). Further, the claim recites: “selecting a skill from the set of skills”. This limitation, under the broadest reasonable interpretation, covers the recitation of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case a judgment. See MPEP § 2106.04(a)(2)(III). Further, the claim recites: “generating a trajectory… the trajectory comprising a sequence of observations over a number of time steps received while the agent interacts with the environment”. This limitation, under the broadest reasonable interpretation, covers the recitation of mathematical concepts, see MPEP §2106.04(a)(2)(I). Further, the claim recites: “processing a relative input comprising (i) an initial observation at an initial time step in the sequence and (ii) a last observation at a last time step in the sequence… process the relative input to generate a relative output that includes a respective relative score corresponding to each skill in the set of skills, each relative score representing an estimated likelihood that the policy neural network was conditioned on the corresponding skill while the trajectory was generated”. This limitation, under the broadest reasonable interpretation, covers the recitation of mathematical concepts, see MPEP §2106.04(a)(2)(I). Further, the claim recites: “processing an absolute input comprising the last observation in the sequence… process the absolute input to generate an absolute output that includes a respective absolute score corresponding to each skill in the set of skills, each absolute score representing an estimated likelihood that the policy neural network was conditioned on the corresponding skill while the trajectory was generated”. This limitation, under the broadest reasonable interpretation, covers the recitation of mathematical concepts, see MPEP §2106.04(a)(2)(I). Further, the claim recites: “generating a reward for the trajectory from the absolute score corresponding to the selected skill and the relative score corresponding to the selected skill, wherein the reward rewards high relative scores and penalizes high absolute scores”. This limitation, under the broadest reasonable interpretation, covers the recitation of mathematical concepts, see MPEP §2106.04(a)(2)(I).
Step 2A Prong 2 Integration into a practical application: This judicial exception is not integrated into a practical application. In particular the claim recites: “a policy neural network for use in controlling an agent interacting with an environment”. This limitation is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Further, the claim recites: “wherein the policy neural network is configured to receive a policy input comprising an input observation characterizing a state of the environment and data identifying a skill from a set of skills”. This limitation amounts to mere data gathering. It is necessary to acquire the data in order to use the recited judicial exception. Therefore, this limitation is insignificant extra-solution activity to the judicial exception, see MPEP §2106.05(g). Further, the claim recites: “by controlling the agent using the policy neural network while the policy neural network is conditioned on the selected skill”, “while controlled using the policy neural network that is conditioned on the selected skill”, “using a relative discriminator neural network that is configured to…”, and “using an absolute discriminator neural network that is configured to…”. These limitations are additional elements that generally link the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Finally, the claim recites: “training the policy neural network on the reward for the trajectory to maximize time discounted expected rewards for generated trajectories”. This is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §2106.05(f).
Step 2B Significantly more: The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements: “a policy neural network for use in controlling an agent interacting with an environment”, “by controlling the agent using the policy neural network while the policy neural network is conditioned on the selected skill”, “while controlled using the policy neural network that is conditioned on the selected skill”, “using a relative discriminator neural network that is configured to…”, and “using an absolute discriminator neural network that is configured to…” generally link the use of the judicial exception to a particular technological environment or field of use. Elements that merely generally link the use of the judicial exception to a particular technological environment or field of use cannot provide an inventive concept. Further, the additional element “wherein the policy neural network is configured to receive a policy input comprising an input observation characterizing a state of the environment and data identifying a skill from a set of skills” is insignificant extra-solution activity to the judicial exception and is directed to receiving or transmitting data over a network which courts have recognized as well-understood, routine, and conventional when they are claimed in a generic manner, see MPEP §2106.05(d)(II). Finally, the additional element “training the policy neural network on the reward for the trajectory to maximize time discounted expected rewards for generated trajectories” amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. Elements that merely amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process cannot provide an inventive concept. The claim is not patent eligible.
Regarding claim 2, the rejection of claim 1 is incorporated, and further, the claim recites: “training the absolute discriminator neural network to optimize an objective function that encourages the absolute score corresponding to the selected skill to be increased”. This limitation is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. Elements that merely amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process cannot provide an inventive concept. The claim is not patent eligible.
Regarding claim 3, the rejection of claim 1 is incorporated, and further, the claim recites: “training the relative discriminator neural network to optimize an objective function that encourages the relative score corresponding to the selected skill to be increased”. This limitation is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. Elements that merely amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process cannot provide an inventive concept. The claim is not patent eligible.
Regarding claim 4, the rejection of claim 1 is incorporated, and further, the claim recites: “wherein the absolute discriminator neural network and the relative discriminator neural network share some parameters”. This limitation is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Elements that merely generally link the use of the judicial exception to a particular technological environment or field of use cannot provide an inventive concept. The claim is not patent eligible.
Regarding claim 5, the rejection of claim 4 is incorporated, and further, the claim recites: “generates encoded representations of received observations”. This limitation recites mathematical concepts in addition to those identified in the rejection of the parent claim. Thus, the claim recites a judicial exception.
Further, the claim recites: “wherein the absolute discriminator neural network and the relative discriminator neural network share an encoder neural network”. This limitation is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Elements that merely generally link the use of the judicial exception to a particular technological environment or field of use cannot provide an inventive concept. The claim is not patent eligible.
Regarding claim 6, the rejection of claim 5 is incorporated, and further, the claim recites: “process the encoded representation of the last observation to generate the absolute output”. This limitation recites mathematical concepts in addition to those identified in the rejection of the parent claim, thus the claim recites a judicial exception.
Further, the claim recites: “wherein the absolute discriminator neural network comprises an absolute decoder neural network configured to…”. This limitation is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Elements that merely generally link the use of the judicial exception to a particular technological environment or field of use cannot provide an inventive concept. The claim is not patent eligible.
Regarding claim 7, the rejection of claim 5 is incorporated, and further, the claim recites: “process a concatenation of the encoded representations of the initial observation and the last observation to generate the relative output”. This limitation recites mathematical concepts in addition to those identified in the rejection of the parent claim, thus the claim recites a judicial exception.
Further, the claim recites: “wherein the relative discriminator neural network comprises a relative decoder neural network configured to…”. This limitation is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Elements that merely generally link the use of the judicial exception to a particular technological environment or field of use cannot provide an inventive concept. The claim is not patent eligible.
Regarding claim 9, the rejection of claim 1 is incorporated, and further, the claim recites: “wherein the reward is equal to or directly proportional to a difference between the relative score corresponding to the selected skill and the absolute score corresponding to the selected skill”. This limitation recites the abstract idea of a mathematical relationship, as directed to “a mathematical relationship is a relationship between variables or numbers. A mathematical relationship may be expressed in words or using mathematical symbols”. See MPEP § 2106.04(a)(2)(I)(A).
The claim does not include any additional elements that amount to an integration of the judicial exception into a practical application, nor to significantly more than the judicial exception. The claim is not patent eligible.
Regarding claim 10, the rejection of claim 1 is incorporated, and further, the claim recites: “wherein the reward is equal to or directly proportional to a difference between a logarithm of the relative score corresponding to the selected skill and a logarithm of the absolute score corresponding to the selected skill”. This limitation recites the abstract idea of a mathematical relationship, as directed to “a mathematical relationship is a relationship between variables or numbers. A mathematical relationship may be expressed in words or using mathematical symbols”. See MPEP § 2106.04(a)(2)(I)(A).
The claim does not include any additional elements that amount to an integration of the judicial exception into a practical application, nor to significantly more than the judicial exception. The claim is not patent eligible.
Regarding claim 11, the rejection of claim 1 is incorporated, and further, the claim recites: “wherein selecting a skill from the set of skills comprises: sampling a skill from a uniform probability distribution over the set of skills”. This limitation recites mathematical concepts in addition to those identified in the rejection of the parent claim, thus the claim recites a judicial exception.
The claim does not include any additional elements that amount to an integration of the judicial exception into a practical application, nor to significantly more than the judicial exception. The claim is not patent eligible.
Regarding claim 12, the rejection of claim 1 is incorporated, and further, the claim recites: “wherein training the policy neural network on the reward for the trajectory comprises training the policy neural network through off-policy reinforcement learning”. This limitation is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Elements that merely generally link the use of the judicial exception to a particular technological environment or field of use cannot provide an inventive concept. The claim is not patent eligible.
Regarding claim 13, the rejection of claim 1 is incorporated, and further, the claim recites: “wherein generating the trajectory comprises generating the trajectory starting from a last state of the environment for a preceding trajectory, and wherein the initial observation characterizes the last state of the environment for the preceding trajectory”. This limitation is a continuation of the “generating a trajectory… the trajectory comprising a sequence of observations over a number of time steps received while the agent interacts with the environment” limitation of the parent claim, and thus the claim recites a judicial exception.
The claim does not include any additional elements that amount to an integration of the judicial exception into a practical application, nor to significantly more than the judicial exception. The claim is not patent eligible.
Regarding claim 14, the rejection of claim 13 is incorporated, and further, the claim recites: “after generating the trajectory, determining whether criteria have been satisfied for resetting the environment; and in response to determining that the criteria are satisfied, selecting, as an initial state for a next trajectory to be generated, a state of the environment from a set of possible initial states of the environment”. This limitation recites mental processes in addition to those identified in the rejection of the parent claim, and thus the claim recites a judicial exception.
The claim does not include any additional elements that amount to an integration of the judicial exception into a practical application, nor to significantly more than the judicial exception. The claim is not patent eligible.
Regarding claim 16: Step 1 Statutory Category: Claim 16 is directed to a machine, which falls under one of the four statutory categories.
Step 2A Prong 1 Judicial Exception: Claim 16 recites, in part, “generate a policy output that defines a control policy for controlling the agent”. This limitation, under the broadest reasonable interpretation, covers the recitation of mathematical concepts, see MPEP §2106.04(a)(2)(I). Further, the claim recites: “selecting a skill from the set of skills”. This limitation, under the broadest reasonable interpretation, covers the recitation of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case a judgment. See MPEP § 2106.04(a)(2)(III). Further, the claim recites: “generating a trajectory… the trajectory comprising a sequence of observations over a number of time steps received while the agent interacts with the environment”. This limitation, under the broadest reasonable interpretation, covers the recitation of mathematical concepts, see MPEP §2106.04(a)(2)(I). Further, the claim recites: “processing a relative input comprising (i) an initial observation at an initial time step in the sequence and (ii) a last observation at a last time step in the sequence… process the relative input to generate a relative output that includes a respective relative score corresponding to each skill in the set of skills, each relative score representing an estimated likelihood that the policy neural network was conditioned on the corresponding skill while the trajectory was generated”. This limitation, under the broadest reasonable interpretation, covers the recitation of mathematical concepts, see MPEP §2106.04(a)(2)(I). Further, the claim recites: “processing an absolute input comprising the last observation in the sequence… process the absolute input to generate an absolute output that includes a respective absolute score corresponding to each skill in the set of skills, each absolute score representing an estimated likelihood that the policy neural network was conditioned on the corresponding skill while the trajectory was generated”. This limitation, under the broadest reasonable interpretation, covers the recitation of mathematical concepts, see MPEP §2106.04(a)(2)(I). Further, the claim recites: “generating a reward for the trajectory from the absolute score corresponding to the selected skill and the relative score corresponding to the selected skill, wherein the reward rewards high relative scores and penalizes high absolute scores”. This limitation, under the broadest reasonable interpretation, covers the recitation of mathematical concepts, see MPEP §2106.04(a)(2)(I).
Step 2A Prong 2 Integration into a practical application: This judicial exception is not integrated into a practical application. In particular the claim recites: “one or more non-transitory computer-readable storage media encoded with instructions that, when executed by one or more computers cause the one or more computers to perform first operations”. This limitation is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §2106.05(f). Further, the claim recites: “a policy neural network for use in controlling an agent interacting with an environment”. This limitation is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Further, the claim recites: “wherein the policy neural network is configured to receive a policy input comprising an input observation characterizing a state of the environment and data identifying a skill from a set of skills”. This limitation amounts to mere data gathering. It is necessary to acquire the data in order to use the recited judicial exception. Therefore, this limitation is insignificant extra-solution activity to the judicial exception, see MPEP §2106.05(g). Further, the claim recites: “by controlling the agent using the policy neural network while the policy neural network is conditioned on the selected skill”, “while controlled using the policy neural network that is conditioned on the selected skill”, “using a relative discriminator neural network that is configured to…”, and “using an absolute discriminator neural network that is configured to…”. These limitations are additional elements that generally link the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Finally, the claim recites: “training the policy neural network on the reward for the trajectory to maximize time discounted expected rewards for generated trajectories”. This is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §2106.05(f).
Step 2B Significantly more: The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements: “a policy neural network for use in controlling an agent interacting with an environment”, “by controlling the agent using the policy neural network while the policy neural network is conditioned on the selected skill”, “while controlled using the policy neural network that is conditioned on the selected skill”, “using a relative discriminator neural network that is configured to…”, and “using an absolute discriminator neural network that is configured to…” generally link the use of the judicial exception to a particular technological environment or field of use. Elements that merely generally link the use of the judicial exception to a particular technological environment or field of use cannot provide an inventive concept. Further, the additional element “wherein the policy neural network is configured to receive a policy input comprising an input observation characterizing a state of the environment and data identifying a skill from a set of skills” is insignificant extra-solution activity to the judicial exception and is directed to receiving or transmitting data over a network which courts have recognized as well-understood, routine, and conventional when they are claimed in a generic manner, see MPEP §2106.05(d)(II). Finally, the additional elements “one or more non-transitory computer-readable storage media encoded with instructions that, when executed by one or more computers cause the one or more computers to perform first operations” and “training the policy neural network on the reward for the trajectory to maximize time discounted expected rewards for generated trajectories” amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. Elements that merely amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process cannot provide an inventive concept. The claim is not patent eligible.
Regarding claim 17: Step 1 Statutory Category: Claim 17 is directed to a machine, which falls under one of the four statutory categories.
Step 2A Prong 1 Judicial Exception: Claim 17 recites, in part, “generate a policy output that defines a control policy for controlling the agent”. This limitation, under the broadest reasonable interpretation, covers the recitation of mathematical concepts, see MPEP §2106.04(a)(2)(I). Further, the claim recites: “selecting a skill from the set of skills”. This limitation, under the broadest reasonable interpretation, covers the recitation of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case a judgment. See MPEP § 2106.04(a)(2)(III). Further, the claim recites: “generating a trajectory… the trajectory comprising a sequence of observations over a number of time steps received while the agent interacts with the environment”. This limitation, under the broadest reasonable interpretation, covers the recitation of mathematical concepts, see MPEP §2106.04(a)(2)(I). Further, the claim recites: “processing a relative input comprising (i) an initial observation at an initial time step in the sequence and (ii) a last observation at a last time step in the sequence… process the relative input to generate a relative output that includes a respective relative score corresponding to each skill in the set of skills, each relative score representing an estimated likelihood that the policy neural network was conditioned on the corresponding skill while the trajectory was generated”. This limitation, under the broadest reasonable interpretation, covers the recitation of mathematical concepts, see MPEP §2106.04(a)(2)(I). Further, the claim recites: “processing an absolute input comprising the last observation in the sequence… process the absolute input to generate an absolute output that includes a respective absolute score corresponding to each skill in the set of skills, each absolute score representing an estimated likelihood that the policy neural network was conditioned on the corresponding skill while the trajectory was generated”. This limitation, under the broadest reasonable interpretation, covers the recitation of mathematical concepts, see MPEP §2106.04(a)(2)(I). Further, the claim recites: “generating a reward for the trajectory from the absolute score corresponding to the selected skill and the relative score corresponding to the selected skill, wherein the reward rewards high relative scores and penalizes high absolute scores”. This limitation, under the broadest reasonable interpretation, covers the recitation of mathematical concepts, see MPEP §2106.04(a)(2)(I).
Step 2A Prong 2 Integration into a practical application: This judicial exception is not integrated into a practical application. In particular the claim recites: “a system comprising one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform first operations”. This limitation is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §2106.05(f). Further, the claim recites: “a policy neural network for use in controlling an agent interacting with an environment”. This limitation is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Further, the claim recites: “wherein the policy neural network is configured to receive a policy input comprising an input observation characterizing a state of the environment and data identifying a skill from a set of skills”. This limitation amounts to mere data gathering. It is necessary to acquire the data in order to use the recited judicial exception. Therefore, this limitation is insignificant extra-solution activity to the judicial exception, see MPEP §2106.05(g). Further, the claim recites: “by controlling the agent using the policy neural network while the policy neural network is conditioned on the selected skill”, “while controlled using the policy neural network that is conditioned on the selected skill”, “using a relative discriminator neural network that is configured to…”, and “using an absolute discriminator neural network that is configured to…”. These limitations are additional elements that generally link the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Finally, the claim recites: “training the policy neural network on the reward for the trajectory to maximize time discounted expected rewards for generated trajectories”. This is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §2106.05(f).
Step 2B Significantly more: The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements: “a policy neural network for use in controlling an agent interacting with an environment”, “by controlling the agent using the policy neural network while the policy neural network is conditioned on the selected skill”, “while controlled using the policy neural network that is conditioned on the selected skill”, “using a relative discriminator neural network that is configured to…”, and “using an absolute discriminator neural network that is configured to…” generally link the use of the judicial exception to a particular technological environment or field of use. Elements that merely generally link the use of the judicial exception to a particular technological environment or field of use cannot provide an inventive concept. Further, the additional element “wherein the policy neural network is configured to receive a policy input comprising an input observation characterizing a state of the environment and data identifying a skill from a set of skills” is insignificant extra-solution activity to the judicial exception and is directed to receiving or transmitting data over a network which courts have recognized as well-understood, routine, and conventional when they are claimed in a generic manner, see MPEP §2106.05(d)(II). Finally, the additional elements “a system comprising one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform first operations” and “training the policy neural network on the reward for the trajectory to maximize time discounted expected rewards for generated trajectories” amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. Elements that merely amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process cannot provide an inventive concept. The claim is not patent eligible.
Regarding claim 18, the rejection of claim 17 is incorporated, and further, claim 18 is substantially similar to claim 2 respectively, and is rejected in the same manner and reasoning applying.
Regarding claim 19, the rejection of claim 17 is incorporated, and further, claim 19 is substantially similar to claim 3 respectively, and is rejected in the same manner and reasoning applying.
Regarding claim 20, the rejection of claim 17 is incorporated, and further, claim 20 is substantially similar to claim 4 respectively, and is rejected in the same manner and reasoning applying.
Regarding claim 21, the rejection of claim 20 is incorporated and further, claim 21 is substantially similar to claim 5 respectively, and is rejected in the same manner and reasoning applying.
Regarding claim 22, the rejection of claim 21 is incorporated, and further, claim 22 is substantially similar to claim 6 respectively, and is rejected in the same manner and reasoning applying.
Response to Arguments
Applicant’s arguments regarding the 35 U.S.C. 101 rejections of the claims have been fully considered but are unpersuasive.
Applicant first argues, in the final paragraph of page 9 of the response, that claim 1 recites an improved method of training a policy neural network to control an agent to perform a set of skills. Applicant specifically points to improving the diversity of the skills, consuming fewer computation resources by training the policy neural network to achieve an acceptable performance over fewer training iterations, and enabling the agent to accomplish tasks more efficiently. Examiner respectfully disagrees. It is important to note that claiming the improved speed or efficiency inherent with applying the abstract idea on a computer does not integrate a judicial exception into a practical application or provide an inventive concept, see MPEP 2106.05(f). Further, while applicant argues improved diversity of skills, it is not clear how this improvement is reflected in the claims, the claim must include the components or steps of the invention that provide the improvement described in the specification, see MPEP 2106.04(d)(1).
Applicant next argues, in the final paragraph of page 10 of the response, that adjustments to parameters of a machine learning model associated with tasks are not subsumed in a mathematical calculation. However, the step of “training” in claim 1 has been identified as an additional element, not a mathematical calculation, in the 35 U.S.C. 101 rejection seen above. Further, each of the limitations identified as mathematical calculations in claim 1 do not merely recite “adjustments to parameters of a machine learning model” but rather detailed steps that amount to mathematical calculations when read in light of the specification and using the broadest reasonable interpretation. For an in depth analysis, see the updated 35 U.S.C. 101 rejection above.
Applicant's arguments regarding the remainder of the claims rely upon the arguments asserted with respect to the independent claims, and are thus unpersuasive.
Conclusion
Claims 1-7, 9-14, and 16-22 have been rejected under 35 U.S.C. 112(b) and 35 U.S.C. 101 only. A complete prior art search was performed for these claims; however, no prior art was uncovered that disclose or fairly suggest the following claimed features:
After detailed search, the cited arts, neither alone nor in combination, teach the claimed
subject matter of claims 1, 16, and 17:
… selecting a skill from the set of skills;
generating a trajectory by controlling the agent using the policy neural network while the policy neural network is conditioned on the selected skill, the trajectory comprising a sequence of observations over a number of time steps received while the agent interacts with the environment while controlled using the policy neural network that is conditioned on the selected skill;
processing a relative input comprising (i) an initial observation at an initial time step in the sequence and (ii) a last observation at a last time step in the sequence using a relative discriminator neural network that is configured to process the relative input to generate a relative output that includes a respective relative score corresponding to each skill in the set of skills, each relative score representing an estimated likelihood that the policy neural network was conditioned on the corresponding skill while the trajectory was generated;
processing an absolute input comprising the last observation in the sequence using an absolute discriminator neural network that is configured to process the absolute input to generate an absolute output that includes a respective absolute score corresponding to each skill in the set of skills, each absolute score representing an estimated likelihood that the policy neural network was conditioned on the corresponding skill while the trajectory was generated;
generating a reward for the trajectory from the absolute score corresponding to the selected skill and the relative score corresponding to the selected skill, wherein the reward rewards high relative scores and penalizes high absolute scores; and
training the policy neural network on the reward for the trajectory to maximize time discounted expected rewards for generated trajectories”
The closest prior art of record includes:
He et al., Skill Discovery of Coordination in Multi-agent Reinforcement Learning, 06/07/2020, https://arxiv.org/pdf/2006.04021 teaches multi-agent skill discovery for discovering skills for coordination patterns of multiple agents. It uses many local discriminators and a global discriminator for processing observations made by the agent. However, all the observations are generated in a single time step by multiple agents, rather than an initial observation being generated at an initial time step and a last observation being generated at a last time step; where the initial and last observations are passed to a relative discriminator and the last observation being passed to an absolute discriminator as required by the claims.
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MOLLY CLARKE SIPPEL whose telephone number is (571)272-3270. The examiner can normally be reached Monday - Friday, 7:30 a.m. - 4:30 p.m. ET..
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kakali Chaki can be reached at (571)272-3719. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/M.C.S./ Examiner, Art Unit 2122
/KAKALI CHAKI/ Supervisory Patent Examiner, Art Unit 2122