Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
Claims 1-8 are rejected under 35 U.S.C. 112(b) as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor regards as the invention.
Taking independent claim 1 as representative, as a first matter, the claim recites the following variables/parameters which do not appear to be defined in the claim: θ, η, ∇, E, and T. Because the claim appears silent as to what these different variable/parameter instances represent, the result is that the equations as recited, and hence the claim when considered as a whole, are rendered vague and indefinite. Hence, under this rationale, the claim is rejected.
As a second matter, the claim recites, in part, a limitation for “updating parameters (θπ) of the policy, wherein the updating is carried out by a Soft Actor Critic (SAC)-style loss.” The Examiner interprets Applicants’ use of the modifier of “style” (bolded) to indicate or express a limitation for not just SAC loss but perhaps teachings/features in the “style” of SAC loss. Said another way, the claim appears directed to not just SAC loss but also teachings/features that are substantially or proximately SAC loss without being exactly or definitively SAC loss. However, the claim language on its face does not provide a sufficient threshold or clarity as to what would sufficiently be in the “style” of SAC loss, the Examiner turns to Applicants’ specification, e.g. [0004], for guidance, and the Examiner finds no teaching by Applicants to more definitively define what would constitute a teaching or feature in the “style” of SAC loss. Applicants’ language appears to feature a relative terminology without any way of defining what constitutes the proper bounds for inclusion into that same terminology. Hence, because the claim on its face is vague as to this definition, and the specification is otherwise silent in terms of providing additional clarification or even a meaningful definition, the effect is that the language as noted, and hence the claim, is vague and indefinite, and therefore rejected.
The other independent claims 5 and 7-8 also recite these same variables/parameters, as noted above by the Examiner per the first matter, without appearing to define them. Similarly, they also recite the same language referenced above, per the second matter. Hence, these additional independent claims are likewise rejected under the same rationale.
The dependent claims 2-4 and 6 depend from one of these aforementioned independent claims. Hence, these claims inherit the deficiency as articulated above with respect to the independent claims and do not appear to otherwise cure it. Accordingly, they are likewise rejected under the same rationale.
Examiner’s Comment
Given the rejection presented above, the Examiner at this time declines to formulate a concrete grounds for a prior art rejection formally. When the Examiner better understands the equations particularly relating to the updating of the second neural network, as recited in Applicants’ pending claims, the Examiner will reconsider whether such a rejection would be appropriate in view of the prior art obtained via the Examiner’s search.
For the sake of the record, the Examiner will provide remarks here that correlate the limiting features in the recited pending claims with teachings in the prior art obtained via search and Applicants’ IDS submission. Though this is not as concrete a formal prior art rejection, the Examiner hopes it will at least provide some value to the record and to further expedite prosecution once the claim is better clarified for the Examiner.
The independent claims broadly involve a reinforcement learning framework directed to learning of a policy for an agent. Involved in this are recited first and second neural network instances, auxiliary parameters, and a policy. Further involve is an iterative process with an undefined/unspecified termination condition, the iterative process including the sampling of state and action pairs and associated rewards and new states, and the sampling of first actions for current states and second actions of the new sampled states, also in Accordance with the policy.
The Examiner believes the basic bones of this claim, as characterized above, are common to many actor-critic reinforcement learning frameworks considered by the Examiner: as but one example, see, e.g., Non-Patent Literature “Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor” (Haarnoja), particularly section 4.2 and its Algorithm 1.
These elements of the claimed invention are also found in the prior art provided by Applicants by way of their IDS, see e.g., Non-Patent Literature “Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems” (Levine) {section 2.1 and algorithms 2-3}, Non-Patent Literature “A Survey on Offline Reinforcement Learning: Taxonomy, Review, and Open Problems” (Prudencio) {section II, A-C}, and Non-Patent Literature “The Fixed Points of Off-Policy TD” (Kolter) {section 2}.
Further the iterative process mentioned above as captured in the independent claims, features are computed from the penultimate layer of the recited first neural network based on the sampled states and actions. Moreover, these features are then used to compute values LA,B and Lg that are in turn used to update “auxiliary parameters” A and B as recited.
The neural networks employed by Haarnoja to model its Q-function and policy would evaluate, in accordance with Algorithm 1 for example, and hence it would expected for features to be computed in their layers, including the penultimate layer of its actor/policy network (as recited).
Non-Patent Literature “A Mathematical Framework for Transformer Circuits” (“Elhage”) establishes the basic foundational idea that logits are the raw values computed in the layers of a neural network, and specifically a logit computed in the penultimate layer would serve as the basis for a final layer activation in a first network’s output, e.g. in a framework such as Haarnoja and Applicants’ invention.
The Examiner reasons that using the logit, rather than the output value corresponding to the logit, would serve in the interest of minimizing or mitigating unwanted effects of distributional shift as is known in the state of the art relating to reinforcement learning generally. See, e.g., Prudencio’s section I, 4th-5th paragraphs; and Levine’s section I, 4th paragraph.
Regarding Applicants’ recitation of “auxiliary parameters” specifically, the Examiner understands this to be parameters that would constrain the framework’s underlying Markov Decision Process, such as taught in Non-Patent Literature “Augmented Lagrangian Method for Instantaneously Constrained Reinforcement Learning Problems” (Li) {section II’s discussion of constraints including A and B matrices examples as provided per its Example 1, and involvement of such constraints into the computing of a Lagrange multiplier per section III}.
The Examiner reasons that these features discussed here are the equivalent of the auxiliary parameters as recited in Applicants’ independent claims, and essentially perform a function for constraining the decision process involved in the evaluation of states and actions in a way that is understood to be preferably bounded.
The Examiner believes the additional references also touch upon this concept of a constrained Markov Decision process that may read on Applicants’ recitation of auxiliary parameters as involved in the claimed invention’s updating aspects:
Non-Patent Literature “Model-based Safe Deep Reinforcement Learning via a Constrained Proximal Policy Optimization Algorithm” (Jayant), Non-Patent Literature “Constrained Model-based Reinforcement Learning with Robust Cross-Entropy Method” (Liu), Non-Patent Literature “Responsive Safety in Reinforcement Learning by PID Lagrangian Methods” (Stooke), Non-Patent Literature “Safe Model-Based Reinforcement Learning with an Uncertainty-Aware Reachability Certificate” (Yu).
The Examiner is particularly interested in the differentiating of Applicants’ invention from Li and the further references noted just above, which related to the constrained Markov Decision Making Process that the Examiner believes includes features that appear to read on Applicants’ recitation of auxiliary parameters per the Independent claims. If Applicants believe an interview is helpful, the Examiner would welcome the opportunity to better understand this aspect.
Further the iterative process mentioned above as captured in the independent claims, a loss is computed to update the first neural network’s parameters, and a soft actor critic loss is used to update the policy’s parameters.
The Examiner reasons that Haarnoja’s section 4.2 teaches both update steps. See, e.g., the section’s 2nd paragraph discussing error minimization as part of its stabilized training approach. The section generally is directed to soft policy iteration inclusive of both updating the Q-function and the policy, which the Examiner reasons are equivalent to the actor and critic roles played by the recited first and second neural networks.
Regarding the dependent claims: Haarnoja’s section 4.1 teaches applying a Bellman operator to update the policy, which the Examiner reasons could read onto claim 2; and Non-Patent Literature “Conservative Q-Learning for Offline Reinforcement Learning” (Kumar, as provided via Applicants’ IDS submission) appears to read onto claim 3.
The Examiner believes claim 4 is likely allowable subject matter but will reevaluate this finding upon Applicants’ Reply with clarifications, remarks, etc. to better understand the claimed invention.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHOURJO DASGUPTA whose telephone number is (571) 272-7207. The examiner can normally be reached M-F 8am-5pm CST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Tamara Kyle can be reached at 571 272 4241. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SHOURJO DASGUPTA/Primary Examiner, Art Unit 2144