DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 1-6 are pending for examination. Claims 1, 5, and 6 are independent.
Priority
Acknowledgment is made of Applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). The certified copy has been filed in parent Application No. JP2022-180115, filed on 11/10/2022.
Information Disclosure Statement
The information disclosure statement (IDS) is submitted on 10/26/2023. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Specification
The disclosure is objected to because of the following informalities:
[0007, 0008, 0009, 0076, 0078, 0082, 0083]: “the shaped reward and a discount factor of the second machine learning model to be leaned; (Examiner believes “leaned” should be corrected to “learned”.)
Appropriate correction is required.
Claim Objections
Claims 1-2 and 4-6 are objected to because of the following informalities:
Claims 1, 5, 6: “updating a policy of a second machine learning model using the shaped reward and a discount factor of the second machine learning model to be leaned;” (Examiner believes “leaned” should be corrected to “learned”.)
Claim 2: "wherein an objective function of (“the student model” lacks an antecedent basis.)
Claim 2: “wherein the processor updates the policy of the student model using the shaped reward, the discount factor, and the inverse temperature, and” (“the policy of the student model” lacks an antecedent basis.)
Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-6 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: The claim is a process, machine, manufacture, or composition of matter.
In the instant application, Claims 1-4 are directed to a machine, Claim 5 is directed to a method, and Claim 6 is directed to a machine. Thus, each of the claims falls within one of the four statutory categories (i.e. process, machine, manufacture, or composition of matter).
With respect to Claim 1:
2A Prong 1: The claim recites an abstract idea, law of nature, or natural phenomenon.
calculate a state value of the next state using the next state and a state value function of a first machine learning model; (This step of calculating a value using a value function is defined in Applicant specification [0019-0020] as using equations and is understood as a mathematical concept (i.e. mathematical calculations) — see MPEP 2106.04(a)(2)(I)(C).)
generate a shaped reward from the state value; (Generating a shaped reward from the state value is defined in Applicant specification [0063] as using an equation and is understood as a mathematical concept (i.e. mathematical calculations) — see MPEP 2106.04(a)(2)(I)(C).)
update a policy of a second machine learning model using the shaped reward and a discount factor of the second machine learning model to be leaned; (Updating a policy using the shaped reward and a discount factor is defined in Applicant specification [0023] as using an equation and is understood as a mathematical concept (i.e. mathematical calculations) — see MPEP 2106.04(a)(2)(I)(C).)
update the discount factor. (Updating a discount factor is defined in Applicant specification [0065] as using a function and is understood as a mathematical concept (i.e. mathematical calculations) — see MPEP 2106.04(a)(2)(I)(C).)
2A Prong 2: The judicial exception is not integrated into a practical application.
A learning device comprising: a memory configured to store instructions; and a processor configured to execute the instructions to: (The memory and processor are understood to be mere instructions to apply the exception using generic computer components — see MPEP 2106.05(f).)
acquire a next state and a reward as a result of an action; (Acquiring information is understood as insignificant extra-solution activity — see MPEP 2106.05(g).)
The additional elements as disclosed above alone or in combination do not integrate the judicial exception into practical application as they are insignificant extra-solution activities in combination with generic implementation of a model and computer component as a tool to perform the abstract idea above.
2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
A learning device comprising: a memory configured to store instructions; and a processor configured to execute the instructions to: (The memory and processor are understood to be mere instructions to apply the exception using generic computer components — see MPEP 2106.05(f).)
acquire a next state and a reward as a result of an action; (Acquiring information is understood as well-understood, routine, and conventional activity, exemplary of receiving or transmitting data over a network — see MPEP 2106.05(d)(II)(i).)
The additional elements as disclosed above alone or in combination do not recite significantly more than a judicial exception as they are well-understood, routine, and conventional activities previously known to the industry in combination with generic implementation of a model and computer components as a tool to perform the abstract idea above.
With respect to Claim 2:
2A Prong 1: The claim recites an abstract idea, law of nature, or natural phenomenon.
wherein an objective function of the student model includes an entropy regularization term; (The entropy regularization term is defined in Applicant specification [0046] as an equation and is understood as a mathematical concept (i.e. mathematical formulas) — see MPEP 2106.04(a)(2)(I)(B).)
wherein the entropy regularization term includes an inverse temperature which is a coefficient indicating a degree of regularization, (The inverse temperature is defined in Applicant specification [0047] as a hyperparameter of the regularization and is understood as a mathematical concept (i.e. mathematical relationships) — see MPEP 2106.04(a)(2)(I)(C).)
wherein the processor updates the policy of the student model using the shaped reward, the discount factor, and the inverse temperature, and (Updating a policy using the shaped reward, the discount factor, and the inverse temperature is defined in Applicant specification [0045-0047] as using an equation and is understood as a mathematical concept (i.e. mathematical calculations) — see MPEP 2106.04(a)(2)(I)(C).)
wherein the processor updates the inverse temperature. (Updating the inverse temperature is defined in Applicant specification [0050] as a continuously changing function factor and is understood as a mathematical concept (i.e. mathematical calculations) — see MPEP 2106.04(a)(2)(I)(C).)
2A Prong 2 & 2B: The claim does not recite additional material.
With respect to Claim 3:
2A Prong 1: The claim recites an abstract idea, law of nature, or natural phenomenon.
wherein the processor optimizes the discount factor so as to approach a predetermined true value. (Optimizing the discount factor is defined in Applicant specification [0048] as a continuously increasing stabilizing factor for the objective function and is understood as a mathematical concept (i.e. mathematical calculations) — see MPEP 2106.04(a)(2)(I)(C).)
2A Prong 2 & 2B: The claim does not recite additional material.
With respect to Claim 4:
2A Prong 1: The claim recites an abstract idea, law of nature, or natural phenomenon.
wherein the processor generates the shaped reward using the true value as the discount factor. (Generating a shaped reward using the true value as the discount factor is defined in Applicant specification [0039] as using an equation and is understood as a mathematical concept (i.e. mathematical calculations) — see MPEP 2106.04(a)(2)(I)(C).)
2A Prong 2 & 2B: The claim does not recite additional material.
With respect to Claim 5:
see the rejection of Claim 1 above. The same rationale applies.
2A Prong 2 & 2B: The claim recites another element “A learning method executed by a computer, comprising:” (The method executed by a computer is understood to be mere instructions to apply the exception using a generic computer or computer components — see MPEP 2106.05(f).)
With respect to Claim 6:
see the rejection of Claim 1 above. The same rationale applies.
2A Prong 2 & 2B: The claim recites another element “A non-transitory computer readable recording medium storing a program, the program causing a computer to execute processing of:” (The non-transitory computer readable recording medium that stores a program is understood to be mere instructions to apply the exception using a generic computer or computer components — see MPEP 2106.05(f).)
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1 and 3-6 are rejected under 35 U.S.C. 103 as being unpatentable over Fang et al. ("Target-driven visual navigation in indoor scenes using reinforcement learning and imitation learning") in view of Ng et al. ("Policy invariance under reward transformations: Theory and application to reward shaping") and in view of Kim et al. ("Adaptive Discount Factor for Deep Reinforcement Learning in Continuing Tasks with Uncertainty").
With respect to Claim 1:
Fang teaches: A learning device comprising: a memory configured to store instructions; and a processor configured to execute the instructions to: ([Pg. 171 Col. 1 ¶1] Fang describes the AI2-THOR environment and the use of TensorFlow which implies the use of a memory and processor.)
acquire a next state and a reward as a result of an action; ([Pg. 168 Col. 2 ¶2, Pg. 169 Fig. 1] Fang describes the Markov decision process (MDP) wherein “the agent considers the current observed images and the goal as input and obtains a new state and relevant reward from the environment by performing an action.”)
calculate a state value of the next state using the next state and a state value function of a first machine learning model; ([Pg. 169 Col. 1 ¶1, Eq. 2, Col. 2 Fig. 1] Fang explains the state-value function that the critic network (first machine learning / teacher model) uses to calculate the state value using the next state of the environment.)
update a policy of a second machine learning model using the ([Pg. 169 Col. 1 Para. 2, Eq. 4-5, Col. 2 Fig. 1] Fang shows how the actor network (second machine learning / student model) updates its policy using reward r and discount factor γ in relation to the critic network (first machine learning / teacher model) learning the value functions of the MDP using the next state from the environment.)
Fang does not teach: generate a shaped reward from the state value;
However, Ng teaches in the same field of endeavor: generate a shaped reward from the state value; ([Pg. 6 §4.3 Para. 1] Ng describes potential-based reward shaping using the state value similarly to Applicant definition in specification [0022])
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Fang’s actor-critic algorithm by further shaping the reward with a potential-based shaping function as taught by Ng. Doing so would hasten the learning process of the student model towards an optimal policy (Ng Pg. 3 Col. 2 ¶1).
Fang in view of Ng does not teach: update the discount factor.
However, Kim teaches in the same field of endeavor: update the discount factor. ([Pg. 7-9 §4.1] Kim discusses the adaptive discount factor in reinforcement learning wherein “we want to define an evaluation function of an agent’s performance according to the current discount factor. Additionally, based on the evaluation function, we want to find an update rule for the discount factor that can improve the agent’s performance.” (Pg. 7¶2))
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Fang in view of Ng’s actor-critic algorithm by gradually updating or adapting the discount factor as taught by Kim. Doing so would allow the student model to reduce overestimations of the reward and to adapt to the variable advantage function (Kim Pg. 7 ¶2-5).
With respect to Claim 3:
Fang in view of Ng and Kim teaches: The learning device according to claim 1,
Kim further teaches: wherein the processor optimizes the discount factor so as to approach a predetermined true value. ([Pg. 8 ¶1-3] Kim discloses how the algorithm increases a discount factor γ1 towards an upper bound γ2. This aligns with Applicant specification [0037-0043] in which Applicant describes the true discount factor γ* as an upper bound to discount factor γ greater than 0. [Pg. 8 ¶4] Kim notes that 0 < γ < 1.)
With respect to Claim 4:
Fang in view of Ng and Kim teaches: The learning device according to claim 1,
further teaches: wherein the processor generates the shaped reward using the true value as the discount factor. (As Applicant specification states that the discount factor γ has an exclusive upper bound of 1 and the true value γ* is higher than 0 and γ, examiner understands the true value to be a highest value of γ closest to 1. [Pg. 5 ¶1] Kim states a goal to “redesign the reward function according to the importance of the reward signal (reward shaping)” and to increase the value of the reward signal in light of the observation that “the optimal policy can very according to the discount factor.” [Pg. 6 ¶2] Kim discusses how the estimated reward sum may be lower when the discount factor is low, further associating the two concepts of reward shaping and discount factor.)
With respect to Claim 5:
Fang in view of Ng and Kim teaches: A learning method executed by a computer, comprising: ([Pg. 168 Col. 1 ¶1] Fang discloses a framework that combines the imitation learning and reinforcement learning methods into one learning method that improves the learning of a robot, implying the execution of the method being performed by a computer.)
Claim 5 is a claim that corresponds to Claim 1 and the remaining limitations are rejected for at least the same reasons therein.
With respect to Claim 6:
Fang in view of Ng and Kim teaches: A non-transitory computer readable recording medium storing a program, the program causing a computer to execute processing of: ([Pg. 171 Col. 1 ¶1] Fang describes the AI2-THOR environment and the use of TensorFlow which are programs and imply the use of a computer-readable recording medium to store those programs.)
Claim 6 is a claim that corresponds to Claim 1 and the remaining limitations are rejected for at least the same reasons therein.
Claim 2 is rejected under 35 U.S.C. 103 as being unpatentable over Fang in view of Ng, in view of Kim, and in view of Liu et al. ("Policy Optimization Reinforcement Learning with Entropy Regularization")
With respect to Claim 2:
Fang in view of Ng and Kim teaches: The learning device according to claim 1,
Fang further teaches: wherein an objective function of the student model includes an entropy regularization term; ([Pg. 169 Col. 2 Eq. 7, ¶1] Fang describes a policy loss function that includes an entropy regularization term. Fang’s policy loss function aligns with the Applicant objective function as it contains [Pg. 169 Col. 1 ¶5] an advantage function A, that, when further expanding the function Q(s,a) within A, is akin to Applicant reward equation (4) in specification [0039].)
wherein the entropy regularization term includes an inverse temperature which is a coefficient indicating a degree of regularization, ([Pg. 169 Col. 2 Eq. 7, ¶1] Fang describes parameter β as a coefficient that adjusts the step length or degree of the entropy regularization term.)
Fang in view of Ng and Kim does not explicitly teach: wherein the processor updates the policy of the student model using the shaped reward, the discount factor, and the inverse temperature, and
wherein the processor updates the inverse temperature.
However, Liu teaches in the same field of endeavor: wherein the processor updates the policy of the student model using the shaped reward, the discount factor, and the inverse temperature, and ([Pg. 2 Last Para, Eq. 2] Liu discloses an optimal model policy function that includes a discount factor γ, reward r(st, at), and inverse temperature coefficient α similarly to the objective function in Applicant specification [0046].)
wherein the processor updates the inverse temperature. ([Pg. 3 ¶1] Liu discusses the tuning or updating of the temperature parameter α that is equivalent to Applicant’s inverse temperature β.)
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Fang in view of Ng and Kim’s actor-critic algorithm by allowing the temperature coefficient of the entropy regularization term to be adjustable as taught by Liu. Doing so would allow the student model policy to further explore more avenues of optimal or near-optimal behavior and to probabilistically decide on an optimal avenue in situations where there are equally optimal policies (Liu Pg. 3 ¶1).
Conclusion
The prior art made of record and not relied upon is considered pertinent to Applicant's disclosure:
Sutton et al. (“Introduction to reinforcement learning”. MIT Press, Cambridge, (2018)) discloses in Section 3 the reinforcement learning problem, Markov Decision processes, and value functions.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to BRIAN D. BUI whose telephone number is (571)270-0463. The examiner can normally be reached Monday - Friday 8:00am - 5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, Applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, ABDULLAH AL KAWSAR can be reached at (571) 270-3169. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/BRIAN D. BUI/Examiner, Art Unit 2127
/ALEXEY SHMATOV/Supervisory Patent Examiner, Art Unit 2123