DETAILED ACTION
1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
2. This communication is in response to the Applicant’s submission filed 25 March 2024, where:
Claims 1-9 have been amended through a preliminary amendment filed 25 March 2024.
Claim 10 has been cancelled through a preliminary amendment filed 25 March 2024.
Claims 1-9 are pending.
Claims 1-9 are rejected.
Information Disclosure Statement
3. Information disclosure statement(s) were submitted on 25 March 2024. The submission complies with the provisions of 37 CFR 1.97. Accordingly, the Examiner considered the information disclosure statement.
Specification
4. The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed.
Claim Rejections - 35 U.S.C. § 112
5. The following is a quotation of 35 U.S.C. § 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
6. Claim 7 is rejected under 35 U.S.C. 112(b) as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor regards as the invention.
Claim 7, line 4, recites “the number of Q-functions.” There is insufficient antecedent basis for this limitation in the claim.
Claim 7, line 5, recites “the number of operation layers.” There is insufficient antecedent basis for this limitation in the claim.
Claim 7, line 8, recites “the number of layer normalization layers.” There is insufficient antecedent basis for this limitation in the claim.
Claim 7, lines 8-9, recites “the output of the previous layer.” There is insufficient antecedent basis for this limitation in the claim.
Claim 7, lines 10-11, recites “the number of times the value propagation calculation is performed according to the operation mode of the control target.” There is insufficient antecedent basis for this limitation in the claim.
Claim Rejections - 35 U.S.C. § 101
7. 35 U.S.C. § 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
8. Claims 1-9 are rejected under 35 U.S.C. § 101 because the claimed invention is directed to an abstract idea without significantly more.
Claim 1 recites a learning device, which is a machine, and thus one of the statutory categories of patentable subject matter. (35 U.S.C. § 101).
However, under Step 2A Prong One, the claim recites the limitations of “calculate each of a plurality of second evaluation values that include noise using a plurality of evaluation models, each of which calculates, on the basis of both a second state resulting from a first action performed by a control target in a first state, and a second action calculated from the second state using a policy model, a second evaluation value obtained by including noise in an index value indicating the result of evaluating the second action in the second state.” The activity of “calculates” contain limitations that can practically be performed in the human mind, including, for example, observations, evaluations, judgments, and opinions, and accordingly, are a mental process, (MPEP § 2106.04(a)(2) sub III), which is one of the groupings of abstract ideas. (MPEP § 2106.04(a)(2)). Thus, claim 1 recites an abstract idea.
Under Step 2A Prong Two, the claim as a whole is not integrated into a practical application, because the additional elements recited in the claim beyond the identified judicial exception include “at least one memory configured to store instructions,” and “at least one processor,” which are recited at a high-level of generality, and accordingly, are generic computer components used to implement the abstract idea, (MPEP § 2106.05(f)), that do not serve to integrate the abstract idea into a practical application. The claim also recites “a plurality of evaluation models,” and “a policy model,” which are recited at a high-level of generality, and accordingly, are generic computer components used to implement the abstract idea, (MPEP § 2106.05(f)), that do not serve to integrate the abstract idea into a practical application.
The claim also recites “update the policy model or parameters of the policy model on the basis of the smallest of the plurality of second evaluation values calculated respectively and a first evaluation value, which is an index value indicating the result of evaluating the first action in the first state.” The activity of “update” is a post-production insignificant extra-solution activity of storing data in memory, (MPEP § 2160.05(g)), that does not serve to integrate the abstract idea into a practical application. Therefore, claim 1 is directed to the abstract idea.
Finally, under Step 2B, the additional elements, taken alone or in combination, do not represent significantly more than the abstract idea itself. The additional elements recited in the claim beyond the identified judicial exception include “at least one memory configured to store instructions,” and “at least one processor,” which are recited at a high-level of generality, and accordingly, are generic computer components used to implement the abstract idea, (MPEP § 2106.05(f)), that do not amount to significantly more than the abstract idea. The claim also recites “a plurality of evaluation models,” and “a policy model,” which are recited at a high-level of generality, and accordingly, are generic computer components used to implement the abstract idea, (MPEP § 2106.05(f)), that do not serve to integrate the abstract idea into a practical application.
The claim also recites “update the policy model or parameters of the policy model on the basis of the smallest of the plurality of second evaluation values calculated respectively and a first evaluation value, which is an index value indicating the result of evaluating the first action in the first state.” The activity of “update” is a well-understood, routine, and conventional activity of storing information in memory, (MPEP § 2106.05(d) sub II.iv), that does not amount to significantly more than the abstract idea. Therefore, claim 1 is subject-matter ineligible.
Claim 2 depends directly or indirectly from claim 1. The claim recites more details or specifics to the abstract idea of “calculate,” where “in an operation result based on information pertaining to the policy model, state information pertaining to the first state, and state information pertaining to the second state, and accordingly, is merely more specific to the abstract idea. Also, the additional elements of the claim does not serve to integrate the abstract idea into integrated into a practical application, (see MPEP § 2106.04(d)), nor do the additional elements amount to significantly more than the abstract idea, (MPEP § 2106.05 sub I; see also MPEP § 2106.05(a) – (h)), and thus, the claim recites no more than the abstract idea. Therefore, claim 2 is subject-matter ineligible.
Claim 3 depends directly or indirectly from claim 1. The claim recites more details or specifics to the abstract idea of “calculate,” to “perform an operation to cause inclusion of the noise and an operation that normalizes the result of the operation to cause inclusion of the noise,” and accordingly, are merely more specific to the abstract idea. Also, the additional elements of the claim does not serve to integrate the abstract idea into integrated into a practical application, (see MPEP § 2106.04(d)), nor do the additional elements amount to significantly more than the abstract idea, (MPEP § 2106.05 sub I; see also MPEP § 2106.05(a) – (h)), and thus, the claim recites no more than the abstract idea. Therefore, claim 3 is subject-matter ineligible.
Claim 4 depends directly or indirectly from claim 1. The claim recites more details or specifics to the abstract idea of “calculate,” to “normalize the result of the operation to cause inclusion of the noise by using a layer normalization layer,” and accordingly, is merely more specific to the abstract idea. Also, the additional elements of the claim does not serve to integrate the abstract idea into integrated into a practical application, (see MPEP § 2106.04(d)), nor do the additional elements amount to significantly more than the abstract idea, (MPEP § 2106.05 sub I; see also MPEP § 2106.05(a) – (h)), and thus, the claim recites no more than the abstract idea. Therefore, claim 4 is subject-matter ineligible.
Claim 5 depends directly or indirectly from claim 1. The claim further recites the limitations of “perform a weighting operation on the policy information pertaining to the first action and the state information pertaining to the first state,” “add noise to the result of the weighting operation,” “normalize the value of the result of the operation to which the noise is added,” “identify the normalized results according to predetermined identification rules,” and “generate an evaluation value indicating the learning status based on the identification results.” The activities of “perform,” “add noise,” “normalize the value,” “identify the normalized results,” and “generate an evaluation value,” contain limitations that can practically be performed in the human mind, including, for example, observations, evaluations, judgments, and opinions, and accordingly, are a mental process, (MPEP § 2106.04(a)(2) sub III), which is one of the groupings of abstract ideas. (MPEP § 2106.04(a)(2)). Also, the additional elements of the claim does not serve to integrate the abstract idea into integrated into a practical application, (see MPEP § 2106.04(d)), nor do the additional elements amount to significantly more than the abstract idea, (MPEP § 2106.05 sub I; see also MPEP § 2106.05(a) – (h)), and thus, the claim recites no more than the abstract idea. Therefore, claim 5 is subject-matter ineligible.
Claim 6 depends directly or indirectly from claim 1. The claim further recites the limitations of “perform a weighting operation on the policy information and the state information pertaining to the first state,” “add noise to the result of the weighting operation,” “normalize the value of the result of the operation to which the noise is added,” “identify the normalized results according to predetermined identification rules,” and “generate an evaluation value indicating the learning status using a Q-function that generates an evaluation value indicating the learning status based on the identification result.” The activities of “perform,” “add noise,” “normalize the value,” “identify the normalized results,” and “generate an evaluation value,” contain limitations that can practically be performed in the human mind, including, for example, observations, evaluations, judgments, and opinions, and accordingly, are a mental process, (MPEP § 2106.04(a)(2) sub III), which is one of the groupings of abstract ideas. (MPEP § 2106.04(a)(2)). Also, the additional elements of the claim does not serve to integrate the abstract idea into integrated into a practical application, (see MPEP § 2106.04(d)), nor do the additional elements amount to significantly more than the abstract idea, (MPEP § 2106.05 sub I; see also MPEP § 2106.05(a) – (h)), and thus, the claim recites no more than the abstract idea. Therefore, claim 6 is subject-matter ineligible.
Claim 7 depends directly or indirectly from claim 1. The claim further recites the limitation of “receive information of at least any one of the following: the number of Q-functions involved in the update; the number operation layers that add the noise in the Q-function; the number of layer normalization layers that normalize the output based on the output of the previous layer in the Q-function; and the number of times the value propagation calculation is performed according to the operation mode of the control target, and wherein the at least one processor is configured to execute the instructions to perform the value propagation operation using the received information.” The activity of “receive information” is a pre-processing insignificant extra-solution activity of receiving data, (MPEP § 2106.05(g)), that does not serve to integrate the abstract idea into a practical application. Also, the activity of “receive information” is a well-understood, routine, and conventional activity of receiving data over a network, (MPEP § 2106.05(d) sub II.i), that does not amount to significantly more than the abstract idea. Therefore, claim 7 is subject-matter ineligible.
Claim 8 recites a control system, which is a machine, and thus one of the statutory categories of patentable subject matter. (35 U.S.C. § 101).
However, under Step 2A Prong One, the claim recites the limitations of “calculate each of a plurality of second evaluation values that include noise using a plurality of evaluation models, each of which calculates, on the basis of both a second state resulting from a first action performed by a control target in a first state, and a second action calculated from the second state using a policy model, a second evaluation value obtained by including noise in an index value indicating the result of evaluating the second action in the second state.” The activity of “calculates” contain limitations that can practically be performed in the human mind, including, for example, observations, evaluations, judgments, and opinions, and accordingly, are a mental process, (MPEP § 2106.04(a)(2) sub III), which is one of the groupings of abstract ideas. (MPEP § 2106.04(a)(2)). Thus, claim 8 recites an abstract idea.
Under Step 2A Prong Two, the claim as a whole is not integrated into a practical application, because the additional elements recited in the claim beyond the identified judicial exception include “at least one memory configured to store instructions,” and “at least one processor,” which are recited at a high-level of generality, and accordingly, are generic computer components used to implement the abstract idea, (MPEP § 2106.05(f)), that do not serve to integrate the abstract idea into a practical application. The claim also recites “a plurality of evaluation models,” and “a policy model,” which are recited at a high-level of generality, and accordingly, are generic computer components used to implement the abstract idea, (MPEP § 2106.05(f)), that do not serve to integrate the abstract idea into a practical application.
The claim also recites “update the policy model or parameters of the policy model on the basis of the smallest of the plurality of second evaluation values calculated respectively and a first evaluation value, which is an index value indicating the result of evaluating the first action in the first state.” The activity of “update” is a post-production insignificant extra-solution activity of storing data in memory, (MPEP § 2160.05(g)), that does not serve to integrate the abstract idea into a practical application. Therefore, claim 8 is directed to the abstract idea.
Finally, under Step 2B, the additional elements, taken alone or in combination, do not represent significantly more than the abstract idea itself. The additional elements recited in the claim beyond the identified judicial exception include “at least one memory configured to store instructions,” and “at least one processor,” which are recited at a high-level of generality, and accordingly, are generic computer components used to implement the abstract idea, (MPEP § 2106.05(f)), that do not amount to significantly more than the abstract idea. The claim also recites “a plurality of evaluation models,” and “a policy model,” which are recited at a high-level of generality, and accordingly, are generic computer components used to implement the abstract idea, (MPEP § 2106.05(f)), that do not serve to integrate the abstract idea into a practical application.
The claim also recites “update the policy model or parameters of the policy model on the basis of the smallest of the plurality of second evaluation values calculated respectively and a first evaluation value, which is an index value indicating the result of evaluating the first action in the first state.” The activity of “update” is a well-understood, routine, and conventional activity of storing information in memory, (MPEP § 2106.05(d) sub II.iv), that does not amount to significantly more than the abstract idea. Therefore, claim 8 is subject-matter ineligible.
Claim 9 recites a learning method, which is a process, and thus one of the statutory categories of patentable subject matter. (35 U.S.C. § 101).
However, under Step 2A Prong One, the claim recites the limitations of “calculate each of a plurality of second evaluation values that include noise using a plurality of evaluation models, each of which calculates, on the basis of both a second state resulting from a first action performed by a control target in a first state, and a second action calculated from the second state using a policy model, a second evaluation value obtained by including noise in an index value indicating the result of evaluating the second action in the second state.” The activity of “calculates” contain limitations that can practically be performed in the human mind, including, for example, observations, evaluations, judgments, and opinions, and accordingly, are a mental process, (MPEP § 2106.04(a)(2) sub III), which is one of the groupings of abstract ideas. (MPEP § 2106.04(a)(2)). Thus, claim 9 recites an abstract idea.
Under Step 2A Prong Two, the claim as a whole is not integrated into a practical application, because the additional elements recited in the claim beyond the identified judicial exception include “a computer,” which is recited at a high-level of generality, and accordingly, is a generic computer component used to implement the abstract idea, (MPEP § 2106.05(f)), that does not serve to integrate the abstract idea into a practical application. The claim also recites “a plurality of evaluation models,” and “a policy model,” which are recited at a high-level of generality, and accordingly, are generic computer components used to implement the abstract idea, (MPEP § 2106.05(f)), that do not serve to integrate the abstract idea into a practical application.
The claim also recites “update the policy model or parameters of the policy model on the basis of the smallest of the plurality of second evaluation values calculated respectively and a first evaluation value, which is an index value indicating the result of evaluating the first action in the first state.” The activity of “update” is a post-production insignificant extra-solution activity of storing data in memory, (MPEP § 2160.05(g)), that does not amount to significantly more than the abstract idea. Therefore, claim 9 is directed to the abstract idea.
Finally, under Step 2B, the additional elements, taken alone or in combination, do not represent significantly more than the abstract idea itself. The additional elements recited in the claim beyond the identified judicial exception include “a computer,” which is recited at a high-level of generality, and accordingly, is a generic computer component used to implement the abstract idea, (MPEP § 2106.05(f)), that does not amount to significantly more than the abstract idea. The claim also recites “a plurality of evaluation models,” and “a policy model,” which are recited at a high-level of generality, and accordingly, are generic computer components used to implement the abstract idea, (MPEP § 2106.05(f)), that do not amount to significantly more than the abstract idea.
The claim also recites “update the policy model or parameters of the policy model on the basis of the smallest of the plurality of second evaluation values calculated respectively and a first evaluation value, which is an index value indicating the result of evaluating the first action in the first state.” The activity of “update” is a well-understood, routine, and conventional activity of storing information in memory, (MPEP § 2106.05(d) sub II.iv), that does not amount to significantly more than the abstract idea. Therefore, claim 9 is subject-matter ineligible.
Claim Rejections – 35 U.S.C. § 103
9. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
10. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
11. Claims 1, 8, and 9 are rejected under 35 U.S.C. 103 as being unpatentable over Wang et al., “An Actor-Critic Algorithm using Cross Evaluation of Value Functions,” IJRA (2018) [hereinafter Wang] in view of Fujimoto et al., “Addressing Function Approximation Error in Actor-Critic Methods,” arXiv (2018) [hereinafter Fujimoto].
Regarding claims 1, 8, and 9, Wang teaches [a] learning device (Wang, Abstract, teaches “learning a global optimal policy [(that is, a learning device )]”) of claim 1, [a] control system (Wang at p. 40, “2. Actor-Critic Method,” teaches a reinforcement learning method [(that is, a control system )] of claim 8, and [a] learning method executed by a computer (Wang at p. 42, “4. An Actor-Critic Algorithm based on Cross Evaluation of Value Functions (DVCAC),” second paragraph, teaches that in “solving reinforcement learning problems,. . . [u]sing the state value function can reduce the amount of storage [(that is, “storage” inherently is at least one memory to store instructions)] and computation [(that is, “computation” is inherently at least one processor configured to execute)] required for parameter updates, but the calculation of the current optimal action is significantly increased”) of claim 9, comprising:
at least one memory configured to store instructions; and at least one processor configured to execute (Wang at p. 42, “4. An Actor-Critic Algorithm based on Cross Evaluation of Value Functions (DVCAC),” second paragraph, teaches that in “solving reinforcement learning problems,. . . [u]sing the state value function can reduce the amount of storage [(that is, “storage” inherently is at least one memory to store instructions)] and computation [(that is, “computation” is inherently at least one processor configured to execute)] required for parameter updates, but the calculation of the current optimal action is significantly increased”) the instructions to:
calculate each of a plurality of second evaluation values that include noise using a plurality of evaluation models (Wang at p. 42, “4. An Actor-Critic Algorithm based on Cross Evaluation of Value Functions (DVCAC),” fourth paragraph, teaches “method for updating the temporal difference of the first set of value functions is shown in equation (10).
PNG
media_image1.png
60
910
media_image1.png
Greyscale
Θ1 represents the policy parameter corresponding to the first set of value functions. Θ2 represents the second one. In equation (10), the value function of the next state is computed by the second set of evaluation values [(that is calculate each of a plurality of second evaluation values)], and the value function of the current state is computed by the first set of evaluation values, so that the first set of value functions will be affected by the second one when they are updated”; Wang at p. 45, “5. Analysis of Experiment Results,” second paragraph, teaches that “[e]ach time the movement in both x axis and y axis will be affected by noise. The noise follows a normal distribution, with the mean value of 0, and the standard deviation of 0.01 [(that is, each of the second evaluation values that include noise using a plurality of evaluation models)]”),
each of which calculates, on the basis of both a second state resulting from a first action performed by a control target in a first state (Wang at p. 39, “1. Introduction,” first paragraph, teaches “maximization is required to find the optimal action of the current state, and based on this action, algorithms continuously update the state-action value functions [(that is, such “maximization” is a control target in a first state)]”; Wang at p. 40, “2. Actor-Critic Method,” second paragraph, teaches “the actor-critic approach consists of both the actor and the critic, which need to maintain value functions and policy at the same time. The actor used to select actions, and the critic can comment on the selected action good or bad. The way that actors choose actions is not based on the current value function, but their own policy. Critics' comments generally take the form of temporal difference errors, which is calculated from the current value function. The temporal difference error is the only output of the critic and drives all the learning between the actor and the critic [19, 20]”; Wang, Algorithm 1 at step 7, teaches “in state x, perform action µ, get [reward] r and the next [(that is, second)] state x’ [(that is, each of which calculates, on the basis of both a second state resulting from a first action performed by a control target in a first state)]), and
a second action calculated from the second state using a policy model (Wang, Algorithm 1, teaches a cross evaluation of value functions [Examiner annotations in dashed-line text boxes]:
PNG
media_image2.png
577
810
media_image2.png
Greyscale
Wang at p. 46, “6. Conclusion,” first paragraph, teaches “[w]hen a set of evaluation functions is updated, another set of evaluation functions is used to evaluate the value function of the next state or the next state action pair [(that is, a second action calculated from the second state using a policy model)]”),
a second evaluation value obtained by including noise (Wang at p. 46, “6. Conclusion,” first paragraph, teaches “[w]hen a set of evaluation functions is updated, another set of evaluation functions is used to evaluate the value function of the next state or the next state action pair [(that is, a second evaluation value)]”; Wang at p. 43, “5. Analysis of Experiment Results,” first paragraph, teaches “In each state, the agent can move in any directions, and the distance between each move is fixed at 0.05. The moves in x axis and y axis are subjected to noise interference. The noise follows normal distribution with average 0 and the standard deviation 0.01 [(that is, a second evaluation value obtained by including noise)]”)) in an index value indicating the result of evaluating the second action in the second state (Wang, Fig. 3, teaches the reward values used as a criterion for evaluating learning algorithms [Examiner annotations in dashed-line text boxes]:
PNG
media_image3.png
424
926
media_image3.png
Greyscale
Wang at p. 44 & Fig. 3, “5. Analysis of Experiment Results,” second paragraph, teaches “The reward value can be used as a criterion for evaluating learning algorithms [(that is, “episodes” are an index value indicating the result of evaluating the second action in the second state)]”); and
update the policy model or parameters of the policy model (Wang at p. 42, “4. An Actor-Critic Algorithm based on Cross Evaluation of Value Functions (DVCAC),” third paragraph, teaches that in “continuous space, state-valued functions are approximated by linear functions. The representation method is shown in equation (7).
PNG
media_image4.png
62
896
media_image4.png
Greyscale
And the updating of the policy parameter is shown in equation (8).
PNG
media_image5.png
53
880
media_image5.png
Greyscale
[(that is, update the policy model or parameters of the policy model)]”), . . .
which is an index value indicating the result of evaluating the first action in the first state (Wang, Fig. 5, teaches accumulated rewards including a random noise following normal distribution in the reward part [Examiner annotations in dashed line text-boxes]:
PNG
media_image6.png
422
734
media_image6.png
Greyscale
Wang at p. 45, “5. Analysis of Experiment Results,” fourth paragraph, teaches that “from Figure 5, in DVCAC, the reward stability is much better than the CACLA algorithm [(that is, the “accumulated reward episodes” which is an index value indicating the result of evaluating the first action in the first state)]”).
Though Wang teaches a temporal difference error that updates a policy model for a reinforcement learning actor-critic variant, Wang, however, does not explicitly teach –
* * *
[update the policy model or parameters of the policy model] on the basis of the smallest of the plurality of second evaluation values calculated respectively and a first evaluation value, . . . .
But Fujimoto teaches –
* * *
[update the policy model or parameters of the policy model] on the basis of the smallest of the plurality of second evaluation values calculated respectively and a first evaluation value (Fujimoto, right column of p. 4, “4.2 Clipped Double Q-Learning for Actor-Critic,” first full paragraph, teaches “As _-1 optimizes with respect to
Q
θ
1
, using an independent estimate in the target update of
Q
θ
1
would avoid the bias introduced by the policy update. . . . [W]e propose to simply upper-bound the less biased value estimate
Q
θ
2
by the biased estimate
Q
θ
1
. This results in taking the minimum between the two estimates, to give the target update of our Clipped Double Q-learning algorithm [(that is, [update the policy model or parameters of the policy model] on the basis of the smallest of the plurality of second evaluation values calculated respectively and a first evaluation value)]”), . . . .
Wang and Fujimoto are from the same or similar field of endeavor. Wang teaches overcome the difficulty of learning a global optimal policy caused by maximization bias in a continuous space, an actor-critic algorithm for cross evaluation of double value function is proposed. Fujimoto teaches an actor-critic reinforcement learning algorithm that builds on Double Q-learning, by taking the minimum value between a pair of critics to limit overestimation.
Thus, it would have been obvious to a person having ordinary skill in the art to a person having ordinary skill in the art as of the effective filing date of the Applicant’s invention to modify Wang pertaining to a temporal difference error for training a policy model with the minimum of two estimates to update the policy model of Fujimoto.
The motivation to do so is to “draw the connection between target networks and overestimation bias, and suggest delaying policy updates to reduce per-update error and further improve performance.” (Fujimoto, Abstract).
Regarding claim 2, the combination of Wang and Fujimoto teaches all of the limitations of claim 1, as described above in detail.
Wang teaches -
wherein the at least one processor is configured to execute the instructions to
calculate the second evaluation value obtained by including the noise in an operation result based on information pertaining to the policy model, state information pertaining to the first state, and state information pertaining to the second state (Wang at p. 41, “2. Actor-Critic Method,” first paragraph, teaches “[t]he actor selects a performing action in the current state based on a lookup table. The lookup table stores the corresponding preference p(x, µ) for each state action pair. x, µ represent state and action. According preference p, the selection probability of each action can be computed under the current state. The action selection method is Gibbs softmax. As shown in Equation 1.
PNG
media_image7.png
115
896
media_image7.png
Greyscale
Critic uses temporal difference error to comment on the quality of actions. Once the critic gets the action, the temporal difference error is calculated from the value function. The temporal difference error δ is calculated with equation (2). V is state value function, ϒ ϵ [0,1) is a discount and r represents an immediate reword [sic].
PNG
media_image8.png
225
1065
media_image8.png
Greyscale
[(that is, the “temporal difference error” is an operation result based on information pertaining to the policy model, state information pertaining to the first state, and state information pertaining to the second state)]”).
Regarding claim 3, the combination of Wang and Fujimoto teaches all of the limitations of claim 1, as described above in detail.
wherein the at least one processor is configured to execute the instructions to
Wang teaches -
perform an operation to cause inclusion of the noise (Wang at p. 43, “5. Analysis of Experiment Results,” first paragraph, teaches that the “moves in x axis and y axis are subjected to noise interference [(that is, perform an operation to cause inclusion of the noise)]”) and
an operation that normalizes the result of the operation to cause inclusion of the noise (Wang at p. 45, “5. Analysis of Experiment Results,” second paragraph, teaches that “[e]ach time the movement in both x axis and y axis will be affected by noise. The noise follows a normal distribution, with the mean value of 0, and the standard deviation of 0.01. If the movement exceeds the boundary, the agent remains on the boundary [(that is, “normal distribution” inherently is an operation that normalizes the result of the operation to cause inclusion of the noise)]”).
12. Claim 4 are rejected under 35 U.S.C. 103 as being unpatentable over Wang et al., “An Actor-Critic Algorithm using Cross Evaluation of Value Functions,” IJRA (2018) [hereinafter Wang] in view of Fujimoto et al., “Addressing Function Approximation Error in Actor-Critic Methods,” arXiv (2018) [hereinafter Fujimoto] and Ma et al., “Multimodal Image Registration with Deep Context Reinforcement Learning,” ResearchGate (2018) [hereinafter Ma].
Regarding claim 4, the combination of Wang and Fujimoto teaches all of the limitations of claim 1, as described above in detail.
Though Wang and Fujimoto teach the feature of noise inclusion through a normal distribution, the combination of Wang and Fujimoto, however, does not explicitly teach –
wherein the at least one processor is configured to execute the instructions to
normalize the result of the operation to cause inclusion of the noise by using a layer normalization layer.
But Ma teaches -
wherein the at least one processor is configured to execute the instructions to
normalize the result of the operation to cause inclusion of the noise by using a layer normalization layer (Ma at p. 2, “1. Introduction,” first full paragraph, teaches “high level features encode large contextual information which are robust against noise and other data variations [(that is, “data variations” are to cause the inclusion of the noise)]”; Ma at p. 4, “3.2 Training the Agent, first paragraph, teaches that “We add batch normalization layer after the input data layer to minimize the effect of intensity distribution discrepancy across different modalities [(that is, the “batch normalization layer” is normalize the result of the operation to cause inclusion of the noise by using a layer normalization layer)]”).
Wang, Fujimoto, and Ma are from the same or similar field of endeavor. Wang teaches overcome the difficulty of learning a global optimal policy caused by maximization bias in a continuous space, an actor-critic algorithm for cross evaluation of double value function is proposed. Fujimoto teaches an actor-critic reinforcement learning algorithm that builds on Double Q-learning, by taking the minimum value between a pair of critics to limit overestimation. Ma teaches a learning-based system derived from deep Q-learning that automatically extracts compact feature representations to reduce the appearance discrepancy between depth and CT data.
Thus, it would have been obvious to a person having ordinary skill in the art to a person having ordinary skill in the art as of the effective filing date of the Applicant’s invention to modify the combination of Wang and Fujimoto pertaining to a temporal difference error for training a policy model having a minimum of two estimates updating the policy model with batch normalization layer of Ma.
The motivation to do so is because “combining deep convolutional neural network with reinforcement learning, the deep reinforcement learning (DRL) has demonstrated superhuman performance in different applications.” (Ma at p. 2, “1. Introduction,” first full paragraph).
13. Claims 5-7 are rejected under 35 U.S.C. 103 as being unpatentable over Wang et al., “An Actor-Critic Algorithm using Cross Evaluation of Value Functions,” IJRA (2018) [hereinafter Wang] in view of Fujimoto et al., “Addressing Function Approximation Error in Actor-Critic Methods,” arXiv (2018) [hereinafter Fujimoto] and Winfried Lotzsch, “Using Deep Reinforcement Learning for the Continuous control of Robotic Arms,” arXiv (2018) [hereinafter Lotzsch].
Regarding claim 5, the combination of Wang and Fujimoto teaches all of the limitations of claim 1, as described above in detail.
Though Wang and Fujimoto teach the feature of noise inclusion through a normal distribution in a reinforcement learning environment, the combination of Wang and Fujimoto, however, does not explicitly teach –
perform a weighting operation on the policy information pertaining to the first action and the state information pertaining to the first state,
add noise to the result of the weighting operation,
normalize the value of the result of the operation to which the noise is added,
identify the normalized results according to predetermined identification rules, and
generate an evaluation value indicating the learning status based on the identification results.
But Lotzsch teaches -
wherein the at least one processor is configured to execute the instructions to:
perform a weighting operation on the policy information pertaining to the first action and the state information pertaining to the first state (Lotzsch. Listing 3.4, line 1, teaches to “randomly initialize critic network Q(s, a|θQ) and actor µ(s|θµ) with weights θQ and θµ [(that is, perform a weighting operation on the policy information pertaining to the first action and the state information pertaining to the first state)]”),
add noise to the result of the weighting operation (Lotzsch, Listing 3.4, line 8, teaches “Select action 𝑎𝑡 = 𝜇(𝑠𝑡|𝜃𝜇) + 𝒩𝑡 according to the current policy and exploration noise [(that is, “random process 𝒩” is to add noise to the result of the weighting operation)]”),
normalize the value of the result of the operation to which the noise is added (Lotzsch at p. 39, “5.1.1 Simulation with Matplotlib,” first full paragraph, teaches “[t]he total reward is normalized to lie in the interval [-1,1] [(that is, normalize the value of the result of the operation to which the noise is added)]”),
identify the normalized results according to predetermined identification rules (Lotzsch at p. 9, “3. Variants of Reinforcement Learning,” first paragraph, teaches that “[t]o form a learning algorithm, the action-value function is redefined with respect to an arbitrary policy 𝜋. The value of 𝑄𝜋(𝑠𝑡, 𝑎𝑡) corresponds to the expected return when taking action 𝑎𝑡 in step 𝑠𝑡 and following policy 𝜋 subsequently. The calculated estimates can then be used to improve the policy. Intuitively, for optimal 𝜋, the action-value function with respect to the policy converges to 𝑄*(𝑠𝑡, 𝑎𝑡). It is possible to estimate 𝑄𝜋(𝑠𝑡, 𝑎𝑡) for example using a Monte-Carlo estimate of the expected return after the end of each episode. This strategy is called offline learning, because it must wait until an episode has finished. However, it is much more practical to store estimates of 𝑄𝜋(𝑠𝑡, 𝑎𝑡) in an array 𝑄(𝑠𝑡, 𝑎𝑡) and reuse existing estimates for future states [(that is, to “reuse existing estimates” is identify the normalized results according to predetermined identification rules, which is called bootstrapping and yields an online learning rule (Sutton and Barto 1998) [(that is, the “online learning rule is according to predetermined identification rules)]”), and
generate an evaluation value indicating the learning status based on the identification results (Lotzsch at p. 20, “3.2.7 Improvements to DQN,” first paragraph, teaches a “loss function used by [Double DQN (D-DQN)] disentangles the action selection from the evaluation of the selected action by using the trained parameters 𝜃 to select future actions instead of those of the target network [(that is, the “loss function” is to generate an evaluation value indicating the learning status based on the identification results)]”).
Wang, Fujimoto, and Lotzsch are from the same or similar field of endeavor. Wang teaches overcome the difficulty of learning a global optimal policy caused by maximization bias in a continuous space, an actor-critic algorithm for cross evaluation of double value function is proposed. Fujimoto teaches an actor-critic reinforcement learning algorithm that builds on Double Q-learning, by taking the minimum value between a pair of critics to limit overestimation. Lotzsch teaches a double deep Q-network in deep reinforcement learning incorporating weights to raise reward impact to temporally close actions.
Thus, it would have been obvious to a person having ordinary skill in the art to a person having ordinary skill in the art as of the effective filing date of the Applicant’s invention to modify the combination of Wang and Fujimoto pertaining to a temporal difference error for training a policy model having a minimum of two estimates updating the policy model with the use of weights to raise reward impact of Lotzsch.
The motivation to do so is to “to reduce training time and eventually help the algorithm to converge.” (Lotzsch, Abstract).
Regarding claim 6, the combination of Wang, Fujimoto and Lotzsch teaches all of the limitations of claim 5, as described above in detail.
Lotzsch teaches -
wherein the at least one processor is configured to execute the instructions to:
perform a weighting operation on the policy information and the state information pertaining to the first state (Lotzsch. Listing 3.4, line 1, teaches to “randomly initialize critic network Q(s, a|θQ) and actor µ(s|θµ) with weights θQ and θµ [(that is, “(s, a) is a state-action pair,” which is to perform a weighting operation on the policy information and the state information pertaining to the first state)]”),
add noise to the result of the weighting operation (Lotzsch, Listing 3.4, line 8, teaches “Select action 𝑎𝑡 = 𝜇(𝑠𝑡|𝜃𝜇) + 𝒩𝑡 according to the current policy and exploration noise [(that is, “random process 𝒩” is to add noise to the result of the weighting operation)]”),
normalize the value of the result of the operation to which the noise is added (Lotzsch at p. 39, “5.1.1 Simulation with Matplotlib,” first full paragraph, teaches “[t]he total reward is normalized to lie in the interval [-1,1] [(that is, normalize the value of the result of the operation to which the noise is added)]”),
identify the normalized results according to predetermined identification rules (Lotzsch at p. 9, “3. Variants of Reinforcement Learning,” first paragraph, teaches that “[t]o form a learning algorithm, the action-value function is redefined with respect to an arbitrary policy 𝜋. The value of 𝑄𝜋(𝑠𝑡, 𝑎𝑡) corresponds to the expected return when taking action 𝑎𝑡 in step 𝑠𝑡 and following policy 𝜋 subsequently. The calculated estimates can then be used to improve the policy. Intuitively, for optimal 𝜋, the action-value function with respect to the policy converges to 𝑄*(𝑠𝑡, 𝑎𝑡). It is possible to estimate 𝑄𝜋(𝑠𝑡, 𝑎𝑡) for example using a Monte-Carlo estimate of the expected return after the end of each episode. This strategy is called offline learning, because it must wait until an episode has finished. However, it is much more practical to store estimates of 𝑄𝜋(𝑠𝑡, 𝑎𝑡) in an array 𝑄(𝑠𝑡, 𝑎𝑡) and reuse existing estimates for future states [(that is, to “reuse existing estimates” is identify the normalized results according to predetermined identification rules, which is called bootstrapping and yields an online learning rule (Sutton and Barto 1998) [(that is, the “online learning rule is according to predetermined identification rules)]”), and
generate an evaluation value indicating the learning status using a Q-function that generates an evaluation value indicating the learning status based on the identification result (Lotzsch at p. 20, “3.2.7 Improvements to DQN,” first paragraph, teaches a “loss function used by [Double DQN (D-DQN)] disentangles the action selection from the evaluation of the selected action by using the trained parameters 𝜃 to select future actions instead of those of the target network [(that is, the “loss function” is to generate an evaluation value indicating the learning status . . . based on the identification results)]”; Lotzsch at p. 13, “3.2.4 Using Neural Networks to Approximate the Q-function,” first paragraph, teaches “Estimating the Q-function can be done by a supervised learning algorithm with the targets for training given by the reinforcement learning algorithm. Therefore a loss function [(that is, using a Q-function that generates an evaluation value)] is introduced that drives the function approximator to output the correct Q-values, where 𝑄(𝑠𝑡, 𝑎; 𝜃) is a function, parametrized by learned parameters 𝜃 [(that is, the “correct Q-values” is indicating the learning status)]”).
Regarding claim 7, the combination of Wang and Fujimoto teaches all of the limitations of claim 1, as described above in detail.
Lotzsch teaches -
wherein the at least one processor is configured to execute the instructions to receive information of at least any one of the following:
the number of Q-functions involved in the update (Lotzsch at p. 48, “Chapter 6, Experimental Results,” second paragraph, teaches that “[f]or comparing the different variants of the DDPG algorithm, the same hyperparameters and network structure were used to obtain a reliable relative performance. We used the following hyperparameters [(that is, a “hyperparameter” is to receive information)] for all experiments:
PNG
media_image9.png
260
788
media_image9.png
Greyscale
[(that is, “update rate” and “maximum number of steps” pertain to the number of Q-functions involved in the update )]”);
the number operation layers that add the noise in the Q-function;
the number of layer normalization layers that normalize the output based on the output of the previous layer in the Q-function; and
the number of times the value propagation calculation is performed according to the operation mode of the control target, and
wherein the at least one processor is configured to execute the instructions to perform the value propagation operation using the received information.
Conclusion
14. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
(US Published Application 20220274255 to Okawa et al.) teaches A control apparatus according to one or more embodiments may calculate a first estimate value of the coordinates of an endpoint of a manipulator based on first sensing data obtained from a first sensor system, calculates a second estimate value of the coordinates of the endpoint of the manipulator based on second sensing data obtained from a second sensor system, and adjust a parameter value for at least one of a first estimation model or a second estimation model to reduce an error between the first estimate value and he second estimate value based on a gradient of the error.
(Hado van Hasselt, “Double Q-Learning,” (2010)) teaches well-known reinforcement learning algorithm Q-learning performs very poorly. This poor performance is caused by large overestimations of action values. These overestimations result from a positive bias that is introduced because Q-learning uses the maximum action value as an approximation for the maximum expected action value. We introduce an alternative way to approximate the maximum expected value for any set of random variables. The obtained double estimator method is shown to sometimes underestimate rather than overestimate the maximum expected value.
15. Any inquiry concerning this communication or earlier communications from the Examiner should be directed to KEVIN L. SMITH whose telephone number is (571) 272-5964. Normally, the Examiner is available on Monday-Thursday 0730-1730.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the Examiner by telephone are unsuccessful, the Examiner’s supervisor, KAKALI CHAKI can be reached on 571-272-3719. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/K.L.S./
Examiner, Art Unit 2122
/KAKALI CHAKI/Supervisory Patent Examiner, Art Unit 2122