Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 06/21/2023 and 05/09/2024 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-9 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
STEP 1
Claim 1-5 are a learning device type claim. Claims 6-9 are a method type claim. Therefore, claims 1-9 are directed to either a process, machine, manufacture or composition of matter.
Regarding claim 1: 2A Prong 1:
estimate a trajectory that minimizes Wasserstein distance, which represents distance between probability distribution of a trajectory of an expert and probability distribution of a trajectory determined based on parameters of the reward function; and (mathematical concept – of estimate a trajectory that minimizes Wasserstein distance. See for example paragraph [0040] and Equation 8 of the instant application (e.g., mathematical calculation)).
update the parameters of the reward function to maximize the Wasserstein distance based on the estimated trajectory (mathematical concept – of update the parameters of the reward function to maximize the Wasserstein distance based on the estimated trajectory. For example, a person can update the parameters using gradient descent method. Further, applicant specification [0057] explicitly state “the parameters of the reward function (cost function) are updated by mathematical optimization” (e.g., mathematical calculation)).
2A Prong 2: This judicial exception is not integrated into a practical application.
Additional elements:
A learning device comprising: a memory storing instructions; and one or more processors configured to execute the instructions to: (This is directed to using computers or other machinery merely as a tool to perform an existing process. See MPEP 2106.05(f)).
accept input of a reward function whose features are set to satisfy a Lipschitz continuity condition; (This is understood to be insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g)).
The additional elements as disclosed above alone or in combination do not integrate the judicial exception into practical application as they are mere insignificant extra solution activity in combination of generic computer functions being implemented with generic computer elements in a high level of generality to perform the disclosed abstract idea above.
2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
A learning device comprising: a memory storing instructions; and one or more processors configured to execute the instructions to: (This is directed to using computers or other machinery merely as a tool to perform an existing process. See MPEP 2106.05(f)).
accept input of a reward function whose features are set to satisfy a Lipschitz continuity condition; ( This is directed to well understood, routine of receiving or transmitting data over a network. See MPEP 2106.05 (d)(II)).
The additional elements as disclosed above in combination of the abstract idea are not sufficient to amount to significantly more than the judicial exception as they are mere insignificant extra solution activity in combination of generic computer functions being implemented with generic computer elements in a high level of generality to perform the disclosed abstract idea above.
Regarding claim 2: Depends on claim 1, thus the rejection of claim 1 is incorporated.2A Prong 1:
...update the parameters of the reward function using a non-expansive mapping gradient method, which is an update rule based on a non-expansive mapping (mathematical concept – of update the parameters of the reward function using a non-expansive mapping gradient method, which is an update rule based on a non-expansive mapping. See for example paragraph [0048] & [0057] of the instant application (e.g., mathematical calculation)).
2A Prong 2 and 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
wherein the processor is configured to execute the instructions to... (This is directed to using computers or other machinery merely as a tool to perform an existing process. See MPEP 2106.05(f)).
Regarding claim 3: Depends on claim 1, thus the rejection of claim 1 is incorporated.2A Prong 1:
...update the parameters of the reward function with a step width less than or equal to a product of a value of a ratio of slope of Wasserstein distance at this update to slope of Wasserstein distance at one previous update and a step width at one previous update so that the Wasserstein distance after parameter update is larger (mathematical concept – of ...update the parameters of the reward function with a step width less than or equal to a product of a value of a ratio of slope of Wasserstein distance at this update to slope of Wasserstein distance at one previous update and a step width at one previous update so that the Wasserstein distance after parameter update is larger. See for example paragraph [0048-49], [0057] and Equations 12-14 of the instant application (e.g., mathematical calculation)).
2A Prong 2 and 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
wherein the processor is configured to execute the instructions to... (This is directed to using computers or other machinery merely as a tool to perform an existing process. See MPEP 2106.05(f)).
Regarding claim 4: Depends on claim 1, thus the rejection of claim 1 is incorporated.2A Prong 1:
...determine whether the Wasserstein distance converges or not; and (mental process – of determine whether the Wasserstein distance converges or not can be performed by the human mind with the help of pen and paper. For example, a human can evaluate the Wasserstein distance and determine if it converged or not (e.g., evolution & judgement )).
in a case where the Wasserstein distance is determined not to be convergent, estimate a trajectory that minimizes Wasserstein distance, which represents distance between probability distribution of a trajectory of an expert and probability distribution of a trajectory determined based on the updated parameters of the reward function, and update the parameters of the reward function so as to maximize the Wasserstein distance (mathematical concept – of estimate a trajectory that minimizes Wasserstein distance, which represents distance between probability distribution of a trajectory of an expert and probability distribution of a trajectory determined based on the updated parameters of the reward function, and update the parameters of the reward function so as to maximize the Wasserstein distance (e.g., mathematical calculation)).
2A Prong 2 and 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
wherein the processor is configured to execute the instructions to: (This is directed to using computers or other machinery merely as a tool to perform an existing process. See MPEP 2106.05(f)).
Regarding claim 5: Depends on claim 1, thus the rejection of claim 1 is incorporated.2A Prong 1: None.
2A Prong 2: This judicial exception is not integrated into a practical application.
Additional elements:
wherein the processor is configured to execute the instructions to... (This is directed to using computers or other machinery merely as a tool to perform an existing process. See MPEP 2106.05(f)).
...accept input of a reward function whose features are set to be linear functions (This is understood to be insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g)).
The additional elements as disclosed above alone or in combination do not integrate the judicial exception into practical application as they are mere insignificant extra solution activity in combination of generic computer functions being implemented with generic computer elements in a high level of generality to perform the disclosed abstract idea above.
2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
wherein the processor is configured to execute the instructions to... (This is directed to using computers or other machinery merely as a tool to perform an existing process. See MPEP 2106.05(f)).
...accept input of a reward function whose features are set to be linear functions ( This is directed to well understood, routine of receiving or transmitting data over a network. See MPEP 2106.05 (d)(II)).
The additional elements as disclosed above in combination of the abstract idea are not sufficient to amount to significantly more than the judicial exception as they are mere insignificant extra solution activity in combination of generic computer functions being implemented with generic computer elements in a high level of generality to perform the disclosed abstract idea above.
Regarding claim 6: is rejected under the same rational of claim 1. Claim 6 only recites the additional elements of A learning method comprising... which is directed to using computers or other machinery merely as a tool to perform an existing process. See MPEP 2106.05(f).
Regarding claim 7: See rejection of claim 2, same rational applies.
Regarding claim 8: is rejected under the same rational of claim 1. Claim 8 only recites the additional elements of A non-transitory computer readable information recording medium storing a learning program causing a computer to perform:... which is directed to using computers or other machinery merely as a tool to perform an existing process. See MPEP 2106.05(f).
Regarding claim 9: See rejection of claim 2, same rational applies.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-2, and 6-9 are rejected under 35 U.S.C. 103 as being unpatentable over Xiao et al. Wasserstein Adversarial Imitation Learning (hereinafter Xiao) in view of Zhang et al. Wasserstein Distance guided Adversarial Imitation Learning with Reward Shape Exploration in further view of Eto et al. US 2022/0343180 A1 (hereinafter Eto).
Regarding claim 1:
Xiao teaches accept input of a reward function whose features are set to satisfy a Lipschitz continuity condition; ( Xiao pg. 7 Algorithm 1 line 1, teaches input accept input of a reward function. Further, Xiao Abstract teaches a reward function with desirable properties (features) and teaches treating the reward function as Kantorovich potentials that is required to be Lipschitz(1)-continuous (pg. 4, sec: From apprenticeship learning to Wasserstein distance, para. 1-3 and pg. 4 Proposition 3.1.)).
estimate a trajectory that minimizes Wasserstein distance, which represents distance between probability distribution of a trajectory of an expert and probability distribution of a trajectory determined based on parameters of the reward function; and ( Xiao pg. 4, sec: Preposition 3.1, para. 1-2 teaches minimizing the Wasserstein distance of two measurements (i.e., probability distribution of a trajectory of an expert and probability distribution of a trajectory) with respect to the ground cost function (i.e.., reward function). In addition, pg. 7 Algorithm 1 teaches sample state action pair from the environment, under the broasted reasonable interpretation this involves an agent making a specific observation of its current state and execute an action, thus estimating a trajectory).
update the parameters of the reward function based on the estimated trajectory (Xiao pg. 6, sec: Policy gradient, para. 1 & pg. 7 Algorithm 1 line 5, teaches update the reward functions parameters
w
via gradient ascent based on the estimated trajectories
(
X
,
Y
)
from lines 3-4 in Algorithm 1).
While Xiao discloses update the parameters of the reward function based on the estimated trajectory, Xiao does not teach or suggest update the parameters of the reward function to maximize the Wasserstein distance based on the estimated trajectory and does not teach or suggest a learning device comprising: a memory storing instructions; and one or more processors configured to execute the instructions to...
However, Zhang analogues in the art teaches the following:
Zhang teaches “minimize the Wasserstein distance” (see pg. 4, left col., para. 4) and further teaches update the parameters of the [discriminator] to maximize the Wasserstein distance based on the estimated trajectory ( Zhang Algorithm 1 lines 7-13 teaches update the discriminator parameters “by maximizing the Wasserstein distance” based on the policy estimated trajectories).
Zhang is also in the same field of endeavor as Xiao (machine learning - Wasserstein Distance guided Adversarial Imitation Learning). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include the functionality of Maximizing and Minimizing the Wasserstein distance, as being disclosed and taught by Zhang, in the system taught by Xiao to yield the predictable results of promoting the performance of imitation learning (Zhang Abstract).
Neither Xiao or Zhang disclose a learning device comprising: a memory storing instructions; and one or more processors configured to execute the instructions to...
Nonetheless Eto teaches the following:
A learning device comprising: ( Eto Fig. 1 element 100 and [0076] teaches a learning device).
a memory storing instructions; and one or more processors configured to execute the instructions to (Eto [0076]).
Eto is also in the same field of endeavor as Xiao and Zhang (machine learning - inverse reinforcement learning.). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include the functionality of learning device and hardware, as being disclosed and taught by Eto, in the system taught by Xiao and Zhang to yield the predictable results of performing inverse reinforcement learning (see Eto [0001]).
Regarding claim 2:
Xiao, Zhang and Eto teach The learning device according to claim 1. Xiao specifically teaches wherein the processor is configured to execute the instructions to update the parameters of the reward function using a non-expansive mapping gradient method, which is an update rule based on a non-expansive mapping ( Xiao Algorithm 1 line 5, teaches update the reward functions parameters w via gradient ascent and pg. 4, sec: Proposition 3.1. para 6, teaches “The constraint on the gradient of reward function implies that the gradient norm at any point x is upperbounded by
1
:
∇
r
2
1
≤
1
. This simple form suggests several ways of computing the Wasserstein distance by enforcing the Lipschitz condition, such as weight clipping [5] and gradient penalty” thus suggesting the parameters of the reward functions are updated using non-expansive mapping (i.e., Lipschitz condition), which is an update rule based on a non-expansive mapping).
Regarding claim 6: is a method claim comprising limitations similar to those of claim 1, therefore is rejected under the same rational of claim 1.
Regarding claim 7: is a method claim comprising limitations similar to those of claim 2, therefore is rejected under the same rational of claim 2.
Regarding claim 8: is a non-transitory computer readable information recording medium claim comprising limitations similar to those of claim 1, therefore is rejected under the same rational of claim 1.
Regarding claim 9: is a non-transitory computer readable information recording medium claim comprising limitations similar to those of claim 2, therefore is rejected under the same rational of claim 2.
Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Xiao, Zhang, Eto in further view of Abbeel et al. Apprenticeship Learning via Inverse Reinforcement Learning (hereinafter Abbeel).
Regarding claim 5:
Xiao, Zhang and Eto teach The learning device according to claim 1. Eto teaches wherein the processor is configured to execute the instructions to... ( Eto Fig. 6 element 1001 teaches a processor).
While Xiao teaches accept as an input a reward function. Neither Xiao, Zhang and Eto explicitly teach accept input of a reward function whose features are set to be linear functions.
Nonetheless, Abbeel analogous in the art teaches the following:
...accept input of a reward function whose features are set to be linear functions (Abbeel pg. 2, left col. sec: Preliminary, para. 2, and right col., para 4 teaches accept input of a reward function whose features are set to be linear functions “Given that the reward R is expressible as a linear combination of the features
ϕ
, the feature expectations for a given policy
π
completely determine the expected sum of discounted rewards for acting according to that policy”).
Abbeel is also in the same field of endeavor as Xiao, Zhang and Eto (machine learning in the field of inverse Reinforcement Learning). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include the functionality of accept input of a reward function whose features are set to be linear functions, as being disclosed and taught by Abbeel, in the system taught by Xiao, Zhang and Eto to yield the predictable results of “correctly recover the expert's true reward function” (see pg. 2, right col., para. 3).
EXAMINER’S NOTE
Claims 3 and 4 have been searched, but no prior art has been uncovered which anticipates nor renders the claim obvious.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to GISEL G FACCENDA whose telephone number is (703)756-1919. The examiner can normally be reached Monday - Friday 8:00 am - 4:00 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Abdullah Al Kawsar can be reached at (571) 270-3169. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/G.G.F./Examiner, Art Unit 2127
/ABDULLAH AL KAWSAR/Supervisory Patent Examiner, Art Unit 2127