DETAILED ACTION
This action is responsive to the application filed on 06/23/2026. Claims 1-17 and 23-25 are pending and have been examined. This action is Final.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Applicant’s claim for the benefit of a prior-filed application under 35 U.S.C. 119(e) or under 35 U.S.C.
120, 121, 365(c), or 386(c) is acknowledged.
Response to Arguments
Argument 1: The applicant argues on pages 1-4 that the 103 obviousness rejection should be withdrawn. The claims stood rejected as obvious over Ho (“Generative Adversarial Imitation Learning”) in view of Blonde, Bousmalis (US 10,970,589), and Ganin (“Domain-adversarial training of neural networks”), with additional references (Peng, Fu, Ding, Baram, and Tobin) applied to various dependent claims. Rather than conceding the merits, the applicant amended independent claim 1 to recite penalizing the discriminator network for classifying the expert and action-trajectory subsets with more than a threshold accuracy, and applying a penalty term to the discriminator’s objective when its classification accuracy on subsets having task-irrelevant characteristics exceeds the threshold. The central argument is that the office action itself conceded (on pages 6, 11, and 13) that the Ho/Blonde/Bousmalis combination does not teach these features, leaving only Ganin cited for them, and that Ganin does not teach them either. Specifically, the applicant alleges that Ganin’s gradient-reversal mechanism applies the penalty to the deep feature extractor, not to the domain classifier, and that Ganin’s domain classifier is actually trained to maximize its own accuracy (by minimizing the domain-classification loss), which is the opposite of the claimed penalty for being too accurate. On that basis, the applicant argues that Ganin not only fails to disclose the amended limitation but affirmatively “teaches away” from it, because it rewards the classifier for accuracy while penalizing the feature extractor. The applicant therefore requests withdrawal of the rejection of claim 1 and its dependent claims, and argues that independent claims 23 and 24 (and their dependents) are allowable for the same reasons.
Examiner Response to Argument 1: The examiner has fully considered the applicant’s arguments in view of the amendment to independent claim 1 and, while the applicant’s arguments directed to Ganin are acknowledged, they are not persuasive to overcome the rejection as newly set forth below. As an initial matter, the applicant’s arguments are directed exclusively to the teachings of Ganin, however, the rejection of claim 1 no longer relies upon Ganin for the amended limitation, and the applicant’s arguments regarding Ganin’s gradient-reversal mechanism and its “teaching away” are therefore moot with respect to the present rejection. Rather, the amended limitation “penalizing the discriminator network for classifying the expert and action trajectory subsets with more than a threshold accuracy by applying a penalty term to the objective of the discriminator network when a classification accuracy of the discriminator network on the expert and action trajectory subsets having the task-irrelevant characteristics exceeds the threshold accuracy” is taught by Peng, which is already of record. Specifically, Peng discloses a variational discriminator bottleneck in which a penalty term, β·(KL(E(z|x)||r(z)) − I_c), is added directly to the objective used to train the discriminator (Peng, Sec 4, Eq. 8), and in which the penalty weight β is adaptively increased by dual gradient descent whenever the mutual information exceeds the information constraint I_c (Peng, Sec 4, Eq. 9), thus applying a penalty term to the discriminator’s own objective when a measured quantity that Peng expressly equates to the discriminator’s classification accuracy exceeds a threshold. Peng further states that “a discriminator that achieves very high accuracy will produce relatively uninformative gradients…by selecting an information constraint I_c < 1, the discriminator is prevented from perfectly differentiating between the distributions” (Peng, Abstract and Sec 4.1), confirming that Peng’s information constraint operates as a threshold on the discriminator’s classification accuracy and that exceeding that constraint results in application of the penalty, precisely as claimed. To the extent the applicant contends that a penalty applied to a feature extractor rather than to the discriminator’s objective fails to teach the amended limitation, that contention is inapposite to the present rejection, because Peng applies the penalty term to the objective of the discriminator itself rather than to a separate feature extractor. Moreover, the applicant’s assertion that the prior art “teaches away” is not commensurate with the standard for teaching away, as a reference that merely does not disclose a limitation does not thereby criticize, discredit, or otherwise discourage the claimed approach (see MPEP 2145). The identification of the expert and action trajectory subsets having task-irrelevant characteristics continues to be taught by Bousmalis, as set forth in the rejection of claim 1 below, and a person of ordinary skill in the art would have been motivated to apply Peng’s accuracy-constraining penalty to those task-irrelevant portions in order to prevent the discriminator from distinguishing expert and agent data on the basis of task-irrelevant characteristics, thereby improving generalization and stabilizing adversarial training. Accordingly, the combination of Ho, Blonde, Bousmalis, and Peng teaches or suggests the amended limitations of claim 1, and the rejection under 35 U.S.C. 103 is maintained as set forth below.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this
Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not
identically disclosed as set forth in section 102, if the differences between the claimed invention and the
prior art are such that the claimed invention as a whole would have been obvious before the effective filing
date of the claimed invention to a person having ordinary skill in the art to which the claimed invention
pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are
summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1, 2, 4-6, 8, 13, and 14, 23-25 is/are rejected under 35 U.S.C. 103 as being unpatentable over NPL reference “Generative Adversarial Imitation Learning” by Ho et. al (referred herein as Ho) in view of NPL reference “Sample-Efficient Imitation Learning via Generative Adversarial Nets” by Blonde et. al (referred herein as Blonde) in view of US10970589B2, by Bousmalis et. al. (referred herein as Bousmalis) further in view of NPL reference “VARIATIONAL DISCRIMINATOR BOTTLENECK: IMPROVING IMITATION LEARNING, INVERSE RL, AND GANS BY CONSTRAINING INFORMATION FLOW” by Peng et. al (referred herein as Peng.)
Regarding claim 1, Ho teaches:
A method of training a neural network to generate action data for controlling an agent to perform a task in an environment, the method comprising: obtaining, for each of a plurality of expert performances of the task, one or more first tuple datasets, each first tuple dataset comprising state data characterizing a state of the environment at a corresponding time during the corresponding expert performance of the task; and ([Ho, page 2 and 7] “In practice, πE will only be provided as a set of trajectories sampled by executing πE in the environment, so the expected cost of πE in Eq. (1) is estimated using these samples [Algorithm 1]”, wherein the examiner interprets the collection of trajectories and the use of a discriminator as detailed in Algorithm 1 to be the same as “obtaining, for each of a plurality of expert performances of the task, one or more first tuple datasets, each first tuple dataset comprising state data characterizing a state of the environment at a corresponding time during the corresponding expert performance of the task”).
Ho does not teach and performing a concurrent process of training the neural network and a discriminator network, the process comprising: (i) a plurality of neural network update steps, each of which comprises: receiving state data characterizing a current state of the environment; using the neural network and the state data to generate action data indicative of an action to be performed by the agent; forming a second tuple dataset comprising the state data; using the second tuple dataset to generate a reward value, wherein the reward value comprises an imitation value generated by the discriminator network based on the second tuple dataset; and training the neural network based on the reward value; (ii) a plurality of discriminator network update steps, each of which comprises: obtaining one or more expert trajectories; obtaining one or more action trajectories; obtaining a plurality of expert trajectory subsets, wherein each expert trajectory subset is a subset of a respective plurality of first tuple datasets corresponding to a respective expert trajectory and includes a portion of the respective plurality of first tuple datasets that has been classified as having task-irrelevant characteristics; obtaining a plurality of action trajectory subsets, wherein each action trajectory subset is a subset of a respective plurality of second tuple datasets corresponding to a respective action trajectory and includes a portion of the respective plurality of second tuple datasets that has been classified as having task-irrelevant characteristics; and training the discriminator network, comprising classifying, by the discriminator network, the pluralities of first and second tuple datasets corresponding to the expert and action trajectory subsets that include, within each action or expert trajectory subset, the portion of the respective plurality of first or second tuple datasets that have been classified as having task-irrelevant characteristics; classifying, by the discriminator network, the pluralities of first and second tuple datasets, and training the discriminator network on an objective that (i) encourages the discriminator network to accurately classify the pluralities of first and second tuple datasets while (ii) penalizing the discriminator network for classifying the expert and action trajectory subsets with more than a threshold accuracy by applying a penalty term to the objective of the discriminator network when a classification accuracy of the discriminator network on the expert and action trajectory subsets having the task-irrelevant characteristics exceeds the threshold accuracy.
Blonde teaches:
performing a concurrent process of training the neural network and a discriminator network, the process comprising: (i) a plurality of neural network update steps, each of which comprises: receiving state data characterizing a current state of the environment; using the neural network and the state data to generate action data indicative of an action to be performed by the agent; forming a second tuple dataset comprising the state data; using the second tuple dataset to generate a reward value, wherein the reward value comprises an imitation value generated by the discriminator network based on the second tuple dataset; and training the neural network based on the reward value; ([Blonde, sec 4] “As an off-policy method, SAM cycles through the following steps: i) the agent uses πθ to interact with M, ii) the agent stores the experienced transitions C in a replay buffer R, iii) updates the reward module φ with an mixture of uniformly sampled state-action pairs from C and τe, iv) updates the reward module φ with an equal mixture of uniformly sampled state-action pairs from R and τe, and v) updates the policy module θ and critic module ψ with transitions sampled from R.”, wherein the examiner interprets the policy module (which corresponds to the neural network that generates actions) being updated from replay buffer (R) samples, thereby generating state-action pairs, to be the same as “receiving state data characterizing a current state of the environment; using the neural network and the state data to generate action data,” because both describe how the neural network (policy module) is trained and captures state data (a.k.a. “second tuple”) and action data, which tuple is fed to a discriminator to obtain an imitation-based reward used to update the neural network.)
(ii) a plurality of discriminator network update steps, each of which comprises: obtaining one or more expert trajectories; obtaining one or more action trajectories; ([Blonde, page 5, Sec 4] “We introduce a reward network…The reward network is trained, each iteration, first on the mini-batch most recently collected by π(theta), then on mini-batches sampled from the replay buffer”, wherein the examiner interprets “mini-batches most recently collected by π” to be the same as “obtaining one or more action trajectories,” and “mini-batches sampled from the replay buffer corresponding to expert trajectories” to be the same as “obtaining one or more expert trajectories,” because both provide state-action samples for discriminator updates.)
obtaining a plurality of expert trajectory subsets, wherein each expert trajectory subset is a subset of a respective plurality of first tuple datasets corresponding to a respective expert trajectory ([Blonde, page 3, sec 3] “Trajectories are traces of interaction between an agent and an MDP. Specifically, we model trajectories as sequences of transitions (st, at, rt, st+1), atomic units of interaction. Demonstrations are provided to the agent through a set of expert trajectories τe, generated by an expert policy πe in M.”, wherein the examiner interprets “provided to the agent through a set of expert trajectories” to be the same as “obtaining a plurality of expert trajectory subsets,” and “trajectories are traces of interaction between an agent and an MDP” (sequences of transitions) to be the same as “a respective plurality of first tuple datasets corresponding to a respective expert trajectory.”)
obtaining a plurality of action trajectory subsets, wherein each action trajectory subset is a subset of a respective plurality of second tuple datasets corresponding to a respective action trajectory and ([Blonde, page 7, Algorithm 1] “Sample uniformly a minibatch Bc of state-action pairs from C. Sample uniformly a minibatch Bc_e of state-action pairs from the expert dataset τe.”, wherein the examiner interprets “Sample uniformly a minibatch Bc of state-action pairs from C” to be the same as obtaining a plurality of action trajectory subsets, and “sequences of transitions (st, at, rt, st+1)” to be the same as a respective plurality of second tuple datasets corresponding to a respective action trajectory.)
training the discriminator network comprising classifying, by the discriminator network, the pluralities of first and second tuple datasets corresponding to the expert and action trajectory subsets; classifying, by the discriminator network, the pluralities of first and second tuple datasets, and training the discriminator network on an objective that (i) encourages the discriminator network to accurately classify the pluralities of first and second tuple datasets ([Blonde, page 3, sec 3] “Generative Adversarial Imitation Learning introduces an extra neural network Dφ to play the role of discriminator… Dφ tries to assert whether a given state-action pair originates from trajectories of πθ or πe” and [Blonde, page 5, sec 5] “The cross-entropy loss used to train the reward network is: Eπθ [- log(1 - Dφ(s, a))] + Eπe [- log Dφ(s, a)]”, wherein the examiner interprets the discriminator asserting whether a state-action pair originates from πθ or πe, trained by the cross-entropy loss, to be the same as classifying the pluralities of first and second tuple datasets and training the discriminator on an objective that encourages accurate classification.)
Ho and Blonde does not teach and includes a portion of the respective plurality of first tuple datasets that has been classified as having task-irrelevant characteristics; … includes a portion of the respective plurality of second tuple datasets that has been classified as having task-irrelevant characteristics, and … while (ii) penalizing the discriminator network for classifying the expert and action trajectory subsets with more than a threshold accuracy by applying a penalty term to the objective of the discriminator network when a classification accuracy of the discriminator network on the expert and action trajectory subsets having the task-irrelevant characteristics exceeds the threshold accuracy.
Bousmalis teaches:
and includes a portion of the respective plurality of first tuple datasets that has been classified as having task-irrelevant characteristics; … includes a portion of the respective plurality of second tuple datasets that has been classified as having task-irrelevant characteristics; and that include, within each action or expert trajectory subset, the portion of the respective plurality of first or second tuple datasets that have been classified as having task-irrelevant characteristics; ([Bousmalis, col 4, lines 27-31], “The private target encoder neural network 210 is specific to the target domain and is configured to receive images from the target domain and to generate, for each received image, a private feature representation of the image.” and [Bousmalis, col 7-8, lines 65-68, 1-7], “The difference loss trains the shared encoder neural network to (i) generate shared feature representations for input images from the target domain that are different from private feature representations for the same input images from the target domain generated by the private target encoder neural network and (ii) generate shared feature representations for input images from the source domain that are different from private feature representations for the same input images from the source domain generated by the private source encoder neural network.”, wherein the examiner interprets the private feature representations generated by the private target encoder neural network to correspond to task-irrelevant characteristics of the input images because they are domain-specific representations separated from the shared feature representations used for the task, such that each first or second tuple dataset includes a portion (the private feature representation) classified as having task-irrelevant characteristics, separated from the task-relevant (shared) portion by the difference loss.)
Ho, Blonde, and Bousmalis does not teach while (ii) penalizing the discriminator network for classifying the expert and action trajectory subsets with more than a threshold accuracy by applying a penalty term to the objective of the discriminator network when a classification accuracy of the discriminator network on the expert and action trajectory subsets having the task-irrelevant characteristics exceeds the threshold accuracy.
Peng teaches while (ii) penalizing the discriminator network for classifying the expert and action trajectory subsets with more than a threshold accuracy by applying a penalty term to the objective of the discriminator network when a classification accuracy of the discriminator network on the expert and action trajectory subsets having the task-irrelevant characteristics exceeds the threshold accuracy. ([Peng, page 4, sec. 4, Eq. 8] “J(D, E) = min over (D,E), max over (β ≥ 0), of E_(x∼p) of E_(z∼E(z|x)) of [− log D(z)] + E_(x∼G) of E_(z∼E(z|x)) of [− log(1 − D(z))] + β·(E_(x∼p̃) of KL(E(z|x) || r(z)) − I_c)” AND [Peng, page 4, sec. 4, Eq. 9] ”we adaptively update β via dual gradient descent to enforce a specific constraint I_c on the mutual information, β ← max(0, β + α_β·(E_(x∼p̃) of KL(E(z|x) || r(z)) − I_c))” AND [Peng, page 1, Abstract-sec. 1] ” However, they suffer from major optimization challenges, one of which is balancing the performance of the generator and discriminator. A discriminator that achieves very high accuracy can produce relatively uninformative gradients,… we can effectively modulate the discriminator’s accuracy” AND [Peng, page 5, sec. 4.1] “Since the minimum amount of information required for binary classification is 1 bit, by selecting an information constraint I_c < 1, the discriminator is prevented from perfectly differentiating between the distributions”, wherein the examiner interprets Peng’s variational discriminator bottleneck term”β·(KL(E(z|x) || r(z)) − I_c)“, which is added to the objective J(D, E) used to train the discriminator D, to be the same as “applying a penalty term to the objective of the discriminator network,” because both add a penalty/constraint term to the objective function used to train the discriminator. The examiner further interprets Peng’s adaptive update of the Lagrange multiplier β by dual gradient descent, “β ← max(0, β + α_β·(KL − I_c)),” which increases the penalty weight β whenever the mutual information exceeds the information constraint I_c, to be the same as “applying a penalty term … when a classification accuracy of the discriminator network … exceeds the threshold accuracy,” because Peng expressly ties the information constraint I_c to the discriminator’s classification accuracy (Peng stating that binary classification requires at least 1 bit of information and that selecting I_c < 1 prevents the discriminator from perfectly differentiating, i.e., from being too accurate), such that the mutual information exceeding I_c corresponds to the discriminator’s classification accuracy exceeding a threshold accuracy, and the resulting increase in β penalizes the discriminator accordingly. The examiner further interprets this penalty to operate on the expert and action trajectory subsets having the task-irrelevant characteristics identified by Bousmalis above, because Peng’s information bottleneck constrains the discriminator to ignore irrelevant information ([ex. Peng, page 3, sec 2-3, “a compressed representation can improve generalization by ignoring irrelevant distractors present in the original input.”), consistent with penalizing accurate classification based on the task-irrelevant (domain-specific/private) portions separated out by Bousmalis.
Ho, Blonde, Bousmalis, Peng, and the instant application are analogous art because they are all directed to training neural networks using adversarial or discriminator-based learning frameworks to improve performance.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the Generative adversarial imitation learning approach disclosed by Ho to include the reward network disclosed by Blonde. One would be motivated to do so to effectively train a discriminator network to distinguish between expert and agent-generated data and provide meaningful reward signals for policy learning, as suggested by Blonde ([Blonde, page 5, sec 5] “The cross-entropy loss used to train the reward network is: Eπθ [- log(1 - Dφ(s, a))] + Eπe [- log Dφ(s, a)].”).
It would have also been obvious to a person of ordinary skill in the art before the effective filing date of the invention to include the private and shared feature representation approach disclosed by Bousmalis. One would be motivated to do so to efficiently separate domain-specific information from task-relevant information within the input data, thereby improving generalization and robustness of the trained model across different domains, as suggested by Bousmalis ([Bousmalis, col 4, lines 27-31] “to generate, for each received image, a private feature representation of the image.”).
It would have also been obvious to a person of ordinary skill in the art before the effective filing date of the invention to include the variational discriminator bottleneck (VDB) penalty term disclosed by Peng. One would be motivated to do so to effectively balance the performance of the discriminator and prevent the discriminator from achieving a very high classification accuracy that produces relatively uninformative gradients, thereby stabilizing adversarial training and improving imitation learning performance, as suggested by Peng ([Peng, Abstract-sec. 1] “a discriminator that achieves very high accuracy will produce relatively uninformative gradients … By enforcing a constraint on the mutual information between the observations and the discriminator’s internal representation, we can effectively modulate the discriminator’s accuracy and maintain useful and informative gradients.”).Claims 23 and 24 are analogous to claim 1, and therefore will face the same rejection set forth above.
Regarding claim 2, Ho, Blonde, Bousmalis, and Peng teaches A method according to claim 1 (see rejection of claim 1).
Ho further teaches in which the first tuple datasets further include corresponding action data generated based on the state of the environment for controlling the agent. ([Ho, page 2] “[the unnormalized] distribution of state-action pairs that an agent encounters when navigating the environment with the policy π, and it allows us to write Eπ[c(s, a)] = <f(s, a)> for any cost function c”, wherein the examiner interprets the inclusion of corresponding action data in the first tuples, which capture state-action pairs encountered by an agent, to be the same as “corresponding action data generated based on the state of the environment for controlling the agent” because both record the state of the environment at a given time during the expert’s performance, and each first tuple captures the action data generated based on that state for controlling the agent.).
Regarding claim 4, Ho, Blonde, Bousmalis, and Peng teaches A method according to claim 1 (see rejection of claim 1).
Ho further teaches in which the training constrains the discriminator network such that, upon receiving any of at least a specified proportion of tuple datasets included in the expert and action trajectory subsets, the discriminator network generates (i) an imitation value below an imitation value threshold if the received tuple dataset is a first tuple dataset, and (ii) an imitation value above the imitation value threshold if the received tuple dataset is a second tuple dataset; ([Ho, page 6, Sec 5] “[GAIL solves Eq. (15)]… Explicitly, we wish to find a saddle point (π,D) of the expression Eπ[log(D(s, a))] + EπE [log(1 − D(s, a))] … with both π and D represented using function approximators: GAIL fits a parameterized policy πθ, with weights θ, and a discriminator network Dw : S × A → (0, 1), with weights w”, wherein the examiner interprets “the discriminator network Dw : S × A → (0,1)” to be the same as the discriminator generating an imitation value below a threshold for action trajectory subsets and above a threshold for expert trajectory subsets, because Dw outputs values near 0 for agent data and near 1 for expert data, with 0.5 serving as the natural threshold.). Claim 25 is analogous to claim 4, and therefore will face the same rejection set forth above.
Regarding claim 5, Ho, Blonde, Bousmalis, and Peng teaches A method according to claim 4 (see rejection of claim 4).
Peng further teaches the objective includes a term which varies inversely dependent with an accuracy parameter, the accuracy parameter (i) taking a higher value if, upon receiving one of the expert trajectory subsets of first tuple datasets, the discriminator network generates with a probability above a probability threshold an imitation value above the imitation value threshold, and (ii) taking a higher value if, upon receiving one of the action trajectory subsets of second tuple datasets, the discriminator network generates with a probability above the probability threshold an imitation value below the imitation value threshold. ([Peng, page 9, Sec 5.1] “Adaptive Constraint:…When β is too small, performance reverts to that achieved by GAIL… Policies trained using dual gradient descent to adaptively update β consistently achieves the best performance overall”, wherein the examiner interprets “adaptive update of β to enforce a desired information constraint” to be the same as “an accuracy parameter that varies inversely” because both are directed to adjusting a parameter such that discriminator confidence (accuracy) increases when correctly classifying expert versus agent data, and the constraint smooths the discriminator landscape when accuracy is too high, thereby functioning as an inverse dependence.)
Ho, Blonde, Bousmalis, Peng, and the instant application are analogous art because they are all directed to modifying parameters of a discriminator network.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method of claim 4 disclosed by Ho, Blonde, Bousmalis, and Peng to include the “adaptive constraint” disclosed by Peng. One would be motivated to do so to effectively enhance the adaptive constraint on the discriminator network, as suggested by Peng ([Peng, page 9, sec 5.1] “Policies trained using dual gradient descent to adaptively update β consistently achieves the best performance overall.”).
Regarding claim 6, Ho, Blonde, Bousmalis, and Peng teaches A method according to claim 1 (see rejection of claim 1).
Ho further teaches (a) in which, for each performance of the task, the corresponding first tuple datasets form an expert sequence of first tuple datasets labelled by a time index which is zero for the first tuple dataset of the expert sequence, and one higher for each successive first tuple dataset than for the preceding one of the expert sequence, and ([Ho, page 2, sec 2] “We work in the γ-discounted infinite horizon setting, and we will use an expectation with respect a policy π ∈ Π to denote an expectation with respect to the trajectory it generates … where s0 ∼ p0, at ∼ π(·|st), and st+1 ∼ P(·|st, at) for t ≥ 0. We will use ˆ Eτ to denote empirical expectation with respect to trajectory samples τ, and we will always refer to the expert policy”, wherein the examiner interprets the value t in the summation, which indexes the state-action pair (st, at), to be a “time index which is zero for the first tuple dataset of the expert sequence, and one higher for each successive first tuple dataset than for the preceding one.”)
Blonde further teaches (b) the neural network update steps are based on one or more action sequences of second tuple datasets, wherein for each action sequence of second tuple datasets: a first second tuple of the action sequence has a time index of zero, and is performed for state data describing the environment in a corresponding initial state, and each of other second tuple datasets of the action sequence has a time index one greater than the preceding second tuple of the action sequence, and is performed for state data describing the environment upon the performance by the agent of the action data generated in the preceding time step. ([Blonde, page 3, sec 3] “Trajectories are traces of interaction between an agent and an MDP. Specifically, we model trajectories as sequences of transitions (st, at, rt, st+1), atomic units of interaction. Demonstrations are provided to the agent through a set of expert trajectories τe, generated by an expert policy πe in M.”, wherein the examiner interprets the ordering of transitions, where the first transition corresponds to an initial state with a time index of zero and each subsequent transition occurs at the next time step, to be the same as the “action sequences of second tuple datasets … has a time index of zero … the action sequence has a time index one greater than the preceding second tuple of the action sequence” as recited.)
Ho, Blonde, Bousmalis, Peng, and the instant application are analogous art because they are all directed to expert sequences of datasets labelled by a sequential time index.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method of claim 1 disclosed by Ho, Blonde, Bousmalis, to include the tuple dataset processing disclosed by Ho and Blonde. One would be motivated to do so to effectively ensure accurate temporal ordering of the expert sequence and to provide action sequences for the tuples, as suggested by Ho and Blonde ([Ho, page 2, sec 2] “We work in the γ-discounted infinite horizon setting, and we will use an expectation with respect a policy π ∈ Π to denote an expectation with respect to the trajectory it generates … where s0 ∼ p0, at ∼ π(·|st), and st+1 ∼ P(·|st, at) for t ≥ 0. We will use ˆ Eτ to denote empirical expectation with respect to trajectory samples τ, and we will always refer to the expert policy” [Blonde, page 3, sec 3] “we model trajectories as sequences of transitions (st, at, rt, st+1), atomic units of interaction.”).
Regarding claim 8, Ho, Blonde, Bousmalis, and Peng teaches A method according to claim 6 (see rejection of claim 6).
Ho further teaches in which all the second tuple datasets employed in each discriminator network update step are tuple datasets for which the corresponding time index is below a third time threshold, the expert sequences employed in the discriminator network update including first tuple datasets having a time index above the third time threshold. (Ho, page 2, sec 2] “We work in the γ-discounted infinite horizon setting, and we will use an expectation with respect a policy π ∈ Π to denote an expectation with respect to the trajectory it generates:” AND [Ho, page 7, sec 6] “a given dataset of state-action pairs is split into 70% training data and 30% validation data.”, wherein the examiner interprets the state-action pairs (s,a) being updated at each time step within the GAIL framework to be the same as “network update step are tuple datasets for which the corresponding time index,” and “dataset of state-action pairs” to be the same as “second tuple datasets” used as input to the discriminator network.)
Regarding claim 13, Ho, Blonde, Bousmalis, and Peng teaches A method according to claim 1 (see rejection of claim 1).
Peng further teaches in which the state data for each tuple dataset comprises image data defining at least one image of the environment. ([Peng, page 7, sec 5], “We evaluate our method on adversarial learning problems in imitation learning, inverse reinforcement learning, and image generation. In the case of imitation learning, we show that the VDB enables agents to learn complex motion skills from a single demonstration, including visual demonstrations provided in the form of video clips. We also show that the VDB improves the performance of inverse RL methods. Inverse RL aims to reconstruct a reward function from a set demonstrations, which can then used to perform the task in new environments, in contrast to imitation learning, which aims to recover a policy directly. Our method is also not limited to control tasks, and we demonstrate its effectiveness for unconditional image generation.”, wherein the examiner interprets “visual demonstrations provided in the form of video clips” to be the same as “image data defining at least one image of the environment,” as both terms are directed to visual representations of the environment since a video clip consist of multiple images of the environment.).
Ho, Blonde, Bousmalis, Peng, and the instant application are analogous art because they are all directed to methods for generating and using image data representing environments in learning or decision-making tasks.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method of claim 1 disclosed by Ho, Blonde, Bousmalis, and Peng to include the imagery-based adversarial learning disclosed by Peng. One would be motivated to do so to effectively enhance adversarial learning using the Variational Discriminator Bottleneck (VDB) method, as suggested by Peng ([Peng, page 7, sec 5] “We also show that the VDB improves the performance of inverse RL methods.”).
Regarding claim 14, Ho, Blonde, Bousmalis, and Peng teaches A method according to claim 13 (see rejection of claim 13).
Peng further teaches in which the state data for each tuple dataset comprises image data defining a plurality of images of the environment. ([Peng, page 7, sec 5] “we show that the VDB enables agents to learn complex motion skills from a single demonstration, including visual demonstrations provided in the form of video clips.”, wherein the examiner interprets visual demonstrations provided as video clips to be the same as “image data defining a plurality of images of the environment,” as a video is a sequence or collection of multiple images.)
Ho, Blonde, Bousmalis, Peng, and the instant application are analogous art because they are all directed to methods for capturing environmental state data comprising multiple images.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method of claim 13 disclosed by Ho, Blonde, Bousmalis, and Peng to include the video-based imagery disclosed by Peng. One would be motivated to do so to effectively enhance adversarial learning using the VDB method with video (i.e., multiple images), as suggested by Peng ([Peng, page 7, sec 5] “visual demonstrations provided in the form of video clips.”).
Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Ho in view of Blonde in view of Bousmalis in view of Peng further in view of NPL reference “Goal-conditioned Imitation Learning” by Ding et. al (referred herein as Ding).
Regarding claim 7, Ho, Blonde, Bousmalis, and Peng teaches A method according to claim 6 (see rejection of claim 6).
Ho, Blonde, Bousmalis, and Peng do not teach in which, in each of the discriminator network update steps, the expert trajectory subset of first tuple datasets are first tuple datasets for which the corresponding time index is below a first time threshold, and the action trajectory subset of second tuple datasets are tuple datasets for which the corresponding time index is below a second time threshold.
Ding teaches in which, in each of the discriminator network update steps, the expert trajectory subset of first tuple datasets are first tuple datasets for which the corresponding time index is below a first time threshold, and the action trajectory subset of second tuple datasets are tuple datasets for which the corresponding time index is below a second time threshold. Ding teaches in which, in each of the discriminator network update steps, the first subset of first tuple datasets are first tuple datasets for which the corresponding time index is below a first time threshold, and the second subset of second tuple datasets are tuple datasets for which the corresponding time index is below a second time threshold. ([Ding, page 4] “The expert trajectories have been collected by asking the expert to reach a specific goal gj. But they are also valid trajectories to reach any other state visited within the demonstration! This is the key motivating insight to propose a new type of relabeling: if we have the transitions
PNG
media_image1.png
41
174
media_image1.png
Greyscale
in a demonstration, we can also consider the transition
PNG
media_image2.png
34
70
media_image2.png
Greyscale
PNG
media_image3.png
48
184
media_image3.png
Greyscale
as also coming from the expert! Indeed that demonstration also went through the state...so if that was the goal, the expert would also have generated this transition. This can be understood as a type of data augmentation leveraging the assumption that the tasks we work on are quasi-static.”, wherein the examiner interprets “considering transitions within a demonstration as separate valid expert transitions” to be the same as “using subsets of tuple datasets whose time indices fall below designated thresholds” because both are directed to relabeling or selecting portions of a trajectory based on temporal progression through visited states).
Ho, Blonde, Bousmalis, Peng, Ding, and the instant application are analogous art because they are all directed to methods for updating discriminator networks based on subsets of datasets selected according to defined temporal index thresholds.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method according to claim 6 disclosed by Ho, Blonde, Bousmalis, and Peng to include the “new type of relabeling” disclosed by Ding. One would be motivated to do so to efficiently augment training data, as suggested by Ding ([Ding, page 4] “data augmentation leveraging the assumption that the tasks we work on are quasi-static.”).
Claims 9-12, 15, and 17 is rejected under 35 U.S.C. 103 as being unpatentable over Ho in view of Blonde in view of Bousmalis in view of Peng further in view of NPL reference “Model-based Adversarial Imitation Learning” by Baram et. al (referred herein as Baram).
Regarding claim 9, Ho, Blonde, Bousmalis, and Peng teaches A method according to claim 6 (see rejection of claim 6).
Ho, Blonde, Bousmalis, and Peng do not teach in which, for each action sequence, a corresponding third time threshold is determined, and the second tuple datasets of the action sequence employed in each discriminator network update step only include second tuple datasets for which the corresponding time index is below a corresponding third time threshold.
Baram teaches in which, for each action sequence, a corresponding third time threshold is determined, and the second tuple datasets of the action sequence employed in each discriminator network update step only include second tuple datasets for which the corresponding time index is below a corresponding third time threshold. ([Baram, page 5, sec 3.3], “We showed that effective imitation learning requires a) to use a model, and b) to process multistep transitions instead of individual state-action pairs. This setup was previously suggested by Shalev-Shwartz et al. [2016] and Heess et al. [2015], who tried to maximize R(π) by expressing it as a multi-step differentiable graph. Our method can be viewed as a variant of their idea when setting: r(s, a) = −D(s, a). This way, instead of maximizing the total reward, we minimize the total discriminator beliefs along a trajectory … Define J(θ) as the discounted sum of discriminator probabilities along a trajectory… Jθ is calculated by applying Eq. 10 and 11 recursively, starting from t = T all the way down to t = 0.”, wherein the examiner interprets recursively calculating discriminator probabilities starting from time T, where T is a threshold, and moving downward, to be the same as “for each action sequence, a corresponding third time threshold is determined,” and using discriminator values only for time indices below that threshold in the discriminator update step).
Ho, Blonde, Bousmalis, Peng, Baram, and the instant application are analogous art because they are all directed to methods for enhancing imitation learning by selectively utilizing discriminator updates based on threshold-based filtering of action sequences.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method according to claim 6 disclosed by Ho, Blonde, Bousmalis, and Peng to include the “process multistep transitions instead of individual state-action pairs” disclosed by Baram. One would be motivated to do so to effectively reduce discriminator uncertainty along action trajectories, as suggested by Baram ([Baram, page 5, sec 3.3] “instead of maximizing the total reward, we minimize the total discriminator beliefs along a trajectory”).
Regarding claim 10, Ho, Blonde, Bousmalis, Peng, and Baram teaches A method according to claim 9 (see rejection of claim 9).
Ho further teaches comprising a step of, for each action sequence, selecting the third time threshold for the action sequence based on imitation values for at least a plurality of the second tuple datasets of the action sequence. ([Ho, page 6, sec 5] “The discriminator network can be interpreted as a local cost function providing learning signal to the policy—specifically, taking a policy step that decreases expected cost with respect to the cost function c(s,a) = logD(s,a) will move toward expert-like regions of state-action space, as classified by the discriminator.” AND ([Ho, page 6, Sec 5] “[GAIL solves Eq. (15)]… Explicitly, we wish to find a saddle point (π,D) of the expression Eπ[log(D(s, a))] + EπE [log(1 − D(s, a))] … with both π and D represented using function approximators: GAIL fits a parameterized policy πθ, with weights θ, and a discriminator network Dw : S × A → (0, 1), with weights w” wherein the examiner interprets “discriminator value” and “saddle point” to be the same as the “imitation value” and “time threshold” as both relate to finding the similarity between tuples of data and a threshold for the action sequence, where D(s,a) represents the discriminator output for a state-action pair (s, a).)
Regarding claim 11, Ho, Blonde, Bousmalis, Peng, and Baram teaches A method according to claim 10 (see rejection of claim 10).
Peng further teaches in which the third time threshold is set as the smallest time index such that a certain number Tpatience of the most recent imitation values is above an imitation quality threshold. ([Peng, page 18, sec C] “When evaluating the performance of the policies, each episode is simulated for a maximum horizon of 20s. Early termination is triggered whenever the character’s torso contacts the ground, leaving the policy a maximum error of π radians for all remaining timesteps.”, wherein the examiner interprets “the smallest time index such that a certain number of recent imitation values is above an imitation quality threshold” to be the same as the stopping condition of Tpatience to ensure imitation values “is above an imitation quality threshold,” as both are checking performance at each time step and terminating once a threshold is met.)
Ho, Blonde, Bousmalis, Peng, Baram, and the instant application are analogous art because they are all directed to evaluating and improving the imitation quality of a learned policy through trajectory analysis over time.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method of claim 10 disclosed by Ho, Blonde, Bousmalis, Peng, and Baram to include the “early termination” trigger disclosed by Peng. One would be motivated to do so to efficiently establish a clear performance-based stopping point during trajectory evaluation, as suggested by Peng ([Peng, page 18, sec C] “leaving the policy a maximum error of π radians for all remaining timesteps.”).
Regarding claim 12, Ho, Blonde, Bousmalis, Peng, and Baram teaches A method according to claim 11 (see rejection of claim 11).
Blonde further teaches in which the imitation quality threshold is based on the imitation values of a plurality of second tuples of that action sequence having a time index below the third time threshold. ([Blonde, page 3, sec 3] “We now introduce additional concepts and notations that will be used in the remainder of this work. The return is the total discounted reward from timestep t, onwards: ... The state-action value, or Q-value, is the expected return after picking action at in state st, and thereafter following policy πθ: Qπθ (st, at) , ... πθ[·] denotes the expectation taken along trajectories generated by πθ in M+ (respectively E>tπe[·] for πe in M) and looking onwards from state st and action at. We want our agent to find a policy πθ that maximizes the expected return from the start state, which constitutes our performance objective...”, wherein the examiner interprets the expectation of returns taken along trajectories generated from state-action pairs occurring after timestep t, where timestep t is chosen that maximizes the expected return and that time step becomes the threshold, to be the same as the “imitation quality threshold” being based on the imitation values of multiple second tuples having “a time index below the third time threshold”).
Ho, Blonde, Bousmalis, Peng, Baram, and the instant application are analogous art because they are all directed to methods for determining imitation thresholds based on evaluating state-action pairs in action sequences along trajectories.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method of claim 11 disclosed by Ho, Blonde, Bousmalis, Peng, and Baram to include the calculation of the “state-action value, or Q-value” disclosed by Blonde. One would be motivated to do so to effectively maximize the expected performance of imitation policies, as suggested by Blonde ([Blonde, page 3, sec 3] “We want our agent to find a policy πθ that maximizes the expected return from the start state, which constitutes our performance objective.”).
Regarding claim 15, Ho, Blonde, Bousmalis, and Peng teaches A method according to claim 13 (see rejection of claim 13).
Ho, Blonde, Bousmalis, and Peng do not teach in which, during at least one of (i) one or more of the neural network update steps, a modified form of the second tuple datasets is generated by making a modification to the state data of the second tuple datasets, and (ii) one or more of the discriminator network update steps, a modified form of the first and/or second tuple datasets is generated by making a modification to the state data of one or more of the first and/or second tuple datasets.
Baram teaches in which, during at least one of (i) one or more of the neural network update steps, a modified form of the second tuple datasets is generated by making a modification to the state data of the second tuple datasets, and (ii) one or more of the discriminator network update steps, a modified form of the first and/or second tuple datasets is generated by making a modification to the state data of one or more of the first and/or second tuple datasets. ([Baram, page 6, sec 3] “we write the derivatives of J over a (s, a, s’) transition in a recursive manner [Eq. 10 and 11] The final gradient Jθ is calculated by applying Eq. 10 and 11 recursively, starting from t = T all the way down to t = 0.”, wherein the examiner interprets recursively adjusting gradients and state information (using terms derived from the forward model) during policy and discriminator updates to be the same as generating a modified form of tuple datasets by making “a modification to the state data” during neural network and discriminator network update steps.)
Ho, Blonde, Bousmalis, Peng, Baram, and the instant application are analogous art because they are all directed to methods for updating neural networks and discriminator networks by modifying datasets based on state data during learning.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method of claim 13 disclosed by Ho, Blonde, Bousmalis, and Peng to include the discriminator update process disclosed by Baram. One would be motivated to do so to efficiently calculate the discounted sum of discriminator probabilities, as suggested by Baram ([Baram, page 6, sec 3] “The final gradient Jθ is calculated by applying Eq. 10 and 11 recursively, starting from t = T all the way down to t = 0.”).
Regarding claim 17, Ho, Blonde, Bousmalis, Peng, and Baram teaches A method according to claim 15 (see rejection of claim 15).
Peng further teaches in which the state data for each tuple dataset comprises image data defining a plurality of images of the environment and in which the modification comprises removing the image data for one or more of the images of the state data. ([Peng, page 2, sec 1] “In this work, we propose a simple regularization technique for adversarial learning, which constrains the information flow from the inputs to the discriminator using a variational approximation to the information bottleneck. By enforcing a constraint on the mutual information between the input observations and the discriminator’s internal representation, we can encourage the discriminator to learn a representation that has heavy overlap between the data and the generator’s distribution, thereby effectively modulating the discriminator’s accuracy and maintaining useful and informative gradients for the generator. Our approach to stabilizing adversarial learning can be viewed as an adaptive variant of instance noise (Salimans et al., 2016; Sønderby et al., 2016; Arjovsky & Bottou, 2017). However, we show that the adaptive nature of this method is critical. Constraining the mutual information between the discriminator’s internal representation and the input allows the regularizer to directly limit the discriminator’s accuracy, which automates the choice of noise magnitude and applies this noise to a compressed representation of the input that is specifically optimized to model the most discerning differences between the generator and data distributions. The main contribution of this work is the variational discriminator bottleneck (VDB), an adaptive stochastic regularization method for adversarial learning that substantially improves performance across a range of different application domains, examples of which are available in Figure 1. Our method can be easily applied to a variety of tasks and architectures. First, we evaluate our method on a suite of challenging imitation tasks, including learning highly acrobatic skills from mocap data with a simulated humanoid character. Our method also enables characters to learn dynamic continuous control skills directly from raw video demonstrations, and drastically improves upon previous work that uses adversarial imitation learning.” wherein the examiner interprets applying noise to a compressed representation of input observations to limit discriminator accuracy, thereby reducing available input information, to be the same as “the modification comprises removing the image data for one or more of the images of the state data,” as both are directed to selectively reducing input information provided to the discriminator to ensure it will focus on the specific tasks desired.)
Ho, Blonde, Bousmalis, Peng, Baram, and the instant application are analogous art because they are all directed to selectively modifying state data to improve adversarial training in imitation learning tasks.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method of claim 15 disclosed by Ho, Blonde, Bousmalis, Peng, and Baram to include the “variational discriminator bottleneck (VDB)” disclosed by Peng. One would be motivated to do so to effectively improve performance across a range of different application domains, as suggested by Peng ([Peng, page 2, sec 1] “substantially improves performance across a range of different application domains”).
Claim 16 is rejected under 35 U.S.C. 103 as being unpatentable over Ho in view of Blonde in view of Bousmalis in view of Peng in view of Baram further in view of NPL reference “Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World” by Tobin et. al (referred herein as Tobin).
Regarding claim 16, Ho, Blonde, Bousmalis, Peng, and Baram teaches A method according to claim 15 (see rejection of claim 15).
Ho, Blonde, Bousmalis, Peng, and Baram do not teach in which the modification comprises applying to the image data one or more modifications selected from a set comprising: brightness changes; contrast changes; saturation changes; cropping; rotation; and addition of noise.
Tobin teaches in which the modification comprises applying to the image data one or more modifications selected from a set comprising: brightness changes; contrast changes; saturation changes; cropping; rotation; and addition of noise. ([Tobin, page 3, sec III] “We randomize the following aspects of the domain for each sample used during training: Number and shape of distractor objects on the table; Position and texture of all objects on the table; Textures of the table, floor, skybox, and robot; Position, orientation, and field of view of the camera; Number of lights in the scene; Position, orientation, and specular characteristics of the lights; Type and amount of random noise added to images.”, wherein the examiner interprets randomizing textures, positions, camera orientation, lighting characteristics, and particularly adding random noise to images to be the same as applying modifications comprising brightness changes, contrast changes, saturation changes, cropping, rotation, and addition of noise.)
Ho, Blonde, Bousmalis, Peng, Baram, Tobin, and the instant application are analogous art because they are all directed to modifying image data to enhance robustness of machine learning models by applying transformations such as brightness, contrast, rotation, and noise addition.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method of claim 15 disclosed by Ho, Blonde, Bousmalis, Peng, and Baram to include the domain randomization disclosed by Tobin. One would be motivated to do so to effectively improve the generalization capabilities of machine learning models, as suggested by Tobin ([Tobin, page 3, sec III] “provide enough simulated variability at training time such that at test time the model is able to generalize to real-world data.”).
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DEVAN KAPOOR whose telephone number is (703)756-1434. The examiner can normally be reached Monday - Friday: 9:00AM - 5:00 PM EST (times may vary).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, David Yi can be reached at (571) 270-7519. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/DEVAN KAPOOR/Examiner, Art Unit 2126
/DAVID YI/Supervisory Patent Examiner, Art Unit 2126