Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
1. This office action is in response to the amendment filed on 06/24/2026. Claim 10 is canceled and claims 1-9 and 11-20 are pending and have been considered below.
2. The rejections of claims 1-20 under 35 U.S. C. 101 as directed to an abstract idea without significantly more are moot pursuant to claims amendments and arguments.
Claim Rejections - 35 USC § 103
3. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-9 and 11-20 are rejected under 35 U.S.C. 102(a)(1) as being unpatentable over Bootstrapped Q-learning with Context Relevant Observation Pruning to Generalize in Text-based Games, Chaudhury et al, hereafter Chaudhury in view of Distilling the Knowledge in a Neural Network, Hinton et al., hereinafter Hinton.
Regarding claim 1, Chaudhury teaches:
training a teacher model on an environment ([Page 1, Column 2, Paragraph 1] teaches CREST, which first trains an overfitted base model on the original observation text in training games using Q-learning.)
generating action scores for actions that are performable within the environment using the teacher model. ([Page 3, Column 2, Paragraph 2] teaches Token Relevance Distribution (TRD): We run inference on the overfitted base model for each training game (indexed by k) and aggregate all the action tokens issued for that particular game as the Episodic Action Token Aggregation (EATA).)
generating pruned states of the environment based on the first soft labels of the actions (Section 1 explains that the base model's action-token information is used for observation pruning and that tokens not semantically related to the base policy's action tokens are removed p. 3, §1….Section 3.2, entitled “Context Relevant Episodic State Truncation (CREST),” explains that the CREST module removes unwanted observation tokens and uses action commands generated by the base policy to determine which tokens are contextually relevant. Id., pp. 4–5, §3.2. calculates a Token Relevance Distribution and uses a threshold to create a hard attention mask for pruning the observation text. Id., p. 5, §3.2. Figure 2(a), p.4, expressly depicts the sequence: Base Model → Episodic Action Tokens → Token Relevance Distribution → threshold → Pruned observation text → Train Bootstrapped Policy.)
training a student model using pruned states of the environment. ([Page 1, Column 2, Paragraph 1] teaches first trains an overfitted base model on the original observation text in training games, apply observation pruning, then re-train a bootstrapped policy on the pruned observation text.)
performing, by the retrained student model, tasks in a new environment, wherein the new environment is different from the environment [Chaudhury expressly identifies the problem of generalization to unseen games. The abstract states that the bootstrapped agent improves generalization in solving unseen TextWorld games p.2. Section 4 reports testing on unseen games and states that the method provides improved out-of-sample generalization. The zero-shot-transfer discussion further describes agents trained on games with 15-room quest lengths being tested on unseen configurations having 20- and 25-room quest lengths without retraining. Id., p.6, §4 (“Zero-shot transfer”))
Chaudhury does not teach:
wherein the action scores correspond to first soft labels; generating, using the teacher model, hard labels based on the first soft labels; generating, using the student model, second soft labels for the actions; determining cross-entropy loss between the second soft labels of the student model and the hard labels of the teacher model; and retraining the student model based on the cross-entropy loss, wherein the retraining of the student model comprises updating parameters of the student model.
However, Hinton discloses wherein the action scores correspond to first soft labels (converting teacher outputs into a soft target distribution. Hinton explains that a transfer set is provided with a soft target distribution generated by the cumbersome/teacher model using a high-temperature softmax. Hinton, p. 2, §2. Hinton further explains that the distilled model is trained against that soft target distribution.); generating, using the teacher model, hard labels based on the first soft labels (Hinton expressly teaches generating correct/hard labels in addition to soft targets and training the distilled model using both. Hinton states that, when correct labels are known, the distilled model may additionally be trained to produce those correct labels p. 2, §2. The reference distinguishes the cross-entropy objective using soft targets from the cross-entropy objective using the correct labels); generating, using the student model, second soft labels for the actions (Hinton teaches that the student produces an output distribution and that the student distribution is trained against the teacher's soft target distribution. p. 2, §2); determining cross-entropy loss between the second soft labels of the student model and the hard labels of the teacher model (Hinton describes a first objective based on cross entropy with soft targets and a second objective based on cross entropy with correct labels p. 2, §2); and retraining the student model based on the cross-entropy loss, wherein the retraining of the student model comprises updating parameters of the student model (Hinton expressly teaches training the distilled model using the cross-entropy objectives p. 2, §2). Therefore, It would have been obvious to one of ordinary skill in the art, at or before the effective filing date of the instant application, to use the feature of Hinton in Chaudhury. One would have been motivated to preserve the teacher's action-selection behavior while training on reduced/pruned observations and facilitate a student policy that retains the teacher's useful action information while benefiting from the reduced observation representation.
Regarding claim 2, the claim recites similar limitation as corresponding to claim 1 and is rejected for similar reasons as claim 1 using similar teachings and rationale. Chaudhury further teaches:
teacher model includes extracting unpruned states from the environment. ([Page 1, Column 2, Paragraph 1] teaches To alleviate this problem, we propose CREST, which first trains an overfitted (unpruned) base (teacher) model on the original observation text in training games (environment) using Q-learning.
Regarding claim 3, the claim recites similar limitation as corresponding to claim 1 and is rejected for similar reasons as claim 1 using similar teachings and rationale. Chaudhury further teaches:
truncating the unpruned states for the generating of the pruned states of the environment. ([Page 1, Column 2, Paragraph 1, section 3.2] teaches we propose CREST, which first trains an overfitted (unpruned) base model on the original observation text in training games using Q-learning. Subsequently, we apply observation pruning (truncating) such that, for each episode of the training games, we remove the observation tokens that are not semantically related to the base policy's action tokens.
Regarding claim 4, the claim recites similar limitation as corresponding to claim 1 and is rejected for similar reasons as claim 1 using similar teachings and rationale. Chaudhury further teaches:
the pruned states of the environment include verb-noun pairs. ([Page 1, Column 2, Paragraph 1] teaches we re-train a bootstrapped policy on the pruned observation text using Q-learning that improves generalization by removing irrelevant tokens. Figure 1 shows an illustrative example of our method. [Page 2, Column 2, Paragraph 1] teaches The action consists of a combination of verb and object output, such as "go north", "take coin", etc.)
PNG
media_image1.png
353
360
media_image1.png
Greyscale
Regarding claim 5, the claim recites similar limitation as corresponding to claim 1 and is rejected for similar reasons as claim 1 using similar teachings and rationale. Chaudhury further teaches:
The retraining of the student model is further based on a distillation loss and the distillation loss uses temperature annealing and the cross-entropy (Hinton expressly teaches knowledge distillation using soft target distributions generated at an elevated temperature and a cross-entropy objective, p. 2, §2. Hinton states that the same high temperature is used in training the distilled model and explains that the soft-target objective is a cross entropy. Hinton further teaches combining the soft-target cross-entropy objective with a hard/correct-label cross-entropy objective, p. 2, §2). One would have been motivated to preserve the teacher's action-selection behavior while training on reduced/pruned observations
Regarding claim 6, the claim recites similar limitation as corresponding to claim 1 and is rejected for similar reasons as claim 1 using similar teachings and rationale. Chaudhury further teaches:
training of the teacher model includes determining weight values of the teacher model that minimize a loss function. ([Page 3, Column 2, Paragraph 3] teaches Token Relevance Distribution (TRD): We run inference on the overfitted base model for each training game (indexed by k) and aggregate all the action tokens issued for that particular game as the Episodic Action Token Aggregation (EATA), Ak. For each token Wi in a given observation text o} at step t for the k th game, we compute the Token Relevance Distribution (TRD) C"..."This relevance score is used to prune irrelevant tokens in the observation text by creating a hard attention mask using a threshold value.)
the loss function includes a policy function that predicts a reward for performing an action given a present state and an expectation value for a state-action pair. ([Page 3, Column 1, Paragraph 1] teaches The parameters of the model are updated, by optimizing the following loss function obtained from the Bellman equation (Sutton et al., 1998),where Q(s, a)(where Q is expectation value, s is present state, and a is action) is obtained as the average of verb and object Q-values, r E (0, 1) is the discount factor. The agent is given a reward of 1 from the environment on completing the objective.)
Regarding claim 7, the claim recites similar limitation as corresponding to claim 1 and is rejected for similar reasons as claim 1 using similar teachings and rationale. Chaudhury further teaches:
teacher model includes a series of long short-term memory neural network layers. ([Page 7, Column 1, Paragraph 2] teaches We use a single LSTM network (teacher) with 100 dimensional hidden units (series of LSTM network layers) in the representation generator. For the action scorer, a single LSTM network with 64-dim hidden unit (for DRQN ) and two MLPs for verb and object Q-values were used. The number of trainable parameters in our policy network is 128,628 for the model with attention and 125,364 without attention.)
Regarding claim 8, the claim recites similar limitation as corresponding to claim 1 and is rejected for similar reasons as claim 1 using similar teachings and rationale. Chaudhury further teaches:
the student model has a same neural network structure as the teacher model [Page 3, Column 2, Paragraph 3] teaches the bootstrapped model is trained on the pruned observation text by removing irrelevant tokens using TRDs. Same model architecture and training methods as the base model are used.
Regarding claim 9, the claim recites similar limitation as corresponding to claim 1 and is rejected for similar reasons as claim 1 using similar teachings and rationale. Chaudhury further teaches:
each of the environment and the new environment correspond to a text game and the actions include commands that an agent can perform within the text game. ([Page 1, Column 1, Paragraph 2, sections 3.1, 4] teaches to interact with the environment, the agent issues text-based action commands ("go west") upon which it receives a reward signal used for training the RL agent.)
Regarding claim 11, the claim recites similar limitation as corresponding claim 1 and is rejected for similar reasons as claim 1 using similar teachings and rationale. Chaudhury also teaches:
Regarding claim 12, the claim recites similar limitation as corresponding claim 2 and is rejected for similar reasons as claim 2 using similar teachings and rationale. Claim 12 also inherits the deficiencies based on dependence of claim 11 and is rejected using similar teachings and rationale.
Regarding claim 13, the claim recites similar limitation as corresponding claim 3 and is rejected for similar reasons as claim 3 using similar teachings and rationale. Claim 13 also inherits the deficiencies based on dependence of claim 11 and is rejected using similar teachings and rationale.
Regarding claim 14, the claim recites similar limitation as corresponding claim 4 and is rejected for similar reasons as claim 4 using similar teachings and rationale. Claim 14 also inherits the deficiencies based on dependence of claim 11 and is rejected using similar teachings and rationale.
Regarding claim 15, the claim recites similar limitation as corresponding claim 5 and is rejected for similar reasons as claim 5 using similar teachings and rationale. Claim 15 also inherits the deficiencies based on dependence of claim 11 and is rejected using similar teachings and rationale.
Regarding claim 16, the claim recites similar limitation as corresponding claim 6 and is rejected for similar reasons as claim 6 using similar teachings and rationale. Claim 16 also inherits the deficiencies based on dependence of claim 11 and is rejected using similar teachings and rationale.
Regarding claim 17, the claim recites similar limitation as corresponding claim 7 and is rejected for similar reasons as claim 7 using similar teachings and rationale. Claim 17 also inherits the deficiencies based on dependence of claim 11 and is rejected using similar teachings and rationale.
Regarding claim 18, the claim recites similar limitation as corresponding claim 8 and is rejected for similar reasons as claim 8 using similar teachings and rationale. Claim 18 also inherits the deficiencies based on dependence of claim 11 and is rejected using similar teachings and rationale.
Regarding claim 19, the claim recites similar limitation as corresponding claim 9 and is rejected for similar reasons as claim 9 using similar teachings and rationale. Claim 19 also inherits the deficiencies based on dependence of claim 11 and is rejected using similar teachings and rationale.
Regarding claim 20, the claim recites similar limitation as corresponding claim 1 and is rejected for similar reasons as claim 1 using similar teachings and rationale. Chaudhury also teaches:
A machine learning system, comprising: a hardware processor; and a memory that stores a computer program. [Page 7, Column 1, Paragraph 2] teaches Our experiments were conducted on a Ubuntul6.04 system with a Titan X (Pascal) GPU. (Where Ubuntu16.04 requires processors, RAM and hard drive space to operate.)
Response to Arguments
4. Applicant’s arguments filed 06/24/2026 have been fully considered but they are moot in light of new ground of rejection(s).
Conclusion
5. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
6. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure (See PTO-892).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Phenuel S. Salomon whose telephone number is (571) 270-1699. The examiner can normally be reached on Mon-Fri 7:00 A.M. to 4:00 P.M. (Alternate Friday Off) EST.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Usmaan Saeed can be reached on (571) 272-4046. The fax phone number for the organization where this application or proceeding is assigned is 571-273-3800.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PHENUEL S SALOMON/Primary Examiner, Art Unit 2146