Prosecution Insights
Last updated: October 02, 2026
Application No. 18/424,687

GENERATING ENVIRONMENT MODELS USING IN-CONTEXT ADAPTATION AND EXPLORATION

Non-Final OA §101§103§112
Filed
Jan 26, 2024
Priority
Jan 26, 2023 — provisional 63/441,425
Examiner
ACOSTA, RILEY SULLIVAN
Art Unit
Tech Center
Assignee
DeepMind Technologies Limited
OA Round
1 (Non-Final)
100%
Grant Probability
Favorable
1-2
OA Rounds
5m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 100% — above average
100%
Career Allowance Rate
1 granted / 1 resolved
+40.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
3y 1m
Avg Prosecution
21 currently pending
Career history
6
Total Applications
across all art units

Statute-Specific Performance

§101
25.0%
-15.0% vs TC avg
§103
54.4%
+14.4% vs TC avg
§102
8.7%
-31.3% vs TC avg
§112
9.8%
-30.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1 resolved cases

Office Action

§101 §103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This action is responsive to the application filed 01/26/2024. Claims 1-21 are presented for examination. Priority Applicant’s claim for the benefit of a provisionally filed application, filed 01/26/2023, is acknowledged. Information Disclosure Statement The information disclosure statement (IDS) submitted 08/01/2024, has been considered by the examiner. Claim Objections Claims 4 & 12 are objected to because of the following informality: Claim 4 recites ‘wherein action selection policy’; however, it should recite - - wherein the action selection policy - -. Claim 12 recites ‘wherein current graph model represents a Markov decision process (MDP) defining transitions between different states the environment’; however, it should recite - - wherein the current graph model represents a Markov decision process (MDP) defining transitions between different states of the environment - -. Appropriate correction is required. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 6 & 8 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. The following claims lack antecedent basis: Claim 6, line 1, the phrase “the action selection policy”; Claim 6, line 2, the phrase “the dynamic programming technique”; Claim 6, line 3, the phrase “the value iteration technique”; Claim 8, line 1, the phrase “the action selection policy”; and Claim 8, line 3, the phrase “the action selection policy”. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-21 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Claim 1 Step 1: The claim recites “A method for controlling an agent interacting with an environment to perform a task, and wherein the method comprises, at a current time step of multiple time steps”; therefore, it is directed to the statutory category of a method. Step 2A Prong 1: The claim recites, inter alia: generating a current graph model that represents the environment, wherein the current graph model comprises vertices that represent states of the environment and edges connecting the vertices, wherein an edge between a first vertex and a second vertex in the graph model indicates it is possible that the environment will transition from a state represented by the first vertex into a state represented by the second vertex as a result of one of a possible set of actions performed by the agent, and wherein generating the current graph model comprises: processing a first Transformer network input that includes the context data using a Transformer neural network to generate a probability distribution over a possible set of edges that can be included in the current graph model: These limitations recite mathematical concepts similar to organizing information and manipulating information through mathematical correlations per MPEP 2106.04(a)(2)(I)(A)(iv). and selecting, in accordance with the probability distribution over the possible set of edges, a subset of the edges to be included in the current graph model: These limitations recite a mentally performable process with the aid of pen and paper of using observation and judgement to select a subset of the edges to be included in the current graph model, in accordance with the probability distribution. selecting, from the possible set of actions and using the current graph model, a current action to be performed by the agent in response to the current observation: These limitations recite a mentally performable process with the aid of pen and paper of using observation and judgement to select, from the possible set of actions and using the current graph model, a current action to be performed by the agent in response to the current observation. Thus, the claim recites a judicial exception. Step 2A Prong 2: This judicial exception is not integrated into a practical application. The additional elements of the claim are as follows: maintaining context data that includes, for each of one or more previous time steps that precede the current time step, (i) data identifying a previous action performed by the agent at the previous time step and (ii) a previous observation characterizing a previous state that the environment transitioned into as a result of the agent performing the previous action: These additional elements amount to insignificant extra-solution activity in the form of mere data gathering per MPEP § 2106.05(g). receiving a current observation characterizing a current state of the environment: These additional elements amount to insignificant extra-solution activity in the form of mere data gathering per MPEP § 2106.05(g). controlling the agent to perform the selected current action to cause the environment to transition from the current state into a new state: These additional elements recite only the idea of controlling the agent to perform the selected current action to cause the environment to transition from the current state into a new state and attempts to cover any implementation of controlling the agent without any restriction as to how the agent is controlled, or specifically what mechanism does the controlling of the agent. Thus, these additional elements do not meaningfully limit the claim and do not integrate the judicial exception into a practical application because this type of recitation is equivalent to the words "apply it". See MPEP 2106.05(f). and updating the context data to include (i) data identifying the selected current action and (ii) a new observation characterizing the new state of the environment: These additional elements amount to insignificant extra-solution activity in the form of mere data gathering per MPEP § 2106.05(g). Thus, the way in which the additional elements use or interact with the judicial exception do not integrate the judicial exception into a practical application. Step 2B: The additional elements from Step 2A Prong 2 include adding words equivalent to "apply it" with the judicial exception and insignificant extra-solution activity of data gathering recited by “maintaining context data that includes, for each of one or more previous time steps that precede the current time step, (i) data identifying a previous action performed by the agent at the previous time step and (ii) a previous observation characterizing a previous state that the environment transitioned into as a result of the agent performing the previous action; receiving a current observation characterizing a current state of the environment; and updating the context data to include (i) data identifying the selected current action and (ii) a new observation characterizing the new state of the environment” which are well-understood routine and conventional activities similar to presenting offers and gathering statistics per MPEP 2106.05(d)(II). Thus, the additional elements, viewed individually or in combination, do not provide an inventive concept or otherwise amount to significantly more than the abstract idea itself. See MPEP § 2106.05. Claim 2 Step 1: A process, as above. Step 2A Prong 1: The claim recites the abstract ideas of claim 1 as well as, inter alia: determining a respective value for each of the possible set of edges using the probability distribution: These limitations recite a mentally performable process with the aid of pen and paper of using observation and judgement to determine a respective value for each of the possible set of edges using the probability distribution. and selecting, from the possible set of edges, one or more edges having determined values that satisfy a threshold value: These limitations recite a mentally performable process with the aid of pen and paper of using observation, judgement, and evaluation to select, from the possible set of edges, one or more edges having determined values that satisfy a threshold value. Thus, the claim recites a judicial exception. Step 2A Prong 2 & Step 2B: There are no additional elements recited so the claim does not provide a practical application and is not considered to be significantly more. As such, the claim is patent ineligible. Claim 3 Step 1: A process, as above. Step 2A Prong 1: The claim recites the abstract ideas of claim 1 as well as, inter alia: generating a reward value for each vertex included in the graph model: These limitations recite mathematical calculations similar to an act of calculating using mathematical methods to determine a variable or number per MPEP 2106.04(a)(2)(I)(C). Thus, the claim recites a judicial exception. Step 2A Prong 2: This judicial exception is not integrated into a practical application. The additional elements of the claim are as follows: by using the Transformer neural network: These additional elements are recited at a high level of generality and amount to invoking computers or other machinery merely as a tool to apply the underlying judicial exception. See MPEP § 2106.05(f). Thus, the way in which the additional elements use or interact with the judicial exception do not integrate the judicial exception into a practical application. Step 2B: The additional elements from Step 2A Prong 2 include invoking generic computer components to apply the underlying judicial exception. Thus, the additional elements, viewed individually or in combination, do not provide an inventive concept or otherwise amount to significantly more than the abstract idea itself. See MPEP § 2106.05. Claim 4 Step 1: A process, as above. Step 2A Prong 1: The claim recites the abstract ideas of claim 1 as well as, inter alia: determining an action selection policy that corresponds to the current graph model: These limitations recite a mentally performable process with the aid of pen and paper of using observation and judgement to determine an action selection policy that corresponds to the current graph model. and using the action selection policy to select the current action: These limitations recite a mentally performable process with the aid of pen and paper of using observation and judgement to select the current action using the action selection policy. Thus, the claim recites a judicial exception. Step 2A Prong 2: This judicial exception is not integrated into a practical application. The additional elements of the claim are as follows: by using a dynamic programming technique: These additional elements recite only the idea of using a dynamic programming technique and attempts to cover any implementation of dynamic programming techniques without any restriction as to the specific dynamic programming technique or how this technique determines an action selection policy. Thus, these additional elements do not meaningfully limit the claim and do not integrate the judicial exception into a practical application because this type of recitation is equivalent to the words "apply it". See MPEP 2106.05(f). wherein action selection policy specifies a mapping from the vertices to the edges included in the current graph model: These additional elements are recited at a high level of generality and merely indicate a field of use or technological environment in which to apply a judicial exception, e.g. the action selection policy, to a particular technological environment or field of use, e.g. specifies a mapping from the vertices to the edges included in the current graph model. See MPEP 2106.05(h). Elements that use or interact with the judicial exception do not integrate the judicial exception into a practical application. Thus, the way in which the additional elements use or interact with the judicial exception do not integrate the judicial exception into a practical application. Step 2B: The additional elements from Step 2A Prong 2 include adding words equivalent to "apply it" with the judicial exception and generally linking the use of the judicial exception to indicate a field of use or technological environment. Thus, the additional elements, viewed individually or in combination, do not provide an inventive concept or otherwise amount to significantly more than the abstract idea itself. See MPEP § 2106.05. Claim 5 Step 1: A process, as above. Step 2A Prong 1: The claim recites the abstract ideas as the judicial exception of claim 4. Step 2A Prong 2: This judicial exception is not integrated into a practical application. The additional elements of the claim are as follows: wherein the dynamic programming technique comprises a value iteration technique: These additional elements are recited at a high level of generality and merely indicate a field of use or technological environment in which to apply a judicial exception, e.g. the dynamic programming technique, to a particular technological environment or field of use, e.g. comprises a value iteration technique. See MPEP 2106.05(h). Elements that use or interact with the judicial exception do not integrate the judicial exception into a practical application. Thus, the way in which the additional elements use or interact with the judicial exception do not integrate the judicial exception into a practical application. Step 2B: The additional elements from Step 2A Prong 2 include generally linking the use of the judicial exception to indicate a field of use or technological environment. Thus, the additional elements, viewed individually or in combination, do not provide an inventive concept or otherwise amount to significantly more than the abstract idea itself. See MPEP § 2106.05. Claim 6 Step 1: A process, as above. Step 2A Prong 1: The claim recites the abstract ideas as the judicial exception of claim 3. Step 2A Prong 2: This judicial exception is not integrated into a practical application. The additional elements of the claim are as follows: wherein determining the action selection policy that corresponds to the current graph model by using the dynamic programming technique comprises: performing the value iteration technique using the edges included in the current graph model and the reward values included in the current graph model: These additional elements are recited at a high level of generality and merely indicate a field of use or technological environment in which to apply a judicial exception, e.g. determining the action selection policy that corresponds to the current graph model by using the dynamic programming technique, to a particular technological environment or field of use, e.g. comprises performing the value iteration technique using the edges included in the current graph model and the reward values included in the current graph model. See MPEP 2106.05(h). Elements that use or interact with the judicial exception do not integrate the judicial exception into a practical application. Thus, the way in which the additional elements use or interact with the judicial exception do not integrate the judicial exception into a practical application. Step 2B: The additional elements from Step 2A Prong 2 include generally linking the use of the judicial exception to indicate a field of use or technological environment. Thus, the additional elements, viewed individually or in combination, do not provide an inventive concept or otherwise amount to significantly more than the abstract idea itself. See MPEP § 2106.05. Claim 7 Step 1: A process, as above. Step 2A Prong 1: The claim recites the abstract ideas of claim 1 as well as, inter alia: processing a second Transformer network input that includes the current observation using the Transformer neural network to generate a probability distribution over the vertices included in the graph model: These limitations recite mathematical concepts similar to organizing information and manipulating information through mathematical correlations per MPEP 2106.04(a)(2)(I)(A)(iv). and selecting, in accordance with the probability distribution over the vertices included in the graph model, a selected vertex as a vertex in the graph model that corresponds to the current observation: These limitations recite a mentally performable process with the aid of pen and paper of using observation and judgement to select, in accordance with the probability distribution over the vertices included in the graph model, a selected vertex as a vertex in the graph model that corresponds to the current observation. Thus, the claim recites a judicial exception. Step 2A Prong 2 & Step 2B: There are no additional elements recited so the claim does not provide a practical application and is not considered to be significantly more. As such, the claim is patent ineligible. Claim 8 Step 1: A process, as above. Step 2A Prong 1: The claim recites the abstract ideas of claim 7 as well as, inter alia: querying the action selection policy using the selected vertex: These limitations recite a mentally performable process with the aid of pen and paper of using observation and judgement to query the action selection policy using the selected vertex. Thus, the claim recites a judicial exception. Step 2A Prong 2 & Step 2B: There are no additional elements recited so the claim does not provide a practical application and is not considered to be significantly more. As such, the claim is patent ineligible. Claim 9 Step 1: A process, as above. Step 2A Prong 1: The claim recites the abstract ideas as the judicial exception of claim 1. Step 2A Prong 2: This judicial exception is not integrated into a practical application. The additional elements of the claim are as follows: wherein the possible set of actions comprise a no-op action: These additional elements are recited at a high level of generality and merely indicate a field of use or technological environment in which to apply a judicial exception, e.g. the possible set of actions, to a particular technological environment or field of use, e.g. comprises a no-op action. See MPEP 2106.05(h). Elements that use or interact with the judicial exception do not integrate the judicial exception into a practical application. Thus, the way in which the additional elements use or interact with the judicial exception do not integrate the judicial exception into a practical application. Step 2B: The additional elements from Step 2A Prong 2 include generally linking the use of the judicial exception to indicate a field of use or technological environment. Thus, the additional elements, viewed individually or in combination, do not provide an inventive concept or otherwise amount to significantly more than the abstract idea itself. See MPEP § 2106.05. Claim 10 Step 1: A process, as above. Step 2A Prong 1: The claim recites the abstract ideas as the judicial exception of claim 1. Step 2A Prong 2: This judicial exception is not integrated into a practical application. The additional elements of the claim are as follows: wherein the current graph models at different time steps include different numbers of edges, or a same number of edges that connect the vertices in different ways: These additional elements are recited at a high level of generality and merely indicate a field of use or technological environment in which to apply a judicial exception, e.g. the current graph models at different time steps, to a particular technological environment or field of use, e.g. includes different numbers of edges, or a same number of edges that connect the vertices in different ways. See MPEP 2106.05(h). Elements that use or interact with the judicial exception do not integrate the judicial exception into a practical application. Thus, the way in which the additional elements use or interact with the judicial exception do not integrate the judicial exception into a practical application. Step 2B: The additional elements from Step 2A Prong 2 include generally linking the use of the judicial exception to indicate a field of use or technological environment. Thus, the additional elements, viewed individually or in combination, do not provide an inventive concept or otherwise amount to significantly more than the abstract idea itself. See MPEP § 2106.05. Claim 11 Step 1: A process, as above. Step 2A Prong 1: The claim recites the abstract ideas as the judicial exception of claim 1. Step 2A Prong 2: This judicial exception is not integrated into a practical application. The additional elements of the claim are as follows: wherein parameter values of the Transformer neural network are fixed during the multiple time steps: These additional elements are recited at a high level of generality and merely indicate a field of use or technological environment in which to apply a judicial exception, e.g. parameter values of the Transformer neural network, to a particular technological environment or field of use, e.g. are fixed during the multiple time steps. See MPEP 2106.05(h). Elements that use or interact with the judicial exception do not integrate the judicial exception into a practical application. Thus, the way in which the additional elements use or interact with the judicial exception do not integrate the judicial exception into a practical application. Step 2B: The additional elements from Step 2A Prong 2 include generally linking the use of the judicial exception to indicate a field of use or technological environment. Thus, the additional elements, viewed individually or in combination, do not provide an inventive concept or otherwise amount to significantly more than the abstract idea itself. See MPEP § 2106.05. Claim 12 Step 1: A process, as above. Step 2A Prong 1: The claim recites the abstract ideas as the judicial exception of claim 1. Step 2A Prong 2: This judicial exception is not integrated into a practical application. The additional elements of the claim are as follows: wherein current graph model represents a Markov decision process (MDP) defining transitions between different states the environment: These additional elements are recited at a high level of generality and merely indicate a field of use or technological environment in which to apply a judicial exception, e.g. current graph model, to a particular technological environment or field of use, e.g. represents a Markov decision process (MDP) defining transitions between different states the environment. See MPEP 2106.05(h). Elements that use or interact with the judicial exception do not integrate the judicial exception into a practical application. Thus, the way in which the additional elements use or interact with the judicial exception do not integrate the judicial exception into a practical application. Step 2B: The additional elements from Step 2A Prong 2 include generally linking the use of the judicial exception to indicate a field of use or technological environment. Thus, the additional elements, viewed individually or in combination, do not provide an inventive concept or otherwise amount to significantly more than the abstract idea itself. See MPEP § 2106.05. Claim 13 Step 1: A process, as above. Step 2A Prong 1: The claim recites the abstract ideas as the judicial exception of claim 1. Step 2A Prong 2: This judicial exception is not integrated into a practical application. The additional elements of the claim are as follows: wherein the agent is a mechanical agent and the environment is a real-world environment: These additional elements are recited at a high level of generality and merely indicate a field of use or technological environment in which to apply a judicial exception, e.g. the agent, to a particular technological environment or field of use, e.g. is a mechanical agent and the environment is a real-world environment. See MPEP 2106.05(h). Elements that use or interact with the judicial exception do not integrate the judicial exception into a practical application. Thus, the way in which the additional elements use or interact with the judicial exception do not integrate the judicial exception into a practical application. Step 2B: The additional elements from Step 2A Prong 2 include generally linking the use of the judicial exception to indicate a field of use or technological environment. Thus, the additional elements, viewed individually or in combination, do not provide an inventive concept or otherwise amount to significantly more than the abstract idea itself. See MPEP § 2106.05. Claim 14 Step 1: A process, as above. Step 2A Prong 1: The claim recites the abstract ideas as the judicial exception of claim 13. Step 2A Prong 2: This judicial exception is not integrated into a practical application. The additional elements of the claim are as follows: wherein the agent is a robot: These additional elements are recited at a high level of generality and merely indicate a field of use or technological environment in which to apply a judicial exception, e.g. the agent, to a particular technological environment or field of use, e.g. is a robot. See MPEP 2106.05(h). Elements that use or interact with the judicial exception do not integrate the judicial exception into a practical application. Thus, the way in which the additional elements use or interact with the judicial exception do not integrate the judicial exception into a practical application. Step 2B: The additional elements from Step 2A Prong 2 include generally linking the use of the judicial exception to indicate a field of use or technological environment. Thus, the additional elements, viewed individually or in combination, do not provide an inventive concept or otherwise amount to significantly more than the abstract idea itself. See MPEP § 2106.05. Claim 15 Step 1: A process, as above. Step 2A Prong 1: The claim recites the abstract ideas as the judicial exception of claim 1. Step 2A Prong 2: This judicial exception is not integrated into a practical application. The additional elements of the claim are as follows: wherein the environment is a real-world environment of a service facility comprising a plurality of items of electronic equipment and the agent is an electronic agent configured to control operation of the service facility: These additional elements are recited at a high level of generality and merely indicate a field of use or technological environment in which to apply a judicial exception, e.g. the environment, to a particular technological environment or field of use, e.g. is a real-world environment of a service facility comprising a plurality of items of electronic equipment and the agent is an electronic agent configured to control operation of the service facility. See MPEP 2106.05(h). Elements that use or interact with the judicial exception do not integrate the judicial exception into a practical application. Thus, the way in which the additional elements use or interact with the judicial exception do not integrate the judicial exception into a practical application. Step 2B: The additional elements from Step 2A Prong 2 include generally linking the use of the judicial exception to indicate a field of use or technological environment. Thus, the additional elements, viewed individually or in combination, do not provide an inventive concept or otherwise amount to significantly more than the abstract idea itself. See MPEP § 2106.05. Claim 16 Step 1: A process, as above. Step 2A Prong 1: The claim recites the abstract ideas as the judicial exception of claim 1. Step 2A Prong 2: This judicial exception is not integrated into a practical application. The additional elements of the claim are as follows: wherein the environment is a real-world manufacturing environment for manufacturing a product and the agent comprises an electronic agent configured to control a manufacturing unit or a machine that operates to manufacture the product: These additional elements are recited at a high level of generality and merely indicate a field of use or technological environment in which to apply a judicial exception, e.g. the environment, to a particular technological environment or field of use, e.g. is a real-world manufacturing environment for manufacturing a product and the agent comprises an electronic agent configured to control a manufacturing unit or a machine that operates to manufacture the product. See MPEP 2106.05(h). Elements that use or interact with the judicial exception do not integrate the judicial exception into a practical application. Thus, the way in which the additional elements use or interact with the judicial exception do not integrate the judicial exception into a practical application. Step 2B: The additional elements from Step 2A Prong 2 include generally linking the use of the judicial exception to indicate a field of use or technological environment. Thus, the additional elements, viewed individually or in combination, do not provide an inventive concept or otherwise amount to significantly more than the abstract idea itself. See MPEP § 2106.05. Claim 17 Step 1: A process, as above. Step 2A Prong 1: The claim recites the abstract ideas as the judicial exception of claim 1. Step 2A Prong 2: This judicial exception is not integrated into a practical application. The additional elements of the claim are as follows: wherein the environment is a simulation of a real-world environment: These additional elements are recited at a high level of generality and merely indicate a field of use or technological environment in which to apply a judicial exception, e.g. the environment, to a particular technological environment or field of use, e.g. is a simulation of a real-world environment. See MPEP 2106.05(h). Elements that use or interact with the judicial exception do not integrate the judicial exception into a practical application. Thus, the way in which the additional elements use or interact with the judicial exception do not integrate the judicial exception into a practical application. Step 2B: The additional elements from Step 2A Prong 2 include generally linking the use of the judicial exception to indicate a field of use or technological environment. Thus, the additional elements, viewed individually or in combination, do not provide an inventive concept or otherwise amount to significantly more than the abstract idea itself. See MPEP § 2106.05. Claim 18 Step 1: A process, as above. Step 2A Prong 1: The claim recites the abstract ideas as the judicial exception of claim 1. Step 2A Prong 2: This judicial exception is not integrated into a practical application. The additional elements of the claim are as follows: wherein the agent is a digital assistant and wherein actions performed by the agent include outputs that are provided by the digital assistant to a user: These additional elements are recited at a high level of generality and merely indicate a field of use or technological environment in which to apply a judicial exception, e.g. the agent, to a particular technological environment or field of use, e.g. is a digital assistant and wherein actions performed by the agent include outputs that are provided by the digital assistant to a user. See MPEP 2106.05(h). Elements that use or interact with the judicial exception do not integrate the judicial exception into a practical application. Thus, the way in which the additional elements use or interact with the judicial exception do not integrate the judicial exception into a practical application. Step 2B: The additional elements from Step 2A Prong 2 include generally linking the use of the judicial exception to indicate a field of use or technological environment. Thus, the additional elements, viewed individually or in combination, do not provide an inventive concept or otherwise amount to significantly more than the abstract idea itself. See MPEP § 2106.05. Claim 19 Step 1: A process, as above. Step 2A Prong 1: The claim recites the abstract ideas as the judicial exception of claim 18. Step 2A Prong 2: This judicial exception is not integrated into a practical application. The additional elements of the claim are as follows: wherein the outputs include one or more of: text displayed to a user in a user interface of the digital assistant; an image displayed to the user in the user interface of the digital assistant; or speech output through one or more speakers of the digital assistant: These additional elements are recited at a high level of generality and merely indicate a field of use or technological environment in which to apply a judicial exception, e.g. the outputs, to a particular technological environment or field of use, e.g. include one or more of: text displayed to a user in a user interface of the digital assistant; an image displayed to the user in the user interface of the digital assistant; or speech output through one or more speakers of the digital assistant. See MPEP 2106.05(h). Elements that use or interact with the judicial exception do not integrate the judicial exception into a practical application. Thus, the way in which the additional elements use or interact with the judicial exception do not integrate the judicial exception into a practical application. Step 2B: The additional elements from Step 2A Prong 2 include generally linking the use of the judicial exception to indicate a field of use or technological environment. Thus, the additional elements, viewed individually or in combination, do not provide an inventive concept or otherwise amount to significantly more than the abstract idea itself. See MPEP § 2106.05. Claim 20 Step 1: This claim is directed to “A system comprising one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations for controlling an agent interacting with an environment to perform a task, and wherein the operations comprise, at a current time step of multiple time steps:”; therefore, it is directed the statutory category of a machine. Step 2A Prong 1: Claim 20 recites the same judicial exception as Claim 1. Step 2A Prong 2: The judicial exception recited in these claims is not integrated into a practical application. The analysis at this step for Claim 20 mirrors that of Claim 1. Step 2B: The additional elements from Step 2A Prong 2 do not contain significantly more than the judicial exception for this claim. The analysis at this step for Claim 20 mirrors that of Claim 1. Claim 21 Step 1: This claim recites "A computer storage medium encoded with instructions that, when executed by one or more computers, cause the one or more computers to perform operations for controlling an agent interacting with an environment to perform a task, and wherein the operations comprise, at a current time step of multiple time steps:"; therefore, it is directed to the statutory category of an article of manufacture. Step 2A Prong 1: Claim 21 recites the same judicial exception as Claim 1. Step 2A Prong 2: The judicial exception recited in these claims are not integrated into a practical application. The only difference between Claim 21 and Claim 1, is that Claim 21 is directed to "A computer storage medium encoded with instructions that, when executed by one or more computers, cause the one or more computers to perform operations for controlling an agent interacting with an environment to perform a task, and wherein the operations comprise, at a current time step of multiple time steps”. However, mere recitation that a judicial exception is to be performed using generic computer equipment in their ordinary capacity, i.e. a computer storage medium encoded with instructions that, when executed by one or more computers, cause the one or more computers to perform operations for controlling an agent interacting with an environment to perform a task, and wherein the operations comprise, at a current time step of multiple time steps, cannot meaningfully integrate the judicial exception into a practical application. See MPEP 2106.05(f). With that exception, the analysis at this step for Claim 21 mirrors that of Claim 1. Step 2B: The additional elements from Step 2A Prong 2 do not contain significantly more than the judicial exception for these claims. The only difference between Claim 21 and Claim 1, is that Claim 21 is directed to "A computer storage medium encoded with instructions that, when executed by one or more computers, cause the one or more computers to perform operations for controlling an agent interacting with an environment to perform a task, and wherein the operations comprise, at a current time step of multiple time steps”. However, mere recitation that a judicial exception is to be performed using generic computer equipment in their ordinary capacity, i.e. a computer storage medium encoded with instructions that, when executed by one or more computers, cause the one or more computers to perform operations for controlling an agent interacting with an environment to perform a task, and wherein the operations comprise, at a current time step of multiple time steps, cannot amount to significantly more than the judicial exception. See MPEP 2106.05(f). With that exception, the analysis at this step for Claim 21 mirrors that of Claim 1. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-6, 10, 12, 17, & 20-21 are rejected under 35 U.S.C. 103 as being unpatentable over Zhu et al. ("VALUE MEMORY GRAPH: A GRAPH-STRUCTURED WORLD MODEL FOR OFFLINE REINFORCEMENT LEARNING", arXiv) (Year: 2022), hereafter Zhu, in view of Fang et al. ("Scene Memory Transformer for Embodied Agents in Long-Horizon Tasks", Stanford, Google, arXiv) (Year: 2019), hereafter Fang, and further in view of Kipf et al. ("Neural Relational Inference for Interacting Systems", Proceedings of the 35th International Conference on Machine Learning, arXiv) (Year: 2018), hereafter Kipf. Regarding independent claim 1, Zhu teaches a method comprising: receiving a current observation characterizing a current state of the environment ([Sec. 3.4 & Alg. 2] discusses receiving a current state of the environment as input); generating a current graph model that represents the environment, wherein the current graph model comprises vertices that represent states of the environment and edges connecting the vertices, wherein an edge between a first vertex and a second vertex in the graph model indicates it is possible that the environment will transition from a state represented by the first vertex into a state represented by the second vertex as a result of one of a possible set of actions performed by the agent ([Sec. 3] discusses generating a graph wherein each vertex represents an observation of the state of the environment; further, two vertices are connected by an edge if they correspond to a transition between graph states, which constitutes a transition from the first vertex state to the second vertex state); selecting, from the possible set of actions and using the current graph model, a current action to be performed by the agent in response to the current observation ([Sec. 3.4 & Alg. 2] discusses selecting the best vertex destination in response to the current observation, using the current graph model to search through the set of possible actions); controlling the agent to perform the selected current action to cause the environment to transition from the current state into a new state ([Sec. 3.4] discusses controlling the agent to convert the graph action into the environment action, causing the environment to transition to a new state). Zhu does not explicitly teach maintaining context data that includes, for each of one or more previous time steps that precede the current time step, (i) data identifying a previous action performed by the agent at the previous time step and (ii) a previous observation characterizing a previous state that the environment transitioned into as a result of the agent performing the previous action; and wherein generating the current graph model comprises: processing a first Transformer network input that includes the context data using a Transformer neural network to generate a probability distribution over a possible set of edges that can be included in the current graph model; and selecting, in accordance with the probability distribution over the possible set of edges, a subset of the edges to be included in the current graph model; and updating the context data to include (i) data identifying the selected current action and (ii) a new observation characterizing the new state of the environment. However, in a similar field of endeavor, Fang teaches a method for controlling an agent wherein context data is maintained ([Abstract] discusses storing each observation to memory, which would include the previous action and the state of the environment); processing a first transformer network input using a Transformer neural network to produce a decision ([Sec. 3.2.2] discusses an attention-based transformer network architecture that processes the accumulated context to produce a decision at each step); and updating the context data ([Sec. 1] discusses updating and storing memory of each observation separately for each time step; thus, the data of the selected action and the observation characterizing the new environment is stored corresponding to its respective time step). Because Zhu teaches receiving a current observation, generating a graph model, wherein the current graph model comprises vertices that represent states of the environment and edges connecting the vertices, wherein an edge between a first vertex and a second vertex in the graph model indicates it is possible that the environment will transition from a state represented by the first vertex into a state represented by the second vertex as a result of one of a possible set of actions performed by the agent, selecting, from the possible set of actions and using the current graph model, a current action to be performed by the agent in response to the current observation, and controlling the agent to perform the selected current action to cause the environment to transition from the current state into a new state; and Fang teaches maintaining context data that includes, for each of one or more previous time steps that precede the current time step, (i) data identifying a previous action performed by the agent at the previous time step and (ii) a previous observation characterizing a previous state that the environment transitioned into as a result of the agent performing the previous action, processing a first Transformer network input that includes the context data using a Transformer neural network to generate a decision, and updating the context data to include (i) data identifying the selected current action and (ii) a new observation characterizing the new state of the environment, accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate maintaining context data that includes, for each of one or more previous time steps that precede the current time step, (i) data identifying a previous action performed by the agent at the previous time step and (ii) a previous observation characterizing a previous state that the environment transitioned into as a result of the agent performing the previous action, processing a first Transformer network input that includes the context data using a Transformer neural network to generate a decision, and updating the context data to include (i) data identifying the selected current action and (ii) a new observation characterizing the new state of the environment as taught by Fang into Zhu’s method, with a reasonable expectation of success, to teach maintaining context data that includes, for each of one or more previous time steps that precede the current time step, (i) data identifying a previous action performed by the agent at the previous time step and (ii) a previous observation characterizing a previous state that the environment transitioned into as a result of the agent performing the previous action; receiving a current observation characterizing a current state of the environment; generating a current graph model that represents the environment, wherein the current graph model comprises vertices that represent states of the environment and edges connecting the vertices, wherein an edge between a first vertex and a second vertex in the graph model indicates it is possible that the environment will transition from a state represented by the first vertex into a state represented by the second vertex as a result of one of a possible set of actions performed by the agent, and wherein generating the current graph model comprises: processing a first Transformer network input that includes the context data using a Transformer neural network to generate a decision; selecting, from the possible set of actions and using the current graph model, a current action to be performed by the agent in response to the current observation; controlling the agent to perform the selected current action to cause the environment to transition from the current state into a new state; and updating the context data to include (i) data identifying the selected current action and (ii) a new observation characterizing the new state of the environment. This combination would have been motivated by the desire to implement the method efficiently on long episodes requiring memory-based policies (Fang [Abstract]). The combination of Zhu and Fang does not explicitly teach using a Transformer neural network to generate a probability distribution over a possible set of edges that can be included in the current graph model; and selecting, in accordance with the probability distribution over the possible set of edges, a subset of the edges to be included in the current graph model. However, in a similar field of endeavor, Kipf teaches a method for computing a probability distribution using a given context over an edge set ([Sec. 3] discusses computing a probability distribution using context input over a set of edges); and selecting a subset of the edges, according to the probability distribution ([Sec. 3] discusses variables are sampled from a SoftMax relaxation of its probability distribution, which constitutes edge selection derived from the distribution over the set of edges). Because the combination of Zhu and Fang teaches maintaining context data that includes, for each of one or more previous time steps that precede the current time step, (i) data identifying a previous action performed by the agent at the previous time step and (ii) a previous observation characterizing a previous state that the environment transitioned into as a result of the agent performing the previous action; receiving a current observation characterizing a current state of the environment; generating a current graph model that represents the environment, wherein the current graph model comprises vertices that represent states of the environment and edges connecting the vertices, wherein an edge between a first vertex and a second vertex in the graph model indicates it is possible that the environment will transition from a state represented by the first vertex into a state represented by the second vertex as a result of one of a possible set of actions performed by the agent, and wherein generating the current graph model comprises: processing a first Transformer network input that includes the context data using a Transformer neural network to generate a decision; selecting, from the possible set of actions and using the current graph model, a current action to be performed by the agent in response to the current observation; controlling the agent to perform the selected current action to cause the environment to transition from the current state into a new state; and updating the context data to include (i) data identifying the selected current action and (ii) a new observation characterizing the new state of the environment; and Kipf teaches using a Transformer neural network to generate a probability distribution over a possible set of edges that can be included in the current graph model, and selecting, in accordance with the probability distribution over the possible set of edges, a subset of the edges to be included in the current graph model, accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate using a Transformer neural network to generate a probability distribution over a possible set of edges that can be included in the current graph model, and selecting, in accordance with the probability distribution over the possible set of edges, a subset of the edges to be included in the current graph model as taught by Kipf into the combination of Zhu and Fang’s method, with a reasonable expectation of success, to teach maintaining context data that includes, for each of one or more previous time steps that precede the current time step, (i) data identifying a previous action performed by the agent at the previous time step and (ii) a previous observation characterizing a previous state that the environment transitioned into as a result of the agent performing the previous action; receiving a current observation characterizing a current state of the environment; generating a current graph model that represents the environment, wherein the current graph model comprises vertices that represent states of the environment and edges connecting the vertices, wherein an edge between a first vertex and a second vertex in the graph model indicates it is possible that the environment will transition from a state represented by the first vertex into a state represented by the second vertex as a result of one of a possible set of actions performed by the agent, and wherein generating the current graph model comprises: processing a first Transformer network input that includes the context data using a Transformer neural network to generate a probability distribution over a possible set of edges that can be included in the current graph model; and selecting, in accordance with the probability distribution over the possible set of edges, a subset of the edges to be included in the current graph model; selecting, from the possible set of actions and using the current graph model, a current action to be performed by the agent in response to the current observation; controlling the agent to perform the selected current action to cause the environment to transition from the current state into a new state; and updating the context data to include (i) data identifying the selected current action and (ii) a new observation characterizing the new state of the environment. This combination would have been motivated by the desire to model dynamics of interacting systems by learning state transition dynamics (Kipf [Sec. 6]). Regarding dependent claim 2, the combination of Zhu, Fang, and Kipf teaches the invention as claimed in claim 1, including wherein selecting the subset of the edges to be included in the current graph model comprises: determining a respective value for each of the possible set of edges using the probability distribution; and selecting, from the possible set of edges, one or more edges having determined values that satisfy a threshold value (Kipf [Sec. 3] discusses producing a per-edge probability score, then the model samples the distribution to obtain an edge decision for each candidate connection; Zhu [Alg. 1] discusses comparing edges of vertices against a threshold, then selecting values that satisfy this comparison). Regarding dependent claim 3, the combination of Zhu, Fang, and Kipf teaches the invention as claimed in claim 1, including wherein selecting the current action to be performed by the agent comprises: generating, by using the Transformer neural network, a reward value for each vertex included in the graph model (Zhu [Sec. 3.3] discusses generating a reward value each iteration for each vertex included; Fang [Sec. 1] discusses the use of a transformer network for action selection; thus, generating a reward value for each vertex included in the graph model, using the Transformer neural network). Regarding dependent claim 4, the combination of Zhu, Fang, and Kipf teaches the invention as claimed in claim 1, including wherein selecting the current action to be performed by the agent comprises: determining an action selection policy that corresponds to the current graph model by using a dynamic programming technique, wherein action selection policy specifies a mapping from the vertices to the edges included in the current graph model; and using the action selection policy to select the current action (Zhu [Sec. 1 & 3.4] discusses an action selection policy of value iteration, which uses dynamic programming, to compute per-vertex values, and then using this to select the action). Regarding dependent claim 5, the combination of Zhu, Fang, and Kipf teaches the invention as claimed in claim 4, including wherein the dynamic programming technique comprises a value iteration technique (Zhu [Sec. 3.4] discusses the dynamic programming technique is a value iteration technique). Regarding dependent claim 6, the combination of Zhu, Fang, and Kipf teaches the invention as claimed in claim 3, including wherein determining the action selection policy that corresponds to the current graph model by using the dynamic programming technique comprises: performing the value iteration technique using the edges included in the current graph model and the reward values included in the current graph model (Zhu [Sec. 3] discusses the value iteration technique is defined over the set of edges included in the graph model and the reward assignment to compute vertex values). Regarding dependent claim 10, the combination of Zhu, Fang, and Kipf teaches the invention as claimed in claim 1, including wherein the current graph models at different time steps include different numbers of edges, or a same number of edges that connect the vertices in different ways (Zhu [Sec. 4.4] discusses as more transitions are incorporated into the graph, the edge count necessarily changes; thus, the graph model at different time steps includes different numbers of edges). Regarding dependent claim 12, the combination of Zhu, Fang, and Kipf teaches the invention as claimed in claim 1, including wherein current graph model represents a Markov decision process (MDP) defining transitions between different states the environment (Zhu [Sec. 3.3] discusses the graph model is a Markov decision process defining transitions between states of the environment). Regarding dependent claim 17, the combination of Zhu, Fang, and Kipf teaches the invention as claimed in claim 1, including wherein the environment is a simulation of a real-world environment (Zhu [Sec. 4 & Conclusion] discusses the environment is a D4RL robotics domain, which is a simulated, offline environment). Regarding claim 20, claim 20 is a system claim that is substantially the same as the method of claim 1. Therefore, claim 20 is rejected for the same reasons as claim 1. Regarding claim 21, claim 21 is a computer-readable storage medium claim that is substantially the same as the method of claim 1. Therefore, claim 21 is rejected for the same reasons as claim 1. Claims 7-8 are rejected under 35 U.S.C. 103 as being unpatentable over Zhu, in view of Fang and Kipf, as applied in claim 1, and further in view of Vinyals et al. ("Pointer Networks", arXiv) (Year: 2017), hereafter Vinyals. Regarding dependent claim 7, the combination of Zhu, Fang, and Kipf teaches the invention as claimed in claim 1, including selecting, processing a second Transformer network input that includes the current observation using the Transformer neural network to generate a decision; and selecting, in accordance with the probability distribution over the vertices included in the graph model, a selected vertex as a vertex in the graph model that corresponds to the current observation (Fang [Sec. 3.2.1-3.2.2] discusses a second network input, used as the query of a distinct attention pass, is built from the current observation; Fang [Sec. 3.2.1-3.2.2] discusses computing a softmax over the memory set; Zhu [Alg. 2] discusses selecting a vertex corresponding to a current observation). The combination of Zhu, Fang, and Kipf does not explicitly teach to generate a probability distribution over the vertices included in the graph model. However, Vinyals teaches a method for generating a softmax distribution wherein a probability distribution over the vertices in the set is generated ([Sec. 2] discusses a softmax computation is a probability distribution over the candidate set; thus, the vertices in the model set). Because the combination of Zhu, Fang, and Kipf teaches processing a second Transformer network input that includes the current observation using the Transformer neural network to generate a decision; and selecting, in accordance with the probability distribution over the vertices included in the graph model, a selected vertex as a vertex in the graph model that corresponds to the current observation; and Vinyals teaches generating a probability distribution over the set of vertices, accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate generating a probability distribution over the set of vertices as taught by Vinyals into the combination of Zhu, Fang, and Kipf’s method, with a reasonable expectation of success, to teach wherein selecting the current action to be performed by the agent comprises: processing a second Transformer network input that includes the current observation using the Transformer neural network to generate a probability distribution over the vertices included in the graph model; and selecting, in accordance with the probability distribution over the vertices included in the graph model, a selected vertex as a vertex in the graph model that corresponds to the current observation. This combination would have been motivated by the desire to implement the method on variable sized inputs comprising different numbers of vertices (Vinyals [Sec. 5]). Regarding dependent claim 8, the combination of Zhu, Fang, Kipf, and Vinyals teaches the invention as claimed in claim 7, including wherein using the action selection policy to select the current action comprises: querying the action selection policy using the selected vertex (Zhu [Sec. 3.4 & Alg. 2] discusses calling an action selection policy using the graph action selected vertex, wherein the chain of action comprises the graph action is converted to the environment action via an action translator). Claims 9 & 11 are rejected under 35 U.S.C. 103 as being unpatentable over Zhu, in view of Fang and Kipf, as applied in claim 1, and further in view of Laskin et al. ("In-context Reinforcement Learning with Algorithm Distillation," CoRR, submitted on October 25, 2022, arXiv:2210.14215v1, 21 pages), hereafter Laskin. Laskin was cited in the IDS submitted 08/01/2024. Regarding dependent claim 9, the combination of Zhu, Fang, and Kipf teaches the invention as claimed in claim 1, including generating a current graph model that represents the environment, wherein the current graph model comprises vertices that represent states of the environment and edges connecting the vertices, wherein an edge between a first vertex and a second vertex in the graph model indicates it is possible that the environment will transition from a state represented by the first vertex into a state represented by the second vertex as a result of one of a possible set of actions performed by the agent (Zhu [Sec. 3] discusses generating a graph wherein each vertex represents an observation of the state of the environment; further, two vertices are connected by an edge if they correspond to a transition between graph states, which constitutes a transition from the first vertex state to the second vertex state); selecting, from the possible set of actions and using the current graph model, a current action to be performed by the agent in response to the current observation (Zhu [Sec. 3.4 & Alg. 2] discusses selecting the best vertex destination in response to the current observation, using the current graph model to search through the set of possible actions). The combination of Zhu, Fang, and Kipf does not explicitly teach wherein the possible set of actions comprise a no-op action. However, in a similar field of endeavor, Laskin teaches a method for reinforcement learning, wherein the possible set of actions comprises a possible no-op action ([Sec. 4] discusses the use of no-op in the possible set of actions). Because the combination of Zhu, Fang, and Kipf teaches a possible set of actions for the agent to perform; and Laskin teaches a possible set of actions comprising a possible no-op action, accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate a possible no-op action within the set of actions as taught by Laskin into the combination of Zhu, Fang, and Kipf’s method, with a reasonable expectation of success, to teach wherein the possible set of actions comprise a no-op action. This combination would have been motivated by the desire to implement the ability to hold the current state, rather than always transitioning to a new one at each time step (Laskin [Sec. 4]). Regarding dependent claim 11, the combination of Zhu, Fang, and Kipf teaches the invention as claimed in claim 1, including processing a first Transformer network input that includes the context data using a Transformer neural network to generate a probability distribution over a possible set of edges that can be included in the current graph model (Fang [Sec. 3.2.2] discusses an attention-based transformer network architecture that processes the accumulated context to produce a decision at each step; Kipf [Sec. 3] discusses computing a probability distribution using context input over a set of edges). The combination of Zhu, Fang, and Kipf does not explicitly teach wherein parameter values of the Transformer neural network are fixed during the multiple time steps. However, in a similar field of endeavor, Laskin teaches a method for reinforcement learning wherein parameters do not need to be updated at each time step ([Sec. 4.3] discusses the RL algorithm is executed in-context without updating the transformer network parameters; thus, the values are fixed during multiple time steps). Because the combination of Zhu, Fang, and Kipf teaches the use of a Transformer neural network for multiple time steps; and Laskin teaches fixed parameter values of the Transformer neural network for multiple time steps, accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate fixed parameter values of the Transformer neural network for multiple time steps as taught by Laskin into the combination of Zhu, Fang, and Kipf’s method, with a reasonable expectation of success, to teach wherein parameter values of the Transformer neural network are fixed during the multiple time steps. This combination would have been motivated by the desire to improve policy in-context, rather than distilling post-learning (Laskin [Sec. 4]). Claims 13 & 14 are rejected under 35 U.S.C. 103 as being unpatentable over Zhu, in view of Fang and Kipf, as applied in claim 1, and further in view of Savinov et al. ("SEMI-PARAMETRIC TOPOLOGICAL MEMORY FOR NAVIGATION", ICLR 2018, arXiv) (Year: 2018), hereafter Savinov. Regarding dependent claim 13, the combination of Zhu, Fang, and Kipf teaches the invention as claimed in claim 1, including maintaining context data that includes, for each of one or more previous time steps that precede the current time step, (i) data identifying a previous action performed by the agent at the previous time step and (ii) a previous observation characterizing a previous state that the environment transitioned into as a result of the agent performing the previous action (Fang [Abstract] discusses storing each observation to memory, which would include the previous action and the state of the environment). The combination of Zhu, Fang, and Kipf does not explicitly teach wherein the agent is a mechanical agent and the environment is a real-world environment. However, in a similar field of endeavor, Savinov teaches a method for an agent navigating an environment, wherein the environment is a real-world environment and the agent is mechanical ([Sec. 2-3] discusses a topological environment, using a physical robot as the agent to navigate the environment). Because the combination of Zhu, Fang, and Kipf teaches the use of an agent and an environment; and Savinov teaches the agent is a mechanical agent and the environment is a real-world environment, accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the agent is a mechanical agent and the environment is a real-world environment as taught by Savinov into the combination of Zhu, Fang, and Kipf’s method, with a reasonable expectation of success, to teach wherein the agent is a mechanical agent and the environment is a real-world environment. This combination would have been motivated by the desire to extend the simulation (Zhu [Sec. 4 & Conclusion]) into a real-world application with robotics (Savinov [Abstract & Sec. 1-2]). Regarding dependent claim 14, the combination of Zhu, Fang, Kipf, and Savinov teaches the invention as claimed in claim 13, including wherein the agent is a robot (Savinov [Sec. 2-3] discusses the agent is a robot). Claim 15 is rejected under 35 U.S.C. 103 as being unpatentable over Zhu, in view of Fang and Kipf, as applied in claim 1, and further in view of Ran et al. ("DeepEE: Joint Optimization of Job Scheduling and Cooling Control for Data Center Energy Efficiency Using Deep Reinforcement Learning", IEEE 39th International Conference on Distributed Computing Systems, IEEE) (Year: 2019), hereafter Ran. Regarding dependent claim 15, the combination of Zhu, Fang, and Kipf teaches the invention as claimed in claim 1, including maintaining context data that includes, for each of one or more previous time steps that precede the current time step, (i) data identifying a previous action performed by the agent at the previous time step and (ii) a previous observation characterizing a previous state that the environment transitioned into as a result of the agent performing the previous action (Fang [Abstract] discusses storing each observation to memory, which would include the previous action by the agent and the state of the environment). The combination of Zhu, Fang, and Kipf does not explicitly teach wherein the environment is a real-world environment of a service facility comprising a plurality of items of electronic equipment and the agent is an electronic agent configured to control operation of the service facility. However, in a similar field of endeavor, Ran teaches a method for using deep reinforcement learning wherein it is used in a service facility ([Abstract & Sec. 1] discusses training a deep-RL electronic agent to control job scheduling and cooling equipment across a real data-center facility, which is a real-world environment of a service facility comprising electronic equipment). Because the combination of Zhu, Fang, and Kipf teaches the use of an agent and an environment; and Ran teaches the environment is a service facility comprising electronic equipment and the agent is an electronic agent controlling operation of the facility, accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the environment as a service facility comprising electronic equipment and the agent as an electronic agent controlling operation of the facility as taught by Ran into the combination of Zhu, Fang, and Kipf’s method, with a reasonable expectation of success, to teach wherein the environment is a real-world environment of a service facility comprising a plurality of items of electronic equipment and the agent is an electronic agent configured to control operation of the service facility. This combination would have been motivated by the desire to extend the method to a real-world application of use in service facilities while saving up to 15% and 10% energy consumption compared to baseline approaches, achieving stable performance gain, and a better tradeoff between energy saving and service quality (Ran [Abstract]). Claim 16 is rejected under 35 U.S.C. 103 as being unpatentable over Zhu, in view of Fang and Kipf, as applied in claim 1, and further in view of Fan et al. ("A Learning Framework for High Precision Industrial Assembly", arXiv) (Year: 2019), hereafter Fan. Regarding dependent claim 16, the combination of Zhu, Fang, and Kipf teaches the invention as claimed in claim 1, including maintaining context data that includes, for each of one or more previous time steps that precede the current time step, (i) data identifying a previous action performed by the agent at the previous time step and (ii) a previous observation characterizing a previous state that the environment transitioned into as a result of the agent performing the previous action (Fang [Abstract] discusses storing each observation to memory, which would include the previous action by the agent and the state of the environment). The combination of Zhu, Fang, and Kipf does not explicitly teach wherein the environment is a real-world manufacturing environment for manufacturing a product and the agent comprises an electronic agent configured to control a manufacturing unit or a machine that operates to manufacture the product. However, in a similar field of endeavor, Fan teaches a method for using deep reinforcement learning wherein it is used in high precision industrial assembly ([Abstract, Sec. 1, & Sec. 3-4] discusses a deep RL-based framework in which an electronic agent is configured to control an industrial assembly, which constitutes a machine that operates to manufacture a product). Because the combination of Zhu, Fang, and Kipf teaches the use of an agent and an environment; and Fan teaches the environment as a real-world manufacturing environment for manufacturing a product and the agent comprises an electronic agent configured to control a manufacturing unit or a machine that operates to manufacture the product, accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the environment as a real-world manufacturing environment for manufacturing a product and the agent comprises an electronic agent configured to control a manufacturing unit or a machine that operates to manufacture the product as taught by Fan into the combination of Zhu, Fang, and Kipf’s method, with a reasonable expectation of success, to teach wherein the environment is a real-world manufacturing environment for manufacturing a product and the agent comprises an electronic agent configured to control a manufacturing unit or a machine that operates to manufacture the product. This combination would have been motivated by the desire to extend the method to a real-world application of use in manufacturing environments while achieving higher efficiency and better stability performance (Fan [Abstract]). Claims 18-19 are rejected under 35 U.S.C. 103 as being unpatentable over Zhu, in view of Fang and Kipf, as applied in claim 1, and further in view of Thomson et al. (US 10810274 B2, published 10/20/2020), hereafter Thomson. Regarding dependent claim 18, the combination of Zhu, Fang, and Kipf teaches the invention as claimed in claim 1, including maintaining context data that includes, for each of one or more previous time steps that precede the current time step, (i) data identifying a previous action performed by the agent at the previous time step (Fang [Abstract] discusses storing each observation to memory, which would include the previous action and the state of the environment). The combination of Zhu, Fang, and Kipf does not explicitly teach wherein the agent is a digital assistant and wherein actions performed by the agent include outputs that are provided by the digital assistant to a user. However, in a similar field of endeavor, Thomson teaches a method for an agent acting as a digital assistant performing actions that are outputted to a user via the digital assistant ([Col. 1-2, Lines 50-2] discusses a digital assistant acting as an agent performing actions within a reinforcement learning space that include outputting the policy actions for presentation to a user). Because the combination of Zhu, Fang, and Kipf teaches the use of an agent and records the actions of the agent; and Thomson teaches the agent is a digital assistant and wherein actions performed by the agent include outputs that are provided by the digital assistant to a user, accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the agent as a digital assistant and wherein actions performed by the agent include outputs that are provided by the digital assistant to a user as taught by Thomson into the combination of Zhu, Fang, and Kipf’s method, with a reasonable expectation of success, to teach wherein the agent is a digital assistant and wherein actions performed by the agent include outputs that are provided by the digital assistant to a user. This combination would have been motivated by the desire to extend the method to a real-world application of use within a user interface space and operate with high accuracy and reliability when performing tasks in response to user requests (Thomson [Col. 1-2, Lines 51-15]). Regarding dependent claim 19, the combination of Zhu, Fang, Kipf, and Thomson teaches the invention as claimed in claim 18, including wherein the outputs include one or more of: text displayed to a user in a user interface of the digital assistant; an image displayed to the user in the user interface of the digital assistant; or speech output through one or more speakers of the digital assistant (Thomson [Col. 11, Lines 10-17] discusses the output is sent to a user interface, displaying on a visual output such as a touch screen, which may include text or graphics and images). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. You et al. ("GraphRNN: Generating Realistic Graphs with Deep Auto-regressive Models", Proceedings of the 35th International Conference on Machine Learning, arXiv) (Year: 2018) ([Abstract] Here we propose GraphRNN, a deep autoregressive model that addresses the above challenges and approximates any distribution of graphs with minimal assumptions about their structure. GraphRNN learns to generate graphs by training on a representative set of graphs and decomposes the graph generation process into a sequence of node and edge formations, conditioned on the graph structure generated so far. In order to quantitatively evaluate the performance of GraphRNN, we introduce a benchmark suite of datasets, baselines and novel evaluation metrics based on Maximum Mean Discrepancy, which measure distances between sets of graphs. Our experiments show that GraphRNN significantly outperforms all baselines, learning to generate diverse graphs that match the structural characteristics of a target set, while also scaling to graphs 50× larger than previous deep models). Any inquiry concerning this communication or earlier communications from the examiner should be directed to RILEY S ACOSTA whose telephone number is (571)272-8714. The examiner can normally be reached Monday-Thursday 6am-4pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jennifer N Welch can be reached at (571)272-7212. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /RILEY S ACOSTA/Examiner, Art Unit 2143 /JENNIFER N WELCH/Supervisory Patent Examiner, Art Unit 2143
Read full office action

Prosecution Timeline

Jan 26, 2024
Application Filed
Sep 03, 2026
Non-Final Rejection mailed — §101, §103, §112 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
100%
Grant Probability
99%
With Interview (+0.0%)
3y 1m (~5m remaining)
Median Time to Grant
Low
PTA Risk
Based on 1 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month