Prosecution Insights
Last updated: October 02, 2026
Application No. 18/834,208

CONTROLLING REINFORCEMENT LEARNING AGENTS USING GEOMETRIC POLICY COMPOSITION

Non-Final OA §101§102§103
Filed
Jul 29, 2024
Priority
Jan 28, 2022 — provisional 63/304,482 +1 more
Examiner
JANSEN II, MICHAEL J
Art Unit
Tech Center
Assignee
DeepMind Technologies Limited
OA Round
1 (Non-Final)
67%
Grant Probability
Favorable
1-2
OA Rounds
2m
Est. Remaining
86%
With Interview

Examiner Intelligence

Grants 67% — above average
67%
Career Allowance Rate
435 granted / 649 resolved
+7.0% vs TC avg
Strong +19% interview lift
Without
With
+19.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 4m
Avg Prosecution
26 currently pending
Career history
693
Total Applications
across all art units

Statute-Specific Performance

§101
1.7%
-38.3% vs TC avg
§103
51.7%
+11.7% vs TC avg
§102
23.5%
-16.5% vs TC avg
§112
19.9%
-20.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 649 resolved cases

Office Action

§101 §102 §103
DETAILED ACTION This is a first office action in response to application 18/834,208 filed 07/29/2024, in which claims 1-22 are presented for examination. A preliminary amendment was filed concurrently therewith providing amendments to claims 3, 6, 8-10, 12-15, 18-22 is acknowledged. Currently claims 1-22 are pending. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-22 are rejected under 35 U.S.C. 101 because the claimed invention is directed to and abstract idea without significantly more. The claim(s) recite(s) a mental process e.g. concepts performed in the human mind (including an observation, evaluation, judgment, opinion) (see MPEP § 2106.04(a)(2), subsection III). This judicial exception is not integrated into a practical application because the generically recited computer elements do not add a meaningful limitation to the abstract idea because they amount to simply implementing the abstract idea on a computer (i.e. a neural network to perform the mental process). The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the claims are generically directed to the observation of the environment in view of a base policy in use and adjusting/changing the base policy to a new policy based upon the observations within the environment. The claims do not appear to expressly provide "significantly more". A human within the environment could perform the aforementioned task and it is noted that the courts consider a mental process (thinking) that "can be performed in the human mind, or by a human using a pen and paper" to be an abstract idea. CyberSource Corp. v. Retail Decisions, Inc., 654 F.3d 1366, 1372, 99 USPQ2d 1690, 1695 (Fed. Cir. 2011). In this case, a human could suggest a policy, monitor or observe environmental conditions, and update/change a policy based on the observation. In regard to Claims 21 and 22 that recite similar language, The Office notes that because both product and process claims may recite a "mental process", the phrase "mental processes" should be understood as referring to the type of abstract idea, and not to the statutory category of the claim. The courts have identified numerous product claims as reciting mental process-type abstract ideas, for instance the product claims to computer systems and computer-readable media in Versata Dev. Group. v. SAP Am., Inc., 793 F.3d 1306, 115 USPQ2d 1681 (Fed. Cir. 2015). See MPEP § 2106.05. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claim(s) 1-4, 6-12, 14-22 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Wierstra et al. U.S. Patent Application Publication No. 2020/0090006 A1 hereinafter Wierstra Consider Claim 1 and similarly recited CRM and System Claims 21-22: Wierstra discloses a computer-implemented method for controlling a reinforcement learning agent in an environment, the method comprising: (Wierstra, [0054-0063], See Abstract.) maintaining data specifying a base policy set comprising a plurality of base policies for controlling the agent; receiving a current observation characterizing a current state of the environment; (Wierstra, [0009], “In one aspect of the present disclosure a neural network system for model-based reinforcement learning is used to select actions to be performed by an agent interacting with an environment, to perform a task in an attempt to achieve a specified result. The system may comprise at least one imagination core which has an input to receive a current observation characterizing a current state of the environment, and optionally historical observations, and which includes a model of the environment.”) generating, for each of one or more of the plurality of base policies, one or more predicted future observations characterizing respective future states of the environment that are subsequent to the current state of the environment by using an environment dynamics neural network that corresponds to the base policy, (Wierstra, [0009], “The imagination core may be configured to output trajectory data in response to the current observation, and/or historical observations, the trajectory data defining a trajectory comprising a sequence of future features of the environment imagined by the imagination core (that is, predicted on the assumption that the agent performs certain actions). The system may also include at least one rollout encoder to encode the sequence of features from the imagination core to provide a rollout embedding for the trajectory. The system may further comprise a reinforcement learning output stage to receive data derived from the rollout embedding and to output action policy data for defining an action policy identifying an action based on the current observation.”) wherein each environment dynamics neural network is configured to receive an environment dynamics network input comprising an input observation characterizing an input state of the environment and a respective action selected by using the corresponding base policy to be performed by the agent in response to the input observation, and to process the environment dynamics network input to generate a predicted future observation characterizing a respective future state of the environment; (Wierstra, [0012], “In some implementations the imagination core comprises a neural environment model coupled to a policy module. The neural environment model receives the current observation and/or a history of observations, and also a current action, and predicts a subsequent observation in response. It may also predict a reward from taking the action. The policy module defines a policy to rollout a sequence of actions and states using the environment model, defining a trajectory. The trajectory may be defined by one or more of predicted observations, predicted actions, predicted rewards, and a predicted sequence termination signal. The neural environment model may predict a subsequent observation in response to a current observation and a history of observations, conditioned on action data from the policy module.”) using the predicted future observations generated for the plurality of base policies to determine a respective estimated value for each composite policy in a composite policy set with respect to the current state of the environment, wherein each composite policy is generated based on the base policy set and, (Wierstra, [0014], “In implementations the neural environment model is a learned model of the environment. Thus a method of training the system may involve pre-training one or more such models for the imagination cores and then training other adaptive components of the system using reinforcement learning. The learned models may be imperfect models of the environment and may be trained on the same or a different environment to the environment in which the RL system operates.”) in each composite policy, each of one or more of the plurality of base policies are subsequently used to select actions to be performed by the agent in response to a corresponding number of consecutive future observations; and (Wierstra, [0040], “To generate a trajectory, a current (actual) observation o.sub.t is input to the unit 2, and input into the IC 1, to generate Ô.sub.t+1, and {circumflex over (r)}.sub.t+1. The prediction ô.sub.t+1 is then input into the IC 1 again, to generate Ô.sub.t+2, and {circumflex over (r)}.sub.t+2. This process is carried out in total r times, to produce a rollout trajectory of τ rollout time-steps. Thus, FIG. 2 shows how the environment model 12 of the IC 1 is used to obtain predictions for multiple time steps into the future, by initializing the rollout with the current actual observation, and subsequently feeding simulated observations into the IC 1, to iteratively create a rollout trajectory custom-character The trajectory custom-character is a sequence of features (f.sub.t+1, . . . , f.sub.t+τ), where for any integer i, f.sub.t+i, denotes the output of the environment model 12 at the i-th step. That is f.sub.t+i comprises ô.sub.t+i, and, in the case that the environment model 12 outputs also reward values, comprises both Ô.sub.t+i and {circumflex over (r)}.sub.t+i.”) selecting, as a current action to be performed by the agent in response to the current observation characterizing the current state of the environment, an action using the respective estimated values for the composite policies. (Wierstra, [0046], “The neural network also includes a policy module (reinforcement learning output stage) 33, which may be another neural network. The policy module 33 receives the imagination code c.sub.ia and the output c.sub.mf of the model-free network 32. The policy model 33 outputs a policy vector π(“action policy”) and an estimated value V. The policy vector π may be a vector characterizing the parameters of a network which uses the current (actual) observation o.sub.t to generate an action a.sub.t and V is value baseline data for the current observation, for determining an advantage of an action defined by the action policy. The neural network system shown in FIG. 3 can be considered as augmenting a model free agent (an agent controlled by a model-free network such as the model-free network 32) by providing additional information from model-based planning. Thus, the neural network system can be considered as an agent with strictly more expressive power than the underlying model-free agent 32.”) Consider Claim 2: Wierstra discloses the method of claim 1, wherein selecting action using the respective estimated values for the composite policies comprises: selecting an action according to the composite policy that has a highest estimated value. (Wierstra, [0035-0040], [0048], [0050], “As mentioned above, the policy module 11 may also be trained in step 52, rather than being pre-set. It was found that it was valuable to train it by “distilling” information from the output of the policy module 33. Specifically, in step 52 a (small) model-free network (which can be denoted as performing the function {circumflex over (π)}=(o.sub.t)) was generated, and then the cost function used in step 52 is augmented by including a cross entropy auxiliary loss between the imagination-augmented policy π=(o.sub.t) produced by the policy module 33 for the current observation o.sub.t, and the policy {circumflex over (π)}=(o.sub.t) for the same observation. The existence of this term means that the policy module 11 will be trained such that the IC 1 tends to produce a rollout trajectory which is similar to the actual trajectories of the agent (i.e. the trajectory of the agent controlled by the neural network system of FIG. 3) in the real environment. It also tends to ensure that the rollout corresponds to trajectories with a relatively high reward. At the same time, the imperfect approximation between the policies results in the rollout policy having a higher entropy, thus striking a balance between exploration and exploitation.”) Consider Claim 3: Wierstra discloses the method of claim 1, wherein in each composite policy, the corresponding number of consecutive future observations in response to which each base policy is used to select actions is determined in accordance with a given switching probability. (Wierstra, [0048], “The two inputs to the environment model are combined (e.g. concatenated) to form structured content. The structured content is input to a convolutional network 43, e.g. comprising one or more convolutional layers. The output of the convolutional network 43 is the predicted observation ô.sub.t+1. This is denoted as a two-dimensional array 44, since it may take the form of a pixel-wise probability distribution for an image. A second (optional) output of the convolution network 43 is a predicted reward {circumflex over (r)}.sub.t+1 which is output as a data structure 45.”) Consider Claim 4: Wierstra discloses the method of claim 3, wherein the given switching probability is a value received from a user that is between zero and one. (Wierstra, [0041], “As explained below, the environment model 12 is formed by training, so it cannot be assumed to be perfect. It might sometimes make erroneous or even nonsensical predictions. For that reason, it is preferred not to rely entirely on the output of the environment model 12. For that reason, the prediction-and-encoding unit includes a rollout encoder 21. The rollout encoder 21 is an adaptive system (e.g. in the form of set of sequential state generation neural networks) which is trained to receive the trajectory, and “encode” it into one or more values referred to as a “rollout embedding”. That is, the encoding interprets the trajectory, i.e. extracts any information useful for the agent's decision. The encoding may include ignoring the trajectory when necessary, e.g. because the rollout encoder 21 is trained to generate an output which has a low (or zero) dependence on the trajectory if the input to the rollout encoder 21 is a trajectory which is likely to be erroneous of nonsensical. The rollout embedding produced by the rollout encoder 21 can be denoted e=ϵ (custom-character), and can be considered a summary of the trajectory produced by the rollout encoder 21.”) Consider Claim 6: Wierstra discloses the method of claim 1, further comprising maintaining a reward estimator associated with the environment, wherein the reward estimator is configured to receive a reward estimator input comprising the current observation characterizing the current state of the environment, and to process the reward estimator input to generate an estimated value of a current reward received by the agent at the current state of the environment. (Wierstra, [0012], “In some implementations the imagination core comprises a neural environment model coupled to a policy module. The neural environment model receives the current observation and/or a history of observations, and also a current action, and predicts a subsequent observation in response. It may also predict a reward from taking the action. The policy module defines a policy to rollout a sequence of actions and states using the environment model, defining a trajectory. The trajectory may be defined by one or more of predicted observations, predicted actions, predicted rewards, and a predicted sequence termination signal. The neural environment model may predict a subsequent observation in response to a current observation and a history of observations, conditioned on action data from the policy module.”) Consider Claim 7: Wierstra discloses the method of claim 6, wherein the current reward is dependent on a previous action performed by the agent in response to a previous observation characterizing a previous state of the environment. (Wierstra, [0012], “In some implementations the imagination core comprises a neural environment model coupled to a policy module. The neural environment model receives the current observation and/or a history of observations, and also a current action, and predicts a subsequent observation in response. It may also predict a reward from taking the action. The policy module defines a policy to rollout a sequence of actions and states using the environment model, defining a trajectory. The trajectory may be defined by one or more of predicted observations, predicted actions, predicted rewards, and a predicted sequence termination signal. The neural environment model may predict a subsequent observation in response to a current observation and a history of observations, conditioned on action data from the policy module.”) Consider Claim 8: Wierstra discloses the method of claim 6, wherein the reward estimator is configured as a deterministic reward estimator. (Wierstra, [0048], “The two inputs to the environment model are combined (e.g. concatenated) to form structured content. The structured content is input to a convolutional network 43, e.g. comprising one or more convolutional layers. The output of the convolutional network 43 is the predicted observation ô.sub.t+1. This is denoted as a two-dimensional array 44, since it may take the form of a pixel-wise probability distribution for an image. A second (optional) output of the convolution network 43 is a predicted reward {circumflex over (r)}.sub.t+1 which is output as a data structure 45.”) Consider Claim 9: Wierstra discloses the method of claim 6, wherein the reward estimator is configured as a machine learning model trained on training data generated as a result of the agent interacting with the environment. (Wierstra, [0039], [0041], [0049-0053], [0049] “We now turn to a description of the training procedure for the neural network, as illustrated in FIG. 5. It includes a first step 51 of training the environment model 12. In a second step 52, the other adaptive components of the network are trained concurrently using a cost function.”) Consider Claim 10: Wierstra discloses the method of claim 6, wherein using the predicted future observations to determine the respective estimated value for each composite policy with respect to the current state of the environment comprises: using the reward estimator to generate a respective estimated value of a future reward for each predicted future observation; and determining, in accordance with a given discount factor, and from the respective estimated values of the future rewards, an estimation of a sum of future rewards received by the agent if the agent were to perform actions selected by using the composite policy beginning from the current state of the environment; and determining a value for the composite policy from the estimation of the sum of the future rewards. (Wierstra, [0012], [0017], [0048], [0040], “To generate a trajectory, a current (actual) observation o.sub.t is input to the unit 2, and input into the IC 1, to generate Ô.sub.t+1, and {circumflex over (r)}.sub.t+1. The prediction ô.sub.t+1 is then input into the IC 1 again, to generate Ô.sub.t+2, and {circumflex over (r)}.sub.t+2. This process is carried out in total r times, to produce a rollout trajectory of τ rollout time-steps. Thus, FIG. 2 shows how the environment model 12 of the IC 1 is used to obtain predictions for multiple time steps into the future, by initializing the rollout with the current actual observation, and subsequently feeding simulated observations into the IC 1, to iteratively create a rollout trajectory custom-character The trajectory custom-character is a sequence of features (f.sub.t+1, . . . , f.sub.t+τ), where for any integer i, f.sub.t+i, denotes the output of the environment model 12 at the i-th step. That is f.sub.t+i comprises ô.sub.t+i, and, in the case that the environment model 12 outputs also reward values, comprises both Ô.sub.t+i and {circumflex over (r)}.sub.t+i.”) Consider Claim 11: Wierstra discloses the method of claim 10, wherein determining the estimation of the sum of future rewards comprises receiving, as the given discount factor, a user-defined value that is between zero and one. (Wierstra, [0047], “FIG. 4 represents a possible structure of the environment model 12 of FIG. 1. A first input to the environment model is the current (actual or imagined) observation o.sub.t or ô.sub.t. In FIG. 4 this is illustrated as a rectangle 41, since it may, for certain embodiments, take the form of a respective value for each of a set of points (pixels) in a two-dimensional array. In another case, the current observation may include multiple values for each pixel, so that it may correspond to multiple two-dimensional arrays. A second input to the environmental model 12 is an action a.sub.t. This may be provided in the form of a vector 42, which has a number of components equal to the number of possible actions. The component corresponding to the action a.sub.t takes a predefined value (e.g. 1), and the other components take other value(s) (e.g. zero).”) Consider Claim 12: Wierstra discloses the method of claim 1, wherein each environment dynamics neural network is trained based on optimizing a cross-entropy temporal- difference (CETD) loss. (Wierstra, [0050], “As mentioned above, the policy module 11 may also be trained in step 52, rather than being pre-set. It was found that it was valuable to train it by “distilling” information from the output of the policy module 33. Specifically, in step 52 a (small) model-free network (which can be denoted as performing the function {circumflex over (π)}=(o.sub.t)) was generated, and then the cost function used in step 52 is augmented by including a cross entropy auxiliary loss between the imagination-augmented policy π=(o.sub.t) produced by the policy module 33 for the current observation o.sub.t, and the policy {circumflex over (π)}=(o.sub.t) for the same observation. The existence of this term means that the policy module 11 will be trained such that the IC 1 tends to produce a rollout trajectory which is similar to the actual trajectories of the agent (i.e. the trajectory of the agent controlled by the neural network system of FIG. 3) in the real environment. It also tends to ensure that the rollout corresponds to trajectories with a relatively high reward. At the same time, the imperfect approximation between the policies results in the rollout policy having a higher entropy, thus striking a balance between exploration and exploitation.”) Consider Claim 14: Wierstra discloses the method of claim 1, wherein maintaining data specifying the base policy set comprising the plurality of base policies for controlling the agent comprises: maintaining a respective base policy neural network that corresponds to each of one or more of the plurality of base policies, wherein each base policy neural network is configured to receive a base policy network input comprising the current observation characterizing the current state of the environment, and to process the base policy network input to generate a base policy network output that specifies an action to be performed by the agent in response to the current observation. (Wierstra, [0012], “In some implementations the imagination core comprises a neural environment model coupled to a policy module. The neural environment model receives the current observation and/or a history of observations, and also a current action, and predicts a subsequent observation in response. It may also predict a reward from taking the action. The policy module defines a policy to rollout a sequence of actions and states using the environment model, defining a trajectory. The trajectory may be defined by one or more of predicted observations, predicted actions, predicted rewards, and a predicted sequence termination signal. The neural environment model may predict a subsequent observation in response to a current observation and a history of observations, conditioned on action data from the policy module.”) Consider Claim 15: Wierstra discloses the method of claim 1, wherein the agent is a mechanical agent or other hardware, the environment is a real-world environment, and the observation comprises data from one or more sensors configured to sense the real-world environment. (Wierstra, [0008], [0009], “The imagination core may be configured to output trajectory data in response to the current observation, and/or historical observations, the trajectory data defining a trajectory comprising a sequence of future features of the environment imagined by the imagination core (that is, predicted on the assumption that the agent performs certain actions). The system may also include at least one rollout encoder to encode the sequence of features from the imagination core to provide a rollout embedding for the trajectory. The system may further comprise a reinforcement learning output stage to receive data derived from the rollout embedding and to output action policy data for defining an action policy identifying an action based on the current observation.”) Consider Claim 16: Wierstra discloses the method of claim 15, wherein the mechanical agent comprises a robot or a vehicle. (Wierstra, [0008], [0009], “The imagination core may be configured to output trajectory data in response to the current observation, and/or historical observations, the trajectory data defining a trajectory comprising a sequence of future features of the environment imagined by the imagination core (that is, predicted on the assumption that the agent performs certain actions). The system may also include at least one rollout encoder to encode the sequence of features from the imagination core to provide a rollout embedding for the trajectory. The system may further comprise a reinforcement learning output stage to receive data derived from the rollout embedding and to output action policy data for defining an action policy identifying an action based on the current observation.”) Consider Claim 17: Wierstra discloses the method of claim 15, wherein the other hardware comprises a heater, a cooler, a humidifier, or other hardware that modifies a property of air in the real-world environment. (Wierstra, [0008], [0009], “The imagination core may be configured to output trajectory data in response to the current observation, and/or historical observations, the trajectory data defining a trajectory comprising a sequence of future features of the environment imagined by the imagination core (that is, predicted on the assumption that the agent performs certain actions). The system may also include at least one rollout encoder to encode the sequence of features from the imagination core to provide a rollout embedding for the trajectory. The system may further comprise a reinforcement learning output stage to receive data derived from the rollout embedding and to output action policy data for defining an action policy identifying an action based on the current observation.”) Consider Claim 18: Wierstra discloses the method of claim 1, wherein the agent is a computer-implemented task manager, the environment is a physical state of a computational device, and the observation comprises data from one or more sensors configured to sense the state of the computational device. (Wierstra, [0008], [0009], “The imagination core may be configured to output trajectory data in response to the current observation, and/or historical observations, the trajectory data defining a trajectory comprising a sequence of future features of the environment imagined by the imagination core (that is, predicted on the assumption that the agent performs certain actions). The system may also include at least one rollout encoder to encode the sequence of features from the imagination core to provide a rollout embedding for the trajectory. The system may further comprise a reinforcement learning output stage to receive data derived from the rollout embedding and to output action policy data for defining an action policy identifying an action based on the current observation.”) Consider Claim 19: Wierstra discloses the method of claim 1, wherein the agent is a software program implemented on one or more computers, the environment is a real-world environment, and the observation comprises data from one or more sensors configured to sense the real-world environment. (Wierstra, [0008], [0009], “The imagination core may be configured to output trajectory data in response to the current observation, and/or historical observations, the trajectory data defining a trajectory comprising a sequence of future features of the environment imagined by the imagination core (that is, predicted on the assumption that the agent performs certain actions). The system may also include at least one rollout encoder to encode the sequence of features from the imagination core to provide a rollout embedding for the trajectory. The system may further comprise a reinforcement learning output stage to receive data derived from the rollout embedding and to output action policy data for defining an action policy identifying an action based on the current observation.”) Consider Claim 20: Wierstra discloses the method of claim 1, further comprising causing the agent to perform the action selected by using the respective estimated values for the composite policies. (Wierstra, [0008], [0009], “The imagination core may be configured to output trajectory data in response to the current observation, and/or historical observations, the trajectory data defining a trajectory comprising a sequence of future features of the environment imagined by the imagination core (that is, predicted on the assumption that the agent performs certain actions). The system may also include at least one rollout encoder to encode the sequence of features from the imagination core to provide a rollout embedding for the trajectory. The system may further comprise a reinforcement learning output stage to receive data derived from the rollout embedding and to output action policy data for defining an action policy identifying an action based on the current observation.”) Claim Rejections - 35 USC § 103 Claim(s) 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wierstra et al. U.S. Patent Application Publication No. 2020/0090006 A1 as applied to claims 1 above, and further in view of Blundell et al. U.S. Patent Application Publication No. 2019/0205757 A1 hereinafter Blundell. Consider Claim 13: Wierstra discloses the use of an encoder as taught method of claim 1, however does not specify wherein each environment dynamics neural network is configured as a respective conditional β-VAE model. Blundell however teaches it was a known technique to provide wherein each environment dynamics neural network is configured as a respective conditional β-VAE model. (Blundell, [0044], “In some other implementations, the system determines the feature representation of the current observation by processing the current observation using a variational auto-encoder (VAE) model to generate a latent representation of the current observation, and using the latent representation of the current observation as the feature representation of the current observation. Generally, a VAE model includes two neural networks: an encoder neural network that receives observations and maps them into corresponding representations of the observations, and a decoder neural network that receives representations and approximately recovers the observations corresponding to the representations. For example, an encoder neural network may include convolutional neural network layers followed by a fully connected neural network layer from which a linear neural network layer outputs representations of the received observations. A decoder neural network generally mirrors the structure of its corresponding encoder neural network. For example, the decoder neural network corresponding to the above-described encoder neural network includes a fully connected neural network layer followed by reverse convolutional neural network layers.”) It therefore would have been obvious to those having ordinary skill in the art before the effective filing date of the invention to use the variable model as this was a known technique as taught by Blundell and would have been recognized by a person of skill in the art to be used for the purpose of generating a latent representation of the current observation. (Blundell, [0044]) Claim Rejections - 35 USC § 103 Claim(s) 5 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wierstra et al. U.S. Patent Application Publication No. 2020/0090006 A1 as applied to claims 1 above, and further in view of Isozaki U.S. Patent Application Publication No. 2012/0303572 A1 hereinafter Isozaki. Consider Claim 5: Wierstra discloses the method of claim 3, wherein in each composite policy, the corresponding number of consecutive future observations in response to which each base policy is used to select actions however does not appear to suggest is determined by evaluating a geometric probability distribution function over the given switching probability. Isozaki however teaches that it was a known technique to those in art for evaluating a geometric probability distribution function over the given switching probability. (Isozaki, [0061-0065], [0062], “When the temperature determination unit 42 calculates the fluctuation of the data, n data items are used as in Equation (3) above. Therefore, the probability function which is calculated based on the n data items and has the highest likelihood, the Bayesian posterior probability function, or the empirical distribution function and a probability calculated based on (n-1) data items may be calculated by the geometric mean of a probability function calculated likewise based on up to j (one in the range of 0.ltoreq.j.ltoreq.n-1) data items. At this time, when j=0, a uniform distribution function can be used. Thus, a deviation from the mean obtained based on the previous data can be specifically calculated as a fluctuation.”) It therefore would have been obvious to those having ordinary skill in the art before the effective filing date of the invention to use geometric probability distribution as this was a known technique in view of Isozaki and would have been recognized in the art to be used for the purpose of a deviation from the mean obtained based on the previous data can be specifically calculated as a fluctuation. Thus, when the recursive calculation is used, the amount of calculation may increase, but the accuracy can be improved. (Isozaki, [0065]) Conclusion Prior art made of record and not relied upon which is still considered pertinent to applicant's disclosure is cited in a current or previous PTO-892. The prior art cited in a current or previous PTO-892 reads upon the applicants claims in part, in whole and/or gives a general reference to the knowledge and skill of persons having ordinary skill in the art before the effective filing date of the invention. Applicant, when responding to this Office action, should consider not only the cited references applied in the rejection but also any additional references made of record. In the response to this office action, the Examiner respectfully requests support be shown for any new or amended claims. More precisely, indicate support for any newly added language or amendments by specifying page, line numbers, and/or figure(s). This will assist The Office in compact prosecution of this application. The Office has cited particular columns, paragraphs, and/or line numbers in the applied rejection of the claims above for the convenience of the applicant. Citations are representative of the teachings in the art and are applied to the specific limitations within each claim, however other passages and figures may apply. Applicant, in preparing a response, should fully consider the cited reference(s) in its entirety and not only the cited portions as other sections of the reference may expand on the teachings of the cited portion(s). Applicant Representatives are reminded of CFR 1.4(d)(2)(ii) which states “A patent practitioner (§ 1.32(a)(1) ), signing pursuant to §§ 1.33(b)(1) or 1.33(b)(2), must supply his/her registration number either as part of the S-signature, or immediately below or adjacent to the S-signature. The number (#) character may be used only as part of the S-signature when appearing before a practitioner’s registration number; otherwise the number character may not be used in an S-signature.” When an unsigned or improperly signed amendment is received the amendment will be listed in the contents of the application file, but not entered. The examiner will notify applicant of the status of the application, advising him or her to furnish a duplicate amendment properly signed or to ratify the amendment already filed. In an application not under final rejection, applicant should be given a two month time period in which to ratify the previously filed amendment (37 CFR 1.135(c) ). Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHAEL J JANSEN II whose telephone number is (571)272-5604. The examiner can normally be reached Normally Available Monday-Friday 9am-4pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Temesghen Ghebretinsae can be reached on 571-272-3017. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Michael J Jansen II/ Primary Examiner, Art Unit 2626
Read full office action

Prosecution Timeline

Jul 29, 2024
Application Filed
Sep 03, 2026
Non-Final Rejection mailed — §101, §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12744000
DISPLAY DEVICE WITH SELF-ADJUSTING POWER SUPPLY
2y 8m to grant Granted Sep 22, 2026
Patent 12743171
DRIVE CIRCUIT, TOUCH DRIVE APPARATUS, AND ELECTRONIC DEVICE
1y 5m to grant Granted Sep 22, 2026
Patent 12734455
INTERACTIVE OBJECT SYSTEMS AND METHODS
2y 1m to grant Granted Sep 15, 2026
Patent 12738238
DATA DRIVER AND DISPLAY DEVICE COMPRISING SAME
1y 9m to grant Granted Sep 15, 2026
Patent 12706025
GATE DRIVER AND DISPLAY DEVICE INCLUDING THE SAME
1y 8m to grant Granted Aug 11, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
67%
Grant Probability
86%
With Interview (+19.0%)
2y 4m (~2m remaining)
Median Time to Grant
Low
PTA Risk
Based on 649 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month