Prosecution Insights
Last updated: October 01, 2026
Application No. 18/698,218

DEMONSTRATION-DRIVEN REINFORCEMENT LEARNING

Non-Final OA §103§112
Filed
Apr 03, 2024
Priority
Oct 05, 2021 — provisional 63/252,605 +1 more
Examiner
MRABI, HASSAN
Art Unit
Tech Center
Assignee
DeepMind Technologies Limited
OA Round
1 (Non-Final)
78%
Grant Probability
Favorable
1-2
OA Rounds
3m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 78% — above average
78%
Career Allowance Rate
297 granted / 381 resolved
+18.0% vs TC avg
Strong +33% interview lift
Without
With
+33.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 9m
Avg Prosecution
19 currently pending
Career history
399
Total Applications
across all art units

Statute-Specific Performance

§101
13.9%
-26.1% vs TC avg
§103
60.2%
+20.2% vs TC avg
§102
9.5%
-30.5% vs TC avg
§112
6.0%
-34.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 381 resolved cases

Office Action

§103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION This Office Action is sent in response to Application’s Communication received on 04/03/2024 for application number 18/698218. The Office hereby acknowledges receipt of the following and placed of record in file: Specification, Drawing, Abstract, Oath/Declaration, and Claims. Claims (1-16), 19 and 20 are presented for examination. Information Disclosure Statement The information disclosure statements (IDS) submitted on 10/01/2024 were filed prior to current Office Action. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Claim Objections Claim 9 is objected to because of the following informalities: the claim recites “The method of any claim 8”, it should be replaced with “the method of claim 8”. Appropriate correction is required. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claim 13 recites the limitation "the distance" and “the threshold distance” in “generating a sparse reward that is equal to one for a respective training observation for the time step for which the distance between the encoded representation of the respective training observation and the encoded representation of the new goal observation is below the threshold distance”. There is insufficient antecedent basis for this limitation in the claim. Claim 14 recites the limitation "the sparse rewards” in “generating the new training sequence that includes (i) the one or more new goal observations determined from the demonstration data, (ii) the sparse rewards, and (iii) the plurality of training observations”. There is insufficient antecedent basis for this limitation in the claim. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 3, 5-11, 13-16, 19-22 are rejected under AIA 35 U.S.C. 103(a) as being unpatentable over Wayne et al. US Patent Application Publication US 20200090042 A1 (hereinafter Wayne) in view of Marcin Andrychowicz et al. 2017, NPL Publication: Hindsight Experience Replay, hereinafter (Andrychowicz) and further in view of Ashvin Nair et al. NPL Publication 2018, Visual Reinforcement Learning with Imagined Goals (hereinafter Nair) and further in view of Yiming Ding et a, NPL Publication 2019, Goal-conditioned Imitation Learning (hereinafter Ding). Regarding claim 1, Wayne teaches A method of training a reinforcement learning system to select actions to be performed by an agent interacting with an environment to perform a particular task, the method comprising ([0010-0011], [0072] wherein Wayne performs actions that are selected by the reinforcement learning, wherein a robot is trained to perform tasks in a simulated environment and the training transferred to a system controlling a real robot) process the policy input in accordance with the policy parameters to generate a policy output that defines an action to be performed by the agent in response to the current observation (Claims 8 and 18 text, [0031], [0060], [0099], [0123], [0126] wherein Wayne teaches a policy generator that may be used to select actions to be performed by an agent interacting with an environment to imitate a state-action trajectory, using the discriminator to discriminate between the imitated state-action trajectory and a reference trajectory, and updating parameters of the policy generator using the reward values conditioned on the target embedding vector) and training the goal-conditioned policy neural network on the new training sequence through reinforcement learning ([0084-0085] wherein Wayne trains the system based on embeddings (latent variables) determined via an encoder, the resulting system is better able to imitate the behavior of the set of trajectories in a robust manner over a wider range of behaviors. As a wider range of behaviors are modelled by the neural network, a smaller number of training trajectories are required to train the neural network, therefore providing a more efficient training method. Furthermore, this method allows for one-shot learning). Wayne does not teach obtaining a training sequence comprising a respective training observation at each of a plurality of time steps, wherein each training observation is received as a result of the agent interacting with the environment controlled using a goal-conditioned policy neural network that has a plurality of policy parameters, wherein the goal-conditioned policy neural network is configured to, at each of the plurality of time steps; generating a new training sequence from the training sequence and the demonstration data, comprising: using the demonstration data to determine one or more new goal observations each characterizing a respective goal state of the environment. However in analogous art of demonstration-driven reinforcement learning, Andrychowicz teaches obtaining a training sequence comprising a respective training observation at each of a plurality of time steps, wherein each training observation is received as a result of the agent interacting with the environment controlled using a goal-conditioned policy neural network that has a plurality of policy parameters, wherein the goal-conditioned policy neural network is configured to, at each of the plurality of time steps (Abstract, Section 1, 2.1, 3, 3.2, 4.2- 4.5 wherein Andrychowicz describes learning formalism consisting of an agent interacting with an environment. Andrychowicz incorporates a technique called Hindsight Experience Replay (HER) which allows the algorithm to perform exactly this kind of reasoning and can be combined with any off-policy RL algorithm. It is applicable whenever there are multiple goals which can be achieved, e.g. achieving each state of the system may be treated as a separate goal. Not only does HER improve the sample efficiency in this setting, but more importantly, it makes learning possible even if the reward signal is sparse and binary. The approach is based on training universal policies which take as input not only the current state, but also a goal state. The pivotal idea behind HER is to replay each episode with a different goal than the one the agent was trying to achieve, e.g. one of the goals which was achieved in the episode) generating a new training sequence from the training sequence and the demonstration data, comprising: using the demonstration data to determine one or more new goal observations each characterizing a respective goal state of the environment (Section 1, 2.1-2.2, 3.1-3.4, 4.4-4.5 wherein Andrychowicz create training and sequence of actions for data demonstration for multiple goals) It would have been obvious to a person in the ordinary skill in the art before the effective filing date of the claimed invention to combine Wayne with Andrychowicz by incorporating the method of obtaining a training sequence comprising a respective training observation at each of a plurality of time steps, wherein each training observation is received as a result of the agent interacting with the environment controlled using a goal-conditioned policy neural network that has a plurality of policy parameters, wherein the goal-conditioned policy neural network is configured to, at each of the plurality of time steps; generating a new training sequence from the training sequence and the demonstration data, comprising: using the demonstration data to determine one or more new goal observations each characterizing a respective goal state of the environment of Andrychowicz into the method of training a reinforcement learning system to select actions to be performed by an agent interacting with an environment to perform a particular task of Wayne for the purpose of incorporating techniques of Hindsight Experience Replay (HER) to improve performance and achieve the goals of interest (Andrychowicz, section 4.3). Wayne does not teach receive a policy input comprising an encoded representation of a current observation characterizing a current state of the environment at the time step and an encoded representation of a goal observation characterizing a goal state of the environment; obtaining demonstration data comprising one or more demonstration sequences, each demonstration sequence comprising a plurality of demonstration observations characterizing. However in analogous art of demonstration-driven reinforcement learning, Nair teaches receive a policy input comprising an encoded representation of a current observation characterizing a current state of the environment at the time step and an encoded representation of a goal observation characterizing a goal state of the environment (Sections 3, 4–4.3 wherein Nair encodes observations and goals, computes latent-space rewards, and relabel transitions using sampled latent goals) obtaining demonstration data comprising one or more demonstration sequences, each demonstration sequence comprising a plurality of demonstration observations characterizing (Sections III-D and IV wherein Nair describes separate demonstration replay, actor–critic training and behavior cloning. Section IV-D uses demonstration goals when starting rollouts). It would have been obvious to a person in the ordinary skill in the art before the effective filing date of the claimed invention to combine Wayne with Nair by incorporating the method of receive a policy input comprising an encoded representation of a current observation characterizing a current state of the environment at the time step and an encoded representation of a goal observation characterizing a goal state of the environment; obtaining demonstration data comprising one or more demonstration sequences, each demonstration sequence comprising a plurality of demonstration observations characterizing of Nair into the method of training a reinforcement learning system to select actions to be performed by an agent interacting with an environment to perform a particular task of Wayne for the purpose embedding the state and goals into a latent space using an encoder to obtain a latent state (Nair, section 4.1). Wayne does not teach generating the new training sequence that includes the respective training observations but indicates that the goal-conditioned policy neural network was conditioned on respective encoded representations of each of the new goal observations at one or more time steps in the training sequence. However in analogous art of demonstration-driven reinforcement learning, Ding teaches generating the new training sequence that includes the respective training observations but indicates that the goal-conditioned policy neural network was conditioned on respective encoded representations of each of the new goal observations at one or more time steps in the training sequence (Sections 1-2, 4, 4.1, 4.3, 5.2, 5.5 wherein Ding maximizes the likelihood of the expert actions under the training agent policy, to Inverse Reinforcement Learning that extracts a reward function from those demonstrations and then trains a policy to maximize it and wherein Ding incorporate demonstrations into Hindsight Experience Replay for training goal-conditioned policies). It would have been obvious to a person in the ordinary skill in the art before the effective filing date of the claimed invention to combine Wayne with Ding by incorporating the method of generating the new training sequence that includes the respective training observations but indicates that the goal-conditioned policy neural network was conditioned on respective encoded representations of each of the new goal observations at one or more time steps in the training sequence of Ding into the method of training a reinforcement learning system to select actions to be performed by an agent interacting with an environment to perform a particular task of Wayne for the purpose incorporating a method that can also be used when the available expert trajectories do not contain the actions or when the expert is suboptimal, which makes it applicable when only kinesthetic, third-person or noisy demonstrations are available. Our code is open-source (Ding, Abstract). Regarding claim 3, Wayne as modified by Andrychowicz, Nair and Ding teach wherein encoded representations of observations of the environment in the policy input are generated by processing the observations using an encoder neural network (Sections 4, wherein Andrychowicz describes algorithm 1 with input that includes processing observations), (Sections III-B–D, wherein Nair 2 describes rendering the observations). Regarding claim 5, Wayne as modified by Andrychowicz, Nair and Ding teach training the goal-conditioned policy neural network on the obtained training sequence through reinforcement learning (Sections III-D and IV wherein Nair describes separate demonstration replay, actor–critic training and behavior cloning. Section IV-D uses demonstration goals when starting rollouts), (Sections 1-2, 4, 4.1, 4.3, 5.2, 5.5 wherein Ding maximizes the likelihood of the expert actions under the training agent policy, to Inverse Reinforcement Learning that extracts a reward function from those demonstrations and then trains a policy to maximize it and wherein Ding incorporate demonstrations into Hindsight Experience Replay for training goal-conditioned policies), ([0084] wherein Wayne describes a moderate number of demonstrations of a variety of different behaviors is available in the form of state-action sequences, or simply sequences of states. The goal is to learn a control policy that can be conditioned on a behavior embedding vector and, when conditioned appropriately, reproduce any behavior from the original set, and, at least to some extent, interpolate between them) Regarding claim 6, Wayne as modified by Andrychowicz, Nair and Ding teach wherein training the goal-conditioned policy neural network through reinforcement learning comprises using a policy gradient technique ([0126] wherein Wayne uses policy gradient algorithms o train the policy by maximizing the discounted sum rewards), (section 2.3, wherein Andrychowicz incorporates policy gradient). Regarding claim 7, Wayne as modified by Andrychowicz, Nair and Ding teach training the goal-conditioned policy neural network through imitation learning using the demonstration data (Section IV-B, equations 6–7, wherein Nair 2 supplies L2 behavior cloning). Regarding claim 8, Wayne as modified by Andrychowicz, Nair and Ding teach wherein the imitation learning optimizes a loss of the goal-conditioned policy neural network that is computed using a L2 loss function, a binary cross entropy loss function, or both (Section 3, wherein Nair incorporates parameterization to train the decoder with cross-entropy loss on normalized pixel values). Regarding claim 9, Wayne as modified by Andrychowicz, Nair and Ding teach using a task progress-based function to downscale the losses associated with actions performed in response to intermediate observations characterizing intermediate states of the environment for the particular task (section III, wherein Nair incorporates Q-Filter that account for the possibility that demonstrations can be suboptimal by applying the behavior cloning loss only to states where the critic Q ( s, a) determines that the demonstrator action is better than the actor action) Regarding claim 10, Wayne as modified by Andrychowicz, Nair and Ding teach maintaining goal observation data comprising (i) demonstration observations characterizing final states of demonstration sequences in the demonstration data and (ii) training observations characterizing final states of training sequences during which the agent successfully interacted with the environment to reach the goal state of the environment as characterized by the goal observation for the training sequence (Sections III-D and IV wherein Nair describes separate demonstration replay, actor–critic training and behavior cloning. Section IV-D uses demonstration goals when starting rollouts) Regarding claim 11, Wayne as modified by Andrychowicz, Nair and Ding teach wherein using the demonstration data to determine the one or more new goal observations each characterizing the respective goal state of the environment comprises: sampling, as the one or more new goal observations, one or more demonstration observations from the goal observation data ([0057], [0060] wherein Wayne describes the reinforcement learning that comprises the encoder of a variational autoencoder neural network, in particular a trained variational autoencoder neural network, the encoder comprising a recurrent neural network configured to encode a probability distribution of trajectories of state-action pairs as an embedding vector defining parameters representing the probability distribution, wherein the reinforcement learning system is configured to determine a target embedding vector for a target trajectory by sampling from the probability distribution encoded for the target trajectory by the encoder, and to train a reinforcement learning neural network using reward values conditioned on the target embedding vector. The system may include a policy generator and a discriminator as previously described. The decoder may comprise an autoregressive neural network to learn state representations). Regarding claim 13, Wayne as modified by Andrychowicz, Nair and Ding teach wherein generating the new training sequence comprises, for each of the new goal observations at the one or more time steps in the training sequence: generating a sparse reward that is equal to one for a respective training observation for the time step for which the distance between the encoded representation of the respective training observation and the encoded representation of the new goal observation is below the threshold distance (sections 1, 4.3 wherein Ding favorite RL algorithm on the reward and uses the off-policy algorithm DDPG to allow for the relabeling techniques. In the goal-conditioned case Ding interpolates between the GAIL reward and an indicator reward. The rewards are pushing the policy towards the goals, so it shouldn't be too conflicting. Furthermore, to avoid any drop in final performance, the weight of the reward coming from GAIL <5cAIL can be annealed. The final proposed algorithm goalGAL, together with the expert relabeling technique is formalized in Algorithm 1). Regarding claim 14, Wayne as modified by Andrychowicz, Nair and Ding teach wherein generating the new training sequence comprises: generating the new training sequence that includes (i) the one or more new goal observations determined from the demonstration data, (ii) the sparse rewards, and (iii) the plurality of training observations (Sections 1-2, 4, 4.1, 4.3, 5.2, 5.5 wherein Ding maximizes the likelihood of the expert actions under the training agent policy, to Inverse Reinforcement Learning that extracts a reward function from those demonstrations and then trains a policy to maximize it and wherein Ding incorporate demonstrations into Hindsight Experience Replay for training goal-conditioned policies). Regarding claim 15, Wayne as modified by Andrychowicz, Nair and Ding teach wherein the particular task comprises a single or dual arm robotic manipulation task (Section 4.6, wherein Andrychowicz describes physical robot applications), (Section 5.2, wherein Nair describes a robot). Regarding claim 16, Wayne as modified by Andrychowicz, Nair and Ding teach wherein the agent is a mechanical agent, the environment is a real-world environment, and the observation comprises data from one or more sensors configured to sense the real-world environment (Abstract, section 1 and 4 wherein Nair teaches sensory input and sample-efficient learning in the real world). Claims 19 and 20 are similar in scope to claim 1 therefore the claims are rejected under similar rationale. Regarding claim 21, Wayne as modified by Andrychowicz, Nair and Ding teach wherein the particular task comprises a single or dual arm robotic manipulation task (Abstract, section 1 and 4 wherein Nair teaches Robot with arms). Regarding claim 22, Wayne as modified by Andrychowicz, Nair and Ding teach wherein the agent is a mechanical agent, the environment is a real-world environment, and the observation comprises data from one or more sensors configured to sense the real-world environment (Abstract, section 1 and 4 wherein Nair teaches sensory input and sample-efficient learning in the real world). Claims 2 and 12 are rejected under AIA 35 U.S.C. 103(a) as being unpatentable over Wayne et al. US Patent Application Publication US 20200090042 A1 (hereinafter Wayne) in view of Marcin Andrychowicz et al. 2017, NPL Publication: Hindsight Experience Replay, hereinafter (Andrychowicz) and further in view of Ashvin Nair et al. NPL Publication 2018, Visual Reinforcement Learning with Imagined Goals (hereinafter Nair) and further in view of Yiming Ding et a, NPL Publication 2019, Goal-conditioned Imitation Learning (hereinafter Ding) and further in view of Nair et al. NPL Publication 2017, Overcoming Exploration in Reinforcement Learning with Demonstrations (hereinafter Nair 2). Regarding claim 2, Wayne as modified by Andrychowicz, Nair and Ding do not teach wherein obtaining the training sequence comprising the respective training observations at each of the plurality of time steps comprises: controlling the agent using the goal-conditioned policy neural network to attempt to cause the environment to transition into the goal state characterized by the goal observation. However in analogous art of demonstration-driven reinforcement learning, Nair 2 teaches wherein obtaining the training sequence comprising the respective training observations at each of the plurality of time steps comprises: controlling the agent using the goal-conditioned policy neural network to attempt to cause the environment to transition into the goal state characterized by the goal observation (Sections III-B–D, wherein Nair 2 describes a method that combines demonstrations with one such method: Deep Deterministic Policy Gradients (DDPG) [23]. DDPG is an off-policy model-free reinforcement learning algorithm for continuous control which can utilize large function approximators such as neural networks. DDPG is an actor-critic method, which bridges the gap between policy gradient methods and value approximation methods for RL. At a high level, DDPG learns an action-value function (critic) by minimizing the Bellman error, while simultaneously learning a policy (actor) by directly maximizing the estimated action-value function with respect to the parameters of the policy. Concretely, DDPG maintains an actor function with parameters, a critic function with parameters and a replay buffer as a set of tuples for each transition experienced). It would have been obvious to a person in the ordinary skill in the art before the effective filing date of the claimed invention to combine Nair 2 with Wayne, Andrychowicz, Nair and Ding by incorporating the method of wherein obtaining the training sequence comprising the respective training observations at each of the plurality of time steps comprises: controlling the agent using the goal-conditioned policy neural network to attempt to cause the environment to transition into the goal state characterized by the goal observation of Nair 2 into the method of training a reinforcement learning system to select actions to be performed by an agent interacting with an environment to perform a particular task of Wayne, Andrychowicz, Nair and Ding for the purpose incorporating a method combines DDPG and demonstrations in several ways to maximally use demonstrations to improve learning (Nair 2, Section D, IV). Regarding claim 12, Wayne as modified by Andrychowicz, Nair, Ding and Nair 2 teach wherein using the demonstration data to determine one or more new goal observations each characterizing the respective goal state of the environment comprises, for each of the plurality of time steps in the training sequence: selecting, as the new goal observation, a demonstration observation from the demonstration observations included in the one or more demonstration sequences for which a distance between an encoded representation of the training observation received at the time step and an encoded representation of the selected demonstration observation is below a threshold distance (Section V, wherein Nair 2 incorporates threshold for observations). Claim 4 is rejected under AIA 35 U.S.C. 103(a) as being unpatentable over Wayne et al. US Patent Application Publication US 20200090042 A1 (hereinafter Wayne) in view of Marcin Andrychowicz et al. 2017, NPL Publication: Hindsight Experience Replay, hereinafter (Andrychowicz) and further in view of Ashvin Nair et al. NPL Publication 2018, Visual Reinforcement Learning with Imagined Goals (hereinafter Nair) and further in view of Yiming Ding et a, NPL Publication 2019, Goal-conditioned Imitation Learning (hereinafter Ding) and further in view of Sermanet et al. NPL Publication 2018, Time-Contrastive Networks: Self-Supervised Learning from Video (hereinafter Sermanet). Regarding claim 4, Wayne as modified by Andrychowicz, Nair and Ding do not teach training the encoder neural network using contrastive learning-based training techniques. However in analogous art of demonstration-driven reinforcement learning, Sermanet teaches training the encoder neural network using contrastive learning-based training techniques (Section III-A, wherein Sermanet describes using time-contrastive networks and using representations to guide robotic imitations of human behaviors and learn to perform new tasks. Wherein the term imitation rather than demonstrations because our models also learn from passive observation of non-demonstration behaviors). It would have been obvious to a person in the ordinary skill in the art before the effective filing date of the claimed invention to combine Sermanet with Wayne, Andrychowicz, Nair and Ding by incorporating the method of training the encoder neural network using contrastive learning-based training techniques of Sermanet into the method of training a reinforcement learning system to select actions to be performed by an agent interacting with an environment to perform a particular task of Wayne, Andrychowicz, Nair and Ding for the purpose incorporating a method of self-supervised approach for learning representations and robotic behaviors entirely from unlabeled videos recorded from multiple viewpoints, and study how this representation can be used in two robotic imitation settings: imitating object interactions from videos of humans, and imitating human poses. (Sermanet, Abstract). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Any inquiry concerning this communication or earlier communications from the examiner should be directed to HASSAN MRABI whose telephone number is (571)272-8875. The examiner can normally be reached on Monday-Friday, 7:30am-5pm. Alt, Friday, EST. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Viker Lamardo can be reached on 571-270-5871. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /HASSAN MRABI/Examiner, Art Unit 2144
Read full office action

Prosecution Timeline

Apr 03, 2024
Application Filed
Sep 23, 2026
Non-Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737539
METHOD AND APPARATUS FOR IMPROVED ANALYSIS OF LEGAL DOCUMENTS
3y 1m to grant Granted Sep 15, 2026
Patent 12716804
GENERATIVE ADVERSARIAL NETWORKS FOR STRUCTURAL DAMAGE DIAGNOSTICS
3y 4m to grant Granted Aug 25, 2026
Patent 12711373
MULTIRESOLUTION HASH ENCODING FOR NEURAL NETWORKS
4y 6m to grant Granted Aug 18, 2026
Patent 12694332
DRIFT-TOLERANT MACHINE LEARNING MODELS
3y 10m to grant Granted Jul 28, 2026
Patent 12664197
EVOLUTION OF TOPICS IN A MESSAGING SYSTEM
4y 10m to grant Granted Jun 23, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
78%
Grant Probability
99%
With Interview (+33.3%)
2y 9m (~3m remaining)
Median Time to Grant
Low
PTA Risk
Based on 381 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month