Prosecution Insights
Last updated: August 17, 2026
Application No. 19/200,502

DATA-DRIVEN ROBOT CONTROL

Non-Final OA §102§DP
Filed
May 06, 2025
Priority
Sep 13, 2019 — provisional 62/900,407 +2 more
Examiner
JACKSON, DANIELLE MARIE
Art Unit
3657
Tech Center
3600 — Transportation & Electronic Commerce
Assignee
GDM Holding LLC
OA Round
1 (Non-Final)
81%
Grant Probability
Favorable
1-2
OA Rounds
1y 4m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 81% — above average
81%
Career Allowance Rate
118 granted / 146 resolved
+28.8% vs TC avg
Strong +27% interview lift
Without
With
+27.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
11 currently pending
Career history
163
Total Applications
across all art units

Statute-Specific Performance

§101
6.7%
-33.3% vs TC avg
§103
52.3%
+12.3% vs TC avg
§102
21.0%
-19.0% vs TC avg
§112
16.9%
-23.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 146 resolved cases

Office Action

§102 §DP
DETAILED ACTION This is the first office action in response to U.S. application 19/200,502. All claims are pending. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claim 1 is rejected under 35 U.S.C. 102(a)(1) as being anticipated by Ide (US 20190272477, IDS). Regarding claim 1, Ide teaches a computer-implemented method (Figs. 5-6) comprising: maintaining robot experience data characterizing robot interactions with an environment, the robot experience data comprising a plurality of experiences that each comprise an observation and an action performed by a respective robot in response to the observation ([0041] discusses sensors supplying sensor data to the input control part 31 where [0045]-[0046] discuss the input control including observation variables which are input into the state estimating part which supplies state information and detects the instructed action where the state information and action are interpreted as an experience and where because the system operates in a cycle as shown by Figs. 5-6 it would maintain a plurality of experiences); obtaining annotation data that assigns, to each experience in a first subset of the experiences in the robot experience data, a respective task-specific reward for a particular task ([0051] discusses the history producing part (annotation data) which updates an action history (tasks) and reward history associated with the actions (task-specific rewards)); training, on the annotation data, a reward model that receives as input an input observation and generates as output a reward prediction that is a prediction of a task-specific reward for the particular task that should be assigned to the input observation ([0049] “The reward estimating part 35 executes estimation of the reward imparted by the user for the action of the information processing apparatus 10 on the basis of the reward model constructed by a reward model learning part 52 and the observation variables based on the input data” with [0054] further discussing the reward model is based on the action/reward history (annotation data)); generating task-specific training data for the particular task that associates each of a plurality of experiences with a task-specific reward for the particular task ([0053]-[0054] discuss the learning part 40 which comprises a motion model learning part 51 and a reward model learning part 52 which generate learning data based on the stored experiences and the reward model being supplied to the reward estimating part), comprising, for each experience in a second subset of the experiences in the robot experience data ([0049] discusses the reward estimating part 35 receiving the observation variables and estimating a reward for the action of the apparatus based on the user input where the user input is interpreted to be the annotation data which is supported by page 8 lines 1-5 of the instant application’s specification which states that the second subset of experiences is associating reward predictions with the experience): processing the observation in the experience using the trained reward model to generate a reward prediction ([0054] “The reward model learning part 52 executes learning of the reward model used in the estimation of the reward to be imparted by the user for the action of the information processing apparatus 10 on the basis of the reward history stored in the storage part 39. The reward model learning part 52 supplies the constructed reward model to the reward estimating part 35”), and associating the reward prediction with the experience ([0049]-[0054] discuss the experience information being updated from the reward estimating part by being stored in the history producing part 38 where the learning part uses the stored associated rewards and action histories); and training a policy neural network on the task-specific training data for the particular task, wherein the policy neural network is configured to receive a network input comprising an observation and to generate a policy output that defines a control policy for a robot performing the particular task ([0045]-[0053] discuss and Fig. 1 shows how using the learning part 40, a reward is associated with an input action and stored to update the learning models of learning part 40 where the motion model learning part 51 transmits the motion model data to the motion producing part 33 where the motion control part 34 then uses this data to control the robot where the motion control is interpreted as a control policy where [0150] discusses the reward model as a neural network). Double Patenting Claim 1 of this application is patentably indistinct from claim 1 of U.S. Patent 12325130 (‘130) and claim 1 of U.S. Patent 11712799 (‘799). Pursuant to 37 CFR 1.78(f), when two or more applications filed by the same applicant or assignee contain patentably indistinct claims, elimination of such claims from all but one application may be required in the absence of good and sufficient reason for their retention during pendency in more than one application. Applicant is required to either cancel the patentably indistinct claims from all but one application or maintain a clear line of demarcation between the applications. See MPEP § 822. The non-statutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A non-statutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on non-statutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b). The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/process/file/efs/guidance/eTD-info-I.jsp. Claim 1 rejected on the ground of non-statutory double patenting as being unpatentable over claim 1 of U.S. Patent 12325130 (‘130) and claim 1 of U.S. Patent 11712799 (‘799). Although the claims at issue are not identical, they are not patentably distinct from each other, as shown in the table below. Claim This Application’s Claim 130’s Claim (1/16/2025) ‘799 Claim (12/19/2022) 1 A computer-implemented method comprising: 1. A computer-implemented method comprising: 1. A computer-implemented method comprising 1 maintaining robot experience data characterizing robot interactions with an environment, the robot experience data comprising a plurality of experiences that each comprise an observation and an action performed by a respective robot in response to the observation; 1. maintaining robot experience data characterizing robot interactions with an environment, the robot experience data comprising a plurality of experiences that each comprise an observation and an action performed by a respective robot in response to the observation; 1. maintaining robot experience data characterizing robot interactions with an environment, the robot experience data comprising a plurality of experiences that each comprise an observation and an action performed by a respective robot in response to the observation; 1 obtaining annotation data that assigns, to each experience in a first subset of the experiences in the robot experience data, a respective task-specific reward for a particular task; 1. obtaining annotation data that assigns, to each experience in a first subset of the experiences in the robot experience data, a respective task-specific reward for a particular task; wherein obtaining the annotation data comprises: providing, for presentation to a user, a representation of one or more of the experiences in the first subset of experience data, comprising providing a video of a robot performing an episode of the particular task; and obtaining, from the user, inputs defining the rewards for the one or more experiences, comprising obtaining, for each of a plurality of frames of the video, a respective input that associates the frame of the video with a respective measure of progress of the robot towards completing the particular task as of the frame 1. obtaining annotation data that assigns, to each experience in a first subset of the experiences in the robot experience data, a respective task-specific reward for a particular task, wherein the first subset of experiences comprises experiences from a plurality of different task episodes of the particular task; 1 training, on the annotation data, a reward model that receives as input an input observation and generates as output a reward prediction that is a prediction of a task-specific reward for the particular task that should be assigned to the input observation; 1. training, on the annotation data, a reward model that receives as input an input observation and generates as output a reward prediction that is a prediction of a task-specific reward for the particular task that should be assigned to the input observation; 1. training, on the annotation data, a reward model that receives as input an input observation and generates as output a reward prediction that is a prediction of a task-specific reward for the particular task that should be assigned to the input observation, 1 generating task-specific training data for the particular task that associates each of a plurality of experiences with a task-specific reward for the particular task, comprising, for each experience in a second subset of the experiences in the robot experience data: 1. generating task-specific training data for the particular task that associates each of a plurality of experiences with a task-specific reward for the particular task, comprising, for each experience in a second subset of the experiences in the robot experience data: 1. generating task-specific training data for the particular task that associates each of a plurality of experiences with a task-specific reward for the particular task, comprising, for each experience in a second subset of the experiences in the robot experience data: 1 processing the observation in the experience using the trained reward model to generate a reward prediction, and associating the reward prediction with the experience; 1. processing the observation in the experience using the trained reward model to generate a reward prediction, and associating the reward prediction with the experience; 1. processing the observation in the experience using the trained reward model to generate a reward prediction, and associating the reward prediction with the experience; 1 and training a policy neural network on the task-specific training data for the particular task, wherein the policy neural network is configured to receive a network input comprising an observation and to generate a policy output that defines a control policy for a robot performing the particular task. 1. and training a policy neural network on the task-specific training data for the particular task, wherein the policy neural network is configured to receive a network input comprising an observation and to generate a policy output that defines a control policy for a robot performing the particular task. 1. and training a policy neural network on the task-specific training data for the particular task, wherein the policy neural network is configured to receive a network input comprising an observation and to generate a policy output that defines a control policy for a robot performing the particular task. ‘799’s claim 1 is the same as the instant application’s claim 1 except that it includes detail on the first subset of experiences comprising experiences from a plurality of different task episodes of a particular task and optimizing the reward model using a loss function. As claim 1 of ‘799 contains all of the claim limitations of the instant application’s claim 1, the instant application has a broader scope than application ‘799 and therefore is not patentably distinct from application ‘799. ‘130’s claim 1 is the same as the instant application’s claim 1 except that it includes detail on the annotation data comprising video frames and associating video frames with the completion of a task. As claim 1 of ‘130 contains all of the claim limitations of the instant application’s claim 1, the instant application has a broader scope than application ‘130 and therefore is not patentably distinct from application ‘130. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: Kim (US 20220032450, IDS) and Luciw (US 20190061147, IDS) teach storing a plurality of experience which includes reward information; Liu (US 20210308863, IDS) teaches episode-based reinforcement learning; Porter (US 10766136, IDS) teaches using a model to determine a level of success for a robotic task; Ozawa (US 20200250490, IDS) teaches updating a reward model using observation variables; Otsuka (US 20190314983, IDS) teaches reinforcement learning using evaluation values as a reward to learn the action model; Lewis (US 20210174245) teaches using reinforcement learning to predict second actions based on a first predicted action and reward; and Xu (US 20200175364) teaches training a reinforcement learning neural network with a reward function. Any inquiry concerning this communication or earlier communications from the examiner should be directed to DANIELLE M JACKSON whose telephone number is (303)297-4364. The examiner can normally be reached Monday-Friday 7:00-4:30 MT. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Abby Lin can be reached at (571) 270-3976. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /D.M.J./ Examiner, Art Unit 3657 /DYLAN M KATZ/ Primary Examiner, Art Unit 3657
Read full office action

Prosecution Timeline

May 06, 2025
Application Filed
Jun 17, 2026
Non-Final Rejection mailed — §102, §DP (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12703099
HAND-EYE CALIBRATION METHODS, SYSTEMS, AND STORAGE MEDIA FOR ROBOTS
2y 6m to grant Granted Aug 11, 2026
Patent 12691897
Trajectory Planning in a Three-Dimensional Search Space with a Space-Time Artificial Potential Field
3y 0m to grant Granted Jul 28, 2026
Patent 12686130
Modeling a Robot Working Environment
2y 0m to grant Granted Jul 21, 2026
Patent 12682756
PARKING SYSTEM AND PARKING METHOD
2y 0m to grant Granted Jul 14, 2026
Patent 12681485
DRONE FLIGHT CONTROL CENTER AND PILOT MONITOR
1y 11m to grant Granted Jul 14, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
81%
Grant Probability
99%
With Interview (+27.0%)
2y 7m (~1y 4m remaining)
Median Time to Grant
Low
PTA Risk
Based on 146 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month