Prosecution Insights
Last updated: August 17, 2026
Application No. 18/020,688

Training a Reinforcement Learning Agent to Control an Autonomous System

Non-Final OA §103
Filed
Feb 10, 2023
Priority
Aug 11, 2020 — DE 10 2020 121 150.3 +1 more
Examiner
CHOU, SHIEN MING
Art Unit
3667
Tech Center
3600 — Transportation & Electronic Commerce
Assignee
Bayerische Motoren Werke Aktiengesellschaft
OA Round
2 (Non-Final)
58%
Grant Probability
Moderate
2-3
OA Rounds
4m
Est. Remaining
87%
With Interview

Examiner Intelligence

Grants 58% of resolved cases
58%
Career Allowance Rate
62 granted / 106 resolved
+6.5% vs TC avg
Strong +28% interview lift
Without
With
+28.3%
Interview Lift
resolved cases with interview
Typical timeline
3y 11m
Avg Prosecution
19 currently pending
Career history
129
Total Applications
across all art units

Statute-Specific Performance

§101
14.9%
-25.1% vs TC avg
§103
49.3%
+9.3% vs TC avg
§102
16.2%
-23.8% vs TC avg
§112
19.0%
-21.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 106 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Respond to Amendment This action is in response to the application filed on ----1/26/2026 for application 18/020,688. Claim 10, 12 – 13, 15 – 18 are pending and have been examined. Claim rejection under 101 and 112 section has been withdrawn in light of the amendment. Claim interpretation has been withdrawn in light of the applicants remarks. Claim 10, 18 are amended. Claim 11, 14 are canceled. Respond to Argument Applicant’s arguments with respect to claim(s) have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claim 10, 12 – 13, 15 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Wang et al., (hereinafter Wang), “A Reinforcement Learning Based Approach for Automated Lane Change Maneuvers” in view of Chen et al., (hereinafter Chen), US20180188039. Claim 10. Wang teaches: A device for training a reinforcement learning agent to control an autonomous system including an automated motor vehicle (sec. I, “autonomous vehicles”, “The focus of the paper is to demonstrate the application of Reinforcement Learning on finding an optimal driving policy”), the device comprising: An electronic control unit (sec. III “In a typical RL problem, the goal is to find an optimal policy”; the system is to perform data processing. A processor (electronic control unit) is inherent) configured to: detect an environment of the autonomous system based on signals received from a plurality of sensors disposed on the automated motor vehicle (sec. III, “sensor inputs … the RL agent can take in hundreds or even millions of features as inputs”; i.e., use sensor to detect its environment); recognize at least one object comparable to the autonomous system in the environment of the autonomous system (sec. III, “the driving environment involves the interaction with other vehicles (at least one object comparable to the autonomous system)”); detect a behavior of the at least one object comparable to the autonomous system (sec. III, “other vehicles whose behaviors may be cooperative or adversarial”); train the reinforcement learning agent in accordance with the detected behavior of the at least one object comparable to the autonomous system to produce a trained reinforcement learning agent (refer to the mapping above & fig. 3, Wang demonstrate the use of reinforcement learning to train an agent based on the detected behavior of the comparable objects); control operation of the automated motor vehicle based on the trained reinforcement learning agent (sec. V, “the RL agent … output a reference guidance for a traditional optimization-based controller to issue a quick and reliable control command to the vehicle”). Wang does not explicitly teach: transform a representation of the at least one object comparable to the autonomous system in the environment of the autonomous system by performing a coordinate transformation in which the representation of the at least one object comparable to the autonomous system in a coordinate system is placed at a position of the autonomous system such that the representation of the at least one object comparable to the autonomous system corresponds to a possible representation of the autonomous system Chen, in the same field of endeavor, explicitly teach: transform a representation of the at least one object comparable to the autonomous system in the environment of the autonomous system by performing a coordinate transformation in which the representation of the at least one object comparable to the autonomous system in a coordinate system is placed at a position of the autonomous system such that the representation of the at least one object comparable to the autonomous system corresponds to a possible representation of the autonomous system (Chen, fig. 9A-B, 0095 – 0098, “vehicle coordinate system illustrated in FIG. 9A such that the positive direction of the X-axis is forward facing direction of the vehicle, the positive direction of the Y-axis is to the left of the vehicle when facing forward, and the positive direction of the Z-axis is upward.”, “LIDAR coordinate system illustrated in FIG. 9B such that the positive direction of the X-axis is to the left of the vehicle when facing forward, the positive direction of the Y-axis is in the forward direction of the vehicle, and the positive direction of the Z-axis is upward.”, “determines the coordinate transform M12c to map the LIDAR coordinate system to the vehicle coordinate system”, “performing the transformation Pcar =Tl2c * Plidar”; i.e., the object detected by sensor/Lidar are to be coordinate-transformed to align with the vehicle coordinate before it is processed for the control of the vehicle) Wang and Chen both teach system for controlling autonomous vehicles and are analogous. It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention with a reasonable likelihood of success to further include the processing details that combine sensor data and map data for the control of the vehicle taught by Chen in the system of Wang to achieve the claimed teaching. One of the ordinary skill in the art would have motivated to make this modification to accommodate sensors of different setup (Chen, 0095 – 0098). Claim 12. Wang teaches all the limitation of Claim 10. Wang further teach: the device is further configured to: train the reinforcement learning in accordance with: the detected behavior of the at least one object comparable to the autonomous system (sec. III, “a joint approach, a RL/ML module can be taken as a mediator component between a perception module and a traditional MPC module, as to take the perception outcome as input and output a reference guidance for the MPC controller”; sec. 3.2, “the decision-making module will issue a command to alter or abort the maneuver”; i.e., the reinforcement training take account the detected behavior of other vehicle); and a reward function for controlling the autonomous system (sec. 3.3.3, “As the longitudinal module and the gap selection module have taken into account the safety concerns, smoothness and efficiency are considered by the lateral controller through the reward function.”). Claim 13. Wang teaches all the limitation of Claim 10. Wang further teach: the device is further configured to: detect a state of the at least one object comparable to the autonomous system; detect an action of the at least one object comparable to the autonomous system, which changes the state of the at least one object comparable to the autonomous system; detect a resulting state of the at least one object comparable to the autonomous system, which is caused by the action of the at least one object comparable to the autonomous system (sec. 3.2, “a time-continuous car-following model for the simulation of highway and urban traffic. It describes the dynamics of relevant vehicles”; i.e., the training data are time series of data. In the case described in sec. I, “drive at high speed by passing a slow vehicle in front”, the distance to the front vehicle at T1 is the state, the speed of the front vehicle is the action and the distance to the front vehicle at T2 is the resulting state ); and train the reinforcement learning agent in accordance with: the state of the at least one object comparable to the autonomous system, the action of the at least one object comparable to the autonomous system, the resulting state of the at least one object comparable to the autonomous system, and a reward function for controlling the autonomous system (refer to the mapping above and Claim 12, the system uses the time series data and a reward function to train the agent). Claim 15. Wang teaches all the limitation of Claim 10. Wang further teach: the device is further configured to: recognize at least two objects comparable to the autonomous system in the environment of the autonomous system; detect the behavior of each of the at least two objects comparable to the autonomous system; and train the reinforcement learning agent in accordance with the respective detected behavior of the at least two objects comparable to the autonomous system (fig. 3 & sec. 3.2, “it may see two leading vehicles in the ego lane and the target lane during the transition … ego-vehicle to adjust its longitudinal acceleration by balancing between its two leaders, if observed, on its ego lane and the target lane. The smaller value will be used to weaken the potential discontinuity in vehicle acceleration incurred from lane change initiation.”; i.e., the model involve observation/detection of multiple vehicles. The training is based on the collective information from such scenario). Claim 18. Claim 18 is the corresponding method of Claim 10. Claim 18 is rejected with same reason. Claim(s) 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wang et al., (hereinafter Wang), “A Reinforcement Learning Based Approach for Automated Lane Change Maneuvers” in view of Chen et al., (hereinafter Chen), US20180188039 as applied to claim 15 above, and further in view of Katare et al., (hereinafter Katare) “Autonomous Embedded System Enabled 3-D Object Detector: (with Point Cloud and Camera)”. Claim 16. Wang and Chen combination teaches all the limitation of Claim 15. The combination does not explicitly teach the device is further configured to: train the reinforcement learning agent simultaneously in accordance with the respective detected behavior of the at least two objects comparable to the autonomous system. Katare, in the same field of endeavor, explicitly teach: train the reinforcement learning agent simultaneously in accordance with the respective detected behavior of the at least two objects comparable to the autonomous system (fig. 1, & “Deep learning based … object detection involves drawing of bounding boxes”; sec. III, “the outputs are bounding boxes on the segmented objects with variables: center, size and yaw angle (cx, cy, cz, l, w, h, ϴ)” the model is trained to detect attributes of multiple objects/vehicles simultaneously). Wang (in view of Chen) and Katare both teach image detection for vehicles and are analogous. It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention with a reasonable likelihood of success to further include the multi-objects simultaneous detection taught by Katare in the system of Wang (in view of Chen) to achieve the claimed teaching. One of the ordinary skill in the art would have motivated to make this modification for its high performance (Katare sec. I). Claim(s) 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wang et al., (hereinafter Wang), “A Reinforcement Learning Based Approach for Automated Lane Change Maneuvers” in view of Chen et al., (hereinafter Chen), US20180188039 as applied to claim 15 above, and further in view of Rosebrock, “YOLO object detection with OpenCV”. Claim 17. Wang and Chen combination teaches all the limitation of Claim 15. The combination does not explicitly teach the device is further configured to: fuse a representation of the at least two objects comparable to the autonomous system in the environment of the autonomous system using a permutation-invariant or using a permutation-equivariant mapping into a common representation; and train the reinforcement learning agent in accordance with the common representation. Rosebrock, in the same field of endeavor, explicitly teach: fuse a representation of the at least two objects comparable to the autonomous system in the environment of the autonomous system using a permutation-invariant or using a permutation-equivariant mapping into a common representation; and train the reinforcement learning agent in accordance with the common representation (Rosebrock, page 14 – 15, The detected objects are appended/added (fused) into a list with same format of x, y, width, height, confidence, classID (common representation). The for loop goes through the detected list one by one, thus permutation-equivariant mapping/appending into the list). Wang (in view of Chen) and Rosebrock both teach object detection for multi objects and are analogous. It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention with a reasonable likelihood of success to further include the output format of the detected results taught by Rosebrock in the system of Wang (in view of Chen) to achieve the claimed teaching. One of the ordinary skill in the art would have motivated to make this modification so the following process can easily read and work on the data (Rosebrock, page 16 – 17, use simple for loop to get all the detected objects in the list). Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHIEN MING CHOU whose telephone number is (571)272-9354. The examiner can normally be reached Monday- Friday 9 am - 5 pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, HITESH PATEL can be reached on 571-270-5442. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SHIEN MING CHOU/Examiner, Art Unit 3667 /ANSHUL SOOD/Primary Examiner, Art Unit 3667
Read full office action

Prosecution Timeline

Feb 10, 2023
Application Filed
Nov 12, 2025
Non-Final Rejection mailed — §103
Jan 26, 2026
Response Filed
May 04, 2026
Final Rejection mailed — §103
Jul 28, 2026
Response after Non-Final Action

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12704846
SYSTEM FOR SYNCING A HARVESTER WITH AN AUTONOMOUS GRAIN CART
3y 0m to grant Granted Aug 11, 2026
Patent 12691867
HYBRID ELECTRIC VEHICLE AND DRIVING CONTROL METHOD THEREFOR
3y 9m to grant Granted Jul 28, 2026
Patent 12680833
METHOD, DEVICE AND SYSTEM FOR PROCESSING A FLIGHT TASK
1y 7m to grant Granted Jul 14, 2026
Patent 12650688
REMOTE ASSISTANCE APPARATUS, METHOD, AND PROGRAM
4y 2m to grant Granted Jun 09, 2026
Patent 12637020
Vehicle Power Architecture, Power Control Module and Associated Method
3y 5m to grant Granted May 26, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

2-3
Expected OA Rounds
58%
Grant Probability
87%
With Interview (+28.3%)
3y 11m (~4m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 106 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month