Prosecution Insights
Last updated: October 01, 2026
Application No. 18/679,182

END-TO-END TRAINABLE ADVANCED DRIVER-ASSISTANCE SYSTEMS

Final Rejection §103
Filed
May 30, 2024
Priority
May 31, 2023 — provisional 63/505,352
Examiner
LU, ZHIYU
Art Unit
2665
Tech Center
2600 — Communications
Assignee
Waymo LLC
OA Round
2 (Final)
49%
Grant Probability
Moderate
3-4
OA Rounds
1y 6m
Est. Remaining
63%
With Interview

Examiner Intelligence

Grants 49% of resolved cases
49%
Career Allowance Rate
381 granted / 779 resolved
-13.1% vs TC avg
Moderate +14% lift
Without
With
+14.1%
Interview Lift
resolved cases with interview
Typical timeline
3y 10m
Avg Prosecution
48 currently pending
Career history
833
Total Applications
across all art units

Statute-Specific Performance

§101
2.8%
-37.2% vs TC avg
§103
67.5%
+27.5% vs TC avg
§102
11.9%
-28.1% vs TC avg
§112
16.6%
-23.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 779 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Applicant’s arguments with respect to claim(s) 1-20 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Applicant's arguments filed 08/03/2026 have been fully considered but they are not persuasive. Regarding amended claim 1, applicant argued that object detection and classification functions in Nehmadi are not apparently based on the lidar data. Applicant furthered argued that Nehmadi does not apparently disclose using the lidar data to generate the object bounding boxes, nor to generate object motion predictions. However, examiner respectfully disagrees. Despite of applicant’s argument, “lidar data” in argued claim is a base for target prediction. The argued claim limit neither “lidar data” alone to output object detection and classification, nor “lidar data” alone to generate object bounding box. Yet, Nehmadi does disclose: [0087] The neural network 112 may be trained during a training phase. This may involve feedforward of data signals to generate the output and then the backpropagation of errors for gradient descent optimization. For example, in some embodiments, the neural network 112 may be trained by using a set of reference images and reference data about the objects and classes of objects in the reference images. That is to say, the neural network 112 is fed a plurality of reference images and is provided with the “ground truth” (i.e., is told what objects and classes of objects are in the reference images and where they appear in the reference images) so that the neural network 112 is trained to recognize (i.e., detect and classify) those objects in images other than the reference images such as the regions of interest supplied by the FLD unit 110. Training results in converging on the set of parameters 150. [0095] The reference sensors 302 may be high-quality sensors that are able to produce accurate RGBDV images covering a wide range of driving scenarios, whereas the production sensors 304 may be lower-cost sensors more suitable for use in a commercial product or real-time environment. As such, the reference sensors 302 are sometimes referred to as “ground truth sensors” and the production sensors 304 are sometimes referred to as “high-volume manufacturing (HMV) sensors” or “test sensors”. In the case of lidar, for example, a lidar sensor that is used as a reference sensor may have a higher resolution, greater field of view, greater precision, better SNR, greater range, higher sensitivity and/or greater power consumption than a production version of this lidar sensor. To take a specific example, the set of reference sensors 302 may include a lidar covering a 360° field of view, with an angular resolution of 0.1°, and range of 200 m, whereas the set of production sensors 304 may include a lidar covering a 120° field of view, with an angular resolution of 0.5° and range of 100 m, together with a radar covering a 120° field of view, with an angular resolution 2° and a range of 200 m. [0084] The indication of the location of one or more objects in a given image, as output by the neural network 112, may include bounding boxes within the image. Each bounding box may surround a corresponding object associated with an object descriptor. The bounding box may be a 2D bounding box or a 3D bounding box, for example. In other cases, the indication of the location of the one or more objects in the image may take the form of a silhouette, cutout or segmented shape. Thus, said prior art is properly applied and maintained. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-4, 6-12, 14-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Nehmadi et al. (US2022/0335729) in view of Fowe (US2019/0187720) and Yang et al. (US2020/0210726). To claim 1, Nehmadi teach a method comprising: generating a set of auto-labeled training data using a first autonomous vehicle (AV) system having multiple sensor modalities (paragraph 0097, set of reference sensors 302 and/or set of production sensors 304 may include various combinations of sensors such as one or more lidar sensors and one or more non-lidar sensors), wherein the auto-labeled training data comprises: a first set of non-lidar data associated with the first AV system (paragraphs 0094, 0097, production sensors 304 may comprises non-lidar sensors), and one or more target predictions for the first set of non-lidar data, the one or more target predictions generated based at least on lidar data associated with the first AV system (paragraphs 0087, 0095, reference sensors 302 are referred to as ground truth; page 12 claim 27, reference sensor may be lidar sensor, wherein such scenario would result non-lidar being labeled/annotated/classified based on lidar data), the one or more target predictions comprising: a bounding box for an object represented in the first set of sensor data (paragraph 0084, output by neural network may including bounding boxes within image), and a motion track of the object detection represented in the first set of sensor data (obvious in paragraph 0075, velocity or motion information); and training, by the processing device and using the auto-labeled training data, an end-to-end perception model of a second AV system lacking a lidar sensor (obvious in paragraphs 0003-0004, objective is to harness advantages of neural networks using available vehicle-grade computing hardware and relatively economic sensors, such that lidar sensor may not be available such as paragraph 0097 showing lidar sensor being optional) to predict presence of one or more objects in a driving environment of the second AV system (Figs. 3A-B; 100, 300 of Fig. 9, two perception systems; Fig. 10; paragraph 0193, embodiments such as particular features, structures or characteristics may be combined in any suitable manner. Different embodiments shows that lidar and/or non-lidar signals may be applied in training and real-time utilization due to limited or available vehicle-grade computing hardware and relatively economic sensors, which combines the ease of use benefits of unsupervised learning with the accuracy and high performance of supervised learning. It would have been obvious to one of ordinary skill in the art to recognize instant claimed invention as one of implementation scenarios in view of embodiments of Nehmadi). Although Nehmadi teach first AV system and second AV system being resided in the same vehicle, instant claim lacks clarification. Moreover, sharing labeled training data among different vehicles would be obvious in the art. In furthering said obviousness, Fowe teach vehicles sharing labeled training data set (Figs. 1-2, paragraphs 0035-0036, 0070). And, Yang teach training a deep neural network in a vehicle system to detect object, to classify the object, to track the object motion, and to generate bounding box on the object within captured image based on lidar data (abstract, Figs. 3A-B; paragraphs 0007-0010, 0040-0047). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate teaching of Fowe and Yang into the method of Nehmadi, in order to improve object detection machine learning classifiers of computer vision system. To claim 9, Nehmadi, Fowe and Yang teach a system (as explained in response to claim 1 above). To claim 17, Nehmadi, Fowe and Yang teach a non-transitory computer-readable storage medium having instructions stored thereon that, when executed by a processing device, cause the processing device to perform operations (as explained in response to claim 1 above). To claim 2, Nehmadi, Fowe and Yang teach claim 1. Nehmadi, Fowe and Yang teach further comprising: obtaining a second set of sensor data using the second AV system, wherein the second set of sensor data (304 of Fig. 9): comprises sensor data of the second plurality of sensor modalities (paragraph 0097), and characterizes one or more objects in a driving environment of the second vehicle; and processing, using the trained end-to-end perception model, the second set of sensor data to obtain one or more inference predictions associated with the one or more objects in the driving environment of the second AV system (as explained in response to claim 1 above, with teachings of Nehmadi, Fowe and Yang). To claim 3, Nehmadi, Fowe and Yang teach claim 2. Nehmadi, Fowe and Yang teach wherein the second set of sensors of the second vehicle comprises at least one camera sensor or at least one radar sensor (Nehmadi, paragraph 0097). To claim 4, Nehmadi, Fowe and Yang teach claim 1. Nehmadi, Fowe and Yang teach wherein the one or more target predictions are generated by one or more models trained using a set of manually-labeled training data (Nehmadi, paragraphs 0078, 0081, 0109, while detection operation may be unsupervised, it also means said detection operation may be supervised; and in initial training stage, and supervised training requires manually-labeled training data, which is also well-known in the art, hence Official Notice is taken). To claim 6, Nehmadi, Fowe and Yang teach claim 1. Nehmadi, Fowe and Yang teach wherein the one or more target predictions comprise at least one of: a bounding box for an object represented in the first set of sensor data, a type of the object detection represented in the first set of sensor data, or a motion track of the object detection represented in the first set of sensor data (Nehmadi, paragraphs 0084-0085). To claim 7, Nehmadi, Fowe and Yang teach claim 1. Nehmadi, Fowe and Yang teach wherein the second set of sensor data comprises at least some of the sensor data of the first set of sensor data (Nehmadi, Figs. 3A-B). To claim 8, Nehmadi, Fowe and Yang teach claim 1. Nehmadi, Fowe and Yang teach wherein the one or more target predictions are further based on: camera data associated with the first AV system, and radar data associated with the first AV system (Nehmadi, paragraph 0097, non-lidar sensors such as radar sensor and camera). To claim 10, Nehmadi, Fowe and Yang teach claim 9. Nehmadi, Fowe and Yang teach wherein the processing device is further configured to: obtain a second set of sensor data using the second AV system, wherein the second set of sensor data: comprises sensor data of the second plurality of sensor modalities, and characterizes one or more objects in a driving environment of the second vehicle; and process, using the trained end-to-end perception model, the second set of sensor data to obtain one or more inference predictions associated with the one or more objects in the driving environment of the second AV system (as explained in response to claim 1 above, wherein training data are shared). To claim 11, Nehmadi, Fowe and Yang teach claim 10. Nehmadi, Fowe and Yang teach wherein the second set of sensors of the second vehicle comprises at least one camera sensor or at least one radar sensor (Nehmadi, paragraph 0097). To claim 12, Nehmadi, Fowe and Yang teach claim 9. Nehmadi, Fowe and Yang teach wherein the one or more target predictions are generated by one or more models trained using a set of manually-labeled training data (as explained in response to claim 4 above). To claim 14, Nehmadi, Fowe and Yang teach claim 9. Nehmadi, Fowe and Yang teach wherein the one or more target predictions comprise at least one of: a bounding box for an object represented in the first set of sensor data, a type of the object detection represented in the first set of sensor data, or a motion track of the object detection represented in the first set of sensor data (as explained in response to claim 6 above). To claim 15, Nehmadi, Fowe and Yang teach claim 9. Nehmadi, Fowe and Yang teach wherein the second set of sensor data comprises at least some of the sensor data of the first set of sensor data (as explained in response to claim 7 above). To claim 16, Nehmadi, Fowe and Yang teach claim 9. Nehmadi, Fowe and Yang teach wherein the one or more target predictions are further based on: camera data associated with the first AV system, and radar data associated with the first AV systems (as explained in response to claim 8 above). To claim 18, Nehmadi, Fowe and Yang teach claim 17. Nehmadi, Fowe and Yang teach wherein the operations further comprise: obtaining a second set of sensor data using the second AV system, wherein the second set of sensor data: comprises sensor data of the second plurality of sensor modalities, and characterizes one or more objects in a driving environment of the second vehicle; and processing, using the trained end-to-end perception model, the second set of sensor data to obtain one or more inference predictions associated with the one or more objects in the driving environment of the second AV system (as explained in response to claim 10 above). To claim 19, Nehmadi, Fowe and Yang teach claim 17. Nehmadi, Fowe and Yang teach wherein the second set of sensors of the second vehicle comprises at least one camera sensor or at least one radar sensor (as explained in response to claim 11 above). To claim 20, Nehmadi, Fowe and Yang teach claim 17. Nehmadi, Fowe and Yang teach wherein the one or more target predictions comprise at least one of: a bounding box for an object represented in the first set of sensor data, a type of the object detection represented in the first set of sensor data, or a motion track of the object detection represented in the first set of sensor data (as explained in response to claim 14 above). Claim(s) 5, 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Nehmadi et al. (US2022/0335729) in view of Fowe (US2019/0187720), Yang et al. (US2020/0210726) and Cohen et al. (US2023/0054575). To claims 5 and 13, Nehmadi, Yang and Fowe teach claims 1 and 9. Despite lack of disclosure, wherein training the end-to-end perception model of the second AV system comprises: pre-training a base model using the set of auto-labeled training data; and fine-tuning the pre-trained base model using at least one set of manually labeled training data would be an obvious-to-try scenario in teachings Nehmadi, Yang and Fowe. In further said obviousness, Cohen teach machine learning in vehicle system, wherein perform supervised learning, unsupervised learning, semi-supervised learning (e.g., supervised pre-training followed by unsupervised fine-tuning, unsupervised pre-training followed by supervised fine-tuning, and/or the like), self-supervised learning, and/or the like (paragraph 0033), which would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate into the method of Nehmadi, Yang and Fowe, in order to implement training by design preference. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to ZHIYU LU whose telephone number is (571)272-2837. The examiner can normally be reached Weekdays: 8:30AM - 5:00PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Stephen R Koziol can be reached at (408) 918-7630. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. ZHIYU . LU Primary Examiner Art Unit 2669 /ZHIYU LU/Primary Examiner, Art Unit 2665 August 31, 2026
Read full office action

Prosecution Timeline

May 30, 2024
Application Filed
Apr 16, 2026
Non-Final Rejection mailed — §103
Jul 01, 2026
Interview Requested
Jul 17, 2026
Interview Requested
Aug 03, 2026
Response Filed
Sep 03, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743905
WEARABLE DEVICE AND BEHAVIOR EVALUATION SYSTEM
2y 3m to grant Granted Sep 22, 2026
Patent 12743790
METHOD FOR MEASURING CHANNEL FLOW BASED ON BIONIC EAGLE-EYE VISION AND APPARATUS THEREOF
1y 5m to grant Granted Sep 22, 2026
Patent 12724106
METHOD, APPARATUS, AND SYSTEM FOR WIRELESS HUMAN AND NON-HUMAN MOTION DETECTION
2y 8m to grant Granted Sep 01, 2026
Patent 12720330
Network-Controlled E-UTRAN Neighbor Cell Measurements
3y 11m to grant Granted Aug 25, 2026
Patent 12711645
WAVEFRONT SENSOR-BASED SYSTEMS FOR CHARACTERIZING OPTICAL ZONE DIAMETER OF AN OPHTHALMIC DEVICE AND RELATED METHODS
5y 5m to grant Granted Aug 18, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
49%
Grant Probability
63%
With Interview (+14.1%)
3y 10m (~1y 6m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 779 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month