Prosecution Insights
Last updated: August 06, 2026
Application No. 18/344,750

SYSTEM AND METHOD TO IMPROVE MULTI-CAMERA MONOCULAR DEPTH ESTIMATION USING POSE AVERAGING

Non-Final OA §103§112
Filed
Jun 29, 2023
Priority
Mar 16, 2021 — provisional 63/161,614 +1 more
Examiner
CHANG, DANIEL CHEOLJIN
Art Unit
2669
Tech Center
2600 — Communications
Assignee
Toyota Technological Institute AT Chicago
OA Round
1 (Non-Final)
88%
Grant Probability
Favorable
1-2
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 88% — above average
88%
Career Allowance Rate
130 granted / 147 resolved
+26.4% vs TC avg
Moderate +14% lift
Without
With
+13.9%
Interview Lift
resolved cases with interview
Typical timeline
2y 5m
Avg Prosecution
10 currently pending
Career history
164
Total Applications
across all art units

Statute-Specific Performance

§101
7.5%
-32.5% vs TC avg
§103
53.0%
+13.0% vs TC avg
§102
14.3%
-25.7% vs TC avg
§112
21.9%
-18.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 147 resolved cases

Office Action

§103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Notice to Applicants This communication is in response to the Application filed on 06/29/2023. Claims 1-16 are pending. Claim Objections Claim 3, 4 and 10 are objected to because of the following informalities: In claim 3, line 2, “target images” should be “the target images”. In claim 4, line 3 and 5, “target images” should be “the target images”. In claim 10, line 3 and 5, “selected target images” should be “the selected target images”. Appropriate correction is required. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claim 2-4 and 10-12 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 2 - the scene (line 2): There is insufficient antecedent basis for this limitation in the claim. Claim 10 - the scene (line 3): There is insufficient antecedent basis for this limitation in the claim. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-6 and 9-14 are rejected under 35 U.S.C. 103 as being unpatentable over GU et al. (U.S. Publication No. 2021/0049371) (hereafter, "GU") in view of Ren et al (U.S. Publication No. 2022/0391632) (hereafter, "Ren") and further in view of Stein et al. (U.S. Publication No. 2022/0237866) (hereafter, "Stein"). Regarding claim 1, GU teaches A method for multi-camera monocular depth estimation using pose averaging, comprising ([0136] The pose graph output from the pose graph construction algorithm or the refined pose graph output from the pose graph optimization algorithm may combined with the depth map output from the mapping-net to produce a 3D point cloud 440; [0045] there is provide an unsupervised deep learning architecture for estimating pose and depth and optionally a point cloud based on image data captured by monocular cameras; [0132]-[0133]); training … an ego-motion estimation model ([0112] the 6 DOF pose output from the tracking-net is compared with the 6 DOF pose calculated by the loss functions and the mapping net is trained to minimise this error via backpropagation) according to the multi-camera PCC photometric loss ([0073] the loss functions may include spatial loss functions and temporal loss functions; [0075] The spatial loss functions may themselves include three subset loss functions. These will be referred to as the spatial photometric consistency loss function, the disparity consistency loss function and the pose consistency loss function). GU does not expressly teach adjusting a multi-camera photometric loss associated with a multi-camera rig of an ego vehicle according to a multi-camera pose consistency constraint (PCC) loss associated with the multi-camera rig of the ego vehicle to form a multi-camera PCC photometric loss; and training a multi-camera depth estimation model … according to the multi-camera PCC photometric loss. However, Ren teaches adjusting a multi-camera photometric loss … according to a multi-camera pose consistency constraint (PCC) loss … to form a multi-camera PCC photometric loss; and ([0125] The unsupervised learning module 114 may generate a 3D point cloud for each of the input images according to the estimated depth for each of the input images, and may calculate an unsupervised photometric loss based on distances between 3D coordinates of aligned feature points in the 3D point clouds of the input images; [0130] the unsupervised learning module 114 may include an inverse projection and calibration image transformer 602 to inverse project each of the input images (Ia, Ib) to a 3D space (e.g., camera coordinates) according to the estimated depth (Da, Db) of the input images (Ia, Ib), and further to world coordinates by an extrinsic matrix ... the inverse projection and calibration image transformer 602 may identify common regions in the input images (Ia, Ib) according to the 3D world coordinates based on the estimated depth (Da, Db), and may calibrate (e.g., may align) the 3D world coordinates of the input images (Ia, Ib) to each other according to the identified common regions; [0131] the unsupervised learning module 114 may calculate an unsupervised photometric loss (Lu) based on distances between the 3D coordinates of the aligned feature points in the 3D point clouds of each of the input images (Ia, Ib); [0077]-[0080]) training a multi-camera depth estimation model … according to the multi-camera PCC photometric loss ([0046] the CV training system 102 may utilize a supervised learning technique (S), an unsupervised learning technique (U), and a weakly supervised learning technique (W) to be trained for various different CV application scenarios. Some non-limiting examples of the CV applications may include monocular depth estimation, stereo matching, image/video enhancement, multi-view depth estimation; [0131] the unsupervised learning module 114 may calculate an unsupervised photometric loss (Lu); [0099]). It would have been obvious before the effective filing date of the claimed invention to one having ordinary skill in the art to modify the device and method of combination of GU to incorporate the step/system of generating a 3D point cloud in world coordinates by camera extrinsic matrix for each input image, aligning the corresponding feature points across these point clouds, computing a photometric loss based on the discrepancies between the 3D coordinates of the aligned feature points and training a multi-view depth estimation … according to the pose consistency photometric loss taught by Ren. The suggestion/motivation for doing so would have been to improve multi-view depth estimation ([0041] accuracy of the supervision output may be further refined (e.g., further improved) during optimization of the unsupervised loss and the weakly supervised loss; [0134] the CV training system 102 may be trained to improve multi-view depth estimation by optimizing the supervised loss function ... the unsupervised loss function ... and the weakly supervised loss function ... concurrently; [0089] the CV training system 102 may be trained to improve monocular depth estimation by optimizing the supervised loss function). Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predicted results. The combination of GU and Ren does not expressly teaches … photometric loss associated with a multi-camera rig of an ego vehicle … loss associated with the multi-camera rig of the ego vehicle. However, Stein teaches photometric loss associated with a multi-camera rig of an ego vehicle … loss associated with the multi-camera rig of the ego vehicle ([0074] FIG. 9 … a multi-modal loss function application engine 950 is configured to supply training data 930 as input to DNN 902. Training data 930 may include various sequences of image frames captured by one or more vehicle-mounted cameras; [0104] A monocular training system from may perform this operation with five different alternative frames to calculate the photometric loss; [0080] the multi-modal loss function application engine 1050 includes four distinct loss function training engines: a photogrammetric loss function training engine 1004, a predicted-image photogrammetric loss function training engine 1006; [0105] The multi-camera technique incorporates two additional frames in addition to the current frame and two previous-in-time frames to the current frame, similar to the five-frame implementation, but exchanges the future frames with frames taken from different cameras, such as the Front Corner Left and Front Corner Right cameras (e.g., camera 2200A and 2200C in FIG. 22); [0019] FIG. 15 illustrates a multiple-camera array on a vehicle). It would have been obvious before the effective filing date of the claimed invention to one having ordinary skill in the art to modify the device and method of combination of GU and Ren to incorporate the step/system of estimating photometric loss associated with multiple-camera array on an autonomous vehicle taught by Stein. The suggestion/motivation for doing so would have been to improve the accuracy of 3D information on the scene surrounding autonomous vehicles ([0076] multiple different loss functions are used to determine accuracy of the output of the DNN 902; [0103] the neural network converges to output accurate 3D information on the scene; [0106] In the improved multi-camera process, different cameras are synchronized, in time, with each other ... To fully use the equations, we need to accurately determine the RT the cameras (e.g., stereo calibration); [0115] There are other advantages to using stereo information from the surround cameras. For example, it may be more accurate at gauging the 3D shape of objects at a distance because of the relatively wide baseline between the surround cameras when comparted to a single camera). Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predicted results. Therefore, it would have been obvious to combine GU and Ren with Stein to obtain the invention as specified in claim 1. Regarding claim 2, the combination of GU and Ren with Stein teaches all the limitations of claim 1 above. Stein teaches further comprising: capturing images of the scene surrounding the ego vehicle using the multi-camera rig of the ego vehicle ([0016] FIG. 12 illustrates the differing outputs from the two neural networks trained via the monocular and surround cameras; [0003] These systems use an array of sensors to continuously observe the vehicle's motion and surroundings; [0114] the neural network is provoked to infer how much the loss between the original five images from the main camera … the loss from the surround cameras differ; [0044] The image sensors 202 may be mounted or affixed to various portions of the vehicle 204, such as a front right location, a front left location, a middle windshield location, a roof location, a rear window location, or the like. Fields of view of some of the images sensors 202 may overlap), in which cameras of the multi-camera rig have a predetermined minimum overlap ([0130] a system may have multiple cameras that have partially or completely overlapping fields of view; [0041] multiple cameras on a vehicle that have overlapping fields of view (FOB) may be used to train the neural network; [0019] FIG. 15 illustrates a multiple-camera array on a vehicle showing predetermined minimum overlap); selecting target images and context images captured by the cameras of the multi-camera rig of the ego vehicle at a same time-step and at different time-steps ([0041] multiple cameras on a vehicle that have overlapping fields of view (FOB) may be used to train the neural network ... the multiple image frames used to train the network may be taken from multiple cameras at one point in time rather than from one camera at multiple points in time; [0078] the left and center images (at time t3) may be processed with a requirement that the gamma-warped images from time t3 are similar photometrically to center image at time t3. A future two pairs of images may be used to set the condition that the gamma inferred from those images is similar, after correcting for camera motion, to the gamma derived using images from times t1 and t2; [0095]); and performing spatial-temporal transformations of the selected target images and the context images to determine the multi-camera photometric loss ([0167] the homography between the image of forward-left field of view 2300A (at time t1) and the image of forward-center field of view 2300B (also at time t1) is derived from the plane normal used for the homography between the image of forward-center field of view 2300A (at time t1) and the image of forward-center field of view (at time t2) and the known position of forward-left camera 2212A and forward-center camera 2212B (external calibration) together with the internal calibration parameters of each camera such as focal length and lens distortion; [0112] input for the inferencing is from a single camera, (e.g., three frames from the main camera), and the surround images are used just for the photometric loss during training). Regarding claim 3, the combination of GU and Ren with Stein teaches all the limitations of claim 2 above. Stein teaches in which performing the spatial-temporal transformations comprises warping target images and the context images captured by different cameras of the multi-camera rig at the different time-steps ([0167] the homography between the image of forward-left field of view 2300A (at time t1) and the image of forward-center field of view 2300B (also at time t1) is derived from the plane normal used for the homography between the image of forward-center field of view 2300A (at time t1) and the image of forward-center field of view (at time t2) and the known position of forward-left camera 2212A and forward-center camera 2212B (external calibration) together with the internal calibration parameters of each camera such as focal length and lens distortion; [0029] aligning pixels between images in sequence (e.g., warping an earlier image to largely match a later image via a homography)) according to a predicted ego-motion of the ego vehicle and known extrinsics of the different cameras ([0078] where multiple cameras and overlapping fields of view are used, the related images from multiple views may be used to achieve geometric loss function training ... the left and center images (at time t3) may be processed with a requirement that the gamma-warped images from time t3 are similar photometrically to center image at time t3; [0081] the images of training data are processed, along with additional available data such as ego-motion corresponding to the images, camera height, epipole, etc., to produce the reference criteria for evaluation of the loss functions). Regarding claim 4, the combination of GU and Ren with Stein teaches all the limitations of claim 2 above. Stein teaches in which performing the spatial-temporal transformations comprises: warping target images and source images captured by a same camera of the multi-camera rig at the different time-steps ([0093] The reference criteria are based on a “future” road structure (e.g., gamma map), which is computed using the DNN. The geometric loss function training engine 1010 uses the “future” ego-motion to warp the “future” road structure to the current road structure 832, or to warp the current road structure 1032 to the “future” road structure using the “future” ego-motion; [0094] the “future” road structure is warped to the current road structure 1032, and a first comparison is made therebetween, and the current road structure 1032 is warped to the “future” road structure, and a second comparison is made therebetween; [0104] A monocular training system from may perform this operation with five different alternative frames to calculate the photometric loss. The five frames are all from the same camera as the current frame; [0038] multiple images from the same camera that were captured at different times are used to train the neural network); and warping the context images and target images between different cameras and captured at the different time-steps ([0167] the homography between the image of forward-left field of view 2300A (at time t1) and the image of forward-center field of view 2300B (also at time t1) is derived from the plane normal used for the homography between the image of forward-center field of view 2300A (at time t1) and the image of forward-center field of view (at time t2) and the known position of forward-left camera 2212A and forward-center camera 2212B (external calibration) together with the internal calibration parameters of each camera such as focal length and lens distortion; [0029] aligning pixels between images in sequence (e.g., warping an earlier image to largely match a later image via a homography)) according to a predicted ego-motion and known extrinsics of the different cameras ([0078] where multiple cameras and overlapping fields of view are used, the related images from multiple views may be used to achieve geometric loss function training ... the left and center images (at time t3) may be processed with a requirement that the gamma-warped images from time t3 are similar photometrically to center image at time t3; [0081] the images of training data are processed, along with additional available data such as ego-motion corresponding to the images, camera height, epipole, etc., to produce the reference criteria for evaluation of the loss functions). Regarding claim 5, the combination of GU and Ren with Stein teaches all the limitations of claim 1 above. Ren teaches further comprising: predicting a transformation from a current frame to a subsequent frame captured by each camera ([0076] the unsupervised learning module 114 may use the estimated depth (Dt) to compensate for a rigid-motion of the object in the input image frames (It−1, It, It+1) in a 3D space. For example, in some embodiments, the unsupervised learning module 114 may include a pose estimator (e.g., a pose estimation network) 302 and a projection and warping image transformer 304) of the multi-camera rig ([0043] The multi-frame/multi-image inputs may be generated from … different sources (e.g., images with different perspectives or different field-of-views from a dual-camera or different cameras)); transforming each predicted transformation to a coordinate frame of a canonical camera ([0079] the projection and warping image transformer 304 may compensate for the rigid-motion of the object in the 2D image of the input image frames (It−1, It, It+1), and may transform the compensated 2D image into a 3D space (e.g., 3D coordinates) according to the estimated depth (Dt)); constraining a translation vector and a rotation matrix of each camera according to the canonical camera ([0077] the pose estimator 302 may determine a rigid-motion of the object from the target frame (t) to the previous frame (t−1), for example, as Mt→t−1, as well as a rigid-motion of the object from the target frame (t) to the next frame (t+1), for example, as Mt→t+1. Here, M may be a motion vector of the object, and each motion vector M may include a rotation (R) and a translation (T)); and generating the multi-camera PCC loss according to a translation loss and a rotation loss of the translation vector and the rotation matrix of each camera ([0079] the projection and warping image transformer 304 may compensate for the rigid-motion of the object in the 2D image of the input image frames (It−1, It, It+1), and may transform the compensated 2D image into a 3D space (e.g., 3D coordinates) according to the estimated depth (Dt); [0080] the unsupervised learning module 114 may calculate an unsupervised photometric loss (Lu) between the input image frames (It−1, It, It+1) according to the 3D coordinates and the rigid-motion compensation … the unsupervised learning module 114 may calculate the unsupervised photometric loss (Lu) according to an unsupervised loss function shown in equation 3 PNG media_image1.png 26 576 media_image1.png Greyscale ; [0081] Mt→t−1 may correspond to a motion vector of the rigid-motion from the target input image frame (It) to the previous input image frame (It−1), Mt→t+1 may correspond to a motion vector of the rigid-motion from the target input image frame (It) to the next input image frame (It+1); [0077] each motion vector M may include a rotation (R) and a translation (T)). Regarding claim 6, the combination of GU and Ren with Stein teaches all the limitations of claim 1 above. Stein teaches in which training comprises enforcing pose consistency constraints to ensure cameras of the multi-camera rig follow a same rigid body motion to train scale-aware models without any ground-truth depth or ego-motion labels ([0032] Although the ANN may be trained to directly determine the depth or the height of the point … given H and the reference plane, it is possible to compute depth Z and then the residual flow … This is also a reason to pre-warp images with a plane model and provide ego-motion (EM); [0168] Since homographies used for alignment generally use a consistent ground plane, the homography from tracking may be decomposed to give the relative motion, and a new homography may be constructed using this motion and the consistent ground plane normal). With respect to claim 9, arguments analogous to those presented for claim 1, are applicable. With respect to claim 10, arguments analogous to those presented for claim 2, are applicable. With respect to claim 11, arguments analogous to those presented for claim 3, are applicable. With respect to claim 12, arguments analogous to those presented for claim 4, are applicable. With respect to claim 13, arguments analogous to those presented for claim 5, are applicable. With respect to claim 14, arguments analogous to those presented for claim 6, are applicable. Claim(s) 7 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over GU et al. (U.S. Publication No. 2021/0049371) (hereafter, "GU") in view of Ren et al (U.S. Publication No. 2022/0391632) (hereafter, "Ren") and further in view of Stein et al. (U.S. Publication No. 2022/0237866) (hereafter, "Stein") and Ferencz et al. (U.S. Publication No. 2023/0341239) (hereafter, "Ferencz"). Regarding claim 7, the combination of GU and Ren with Stein teaches all the limitations of claim 1 above. Stein teaches in which training the depth estimation model and the ego-motion estimation model comprises ([0032] Although the ANN may be trained to directly determine the depth or the height of the point … given H and the reference plane, it is possible to compute depth Z and then the residual flow … This is also a reason to pre-warp images with a plane model and provide ego-motion (EM)) leveraging cross-camera temporal contexts via spatio-temporal photometric constraints ([0167] the homography between the image of forward-left field of view 2300A (at time t1) and the image of forward-center field of view 2300B (also at time t1) is derived from the plane normal used for the homography between the image of forward-center field of view 2300A (at time t1) and the image of forward-center field of view (at time t2) and the known position of forward-left camera 2212A and forward-center camera 2212B (external calibration) together with the internal calibration parameters of each camera such as focal length and lens distortion). Stein does not expressly teach to increase an amount of overlap between cameras of the multi-camera rig using a predicted ego-motion of the ego vehicle. However, Ferencz teaches to increase an amount of overlap between cameras of the multi-camera rig using a predicted ego-motion of the ego vehicle ([0370] as a result of the ego motion of a host vehicle and a frame capture rate of a camera mounted on the host vehicle, a first top view image 3110 may overlap with a second top view image 3120 in an overlap region 3130. The amount of overlap may increase with higher frame capture rates and/or a slower speed of the host vehicle—i.e., a higher density along a road segments of captured frame view images may result in larger overlap regions between corresponding converted top view images). It would have been obvious before the effective filing date of the claimed invention to one having ordinary skill in the art to modify the device and method of combination of Stein to incorporate the step/system of increasing the amount of overlap by using ego motion of an autonomous vehicle taught by Ferencz. The suggestion/motivation for doing so would have been to improve the accuracy of depth ([0365] Due to the relatively low accuracy of GPS positioning (e.g., 10 m accuracy, 5 m accuracy, etc.) the use of ego motion may help refine actual camera positions and improve the accuracy of depth d). Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predicted results. Therefore, it would have been obvious to combine Stein with Ferencz to obtain the invention as specified in claim 7. With respect to claim 15, arguments analogous to those presented for claim 7, are applicable. Claim(s) 8 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over GU et al. (U.S. Publication No. 2021/0049371) (hereafter, "GU") in view of Ren et al (U.S. Publication No. 2022/0391632) (hereafter, "Ren") and further in view of Stein et al. (U.S. Publication No. 2022/0237866) (hereafter, "Stein") and Huang et al. (U.S. Publication No. 2019/0004533) (hereafter, "Huang"). Regarding claim 8, the combination of GU and Ren with Stein teaches all the limitations of claim 1 above. GU teaches the ego-motion estimation model ([0136] The pose graph output from the pose graph construction algorithm or the refined pose graph output from the pose graph optimization algorithm may combined with the depth map output from the mapping-net to produce a 3D point cloud 440; [0067]; [0112] the 6 DOF pose output from the tracking-net is compared with the 6 DOF pose calculated by the loss functions and the mapping net is trained to minimise this error via backpropagation). GU does not expressly teach further comprising: predicting a 360° point cloud of a scene surrounding the ego vehicle according to the trained multi-camera depth estimation model; and planning a trajectory of the ego vehicle according to the 360° point cloud of the scene surrounding the ego vehicle. However, Huang teaches further comprising: predicting a 360° point cloud of a scene surrounding the ego vehicle ([0027] Machine learning (deep learning) techniques are used to combine a low resolution LIDAR unit with a calibrated multi-camera system to realize a functional equivalent of a high resolution LIDAR unit to generate 3-D point clouds. The multi-camera system is designed to output a wide-angle (e.g., 360 degree) mono or stereo color (e.g., red, green, and blue, or RGB) panorama images); [0035] Cameras 211 may include one or more devices to capture images of the environment surrounding the autonomous vehicle) according to the trained multi-camera depth estimation model ([0027] a high resolution 3-D point cloud can be generated from the wide-angle panorama depth map; [0029] the second depth map represents a second point cloud utilized to perceive the driving environment surround the ADV); and planning a trajectory of the ego vehicle according to the 360° point cloud of the scene surrounding the ego vehicle ([0058]Decision module 304/planning module 305 may include a navigation system or functionalities of a navigation system to determine a driving path for the autonomous vehicle ... the navigation system may determine a series of speeds and directional headings to effect movement of the autonomous vehicle along a path that substantially avoids perceived obstacles while generally advancing the autonomous vehicle along a roadway-based path leading to an ultimate destination; [0027] generate 3-D point clouds. The multi-camera system is designed to output a wide-angle (e.g., 360 degree) mono or stereo color (e.g., red, green, and blue, or RGB) panorama images). It would have been obvious before the effective filing date of the claimed invention to one having ordinary skill in the art to modify the device and method of combination of GU to incorporate the step/system of predicting a 360° point cloud of the environment surrounding the autonomous vehicle according to the multi-camera depth estimation and determining a driving path for the autonomous vehicle taught by Huang. The suggestion/motivation for doing so would have been to improve the efficiency of route for autonomous vehicles ([0041] planning system 110 can plan an optimal route and drive vehicle 101, for example, via control system 111, according to the planned route to reach the specified destination safely and efficiently). Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predicted results. Therefore, it would have been obvious to combine GU with Huang to obtain the invention as specified in claim 8. With respect to claim 16, arguments analogous to those presented for claim 8, are applicable. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to DANIEL C. CHANG whose telephone number is (571)270-1277. The examiner can normally be reached Monday-Thursday and Alternate Fridays 8:00-5:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chan S. Park can be reached at (571) 272-7409. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /DANIEL C CHANG/Examiner, Art Unit 2669 /CHAN S PARK/Supervisory Patent Examiner, Art Unit 2669
Read full office action

Prosecution Timeline

Jun 29, 2023
Application Filed
Aug 13, 2025
Non-Final Rejection mailed — §103, §112
Mar 17, 2026
Response after Non-Final Action

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12675896
IMAGE PROCESSING APPARATUS, METHOD, AND PROGRAM
2y 9m to grant Granted Jul 07, 2026
Patent 12670611
AIRCRAFT STEERING ANGLE DETERMINATION
3y 4m to grant Granted Jun 30, 2026
Patent 12657749
IMAGE DEPTH PREDICTION METHOD, ELECTRONIC DEVICE, AND NON-TRANSITORY STORAGE MEDIUM
2y 11m to grant Granted Jun 16, 2026
Patent 12657661
LEAST SIGNIFICANT BIT (LSB) INFORMATION PRESERVED SIGNAL INTERPOLATION WITH LOW BIT RESOLUTION PROCESSORS
2y 9m to grant Granted Jun 16, 2026
Patent 12614285
Image Analysis Method and Apparatus, Computer Device, and Readable Storage Medium
2y 9m to grant Granted Apr 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
88%
Grant Probability
99%
With Interview (+13.9%)
2y 5m (~0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 147 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month