Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Notice to Applicants
This communication is in response to the Application filed on 3/5/2025.
Claims 88-107 are pending. Claims 1-87 have been cancelled.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 99-102 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 99 recites the limitation, “an associated uncertainty estimate” (line 2). It is unclear if “an associated uncertainty estimate” is referring back to “an associated uncertainty estimate” (line 8 in claim 88), an additional associated uncertainty estimate which is different from “an associated uncertainty estimate” or something else. Clarification/explanation is required.
Claim 99 (Claim 100 & 101) recites the limitation, “the perception outputs” (line 3). It is unclear if “the perception outputs” is referring back to “historical perception outputs” (line 10 in claim 88), “a perception output (line 5 in claim 88) and an additional perception output (line 1-2 in claim 99)” or something else. Clarification/explanation is required.
Claim 100 recites the limitation, “the uncertainty estimate associated” (line 1). There is insufficient antecedent basis for these limitations in the claim. It is unclear if “the uncertainty estimate associated” is referring back to “an associated uncertainty estimate” (line 8 in claim 88) or something else. Clarification/explanation is required.
Claim 100 recites the limitations, “their defined probabilistic uncertainty distributions” (line 4). These limitations lack antecedent basis because claim 100 depends on claim 99 (which does not depend on claim 98), but “defining a probabilistic uncertainty distribution” are cited in claim 98. Claim 99 should depend on claim 98. Clarification/explanation is required.
Claim 101 recites the limitation, “the uncertainty estimate associated” (line 1 & 4-5). There is insufficient antecedent basis for these limitations in the claim. It is unclear if “the uncertainty estimate associated” is referring back to “an associated uncertainty estimate” (line 8 in claim 88) or something else. Clarification/explanation is required.
Claim 101 recites the limitations, “their defined probabilistic uncertainty distributions” (line 4). These limitations lack antecedent basis because claim 101 depends on claim 99 (which does not depend on claim 98), but “defining a probabilistic uncertainty distribution” are cited in claim 98. Claim 99 should depend on claim 98. Clarification/explanation is required.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim 88, 95, 98-102 and 104-107 are rejected under 35 U.S.C. 103 as being unpatentable over Ozog (U.S. Publication No. 2020/0005068) in view of Gulino et al. (U.S. Publication No. 2020/0004259) (hereafter, "Gulino").
Regarding claim 88, Ozog teaches a computer-implemented method of perceiving structure in an environment, the method comprising ([0010] a method includes, in response to acquiring sensor data from at least one sensor representing a surrounding environment of a robotic device, extracting a feature representation of an observed line feature in the sensor data by providing a probability distribution that is defined based, at least in part, on feature parameters that overparameterize the observed line feature. The method includes converting the feature representation of the observed line feature into reduced parameters to avoid the feature parameters overparameterizing the observed line feature): receiving at least one structure observation input pertaining to the environment ([0028] the sensor module 220 generally includes instructions that function to control the processor 110 to receive data inputs from one or more sensors of the vehicle 100 in the form of the sensor data 250. The sensor data 250 is ... perceived information that embodies observations of a surrounding environment of the vehicle 100 and, thus, observations of different objects including dynamic and static objects/surfaces within the surrounding environment; [0017] electronic systems generally represent various features observed in a surrounding environment through the use of lines that approximate the features themselves (e.g., lane markers, poles, etc.) or that approximate aspects of the features (e.g., building corners, curbs, etc.); [0044] At 310, the sensor module 220 acquires from at least one sensor of the vehicle 100 the sensor data 250; [0032] the sensor module 220 ... analyzes the sensor data 250 upon receipt to identify line features); processing the at least one structure observation input in a perception pipeline to compute a perception output ([0046] At 320, the sensor module 220 extracts a feature representation of an observed line feature in the sensor data 250 ... the sensor module 220 includes a line segment detector that employs one or more line segment detection techniques as may be known in the art to identify line features within the sensor data 250 and to then extract the line features; [0026] the vehicle 100 includes a translation system 170 that is implemented to perform methods and other functions ... relating to acquiring observations of line features); determining one or more uncertainty source inputs pertaining to the at least one structure observation input ([0035] the transform module 230 initially samples from the probability distribution to generate intermediate representations of the observed line feature using the reduced parameters. That is, the intermediate representations are provided in a form that conforms with the reduced parameters (i.e., four total parameters with two representing azimuth and elevation and two representing position); [0008] The reduced parameters include an observation uncertainty for the line feature that is based, at least part, on the probability distribution; [0050] the transform module 230 selectively samples from the probability distribution representing the observed line feature. Thus, the transform module 230 may generate two separate intermediate representations of the observed line feature that are provided according to the feature parameters ... the two intermediate representations are selective samples of the line feature that produce two intermediate lines. The transform module 230 then converts the representations into vectors of the reduced parameters before proceeding with the functions of the unscented transform; [0021]); determining for the perception output an associated uncertainty estimate ([0039] the transform module 230 applies an unscented transform to the samples ... the transform module 230 computes an error and a mean over the samples to convert the feature representation of the observed line feature into the reduced parameters and including the observation uncertainty for the observed line feature as a four-by-four covariance matrix; [0052] the transform module 230 generates an electronic output in the form of a mean vector representing the observed lined feature according to the reduced parameters and a full-rank covariance matrix that is four-by-four representing the observation uncertainty) by applying, to the one or more uncertainty source inputs, an uncertainty estimation function ([0041] the transform module 230 produces the observation uncertainty for the observed line feature according to the reduced parameters. Thus, the transform module 230 ... converts the original observation provided using the overparameterization into the reduced parameters as shown using the unscented transform to produce a covariance matrix that is four-by-four and full-rank; [0058] the transform module 230 applies an unscented transform to the intermediate sampled lines to generate a mean vector that represents the line feature 410 using the noted reduced parameters and also a covariance matrix that represents the observation uncertainty of the line feature 410. By producing the representation of the line feature 410 using the reduced parameters the corresponding covariance matrix is full-rank and thus improves subsequent use within estimation frameworks) … executing a robotic decision-making process in dependence on the perception output and ([0019] the translation system may be embedded within a robotic device such as an autonomous vehicle or other machine that includes one or more sensors to perceive aspects surrounding the device; [0077] the autonomous driving module(s) 160 can be in communication to send and/or receive information from the various vehicle systems 140 to control the movement, speed, maneuvering, heading, direction, etc. of the vehicle 100; [0059] assuming that the illustrated line segments 410, 420, 430, and 440 are represented according to the reduced parameters as previously explained, in one estimation framework approach, corresponding mapped line features can be selected from the map 260 and correlated according to, for example, Kalman filtering in order to localize the vehicle 100 within the illustrated environment. In this way, the translation system 170 provides an improved representation for line features that improves integration within subsequent estimation frameworks to extrapolate a probabilistic approach without compensating for overparameterized representations using inefficient processes) the associated uncertainty estimate to plan at least one mobile robot path in dependence on the perception output and the associated uncertainty estimate ([0054] the observation uncertainty … in the full-rank covariance matrix provides for implementing the estimation framework using a fully probabilistic approach without coarsely estimating the correlation using weighting or another approach; [0055] At 360, the transform module 230 provides an electronic output of a detection distribution that results from the sensor fusion … the detection distribution indicates an estimation uncertainty for a correlation between the mapped line feature and the observed line feature that is a function of the observation uncertainty. In either case, the detection distribution is a modeled distribution that provides information about the surrounding environment according to the observed line feature and is provided according to the reduced parameters. Consequently, the translation system 170 … controls one or more aspects of the vehicle 100 according to the detection distribution ... the detection distribution is used within a simultaneous localization and mapping (SLAM) framework to model observed line features and determine a location of the vehicle 100 … the localization and/or mapping provided for by the detection distribution influence operation of the autonomous driving module 160, and by extension the vehicle 100 by factoring into determinations of path selection, vehicle controls).
Ozog does not expressly teach … learned from statistical analysis of historical perception outputs; and.
However, Gulino teaches determining for the perception output an associated uncertainty estimate ([0028] The uncertainty system can determine an uncertainty associated with the object prediction ... the prediction system may calculate two or more predicted trajectories for an object and provide probability distributions for each of the predicted trajectories to the motion planning system) by applying, to the one or more uncertainty source inputs, an uncertainty estimation function learned from statistical analysis of historical perception outputs; and ([0115] One or more machine-learned models trained to generate uncertainty data … may provide a continuous probability of uncertainty to the prediction system 104. Such a model may be able to process the state, history, local, and global information inside of one or more fusing layers; [0029] A tracking component of the perception system may provide an output including state and/or uncertainty data associated with the state of an object. The state information can be used with history data associated with object detections to determine an uncertainty. History information such as one or more previous object detections may be used to determine uncertainty associated with an object detection).
It would have been obvious before the effective filing date of the claimed invention to one having ordinary skill in the art to modify the device and method of Ozog to incorporate the step/system of using trained machine-learning models on history and state data to find object uncertainty taught by Gulino.
The suggestion/motivation for doing so would have been to improve the object prediction and motion planning for robots ([0021] the present disclosure is directed to systems and methods that generate and share uncertainty data between components of an autonomous vehicle computing system to improve object prediction and motion planning for the autonomous vehicle; [0045] The machine-learned models configured for motion planning can use the data indicative of uncertainty to generate improved motion plans for the autonomous vehicle; [0072] Prediction system 104 can utilize the uncertainty data along with the object detection data, state data and classification data in order to generate more accurate and reliable predictions associated with the detected objects.). Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predicted results. Therefore, it would have been obvious to combine Ozog with Gulino to obtain the invention as specified in claim 1.
Regarding claim 95, the combination of Ozog and Gulino teaches all the limitations of claim 88 above. Ozog teaches wherein the perception output is computed based on at least one additional structure observation input ([0026] the vehicle 100 includes a translation system 170 that is implemented to perform methods and other functions as disclosed herein relating to acquiring observations of line features and converting the observations into a representation that is not overparameterized and permits an efficient use within additional application frameworks).
Regarding claim 98, the combination of Ozog and Gulino teaches all the limitations of claim 88 above. Ozog teaches wherein the associated uncertainty estimate is outputted by the uncertainty estimation function in a form of a covariance or other distribution parameter(s) ([0039] the transform module 230 applies an unscented transform to the samples ... the transform module 230 computes an error and a mean over the samples to convert the feature representation of the observed line feature into the reduced parameters and including the observation uncertainty for the observed line feature as a four-by-four covariance matrix; [0052] the transform module 230 generates an electronic output in the form of a mean vector representing the observed lined feature according to the reduced parameters and a full-rank covariance matrix that is four-by-four representing the observation uncertainty) defining a probabilistic uncertainty distribution for the perception output ([0041] the transform module 230 outputs a the full-rank covariance matrix for use in, for example, one or more different estimation frameworks (e.g., Kalman filtering, least squares, factor-graphs, Markov Random fields, etc.) to provide for concise determinations of line position, line correlations with map features, and so on in a probabilistic form; [0034] a covariance matrix (e.g., a six-by-six matrix) representing the probability distribution of the observed line feature is based, at least in part, on the feature parameters), whereby the covariance or other distribution parameter(s) varies in dependence on the one or more uncertainty source inputs ([0039] the transform module 230 computes an error and a mean over the samples to convert the feature representation of the observed line feature into the reduced parameters and including the observation uncertainty for the observed line feature as a four-by-four covariance matrix).
Regarding claim 99, the combination of Ozog and Gulino teaches all the limitations of claim 88 above. Ozog teaches further comprising: receiving at least one additional perception output and an associated uncertainty estimate; and ([0053] At 340, the transform module 230 identifies a mapped line feature from the map 260 that potentially corresponds with the observed line feature; [0054] the transform module 230 applies an estimation framework to perform the sensor fusion that uses Kalman filtering, particle filtering, factor-graphs, Markov random fields, or another suitable framework to estimate the correlation. Whichever particular process is employed, the observation uncertainty from 330 embodied in the full-rank covariance matrix provides for implementing the estimation framework using a fully probabilistic approach without coarsely estimating the correlation using weighting or another approach) computing a fused perception output based on the perception outputs according to their associated uncertainty estimates ([0054] At 350, the transform module 230 performs sensor fusion between the observed line feature and the mapped line feature … the observation uncertainty from 330 embodied in the full-rank covariance matrix provides for implementing the estimation framework using a fully probabilistic approach; [0055] At 360, the transform module 230 provides an electronic output of a detection distribution that results from the sensor fusion as discussed at block 350 ... the detection distribution indicates an estimation uncertainty for a correlation between the mapped line feature and the observed line feature that is a function of the observation uncertainty).
Regarding claim 100, the combination of Ozog and Gulino teaches all the limitations of claim 99 above. Ozog teaches wherein the uncertainty estimate associated with the at least one additional perception output is in a form of a distribution parameter(s) and ([0054] At 350, the transform module 230 performs sensor fusion between the observed line feature and the mapped line feature ... the observation uncertainty from 330 embodied in the full-rank covariance matrix provides for implementing the estimation framework using a fully probabilistic approach) the fused perception output is computed by applying a Bayes filter to the perception outputs according to their defined probabilistic uncertainty distributions ([0054] At 350, the transform module 230 performs sensor fusion between the observed line feature and the mapped line feature ... the transform module 230 applies an estimation framework to perform the sensor fusion that uses Kalman filtering ... the observation uncertainty from 330 embodied in the full-rank covariance matrix provides for implementing the estimation framework using a fully probabilistic approach; [0055] At 360, the transform module 230 provides an electronic output of a detection distribution that results from the sensor fusion as discussed at block 350 … the detection distribution indicates an estimation uncertainty for a correlation between the mapped line feature and the observed line feature that is a function of the observation uncertainty; A Kalman filter is a specific, well-known mathematical implementation of a Bayes filter).
Regarding claim 101, the combination of Ozog and Gulino teaches all the limitations of claim 99 above. Ozog teaches wherein the uncertainty estimate associated with the at least one additional perception output is in a form of a distribution parameter(s), and ([0054] At 350, the transform module 230 performs sensor fusion between the observed line feature and the mapped line feature ... the observation uncertainty from 330 embodied in the full-rank covariance matrix provides for implementing the estimation framework using a fully probabilistic approach) the fused perception output is computed by applying a Bayes filter to the perception outputs according to their defined probabilistic uncertainty distributions, and ([0054] At 350, the transform module 230 performs sensor fusion between the observed line feature and the mapped line feature ... the transform module 230 applies an estimation framework to perform the sensor fusion that uses Kalman filtering ... the observation uncertainty from 330 embodied in the full-rank covariance matrix provides for implementing the estimation framework using a fully probabilistic approach; [0055] At 360, the transform module 230 provides an electronic output of a detection distribution that results from the sensor fusion as discussed at block 350 … the detection distribution indicates an estimation uncertainty for a correlation between the mapped line feature and the observed line feature that is a function of the observation uncertainty; A Kalman filter is a specific, well-known mathematical implementation of a Bayes filter) wherein the uncertainty estimate associated with the at least one additional perception output is a covariance and ([0042] the transform module 230 identifies mapped line feature according to one or more known approaches. In either case, the transform module 230 can then use the observed uncertainty distribution of the covariance matrix to perform sensor fusion between the corresponding/mapped line feature and the observed line feature) the fused perception output is computed by applying a Kalman filter to the perception outputs according to their respective covariances ([0054] the transform module 230 performs sensor fusion between the observed line feature and the mapped line feature. In one embodiment, the transform module 230 applies an estimation framework to perform the sensor fusion that uses Kalman filtering; [0041] the transform module 230 outputs a the full-rank covariance matrix for use in, for example, one or more different estimation frameworks (e.g., Kalman filtering, least squares, factor-graphs, Markov Random fields, etc.) to provide for concise determinations of line position, line correlations with map features, and so on in a probabilistic form).
Regarding claim 102, the combination of Ozog and Gulino teaches all the limitations of claim 99 above. Ozog teaches wherein the robotic decision-making process is executed based on the fused perception output ([0019] the translation system may be embedded within a robotic device such as an autonomous vehicle or other machine that includes one or more sensors to perceive aspects surrounding the device; [0055] the transform module 230 provides an electronic output of a detection distribution that results from the sensor fusion ... the translation system 170 … controls one or more aspects of the vehicle 100 according to the detection distribution ... the detection distribution is used within a simultaneous localization and mapping (SLAM) framework to model observed line features and determine a location of the vehicle 100 therefrom ... the detection distribution maps line features within the environment … the localization and/or mapping provided for by the detection distribution influence operation of the autonomous driving module 160, and by extension the vehicle 100 by factoring into determinations of path selection, vehicle controls).
Regarding claim 104, the combination of Ozog and Gulino teaches all the limitations of claim 88 above. Ozog teaches wherein the one or more uncertainty source inputs comprise at least one of: an input derived from the at least one structure observation input, and an input derived from the perception output ([0050] the transform module 230 may generate two separate intermediate representations of the observed line feature that are provided according to the feature parameters (e.g., 6 value representation). Thus, the two intermediate representations are selective samples of the line feature that produce two intermediate lines).
Regarding claim 105, the combination of Ozog and Gulino teaches all the limitations of claim 88 above. Ozog teaches wherein the one or more uncertainty source inputs comprise an uncertainty estimate provided for the at least one structure observation input ([0052] the transform module 230 generates an electronic output in the form of a mean vector representing the observed lined feature according to the reduced parameters and a full-rank covariance matrix that is four-by-four representing the observation uncertainty. Thus, the full-rank covariance matrix quantifies innovation of the observed line feature according to the reduced parameters).
With respect to claim 106, arguments analogous to those presented for claim 88, are applicable.
With respect to claim 107, arguments analogous to those presented for claim 88, are applicable.
Claim 89 is rejected under 35 U.S.C. 103 as being unpatentable over Ozog (U.S. Publication No. 2020/0005068) in view of Gulino et al. (U.S. Publication No. 2020/0004259) (hereafter, "Gulino") and further in view of SORIN et al. (U.S. Publication No. 2010/0163191) (hereafter, "SORIN").
Regarding claim 89, the combination of Ozog and Gulino teaches all the limitations of claim 88 above. Ozog teaches wherein the perception output and the associated uncertainty estimate are used to estimate … for the at least one mobile robot path ([0054] the observation uncertainty … in the full-rank covariance matrix provides for implementing the estimation framework using a fully probabilistic approach without coarsely estimating the correlation using weighting or another approach; [0055] At 360, the transform module 230 provides an electronic output of a detection distribution that results from the sensor fusion … the detection distribution indicates an estimation uncertainty for a correlation between the mapped line feature and the observed line feature that is a function of the observation uncertainty. In either case, the detection distribution is a modeled distribution that provides information about the surrounding environment according to the observed line feature and is provided according to the reduced parameters. Consequently, the translation system 170 … controls one or more aspects of the vehicle 100 according to the detection distribution ... the detection distribution is used within a simultaneous localization and mapping (SLAM) framework to model observed line features and determine a location of the vehicle 100 … the detection distribution maps line features within the environment. In either case, the localization and/or mapping provided for by the detection distribution influence operation of the autonomous driving module 160, and by extension the vehicle 100 by factoring into determinations of path selection, vehicle controls).
Ozog does not expressly teach … a risk of collision.
However, SORIN teaches the perception output and the associated uncertainty estimate are used to estimate a risk of collision for the at least one mobile robot path ([0027] “risk-aware” refers to the inclusion of probabilistic estimates of future trajectories of dynamic obstacles that may appear with some uncertainty in an operating environment of the vehicle or other robot; [0049] The hardware processor at the motion planning module 205 samples the trajectories output at the object tracker 203 to adjust the cost/risk … the motion planning module 205 adjusts a probability of collision, which can include a cost and/or risk value, along each edge of the graph to account for the sample trajectories. The sample trajectories can be mapped to voxels (or other representations of an obstacle) and used to perform collision detection).
It would have been obvious before the effective filing date of the claimed invention to one having ordinary skill in the art to modify the device and method of Ozog to incorporate the step/system of using probabilistic object tracker outputs (perception and uncertainty) to adjust collision probabilities for estimating risk for robot paths taught by SORIN.
The suggestion/motivation for doing so would have been to improve the obstacle detection accuracy for autonomous robots ([0015] Depth perception of drivable surfaces or regions is an important aspect of allowing for and improving operation of autonomous vehicle or driver assistance features. For example, a vehicle must know precisely where obstacles or drivable surfaces are located to navigate safely around objects). Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predicted results. Therefore, it would have been obvious to combine Ozog, Gulino and SORIN to obtain the invention as specified in claim 89.
Claim 90-92 are rejected under 35 U.S.C. 103 as being unpatentable over Ozog (U.S. Publication No. 2020/0005068) in view of Gulino et al. (U.S. Publication No. 2020/0004259) (hereafter, "Gulino") and further in view of Chakravarty et al. (U.S. Publication No. 2019/032559) (hereafter, "Chakravarty").
Regarding claim 90, the combination of Ozog and Gulino teaches all the limitations of claim 88 above. The combination of Ozog and Gulino does not teach wherein the at least one structure observation input comprises a depth estimate.
However, Chakravarty teaches wherein the at least one structure observation input comprises a depth estimate ([0036] a process 200 for training a GAN (generative adversarial network) generator 104 for depth perception ... The system 200 includes stereo camera images 202 and depth maps 220 generated by the GAN generator 104 that are based on the stereo camera images 202 ...The stereo camera images 202 include stereo camera images, wherein each stereo camera image comprises an image pair having a right image 212, 214, 216 and a left image 206, 208, 210 … The images, including image pair A 206, 212, image pair B 208, 214, and image pair C 210, 216 form a stereo image sequence wherein each pair of images are captured by stereo cameras in sequence as the stereo cameras move through an environment).
It would have been obvious before the effective filing date of the claimed invention to one having ordinary skill in the art to modify the device and method of Ozog and Gulino to incorporate the step/system of generating a depth map (which inherently constitutes a depth estimate) from stereo camera image sequences taught by Chakravarty.
The suggestion/motivation for doing so would have been to improve the obstacle detection accuracy for autonomous robots ([0015] Depth perception of drivable surfaces or regions is an important aspect of allowing for and improving operation of autonomous vehicle or driver assistance features. For example, a vehicle must know precisely where obstacles or drivable surfaces are located to navigate safely around objects). Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predicted results. Therefore, it would have been obvious to combine Ozog, Gulino and Chakravarty to obtain the invention as specified in claim 90.
Regarding claim 91, the combination of Ozog, Gulino and Chakravarty teaches all the limitations of claim 90 above. Chakravarty teaches wherein the depth estimate is a stereo depth estimate computed from a stereo image pair ([0036] a process 200 for training a GAN (generative adversarial network) generator 104 for depth perception ... The system 200 includes stereo camera images 202 and depth maps 220 generated by the GAN generator 104 that are based on the stereo camera images 202 … The stereo camera images 202 include stereo camera images, wherein each stereo camera image comprises an image pair having a right image 212, 214, 216 and a left image 206, 208, 210 … The images, including image pair A 206, 212, image pair B 208, 214, and image pair C 210, 216 form a stereo image sequence wherein each pair of images are captured by stereo cameras in sequence as the stereo cameras move through an environment).
Regarding claim 92, the combination of Ozog, Gulino and Chakravarty teaches all the limitations of claim 90 above. Chakravarty teaches wherein the depth estimate is in a form of a depth map ([0036] The system 200 includes stereo camera images 202 and depth maps 220 generated by the GAN generator 104 that are based on the stereo camera images 202 ... The system includes depth maps 220 corresponding to the stereo camera images 202 … depth map A 222 is generated based on image pair A 206, 212; depth map B 224 is generated based on image pair B 208, 214; and depth map C 226 is generated based on image pair C 210, 216).
Claim 93 and 94 are rejected under 35 U.S.C. 103 as being unpatentable over Ozog (U.S. Publication No. 2020/0005068) in view of Gulino et al. (U.S. Publication No. 2020/0004259) (hereafter, "Gulino") and further in view of Chakravarty et al. (U.S. Publication No. 2019/032559) (hereafter, "Chakravarty") and QI et al. (U.S. Publication No. 2019/0147245) (hereafter, "QI").
Regarding claim 93, the combination of Ozog, Gulino and Chakravarty teaches all the limitations of claim 90 above. The combination of Ozog, Gulino and Chakravarty does not teach wherein the perception output is a 3D object localization output computed by applying 3D object localization processing to the depth estimate.
However, QI teaches wherein the perception output is a 3D object localization output computed by applying 3D object localization processing to the depth estimate ([0122] MV3D projects LiDAR point cloud to bird's eye view and trains a region proposal network (RPN) for 3D bounding box proposal; [0123] Some 3D based methods train 3D object classifiers by SVMs on hand-designed geometry features extracted from point cloud and then localize objects using sliding-window search; [0125] The depth data, obtained from LiDAR or indoor depth sensors, is represented as a point cloud in the RGB camera coordinate).
It would have been obvious before the effective filing date of the claimed invention to one having ordinary skill in the art to modify the device and method of Ozog, Gulino and Chakravarty to incorporate the step/system of computing 3D localization from depth/point cloud data to output 3D bounding box proposals taught by QI.
The suggestion/motivation for doing so would have been to improve the efficiency of 3D perception in autonomous robotics ([0002] the necessity to efficiently and accurately analyze the 3D data created by such sensors to perform object detection, classification, and localization becomes of utmost importance; [0047] The methods of implementing three-dimensional perception in an autonomous robotic system herein are more effective and efficient, because extracting a three-dimensional frustum from the three-dimensional depth data using the attention region reduces the amount of data used in the analysis and learning processes). Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predicted results. Therefore, it would have been obvious to combine Ozog, Gulino and Chakravarty with QI to obtain the invention as specified in claim 93.
Regarding claim 94, the combination of Ozog, Gulino and Chakravarty with QI teaches all the limitations of claim 93 above. QI teaches wherein the 3D object localization processing comprises at least one of: orientation detection, size detection, or 3D template fitting ([0122] MV3D projects LiDAR point cloud to bird's eye view and trains a region proposal network (RPN) for 3D bounding box proposal; [0123] Some 3D based methods train 3D object classifiers by SVMs on hand-designed geometry features extracted from point cloud and then localize objects using sliding-window search; [0125] Each object is represented by a class (one among k predefined classes) and an amodal 3D bounding box. The amodal box bounds the complete object even if part of the object is occluded or truncated. The 3D box is parameterized by its size h, w, l, center cx, ci, cz, and orientation θ, φ, ψ relative to a predefined canonical pose for each category).
Claim 96 is rejected under 35 U.S.C. 103 as being unpatentable over Ozog (U.S. Publication No. 2020/0005068) in view of Gulino et al. (U.S. Publication No. 2020/0004259) (hereafter, "Gulino") and further in view of QI et al. (U.S. Publication No. 2019/0147245) (hereafter, "QI").
Regarding claim 96, the combination of Ozog and Gulino teaches all the limitations of claim 95 above. The combination of Ozog and Gulino does not expressly teach wherein the at least one additional structure observation input is a 2D object detection result.
However, QI teaches wherein the at least one additional structure observation input is a 2D object detection result ([0044] the projection matrix is known so that the 3D frustum can be extracted from the 2D image; [0118] to take advantage of mature 2D object detectors. First, extraction is performed of the 3D bounding frustum of an object by extruding 2D bounding boxes from image region detectors; [0127] the systems, methods, and devices herein leverage mature 2D object detector to propose 2D object regions in RGB images as well as to classify objects).
It would have been obvious before the effective filing date of the claimed invention to one having ordinary skill in the art to modify the device and method of Ozog and Gulino to incorporate the step/system of using 2D object detectors to propose 2D object regions and bounding boxes from images taught by QI.
Motivation for this combination has been stated in claim 93.
Claim 97 is rejected under 35 U.S.C. 103 as being unpatentable over Ozog (U.S. Publication No. 2020/0005068) in view of Gulino et al. (U.S. Publication No. 2020/0004259) (hereafter, "Gulino") and further in view of QI et al. (U.S. Publication No. 2019/0147245) (hereafter, "QI") and Ganguli et al. (U.S. Publication No. 2021/0174530) (hereafter, "Ganguli").
Regarding claim 97, the combination of Ozog, Gulino and QI teaches all the limitations of claim 96 above. The combination of Ozog, Gulino and QI does not expressly teach wherein the at least one structure observation input comprises a depth estimate, and wherein the depth estimate is a stereo depth estimate computed from a stereo image pair, and the 2D object detection result is computed by applying 2D object recognition to at least one image of the stereo image pair.
However, Ganguli teaches wherein the at least one structure observation input comprises a depth estimate ([0032] the multiple stereo based depth estimation mechanism 300 comprises a dual stereo pairs 310 including stereo pair 1 310-1 with a wider field of view and stereo pair 2 310-2 with a narrow field of view, a left image fusion unit 320, a right image fusion unit 330, an image registration unit 350, a stereo based disparity estimator 360, and a multi-stereo based depth estimator 370), and wherein the depth estimate is a stereo depth estimate computed from a stereo image pair, and ([0032] To perform depth estimation, two pairs of stereo images, one with a wider field of view and the other with a narrow field of view, are obtained, at 410, via the stereo pair 1 image acquisition unit 310-1 and the second pair 2 image acquisition unit 310-2, respectively. Each pair of such obtained stereo images includes a left image and a right image; [0041] To achieve adaptive usage of available stereo pairs based on need, the mechanism 800 also comprises an object recognition unit 820, an adaptive image acquisition unit 830, left and right image fusion units 840-1 and 840-2, an image registration unit 850, a stereo based disparity estimator 860, and a multi-stereo based depth estimator 870) the 2D object detection result is computed by applying 2D object recognition to at least one image of the stereo image pair ([0040] the mechanism 800 includes a plurality of N (N>2) stereo pairs, e.g., 810-1, . . . , 810-N, one directed to acquiring stereo images with a wider field of view and rest directed to acquiring stereo images with a narrow field of view; [0041] the mechanism 800 also comprises an object recognition unit 820 ... The object recognition unit 820 is provided to detect relevant objects in a scene (e.g., nearby vehicles or other objects) based on object detection models 825. The detection may be performed based on an image acquired via a camera with a wider field of view so that it can provide not only the entire scene but also efficiency in the detection due to its lower resolution).
It would have been obvious before the effective filing date of the claimed invention to one having ordinary skill in the art to modify the device and method of Ozog, Gulino and QI to incorporate the step/system of outputting a stereo depth estimate from stereo image pair by multi-stereo based depth estimator and detecting objects based on a wider field of view image among stereo image pair by object recognition taught by Ganguli.
The suggestion/motivation for doing so would have been to improve the depth estimation in autonomous vehicle ([0005] there is a need to provide an improved solution for estimating the depth information in autonomous driving; [0030] Such an arrangement enables accurate combination of the information from the two stereo pairs. For example, images from the first stereo pair and the second stereo pair can be fused to achieve improved combined resolution which may enable enhanced performance in depth estimation). Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predicted results. Therefore, it would have been obvious to combine Ozog, Gulino and QI with Ganguli to obtain the invention as specified in claim 97.
Claim 103 is rejected under 35 U.S.C. 103 as being unpatentable over Ozog (U.S. Publication No. 2020/0005068) in view of Gulino et al. (U.S. Publication No. 2020/0004259) (hereafter, "Gulino") and further in view of LIM et al. (KR 20150105129) (hereafter, "LIM").
Regarding claim 103, the combination of Ozog and Gulino teaches all the limitations of claim 88 above. Ozog teaches wherein the uncertainty estimation function ([0039] the transform module 230 applies an unscented transform … the transform module 230 computes an error and a mean over the samples to convert the feature representation of the observed line feature into the reduced parameters and including the observation uncertainty for the observed line feature).
Ozog does not expressly teach … is embodied as a lookup table containing one or more uncertainty estimation parameters learned from the statistical analysis.
However, LIM teaches wherein the uncertainty estimation function is embodied as a lookup table containing one or more uncertainty estimation parameters learned from the statistical analysis ([0069] The information storage unit stores information of the learned sub-region, weight values of the selected LUTs, information of the selected LUTs, etc. The weights of the LUT correspond to the reliability of the LUT and can be determined based on the error probability when determining the value of each bin of the LUT. For example, if the error probability is low, it means the LUT has high reliability, so the weights of the LUT can have large values).
It would have been obvious before the effective filing date of the claimed invention to one having ordinary skill in the art to modify the device and method of Ozog, Gulino and QI to incorporate the step/system of determining/learning the weights of LUT (uncertainty estimation parameters) based on the error probability (uncertainty estimation) and storing the weights into LUT taught by Ganguli.
The suggestion/motivation for doing so would have been to improve the efficiency and accuracy of detecting objects ([0004] The problem that the present invention aims to solve is to provide an object detection method and apparatus capable of rapidly detecting objects with a high classification rate). Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predicted results. Therefore, it would have been obvious to combine Ozog, Gulino with LIM to obtain the invention as specified in claim 103.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DANIEL C. CHANG whose telephone number is (571)270-1277. The examiner can normally be reached Monday-Thursday and Alternate Fridays 8:00-5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chan S. Park can be reached at (571) 272-7409. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/DANIEL C CHANG/Examiner, Art Unit 2669 /CHAN S PARK/Supervisory Patent Examiner, Art Unit 2669