DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1 – 4, 9, 12 – 14, 17 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Bagwell (US 11010907 B1; hereafter referred to as Bagwell) in view of Tariq (US 10922574 B1; hereafter referred to as Tariq).
Regarding Claim 1, Bagwell teaches:
A method of controlling a vehicle, the method comprising:
capturing, by one or more processors, using a first sensor, first sensor data including a first object (Bagwell, col. 3, line 8 – 10, “the system may capture sensor data with one or more sensors of the system, such as LIDAR sensors … image sensors, depth sensors. A perception system may process the sensor data to generate multiple bounding boxes for objects represented in the sensor data.”; col. 5, lines 43 – 44, “the sensor data 106 represents objects in an urban environment, such as cars, trucks, roads, buildings, bikes, pedestrians, etc.”);
capturing, by the one or more processors, using the first sensor, second sensor data including a second object (Bagwell, col. 2, line 9 – 11, “sensor data previously captured by a vehicle to annotate the sensor data with ground truth bounding boxes for objects represented in the sensor data”; Bagwell, col. 3, line 8 – 10, “the system may capture sensor data with one or more sensors of the system, such as LIDAR sensors … image sensors, depth sensors. A perception system may process the sensor data to generate multiple bounding boxes for objects represented in the sensor data”; Bagwell, col. 3, line 36 – 38, “a computing device receives annotated data from a user for sensor data previously captured by a vehicle”);
determining, by the one or more processors, a distance to the first object from the vehicle based on the first sensor data (Bagwell, col. 2 lines 47 – 52, “the computing device may determine one or more characteristics associated with the situation in which the sensor data was captured, such … a distance from the vehicle to an object when the sensor data was captured”; Bagwell, col. 27, line 35 – 38, “ the one or more processors to input, into the machine learned model, data indicating at least one of a distance from the vehicle to the object”);
determining, by the one or more processors, whether the distance to the first object is beyond a range of a second sensor (Bagwell, col. 10, line 63 - col. 11, line 7 “the training component 212 may train the machine learned model 220 to output a distance from a vehicle to an object is at a distance, above/below a threshold distance, and/or within a range (e.g., at 60 feet, between 100 and 120 feet, below 30 feet, etc.)”);
when the distance to the first object is determined to be beyond the range of the second sensor, comparing, by the one or more processors, a three-dimensional location based on the distance to the first object and a three-dimensional location based on a distance to the second object (Bagwell, col. 10, line 63 – col. 11, line 10, “the training component 212 may train the machine learned model 220 to output a particular bounding box when a proximity of a vehicle or the object to a road feature is a distance, above/below a threshold distance, and/or within a range (e.g., an intersection, parking lane, etc. is less/more than a threshold distance away)”; Bagwell, col.15, line 30 – 36, “To determine a depth (e.g., distance) between the image sensor and the object contact point, the ray may be unprojected onto a three-dimensional surface mesh, and an intersection point between the ray and the three-dimensional surface mesh may be used as an initial estimate for the projected location of the object contact point”; Bagwell, col. 15, line 61 – 63, “the association component 306 may associate a track with a bounding box and LIDAR data for an object”; Bagwell, col. 16, line 3 – 5, “The comparison may be based on a size, orientation, velocity, and/or position of a detected object and a track”);
While Bagwell teaches determining similarity score of the ground truth bounding boxes and the tracks and uses the bounding box output to track the object and perform operations on controlling the vehicle (Bagwell, col. 16, line 11 – 19, “the association component 306 may determine whether a detected object (e.g., bounding box for the object) is within a threshold distance of a previous position of the object associated with a track, whether the detect object has a threshold amount of velocity to a previous velocity of the object associated with the track, whether the detected object has a threshold amount of similarity in orientation to a previous orientation of the object associated with the track”), it does not explicitly teach:
when the distance to the first object is determined to be beyond the range of the second sensor, determining, by the one or more processors, visual similarity for the first object and the second object;
determining, by the one or more processors, whether the first object is the second object based on the visual similarity and the comparison; and
controlling, by the one or more processors, the vehicle in an autonomous driving mode based on the determination of whether the first object is the second object.
In the same field of endeavor, Tariq teaches:
when the distance to the first object is determined to be beyond the range of the second sensor, determining, by the one or more processors, visual similarity for the first object and the second object (Tariq, col. 2, line 66 – col. 3, line 4, “the computing system may determine, based at least in part on the distances satisfying the threshold distance, that object detections associated with the images is associated with a same object (e.g., a same bicycle), a same class of object, or a different class of object”; Tariq, col. 5, line 60 - 65, “object identifying and/or matching component(s) 116 may determine, based at least in part on the distances satisfying a threshold distance, that an object detection associated with the pixel embeddings of the object and an object detection associated with the bounding box embedding are associated with a same object”);
determining, by the one or more processors, whether the first object is the second object based on the visual similarity and the comparison (Tariq, col. 5, line 55 – col. 6, line 5, “object identifying and/or matching component(s) 116 may determine that distances, in a spatial representation of embedding space, between points associated with pixels of an object and a bounding box of the object may satisfy a threshold distance (e.g., a distance that is close to, or equal to, zero). Accordingly, object identifying and/or matching component(s) 116 may determine, based at least in part on the distances satisfying a threshold distance, that an object detection associated with the pixel embeddings of the object and an object detection associated with the bounding box embedding are associated with a same object… For example, object identifying and/or matching component(s) 116 may determine that an object detection associated with the pixel embeddings of object 122 and an object detection associated with an embedding for bounding box 124 are associated with a same object (e.g., object 122)”; Tariq, col. 6, lines 48 – 50, “Embeddings may be represented by embedding point(s)/location(s) 210 in a spatial representation 212”)); and
controlling, by the one or more processors, the vehicle in an autonomous driving mode based on the determination of whether the first object is the second object (Tariq, col. 3, line 10 – 18, “the computing system of a vehicle, for example, may be able to improve its tracking of objects (e.g., obstacles) and its trajectory and/or route planning, e.g., to control movement of the vehicle to avoid colliding with obstacles…determining whether an object corresponds to another object (e.g., the two objects are the same object) or whether the object corresponds to a bounding box of the object may affect how the vehicle is controlled”).
Bagwell and Tariq are considered analogous art as they are reasonably pertinent to the same field of endeavor of image processing. Therefore, it would have been obvious to one of the ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Bagwell with the invention of Tariq to make the invention that determines whether the first object is the second object based on the visual similarity when the distance to the first object is determined to be beyond the range of the second sensor; whether the first object is the second object based on the visual similarity and the comparison; and controlling the vehicle in an autonomous driving mode based on the determination of whether the first object is the second object; doing so can use the image data and/or other sensor data to efficiently distinguish between objects, accurately identify and track the objects and navigating autonomous vehicles (Tariq, Col. 1, 5 – 20); thus, one of the ordinary skill in the art would have been motivated to combine the references.
Regarding Claim 2, Bagwell in view of Tariq teaches the method of claim 1, wherein the first sensor is a camera, and the second sensor is a LIDAR sensor, such that the visual similarity is determined when the first object is beyond a range of the LIDAR sensor in the first sensor data (Tariq, col. 2, line 44 – 48, “ one or more sensors (e.g., one or more image sensors, one or more lidar sensors, one or more radar sensors, and/or one or more time-of-flight sensors, etc.) of a vehicle (e.g., an autonomous vehicle) may capture images or data of objects”; Tariq, col. 2, line 66 – col. 3, line 4, “the computing system may determine, based at least in part on the distances satisfying the threshold distance, that object detections associated with the images is associated with a same object (e.g., a same bicycle), a same class of object, or a different class of object”; Tariq, col. 5, line 60 - 65, “object identifying and/or matching component(s) 116 may determine, based at least in part on the distances satisfying a threshold distance, that an object detection associated with the pixel embeddings of the object and an object detection associated with the bounding box embedding are associated with a same object”).
Regarding Claim 3, Bagwell in view of Tariq teaches the method of claim 1, wherein the first sensor data and the second sensor data are camera images (Bagwell, col. 3, lines 8 – 10, “the system may capture sensor data with one or more sensors of the system, such as, image sensors”).
Regarding Claim 4, Bagwell in view of Tariq teaches the method of claim 1, wherein determining whether the first object is the second object includes inputting the visual similarity and a result of the comparison into a model (Tariq, col. 16, line 60 – col. 17, line 3, “ the bounding box is associated with the first object, and the second object is at least partially within an area of the bounding box, the machine learning model is further trained to output a second bounding box associated with the second object and having second bounding box parameters”).
Regarding Claim 9, Bagwell in view of Tariq teaches the method of claim 4, wherein a result of the model is a value indicative of similarity of the first object and the second object, and the method further comprises comparing the value to a threshold, and wherein controlling the vehicle is further based on the comparison of the value to the threshold (Tariq, col. 2, line 63, col. 3, line 5, “ the computing system may determine that a distance between embeddings associated with an image may satisfy a threshold distance (e.g., a distance that is close to, or equal to, zero). Furthermore, the computing system may determine, based at least in part on the distances satisfying the threshold distance, that object detections associated with the images is associated with a same object (e.g., a same bicycle), a same class of object, or a different class of object”; Tariq, col. 3, line 14 – 18 – “ determining whether an object corresponds to another object (e.g., the two objects are the same object) or whether the object corresponds to a bounding box of the object may affect how the vehicle is controlled”).
Regarding Claim 12, Bagwell teaches:
A system for controlling a vehicle, the system comprising one or more processors configured to:
capture, using a first sensor, first sensor data including a first object (Bagwell, col. 3, line 8 – 10, “the system may capture sensor data with one or more sensors of the system, such as LIDAR sensors … image sensors, depth sensors. A perception system may process the sensor data to generate multiple bounding boxes for objects represented in the sensor data.”; col. 5, lines 43 – 44, “the sensor data 106 represents objects in an urban environment, such as cars, trucks, roads, buildings, bikes, pedestrians, etc.”);
capture, using the first sensor, second sensor data including a second object (Bagwell, col. 2, line 9 – 11, “sensor data previously captured by a vehicle to annotate the sensor data with ground truth bounding boxes for objects represented in the sensor data”; Bagwell, col. 3, line 8 – 10, “the system may capture sensor data with one or more sensors of the system, such as LIDAR sensors … image sensors, depth sensors. A perception system may process the sensor data to generate multiple bounding boxes for objects represented in the sensor data”; Bagwell, col. 3, line 36 – 38, “a computing device receives annotated data from a user for sensor data previously captured by a vehicle”);
determine, by the one or more processors, a distance to the first object from the vehicle based on the first sensor data (Bagwell, col. 2 lines 47 – 52, “the computing device may determine one or more characteristics associated with the situation in which the sensor data was captured, such … a distance from the vehicle to an object when the sensor data was captured”; Bagwell, col. 27, line 35 – 38, “ the one or more processors to input, into the machine learned model, data indicating at least one of a distance from the vehicle to the object”);
determine, by the one or more processors, whether the distance to the first object is beyond a range of a second sensor (Bagwell, col. 10, line 63 - col. 11, line 7 “the training component 212 may train the machine learned model 220 to output a distance from a vehicle to an object is at a distance, above/below a threshold distance, and/or within a range (e.g., at 60 feet, between 100 and 120 feet, below 30 feet, etc.)”);
when the distance to the first object is determined to be beyond the range of the second sensor, comparing, by the one or more processors, a three-dimensional location based on the distance to the first object and a three-dimensional location based on a distance to the second object (Bagwell, col. 10, line 63 – col. 11, line 10, “the training component 212 may train the machine learned model 220 to output a particular bounding box when a proximity of a vehicle or the object to a road feature is a distance, above/below a threshold distance, and/or within a range (e.g., an intersection, parking lane, etc. is less/more than a threshold distance away)”; Bagwell, col.15, line 30 – 36, “To determine a depth (e.g., distance) between the image sensor and the object contact point, the ray may be unprojected onto a three-dimensional surface mesh, and an intersection point between the ray and the three-dimensional surface mesh may be used as an initial estimate for the projected location of the object contact point”; Bagwell, col. 15, line 61 – 63, “the association component 306 may associate a track with a bounding box and LIDAR data for an object”; Bagwell, col. 16, line 3 – 5, “The comparison may be based on a size, orientation, velocity, and/or position of a detected object and a track”);
While Bagwell teaches determining similarity score of the ground truth bounding boxes and the tracks and uses the bounding box output to track the object and perform operations on controlling the vehicle (Bagwell, col. 16, line 11 – 19, “the association component 306 may determine whether a detected object (e.g., bounding box for the object) is within a threshold distance of a previous position of the object associated with a track, whether the detect object has a threshold amount of velocity to a previous velocity of the object associated with the track, whether the detected object has a threshold amount of similarity in orientation to a previous orientation of the object associated with the track”), it does not explicitly teach:
when the distance to the first object is determined to be beyond the range of the second sensor, determining, by the one or more processors, visual similarity for the first object and the second object;
determine, by the one or more processors, whether the first object is the second object based on the visual similarity and the comparison; and
control, by the one or more processors, the vehicle in an autonomous driving mode based on the determination of whether the first object is the second object.
In the same field of endeavor, Tariq teaches:
when the distance to the first object is determined to be beyond the range of the second sensor, determining, by the one or more processors, visual similarity for the first object and the second object (Tariq, col. 2, line 66 – col. 3, line 4, “the computing system may determine, based at least in part on the distances satisfying the threshold distance, that object detections associated with the images is associated with a same object (e.g., a same bicycle), a same class of object, or a different class of object”; Tariq, col. 5, line 60 - 65, “object identifying and/or matching component(s) 116 may determine, based at least in part on the distances satisfying a threshold distance, that an object detection associated with the pixel embeddings of the object and an object detection associated with the bounding box embedding are associated with a same object”);
determine, by the one or more processors, whether the first object is the second object based on the visual similarity and the comparison (Tariq, col. 5, line 55 – col. 6, line 5, “object identifying and/or matching component(s) 116 may determine that distances, in a spatial representation of embedding space, between points associated with pixels of an object and a bounding box of the object may satisfy a threshold distance (e.g., a distance that is close to, or equal to, zero). Accordingly, object identifying and/or matching component(s) 116 may determine, based at least in part on the distances satisfying a threshold distance, that an object detection associated with the pixel embeddings of the object and an object detection associated with the bounding box embedding are associated with a same object… For example, object identifying and/or matching component(s) 116 may determine that an object detection associated with the pixel embeddings of object 122 and an object detection associated with an embedding for bounding box 124 are associated with a same object (e.g., object 122)”; Tariq, col. 6, lines 48 – 50, “Embeddings may be represented by embedding point(s)/location(s) 210 in a spatial representation 212”)); and
control, by the one or more processors, the vehicle in an autonomous driving mode based on the determination of whether the first object is the second object (Tariq, col. 3, line 10 – 18, “the computing system of a vehicle, for example, may be able to improve its tracking of objects (e.g., obstacles) and its trajectory and/or route planning, e.g., to control movement of the vehicle to avoid colliding with obstacles…determining whether an object corresponds to another object (e.g., the two objects are the same object) or whether the object corresponds to a bounding box of the object may affect how the vehicle is controlled”).
Bagwell and Tariq are considered analogous art as they are reasonably pertinent to the same field of endeavor of image processing. Therefore, it would have been obvious to one of the ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Bagwell with the invention of Tariq to make the invention that determines whether the first object is the second object based on the visual similarity when the distance to the first object is determined to be beyond the range of the second sensor; whether the first object is the second object based on the visual similarity and the comparison; and controlling the vehicle in an autonomous driving mode based on the determination of whether the first object is the second object; doing so can use the image data and/or other sensor data to efficiently distinguish between objects, accurately identify and track the objects and navigating autonomous vehicles (Tariq, Col. 1, 5 – 20); thus, one of the ordinary skill in the art would have been motivated to combine the references.
Regarding Claim 13, Bagwell in view of Tariq teaches the system of claim 12, wherein the first sensor is a camera, and the second sensor is a LIDAR sensor, such that the visual similarity is determined when the first object is beyond a range of the LIDAR sensor in the first sensor data (Tariq, col. 2, line 44 – 48, “ one or more sensors (e.g., one or more image sensors, one or more lidar sensors, one or more radar sensors, and/or one or more time-of-flight sensors, etc.) of a vehicle (e.g., an autonomous vehicle) may capture images or data of objects”; Tariq, col. 2, line 66 – col. 3, line 4, “the computing system may determine, based at least in part on the distances satisfying the threshold distance, that object detections associated with the images is associated with a same object (e.g., a same bicycle), a same class of object, or a different class of object”; Tariq, col. 5, line 60 - 65, “object identifying and/or matching component(s) 116 may determine, based at least in part on the distances satisfying a threshold distance, that an object detection associated with the pixel embeddings of the object and an object detection associated with the bounding box embedding are associated with a same object”).
Regarding Claim 14, Bagwell in view of Tariq teaches the system of claim 12, wherein the one or more processors are further configured to determine whether the first object is the second object includes inputting the visual similarity and a result of the comparison into a model (Tariq, col. 16, line 60 – col. 17, line 3, “ the bounding box is associated with the first object, and the second object is at least partially within an area of the bounding box, the machine learning model is further trained to output a second bounding box associated with the second object and having second bounding box parameters”).
Regarding Claim 17, Bagwell in view of Tariq teaches the system of claim 14, wherein a result of the model is a value indicative of similarity of the first object and the second object, and the method further comprises comparing the value to a threshold, and wherein controlling the vehicle is further based on the comparison of the value to the threshold (Tariq, col. 2, line 63, col. 3, line 5, “ the computing system may determine that a distance between embeddings associated with an image may satisfy a threshold distance (e.g., a distance that is close to, or equal to, zero). Furthermore, the computing system may determine, based at least in part on the distances satisfying the threshold distance, that object detections associated with the images is associated with a same object (e.g., a same bicycle), a same class of object, or a different class of object”; Tariq, col. 3, line 14 – 18 – “ determining whether an object corresponds to another object (e.g., the two objects are the same object) or whether the object corresponds to a bounding box of the object may affect how the vehicle is controlled”).
Regarding Claim 20, Bagwell in view of Tariq teaches the system of claim 12, further comprising the vehicle (Bagwell, col. 1, line 47 – 49, “a system, such as an autonomous vehicle, may implement various techniques that independently (or in cooperation) capture and/or process sensor data”).
Claims 5 – 8 and 15 – 16 are rejected under 35 U.S.C. 103 as being unpatentable over Bagwell (US 11010907 B1; hereafter referred to as Bagwell) in view of Tariq (US 10922574 B1; hereafter referred to as Tariq) further in view of Das et al. (US 20210181758 A1; hereafter referred to as Das).
Regarding Claim 5, Bagwell in view of Tariq teaches the method of claim 4, wherein determining the visual similarity includes inputting the first sensor data and the second sensor data into a second model (Bagwell, col. 26, line 23 – 31, “receiving sensor data from one or more sensors associated with an autonomous vehicle, the sensor data representative of an object in an environment; determining, based at least in part on the sensor data, a first three-dimensional bounding box representing the object and a second three-dimensional bounding box representing the object; inputting the first and second three-dimensional bounding boxes into a machine learned model; receiving, from the machine learned model”), but does not explicitly teach:
wherein determining the visual similarity includes inputting the first sensor data and the second sensor data into a second model different from the model;
In the same field of endeavor, Das teaches:
wherein determining the visual similarity includes inputting the first sensor data and the second sensor data into a second model different from the model (Das, [0054] “the localization component 226, the perception component 228, the planning component 230, the tracking component 232, and/or other components of the system 200 may comprise one or more ML models… the localization component 226, the perception component 228, the planning component 230, and/or the tracking component 232 may each comprise different ML model pipelines”)
Bagwell, Tariq and Das and considered analogous art as they are reasonably pertinent to the same field of endeavor of image processing. Therefore, it would have been obvious to one of the ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Bagwell in view of Tariq with the invention of Das to make the invention that determines the visual similarity by inputting the first sensor data and the second sensor data into a second model different from the model; doing so can improve the efficiency of navigating a vehicle safely and training machine learning models (Das, [0002]); thus, one of the ordinary skill in the art would have been motivated to combine the references.
Regarding Claim 6, Bagwell in view of Tariq further in view of Das teaches the method of claim 5, wherein the model is a convolutional neural network, and the second model is a second multilayer perceptron model (Bagwell, col.22, line 44 – 65, “machine learning algorithms may include, Convolutional Neural Network (CNN)”) … association rule learning algorithms (e.g., perceptron)”; Das [0055] “machine-learning algorithms can include … association rule learning algorithms (e.g., perceptron)”).
Regarding Claim 7, Bagwell in view of Tariq teaches the method of claim 4, wherein the comparing includes inputting the first sensor data and the second sensor data into a second model (Bagwell, col. 19, line 49 – 54, “determine using a machine learned model, an output bounding box. In examples, this may include inputting characteristic data (e.g., for the sensor data) and/or bounding boxes that are associated with a track into the machine learned model and receiving the output bounding box from the machine learned model”), but does not explicitly teach:
wherein the comparing includes inputting the first sensor data and the second sensor data into a second model different from the model.
In the same field of endeavor, Das teaches:
wherein the comparing includes inputting the first sensor data and the second sensor data into a second model different from the model (Das, [0022] “multiple sensor modalities to track objects may require comparing the object detection from each sensor modality to the previous track, whereas the instant techniques comprise comparing the estimated object detection determined by the ML model with the previous track”; Das, [0054] “the localization component 226, the perception component 228, the planning component 230, the tracking component 232, and/or other components of the system 200 may comprise one or more ML models… the localization component 226, the perception component 228, the planning component 230, and/or the tracking component 232 may each comprise different ML model pipelines”; Das, [0059] “a vision pipeline 302 may output an environment representation 308 based at least in part on vision data 310 (e.g., sensor data comprising one or more RGB images)”).
Bagwell, Tariq and Das and considered analogous art as they are reasonably pertinent to the same field of endeavor of image processing. Therefore, it would have been obvious to one of the ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Bagwell in view of Tariq with the invention of Das to make the invention that compares by inputting the first sensor data and the second sensor data into a second model different from the model; doing so can improve the efficiency of navigating a vehicle safely and training machine learning models (Das, [0002]); thus, one of the ordinary skill in the art would have been motivated to combine the references.
Regarding Claim 8, Bagwell in view of Tariq further in view of Das teaches the method of claim 7, wherein the model is a first multilayer perceptron model, and the second model is a second multilayer perceptron model (Bagwell, col.22, line 44 – 65, “machine learning algorithms may include ... association rule learning algorithms (e.g., perceptron)”; Das [0055] “machine-learning algorithms can include … association rule learning algorithms (e.g., perceptron)”).
Regarding Claim 15, Bagwell in view of Tariq teaches the system of claim 14, wherein the one or more processors are further configured to determine the visual similarity includes inputting the first sensor data and the second sensor data into a second model (Bagwell, col. 26, line 23 – 31, “receiving sensor data from one or more sensors associated with an autonomous vehicle, the sensor data representative of an object in an environment; determining, based at least in part on the sensor data, a first three-dimensional bounding box representing the object and a second three-dimensional bounding box representing the object; inputting the first and second three-dimensional bounding boxes into a machine learned model; receiving, from the machine learned model”), but does not explicitly teach:
determine the visual similarity includes inputting the first sensor data and the second sensor data into a second model different from the model;
In the same field of endeavor, Das teaches:
determine the visual similarity includes inputting the first sensor data and the second sensor data into a second model different from the model (Das, [0054] “the localization component 226, the perception component 228, the planning component 230, the tracking component 232, and/or other components of the system 200 may comprise one or more ML models… the localization component 226, the perception component 228, the planning component 230, and/or the tracking component 232 may each comprise different ML model pipelines”)
Bagwell, Tariq and Das and considered analogous art as they are reasonably pertinent to the same field of endeavor of image processing. Therefore, it would have been obvious to one of the ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Bagwell in view of Tariq with the invention of Das to make the invention that determines the visual similarity by inputting the first sensor data and the second sensor data into a second model different from the model; doing so can improve the efficiency of navigating a vehicle safely and training machine learning models (Das, [0002]); thus, one of the ordinary skill in the art would have been motivated to combine the references.
Regarding Claim 16, Bagwell in view of Tariq teaches the system of claim 14, wherein the one or more processors are further configured to perform the comparison by inputting the first sensor data and the second sensor data into a second model (Bagwell, col. 19, line 49 – 54, “determine using a machine learned model, an output bounding box. In examples, this may include inputting characteristic data (e.g., for the sensor data) and/or bounding boxes that are associated with a track into the machine learned model and receiving the output bounding box from the machine learned model”), but does not explicitly teach:
comparison by inputting the first sensor data and the second sensor data into a second model different from the model.
In the same field of endeavor, Das teaches:
comparison by inputting the first sensor data and the second sensor data into a second model different from the model (Das, [0022] “multiple sensor modalities to track objects may require comparing the object detection from each sensor modality to the previous track, whereas the instant techniques comprise comparing the estimated object detection determined by the ML model with the previous track”; Das, [0054] “the localization component 226, the perception component 228, the planning component 230, the tracking component 232, and/or other components of the system 200 may comprise one or more ML models… the localization component 226, the perception component 228, the planning component 230, and/or the tracking component 232 may each comprise different ML model pipelines”; Das, [0059] “a vision pipeline 302 may output an environment representation 308 based at least in part on vision data 310 (e.g., sensor data comprising one or more RGB images)”).
Bagwell, Tariq and Das and considered analogous art as they are reasonably pertinent to the same field of endeavor of image processing. Therefore, it would have been obvious to one of the ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Bagwell in view of Tariq with the invention of Das to make the invention that compares by inputting the first sensor data and the second sensor data into a second model different from the model; doing so can improve the efficiency of navigating a vehicle safely and training machine learning models (Das, [0002]); thus, one of the ordinary skill in the art would have been motivated to combine the references.
Claims 10 – 11 and 18 – 19 are rejected under 35 U.S.C. 103 as being unpatentable over Bagwell (US 11010907 B1; hereafter referred to as Bagwell) in view of Tariq (US 10922574 B1; hereafter referred to as Tariq) further in view of Gesch et al. (US 20200202714 A1; hereafter referred to as Gesch).
Regarding Claim 10, Bagwell in view of Tariq teaches the method of claim 1, further comprising:
generating a track for an object using the first sensor data and the second sensor data, the track identifying changes in the object’s location over time (Bagwell, col. 2, line 15 – 21, “the computing device may process the sensor data with an existing perception system to determine tracks for objects represented in the sensor data and bounding boxes for the objects. A track of an object may represent a current or previous position, velocity, acceleration, orientation, and/or heading of the object over a period of time (e.g., 5 seconds)”);
However, Bagwell in view of Tariq does not explicitly teach:
determining whether the object is a stopped vehicle on a shoulder area using the track, wherein the controlling is further based on the determination of whether the object is a stopped vehicle on the shoulder area.
In the same field of endeavor, Gesch teaches:
determining whether the object is a stopped vehicle on a shoulder area using the track, wherein the controlling is further based on the determination of whether the object is a stopped vehicle on the shoulder area (Gesh, [0019] “The vehicle control system, in one embodiment, detects a stopped vehicle and automatically classifies the stopped vehicle as being either an abandoned vehicle or valid vehicle. A valid vehicle can be, for example, a police car, an ambulance, a road maintenance vehicle, a tow truck, a privately-owned vehicle, etc., that has a pedestrian within a threshold proximity, for example, fifty feet”).
Bagwell, Tariq and Gesch and considered analogous art as they are reasonably pertinent to the same field of endeavor of image processing. Therefore, it would have been obvious to one of the ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Bagwell in view of Tariq with the invention of Gesch to make the invention that determines whether the object is a stopped vehicle on a shoulder area using the track, wherein the controlling is further based on the determination of whether the object is a stopped vehicle on the shoulder area; doing so can control a moving vehicle in response to identifying a stationary vehicle along a roadside and improve safety by avoiding accidents (Gesch, [0002] ); thus, one of the ordinary skill in the art would have been motivated to combine the references.
Regarding Claim 11, Bagwell in view of Tariq further in view of Gesch teaches the method of claim 10, wherein controlling the vehicle includes causing the vehicle to change lanes to move away from the stopped vehicle (Gesch, [0020] “In an autonomous vehicle, when the vehicle control system classifies a stopped vehicle as a valid vehicle, the vehicle control system can: 1) determine a safety score based on contextual information, 2) determine an appropriate modification to the current trajectory and speed of the autonomous vehicle based on the safety score, 3) provide a notification to a user of the autonomous vehicle indicating that the autonomous vehicle is approaching a valid vehicle and about to modify the trajectory and speed of the autonomous vehicle, and 4) execute the modification”; Gesch, [0040] “The recommended action can include, for example, any of a range of trajectory modifications such as slowing down, changing lanes, shifting to an edge of a lane without changing lanes, or different combinations thereof. The drive control module 150 can attempt to execute a trajectory modification based on the safety score”).
Regarding Claim 18, Bagwell in view of Tariq teaches the system of claim 11, further configured to:
generate a track for an object using the first sensor data and the second sensor data, the track identifying changes in the object’s location over time (Bagwell, col. 2, line 15 – 21, “the computing device may process the sensor data with an existing perception system to determine tracks for objects represented in the sensor data and bounding boxes for the objects. A track of an object may represent a current or previous position, velocity, acceleration, orientation, and/or heading of the object over a period of time (e.g., 5 seconds)”);
However, Bagwell in view of Tariq does not explicitly teach:
determine whether the object is a stopped vehicle on a shoulder area using the track, wherein the controlling is further based on the determination of whether the object is a stopped vehicle on the shoulder area.
In the same field of endeavor, Gesch teaches:
determine whether the object is a stopped vehicle on a shoulder area using the track, wherein the controlling is further based on the determination of whether the object is a stopped vehicle on the shoulder area (Gesh, [0019] “The vehicle control system, in one embodiment, detects a stopped vehicle and automatically classifies the stopped vehicle as being either an abandoned vehicle or valid vehicle. A valid vehicle can be, for example, a police car, an ambulance, a road maintenance vehicle, a tow truck, a privately-owned vehicle, etc., that has a pedestrian within a threshold proximity, for example, fifty feet”).
Bagwell, Tariq and Gesch and considered analogous art as they are reasonably pertinent to the same field of endeavor of image processing. Therefore, it would have been obvious to one of the ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Bagwell in view of Tariq with the invention of Gesch to make the invention that determines whether the object is a stopped vehicle on a shoulder area using the track, wherein the controlling is further based on the determination of whether the object is a stopped vehicle on the shoulder area; doing so can control a moving vehicle in response to identifying a stationary vehicle along a roadside and improve safety by avoiding accidents (Gesch, [0002] ); thus, one of the ordinary skill in the art would have been motivated to combine the references.
Regarding Claim 19, Bagwell in view of Tariq further in view of Gesch teaches the system of claim 18, wherein the one or more processors are further configured to control the vehicle by causing the vehicle to change lanes to move away from the stopped vehicle (Gesch, [0020] “In an autonomous vehicle, when the vehicle control system classifies a stopped vehicle as a valid vehicle, the vehicle control system can: 1) determine a safety score based on contextual information, 2) determine an appropriate modification to the current trajectory and speed of the autonomous vehicle based on the safety score, 3) provide a notification to a user of the autonomous vehicle indicating that the autonomous vehicle is approaching a valid vehicle and about to modify the trajectory and speed of the autonomous vehicle, and 4) execute the modification”; Gesch, [0040] “The recommended action can include, for example, any of a range of trajectory modifications such as slowing down, changing lanes, shifting to an edge of a lane without changing lanes, or different combinations thereof. The drive control module 150 can attempt to execute a trajectory modification based on the safety score”).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
US 20220101020 A1 TRACKING OBJECTS USING SENSOR DATA SEGMENTATIONS AND/OR REPRESENTATIONS Techniques are disclosed for tracking objects in sensor data, such as multiple images or multiple LIDAR clouds. The techniques may include comparing segmentations of sensor data such as by, for example, determining a similarity of a first segmentation of first sensor data and a second segmentation of second sensor data. Comparing the similarity may comprise determining a first embedding associated with the first segmentation and a second embedding associated with the second segmentation and determining a distance between the first embedding and the second embedding.
Contact Information
Any inquiry concerning this communication or earlier communications from the examiner should be directed to VAISALI RAO KOPPOLU whose telephone number is (571)270-0273. The examiner can normally be reached Monday - Friday 8:30 - 5.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jennifer Mehmood can be reached at (571) 272-2976. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format.
For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
VAISALI RAO. KOPPOLU
Examiner
Art Unit 2664
/VAISALI RAO KOPPOLU/Examiner of Art Unit 2664