DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 7/17/2025 has been entered.
Response to Amendment
This action is in response to amendments and remarks filed on 11/03/2025. The examiner notes the following adjustments to the claims by the applicant: (i) Claims 1, 2, 7, 8, 9, 11, 18, 19, 21, 22, and 25 are amended; and (ii) Claim 10 is cancelled. Therefore, Claims 1-9, 11 and 17-25 are pending examination, in which Claims 1, 9 and 18 are independent claims.
In light of the instant amendments and arguments:
The objection to Fig. 7 is withdrawn.
Claims 11, 22 and 24 and 25 are rejected under 35 U.S.C. § 112(b), as lacking proper antecedent basis, due to dependence from canceled Claim 10.
Further examination resulted in a new rejection of Claims 1-9, 11 and 17-25 under 35 U.S.C. § 102, as detailed below.
THIS ACTION IS MADE FINAL. Necessitated by amendment.
Response to Arguments
Applicant presents the following arguments regarding the previous office action:
To overcome the 35 U.S.C. § 102 rejection, the applicant has amended, for example, Claim 1 to include the additional underlined limitations: "determining, based at least on the image data, first boundary locations of a boundary associated with the intersection as represented by the images an individual first boundary location of the first boundary locations being associated with a respective image of the images; projecting, based at least on locations associated with the one or more second vehicles when obtaining the image data";
“Applicant respectfully asserts that the Office has not shown that Mohammadabadi teaches or suggests, at least, "projecting, based at least on locations associated with the one or more second vehicles when obtaining the image data, the first boundary locations from the images to second boundary locations of the boundary on the map; [and] determining, based at least on a grouping of the second boundary locations on the map as being associated with the boundary, a final boundary location of the boundary on the map," as amended claim 1 recites.”;
“As shown, Mohammadabadi describes determining bounding shapes within images that represent intersection coverage maps. Id., paras. [0025] and [0047]. Additionally, as discussed above, the Office cites the left sides of the bounding shapes in Mohammadabadi as allegedly teaching the "second boundary locations" of independent claim 1. However, the bounding shapes in Mohammadabadi indicate the locations of the intersection boundaries within the images. The Office has not shown that Mohammadabadi teaches or suggest "second boundary locations of the boundary on [a] map." For instance, in Mohammadabadi, the left sides of the bounding shapes and/or any other sides of the bounding shapes again indicate locations in the images, but not a map.”;
“the Office has not shown that Mohammadabadi teaches or suggest "grouping" multiple locations of a boundary of the intersection on a map. Consequently, the Office has not shown that Mohammadabadi teaches or suggests "projecting, based at least on locations associated with the one or more second vehicles when obtaining the image data, the first boundary locations from the images to second boundary locations of the boundary on the map; [and] determining, based at least on a grouping of the second boundary locations on the map as being associated with the boundary, a final boundary location of the boundary on the map," as amended claim 1 recites. Additionally, and for similar reasons, the Office has not shown that Mohammadabadi teaches or suggest, "updating, based at least on the final boundary location, the map to indicate an intersection location associated with the intersection," as amended claim 1 recites.”.
Applicant's arguments A., B., C. and D. appear to be directed to the instantly amended subject matter. Accordingly, they have been addressed in the rejections below.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 11, 22 and 24-25 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claims 11, 22 and 24-25 depend from canceled Claim 10, and thus lack sufficient antecedent basis. [For the purposes of examination, the examiner treated Claims 11, 22 and 24-25 as depending from Claim 9.]
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1, 9 and 11 and 17-25 are rejected under 35 U.S.C. §102 as being unpatentable over Mohammadabadi et al. (US 2024/0101118 A, henceforth Mohammadabadi).
Regarding Claim 1, Mohammadabadi explicitly recites the limitations: causing a first vehicle to navigate through an intersection {perception processing layer: “The decoded output(s) 504 may be used to perform one or more operations by a control component(s) 518 of the vehicle 700…a control layer may use the information for determining controls when approaching, navigating through, and/or exiting the intersection(s) (e.g., based on attributes such as wait conditions, size of the intersection, distance to the intersection, etc.).”, ¶[0078]} based at least on map data corresponding to a map {intersection coverage maps 110 and 126, Fig. 1}, wherein the map is generated {generation of coverage maps: “The labeled bounding boxes and the semantic information may then be used by a ground truth encoder to generate intersection locations, coverage maps, confidence values, attributes, distances, distance coverage maps, and/or other information corresponding to the intersection as determined from the annotations. Locations of intersection bounding shapes may be encoded to generate one or more intersection coverage maps.”, ¶[0027] and ¶[0028]}, at least, by: receiving image data obtained using one or more sensors {“Referring now to FIG. 5, FIG. 5 is a data flow diagram illustrating an example process 500 for detecting and classifying intersections using outputs from sensors of a vehicle in real-time or near real-time”, ¶[0072]} of one or more second vehicles {crowd-sourcing data to update map: “The updates to the map information 794 may include updates for the HD map 722…the map information 794 may have resulted from new training and/or experiences represented in data received from any number of vehicles in the environment, and/or based on training performed at a datacenter (e.g., using the server(s) 778 and/or other servers).”, ¶[0191]}, the image data representative of images depicting one or more portions of the intersection {¶[0025] teaches of a multiple image capturing devices, producing a continuous stream of images of an intersection; with respect to Fig. 2A, vehicle sensors capture bounding shape 204 intersection 206, thus providing estimates of the location of all 4 sides of this 4-way intersection; based on the aforementioned crowd-sourcing aspect of data collection (¶[0191]), estimates of all 4 side of the intersection are gathered for each vehicle approaching the intersection}; determining, based at least on the image data, first boundary locations of a boundary associated with the intersection as represented by the images {“using live perception to generate an understanding of each intersection”, ¶[0024], which includes “Receiving image data generated using an image sensor”, B402 in Fig. 4} an individual first boundary location of the first boundary locations being associated with a respective image of the images {“the updated coverage map 328 may include a bottom edge 330 that is the same as the initial intersection coverage map 326”, ¶[0049] and Fig. 5}; projecting, based at least on locations associated with the one or more second vehicles when obtaining the image data, the first boundary locations from the images to second boundary locations of the boundary as represented by the map {crowd-sourcing data, ¶[0191], provides multiple images from multiple vehicle, as will be appreciated by one skilled in the art, which may, for example, be the top edge, bottom edge, left-side or right-side of a bounding shape, such as 204 in Fig. 2A, such that the bottom edge for a first vehicle is associated with the top edge of a second with a vehicle approaching from the opposite direction, i.e., one skilled in the art will appreciate that different portions of the box can be provided by different images provide different vehicles, related to the aforementioned crowd-sourcing of data: “The sensor data may be applied to a neural network…that is trained to identify areas of interests pertaining to intersections…the neural network may be designed to compute data representative of intersection locations, classifications, and/or distances to intersections. For example, the computed outputs may be used to determine locations of a bounding shape(s) corresponding to an intersection(s)…the computed location ¶[0025]}; determining, based at least on a grouping of the second boundary locations as being associated with the boundary {the crowd-sourcing of intersection data, in ¶[0191]), corresponds to each vehicle approaching the intersection providing a top-edge, bottom-edge, left-edge or right-edge of the bounding box 322, Fig. 3B, to a compiled estimate of each intersection side}, a final boundary location of the boundary as represented by the map {final bounding box and intersection coordinates, ¶[0025], are based on multiple images feed into a neural network, ¶[0025], one skilled in the will appreciate that this will lead to a compilation of data for area/location/position of an intersection, for example, that will be distilled by the neural network for a single output for each area/location/position of the intersection: “determining final bounding box locations of final bounding boxes corresponding to the one or more intersections based at least in part on the first data and the second data.”, ¶[0085] and Fig. 5}; and updating, based at least on the final boundary location, the map {crowd-sourcing data to update map: “The updates to the map information 794 may include updates for the HD map 722...the neural networks 792, the updated neural networks 792, and/or the map information 794 may have resulted from new training and/or experiences represented in data received from any number of vehicles in the environment, and/or based on training performed at a datacenter (e.g., using the server(s) 778 and/or other servers).”, ¶[0191]} to indicate an intersection location associated with the intersection {updated intersection maps: “FIGS. 3A-3B, the white portions of intersection coverage maps (e.g., initial intersection coverage maps 304, 326 and update intersection coverage maps 306, 328)”, ¶[0047], resulting from the exchange of map data between the perception layer of the vehicle’s control system - “a perception layer of an autonomous driving software stack may update information about the environment based on the intersection information, a world model manager may update the world model to reflect the location”, ¶[0078], and “The outputs may include information such as….information about objects and status of objects as perceived by the controller(s) 736, etc.”, ¶[0095] - and external server(s) - “The server(s) 778 may transmit, over the network(s) 790 and to the vehicles…updated…map information 794”, ¶[0191]}.
Regarding Claim 2, Mohammadabadi discloses the limitations of Claim 1. In addition, Mohammadabadi explicitly recites the limitations: wherein the map is further generated by: receiving sensor data obtained using one or more second sensors {102, Fig. 1 and B402, Fig. 4} of the one or more second vehicles {“the updated…map information 794 may have resulted from new training and/or experiences represented in data received from any number of vehicles in the environment,”, ¶[0191]}, the sensor data representative of the locations associated with the one or more second vehicles when obtaining the image data {¶[0025] teaches of a multiple image capturing devices, producing a continuous stream of images of an intersection: “sensor data (e.g., image data, LIDAR data, RADAR data, etc.) may be received and/or generated using sensors (e.g., cameras, RADAR sensors, LIDAR sensors, etc.) located or otherwise disposed on an autonomous or semi-autonomous vehicle…. identify areas of interests pertaining to intersections (e.g., raised pavement markers, rumble strips, colored lane dividers, sidewalks, cross-walks, turn-offs, etc.) represented by the sensor data”}, and determining the locations associated with the one or more second vehicles based at least on the sensor data {crowd-sourcing data to update map: “The updates to the map information 794 may include updates for the HD map 722…the map information 794 may have resulted from new training and/or experiences represented in data received from any number of vehicles in the environment, and/or based on training performed at a datacenter (e.g., using the server(s) 778 and/or other servers).”, ¶[0191]}.
Regarding Claim 3, Mohammadabadi discloses the limitations of Claim 1. In addition, Mohammadabadi explicitly recites the limitations: wherein the determining the first boundary locations of the boundary comprises: determining, based at least on the image data, bounding shapes indicating areas of the images that depict the intersection {322, Fig. 3B; “determining final bounding box locations of final bounding boxes corresponding to the one or more intersections based at least in part on the first data and the second data.”, ¶[0085] and Fig. 5}; and determining, based at least on the bounding shapes, the first boundary locations of the boundary associated with the intersection {the crowd-sourcing of intersection data, in ¶[0191]), corresponds to each vehicle approaching the intersection providing a top-edge, bottom-edge, left-edge or right-edge of the bounding box 322, Fig. 3B, to a compiled estimate of each intersection side, with the compile performed by a neural network: “The sensor data may be applied to a neural network...is trained to identify areas of interests pertaining to intersections (e.g., raised pavement markers, rumble strips, colored lane dividers, sidewalks, cross-walks, turn-offs, etc.) represented by the sensor data… the computed outputs may be used to determine locations of a bounding shape(s) corresponding to an intersection(s),”, ¶[0025]; also, ¶[0049] describes using the bottom edge of bounding box to update the map (i.e., “the updated coverage map 328 may include a bottom edge 330 that is the same as the initial intersection coverage map 326”); also, ¶[0068]}.
Regarding Claim 4, Mohammadabadi discloses the limitations of Claim 1. In addition, Mohammadabadi explicitly recites the limitations: wherein the map is further generated by: receiving second image data obtained using the one or more sensors of the one or more second vehicles {¶[0025] teaches of a multiple image capturing devices, producing a continuous stream of images of an intersection: “sensor data (e.g., image data, LIDAR data, RADAR data, etc.) may be received and/or generated using sensors (e.g., cameras, RADAR sensors, LIDAR sensors, etc.) located or otherwise disposed on an autonomous or semi-autonomous vehicle”}, the second image data representative of one or more second images depicting one or more second portions of the intersection {“The sensor data may be applied to a neural network…the computed outputs may be used to determine locations of a bounding shape(s) corresponding to an intersection(s),”, ¶[0025]}; and determining, based at least on the second image data, one or more third boundary locations of a second boundary associated with the intersection {the crowd-sourcing of intersection data, in ¶[0191]), corresponds to each vehicle providing a top-edge, bottom-edge, left-edge or right-edge of the bounding box 322, Fig. 3B, to a compiled estimate of each intersection side; also see bounding box 322 in Fig. 3B surrounding a 4-way intersection; wherein ¶[0049] describes using the bottom edge of bounding box to update the map (i.e., “the updated coverage map 328 may include a bottom edge 330 that is the same as the initial intersection coverage map 326”), one skilled in the art will appreciate that sides and top of box 322 can be used as the relevant data to update a map}, wherein the updating the map {¶[0191]} to indicate the intersection location associated with the intersection is further based at least on the one or more third boundary locations of the second boundary {final bounding box and intersection coordinates, ¶[0085], are based on multiple images feed into a neural network, ¶[0025]; “determining final bounding box locations of final bounding boxes corresponding to the one or more intersections based at least in part on the first data and the second data.”, ¶[0085] and Fig. 5; “neural network may be designed to compute data representative of intersection locations, classifications, and/or distances to intersections. For example, the computed outputs may be used to determine locations of a bounding shape(s) corresponding to an intersection(s), confidence maps for whether or not pixels correspond to an intersection, distances to the intersection(s), classification and semantic information (e.g., attributes, wait conditions), and/or other information. In some examples, the computed location information (e.g., pixel distance to left edge, right edge, top edge, bottom edge of corresponding bounding box) for an intersection may be represented as a pixel-based coverage map with each pixel corresponding to an intersection as uniformly weighted.”, ¶[0025]}.
Regarding Claim 5, Mohammadabadi discloses the limitations of Claim 1. In addition, Mohammadabadi explicitly recites the limitations: wherein the map is further generated by: determining, based at least on the grouping of the second boundary locations {the crowd-sourcing of intersection data, in ¶[0191]), corresponds to each vehicle providing a top-edge, bottom-edge, left-edge or right-edge of the bounding box 322, Fig. 3B, to a compiled estimate of each intersection side}, a bounding shape associated with the boundary as represented by the map {generation of coverage maps includes relating bounding shapes, location and bounding boxes: “The labeled bounding boxes and the semantic information may then be used by a ground truth encoder to generate intersection locations, coverage maps, confidence values, attributes, distances, distance coverage maps, and/or other information corresponding to the intersection as determined from the annotations. Locations of intersection bounding shapes may be encoded to generate one or more intersection coverage maps.”, ¶[0027] and ¶[0028]}, wherein the determining the final boundary location of the boundary is based at least on the bounding shape {final bounding box and intersection coordinates, ¶[0085], are based on multiple images feed into a neural network, ¶[0025]; “determining final bounding box locations of final bounding boxes corresponding to the one or more intersections based at least in part on the first data and the second data.”, ¶[0085] and Fig. 5; in addition, ¶[0025] describes determining the location of “bounding shape(s) corresponding to an intersection(s)”}.
Regarding Claim 6, Mohammadabadi discloses the limitations of Claim 1. In addition, Mohammadabadi explicitly recites the limitations: wherein the map is further generated by: determining a filtered set of boundary locations of the boundary by removing one or more second boundary locations from the grouping of the second boundary locations {Figs. 3A-3B represent the process of creating a bounding box that contains the field-of-view/FOV limited to an intersection, wherein filtering occurs from the initial bounding box estimate of the FOV, i.e., 304/326, to the final bounding box estimate, i.e., 306/328: “With respect to FIGS. 3A-3B, the white portions of intersection coverage maps (e.g., initial intersection coverage maps 304, 326 and update intersection coverage maps 306, 328) may correspond to confidence values above a threshold (e.g., 0.8, 0.9, 1, etc.) that indicate an intersection—or bounding box corresponding thereto—is associated with the pixels therein.”, ¶[0047]}; wherein the determining the final boundary location of the boundary is based at least on the filtered set of boundary locations {final bounding box and intersection coordinates, ¶[0085], are based on multiple images feed into a neural network, ¶[0025]; “determining final bounding box locations of final bounding boxes corresponding to the one or more intersections based at least in part on the first data and the second data.”, ¶[0085] and Fig. 5; one skilled in the art will appreciate that filtering to obtain the best representative data is a feature of using neural networks to handle large amounts of data}.
Regarding Claim 7, Mohammadabadi discloses the limitations of Claim 1. In addition, Mohammadabadi explicitly recites the limitations: wherein the determining the final boundary location comprises determining the final boundary location {final bounding box and intersection coordinates, ¶[0085], are based on multiple images feed into a neural network, ¶[0025]; “determining final bounding box locations of final bounding boxes corresponding to the one or more intersections based at least in part on the first data and the second data.”, ¶[0085] and Fig. 5} as an average location of the second boundary locations included in the grouping {Figs. 3A-3B represent the process of creating a bounding box that contains the field-of-view/FOV limited to an intersection, wherein filtering occurs from the initial bounding box estimate of the FOV, i.e., 304/326, to the final bounding box estimate, i.e., 306/328: “With respect to FIGS. 3A-3B, the white portions of intersection coverage maps (e.g., initial intersection coverage maps 304, 326 and update intersection coverage maps 306, 328) may correspond to confidence values above a threshold (e.g., 0.8, 0.9, 1, etc.) that indicate an intersection—or bounding box corresponding thereto—is associated with the pixels therein.”, ¶[0047], and averaging is inherent in the Deep neural network calculations used to determine intersection bounding box, which includes a variety of data provided to it: “The sensor data may be applied to a neural network…is trained to identify areas of interests pertaining to intersections…More specifically, the neural network may be designed to compute data representative of intersection locations, classifications, and/or distances to intersections. For example, the computed outputs may be used to determine locations of a bounding shape(s) corresponding to an intersection(s)…the computed location information (e.g., pixel distance to left edge, right edge, top edge, bottom edge of corresponding bounding box) for an intersection may be represented as a pixel-based coverage map with each pixel corresponding to an intersection as uniformly weighted. In addition, the distance (e.g., distance of the vehicle to the bottom edge of the intersection bounding box) may be represented as a pixel-based distance coverage map.”, ¶[0025]}.
Regarding Claim 8, Mohammadabadi discloses the limitations of Claim 1. In addition, Mohammadabadi explicitly recites the limitations: wherein the updating the map {¶[0191]} is further by: receiving sensor data obtained using the one or more second vehicles {the equivalent of edge 330, Fig. 3B, is determined for each crowd-sourced vehicle ¶[0191] approaching the intersection}, the sensor data representative of the locations of the one or more second vehicles {¶[0025] teaches of a multiple image capturing devices, producing a continuous stream of images of an intersection: “sensor data (e.g., image data, LIDAR data, RADAR data, etc.) may be received and/or generated using sensors (e.g., cameras, RADAR sensors, LIDAR sensors, etc.) located or otherwise disposed on an autonomous or semi-autonomous vehicle…. identify areas of interests pertaining to intersections (e.g., raised pavement markers, rumble strips, colored lane dividers, sidewalks, cross-walks, turn-offs, etc.) represented by the sensor data”}, while the one or more second vehicles obtained the image data {perception data used to control vehicle: “ a perception layer of an autonomous driving software stack may update information about the environment based on the intersection information, a world model manager may update the world model to reflect the location, distance, attributes, and/or other information about the intersection(s), a control layer may use the information for determining controls when approaching, navigating through, and/or exiting the intersection(s) (e.g., based on attributes such as wait conditions, size of the intersection, distance to the intersection, etc.).”, ¶[0078]}; and localizing, based at least on the sensor data, the one or more second vehicles with respect to the map {Lidar used for localization and establishing bounding shape location: “ the current systems and methods provide techniques to detect and classify intersections using outputs from sensors (e.g., cameras, RADAR sensors, LIDAR sensors, etc.) of a vehicle in real-time or near real-time.”, ¶[0022], and “The method 400, at block B406, includes associating location information corresponding to a location of the bounding shape with at least one pixel within the bounding shape based at least in part on the first data. For example, bounding box(es) 118A may include location information corresponding to the edges of the bounding box(es) 118A, vertices of the bounding box(es) 118A, etc.”, ¶[0068]}, wherein the projecting the first boundary locations is based at least on the localizing {“determining final bounding box locations of final bounding boxes corresponding to the one or more intersections based at least in part on the first data and the second data.”, ¶[0085] and Fig. 5}.
Regarding Claim 9, Mohammadabadi explicitly recites the limitations: one or more processors {“Now referring to FIG. 4, each block of method 400, described herein, comprises a computing process that may be performed using any combination of hardware, firmware, and/or software. For instance, various functions may be carried out by a processor executing instructions stored in memory. The method 400 may also be embodied as computer-usable instructions stored on computer storage media.”, ¶[0065]; 502, Fig. 5} to: determine one or more vehicle locations of one or more vehicles within an environment {Lidar used for localization and establishing bounding shape location: “ the current systems and methods provide techniques to detect and classify intersections using outputs from sensors (e.g., cameras, RADAR sensors, LIDAR sensors, etc.) of a vehicle in real-time or near real-time.”, ¶[0022], and “in order to accurately classify two or more intersections that may visually overlap in image-space, the machine learning model(s) may be trained to compute, for pixels within the bounding shape, pixel distances corresponding to edges of a corresponding bounding shape such that the bounding shape can be generated”, ¶[0023]}; receive image data obtained using the one or more vehicles {“Referring now to FIG. 5 , FIG. 5 is a data flow diagram illustrating an example process 500 for detecting and classifying intersections using outputs from sensors of a vehicle in real-time or near real-time”, ¶[0072]}, the image data representative of one or more images depicting an intersection within the environment {“receiving image data generated using an image sensor. For example, an instance of sensor data 102 (e.g., an image) may be received and/or generated, where the instance of the sensor data 102 depicts a field of view and/or a sensor field of a sensor of the vehicle 700 that includes one or more intersections.”, ¶[0066]}; determine, based at least on the image data {“using live perception to generate an understanding of each intersection”, ¶[0024], which includes “Receiving image data generated using an image sensor”, B402 in Fig. 4}, first locations of the boundaries of the intersection as represented by the one or more images {determining a bounding box to isolate the portion of an image/sensor data corresponding to an intersection: “The sensor data may be applied to a neural network…that is trained to identify areas of interests pertaining to intersections…More specifically, the neural network may be designed to compute data representative of intersection locations, classifications, and/or distances to intersections. For example, the computed outputs may be used to determine locations of a bounding shape(s) corresponding to an intersection(s)…the computed location information (e.g., pixel distance to left edge, right edge, top edge, bottom edge of corresponding bounding box)”, ¶[0025]}; localize the one or more vehicles with respect to a map based at least on the one or more vehicle locations {Lidar used for localization and establishing bounding shape location, ¶[0022], and “in order to accurately classify two or more intersections that may visually overlap in image-space, the machine learning model(s) may be trained to compute, for pixels within the bounding shape, pixel distances corresponding to edges of a corresponding bounding shape such that the bounding shape can be generated”, ¶[0023]}; project {updating server with map data and receiving updates back from the server: “a perception layer of an autonomous driving software stack may update information about the environment based on the intersection information, a world model manager may update the world model to reflect the location,” ¶[0078], “The outputs may include information such as vehicle velocity, speed, time, map data (e.g., the HD map 722 of FIG. 7C), location data (e.g., the vehicle's 700 location, such as on a map), direction, location of other vehicles (e.g., an occupancy grid), information about objects and status of objects as perceived by the controller(s) 736, etc.”, ¶[0095]}, using the one or more vehicles as localized with respect to the map, the first locations of the boundaries of the intersection as represented by the one or more images to second locations of the boundaries of the intersection on the map {the crowd-sourcing of intersection data, in ¶[0191]), corresponds to each vehicle approaching the intersection providing a top-edge, bottom-edge, left-edge or right-edge of the bounding box 322, Fig. 3B, to a compiled estimate of each intersection side, with the compile performed by a neural network: “The sensor data may be applied to a neural network...is trained to identify areas of interests pertaining to intersections (e.g., raised pavement markers, rumble strips, colored lane dividers, sidewalks, cross-walks, turn-offs, etc.) represented by the sensor data… the computed outputs may be used to determine locations of a bounding shape(s) corresponding to an intersection(s),”, ¶[0025]; also, ¶[0049] describes using the bottom edge of bounding box to update the map (i.e., “the updated coverage map 328 may include a bottom edge 330 that is the same as the initial intersection coverage map 326”); also, ¶[0068]}; generate, by at least connecting the second boundary locations on the map, a bounding shape indicating an intersection location on the map {determining a final bounding box: “determining final bounding box locations of final bounding boxes corresponding to the one or more intersections based at least in part on the first data and the second data.”, ¶[0085] and Fig. 5; one skilled in the art will appreciate that an overall representation of the bounding area of a four-way intersection begin by first combining the boundary of a first intersection side with a second intersection side}; and update the map {crowd-sourcing data to update map: “The updates to the map information 794 may include updates for the HD map 722...the neural networks 792, the updated neural networks 792, and/or the map information 794 may have resulted from new training and/or experiences represented in data received from any number of vehicles in the environment, and/or based on training performed at a datacenter (e.g., using the server(s) 778 and/or other servers).”, ¶[0191]} to include the boundary shape indicating the intersection location associated with the intersection {determining a final bounding box: “determining final bounding box locations of final bounding boxes corresponding to the one or more intersections based at least in part on the first data and the second data.”, ¶[0085] and Fig. 5}.
Regarding Claim 11, Mohammadabadi discloses the limitations of Claim {¶[0025] teaches of a multiple image capturing devices, producing a continuous stream of images of an intersection; with respect to Fig. 2A, vehicle sensors capture bounding shape 204 intersection 206, thus providing estimates of the location of all 4 sides of this 4-way intersection; based on the aforementioned crowd-sourcing aspect of data collection (¶[0191]), estimates of all 4 side of the intersection are gathered for each vehicle approaching the intersection}; and determining, based at least on the one or more second bounding shapes, the first locations of the boundaries of the intersection {the crowd-sourcing of intersection data, in ¶[0191]), corresponds to each vehicle approaching the intersection providing a top-edge, bottom-edge, left-edge or right-edge of the bounding box 322, Fig. 3B, to a compiled estimate of each intersection side, with the compile performed by a neural network: “The sensor data may be applied to a neural network...is trained to identify areas of interests pertaining to intersections (e.g., raised pavement markers, rumble strips, colored lane dividers, sidewalks, cross-walks, turn-offs, etc.) represented by the sensor data… the computed outputs may be used to determine locations of a bounding shape(s) corresponding to an intersection(s),”, ¶[0025]; also, ¶[0049] describes using the bottom edge of bounding box to update the map (i.e., “the updated coverage map 328 may include a bottom edge 330 that is the same as the initial intersection coverage map 326”); also, ¶[0068]}.
Regarding Claim 17, Mohammadabadi discloses the limitations of Claim 9. In addition, Mohammadabadi explicitly recites the limitations: a control system for an autonomous or semi-autonomous machine; a perception system {perception processing layer: “The decoded output(s) 504 may be used to perform one or more operations by a control component(s) 518 of the vehicle 700. For example, a perception layer of an autonomous driving software stack may update information about the environment based on the intersection information, a world model manager may update the world model to reflect the location, distance, attributes, and/or other information about the intersection(s), a control layer may use the information for determining controls when approaching, navigating through, and/or exiting the intersection(s) (e.g., based on attributes such as wait conditions, size of the intersection, distance to the intersection, etc.).”, ¶[0078]} for an autonomous or semi-autonomous machine {“the intersection becomes critical to safe and effective autonomous and/or semi-autonomous driving”, ¶[0003]}; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for performing conversational AI operations; a system for generating synthetic data; a system incorporating one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.
Regarding Claim 18, Mohammadabadi explicitly recites the limitations: a processor {“Now referring to FIG. 4, each block of method 400, described herein, comprises a computing process that may be performed using any combination of hardware, firmware, and/or software. For instance, various functions may be carried out by a processor executing instructions stored in memory. The method 400 may also be embodied as computer-usable instructions stored on computer storage media.”, ¶[0065]; 502, Fig. 5} comprising: processing circuitry to: cause a first vehicle to perform one or more operations {perception processing layer: “The decoded output(s) 504 may be used to perform one or more operations by a control component(s) 518 of the vehicle 700…a control layer may use the information for determining controls when approaching, navigating through, and/or exiting the intersection(s) (e.g., based on attributes such as wait conditions, size of the intersection, distance to the intersection, etc.).”, ¶[0078]} based at least on map data corresponding to a map {intersection coverage maps 110 and 126, Fig. 1}, wherein the map is generated {generation of coverage maps: “The labeled bounding boxes and the semantic information may then be used by a ground truth encoder to generate intersection locations, coverage maps, confidence values, attributes, distances, distance coverage maps, and/or other information corresponding to the intersection as determined from the annotations. Locations of intersection bounding shapes may be encoded to generate one or more intersection coverage maps.”, ¶[0027] and ¶[0028]}, at least, by: determining, based at least on first image data {“using live perception to generate an understanding of each intersection”, ¶[0024], which includes “Receiving image data generated using an image sensor”, B402 in Fig. 4} obtained using a second vehicle {crowd-sourcing data to update map: “The updates to the map information 794 may include updates for the HD map 722…the map information 794 may have resulted from new training and/or experiences represented in data received from any number of vehicles in the environment, and/or based on training performed at a datacenter (e.g., using the server(s) 778 and/or other servers).”, ¶[0191]}, a first location of a boundary of an intersection as represented by the first image data {“Referring now to FIG. 5, FIG. 5 is a data flow diagram illustrating an example process 500 for detecting and classifying intersections using outputs from sensors of a vehicle in real-time or near real-time”, ¶[0072]; one skilled in the art will appreciate that “first image data” can be obtained from any vehicle approaching the intersection from any direction; see also ¶[0025]}; determining, based at least on the first location of the boundary, a first corresponding location of the boundary on the map {“the updated coverage map 328 may include a bottom edge 330 that is the same as the initial intersection coverage map 326”, ¶[0049] and Fig. 5}; determining, based at least on second image data obtained using a third vehicle, a second location of the boundary of the intersection as represented by the second image data {crowd-sourcing data, ¶[0191], provides multiple images from multiple vehicle, as will be appreciated by one skilled in the art, which may, for example, be the top edge, bottom edge, left-side or right-side of a bounding shape, such as 204 in Fig. 2A, such that the bottom edge for a first vehicle is associated with the top edge of a second with a vehicle approaching from the opposite direction, i.e., one skilled in the art will appreciate that different portions of the box can be provided by different images provide different vehicles, related to the aforementioned crowd-sourcing of data: “The sensor data may be applied to a neural network…that is trained to identify areas of interests pertaining to intersections…the neural network may be designed to compute data representative of intersection locations, classifications, and/or distances to intersections. For example, the computed outputs may be used to determine locations of a bounding shape(s) corresponding to an intersection(s)…the computed location ¶[0025]}; determining, based at least on the second location of the boundary, a second corresponding location of the boundary on the map {the equivalent of edge 330, Fig. 3B, is determined for each crowd-sourced vehicle ¶[0191] approaching the intersection, and ¶[0025] teaches of a multiple image capturing devices, producing a continuous stream of images of an intersection; with respect to Fig. 2A, vehicle sensors capture bounding shape 204 intersection 206, thus providing estimates of the location of all 4 sides of this 4-way intersection; based on the aforementioned crowd-sourcing aspect of data collection (¶[0191]), estimates of all 4 side of the intersection are gathered for each vehicle approaching the intersection}; determining, based at least on second image data obtained using a third vehicle {crowd-sourcing data to update map: (i.e., “data received from any number of vehicles in the environment”, ¶[0191]}, a second location of the boundary of the intersection as represented by the second image data {the crowd-sourcing of intersection data, in ¶[0191]), corresponds to each vehicle approaching the intersection providing a top-edge, bottom-edge, left-edge or right-edge of the bounding box 322, Fig. 3B, to a compiled estimate of each intersection side, with the compile performed by a neural network (¶[0025] and¶[0068]); wherein ¶[0049] describes using the bottom edge of bounding box to update the map (i.e., “the updated coverage map 328 may include a bottom edge 330 that is the same as the initial intersection coverage map 326”)}; determining, based at least on the first corresponding location on the map and the second corresponding location on the map, a final location of the boundary of the intersection as represented by the map {“bounding box(es) 118A may include location information corresponding to the edges of the bounding box(es) 118A, vertices of the bounding box(es) 118A, etc.”, ¶[0068] and Fig. 1; one skilled in the art will appreciate that crowd-sourced vehicle entering the intersection from different directions, each providing the aforementioned “bottom” edges of bounding box 118A, provide the data for a multi-sided representation of the intersection, for example, in rectangular form}; and updating the map {updated intersection maps: “FIGS. 3A-3B, the white portions of intersection coverage maps (e.g., initial intersection coverage maps 304, 326 and update intersection coverage maps 306, 328)”, ¶[0047], resulting from the exchange of map data between the perception layer of the vehicle’s control system - “a perception layer of an autonomous driving software stack may update information about the environment based on the intersection information, a world model manager may update the world model to reflect the location”, ¶[0078], and “The outputs may include information such as….information about objects and status of objects as perceived by the controller(s) 736, etc.”, ¶[0095] - and external server(s) - “The server(s) 778 may transmit, over the network(s) 790 and to the vehicles…updated…map information 794”, ¶[0191]} to indicate the final location of the intersection {bounding box 322 in Fig. 3B}.
Regarding Claim 19, Mohammadabadi discloses the limitations of Claim 18. In addition, Mohammadabadi explicitly recites the limitations: wherein the determining the first corresponding location comprises projecting {updated intersection maps: “FIGS. 3A-3B, the white portions of intersection coverage maps (e.g., initial intersection coverage maps 304, 326 and update intersection coverage maps 306, 328)”, ¶[0047], resulting from the exchange of map data between the perception layer of the vehicle’s control system - “a perception layer of an autonomous driving software stack may update information about the environment based on the intersection information, a world model manager may update the world model to reflect the location”, ¶[0078], and “The outputs may include information such as….information about objects and status of objects as perceived by the controller(s) 736, etc.”, ¶[0095] - and external server(s) - “The server(s) 778 may transmit, over the network(s) 790 and to the vehicles…updated…map information 794”, ¶[0191]}, based at least on localizing {“The method 400, at block B406, includes associating location information corresponding to a location of the bounding shape with at least one pixel within the bounding shape based at least in part on the first data. For example, bounding box(es) 118A may include location information corresponding to the edges of the bounding box(es) 118A, vertices of the bounding box(es) 118A, etc.”, ¶[0068]} the second vehicle {Lidar used for localization and establishing bounding shape location: “the current systems and methods provide techniques to detect and classify intersections using outputs from sensors (e.g., cameras, RADAR sensors, LIDAR sensors, etc.) of a vehicle in real-time or near real-time.”, ¶[0022], and “in order to accurately classify two or more intersections that may visually overlap in image-space, the machine learning model(s) may be trained to compute, for pixels within the bounding shape, pixel distances corresponding to edges of a corresponding bounding shape such that the bounding shape can be generated”, ¶[0023]; final bounding box and final intersection coordinate described in ¶[0085]}, the first location to the first corresponding location of the boundary as represented by the map {“The method 400, at block B406, includes associating location information corresponding to a location of the bounding shape…may include location information corresponding to the edges of the bounding box(es) 118A, vertices of the bounding box(es) 118A, etc.”, ¶[0068]; Lidar used for localization and establishing bounding shape location: “ the current systems and methods provide techniques to detect and classify intersections using outputs from sensors (e.g., cameras, RADAR sensors, LIDAR sensors, etc.) of a vehicle in real-time or near real-time.”, ¶[0022]}; and the determining the second corresponding location comprises projecting, based at least on localizing the third vehicle {crowd-sourcing data to update map: “The updates to the map information 794 may include updates for the HD map 722…the map information 794 may have resulted from new training and/or experiences represented in data received from any number of vehicles in the environment, and/or based on training performed at a datacenter (e.g., using the server(s) 778 and/or other servers).”, ¶[0191]}, the second location to the second corresponding location of the boundary as represented by the map {the crowd-sourcing of intersection data, in ¶[0191]), corresponds to each vehicle approaching the intersection providing a top-edge, bottom-edge, left-edge or right-edge of the bounding box 322, Fig. 3B, to a compiled estimate of each intersection side, with the compile performed by a neural network: “The sensor data may be applied to a neural network...is trained to identify areas of interests pertaining to intersections (e.g., raised pavement markers, rumble strips, colored lane dividers, sidewalks, cross-walks, turn-offs, etc.) represented by the sensor data… the computed outputs may be used to determine locations of a bounding shape(s) corresponding to an intersection(s),”, ¶[0025]; also ¶[0068]}.
Regarding Claim 20, Mohammadabadi discloses the limitations of Claim 18. In addition, Mohammadabadi explicitly recites the limitations: wherein the process is comprised in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system {perception processing layer: “The decoded output(s) 504 may be used to perform one or more operations by a control component(s) 518 of the vehicle 700. For example, a perception layer of an autonomous driving software stack may update information about the environment based on the intersection information, a world model manager may update the world model to reflect the location, distance, attributes, and/or other information about the intersection(s), a control layer may use the information for determining controls when approaching, navigating through, and/or exiting the intersection(s) (e.g., based on attributes such as wait conditions, size of the intersection, distance to the intersection, etc.).”, ¶[0078]} for an autonomous or semi-autonomous machine {“the intersection becomes critical to safe and effective autonomous and/or semi-autonomous driving”, ¶[0003]}; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for performing conversational AI operations; a system for generating synthetic data; a system incorporating one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.
Regarding Claim 21, Mohammadabadi discloses the limitations of Claim 18. In addition, Mohammadabadi explicitly recites the limitations: wherein the map is further generated by: determining one or more third boundary locations of one or more second boundaries of the intersection {the crowd-sourcing of intersection data, in ¶[0191]), corresponds to each vehicle providing a top-edge, bottom-edge, left-edge or right-edge of the bounding box 322, Fig. 3B, to a compiled estimate of each intersection side; also see bounding box 322 in Fig. 3B surrounding a 4-way intersection; wherein ¶[0049] describes using the bottom edge of bounding box to update the map (i.e., “the updated coverage map 328 may include a bottom edge 330 that is the same as the initial intersection coverage map 326”), one skilled in the art will appreciate that sides and top of box 322 can be used as the relevant data to update a map}; determining, based at least on connecting the final boundary location with the one or more third boundary locations, a bounding shape associated with the intersection location of the intersection {final bounding box and intersection coordinates, ¶[0085], are based on multiple images feed into a neural network, ¶[0025]; “determining final bounding box locations of final bounding boxes corresponding to the one or more intersections based at least in part on the first data and the second data.”, ¶[0085] and Fig. 5; and ¶[0025]}, wherein the updating the map to indicate the intersection final location of the intersection comprises updating the map to indicate the boundary shape {updated intersection maps: “FIGS. 3A-3B, the white portions of intersection coverage maps (e.g., initial intersection coverage maps 304, 326 and update intersection coverage maps 306, 328)”, ¶[0047], resulting from the exchange of map data between the perception layer of the vehicle’s control system - “a perception layer of an autonomous driving software stack may update information about the environment based on the intersection information, a world model manager may update the world model to reflect the location”, ¶[0078], and “The outputs may include information such as….information about objects and status of objects as perceived by the controller(s) 736, etc.”, ¶[0095] - and external server(s) - “The server(s) 778 may transmit, over the network(s) 790 and to the vehicles…updated…map information 794”, ¶[0191]}.
Regarding Claim 22, Mohammadabadi discloses the limitations of Claim {final bounding box and intersection coordinates, ¶[0085], are based on multiple images feed into a neural network, ¶[0025]: “neural network may be designed to compute data representative of intersection locations, classifications, and/or distances to intersections. For example, the computed outputs may be used to determine locations of a bounding shape(s) corresponding to an intersection(s), confidence maps for whether or not pixels correspond to an intersection, distances to the intersection(s), classification and semantic information (e.g., attributes, wait conditions), and/or other information. In some examples, the computed location information (e.g., pixel distance to left edge, right edge, top edge, bottom edge of corresponding bounding box) for an intersection may be represented as a pixel-based coverage map with each pixel corresponding to an intersection as uniformly weighted.”}, wherein the bounding shape is generated by at least connecting the final locations of the boundaries on the map {determining a final bounding box: “determining final bounding box locations of final bounding boxes corresponding to the one or more intersections based at least in part on the first data and the second data.”, ¶[0085] and Fig. 5}.
Regarding Claim 23, Mohammadabadi discloses the limitations of Claim 22. In addition, Mohammadabadi explicitly recites the limitations: wherein the determination of the final locations of the boundaries of the intersection comprises: determining sets of the second locations that are associated with individual boundaries of the boundaries of the intersection {final bounding box and intersection coordinates, ¶[0085], are based on multiple images feed into a neural network, ¶[0025], one skilled in the will appreciate that this will lead to a compilation of data for area/location/position of an intersection, for example, that will be distilled by the neural network for a single output for each area/location/position of the intersection: “determining final bounding box locations of final bounding boxes corresponding to the one or more intersections based at least in part on the first data and the second data.”, ¶[0085] and Fig. 5}; and averaging the sets of the second locations to determine the final locations of the boundaries of the intersection {Figs. 3A-3B represent the process of creating a bounding box that contains the field-of-view/FOV limited to an intersection, wherein filtering occurs from the initial bounding box estimate of the FOV, i.e., 304/326, to the final bounding box estimate, i.e., 306/328: “With respect to FIGS. 3A-3B, the white portions of intersection coverage maps (e.g., initial intersection coverage maps 304, 326 and update intersection coverage maps 306, 328) may correspond to confidence values above a threshold (e.g., 0.8, 0.9, 1, etc.) that indicate an intersection—or bounding box corresponding thereto—is associated with the pixels therein.”, ¶[0047], and averaging is inherent in the Deep neural network calculations used to determine intersection bounding box, which includes a variety of data provided to it: “The sensor data may be applied to a neural network…that is trained to identify areas of interests pertaining to intersections…the neural network may be designed to compute data representative of intersection locations, classifications, and/or distances to intersections. For example, the computed outputs may be used to determine locations of a bounding shape(s) corresponding to an intersection(s)…the computed location information (e.g., pixel distance to left edge, right edge, top edge, bottom edge of corresponding bounding box) for an intersection may be represented as a pixel-based coverage map with each pixel corresponding to an intersection as uniformly weighted. In addition, the distance (e.g., distance of the vehicle to the bottom edge of the intersection bounding box) may be represented as a pixel-based distance coverage map.”, ¶[0025]}.
Regarding Claim 24, Mohammadabadi discloses the limitations of Claim {322, Fig. 3B}.
Regarding Claim 25, Mohammadabadi discloses the limitations of Claim {322, Fig. 3B} comprises at least: identifying individual second boundary locations of the second boundary locations that are associated with individual boundaries of the boundaries of the intersection {final bounding box and intersection coordinates, ¶[0085], are based on multiple images feed into a neural network, ¶[0025], one skilled in the will appreciate that this will lead to a compilation of data for area/location/position of an intersection, for example, that will be distilled by the neural network for a single output for each area/location/position of the intersection: “determining final bounding box locations of final bounding boxes corresponding to the one or more intersections based at least in part on the first data and the second data.”, ¶[0085] and Fig. 5}; and connecting the individual boundary locations together {one skilled in the art will appreciate that generating a final bounding box is a compilation of individual sides: “determining final bounding box locations of final bounding boxes corresponding to the one or more intersections based at least in part on the first data and the second data.”, ¶[0085] and Fig. 5}.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
US 10,936,902 B1 – Use of the bounding-box technique to identify objects after training a machine learning algorithm.
WO 2018/184963 A2 – Use of a neural network and the bounding-box technique to detect vehicles in the field-of-view of an on-board camera.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to RICHARD EDWIN GEIST whose telephone number is (703)756-5854. The examiner can normally be reached Monday-Friday, 9am-6pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Christian Chace can be reached at (571) 272-4190. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/R.E.G./Examiner, Art Unit 3665
/CHRISTIAN CHACE/ Supervisory Patent Examiner, Art Unit 3665