DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Upon further review of the recent submission of the IDS, the claims are rejected at least by Cho et al. (US 2026/0109373).
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claim(s) 1 – 3 and 9 – 20 is/are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Cho et al. (US 2026/0109373).
Regarding independent claim 1, Cho teaches a method (Figures 3A – 3H and 7) comprising:
identifying, based at least on sensor data generated by a plurality of ego-machines in an environment (paragraph 23: the computing system on the ego can acquire sensor data (e.g., optical camera images and LiDAR) as well as map data (e.g., navigation map) of the surroundings of the ego), {1}one or more points{1} (paragraph 77: The set of encodings of the vector space encoding 302 can be generated, determined, or otherwise {1}derived from sensor data{1} (e.g., from cameras and other sensors) acquired by an ego and map data (e.g., the navigation map 235) defining a topology of an environment surrounding the ego; paragraph 62: In operation, as the one or more egos 140 navigate, their sensors collect data and transmit the data to the analytics server 110a, as depicted in the data stream 172) associated with one or more cells of a two-dimensional (2D) representation of the environment (paragraph 78: With the identification, the computing system can find, determine, or otherwise identify at least one point 306A within a set of grid points in a first grid 308 defined over the environment represented by a map 310);
for each cell of at least one of the one or more cells of the 2D representation of the environment, generating {2}an encoded representation{2} of {1}a set of the one or more points{1} that are associated with the cell (paragraph 77: {2}The set of encodings of the vector space encoding 302{2} can be generated, determined, or otherwise {1}derived from sensor data{1} (e.g., from cameras and other sensors) acquired by an ego and map data (e.g., the navigation map 235) defining a topology of an environment surrounding the ego) using a first encoder of one or more neural networks (paragraph 79: The first point predictor unit 310 can correspond to a portion of a machine learning (ML) model (e.g., the lane encoder 260 as discussed herein), and can include a set of weights arranged in accordance with a cross-attention, a self-attention, and a transformer, among others); and
generating, based at least on applying the encoded representations of the set of the one or more points associated with the cell to a second encoder of the one or more neural networks (paragraph 92: the computing system can process the portion of the vector space encoding 302 in accordance with the weights of the second point predictor unit 312 to calculate, determine, or otherwise generate at least one coordinate 332B corresponding to the point 306B. Using the point, the computing system can use the second point predictor unit 312 to calculate, produce, or otherwise determine a first index value 332'B for the point 306B), a representation of one or more consolidated lane lines observed by the plurality of ego-machines (paragraph 95: the computing system can apply an output (e.g., embeddings derived from the first index value 330'8) of the first point predictor unit 310, an output (e.g., embeddings derived from the second index value 332'8) of the second point predictor unit 312, or an output (e.g., embeddings derived from the topology type 334'8) into the point attribute predictor 316; paragraph 96: By applying the point attribute predictor 316, the computing system can generate or determine a set of point attributes 336B for the point 306B. The set of point attributes 336B can identify or include, for example: at least one fork point identifying an index value of a previous point from which the current point 306B forms a fork; at least one merge point identifying an index value of a previous point from which the current point 306B forms a merge; and a set of spline coefficients 338B, among others).
Regarding dependent claim 2, Cho teaches wherein the sensor data represented in the encoded representation of the set of the one or more points includes data obtained by at least one of a LiDAR sensor (paragraph 22: the ego can acquire sensor data (e.g., Light Detection and Ranging (LiDAR) and optical images) and apply an image segmentation model to detect lines along a road surface to recognize lane segments on the road), a RADAR sensor (paragraph 55: a radar 170n and ultrasound sensors 170p may be configured to monitor the distance of the egos 140 to other objects), or a camera for each of the plurality of ego-machines (paragraph 23: the computing system on the ego can acquire sensor data (e.g., optical camera images and LiDAR) as well as map data (e.g., navigation map) of the surroundings of the ego).
Regarding dependent claim 3, Cho teaches wherein the one or more points include points sampled from a polyline positioned within the one or more cells of the 2D representation of the environment (paragraph 25: The computing system can also calculate spline coefficients defining connectivity between the two points to define a corresponding lane segment through the environment).
Regarding dependent claim 9, Cho teaches wherein the first encoder includes a cross-attention layer and a self-attention layer (paragraph 79: The first point predictor unit 310 can correspond to a portion of a machine learning (ML) model (e.g., the lane encoder 260 as discussed herein), and can include a set of weights arranged in accordance with a cross-attention, a self-attention, and a transformer, among others).
Regarding dependent claim 10, Cho teaches for each cell of the at least one of the one or more cells of the 2D representation of the environment, including a representation of a cell position indicating a relative position of the cell, among the one or more cells, in the encoded representation generated using the first encoder (paragraph 85: The computing system may correspond to the ego computing device 141 on an ego 140 or a centralized service such as the analytics server 110a. In addition, the computing system can apply an output (e.g., embeddings derived from the first index value 330'A) of the first point predictor unit 310, an output (e.g., embeddings derived from the second index value 332'A) of the second point predictor unit 312, or an output (e.g., embeddings derived from the topology type 334'A) into the point attribute predictor 316. The portion of the vector space encoding 302 inputted into the point attribute predictor 316 can be the same or can differ from the portion of the vector space encoding 302 inputted into the first point prediction unit 310, the second point prediction unit 312, or the topology predictor unit 314).
Regarding dependent claim 11, Cho teaches wherein the second encoder includes one or more self-attention layers and a multilayer perceptron (paragraph 82: The second point predictor unit 312 can correspond to a portion of a machine learning (ML) model (e.g., the lane encoder 260 as discussed herein), and can include a set of weights arranged in accordance with a cross-attention, a self-attention, and a transformer, among others; paragraph 69: The ML models in the visual component 205 can include, for example, a set of self-regulated networks (RegNets) 210A-C (hereinafter generally referred to as residual networks 210), a set of feature pyramid networks (FPNs) 215A-C (hereinafter generally referred to as feature pyramid networks 215), at least one transformer 220, and at least one video module 225, among others).
Regarding dependent claim 12, Cho teaches applying the representation of the one or more consolidated lane lines as input to a decoder to infer one or more lanes (paragraph 86: By applying the point attribute predictor 316, the computing system can generate or determine a set of point attributes 336A for the point 306A. The set of point attributes 336A can identify or include, for example: at least one fork point identifying an index value of a previous point from which the current point 306A forms a fork; at least one merge point identifying an index value of a previous point from which the current point 306A forms a merge; and a set of spline coefficients 338A, among others).
Regarding dependent claim 13, Cho teaches generating a lane graph based at least on the one or more consolidated lane lines (paragraph 75: The lane instances 265 and the adjacency matrix 270 can collectively be referred to as a graph or a language of lanes).
Regarding dependent claim 14, Cho teaches wherein the method is performed by at least one of: a control system for an autonomous or semi-autonomous machine (paragraphs 74: The lane guidance module 240 can retrieve, obtain, or otherwise identify the navigation map 235 defining the topology surrounding the ego); a perception system for an autonomous or semi-autonomous machine (paragraphs 74: The lane guidance module 240 can retrieve, obtain, or otherwise identify the navigation map 235 defining the topology surrounding the ego); a system for performing simulation operations (paragraph 63: The analytics server 110a may generate a training dataset using data collected from the egos 140 (e.g., camera feed received from the egos 140)); a system for performing real-time streaming (paragraph 29: the analytics server 110a can use the methods discussed herein to train the AI model(s) 110c using data retrieved from the egos 140 ( e.g., by using data streams 172 and 174)); a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; a system for performing digital twin operations; a system for performing deep learning operations (paragraph 69: The visual component 205 can include artificial intelligence (AI) algorithms or machine learning (ML) models to process the sensor data from the set of cameras and other sensors); a system implemented using an edge device (paragraph 28: the network 130 may also include communications over a cellular network, including, for example, a GSM (Global System for Mobile Communications), CDMA (Code Division Multiple Access), or an EDGE (Enhanced Data for Global Evolution) network); a system implemented using a robot (paragraph 34: The egos 140 are not limited to being vehicles and may include robotic devices as well); a system incorporating one or more virtual machines (VMs); a system implemented at least partially in a data center (paragraph 26: The system 100 may include an analytics server 110a, a system database 110b, an administrator computing device 120, egos l40a-b (collectively ego(s) 140), ego computing devices 14la-c (collectively ego computing devices 141), and a server 160); a system for performing light transport simulation (paragraph 63: The analytics server 110a may generate a training dataset using data collected from the egos 140 (e.g., camera feed received from the egos 140)); a system for performing collaborative content creation for 3D assets; a system for generating synthetic data (paragraph 29: the vehicle 140a having the ego computing device 141a may transmit its camera feed to the trained AI model(s) 110c and may generate a graph defining lane segments in the environment (e.g., data stream 174)); or a system implemented at least partially using cloud computing resources (paragraph 33: While the system 100 includes a single analytics server 110a, the system 100 may include any number of computing devices operating in a distributed computing environment, such as a cloud environment).
Regarding independent claim 15, Cho teaches one or more processors comprising one or more processing units (Figures 1A, 1B, and 2) to:
generate, based at least on using a first encoder of one or more neural networks (paragraph 79: The first point predictor unit 310 can correspond to a portion of a machine learning (ML) model (e.g., the lane encoder 260 as discussed herein), and can include a set of weights arranged in accordance with a cross-attention, a self-attention, and a transformer, among others), {2}one or more cell latent representations{2} of {1}a set of sensor data{1} (paragraph 77: {2}The vector space encoding 302 (sometimes herein referred to as a tensor){2} can identify or include a set of encodings (sometimes herein referred as a set of embeddings or feature maps). The set of encodings of the vector space encoding 302 can be generated, determined, or otherwise {1}derived from sensor data{1} (e.g., from cameras and other sensors) acquired by an ego and map data (e.g., the navigation map 235) defining a topology of an environment surrounding the ego) generated by a plurality of ego-machines (paragraph 23: the computing system on the ego can acquire sensor data (e.g., optical camera images and LiDAR) as well as map data (e.g., navigation map) of the surroundings of the ego; paragraph 62: In operation, as the one or more egos 140 navigate, their sensors collect data and transmit the data to the analytics server 110a, as depicted in the data stream 172) and associated with one or more cells of a two-dimensional (2D) representation of a region in an environment (paragraph 78: With the identification, the computing system can find, determine, or otherwise identify at least one point 306A within a set of grid points in a first grid 308 defined over the environment represented by a map 310); and
generate, based at least on applying the one or more cell latent representations of the set of sensor data to a second encoder of the one or more neural networks (paragraph 92: the computing system can process the portion of the vector space encoding 302 in accordance with the weights of the second point predictor unit 312 to calculate, determine, or otherwise generate at least one coordinate 332B corresponding to the point 306B. Using the point, the computing system can use the second point predictor unit 312 to calculate, produce, or otherwise determine a first index value 332'B for the point 306B), a region latent representation of one or more consolidated lane lines in the region of the environment (paragraph 95: the computing system can apply an output (e.g., embeddings derived from the first index value 330'8) of the first point predictor unit 310, an output (e.g., embeddings derived from the second index value 332'8) of the second point predictor unit 312, or an output (e.g., embeddings derived from the topology type 334'8) into the point attribute predictor 316; paragraph 96: By applying the point attribute predictor 316, the computing system can generate or determine a set of point attributes 336B for the point 306B. The set of point attributes 336B can identify or include, for example: at least one fork point identifying an index value of a previous point from which the current point 306B forms a fork; at least one merge point identifying an index value of a previous point from which the current point 306B forms a merge; and a set of spline coefficients 338B, among others).
Regarding dependent claim 16, Cho teaches wherein the one or more processing units are further to: provide the region latent representation, as input, to a decoder to infer lane data; infer, via the decoder, the lane data associated with one or more lanes based at least on the region latent representation of the one or more consolidated lane lines; and generate a lane graph based at least one the inferred one or more lanes (paragraph 86: By applying the point attribute predictor 316, the computing system can generate or determine a set of point attributes 336A for the point 306A. The set of point attributes 336A can identify or include, for example: at least one fork point identifying an index value of a previous point from which the current point 306A forms a fork; at least one merge point identifying an index value of a previous point from which the current point 306A forms a merge; and a set of spline coefficients 338A, among others).
Regarding independent claim 18, Cho teaches a system comprising one or more processing units (Figures 1A, 1B, and 2) to:
for each cell of one or more cells of a 2D representation of an environment (paragraph 78: With the identification, the computing system can find, determine, or otherwise identify at least one point 306A within a set of grid points in a first grid 308 defined over the environment represented by a map 310), generate, using a first encoder of a transformer machine learning model (paragraph 79: The first point predictor unit 310 can correspond to a portion of a machine learning (ML) model (e.g., the lane encoder 260 as discussed herein), and can include a set of weights arranged in accordance with a cross-attention, a self-attention, and a transformer, among others), {2}an encoded representation{2} of a set of one or more points that are associated with the cell and that are identified based at least on {1}sensor data{1} (paragraph 77: {2}The vector space encoding 302 (sometimes herein referred to as a tensor){2} can identify or include a set of encodings (sometimes herein referred as a set of embeddings or feature maps). The set of encodings of the vector space encoding 302 can be generated, determined, or otherwise {1}derived from sensor data{1} (e.g., from cameras and other sensors) acquired by an ego and map data (e.g., the navigation map 235) defining a topology of an environment surrounding the ego) generated by a plurality of ego-machines in the environment (paragraph 23: the computing system on the ego can acquire sensor data (e.g., optical camera images and LiDAR) as well as map data (e.g., navigation map) of the surroundings of the ego; paragraph 62: In operation, as the one or more egos 140 navigate, their sensors collect data and transmit the data to the analytics server 110a, as depicted in the data stream 172); and
generate, based at least on applying the encoded representations of the set of the one or more points associated with the one or more cells to a second encoder of the transformer machine learning model (paragraph 92: the computing system can process the portion of the vector space encoding 302 in accordance with the weights of the second point predictor unit 312 to calculate, determine, or otherwise generate at least one coordinate 332B corresponding to the point 306B. Using the point, the computing system can use the second point predictor unit 312 to calculate, produce, or otherwise determine a first index value 332'B for the point 306B), a representation of one or more consolidated lane lines observed by the plurality of ego-machines (paragraph 95: the computing system can apply an output (e.g., embeddings derived from the first index value 330'8) of the first point predictor unit 310, an output (e.g., embeddings derived from the second index value 332'8) of the second point predictor unit 312, or an output (e.g., embeddings derived from the topology type 334'8) into the point attribute predictor 316; paragraph 96: By applying the point attribute predictor 316, the computing system can generate or determine a set of point attributes 336B for the point 306B. The set of point attributes 336B can identify or include, for example: at least one fork point identifying an index value of a previous point from which the current point 306B forms a fork; at least one merge point identifying an index value of a previous point from which the current point 306B forms a merge; and a set of spline coefficients 338B, among others).
Regarding dependent claim 19, Cho teaches wherein the one or more processing units are further to generate a cell position for each cell of the one or more cells and include the cell position with the corresponding encoded representation (paragraph 85: The computing system may correspond to the ego computing device 141 on an ego 140 or a centralized service such as the analytics server 110a. In addition, the computing system can apply an output (e.g., embeddings derived from the first index value 330'A) of the first point predictor unit 310, an output (e.g., embeddings derived from the second index value 332'A) of the second point predictor unit 312, or an output (e.g., embeddings derived from the topology type 334'A) into the point attribute predictor 316. The portion of the vector space encoding 302 inputted into the point attribute predictor 316 can be the same or can differ from the portion of the vector space encoding 302 inputted into the first point prediction unit 310, the second point prediction unit 312, or the topology predictor unit 314).
Regarding claims 17 and 20, claims 17 and 20 are similar in scope as to claim 14, thus the rejections for claim 14 hereinabove are applicable to claims 17 and 20.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claim(s) 4 – 6 is/are rejected under 35 U.S.C. 103 as being unpatentable over Cho et al. (US 2026/0109373) in view of Official Notice.
Regarding dependent claim 4, Cho teaches wherein generating the encoded representation of the set of the one or more points is based at least on using the first encoder to process one or more cell representations, wherein a cell representation of the one or more cell representations is generated by identifying, from a point cloud, the set of the one or more points that are associated with the cell, however Cho does disclose sensor data is collected from LIDAR (paragraphs 22, 23). Examiner takes Official Notice that the concept of identifying point using point cloud data from LIDAR sensors are well known and expected in the art. It would have been obvious for one of ordinary skill in the art at the time of the invention (pre-AIA ) or at the time of the effective filing date of the application (AIA ) to modify Cho's system to achieve a predictable result of collecting point cloud data points through LIDAR by incorporating system of Cho that collects data points through various sensors like LIDAR and to replace the data points from LIDAR with point cloud data points, and the result would have been predictable.
Regarding dependent claim 5, Cho teaches wherein the cell representation of the one or more cell representations is further generated by determining a local coordinate position for each point of the set of the one or more points that are associated with the cell and determining an encoded position for each point of the set of the one or more points that are associated with the cell (paragraph 85: The computing system may correspond to the ego computing device 141 on an ego 140 or a centralized service such as the analytics server 110a. In addition, the computing system can apply an output (e.g., embeddings derived from the first index value 330'A) of the first point predictor unit 310, an output (e.g., embeddings derived from the second index value 332'A) of the second point predictor unit 312, or an output (e.g., embeddings derived from the topology type 334'A) into the point attribute predictor 316. The portion of the vector space encoding 302 inputted into the point attribute predictor 316 can be the same or can differ from the portion of the vector space encoding 302 inputted into the first point prediction unit 310, the second point prediction unit 312, or the topology predictor unit 314).
Regarding dependent claim 6, Cho teaches wherein the cell representation of the one or more cell representations is further generated by generating, for each point of the set of the one or more points that are associated with the cell, a vector embedding that includes the corresponding local coordinate position and the encoded position (paragraph 77: The vector space encoding 302 (sometimes herein referred to as a tensor) can identify or include a set of encodings (sometimes herein referred as a set of embeddings or feature maps); paragraph 78: the computing system can find, determine, or otherwise identify at least one point 306A within a set of grid points in a first grid 308 defined over the environment represented by a map 310. The point 306A can correspond to a starting point from which the ego is to navigate through the environment as defined by the map 310. The grid 308 can specify or define the set of grid points (or coordinates) at a resolution coarser or lower than an original resolution of the grid point as defined by the map 310) and aggregating the vector embeddings for the points of the set of the one or more points to generate the cell representation (paragraph 72: The transformer 220 can receive, collect, or otherwise aggregate the set of embeddings generated by the feature pyramid networks 215 and by extension the self-regulated networks 210 from the sensor data).
Claim(s) 7 is/are rejected under 35 U.S.C. 103 as being unpatentable over Cho et al. (US 2026/0109373) in view of Official Notice.
Regarding dependent claim 7, Cho teaches wherein the set of the one or more points associated with the cell for which the encoded representation is generated using the first encoder of the one or more neural networks comprise randomly sampled points from the one or more points associated with the one or more cells. Examiner takes Official Notice that the concept of randomly sampling points from similar locations between images/sensor data and the advantage of reducing the number of data points to process are well known and expected in the art. It would have been obvious for one of ordinary skill in the art at the time of the invention (pre-AIA ) or at the time of the effective filing date of the application (AIA ) to modify Cho's system to randomly sample points of the sensor data / image data from the vehicles. One would be motivated to do so because this would reduce the number of data points to process to produce near accurate lane line data.
Allowable Subject Matter
Claim 8 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JEFFREY J CHOW whose telephone number is (571)272-8078. The examiner can normally be reached 11AM-7PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Devona Faulk can be reached at 571-272-7515. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JEFFREY J CHOW/Primary Examiner, Art Unit 2615