Prosecution Insights
Last updated: October 01, 2026
Application No. 18/606,899

FEATURE GENERATION OF DASHED LINE COMPONENTS FOR AUTONOMOUS SYSTEMS AND APPLICATIONS

Final Rejection §103
Filed
Mar 15, 2024
Examiner
AZARIAN, SEYED H
Art Unit
2675
Tech Center
2600 — Communications
Assignee
NVIDIA Corporation
OA Round
2 (Final)
90%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 90% — above average
90%
Career Allowance Rate
814 granted / 909 resolved
+27.5% vs TC avg
Moderate +12% lift
Without
With
+12.0%
Interview Lift
resolved cases with interview
Fast prosecutor
2y 1m
Avg Prosecution
7 currently pending
Career history
914
Total Applications
across all art units

Statute-Specific Performance

§101
18.4%
-21.6% vs TC avg
§103
25.5%
-14.5% vs TC avg
§102
38.3%
-1.7% vs TC avg
§112
7.8%
-32.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 909 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . RESPONSE TO AMENDMENT Based on applicants’ amendment, filed on 8/3/2026, see pages 2 through 14 of remark, with respect to canceled claims 6, 11 and amended claims 1, 10, 19 and new claims 20-21, have been fully considered, and upon further consideration, they are moot in view of the new ground (s) of rejection as necessitated by applicant’s amendment is made in view of Takagi et al (U.S. Pub No: 2024/0004042 A1). Contrary to the applicant’s assertion, as he traverses, limitations in the amended claim, that Pham fails to disclose the recited “determining a frequency associated with the dashed line, or determining, a relationship between portion of the dashed line based on the frequency”. The Examiner has thoroughly reviewed Applicant's argument and respectfully wants to point out, regarding claim 1, Pham reference discloses: (see pages 2-3, paragraphs, [0026-0029], sensor data (e.g., images, videos, etc.) may be received and/or generated using sensors (e.g., cameras, RADAR sensors, LIDAR sensors, etc.) located or otherwise disposed on an autonomous or semi-autonomous vehicle. The sensor data may be applied to a neural network (e.g., a deep neural network (DNN), such as a convolutional neural network (CNN)) that is trained to “identify areas of interest” pertaining to “road markings”, road boundaries, intersections, and/or the like (e.g., raised pavement markers, rumble strips, colored lane dividers, sidewalks, cross-walks, turn-offs, etc.) represented by the sensor data, as well as semantic and/or directional information pertaining thereto. More specifically, the neural network may be designed to compute key points corresponding to segments of an intersection (e.g., corresponding to lanes, bike paths, etc., and/or corresponding to features therein such as cross walks, intersection entry points, intersection exit points, etc.), and to generate outputs identifying, for each key point, a width of a lane corresponding to the key point, a directionality of the lane, a heading corresponding to the lane, semantic information corresponding to the key. For example, the channels may represent key point heat maps, offset vectors, directional vectors, heading vectors, number of lanes and lane types, width of lanes, and/or other channels. During training, the DNN may be trained with images or other sensor data representations labeled or annotated with line segments representing lanes, crosswalks, entry-lines, exit-lines, bike lanes, etc., and may further include semantic information corresponding thereto. The labeled line segments and semantic information may then be used by a ground truth encoder to generate key point heat maps, offset vector maps, directional vector maps, heading vector maps, lane width “intensity maps”, lane counts, lane classifications, and/or other information corresponding to the intersection as determined from the annotations. In some examples, key points may also include corner or end points of the line segments corresponding to lanes, which may be inferred from the center points, the line direction vectors, and/or width information, or may be computed directly using one or more vector fields corresponding to the end point key points. As a result, the intersection structure may be encoded using these key points with limited labeling required, as the information may be determined using the line segment annotations and associated semantic information. In addition to heat maps for key point locations, two-dimensional (2D) vector fields may be used to encode heading directions and line directions. The lane widths may be encoded by assigning an intensity value equal to the lane width normalized by image width, and the same intensity value may be assigned to other pixels within a defined radius. In some examples, the end point key points may be encoded using an offset vector field, such that the DNN learns to compute the location of not only the center key points, but also the end point key points. Using this information, the width of the lanes may be directly determined from the end point key points (e.g., a distance from one end point to the other), the directionality of the line segments extending across the lanes may be determined from the end point key points (e.g., the direction may be defined by a line from one end point to the other end point), and/or the heading direction may be determined from the end point key points (e.g., taking a normal from the directionality to define an angle, and using semantic information to determine the heading). Also pages 5 and 7, paragraphs, [0045] and [0056], further, as indicated by arrows 202A-202V. The heading direction may represent a direction of the traffic pertaining to a certain lane. In some examples, the heading directions may be associated with a center (or key) point of its corresponding lane label. For example, heading direction 202S may be associated with a center point of lane 204. The different classification labels may be represented in FIG. 2A by different line types e.g., solid lines, “dashed lines”, etc. to represent different classifications. However, this is not intended to be limiting, and any visualization of the lane labels and their classifications may include different shapes, patterns, fills, colors, symbols, and/or other identifiers to illustrate differences in classification labels for features (e.g., lanes) in the images. In some embodiments, the intensity map(s) 128 may be implemented to encode lane widths—as determined from the lane label(s) 118A corresponding to segments of the lane(s). For example, once the lane widths are determined, the lane width for the lane segment corresponding to each key point may be encoded by assigning an intensity value equal to the lane width (e.g., in image-space) normalized by image width (e.g., also in image-space) to the key point. In some examples, the same intensity value may be assigned to other pixels within a defined radius of the associated key point, similar to as described herein with respect to the direction vector fields 126A and the heading vector fields 126B. Finally, pages 11 and 16, paragraphs, [0088] and [0142], the paths through the intersection may be used to perform one or more operations by a control component(s) 518 of the vehicle 800. In some examples, a lane graph may be augmented with the information related to the paths. The lane graph may be input to one or more control component(s) 518 to perform various planning and control tasks. For example, a world model manager may update the world model for aid in navigating the intersection, a path planning layer of an autonomous driving software stack may use the intersection information to determine the path through the intersection (e.g., along one of the determined potential paths), and/or a control component may determine controls of the vehicle for navigating the intersection according to a determined path. In some examples, the PVA may be used to perform dense optical flow. According to process raw RADAR data (e.g., using a 4D Fast Fourier Transform) to provide Processed RADAR. In other examples, the PVA is used for time of flight depth processing, by processing raw time of flight data to provide processed time of flight data, for example. However, based on amended claim Pham, does not explicitly state “determining a frequency associated with the representation”. On the other hand, Takagi, in the same field of “a measuring device using the LiDAR technology includes a light source, an interference optical system that separates light into reference light and irradiation light, generates interference light by causing interference between the reference light and reflected light corresponding to an intensity of the interference light”, teaches (see page 5, paragraphs, [0068-0070] FIG. 2A is a graph schematically illustrating time changes in the reference light 20L.sub.1 and the reflected light 20L.sub.3 with respect to the frequency when the object 10 is stationary. The solid line represents the reference light and the dashed line represents the reflected light. The frequency of the reference light 20L.sub.1 illustrated in FIG. 2A repeatedly changes over time in a triangular waveform. More specifically, the frequency of the reference light 20L.sub.1 repeats up-chirping and down-chirping. The increase in frequency in up-chirp and the decrease in frequency in down-chirp are equal to each other. The frequency of the reflected light 20L.sub.3 shifts along the time axis compared to the frequency of the reference light 20L.sub.1. The amount of time shift in the reflected light 20L.sub.3 is equal to the time required for the irradiation light 20L.sub.2 to be emitted from the measuring device 100 and reflected by the object 10 to return as the reflected light 20L.sub.3. As a result, the interference light 20L.sub.4 obtained by superimposing the reference light 20L.sub.1 and the reflected light 20L.sub.3 so as to interfere with each other has a frequency corresponding to the difference between the frequency of the reflected light 20L.sub.3 and the frequency of the reference light 20L.sub.1. The double-headed arrows illustrated in FIG. 2A represent the differences between the two frequencies. The photodetector 60 outputs a signal indicating the intensity of the interference light 20L.sub.4. The signal is called a beat signal. The frequency of the beat signal, that is, the beat frequency, is equal to the difference in frequency described above. The processing circuit 70 can generate data regarding the distance of the object 10 from the beat frequency. FIG. 2B is a graph schematically illustrating time changes in the reference light 20L.sub.1 and the reflected light 20L.sub.3 with respect to the frequency when the object 10 approaches the measuring device 100. When the object 10 approaches, the frequency of the reflected light 20L.sub.3 shifts in an increasing direction along the frequency axis due to Doppler shift, compared to when the object 10 is stationary. The amount of shift in the frequency of the reflected light 20L.sub.3 depends on a component of a velocity vector of the object 10 projected in the direction of the reflected light 20L.sub.3. The beat frequency differs between the up-chirp period and the down-chirp period of the reference light 20L.sub.1 and the reflected light 20L.sub.3. In the example illustrated in FIG. 2B, the beat frequency in the down-chirp period of the both is higher than the beat frequency in the up-chirp period thereof. The processing circuit 70 can generate data on the velocity of object from the difference in beat frequency due to Doppler shift. The processing circuit can also generate data on the distance of the object 10 from the average value of the beat frequency during the up-chirp and down-chirp periods. Next, the relationship between the distance from the measuring device to the object 10 and the beat frequency will be described with reference to FIG. 3. Assuming that the frequency range in the up-chirp period or down-chirp period is Δf, the time required to change Δf is Δt, the speed of light is c, and the distance from the measuring device to the object is d, the beat frequency f.sub.beat is expressed by the following Equation (1). Also page 5, paragraphs, [0072-0074] FIG. 3 is a graph illustrating the relationship between the distance d from the measuring device 100 to the object 10 and the beat frequency f.sub.beat according to Equation (1). The thick “solid line” represents a mode with Δf=12.5 GHz and Δt=10 ρsec, that is, the time rate of change in frequency is Δf/Δt=1.25×10.sup.15 Hz/sec. The thick “dashed line” represents a mode with Δf=1.12 GHz and Δt=10 ρsec, that is, the time rate of change in frequency is Δf/Δt=1.12×10.sup.14 Hz/sec. The thin horizontal dashed line is an example of a maximum measurable value of the beat frequency f.sub.beat, which is 75 MHz. According to Equation (1), the ranging range d increases as the time rate of change Δf/Δt of the frequency decreases. In the example illustrated in FIG. 3, considering the maximum measurable value of the beat frequency f.sub.beat, the ranging range d is 9 m in a mode with a high time rate of change in frequency and is 100 m in a mode with a low time rate of change in frequency. The mode with a high time rate of change in frequency is a mode with a narrow ranging range, that is, a narrow range mode. The mode with a low time rate of change in frequency is a mode with a wide ranging range, that is, a wide range mode. The accuracy of ranging is improved as the time rate of change in frequency increases. This is because the higher the time rate of change in frequency Δf/Δt, the greater the amount of change in the beat frequency f.sub.beat with respect to the amount of change in the distance d. The beat frequency fb eat is obtained by Fourier transforming a beat signal with respect to time. The greater the amount of change in the beat frequency f.sub.beat compared to the frequency resolution of the Fourier transform, the higher the accuracy of ranging. In the example illustrated in FIG. 3, when ranging is performed with a frequency resolution of 800 Hz, the ranging accuracy is about 0.1 mm in the narrow range mode and is about several mm in the wide range mode. The ranging accuracy is 50 repeat accuracy. Therefore, it would have been obvious to one having ordinary skill in the art at the time the invention was made to modify the Pham invention according to the teaching of Takagi because to combine, Pham reference that uses LIDAR sensor intensity information and key points of dashed lines. The different classification labels may be represented in FIG. 2A by different line types e.g., solid lines, dashed lines, etc. to represent different classifications, and any visualization of the lane labels and their classifications may include different shapes, patterns, with the Takagi invention that teaches using a Fourier transform, reflected light 20L.sub.3 with respect to the frequency when the object 10 is stationary. The solid line represents the reference light and the dashed line represents the reflected light. The frequency of the reference light, illustrated in FIG. 2A, for a system and method of vehicle localization and understanding it’s environment. Contrary to the applicant’s assertion, as he traverses, limitations in the amended claim 10, that Pham fails to disclose the recited “pattern associated with variations in the intensity values”. The Examiner respectfully wants to point out, regarding claim 10, Pham reference discloses: See above, also page 2, paragraphs, [0026-0027] in deployment, sensor data (e.g., images, videos, etc.) may be received and/or generated using sensors (e.g., cameras, RADAR sensors, LIDAR sensors, etc.) located or otherwise disposed on an autonomous or semi-autonomous vehicle. The sensor data may be applied to a neural network (e.g., a deep neural network (DNN), such as a convolutional neural network (CNN)) that is trained to identify areas of interest pertaining to “road markings, road boundaries, intersections, and/or the like (e.g., raised pavement markers, rumble strips, colored lane dividers”, sidewalks, cross-walks, turn-offs, etc.) represented by the sensor data, as well as semantic and/or directional information pertaining thereto. More specifically, the neural network may be designed to compute key points corresponding to segments of an intersection (e.g., corresponding to lanes, bike paths, etc., and/or corresponding to “features” therein such as cross walks, intersection entry points, intersection exit points, etc.), and to generate outputs identifying, for each key point, a width of a lane corresponding to the key point, a directionality of the lane, a heading corresponding to the lane, semantic information corresponding to the key point (e.g., crosswalk, crosswalk entry, crosswalk exit, intersection entry, intersection exit, etc., or a combination thereof), and/or other information. In some examples, the computed key points may be denoted by pixels represented by the sensor data where center and/or end points of intersection features are located, such as pedestrian crossing (e.g., crosswalks) entry lines, intersection entrance lines, intersection exit lines, and pedestrian crossing exit lines. The DNN may be trained to predict any number of different information e.g., via any number of channels that correspond to the intersection structure, orientation, and pose. For example, the channels may represent key point heat maps, offset vectors, directional vectors, heading vectors, number of lanes and lane types, width of lanes, and/or other channels. During training, the DNN may be trained with images or other sensor data representations labeled or annotated with line segments representing lanes, crosswalks, entry-lines, exit-lines, bike lanes, etc., and may further include semantic information corresponding thereto. The labeled line segments and semantic information may then be used by a ground truth encoder to generate key point heat maps, offset vector maps, directional vector maps, heading vector maps, lane width intensity maps. Also page 5, paragraphs, [0045-0046] for example, heading direction 202S may be associated with a center point of lane 204. The different classification labels may be represented in FIG. 2A by different line types e.g., solid lines, dashed lines, etc. to represent different classifications. However, this is not intended to be limiting, and any visualization of the lane labels and their classifications may include different shapes, “patterns”, fills, colors, symbols, and/or other identifiers to illustrate differences in classification labels for features (e.g., lanes) in the images. As illustrated, lanes 222A and 222B may be classified as intersection entry line stop line. In this way, similarly classified features of the image may be annotated in a similar manner. Further, it should be noted that classification(s) 118B may be compound nouns. The different classification labels may be represented in FIG. 2B by solid lines, dashed lines, etc. to represent different classifications. Further, the different classification labels may be nouns and/or compound nouns. This is not intended to be limiting, and any naming convention for classifications may be used to illustrate differences in classification labels for features (e.g., lanes) in the images). DETAILED ACTION Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103(a) which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-4, 9, 10, 12, 17-20 and 22 are rejected under 35 U.S.C. 103(a) as being unpatentable over Pham et al (U.S. Pub No: 2020/0341466 A1) in view of Takagi et al (U.S. Pub No: 2024/0004042 A1). Regarding claim 1, Pham discloses a method comprising: obtaining input data representing a dashed line associated with a drivable surface (see page 2, paragraph, [0026] In deployment, sensor data (e.g., images, videos, etc.) received and/or generated (input data), using sensors (e.g., cameras, RADAR sensors, LIDAR sensors, etc.) located or otherwise disposed on an autonomous or semi-autonomous vehicle. The sensor data may be applied to a neural network (e.g., a deep neural network (DNN), such as a convolutional neural network (CNN)) that is trained to identify areas of interest pertaining to “road markings”, road (drivable surface) boundaries, intersections, and/or the like (e.g., raised pavement markers, rumble strips, colored lane dividers, sidewalks, cross-walks, turn-offs, etc.) represented by the sensor data, as well as semantic and/or directional information pertaining thereto. More specifically, the neural network may be designed to compute key points corresponding to segments of an intersection (e.g., corresponding to lanes, bike paths, etc., and/or corresponding to features therein—such as cross walks, intersection entry points, intersection exit points, etc.), and to generate outputs identifying, for each key point, a width of a lane corresponding to the key point, a directionality of the lane, a heading corresponding to the lane, semantic information corresponding to the key point (e.g., crosswalk, crosswalk entry, crosswalk exit, intersection entry, intersection exit, etc., or a combination thereof), and/or other information. In some examples, the computed key points may be denoted by pixels represented by the sensor data where center and/or end points of intersection features are located. Also, page 5, paragraph, [0045] further, labels for the lanes 204, 206, and 208 may further be annotated with a corresponding heading direction, as indicated by arrows 202A-202V. The heading direction may represent a direction of the traffic pertaining to a certain lane. In some examples, the heading directions may be associated with a center (or key) point of its corresponding lane label. For example, heading direction may be associated with a center point of lane. The different classification labels may be represented in FIG. 2A by different line types e.g., solid lines, “dashed lines”, etc. to represent different classifications. However, this is not intended to be limiting, and any visualization of the lane labels and their classifications may include different shapes, patterns, fills, colors, symbols, and/or other identifiers to illustrate differences in classification labels for features (e.g., lanes) in the images); generating a representation associated with the dashed line based at least on intensity values associated with points corresponding to the input data; determining a frequency associated with the representation; (see page 3, paragraphs, [0027] and [0029], during training, the DNN may be trained with images or other sensor data representations labeled or annotated with line segments representing lanes, crosswalks, entry-lines, exit-lines, bike lanes, etc., and may further include semantic information corresponding thereto. The labeled “line segments” (dashed lines), and semantic information may then be used by a ground truth encoder to generate key point heat maps, offset vector maps, directional vector maps, heading vector maps, lane width “intensity maps”, lane counts, lane classifications, and/or other information corresponding to the intersection as determined from the annotations. In some examples, key points may also include corner or end points of the line segments corresponding to lanes, which may be inferred from the center points, the line direction vectors, and/or width information, or may be computed directly using one or more vector fields corresponding to the end point key points. [0029] the 2D line directional vector fields and the 2D heading direction vector fields may be encoded by assigning a directional vector to the center key point, and then assigning the same directional vector to each pixel within a defined radius of the location of the center key point. The lane widths may be encoded by assigning an “intensity value” equal to the lane width normalized by image width, and the same intensity value may be assigned to other pixels within a defined radius. In some examples, the end point key points may be encoded using an offset vector field, such that the DNN learns to compute the location of not only the center key points, but also the end point key points. Using this information, the width of the lanes may be directly determined from the end point key points (e.g., a distance from one end point to the other), the directionality of the line segments extending across the lanes may be determined from the end point key points (e.g., the direction may be defined by a line from one end point to the other end point); determining, based at least on the frequency representation, a relationship between at least a first portion of the dashed line and a second portion of the dashed line; and determining, based at least on the relationship, information associated with one or more components of the dashed line, the information including at least one or more locations associated with the one or more components (see above, also page 5, paragraph, [0041] The lane (or line) label(s) may include annotations, or other label types, corresponding to features or “areas of interest” corresponding to the intersection. In some examples, an intersection structure may be defined as a set of “line segments” corresponding to lanes, crosswalks, entry-lines, exit-lines, bike lanes, etc., in the sensor data. The line segments may be generated as polylines, with a center of each polyline defined as the center for the corresponding line segment. The classification(s) may be generated for each of the images (or other data representations) and/or for each one or more of the line segments and centers in the images represented by the sensor data used for training the machine learning model(s). The number of classification(s) may correspond to the number and/or types of features that the machine learning model(s) is trained to predict, or to the number of lanes and/or types of features in the respective image (dashed lines). Depending on the embodiment, the classification(s) may correspond to classifications or tags corresponding to the feature type, such as but not limited to, crosswalk, crosswalk entry, crosswalk exit, intersection entry. Also pages 10-11, paragraphs, [0083-0084] and [0086-0089] the lane shapes may be determined by decoding the intensity maps to determine the width of the lanes and/or decoding the direction vector fields 112A to determine a directionality of a line segment corresponding to a center key point. As such, the location of the center key point, the directionality, and the width may be leveraged to determine the lane shapes (e.g., by extending the line segment from the center key point according to the directionality and up to a width of the line segment such as half the width from the center key point along the directionality in one direction, and then half the width from the center key point along the directionality in the opposite direction). For example, the lane headings may be determined as the normal to a line segment generated “between” the known locations of the left edge key point and the right edge key point. In other examples, the lane headings may be determined by decoding the heading vector fields 112B to determine the angle corresponding to the heading direction of the lane. In some examples, the lane headings may leverage the lane types 510 and/or other semantic information output by the network, as described herein. In some examples, the path generator 516 may implement curve fitting in order to determine final shapes that most accurately reflect a natural curve of the potential paths. Any known curve fitting algorithms may be used, such as but not limited to, polyline fitting, polynomial fitting, and/or clothoid fitting. The shape of the potential paths may be determined based on the locations of the key points 506 and the corresponding heading vectors (e.g., angles) (e.g., lane headings 514) associated with the key points to be connected. In some examples, the shape of a potential path may be aligned with a tangent of the heading vector at the location of the key points to be connected. The curve fitting process may be repeated for all key points that may potentially be connected to each other to generate all possible paths the vehicle 800 may take to navigate the intersection. In some examples, non-feasible paths may be removed from consideration based on traffic rules and physical restrictions associated with such paths. The remaining potential paths may be determined to be feasible 3D paths or trajectories that the vehicle 800 may take to traverse the intersection. In some embodiments, the path generator 516 may use a matching algorithm to connect the key points 506 and generate the potential paths for the vehicle to navigate the intersection. In such examples, matching scores may be determined for each pair of key points based on the location of the key points, lane headings 514 corresponding to the key points (e.g., two key points corresponding to different directions of travel will not be connected), and the shape of the fitted curve between the pair of key points. Each key point corresponding to an intersection entry may be connected to multiple key points corresponding to an intersection exit thereby generating a plurality of potential paths for the vehicle. In some examples, a linear matching algorithm such as Hungarian matching algorithm may be used. In other examples, a non-linear matching algorithm such as spectral matching algorithm may be used to connect a pair of key points. Now referring to FIG. 6A, FIG. 6A illustrates an example intersection structure prediction 600A (e.g., output(s) 106 of FIG. 5) generated using a neural network (e.g., machine learning model(s) 104 of FIG. 5), in accordance with some embodiments of the present disclosure. The prediction 600A includes a visualization of predicted line segments corresponding to lane classifications 604, 606, 608, 610, and 612, for each lane detected in the sensor data (e.g., sensor data 102 of FIG. 5). For example, the lane classification 604 may correspond to an entrance to a pedestrian crossing lane type, the lane classification 606 may correspond to an entrance to an intersection and/or an exit from a pedestrian crossing lane type, the lane classification 608 may correspond to an exit from a pedestrian crossing lane type, the lane classification 610 may correspond to an exit from an intersection and/or an entrance to a pedestrian crossing lane type, and the lane classifications 612 may correspond to a non-drivable lane type. Each line segment may also be associated with a center key point 506 and/or a corresponding heading direction(s) (e.g., key points and associated vectors 602A-602V). As such, the intersection structure and pose may be represented by a set of line segments with corresponding line classifications, key points, and/or heading directions). Regarding claim 2, Pham discloses the method of claim 1, wherein the input data is first feature data obtained from mapping data, the method further comprising: generating second feature data associated with the one or more components, the second feature data including the one or more locations; and causing the mapping data to be updated to include the second feature data (see claim 1, also abstract, in various examples, live perception from sensors of a vehicle may be leveraged to generate potential paths for the vehicle to “navigate” an intersection in real-time or near real-time. For example, a deep neural network (DNN) may be trained to compute various outputs such as heat maps corresponding to key points associated with the intersection, vector fields corresponding to directionality, heading, and offsets with respect to lanes, “intensity maps” corresponding to widths of lanes, and/or classifications corresponding to line segments of the intersection. The outputs may be decoded and/or otherwise post-processed to reconstruct an intersection or key points corresponding thereto and to determine proposed or potential paths for navigating the vehicle through the intersection. Also, page 7, paragraphs, [0055-0057] in some non-limiting embodiments, the offset 126 vector fields may be encoded by assigning, for each pixel in an offset vector field, a vector pointing to the closest ground truth key “point pixel location”. In this way, smaller encoding channels may be used even when the 2D encoding channels have a different spatial resolution than the input image. This allows the machine learning model(s) to train and predict intersection structure and pose in a computationally less expensive manner because the smaller encoding channels may be used without losing information due to down-sampling of images during processing by the machine learning model(s). In some examples, ground truth data 122 for a “number of features” (e.g., lanes) per classification(s) may be encoded directly using a simple count. The machine learning model(s) may then be trained to predict the number of “features per classification”(s) directly. [0056] In some embodiments, the “intensity map”, may be implemented to encode lane widths as determined from the lane label(s) 118A corresponding to segments of the lane(s). For example, once the lane widths are determined, the lane width for the lane segment corresponding to each key point may be encoded by assigning an intensity value equal to the lane width normalized by image width (e.g., also in image-space) to the key point. In some examples, the same intensity value may be assigned to other pixels within a defined radius of the associated key point, similar to as described herein with respect to the direction vector fields and the heading vector fields. Once the ground truth data is generated for each instance of the sensor data (e.g., for each image where the sensor data includes image data), the machine learning model(s) may be trained using the ground truth data. For example, the machine learning model(s) may generate output(s), and the output(s) may be compared using the loss function(s) to the ground truth data corresponding to the respective instance of the sensor data. As such, feedback from the loss function(s) 130 may be used to “update” parameters (e.g., weights and biases) of the machine learning model(s) 104 in view of the ground truth data 122 until the machine learning model(s) 104 converges to an acceptable or desirable accuracy. Using the process 100, the machine learning model(s) 104 may be trained to accurately predict the output(s) (and/or associated classifications) from the sensor data using the loss function(s) and the ground truth data. In some examples, different loss functions may be used to train the machine learning model(s) to predict different outputs. For example, a first loss function 130 may be used for comparing the heat map(s) and and a second loss function may be used for comparing the intensity maps 128 and the intensity maps 114. As such, in non-limiting embodiments, one or more of the output channels be trained using a different loss function 130 than another of the output channels). Regarding claim 3, Pham discloses the method of claim 1, further comprising causing, based at least on the one or more locations associated with the one or more components, a machine to perform one or more operations (see claim 1, also page 1, paragraph, [0006] in contrast to conventional systems, such as those described above, the current system may use live perception of the vehicle to detect the intersection pose and generate paths for navigating the intersection. Key points (e.g., center points and/or end points) of line segments corresponding to features of an intersection such as lanes, crosswalks, intersection entry or exit lines, bike paths, etc. may be leveraged to generate potential paths for a vehicle to navigate an intersection. For example, machine learning algorithm(s) such as deep neural networks (DNNs) may be trained to compute information corresponding to an intersection—such as key points, heading directions, widths of lanes, number of lanes, etc. and this information may be used to connect together center key points (e.g., key points corresponding to centers of line segments) to generate paths and/or trajectories for the vehicle to effectively and accurately navigate the intersection. As such, semantic information associated with the predicted key points such as directionality, heading, width, and/or classification information corresponding to segments of the intersection—may be computed and leveraged in order to gain an understanding of the intersection pose. For example, the outputs of the DNN may be used to directly or indirectly (e.g., via decoding) determine: a location of each lane, bike path, cross-walk, and/or the like; a number of lanes associated with the intersection; a geometry of the lanes, bike paths, crosswalks, and/or the like; a direction of travel (or heading direction) corresponding to each lane; and/or other intersection structure information). Regarding claim 4, Pham discloses the method of claim 1, wherein the one or more locations associated with the one or more components corresponds with at least one of one or more start points or one or more ends points associated with the one or more components (see claim 1, also page 1, paragraph, [0006] in contrast to conventional systems, such as those described above, the current system may use live perception of the vehicle to detect the intersection pose and generate paths for navigating the intersection. Key points (e.g., center points and/or end points) of line segments corresponding to features of an intersection such as lanes, crosswalks, intersection entry or exit lines, bike paths, etc. may be leveraged to generate potential paths for a vehicle to navigate an intersection. For example, machine learning algorithm(s) such as deep neural networks (DNNs) may be trained to compute information corresponding to an intersection such as key points, heading directions, widths of lanes, number of lanes, etc. and this information may be used to connect together center key points (e.g., key points corresponding to centers of line segments) to generate paths and/or trajectories for the vehicle to effectively and accurately navigate the intersection. As such, semantic information associated with the predicted key points such as directionality, heading, width, and/or classification information corresponding to segments of the intersection may be computed and leveraged in order to gain an understanding of the intersection pose. For example, the outputs of the DNN may be used to directly or indirectly (e.g., via decoding) determine: a location of each lane, bike path, cross-walk, and/or the like; a number of lanes associated with the intersection; a geometry of the lanes, bike paths, crosswalks, and/or the like; a direction of travel (or heading direction) corresponding to each lane; and/or other intersection structure information). Regarding claim 9, Pham discloses the method of claim 1, wherein the input data is an intensity image generated based at least on LiDAR data, the intensity image representing the dashed line associated with the drivable surface from a top-down perspective ( see claim 1, also page 19, paragraph, [0175] LIDAR technologies, such as 3D flash LIDAR, may also be used. 3D Flash LIDAR uses a flash of a laser as a transmission source, to illuminate vehicle surroundings up to approximately 200 m. A flash LIDAR unit includes a receptor, which records the laser pulse transit time and the reflected light on each pixel, which in turn corresponds to the range from the vehicle to the objects. Flash LIDAR may allow for highly accurate and distortion-free images of the surroundings to be generated with every laser flash. In some examples, four flash LIDAR sensors may be deployed, one at each side of the vehicle 800. Available 3D flash LIDAR systems include a solid-state 3D staring array LIDAR camera with no moving parts other than a fan (e.g., a non-scanning LIDAR device). The flash LIDAR device may use a 5-nanosecond class I (eye-safe) laser pulse per frame and may capture the reflected laser light in the form of 3D range point clouds and co-registered intensity data. By using flash LIDAR, and because flash LIDAR is a solid-state device with no moving parts, the LIDAR sensor(s) 864 may be less susceptible to motion blur, vibration, and/or shock). Regarding claim 12, Pham discloses the system of claim 10, wherein the input data is an intensity image generated based at least on LiDAR data, the intensity image representing the feature associated with the drivable surface from a top-down perspective (see claim 1, also page 19, paragraph, [0175] in some examples, LIDAR technologies, such as 3D flash LIDAR, may also be used. 3D Flash LIDAR uses a flash of a laser as a transmission source, to illuminate vehicle surroundings up to approximately 200 m. A flash LIDAR unit includes a receptor, which records the laser pulse transit time and the reflected light on each pixel, which in turn corresponds to the range from the vehicle to the objects. Flash LIDAR may allow for highly accurate and distortion-free images of the surroundings to be generated with every laser flash. In some examples, four flash LIDAR sensors may be deployed, one at each side of the vehicle 800. Available 3D flash LIDAR systems include a solid-state 3D staring array LIDAR camera with no moving parts other than a fan (e.g., a non-scanning LIDAR device). The flash LIDAR device may use a 5-nanosecond class I (eye-safe) laser pulse per frame and may capture the reflected laser light in the form of 3D range point clouds and co-registered intensity data. By using flash LIDAR, and because flash LIDAR is a solid-state device with no moving parts, the LIDAR sensor(s) 864 may be less susceptible to motion blur, vibration, and/or shock). Regarding claim 17, Pham discloses the method of claim 10, wherein the one or more locations associated with the one or more components corresponds with at least one of one or more start points or one or more ends points associated with the one or more components (see claim 1, also page 1, paragraph, [0006] in contrast to conventional systems, such as those described above, the current system may use live perception of the vehicle to detect the intersection pose and generate paths for navigating the intersection. Key points (e.g., center points and/or end points) of line segments corresponding to features of an intersection such as lanes, crosswalks, intersection entry or exit lines, bike paths, etc. may be leveraged to generate potential paths for a vehicle to navigate an intersection. For example, machine learning algorithm(s) such as deep neural networks (DNNs) may be trained to compute information corresponding to an intersection such as key points, heading directions, widths of lanes, number of lanes, etc. and this information may be used to connect together center key points (e.g., key points corresponding to centers of line segments) to generate paths and/or trajectories for the vehicle to effectively and accurately navigate the intersection. As such, semantic information associated with the predicted key points such as directionality, heading, width, and/or classification information corresponding to segments of the intersection may be computed and leveraged in order to gain an understanding of the intersection pose. For example, the outputs of the DNN may be used to directly or indirectly (e.g., via decoding) determine: a location of each lane, bike path, cross-walk, and/or the like; a number of lanes associated with the intersection; a geometry of the lanes, bike paths, crosswalks, and/or the like; a direction of travel (or heading direction) corresponding to each lane; and/or other intersection structure information). Regarding claim 22, Pham discloses the system of claim 10, wherein the determination of the one or more locations associated with the one or more components of the feature is further based at least on determining, based at least on the pattern associated with the variations in the intensity values, one or more starting locations and one or more ending locations associated with the one or more components (see claim 1, also see abstract, in various examples, live perception from sensors of a vehicle may be leveraged to generate potential paths for the vehicle to navigate an intersection in real-time or near real-time. For example, a deep neural network (DNN) may be trained to compute various outputs such as heat maps corresponding to key points associated with the intersection, vector fields corresponding to directionality, heading, and offsets with respect to lanes, intensity maps corresponding to widths of lanes, and/or classifications corresponding to line segments of the intersection. The outputs may be decoded and/or otherwise post-processed to reconstruct an intersection or key points corresponding thereto and to determine proposed or potential paths for navigating the vehicle through the intersection. Also, page 7, paragraphs, [0055-0057] in some non-limiting embodiments, offset vectors may be determined and encoded to generate the offset vector field 126C. The offset 126 vector fields may be encoded by assigning, for each pixel in an offset vector field, a vector pointing to the closest ground truth key point pixel location. In this way, smaller encoding channels may be used even when the 2D encoding channels have a different spatial resolution than the input image (e.g., a down-sampled image). This allows the machine learning model(s) 104 to train and predict intersection structure and pose in a computationally less expensive manner because the smaller encoding channels may be used without losing information due to down-sampling of images during processing by the machine learning model(s) 104. In some examples, ground truth data 122 for a number of features (e.g., lanes) per classification(s) 118B may be encoded directly using a simple count. The machine learning model(s) 104 may then be trained to predict the number of features per classification(s) 118B directly. In some embodiments, the intensity map(s) 128 may be implemented to encode lane widths—as determined from the lane label(s) 118A corresponding to segments of the lane(s). For example, once the lane widths are determined, the lane width for the lane segment corresponding to each key point may be encoded by assigning an intensity value equal to the lane width (e.g., in image-space) normalized by image width (e.g., also in image-space) to the key point. In some examples, the same intensity value may be assigned to other pixels within a defined radius of the associated key point, similar to as described herein with respect to the direction vector fields 126A and the heading vector fields 126B. Once the ground truth data 122 is generated for each instance of the sensor data 102 (e.g., for each image where the sensor data 102 includes image data), the machine learning model(s) 104 may be trained using the ground truth data 122. For example, the machine learning model(s) 104 may generate output(s), and the output(s) may be compared using the loss function(s) to the ground truth data corresponding to the respective instance of the sensor data 102. As such, feedback from the loss function(s) 130 may be used to update parameters (e.g., weights and biases) of the machine learning model(s) 104 in view of the ground truth data 122 until the machine learning model(s) 104 converges to an acceptable or desirable accuracy. Using the process 100, the machine learning model(s) 104 may be trained to accurately predict the output(s) 106 (and/or associated classifications) from the sensor data 102 using the loss function(s) 130 and the ground truth data 122. In some examples, different loss functions 130 may be used to train the machine learning model(s) 104 to predict different outputs 106. For example, a first loss function 130 may be used for comparing the heat map(s) 108 and 124 and a second loss function 130 may be used for comparing the intensity maps 128 and the intensity maps 114. As such, in non-limiting embodiments, one or more of the output channels be trained using a different loss function 130 than another of the output channels). With regard to claims 10 and 18-20, the arguments analogous to those presented above for claims 1, 2, 3, 4, 9, 11, 12, 17 and 22, are respectively applicable to claims 10 and 18-20. Allowable Subject Matter Claims 5-8, 13-16 and 21 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action. Contact Information Any inquiry concerning this communication or earlier communications from the examiner should be directed to Seyed Azarian whose telephone number is (571) 272-7443. The examiner can normally be reached on Monday through Thursday from 6:00 a.m. to 7:30 p.m. If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Matthew Bella, can be reached at (571) 272-7778. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application information Retrieval (PAIR) system. Status information for published application may be obtained from either Private PAIR or Public PAIR. Status information about the PAIR system, see http:// pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). /SEYED H AZARIAN/Primary Examiner, Art Unit 2667 August 23, 2026
Read full office action

Prosecution Timeline

Mar 15, 2024
Application Filed
May 22, 2026
Non-Final Rejection mailed — §103
Jul 12, 2026
Interview Requested
Jul 30, 2026
Examiner Interview Summary
Jul 30, 2026
Response Filed
Jul 30, 2026
Applicant Interview (Telephonic)
Sep 10, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737835
INFORMATION PROCESSING APPARATUS, METHOD OF CONTROLLING INFORMATION PROCESSING APPARATUS, AND STORAGE MEDIUM
2y 6m to grant Granted Sep 15, 2026
Patent 12738097
ACTION RECOGNITION APPARATUS AND METHOD
2y 3m to grant Granted Sep 15, 2026
Patent 12731432
LOW-RESOLUTION EMBEDDED FACE IDENTIFICATION SENSOR WITH MACHINE LEARNING
2y 3m to grant Granted Sep 08, 2026
Patent 12718539
MULTI-MODAL UNDERSTANDING OF EMOTIONS IN VIDEO CONTENT
3y 9m to grant Granted Aug 25, 2026
Patent 12718362
MEDICAL IMAGE PROCESSING APPARATUS, ENDOSCOPE SYSTEM, MEDICAL IMAGE PROCESSING METHOD, AND MEDICAL IMAGE PROCESSING PROGRAM
2y 12m to grant Granted Aug 25, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
90%
Grant Probability
99%
With Interview (+12.0%)
2y 1m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 909 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month