Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim(s) 1 - 20 are pending for examination.
This Action is made FINAL.
Response to Arguments
Applicant's arguments with respect to the previous rejection of claims 1 - 20 under 35 U.S.C. 103 have been considered but are not persuasive.
First applicant argues: “Without agreeing with the rejection, for added clarity only, claim 1 is amended to recite that map queries are generated from BEV features (which have been extracted based on the multi-view images). Liu does not generate its queries based on its BEV features. Liu's queries are not dependent on the image inputs nor are they generated from any information derived from its input images. Rather, Liu's queries are learnable reusable scene-independent queries that represent categories of scene elements ("The detector uses learnable element queries ... Element queries are similar to object queries used in Detection Transformer ... a query represents an object. In our case, an element query represents a map element [in the abstract]"). Moreover, Liu extracts keypoint embeddings from its BEV features, not queries, as shown in Figure 2”
Examiner disagrees. Even though Liu generates “keypoint embeddings” it can be seen in fig.2 that these are the inputs for the polyline generator and thus can be considered queries. Additionally, section 3.4 of Liu states “Each polyline’s keypoint coordinates and class label are tokenized and fed in as the query inputs of the transformer decoder. Then a sequence of vertex tokens are fed into the transformer iteratively, integrating BEV features with cross-attention, and decoded as polyline vertices.”
Second applicant argues: “Claim 1 also recites generating a vectorized map "based on first memory tokens stored in a memory corresponding to queries of previously-processed image frames of a previous vectorized map". The rejection cites only Liu's Section 3.4. The only mention of tokens in Liu is the tokenization of each "[e]ach polyline's keypoint coordinates and class label" (second to last par. of Section 3.4). Claim 1 recites that the tokens used to generate the claimed vectorized map "correspond[] to queries of previously-processed image frames". A principle of claim interpretation is that different claim terms are assumed to have different meanings. As noted in CAE Screenplates Inc. v. Heinrich Fiedler GmbH & Co. KG, 224 F.3d 1308 (Fed. Cir. 2000), "In the absence of any evidence to the contrary, we must presume that the use of these different terms in the claims connotes different meanings."). Therefore, the recited previously-processed image frames are presumed to not refer to the multi-view images also recited in claim 1, from which the BEV features are derived (and based on which the map queries are generated).
Further regarding the tokens of the previously-processed image frames recited in claim 1 (for a previous vectorized map), Liu generally lacks any notion of using any pieces of information that correspond to historical or previous processing of a previous vectorized map, and Afshar is not cited as teaching this feature.”
Applicant’s argument is moot in view of new ground of rejection.
Third Applicant argues: “Claim 1 also recites that the multi-view images (from which the BEV features are extracted) are frames "at consecutive time points", each comprised of image frames for its corresponding time point. The rejection finds that Liu lacks this teaching, and that it "would have been prima facie obvious ... to have modified Liu to incorporate the teachings of Afshar to use the vectorized map generated from images for autonomous driving". The proposed combination fails to meet the language of claim 1. Claim 1 recites that its multi-view images (each comprising image frames) are of the vehicle's environment at consecutive time points. Ashfar has no such teaching. Ashfar teaches inputting only a single image to its CNN ("perception system 402 provides the data associated with the image to CNN 440, where the image is a greyscale image represented as values stored in a two-dimensional (2D) array.", par. 0085).”
Examiner disagrees. Liu already teaches the Multiview image. It would be trivial for one of ordinary skill in the art to use a Multiview image instead of a “single image” a Multiview image can merely be a stitched together single image frames. Thus the processing the neural network would have to do is close to identical between a Multiview image and a single image. It would be also trivial for “Liu to record the environment continuously as in Ashfar instead of just taking a single frame. It should be noted that Liu may indeed continuously record surround images, Liu is just silent on the matter of the frequency of recording and processing.
Fourth Applicant argues: “In addition, the rejection is traversed because Ashfar's applicability to Liu, if any, comes well after the relevant teachings of Liu. Liu has no explicit teaching of path trajectories (the cited motivation in Ashfar). At best, Liu arguably implicitly teaches using path trajectories in its vehicle-control phase, which uses Liu's vector maps. Improving route planning (trajectory path), as taught by Ashfar, has no relevance to the techniques that Liu uses to generate a vector map.”
Examiner disagrees. Applicant’s recitation of control using a vectorized map is generic. Liu teaches the vectorized map generation Ashfar teaches the controlling a vehicle using a vectorized map as input. Using vectorized maps for control has the benefit of improving reliability of vehicle control as discussed in Ashfar.
Lastly Applicant argues: “Finally, while paragraph 0158 of Ashfar discusses periodically predicting a trajectory, this involves only repeating the prediction process 900 at different times for different respective images. This has no relevance to using multiple image frames of consecutive time points to produce a single vectorized map, as recited in claim 1.”
Applicant’s argument is moot in view of new ground of rejection.
Applicant presents no new arguments with regard to the dependent claims. Applicant’s arguments directed toward the pending dependent claims are that due to their dependency on what the applicant believes are allowable independent claims, the pending dependent claims should also be allowable. However as previously stated, examiner has not found applicant’s arguments directed towards the independent claims persuasive.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim(s) 20 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 20 recites the limitation "the image frames" in line 7. There is insufficient antecedent basis for this limitation in the claim.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Liu et al. (VectorMapNet: End-to-end Vectorized HD Map Learning, 2023, hereinafter known as Liu) in view of Afshar et al. (US 20240132112 A1, hereinafter known as Afshar) and Yasarla et al. (US 20240303841 A1; hereinafter Yasarla).
Liu and Afshar were cited in a previous office action.
Regarding Claim 1, Liu teaches receiving multi-view images
{Section 3 “Similar to HDMapNet (Li et al., 2021), our task is to vectorize map elements using data from onboard sensors of autonomous vehicle, such as RGB cameras and/or LiDARs.”
Section 3.2 “The objective of BEV feature extractor is to lift various modality inputs into a canonical feature space and aggregates and align features these features into a canonical representation termed BEV features FBEV ∈ RW× H× (C1+C2) based on their coordinates, where W and H represent the width and height of the BEV feature, respectively; C1 and C2 represent the output channels of the BEV feature extracted from the two common modalities: surrounding camera images I and LiDAR points P”
A surrounding view image is now to be stitched together individual image frames from different perspectives. Fig. 1 shows an example of a surround view image.
}
extracting bird's-eye view (BEV) features respectively corresponding to the consecutive time points based on each of the multi-view images comprising the image frames, and
{ Section 3.2 “The objective of BEV feature extractor is to lift various modality inputs into a canonical feature space and aggregates and align features these features into a canonical representation termed BEV features FBEV ∈ RW× H× (C1+C2) based on their coordinates, where W and H represent the width and height of the BEV feature, respectively; C1 and C2 represent the output channels of the BEV feature extracted from the two common modalities: surrounding camera images I and LiDAR points P”
}
generating map queries corresponding to the consecutive time points, based on the BEV features;
{ Section 3.3 “After extracting the birds-eye view (BEV) features, VectorMapNet have to identify and abstractly represent map elements using these features. We employ a hierarchical representation for this purpose, specifically through element queries and keypoint queries, enabling us to model the nonlocal shape of map elements effectively. We leverage a variant of transformer set prediction detector (Carion et al., 2020) to achieve this goal, as it is a robust detector that eliminates the need for extra post-processing. Specifically, the detector represents map elements locations and categories by predicting their element keypoints A and class labels L from the BEV features FBEV.”
see figure 2 where element keypoints become the queries for the polyline generator.
Section 3.4 “Each polyline’s keypoint coordinates and class label are tokenized and fed in as the query inputs of the transformer decoder. Then a sequence of vertex tokens are fed into the transformer iteratively, integrating BEV features with cross-attention, and decoded as polyline vertices.”
}
generating a vectorized map by predicting and vectorizing map elements represented in the image frames, the generating based on first memory tokens stored in a memory corresponding to queries
{Fig. 2 and See all of section 3.4
}
providing information that can be used for the driving of the vehicle based on the vectorized map.
{section 4.2 “To further demonstrate this flexibility, we expand Vector Map Net to predict the centerline, an imaginary line commonly used as a reference for driving direction, vehicle positioning, and navigation.”
}
Liu does not teach, receiving multi-view images respectively corresponding to consecutive time points
and
the generating based on first memory tokens stored in a memory corresponding to queries
and
controlling the driving of the vehicle based on the vectorized map.
However, Afshar receiving
{Para [0069] “In some embodiments, perception system 402 receives data associated with at least one physical object (e.g., data that is used by perception system 402 to detect the at least one physical object) in an environment and classifies the at least one physical object. In some examples, perception system 402 receives image data captured by at least one camera (e.g., cameras 202a), the image associated with (e.g., representing) one or more physical objects within a field of view of the at least one camera. In such an example, perception system 402 classifies at least one physical object based on one or more groupings of physical objects (e.g., bicycles, vehicles, traffic signs, pedestrians, and/or the like). In some embodiments, perception system 402 transmits data associated with the classification of the physical objects to planning system 404 based on perception system 402 classifying the physical objects.”
Para [0158] “In some embodiments, the process 900 includes: periodically predicting a future trajectory of an agent in a current environment of the vehicle based on at least one reference path determined for the agent. For example, the perception system can perform the process 900 in a period, e.g., every 10 seconds, 20 seconds, 30 seconds, or 1 minute. In some embodiments, the perception system performs the process 900 continuously. For example, once a round of the process 900 ends, the process 900 restarts or reiterates. In some embodiments, the perception system performs the process 900 in response to a triggering event, e.g., an input from a driver.”
}
controlling the driving of the vehicle based on the vectorized map.
{Para [0030] “For each agent, the path-based trajectory prediction can include multiple operations: 1) vectorizing map into connected lane segments; 2) sampling the vectorized map for candidate reference paths (e.g., in 8 seconds) with reachable lane segments or reachable targets (e.g., end points) of the candidate reference paths; 3) classifying a set of candidate reference paths (e.g., by predicting a discrete probability distribution over the candidate reference paths) based on defined feature vectors, including scene feature vector (e.g., agent behavior) and path feature vector (e.g., first point, middle point, last point, direction, and length of each candidate reference path); 4) making trajectory prediction with respect to one or more selected reference paths in the Frenet frame using agents feature map augmented with path information; and 5) transforming the predicted trajectories back to Cartesian co-ordinates relative to the agent to obtain multimodal predictions.”
Para [0031] “Some of the advantages of these techniques are as follows. For example, the techniques predict trajectories conditioned on feature descriptors of a complete reference path from the agent's current location to the agent's goal instead of just its goal locations. This is a much more informative feature descriptor and leads to more map compliant trajectories over longer prediction horizons compared to goal based prediction. Also, the techniques use reference paths, which allow to predict trajectories in the path relative Frenet frame relative to each sampled path. Compared to the Cartesian frame with varying lane locations and curvatures, predictions in the Frenet frame can have much lower variance. This again leads to more map compliant trajectories that better generalize to novel scene layouts. Moreover, compared to using a rasterized HD map for its scene and reference path encoders, the techniques directly encode the scene and reference paths using polylines, making the encoders more efficient. The techniques can sample and classify variable length reference paths along each lane centerline, which provides trajectory prediction with more flexibility to predict different motion profiles along lanes. The techniques can improve path prediction and path compliance, e.g., using agent past trajectory history in the prediction. The techniques can enhance performance of prediction in multi-lane turns with better path classifier and scene upsampling. In addition to standard metrics for multimodal prediction, the techniques can enhance two map compliance metrics of the predicted trajectories (e.g., commonly used drivable area compliance metric and a new lane deviation metric), for example, by utilizing map prior knowledge (e.g., high likelihood drivable areas). Further, the techniques can improve interaction reasoning in path encoder and improve the map and agents interaction graph. The techniques can improve reaction of autonomous vehicles to surrounding environments (e.g., periodically or continuously) to achieve reliable and accurate prediction for their own route/trajectory or operation planning, which realizes safe and reliable driving.”
}
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Liu to incorporate the teachings of Afshar to use the vectorized map generated from images for autonomous driving because Para [0030] “. The techniques can improve reaction of autonomous vehicles to surrounding environments (e.g., periodically or continuously) to achieve reliable and accurate prediction for their own route/trajectory or operation planning, which realizes safe and reliable driving.”
Liu in view of Afshar does not teach generating a vectorized map by predicting and vectorizing map elements represented in the image frames, the generating based on first memory tokens stored in a memory corresponding to queries
However Yasarla teaches generating a spatial representation by predicting spatial representation
{Para [0066] “A stream of images 420 may be provided by a monocular imaging system (e.g., the image capture devices 302) at a frame rate (e.g., 30 Hz, 60 Hz, etc.) and each image is encoded in the encoder 402. In some aspects, the encoder 402 is configured to identify features in each image, and the features are often referred to as vectors or tokens. In one aspect, the encoder 402 represents the image as query tokens, which are potential features of interest in the scene. As described in FIG. 6, the query tokens are provided to the depth estimator 404 and features within the query tokens are identified by the feature engine 408. In some aspects, the decoder 406 includes memory tokens, which are tokens that identify relevant features within the scene and are stored for use in connection with later images. In some cases, the memory tokens may be referred to as accumulated query information and represent relevant features of previous images. The feature engine 408 is configured to update and maintain the memory tokens based on the image 420. The memory tokens summarize and store key past information as features and the depth estimator 404 can cross-reference relevant features from previous frames when inferring depth on the image 420. The depth estimator 404 updates the memory tokens each image and stores the most relevant information.”
}
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Liu in view of Afshar to incorporate the teachings of Yasarla to use memory tokens and information from previous processing because it improves accuracy Para [0075] “In some aspects, the additional decoder information provided between decoding iterations of the decoder 506 may improve accuracy.”
Regarding Claim 2, Lui in view of Afshar and Yasarla teaches The method of claim 1. Liu further teaches wherein the extracting of the BEV features and the map queries comprises: extracting image features of a perspective view (PV) corresponding to the image frames using a backbone network; transforming the image features of the PV into the BEV features; extracting the map queries at a frame level used to construct the vectorized map based on the BEV features and a query corresponding to the image frames; and outputting the BEV features and the map queries.
{See section 3.2 and section 3.3
Also see Section 3 “The inputs and outputs of the mapping problem are not perfectly aligned. They exist in different view spaces (e.g. camera data is in perspective
view and map elements are in BEV)”
}
Regarding Claim 3, Lui in view of Afshar and Yasarla teaches The method of claim 1. Liu further teaches wherein the generating of the vectorized map comprises: reading the first memory tokens; based on the map queries, the BEV features, and the first memory tokens, generating map tokens comprising the map elements comprised in the vectorized map and/or clip tokens comprising vectorized features corresponding to the image frames; and generating the vectorized map based on the map tokens.
{See figure 2 and Section 3.3 and Section 3.4
}
Regarding Claim 4, Lui in view of Afshar and Yasarla teaches The method of claim 3. Liu further teaches wherein the generating of the map tokens and/or the clip tokens comprises: generating, from the map queries and the first memory tokens, the clip tokens comprising cues for the map elements in a feature space corresponding to the image frames; updating the BEV features using the clip tokens such that the BEV features comprise hidden map elements; and generating the map tokens using the updated BEV features and the map queries.
{See figure 2 and Section 3.3 and Section 3.4
}
Regarding Claim 5, Lui in view of Afshar and Yasarla teaches The method of claim 3. Liu further teaches wherein sizes of the map queries are determined based on sizes of the clip tokens, a number of the map elements, or a number of points for each of the map elements.
{ See figure 2 and Section 3.3 and Section 3.4
}
Regarding Claim 6, Lui in view of Afshar and Yasarla teaches The method of claim 4. Liu further teaches wherein the updating of the BEV features comprises: extracting a query from the BEV features; extracting a key and a value from the clip tokens; and updating the BEV features via a cross-attention network and a feed-forward network using the query, the key, and the value.
{ See figure 2 and Section 3.3 and Section 3.4
}
Regarding Claim 7, Lui in view of Afshar and Yasarla teaches The method of claim 4. Liu further teaches wherein the generating of the map tokens comprises generating the map tokens from the map queries and the updated BEV features using a deformable attention network, a decoupled self-attention network, and a feed-forward network.
{ See figure 2 and Section 3.3 and Section 3.4
}
Regarding Claim 8, Lui in view of Afshar and Yasarla teaches The method of claim 7. Liu further teaches wherein the generating of the map tokens comprises generating the map tokens by extracting the queries from the map queries using the deformable attention network and obtaining a value from the updated BEV features.
{ See figure 2 and Section 3.3 and Section 3.4
}
Regarding Claim 9, Lui in view of Afshar and Yasarla teaches The method of claim 1. Liu further teaches wherein the generating of the vectorized map comprises generating the vectorized map by predicting the map elements represented in the image frames by a pre-trained neural network and vectorizing the map elements for each instance, and the pre-trained neural network comprises at least one of: a (2-1)-th neural network configured to read the first memory tokens from the memory or write second memory tokens to the memory; and a (2-2)-th neural network configured to generate the vectorized map corresponding to a current frame among the image frames based on the map queries, the BEV features, and the first memory tokens.
{See figure 2 and Section 3.3 and Section 3.4 and Section 3.5}
Regarding Claim 10, Lui in view of Afshar and Yasarla teaches The method of claim 9. Liu further teaches wherein the generating of the vectorized map comprises: writing the map tokens to the memory by the (2-1)-th neural network; and generating the vectorized map as a map token corresponding to the current frame among the map tokens passes through a prediction head.
{ See figure 2 and Section 3.3 and Section 3.4 and Section 3.5
}
Regarding Claim 11, Lui in view of Afshar and Yasarla teaches The method of claim 9. Liu further teaches further comprising: generating the second memory tokens by writing the map tokens and the clip tokens to the memory using the (2-1)-th neural network; and outputting the second memory tokens.
{ See figure 2 and Section 3.3 and Section 3.4 and Section 3.5
}
Regarding Claim 12, Lui in view of Afshar and Yasarla teaches The method of claim 9. Liu further teaches wherein the (2-1)-th neural network is configured to preserve time information corresponding to the previous image frames by reading first memory tokens corresponding to the previous image frames to propagate the first memory tokens as an input for the (2-2)-th neural network.
{ See figure 2 and Section 3.3 and Section 3.4 and Section 3.5
}
Regarding Claim 13, Lui in view of Afshar and Yasarla teaches The method of claim 9. Liu further teaches wherein the (2-1)-th neural network is configured to set intra-clip associations between the map elements by associating inter-clip information through propagation of clip tokens generated in the (2-2)-th neural network.
{ See figure 2 and Section 3.3 and Section 3.4 and Section 3.5
}
Regarding Claim 14, Lui in view of Afshar and Yasarla teaches The method of claim 9. Liu further teaches wherein the (2-1)-th neural network is configured to generate the second memory tokens comprising global map information through embedding of a learnable frame and store the second memory tokens in the memory, based on the map tokens and the clip tokens generated in the (2-2)-th neural network.
{ See figure 2 and Section 3.3 and Section 3.4 and Section 3.5
}
Regarding Claim 15, Lui in view of Afshar and Yasarla teaches The method of claim 9. Liu further teaches wherein the (2-1)-th neural network is configured to generate the second memory tokens by combining clip tokens, the map tokens, and the first memory tokens together.
{ See figure 2 and Section 3.3 and Section 3.4 and Section 3.5
}
Regarding Claim 16, Lui in view of Afshar and Yasarla teaches The method of claim 9. Liu further teaches wherein the (2-2)-th neural network is configured to generate the vectorized map by outputting a map token corresponding to a current frame having a predetermined time window corresponding to lengths of the image frames, based on the first memory tokens, the BEV feature, and the map queries.
{ See figure 2 and Section 3.3 and Section 3.4 and Section 3.5
}
Regarding Claim 17, Lui in view of Afshar and Yasarla teaches The method of claim 1. Liu further teaches wherein the map elements comprise a crosswalk, a road, a lane, a lane boundary, a building, a curbstone, or traffic lights comprised in the driving environment.
{Section 3 “Similar to HDMapNet (Li et al., 2021), our task is to vectorize map elements using data from onboard sensors of autonomous vehicle, such as RGB cameras and/or LiDARs. These map elements include but are not limited to: Road boundaries (boundaries of roads separating roads and sidewalks, typically irregularly-shaped curves of arbitrary lengths), Lane dividers (boundaries dividing lanes on the road, usually straight lines), and Pedestrian crossings (regions with white markings indicating legal pedestrian crossing points, typically represented as polygons).”
}
Regarding Claim 18, Lui in view of Afshar and Yasarla teaches The method of claim 1. Afshar further teaches A non-transitory computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform the method
{Para [0063] “In some embodiments, device 300 performs one or more processes described herein. Device 300 performs these processes based on processor 304 executing software instructions stored by a computer-readable medium, such as memory 305 and/or storage component 308. A computer-readable medium (e.g., a non-transitory computer readable medium) is defined herein as a non-transitory memory device. A non-transitory memory device includes memory space located inside a single physical storage device or memory space spread across multiple physical storage devices.”
}
Regarding Claim 19, Liu teaches An apparatus for controlling driving of a vehicle, the apparatus comprising: memory configured to store multi-view images of a driving environment of the vehicle, the multi-view image
{Section 3 “Similar to HDMapNet (Li et al., 2021), our task is to vectorize map elements using data from onboard sensors of autonomous vehicle, such as RGB cameras and/or LiDARs.”
Section 3.2 “The objective of BEV feature extractor is to lift various modality inputs into a canonical feature space and aggregates and align features these features into a canonical representation termed BEV features FBEV ∈ RW× H× (C1+C2) based on their coordinates, where W and H represent the width and height of the BEV feature, respectively; C1 and C2 represent the output channels of the BEV feature extracted from the two common modalities: surrounding camera images I and LiDAR points P”
A surrounding view image is now to be stitched together individual image frames from different perspectives. Fig. 1 shows an example of a surround view image.
}
a first neural network configured to extract bird's-eye view (BEV) features respectively corresponding to consecutive time points based on each of the multi-view images comprising the image frames, and configured to extract, based on the BEV feature, map queries respectively corresponding to the consecutive time points for each of the image frames;
{ Section 3.2 “The objective of BEV feature extractor is to lift various modality inputs into a canonical feature space and aggregates and align features these features into a canonical representation termed BEV features FBEV ∈ RW× H× (C1+C2) based on their coordinates, where W and H represent the width and height of the BEV feature, respectively; C1 and C2 represent the output channels of the BEV feature extracted from the two common modalities: surrounding camera images I and LiDAR points P”
It is also discussed in Section 3.2 that a convolutional neural network is used
Section 3.3 “After extracting the birds-eye view (BEV) features, VectorMapNet have to identify and abstractly represent map elements using these features. We employ a hierarchical representation for this purpose, specifically through element queries and keypoint queries, enabling us to model the nonlocal shape of map elements effectively. We leverage a variant of transformer set prediction detector (Carion et al., 2020) to achieve this goal, as it is a robust detector that eliminates the need for extra post-processing. Specifically, the detector represents map elements locations and categories by predicting their element keypoints A and class labels L from the BEV features FBEV.”
see figure 2 where element keypoints become the queries for the polyline generator.
Section 3.4 “Each polyline’s keypoint coordinates and class label are tokenized and fed in as the query inputs of the transformer decoder. Then a sequence of vertex tokens are fed into the transformer iteratively, integrating BEV features with cross-attention, and decoded as polyline vertices.”
}
a second neural network configured to generate a vectorized map by predicting and vectorizing map elements represented in the image frames, the generating based on first memory tokens stored in a memory corresponding to queries
{Fig. 2 and See all of section 3.4
}
a processor configured toproviding information that can be used for driving of the vehicle based on the vectorized map.
{section 4.2 “To further demonstrate this flexibility, we expand Vector Map Net to predict the centerline, an imaginary line commonly used as a reference for driving direction, vehicle positioning, and navigation.”
}
Liu does not teach, the multi-view image respectively corresponding to consecutive time points
And
generating based on first memory tokens stored in a memory corresponding to queries of previously-processed images of a previous vectorized map
and and a processor configured to control driving of the vehicle based on the vectorized map.
However, Afshar teaches the multi-view image respectively corresponding to consecutive time points
{Para [0069] “In some embodiments, perception system 402 receives data associated with at least one physical object (e.g., data that is used by perception system 402 to detect the at least one physical object) in an environment and classifies the at least one physical object. In some examples, perception system 402 receives image data captured by at least one camera (e.g., cameras 202a), the image associated with (e.g., representing) one or more physical objects within a field of view of the at least one camera. In such an example, perception system 402 classifies at least one physical object based on one or more groupings of physical objects (e.g., bicycles, vehicles, traffic signs, pedestrians, and/or the like). In some embodiments, perception system 402 transmits data associated with the classification of the physical objects to planning system 404 based on perception system 402 classifying the physical objects.”
Para [0158] “In some embodiments, the process 900 includes: periodically predicting a future trajectory of an agent in a current environment of the vehicle based on at least one reference path determined for the agent. For example, the perception system can perform the process 900 in a period, e.g., every 10 seconds, 20 seconds, 30 seconds, or 1 minute. In some embodiments, the perception system performs the process 900 continuously. For example, once a round of the process 900 ends, the process 900 restarts or reiterates. In some embodiments, the perception system performs the process 900 in response to a triggering event, e.g., an input from a driver.”
}
and a processor configured to control driving of the vehicle based on the vectorized map.
{Para [0030] “For each agent, the path-based trajectory prediction can include multiple operations: 1) vectorizing map into connected lane segments; 2) sampling the vectorized map for candidate reference paths (e.g., in 8 seconds) with reachable lane segments or reachable targets (e.g., end points) of the candidate reference paths; 3) classifying a set of candidate reference paths (e.g., by predicting a discrete probability distribution over the candidate reference paths) based on defined feature vectors, including scene feature vector (e.g., agent behavior) and path feature vector (e.g., first point, middle point, last point, direction, and length of each candidate reference path); 4) making trajectory prediction with respect to one or more selected reference paths in the Frenet frame using agents feature map augmented with path information; and 5) transforming the predicted trajectories back to Cartesian co-ordinates relative to the agent to obtain multimodal predictions.”
Para [0031] “Some of the advantages of these techniques are as follows. For example, the techniques predict trajectories conditioned on feature descriptors of a complete reference path from the agent's current location to the agent's goal instead of just its goal locations. This is a much more informative feature descriptor and leads to more map compliant trajectories over longer prediction horizons compared to goal based prediction. Also, the techniques use reference paths, which allow to predict trajectories in the path relative Frenet frame relative to each sampled path. Compared to the Cartesian frame with varying lane locations and curvatures, predictions in the Frenet frame can have much lower variance. This again leads to more map compliant trajectories that better generalize to novel scene layouts. Moreover, compared to using a rasterized HD map for its scene and reference path encoders, the techniques directly encode the scene and reference paths using polylines, making the encoders more efficient. The techniques can sample and classify variable length reference paths along each lane centerline, which provides trajectory prediction with more flexibility to predict different motion profiles along lanes. The techniques can improve path prediction and path compliance, e.g., using agent past trajectory history in the prediction. The techniques can enhance performance of prediction in multi-lane turns with better path classifier and scene upsampling. In addition to standard metrics for multimodal prediction, the techniques can enhance two map compliance metrics of the predicted trajectories (e.g., commonly used drivable area compliance metric and a new lane deviation metric), for example, by utilizing map prior knowledge (e.g., high likelihood drivable areas). Further, the techniques can improve interaction reasoning in path encoder and improve the map and agents interaction graph. The techniques can improve reaction of autonomous vehicles to surrounding environments (e.g., periodically or continuously) to achieve reliable and accurate prediction for their own route/trajectory or operation planning, which realizes safe and reliable driving.”
Para [0063] “In some embodiments, device 300 performs one or more processes described herein. Device 300 performs these processes based on processor 304 executing software instructions stored by a computer-readable medium, such as memory 305 and/or storage component 308. A computer-readable medium (e.g., a non-transitory computer readable medium) is defined herein as a non-transitory memory device. A non-transitory memory device includes memory space located inside a single physical storage device or memory space spread across multiple physical storage devices.”
}
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Liu to incorporate the teachings of Afshar to use the vectorized map generated from images for autonomous driving because Para [0030] “. The techniques can improve reaction of autonomous vehicles to surrounding environments (e.g., periodically or continuously) to achieve reliable and accurate prediction for their own route/trajectory or operation planning, which realizes safe and reliable driving.”
Liu in view of Afshar does not teach generating based on first memory tokens stored in a memory corresponding to queries of previously-processed images of a previous vectorized map
However Yasarla teaches a second neural network configured to generate a spatial representation by predicting generating based on first memory tokens stored in a memory corresponding to queries of previously-processed images of a previous spatial representation
{Para [0066] “A stream of images 420 may be provided by a monocular imaging system (e.g., the image capture devices 302) at a frame rate (e.g., 30 Hz, 60 Hz, etc.) and each image is encoded in the encoder 402. In some aspects, the encoder 402 is configured to identify features in each image, and the features are often referred to as vectors or tokens. In one aspect, the encoder 402 represents the image as query tokens, which are potential features of interest in the scene. As described in FIG. 6, the query tokens are provided to the depth estimator 404 and features within the query tokens are identified by the feature engine 408. In some aspects, the decoder 406 includes memory tokens, which are tokens that identify relevant features within the scene and are stored for use in connection with later images. In some cases, the memory tokens may be referred to as accumulated query information and represent relevant features of previous images. The feature engine 408 is configured to update and maintain the memory tokens based on the image 420. The memory tokens summarize and store key past information as features and the depth estimator 404 can cross-reference relevant features from previous frames when inferring depth on the image 420. The depth estimator 404 updates the memory tokens each image and stores the most relevant information.”
}
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Liu in view of Afshar to incorporate the teachings of Yasarla to use memory tokens and information from previous processing because it improves accuracy Para [0075] “In some aspects, the additional decoder information provided between decoding iterations of the decoder 506 may improve accuracy.”
Regarding Claim 20, Liu teaches A vehicle comprising: sensors configured to capture multi-view images
{Section 3 “Similar to HDMapNet (Li et al., 2021), our task is to vectorize map elements using data from onboard sensors of autonomous vehicle, such as RGB cameras and/or LiDARs.”
Section 3.2 “The objective of BEV feature extractor is to lift various modality inputs into a canonical feature space and aggregates and align features these features into a canonical representation termed BEV features FBEV ∈ RW× H× (C1+C2) based on their coordinates, where W and H represent the width and height of the BEV feature, respectively; C1 and C2 represent the output channels of the BEV feature extracted from the two common modalities: surrounding camera images I and LiDAR points P”
}
a neural network configured to extract bird's-eye view (BEV) features respectively corresponding to the consecutive time points based on each of the multi-view images comprising the image frames, to extract map queries from the BEV features,
{ Section 3.2 “The objective of BEV feature extractor is to lift various modality inputs into a canonical feature space and aggregates and align features these features into a canonical representation termed BEV features FBEV ∈ RW× H× (C1+C2) based on their coordinates, where W and H represent the width and height of the BEV feature, respectively; C1 and C2 represent the output channels of the BEV feature extracted from the two common modalities: surrounding camera images I and LiDAR points P”
It is also discussed in Section 3.2 that a convolutional neural network is used
Section 3.3 “After extracting the birds-eye view (BEV) features, VectorMapNet have to identify and abstractly represent map elements using these features. We employ a hierarchical representation for this purpose, specifically through element queries and keypoint queries, enabling us to model the nonlocal shape of map elements effectively. We leverage a variant of transformer set prediction detector (Carion et al., 2020) to achieve this goal, as it is a robust detector that eliminates the need for extra post-processing. Specifically, the detector represents map elements locations and categories by predicting their element keypoints A and class labels L from the BEV features FBEV.
The detector uses learnable element … as its inputs, where d represents the hidden embedding size and Nmax is a preset constant, which is much greater than the number of map elements N in the scene. The i-th element query is composed of k element keypoint embeddings Element queries are similar to object queries used in Detection Transformer (DETR) (Carion et al., 2020), where a query represents an object. In our case, an element query represents a map element.”
}
and configured to generate a vectorized map by predicting and vectorizing map elements represented in the image frames based on first memory tokens stored in a memory corresponding to queries
{Fig. 2 and see all of section 3.4
}
a processor configured toproviding information that can be used for driving of the vehicle based on the vectorized map.
{section 4.2 “To further demonstrate this flexibility, we expand Vector Map Net to predict the centerline, an imaginary line commonly used as a reference for driving direction, vehicle positioning, and navigation.”
}
Liu does not teach, images respectively corresponding to consecutive time points
and
based on first memory tokens stored in a memory corresponding to queries of previous image frames of the image frames used to generate a previous vectorized map
and a processor configured to generate a control signal for driving the vehicle based on the vectorized map.
However, Afshar teaches images respectively corresponding to consecutive time points
{Para [0069] “In some embodiments, perception system 402 receives data associated with at least one physical object (e.g., data that is used by perception system 402 to detect the at least one physical object) in an environment and classifies the at least one physical object. In some examples, perception system 402 receives image data captured by at least one camera (e.g., cameras 202a), the image associated with (e.g., representing) one or more physical objects within a field of view of the at least one camera. In such an example, perception system 402 classifies at least one physical object based on one or more groupings of physical objects (e.g., bicycles, vehicles, traffic signs, pedestrians, and/or the like). In some embodiments, perception system 402 transmits data associated with the classification of the physical objects to planning system 404 based on perception system 402 classifying the physical objects.”
Para [0158] “In some embodiments, the process 900 includes: periodically predicting a future trajectory of an agent in a current environment of the vehicle based on at least one reference path determined for the agent. For example, the perception system can perform the process 900 in a period, e.g., every 10 seconds, 20 seconds, 30 seconds, or 1 minute. In some embodiments, the perception system performs the process 900 continuously. For example, once a round of the process 900 ends, the process 900 restarts or reiterates. In some embodiments, the perception system performs the process 900 in response to a triggering event, e.g., an input from a driver.”
}
and a processor configured to control driving of the vehicle based on the vectorized map.
{Para [0030] “For each agent, the path-based trajectory prediction can include multiple operations: 1) vectorizing map into connected lane segments; 2) sampling the vectorized map for candidate reference paths (e.g., in 8 seconds) with reachable lane segments or reachable targets (e.g., end points) of the candidate reference paths; 3) classifying a set of candidate reference paths (e.g., by predicting a discrete probability distribution over the candidate reference paths) based on defined feature vectors, including scene feature vector (e.g., agent behavior) and path feature vector (e.g., first point, middle point, last point, direction, and length of each candidate reference path); 4) making trajectory prediction with respect to one or more selected reference paths in the Frenet frame using agents feature map augmented with path information; and 5) transforming the predicted trajectories back to Cartesian co-ordinates relative to the agent to obtain multimodal predictions.”
Para [0031] “Some of the advantages of these techniques are as follows. For example, the techniques predict trajectories conditioned on feature descriptors of a complete reference path from the agent's current location to the agent's goal instead of just its goal locations. This is a much more informative feature descriptor and leads to more map compliant trajectories over longer prediction horizons compared to goal based prediction. Also, the techniques use reference paths, which allow to predict trajectories in the path relative Frenet frame relative to each sampled path. Compared to the Cartesian frame with varying lane locations and curvatures, predictions in the Frenet frame can have much lower variance. This again leads to more map compliant trajectories that better generalize to novel scene layouts. Moreover, compared to using a rasterized HD map for its scene and reference path encoders, the techniques directly encode the scene and reference paths using polylines, making the encoders more efficient. The techniques can sample and classify variable length reference paths along each lane centerline, which provides trajectory prediction with more flexibility to predict different motion profiles along lanes. The techniques can improve path prediction and path compliance, e.g., using agent past trajectory history in the prediction. The techniques can enhance performance of prediction in multi-lane turns with better path classifier and scene upsampling. In addition to standard metrics for multimodal prediction, the techniques can enhance two map compliance metrics of the predicted trajectories (e.g., commonly used drivable area compliance metric and a new lane deviation metric), for example, by utilizing map prior knowledge (e.g., high likelihood drivable areas). Further, the techniques can improve interaction reasoning in path encoder and improve the map and agents interaction graph. The techniques can improve reaction of autonomous vehicles to surrounding environments (e.g., periodically or continuously) to achieve reliable and accurate prediction for their own route/trajectory or operation planning, which realizes safe and reliable driving.”
Para [0063] “In some embodiments, device 300 performs one or more processes described herein. Device 300 performs these processes based on processor 304 executing software instructions stored by a computer-readable medium, such as memory 305 and/or storage component 308. A computer-readable medium (e.g., a non-transitory computer readable medium) is defined herein as a non-transitory memory device. A non-transitory memory device includes memory space located inside a single physical storage device or memory space spread across multiple physical storage devices.”
}
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Liu to incorporate the teachings of Afshar to use the vectorized map generated from images for autonomous driving because Para [0030] “. The techniques can improve reaction of autonomous vehicles to surrounding environments (e.g., periodically or continuously) to achieve reliable and accurate prediction for their own route/trajectory or operation planning, which realizes safe and reliable driving.”
Liu in view of Afshar does not teach based on first memory tokens stored in a memory corresponding to queries of previous image frames of the image frames used to generate a previous vectorized map
However Yasarla teaches generate a spatial representation by predicting based on first memory tokens stored in a memory corresponding to queries of previous image frames of the image frames used to generate a previous spatial representation
{Para [0066] “A stream of images 420 may be provided by a monocular imaging system (e.g., the image capture devices 302) at a frame rate (e.g., 30 Hz, 60 Hz, etc.) and each image is encoded in the encoder 402. In some aspects, the encoder 402 is configured to identify features in each image, and the features are often referred to as vectors or tokens. In one aspect, the encoder 402 represents the image as query tokens, which are potential features of interest in the scene. As described in FIG. 6, the query tokens are provided to the depth estimator 404 and features within the query tokens are identified by the feature engine 408. In some aspects, the decoder 406 includes memory tokens, which are tokens that identify relevant features within the scene and are stored for use in connection with later images. In some cases, the memory tokens may be referred to as accumulated query information and represent relevant features of previous images. The feature engine 408 is configured to update and maintain the memory tokens based on the image 420. The memory tokens summarize and store key past information as features and the depth estimator 404 can cross-reference relevant features from previous frames when inferring depth on the image 420. The depth estimator 404 updates the memory tokens each image and stores the most relevant information.”
}
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Liu in view of Afshar to incorporate the teachings of Yasarla to use memory tokens and information from previous processing because it improves accuracy Para [0075] “In some aspects, the additional decoder information provided between decoding iterations of the decoder 506 may improve accuracy.”
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALEXANDER MATTA whose telephone number is (571)272-4296. The examiner can normally be reached Mon - Fri 10:00-6:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, James Lee can be reached at (571) 270-5965. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/A.G.M./Examiner, Art Unit 3668
/JAMES J LEE/Supervisory Patent Examiner, Art Unit 3668