DETAILED ACTION
Contents
Notice of Pre-AIA or AIA Status 2
Claim Rejections - 35 USC § 102 2
Claim Rejections - 35 USC § 103 10
Conclusion 14
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This action is responsive to applicant’s claim set received on 10/7/24. Claims 1-20 are currently pending.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless - (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention.(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1, 2, 4-5, 13, 15-16, 20 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Garimella et al (US 2022/0274625 A1). Regarding claim 1, Garimella discloses a graph-based driving environment perception method of a computer device comprising at least one processor, the graph-based driving environment perception method comprising: detecting, by the at least one processor, driving environment objects on the road in a vector form (see 0018, 0041-00043, 0019; The vectorization component 230 can include functionality for vectorizing and storing node representations of objects in the environment. As noted above, the vectorization component 230 may vectorize various map elements stored within the map(s) 224 (e.g., roads, lanes, crosswalks, buildings, trees, curbs, medians, street signs, traffic signals, etc.), as well as any entities or other objects detected and classified by the perception component 222. To vectorize a map elements or an entity, the vectorization component 230 may define one or more polylines representing the object, where each polyline includes one or more line segments. The polylines may convey the size, shape, and position of an object (e.g., a map element or an entity perceived in the environment), and the vectorization component 230 may encode and aggregate the polylines into a node data structure (or node) representing the object.
[0042] In some examples, nodes representing the objects in an environment also may include attributes associated with the objects, and the vectorization component 230 may store the associated attributes within the node. For map element objects, the vectorization component 230 may determine various attributes from the map(s) 224 and/or external map servers, including the map element type (e.g., such road, lane, crosswalk, sidewalk, street sign, etc.) and/or subtypes (e.g., lane type, street sign type, etc.), and additional elements such as directionality (e.g., intended direction of travel), permissibility (e.g., types of entities permitted to travel on the map element), and the like. For entity objects perceived in the environment, the vectorization component 230 can receive data from the perception component 222 to determine attribute information of the entity objects over time.
[0043] In some examples, attributes of an entity (e.g., a car, a bicycle, a pedestrian, etc.) can be determined based on sensor data captured over time, and can include, but are not limited to, one or more of a position of the entity at a time (e.g., wherein the position can be represented in the frame of reference discussed above), a velocity of the entity at the time (e.g., a magnitude and/or angle with respect to the first axis (or other reference line)), an acceleration of the entity at the time, an indication of whether the entity is in a drivable area (e.g., whether the entity is on a sidewalk or a road), an indication of whether the entity is in a crosswalk region, an entity context (e.g., a presence of vehicles or other objects in the environment of the entity and attribute(s) associated with the entity), an object association (e.g., whether the entity is travelling in a group of entities of the same type), and the like).
Regarding claim 13, Garimella discloses a computer device comprising: at least one processor configured to execute computer-readable instructions wherein the at least one processor causes the computer device to (see 0033, 0089-0090; Process 700 is illustrated as a collection of blocks in a logical flow diagram, which represent a sequence of operations, some or all of which can be implemented in hardware, software, or a combination thereof. In the context of software, the blocks represent computer-executable instructions stored on one or more computer-readable media that, which when executed by one or more processors, perform the recited operations. Generally, computer-executable instructions include routines, programs, objects, components, encryption, deciphering, compressing, recording, data structures and the like that perform particular functions or implement particular abstract data types. The order in which the operations are described should not be construed as a limitation. Any number of the described blocks can be combined in any order and/or in parallel to implement the processes, or alternative processes, and not all of the blocks need be executed. For discussion purposes, the processes herein are described with reference to the frameworks, architectures and environments described in the examples herein, although the processes may be implemented in a wide variety of other frameworks, architectures or environments), detect driving environment objects on the road in a vector form (see 0018, 0041-00043, 0019; The vectorization component 230 can include functionality for vectorizing and storing node representations of objects in the environment. As noted above, the vectorization component 230 may vectorize various map elements stored within the map(s) 224 (e.g., roads, lanes, crosswalks, buildings, trees, curbs, medians, street signs, traffic signals, etc.), as well as any entities or other objects detected and classified by the perception component 222. To vectorize a map elements or an entity, the vectorization component 230 may define one or more polylines representing the object, where each polyline includes one or more line segments. The polylines may convey the size, shape, and position of an object (e.g., a map element or an entity perceived in the environment), and the vectorization component 230 may encode and aggregate the polylines into a node data structure (or node) representing the object.
[0042] In some examples, nodes representing the objects in an environment also may include attributes associated with the objects, and the vectorization component 230 may store the associated attributes within the node. For map element objects, the vectorization component 230 may determine various attributes from the map(s) 224 and/or external map servers, including the map element type (e.g., such road, lane, crosswalk, sidewalk, street sign, etc.) and/or subtypes (e.g., lane type, street sign type, etc.), and additional elements such as directionality (e.g., intended direction of travel), permissibility (e.g., types of entities permitted to travel on the map element), and the like. For entity objects perceived in the environment, the vectorization component 230 can receive data from the perception component 222 to determine attribute information of the entity objects over time.
[0043] In some examples, attributes of an entity (e.g., a car, a bicycle, a pedestrian, etc.) can be determined based on sensor data captured over time, and can include, but are not limited to, one or more of a position of the entity at a time (e.g., wherein the position can be represented in the frame of reference discussed above), a velocity of the entity at the time (e.g., a magnitude and/or angle with respect to the first axis (or other reference line)), an acceleration of the entity at the time, an indication of whether the entity is in a drivable area (e.g., whether the entity is on a sidewalk or a road), an indication of whether the entity is in a crosswalk region, an entity context (e.g., a presence of vehicles or other objects in the environment of the entity and attribute(s) associated with the entity), an object association (e.g., whether the entity is travelling in a group of entities of the same type), and the like), and model the driving environment objects from the vector form to a graph representation (see 0020, 0045, figs. 1, 4-5; The graph neural network (GNN) component 232 can include functionality to generate and maintain a GNN including nodes representing map element objects and entity objects within the environment of the vehicle 202. As noted above, the vectorization component 230 may generate vectorized representations of map elements within the map(s) 224, and entities or other objects perceived by the perception component 222, and the GNN component 232 may store the vectorized representations as nodes in a graph structure. Additionally, the GNN component 232 may include functionality to execute inference operations on the graph structure, using machine-learning techniques to determine inferred updated states of the nodes and edge features within the graph structure… At operation 132, the autonomous vehicle generates and/or updates a GNN (or other graph structure) to include the vectorized map element nodes and/or the vectorized entity nodes described above. An example graph structure 136 of a GNN is depicted in box 134. In some cases, a GNN component within the autonomous vehicle may receive vectorized representations of objects (e.g., map elements and/or entities) from the vectorization component, and may create new nodes within the GNN, remove nodes from the GNN, and/or modify existing nodes of the GNN based on the received map data and/or entity data. Additionally, the GNN component may create and maintain edge features associated with node-pairs in the GNN. As noted above, the nodes in the GNN may store sets of attributes representing an object (e.g., a map element node or an entity node), and the edge features may include data indicating the relative positions and/or the relative yaw values of pairs of nodes. Though not depicted in FIG. 1 for clarity of illustration, in some examples, the GNN may be a fully connected structure in which each distinct pair of nodes is associated with a unique edge feature and/or edge data).
Regarding claim 20, Garimella discloses a non-transitory computer-readable recording medium storing instructions that, when executed by a processor, (see 0033, 0089-0090; Process 700 is illustrated as a collection of blocks in a logical flow diagram, which represent a sequence of operations, some or all of which can be implemented in hardware, software, or a combination thereof. In the context of software, the blocks represent computer-executable instructions stored on one or more computer-readable media that, which when executed by one or more processors, perform the recited operations. Generally, computer-executable instructions include routines, programs, objects, components, encryption, deciphering, compressing, recording, data structures and the like that perform particular functions or implement particular abstract data types. The order in which the operations are described should not be construed as a limitation. Any number of the described blocks can be combined in any order and/or in parallel to implement the processes, or alternative processes, and not all of the blocks need be executed. For discussion purposes, the processes herein are described with reference to the frameworks, architectures and environments described in the examples herein, although the processes may be implemented in a wide variety of other frameworks, architectures or environments) cause the processor to execute a graph-based driving environment perception method comprising: detecting driving environment objects on the road in a vector form (see 0018, 0041-00043, 0019; The vectorization component 230 can include functionality for vectorizing and storing node representations of objects in the environment. As noted above, the vectorization component 230 may vectorize various map elements stored within the map(s) 224 (e.g., roads, lanes, crosswalks, buildings, trees, curbs, medians, street signs, traffic signals, etc.), as well as any entities or other objects detected and classified by the perception component 222. To vectorize a map elements or an entity, the vectorization component 230 may define one or more polylines representing the object, where each polyline includes one or more line segments. The polylines may convey the size, shape, and position of an object (e.g., a map element or an entity perceived in the environment), and the vectorization component 230 may encode and aggregate the polylines into a node data structure (or node) representing the object.
[0042] In some examples, nodes representing the objects in an environment also may include attributes associated with the objects, and the vectorization component 230 may store the associated attributes within the node. For map element objects, the vectorization component 230 may determine various attributes from the map(s) 224 and/or external map servers, including the map element type (e.g., such road, lane, crosswalk, sidewalk, street sign, etc.) and/or subtypes (e.g., lane type, street sign type, etc.), and additional elements such as directionality (e.g., intended direction of travel), permissibility (e.g., types of entities permitted to travel on the map element), and the like. For entity objects perceived in the environment, the vectorization component 230 can receive data from the perception component 222 to determine attribute information of the entity objects over time.
[0043] In some examples, attributes of an entity (e.g., a car, a bicycle, a pedestrian, etc.) can be determined based on sensor data captured over time, and can include, but are not limited to, one or more of a position of the entity at a time (e.g., wherein the position can be represented in the frame of reference discussed above), a velocity of the entity at the time (e.g., a magnitude and/or angle with respect to the first axis (or other reference line)), an acceleration of the entity at the time, an indication of whether the entity is in a drivable area (e.g., whether the entity is on a sidewalk or a road), an indication of whether the entity is in a crosswalk region, an entity context (e.g., a presence of vehicles or other objects in the environment of the entity and attribute(s) associated with the entity), an object association (e.g., whether the entity is travelling in a group of entities of the same type), and the like); and modeling the driving environment objects from the vector form to a graph representation (see 0020, 0045, figs. 1, 4-5; The graph neural network (GNN) component 232 can include functionality to generate and maintain a GNN including nodes representing map element objects and entity objects within the environment of the vehicle 202. As noted above, the vectorization component 230 may generate vectorized representations of map elements within the map(s) 224, and entities or other objects perceived by the perception component 222, and the GNN component 232 may store the vectorized representations as nodes in a graph structure. Additionally, the GNN component 232 may include functionality to execute inference operations on the graph structure, using machine-learning techniques to determine inferred updated states of the nodes and edge features within the graph structure… At operation 132, the autonomous vehicle generates and/or updates a GNN (or other graph structure) to include the vectorized map element nodes and/or the vectorized entity nodes described above. An example graph structure 136 of a GNN is depicted in box 134. In some cases, a GNN component within the autonomous vehicle may receive vectorized representations of objects (e.g., map elements and/or entities) from the vectorization component, and may create new nodes within the GNN, remove nodes from the GNN, and/or modify existing nodes of the GNN based on the received map data and/or entity data. Additionally, the GNN component may create and maintain edge features associated with node-pairs in the GNN. As noted above, the nodes in the GNN may store sets of attributes representing an object (e.g., a map element node or an entity node), and the edge features may include data indicating the relative positions and/or the relative yaw values of pairs of nodes. Though not depicted in FIG. 1 for clarity of illustration, in some examples, the GNN may be a fully connected structure in which each distinct pair of nodes is associated with a unique edge feature and/or edge data)
Regarding claims 2, 4, 5, Garimella discloses modeling, by the at least one processor, the driving environment objects from the vector form to a graph representation (see 0020, 0045);
detecting geometric information of the driving environment objects, semantic information indicating object types, and instance information indicating object classification through graph representation (see 0035, 0019, 0041, 0043, 0042, 0074, 0077, 0080, 0020-0022, 0045);
transforming points and lines constituting a vector of the driving environment objects to the graph representation expressed with nodes and edges that represent connectivity of the nodes (see 0015, 0067, 0073, 0070, 0021, 0074-0076).
Regarding claims 15-16, Garimella discloses detect geometric information of the driving environment objects, semantic information indicating object types, and instance information indicating object classification through graph representation (see 0019, 0035-0043, 0121, 0080, 0045, 0074, 0077, 0080);
transform points and lines constituting a vector of the driving environment objects to the graph representation expressed with nodes and edges that represents connectivity of the nodes (see 0015-0016, 0020-0021, 0045, 0070, 0073-0076).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimedinvention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 3, 7, 14, 18 are rejected under 35 U.S.C. 103 as being unpatentable over Garimella et al (US 2022/0274625 A1) in view of Li et al (CV: “Graph-based Topology Reasoning for Driving Scenes”).
Regarding claims 3, 7, Garimella teaches all elements as mentioned above in claim 1. Garimella does not teach expressly extracting an image feature map from multi-view camera input; and transforming the image feature map to a bird's-eye-view (BEV) space of a vehicular coordinate system using at least one of a convolutional neural network (CNN), a multi-layer perceptron (MLP), and a cross-attention; detecting vertices and edges that constitute a graph from a BEV feature map extracted from multi-view camera input; and computing a graph adjacency matrix using the vertices and the edges.
Li, in the same field of endeavor, teaches extracting an image feature map from multi-view camera input; and transforming the image feature map to a bird's-eye-view (BEV) space of a vehicular coordinate system using at least one of a convolutional neural network (CNN), a multi-layer perceptron (MLP), and a cross-attention (see section 3.1, 4.1.1, 3.2, 3.1, 4.1.1, 4.2);
detecting vertices and edges that constitute a graph from a BEV feature map extracted from multi-view camera input; and computing a graph adjacency matrix using the vertices and the edges (see section 3.1, 3.4.1, 3.2, 3.3.3, 3.4.2, 4.1.3).
It would have been obvious (before the effective filing date of the claimed invention) or (at the time the invention was made) to one of ordinary skill in the art to modify Garimella to utilize the cited limitations as suggested by Li. The suggestion/motivation for doing so would have been to enhance performance by perceptual and topological metrics (see abstract). Furthermore, the prior art collectively includes each element claimed (though not all in the same reference), and one of ordinary skill in the art could have combined the elements in the manner explained above using known engineering design, interface and/or programming techniques, without changing a “fundamental” operating principle of Garimella, while the teaching of Li continues to perform the same function as originally taught prior to being combined, in order to produce the repeatable and predictable result. It is for at least the aforementioned reasons that the examiner has reached a conclusion of obviousness with respect to the claim in question.
Regarding claims 14, 18, the claims are analyzed as a device that implements the limitations of claims 3, 7 (see rejection of claims 3 and 7).
Claim 6, 17 are rejected under 35 U.S.C. 103 as being unpatentable over Garimella et al (US 2022/0274625 A1) in view of Xu et al (CV: “CenterLineDet: CenterLine Graph Detection for Road Lanes with Vehicle-mounted Sensors by Transformer for HD Map Generation”).
Regarding claim 6, Garimella teaches all elements as mentioned above in claim 2. Garimella does not teach expressly modeling, by the at least one processor, the driving environment objects from the vector form to a graph representation.
Xu, in the same field of endeavor, teaches modeling, by the at least one processor, the driving environment objects from the vector form to a graph representation (see section III).
It would have been obvious (before the effective filing date of the claimed invention) or (at the time the invention was made) to one of ordinary skill in the art to modify Garimella to utilize the cited limitations as suggested by Xu. The suggestion/motivation for doing so would have been to enhance detection of centerlines for automatic HDF map generation (see abstract). Furthermore, the prior art collectively includes each element claimed (though not all in the same reference), and one of ordinary skill in the art could have combined the elements in the manner explained above using known engineering design, interface and/or programming techniques, without changing a “fundamental” operating principle of Garimella, while the teaching of Xu continues to perform the same function as originally taught prior to being combined, in order to produce the repeatable and predictable result. It is for at least the aforementioned reasons that the examiner has reached a conclusion of obviousness with respect to the claim in question.
Regarding claim 17, the claim is analyzed as a device that implements the limitations of claim 6 (see rejection of claim 6).
Allowable Subject Matter
Claims 8-12, 19 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Regarding claim 8, none of the references of record alone or in combination suggest or fairly teach wherein the detecting of the vertices and the edges comprises detecting the vertices and the edges that map elements consist of, through a CNN decoder using the BEV feature map.
Regarding claim 9-12, none of the references of record alone or in combination suggest or fairly teach representing graph node embeddings by combining the vertices and the edges; and predicting connection between nodes as an adjacency matrix with likelihood, based on similarity between the nodes.
Regarding claim 19, none of the references of record alone or in combination suggest or fairly teach detect the vertices and the edges that are map elements through a convolutional neural network (CNN) decoder using the BEV feature map, represent graph node embeddings by combining the vertices and the edges, and predict connection between nodes as an adjacency matrix with likelihood, based on similarity between the nodes.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to EDWARD PARK. The examiner’s contact information is as follows:
Telephone: (571)270-1576 | Fax: 571.270.2576 | Edward.Park@uspto.gov
For email communications, please notate MPEP 502.03, which outlines procedures pertaining to communications via the internet and authorization. A sample authorization form is cited within MPEP 502.03, section II.
The examiner can normally be reached on M-F 9-6 CST.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, John M. Villecco, can be reached on (571) 272-7319. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/EDWARD PARK/Primary Examiner, Art Unit 2661