DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Drawings
The drawings are objected to because Figure 4 includes a typographical error with the term “AverageAttributes in Samples” and the terms “Sequency” and because they include the following reference character(s) not mentioned in the description: 192 in Figure 1C. Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. The figure or figure number of an amended drawing should not be labeled as “amended.” If a drawing figure is to be canceled, the appropriate figure must be removed from the replacement sheet, and where necessary, the remaining figures must be renumbered and appropriate changes made to the brief description of the several views of the drawings for consistency. Additional replacement sheets may be necessary to show the renumbering of the remaining figures. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-2, 5-8, 14-15, and 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over BURLINA et al. (US 12525025 B1), hereinafter referenced as BURLINA, in view of YAN (US 20230086735 A1), hereinafter referenced as YAN.
Regarding claim 1, BURLINA explicitly teaches an apparatus for data collection (Fig. 5, #502 called vehicle with #506 called sensor system and #522 called memory. Col. 23, Line [4-10]-BURLINA discloses the vehicle 502 may include a vehicle computing device(s) 504, one or more sensor system(s) 506, emitter(s) 508, network interfaces 510, drive system(s) 512, and/or at least one direct connection 514. The vehicle computing device(s) 504 may represent the vehicle computing system(s) 122 and the sensor system(s) 506 may represent the sensor system(s) 116.), comprising:
at least one memory (Fig. 5, #522 called memory. Col. 25, Line [5-10]-BURLINA discloses memory 522 and/or 526 may be examples of non-transitory computer-readable media. The memory 522 and/or 526 may store an operating system and one or more software applications, instructions, programs, and/or data to implement the methods described herein and the functions attributed to the various systems.); and
at least one processor (Fig. 5, #520 called processor.) coupled to the at least one memory and configured to (Fig. 5, #520 called processor and #522 called memory. Col. 24, Lines [56-62]-BURLINA discloses the vehicle computing device(s) 504 may include processor(s) 520 and memory 522 communicatively coupled with the processor(s) 520. The computing device(s) 518 may also include processor(s) 524, and/or memory 526. The processor(s) 520 and/or 524 may be any suitable processor capable of executing instructions to process data and perform operations as described herein.):
detect a set of objects in an environment (Fig. 1. Col. 8, Lines [61-65]-BURLINA discloses the vehicle computing system(s) 122 may comprise an object detection component 124 configured to detect one or more objects (e.g., the objects 110, 112, and 114) in the environment 100 based on data from the sensor system(s) 116.) based on an obtained set of multimodal data from a plurality of sensors (Fig. 1. Col. 7, Lines [39-53]-BURLINA discloses the sensor data 120 may comprise multispectral data spanning several discrete spectral bands (e.g., ultra-violet to visible light, visible light to infra-red), and/or hyperspectral sensor(s) which may capture nearly continuous wavelengths spanning a wide range of the electromagnetic spectrum. Additionally, the sensor system(s) may include one or more LiDAR sensors, inertial sensors (e.g., inertial measurement units, accelerometers, magnetometers, gyroscopes, etc.), location sensors (e.g., a global positioning system (GPS)), depth sensors (e.g., stereo cameras and range cameras), radar sensors, time-of-flight sensors, sonar sensors, thermal imaging sensors, or any other sensor modalities, and the sensor data 120 may comprise data captured by any of these modalities.);
generate a scene graph (Fig. 4, #420 called graph. Col. 20, Lines [1-6]-BURLINA discloses in an example graph 420 illustrated in FIG. 4, which may represent at least the portion of the scene, each node 420A, 420B, 420C is associated with a spectral response as shown. In the example graph 420, the nodes 420A, 420B, 420C are connected by edges 424 indicating an adjacency or proximity relationship between the nodes.) based on the set of objects (Fig. 4. Col. 20, Lines [11-17]-BURLINA discloses a candidate region, which may contain a potential object, may be identified based on one or more sensor modalities, and the graph 420 may represent the candidate region. For example, the candidate region may be based on a point cloud from LiDAR sensors, and correspond to a cluster of points located a threshold height above a driving surface which may indicate a potential object.);
BURLINA fails to explicitly teach receive a query scene graph, wherein the query scene graph describes a scenario of interest; match the scene graph with the query scene graph; and output the scene graph based on a successful match between the scene graph and the query scene graph.
However, YAN explicitly teaches receive a query scene graph (Fig. 3, illustrates receiving a query scene graph in the scene graph matching block. Paragraph [0072]-YAN discloses the visual relationship system 102 can utilize the extracted objects 306 and relationship features 308 that are defined in the terms of the query 302 to perform query graph generation 310. A query graph 312 can be generated where objects 306 and relationship features 308 extracted from the terms of the query 302 are utilized as nodes 314 and edges 316 between nodes, respectively. Further in paragraph [0073]-YAN discloses the visual relationship system 102 can perform scene graph matching 318 between query graph 312 and scene graphs 202 from scene graph database 118.), wherein the query scene graph describes a scenario of interest (Fig. 3. Paragraph [0070]-YAN discloses a query 302 including terms descriptive of a visual relationship can be provided to the visual relationship system 102. Further in paragraph [0072]-YAN discloses the visual relationship system 102 can utilize the extracted objects 306 and relationship features 308 that are defined in the terms of the query 302 to perform query graph generation 310. A query graph 312 can be generated where objects 306 and relationship features 308 extracted from the terms of the query 302 are utilized as nodes 314 and edges 316 between nodes, respectively.);
match the scene graph with the query scene graph (Fig. 3. Paragraph [0073]-YAN discloses the visual relationship system 102 can perform scene graph matching 318 between query graph 312 and scene graphs 202 from scene graph database 118. In some implementations, the matching, which is further described below, between query graph 312 and scene graphs 202 from scene graph database 118 includes searching a scene graph index 216 to retrieve key frames 115 corresponding relevant videos 114 that are responsive to query 120.); and
output the scene graph based on a successful match between the scene graph and the query scene graph (Fig. 3. Paragraph [0073]-YAN discloses a set of scene graphs 202 that match the query graph 312 are selected from the scene graphs 202 in the scene graph database 118. The query graph 312 can be matched with indexes in the scene graph database 118 for retrieving relevant videos 114 and key frames 115 including respective timestamps 207 associated with the key frames 115 as query results (wherein the scene graphs 202 are the output scene graphs).).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of BURLINA of an apparatus for data collection, comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: detect a set of objects in an environment based on an obtained set of multimodal data from a plurality of sensors; generate a scene graph based on the set of objects; with the teachings of YAN of receive a query scene graph, wherein the query scene graph describes a scenario of interest; match the scene graph with the query scene graph; and output the scene graph based on a successful match between the scene graph and the query scene graph.
Wherein having BURLINA’s apparatus for receiving and processing data from a vehicle having receive a query scene graph, wherein the query scene graph describes a scenario of interest; match the scene graph with the query scene graph; and output the scene graph based on a successful match between the scene graph and the query scene graph.
The motivation behind the modification would have been to obtain an apparatus/method for receiving and processing data from a vehicle that enhances efficiently and accurately detects objects surrounding a vehicle and planning navigation for the vehicle. Since both BURLINA and YAN relate to detecting objects and generating scene graphs of the detected objects, wherein BURLINA using data from multiple spectral bands may allow for more robust detection of objects, resulting in improved collision avoidance with animals, debris, and/or other relatively smaller objects on driving surfaces, while YAN an advantage of this technology is that it can facilitate efficient and accurate discovery of videos and key frames within videos using natural language descriptions of visual relationships between objects depicted in the key frames, and may reduce a number of queries required to be entered by a user in order to find a particular video of interest. This in turn reduces the number of computer resources required to execute multiple queries until the appropriate video has been identified. Please see BURLINA et al. (US 12525025 B1), Col. 2, Line [3-26], and YAN (US 20230086735 A1), Paragraph [0014].
Regarding claim 2, BURLINA in view of YAN explicitly teach the apparatus of claim 1,
BURLINA further explicitly teaches wherein the at least one processor is configured to (Fig. 5, #520 called processor and #522 called memory. Col. 24, Lines [56-62]-BURLINA discloses the vehicle computing device(s) 504 may include processor(s) 520 and memory 522 communicatively coupled with the processor(s) 520. The computing device(s) 518 may also include processor(s) 524, and/or memory 526. The processor(s) 520 and/or 524 may be any suitable processor capable of executing instructions to process data and perform operations as described herein.):
BURLINA fails to explicitly teach receive a first description of a first scenario of interest, wherein the first description comprises a textual description of the first scenario of interest; and output the textual description of the first scenario of interest.
However, YAN explicitly teaches receive a first description of a first scenario of interest (Fig. 3. Paragraph [0070]-YAN discloses a query 302 including terms descriptive of a visual relationship can be provided to the visual relationship system 102. In some implementations, query 302 is a textual query that is generated by the speech-to-text converter 106 from a query 120 received by the visual relationship system 102 from a user on a user device 104.),
wherein the first description comprises a textual description of the first scenario of interest (Fig. 3. Paragraph [0071]-YAN discloses a query 302 is “I want a boy holding a ball” where the object-terms are determined as “boy” and “ball” and relationship feature-terms are determined as “holding.”); and
output the textual description of the first scenario of interest (Fig. 3, illustrates outputting the textual description (called parsed query #302) into Visual Relationship System #102. Paragraph [0071]-YAN discloses visual relationship system 102 can receive the query 302 as input and perform feature/object extraction 304 on the query 302 to determine terms of the query 302 defining objects 306 and relationship features 308.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of BURLINA in view of YAN of an apparatus for data collection, comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: detect a set of objects in an environment based on an obtained set of multimodal data from a plurality of sensors; generate a scene graph based on the set of objects; with the teachings of YAN of receive a first description of a first scenario of interest, wherein the first description comprises a textual description of the first scenario of interest; and output the textual description of the first scenario of interest.
Wherein having BURLINA’s apparatus for receiving and processing data from a vehicle having receive a first description of a first scenario of interest, wherein the first description comprises a textual description of the first scenario of interest; and output the textual description of the first scenario of interest.
The motivation behind the modification would have been to obtain an apparatus/method for receiving and processing data from a vehicle that enhances efficiently and accurately detects objects surrounding a vehicle and planning navigation for the vehicle. Since both BURLINA and YAN relate to detecting objects and generating scene graphs of the detected objects, wherein BURLINA using data from multiple spectral bands may allow for more robust detection of objects, resulting in improved collision avoidance with animals, debris, and/or other relatively smaller objects on driving surfaces, while YAN an advantage of this technology is that it can facilitate efficient and accurate discovery of videos and key frames within videos using natural language descriptions of visual relationships between objects depicted in the key frames, and may reduce a number of queries required to be entered by a user in order to find a particular video of interest. This in turn reduces the number of computer resources required to execute multiple queries until the appropriate video has been identified. Please see BURLINA et al. (US 12525025 B1), Col. 2, Line [3-26], and YAN (US 20230086735 A1), Paragraph [0014].
Regarding claim 5, BURLINA in view of YAN explicitly teach the apparatus of claim 1,
BURLINA further explicitly teaches wherein the at least one processor is configured to (Fig. 5, #520 called processor and #522 called memory. Col. 24, Lines [56-62]-BURLINA discloses the vehicle computing device(s) 504 may include processor(s) 520 and memory 522 communicatively coupled with the processor(s) 520. The computing device(s) 518 may also include processor(s) 524, and/or memory 526. The processor(s) 520 and/or 524 may be any suitable processor capable of executing instructions to process data and perform operations as described herein.):
BURLINA fails to explicitly teach determine a relationship between a first object in the set of objects and a second object in the set of objects; and encode the relationship in the scene graph.
However, YAN explicitly teaches determine a relationship between a first object in the set of objects and a second object in the set of objects (Fig. 2A. Paragraph [0061]-YAN discloses each relationship feature 212 defines a relationship between a first object and a second, different object. For example, a relationship feature 212 can be “holding,” where the relationship feature 212 defines a relationship between a first object “boy” and a second object “ball,” to define a visual relationship of “boy” “holding” “ball.”); and
encode the relationship in the scene graph (Fig. 2A. Paragraph [0051]-YAN discloses a scene graph 202 includes a set of nodes 204 and a set of edges 206 that interconnect a subset of nodes in the set of nodes. Each scene graph 202 can define a set of objects that are represented by respective nodes 204, e.g., where a first object is represented by a first node from the set of nodes, and a second object is represented by a second node from the set of nodes. The first node and the second node can be connected by an edge representing a relationship feature that is defining of a relationship between the two objects.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of BURLINA in view of YAN of an apparatus for data collection, comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: detect a set of objects in an environment based on an obtained set of multimodal data from a plurality of sensors; generate a scene graph based on the set of objects; with the teachings of YAN of determine a relationship between a first object in the set of objects and a second object in the set of objects; and encode the relationship in the scene graph.
Wherein having BURLINA’s apparatus for receiving and processing data from a vehicle having determine a relationship between a first object in the set of objects and a second object in the set of objects; and encode the relationship in the scene graph.
The motivation behind the modification would have been to obtain an apparatus/method for receiving and processing data from a vehicle that enhances efficiently and accurately detects objects surrounding a vehicle and planning navigation for the vehicle. Since both BURLINA and YAN relate to detecting objects and generating scene graphs of the detected objects, wherein BURLINA using data from multiple spectral bands may allow for more robust detection of objects, resulting in improved collision avoidance with animals, debris, and/or other relatively smaller objects on driving surfaces, while YAN an advantage of this technology is that it can facilitate efficient and accurate discovery of videos and key frames within videos using natural language descriptions of visual relationships between objects depicted in the key frames, and may reduce a number of queries required to be entered by a user in order to find a particular video of interest. This in turn reduces the number of computer resources required to execute multiple queries until the appropriate video has been identified. Please see BURLINA et al. (US 12525025 B1), Col. 2, Line [3-26], and YAN (US 20230086735 A1), Paragraph [0014].
Regarding claim 6, BURLINA in view of YAN explicitly teach the apparatus of claim 5,
BURLINA fails to explicitly teach wherein the relationship comprises at least one of a distance between the first object and the second object, or an intent of the first object with respect to the second object.
However, YAN explicitly teaches wherein the relationship comprises at least one of a distance between the first object and the second object (Fig. 2A. Paragraph [0061]-YAN discloses relationships can be determined by the visual relationship model 108, for example, based in part on proximity/spatial distances between objects, known relationships between categories of objects, user-defined relationships between particular objects and/or categories of objects, or the like.), or an intent of the first object with respect to the second object (Fig. 2A. Paragraph [0061]-YAN discloses a relationship feature 212 can be “holding,” where the relationship feature 212 defines a relationship between a first object “boy” and a second object “ball,” to define a visual relationship of “boy” “holding” “ball” (wherein the intent is holding).).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of BURLINA in view of YAN of an apparatus for data collection, comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: detect a set of objects in an environment based on an obtained set of multimodal data from a plurality of sensors; generate a scene graph based on the set of objects; with the teachings of YAN of wherein the relationship comprises at least one of a distance between the first object and the second object, or an intent of the first object with respect to the second object.
Wherein having BURLINA’s apparatus for receiving and processing data from a vehicle having wherein the relationship comprises at least one of a distance between the first object and the second object, or an intent of the first object with respect to the second object.
The motivation behind the modification would have been to obtain an apparatus/method for receiving and processing data from a vehicle that enhances efficiently and accurately detects objects surrounding a vehicle and planning navigation for the vehicle. Since both BURLINA and YAN relate to detecting objects and generating scene graphs of the detected objects, wherein BURLINA using data from multiple spectral bands may allow for more robust detection of objects, resulting in improved collision avoidance with animals, debris, and/or other relatively smaller objects on driving surfaces, while YAN an advantage of this technology is that it can facilitate efficient and accurate discovery of videos and key frames within videos using natural language descriptions of visual relationships between objects depicted in the key frames, and may reduce a number of queries required to be entered by a user in order to find a particular video of interest. This in turn reduces the number of computer resources required to execute multiple queries until the appropriate video has been identified. Please see BURLINA et al. (US 12525025 B1), Col. 2, Line [3-26], and YAN (US 20230086735 A1), Paragraph [0014].
Regarding claim 7, BURLINA in view of YAN explicitly teach the apparatus of claim 5,
BURLINA further explicitly teaches wherein the at least one processor is configured to (Fig. 5, #520 called processor and #522 called memory. Col. 24, Lines [56-62]-BURLINA discloses the vehicle computing device(s) 504 may include processor(s) 520 and memory 522 communicatively coupled with the processor(s) 520. The computing device(s) 518 may also include processor(s) 524, and/or memory 526. The processor(s) 520 and/or 524 may be any suitable processor capable of executing instructions to process data and perform operations as described herein.):
BURLINA fails to explicitly teach encode the first object as a first node in the scene graph; encode the second object as a second node in the scene graph; and encode the relationship as an edge between the first node and the second node.
However, YAN explicitly teaches encode the first object as a first node in the scene graph (Fig. 2A. Paragraph [0051]-YAN discloses each scene graph 202 can define a set of objects that are represented by respective nodes 204, e.g., where a first object is represented by a first node from the set of nodes, and a second object is represented by a second node from the set of nodes. The first node and the second node can be connected by an edge representing a relationship feature that is defining of a relationship between the two objects.);
encode the second object as a second node in the scene graph (Fig. 2A. Paragraph [0051]-YAN discloses each scene graph 202 can define a set of objects that are represented by respective nodes 204, e.g., where a first object is represented by a first node from the set of nodes, and a second object is represented by a second node from the set of nodes. The first node and the second node can be connected by an edge representing a relationship feature that is defining of a relationship between the two objects.); and
encode the relationship as an edge between the first node and the second node (Fig. 2A. Paragraph [0051]-YAN discloses each scene graph 202 can define a set of objects that are represented by respective nodes 204, e.g., where a first object is represented by a first node from the set of nodes, and a second object is represented by a second node from the set of nodes. The first node and the second node can be connected by an edge representing a relationship feature that is defining of a relationship between the two objects.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of BURLINA in view of YAN of an apparatus for data collection, comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: detect a set of objects in an environment based on an obtained set of multimodal data from a plurality of sensors; generate a scene graph based on the set of objects; with the teachings of YAN of encode the first object as a first node in the scene graph; encode the second object as a second node in the scene graph; and encode the relationship as an edge between the first node and the second node.
Wherein having BURLINA’s apparatus for receiving and processing data from a vehicle having encode the first object as a first node in the scene graph; encode the second object as a second node in the scene graph; and encode the relationship as an edge between the first node and the second node.
The motivation behind the modification would have been to obtain an apparatus/method for receiving and processing data from a vehicle that enhances efficiently and accurately detects objects surrounding a vehicle and planning navigation for the vehicle. Since both BURLINA and YAN relate to detecting objects and generating scene graphs of the detected objects, wherein BURLINA using data from multiple spectral bands may allow for more robust detection of objects, resulting in improved collision avoidance with animals, debris, and/or other relatively smaller objects on driving surfaces, while YAN an advantage of this technology is that it can facilitate efficient and accurate discovery of videos and key frames within videos using natural language descriptions of visual relationships between objects depicted in the key frames, and may reduce a number of queries required to be entered by a user in order to find a particular video of interest. This in turn reduces the number of computer resources required to execute multiple queries until the appropriate video has been identified. Please see BURLINA et al. (US 12525025 B1), Col. 2, Line [3-26], and YAN (US 20230086735 A1), Paragraph [0014].
Regarding claim 8, BURLINA in view of YAN explicitly teach the apparatus of claim 1,
BURLINA further explicitly teaches wherein the obtained set of multimodal data includes at least one of an image (Fig. 5, #506 called sensor system. Col. 23, Line [11-17]-BURLINA discloses the sensor system(s) 506 may include LiDAR sensors, radar sensors, ultrasonic transducers, sonar sensors, location sensors (e.g., global positioning system (GPS), compass, etc.), inertial sensors (e.g., inertial measurement units (IMUs), accelerometers, magnetometers, gyroscopes, etc.), image sensors (e.g., red-green-blue (RGB) cameras, intensity, depth, time of flight cameras, etc.). Further in Col. 23, Line [27-29]-BURLINA discloses the sensor system(s) 506 may provide sensor data to the vehicle computing device(s) 504.), a light detection and ranging (LIDAR) data (Fig. 5, #506 called sensor system. Col. 23, Line [11-17]-BURLINA discloses the sensor system(s) 506 may include LiDAR sensors, radar sensors, ultrasonic transducers, sonar sensors, location sensors (e.g., global positioning system (GPS), compass, etc.), inertial sensors (e.g., inertial measurement units (IMUs), accelerometers, magnetometers, gyroscopes, etc.), image sensors (e.g., red-green-blue (RGB) cameras, intensity, depth, time of flight cameras, etc.). Further in Col. 23, Line [27-29]-BURLINA discloses the sensor system(s) 506 may provide sensor data to the vehicle computing device(s) 504.), or radio detection and ranging (RADAR) data (Fig. 5, #506 called sensor system. Col. 23, Line [11-17]-BURLINA discloses the sensor system(s) 506 may include LiDAR sensors, radar sensors, ultrasonic transducers, sonar sensors, location sensors (e.g., global positioning system (GPS), compass, etc.), inertial sensors (e.g., inertial measurement units (IMUs), accelerometers, magnetometers, gyroscopes, etc.), image sensors (e.g., red-green-blue (RGB) cameras, intensity, depth, time of flight cameras, etc.). Further in Col. 23, Line [27-29]-BURLINA discloses the sensor system(s) 506 may provide sensor data to the vehicle computing device(s) 504.).
Regarding claim 14, BURLINA explicitly teaches a method for data collection (Fig. 4. Col. 17, Line [46-48]-BURLINA discloses FIG. 4 includes textual and visual flowcharts to illustrate an example process 400 for determining a presence of an object from multispectral data.), comprising:
detecting a set of objects in an environment (Fig. 1. Col. 8, Lines [61-65]-BURLINA discloses the vehicle computing system(s) 122 may comprise an object detection component 124 configured to detect one or more objects (e.g., the objects 110, 112, and 114) in the environment 100 based on data from the sensor system(s) 116.) based on an obtained set of multimodal data from a plurality of sensors (Fig. 1. Col. 7, Lines [39-53]-BURLINA discloses the sensor data 120 may comprise multispectral data spanning several discrete spectral bands (e.g., ultra-violet to visible light, visible light to infra-red), and/or hyperspectral sensor(s) which may capture nearly continuous wavelengths spanning a wide range of the electromagnetic spectrum. Additionally, the sensor system(s) may include one or more LiDAR sensors, inertial sensors (e.g., inertial measurement units, accelerometers, magnetometers, gyroscopes, etc.), location sensors (e.g., a global positioning system (GPS)), depth sensors (e.g., stereo cameras and range cameras), radar sensors, time-of-flight sensors, sonar sensors, thermal imaging sensors, or any other sensor modalities, and the sensor data 120 may comprise data captured by any of these modalities.);
generating a scene graph (Fig. 4, #420 called graph. Col. 20, Lines [1-6]-BURLINA discloses in an example graph 420 illustrated in FIG. 4, which may represent at least the portion of the scene, each node 420A, 420B, 420C is associated with a spectral response as shown. In the example graph 420, the nodes 420A, 420B, 420C are connected by edges 424 indicating an adjacency or proximity relationship between the nodes.) based on the set of objects (Fig. 4. Col. 20, Lines [11-17]-BURLINA discloses a candidate region, which may contain a potential object, may be identified based on one or more sensor modalities, and the graph 420 may represent the candidate region. For example, the candidate region may be based on a point cloud from LiDAR sensors, and correspond to a cluster of points located a threshold height above a driving surface which may indicate a potential object.);
BURLINA fails to explicitly teach receiving a query scene graph, wherein the query scene graph describes a scenario of interest; matching the scene graph with the query scene graph; and outputting the scene graph based on a successful match between the scene graph and the query scene graph.
However, YAN explicitly teaches receiving a query scene graph (Fig. 3, illustrates receiving a query scene graph in the scene graph matching block. Paragraph [0072]-YAN discloses the visual relationship system 102 can utilize the extracted objects 306 and relationship features 308 that are defined in the terms of the query 302 to perform query graph generation 310. A query graph 312 can be generated where objects 306 and relationship features 308 extracted from the terms of the query 302 are utilized as nodes 314 and edges 316 between nodes, respectively. Further in paragraph [0073]-YAN discloses the visual relationship system 102 can perform scene graph matching 318 between query graph 312 and scene graphs 202 from scene graph database 118.),
wherein the query scene graph describes a scenario of interest (Fig. 3. Paragraph [0070]-YAN discloses a query 302 including terms descriptive of a visual relationship can be provided to the visual relationship system 102. Further in paragraph [0072]-YAN discloses the visual relationship system 102 can utilize the extracted objects 306 and relationship features 308 that are defined in the terms of the query 302 to perform query graph generation 310. A query graph 312 can be generated where objects 306 and relationship features 308 extracted from the terms of the query 302 are utilized as nodes 314 and edges 316 between nodes, respectively.);
matching the scene graph with the query scene graph (Fig. 3. Paragraph [0073]-YAN discloses the visual relationship system 102 can perform scene graph matching 318 between query graph 312 and scene graphs 202 from scene graph database 118. In some implementations, the matching, which is further described below, between query graph 312 and scene graphs 202 from scene graph database 118 includes searching a scene graph index 216 to retrieve key frames 115 corresponding relevant videos 114 that are responsive to query 120.); and
outputting the scene graph based on a successful match between the scene graph and the query scene graph (Fig. 3. Paragraph [0073]-YAN discloses a set of scene graphs 202 that match the query graph 312 are selected from the scene graphs 202 in the scene graph database 118. The query graph 312 can be matched with indexes in the scene graph database 118 for retrieving relevant videos 114 and key frames 115 including respective timestamps 207 associated with the key frames 115 as query results (wherein the scene graphs 202 are the output scene graphs).).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of BURLINA of a method for data collection, comprising: detecting a set of objects in an environment based on an obtained set of multimodal data from a plurality of sensors; generating a scene graph based on the set of objects; with the teachings of YAN of receiving a query scene graph, wherein the query scene graph describes a scenario of interest; matching the scene graph with the query scene graph; and outputting the scene graph based on a successful match between the scene graph and the query scene graph.
Wherein having BURLINA’s apparatus for receiving and processing data from a vehicle having receiving a query scene graph, wherein the query scene graph describes a scenario of interest; matching the scene graph with the query scene graph; and outputting the scene graph based on a successful match between the scene graph and the query scene graph.
The motivation behind the modification would have been to obtain an apparatus/method for receiving and processing data from a vehicle that enhances efficiently and accurately detects objects surrounding a vehicle and planning navigation for the vehicle. Since both BURLINA and YAN relate to detecting objects and generating scene graphs of the detected objects, wherein BURLINA using data from multiple spectral bands may allow for more robust detection of objects, resulting in improved collision avoidance with animals, debris, and/or other relatively smaller objects on driving surfaces, while YAN an advantage of this technology is that it can facilitate efficient and accurate discovery of videos and key frames within videos using natural language descriptions of visual relationships between objects depicted in the key frames, and may reduce a number of queries required to be entered by a user in order to find a particular video of interest. This in turn reduces the number of computer resources required to execute multiple queries until the appropriate video has been identified. Please see BURLINA et al. (US 12525025 B1), Col. 2, Line [3-26], and YAN (US 20230086735 A1), Paragraph [0014].
Regarding claim 15, BURLINA in view of YAN explicitly teach the method of claim 14, further comprising:
BURLINA fails to explicitly teach receiving a first description of a first scenario of interest, wherein the first description comprises a textual description of the first scenario of interest; and outputting the textual description of the first scenario of interest.
However, YAN explicitly teaches receiving a first description of a first scenario of interest (Fig. 3. Paragraph [0070]-YAN discloses a query 302 including terms descriptive of a visual relationship can be provided to the visual relationship system 102. In some implementations, query 302 is a textual query that is generated by the speech-to-text converter 106 from a query 120 received by the visual relationship system 102 from a user on a user device 104.),
wherein the first description comprises a textual description of the first scenario of interest (Fig. 3. Paragraph [0071]-YAN discloses a query 302 is “I want a boy holding a ball” where the object-terms are determined as “boy” and “ball” and relationship feature-terms are determined as “holding.”); and
outputting the textual description of the first scenario of interest (Fig. 3, illustrates outputting the textual description (called parsed query #302) into Visual Relationship System #102. Paragraph [0071]-YAN discloses visual relationship system 102 can receive the query 302 as input and perform feature/object extraction 304 on the query 302 to determine terms of the query 302 defining objects 306 and relationship features 308.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of BURLINA in view of YAN of a method for data collection, comprising: detecting a set of objects in an environment based on an obtained set of multimodal data from a plurality of sensors; generating a scene graph based on the set of objects; with the teachings of YAN of receiving a first description of a first scenario of interest, wherein the first description comprises a textual description of the first scenario of interest; and outputting the textual description of the first scenario of interest.
Wherein having BURLINA’s apparatus for receiving and processing data from a vehicle having receiving a first description of a first scenario of interest, wherein the first description comprises a textual description of the first scenario of interest; and outputting the textual description of the first scenario of interest.
The motivation behind the modification would have been to obtain an apparatus/method for receiving and processing data from a vehicle that enhances efficiently and accurately detects objects surrounding a vehicle and planning navigation for the vehicle. Since both BURLINA and YAN relate to detecting objects and generating scene graphs of the detected objects, wherein BURLINA using data from multiple spectral bands may allow for more robust detection of objects, resulting in improved collision avoidance with animals, debris, and/or other relatively smaller objects on driving surfaces, while YAN an advantage of this technology is that it can facilitate efficient and accurate discovery of videos and key frames within videos using natural language descriptions of visual relationships between objects depicted in the key frames, and may reduce a number of queries required to be entered by a user in order to find a particular video of interest. This in turn reduces the number of computer resources required to execute multiple queries until the appropriate video has been identified. Please see BURLINA et al. (US 12525025 B1), Col. 2, Line [3-26], and YAN (US 20230086735 A1), Paragraph [0014].
Regarding claim 18, BURLINA in view of YAN explicitly teach the method of claim 14, further comprising:
BURLINA fails to explicitly teach determining a relationship between a first object in the set of objects and a second object in the set of objects; and encoding the relationship in the scene graph.
However, YAN explicitly teaches determining a relationship between a first object in the set of objects and a second object in the set of objects (Fig. 2A. Paragraph [0061]-YAN discloses each relationship feature 212 defines a relationship between a first object and a second, different object. For example, a relationship feature 212 can be “holding,” where the relationship feature 212 defines a relationship between a first object “boy” and a second object “ball,” to define a visual relationship of “boy” “holding” “ball.”); and
encoding the relationship in the scene graph (Fig. 2A. Paragraph [0051]-YAN discloses a scene graph 202 includes a set of nodes 204 and a set of edges 206 that interconnect a subset of nodes in the set of nodes. Each scene graph 202 can define a set of objects that are represented by respective nodes 204, e.g., where a first object is represented by a first node from the set of nodes, and a second object is represented by a second node from the set of nodes. The first node and the second node can be connected by an edge representing a relationship feature that is defining of a relationship between the two objects.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of BURLINA in view of YAN of a method for data collection, comprising: detecting a set of objects in an environment based on an obtained set of multimodal data from a plurality of sensors; generating a scene graph based on the set of objects; with the teachings of YAN of determining a relationship between a first object in the set of objects and a second object in the set of objects; and encoding the relationship in the scene graph.
Wherein having BURLINA’s apparatus for receiving and processing data from a vehicle having determining a relationship between a first object in the set of objects and a second object in the set of objects; and encoding the relationship in the scene graph.
The motivation behind the modification would have been to obtain an apparatus/method for receiving and processing data from a vehicle that enhances efficiently and accurately detects objects surrounding a vehicle and planning navigation for the vehicle. Since both BURLINA and YAN relate to detecting objects and generating scene graphs of the detected objects, wherein BURLINA using data from multiple spectral bands may allow for more robust detection of objects, resulting in improved collision avoidance with animals, debris, and/or other relatively smaller objects on driving surfaces, while YAN an advantage of this technology is that it can facilitate efficient and accurate discovery of videos and key frames within videos using natural language descriptions of visual relationships between objects depicted in the key frames, and may reduce a number of queries required to be entered by a user in order to find a particular video of interest. This in turn reduces the number of computer resources required to execute multiple queries until the appropriate video has been identified. Please see BURLINA et al. (US 12525025 B1), Col. 2, Line [3-26], and YAN (US 20230086735 A1), Paragraph [0014].
Regarding claim 19, BURLINA in view of YAN explicitly teach the method of claim 18,
BURLINA fails to explicitly teach wherein the relationship comprises at least one of a distance between the first object and the second object, or an intent of the first object with respect to the second object.
However, YAN explicitly teaches wherein the relationship comprises at least one of a distance between the first object and the second object (Fig. 2A. Paragraph [0061]-YAN discloses relationships can be determined by the visual relationship model 108, for example, based in part on proximity/spatial distances between objects, known relationships between categories of objects, user-defined relationships between particular objects and/or categories of objects, or the like.), or an intent of the first object with respect to the second object (Fig. 2A. Paragraph [0061]-YAN discloses a relationship feature 212 can be “holding,” where the relationship feature 212 defines a relationship between a first object “boy” and a second object “ball,” to define a visual relationship of “boy” “holding” “ball” (wherein the intent is holding).).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of BURLINA in view of YAN of a method for data collection, comprising: detecting a set of objects in an environment based on an obtained set of multimodal data from a plurality of sensors; generating a scene graph based on the set of objects; with the teachings of YAN of wherein the relationship comprises at least one of a distance between the first object and the second object, or an intent of the first object with respect to the second object.
Wherein having BURLINA’s apparatus for receiving and processing data from a vehicle having wherein the relationship comprises at least one of a distance between the first object and the second object, or an intent of the first object with respect to the second object.
The motivation behind the modification would have been to obtain an apparatus/method for receiving and processing data from a vehicle that enhances efficiently and accurately detects objects surrounding a vehicle and planning navigation for the vehicle. Since both BURLINA and YAN relate to detecting objects and generating scene graphs of the detected objects, wherein BURLINA using data from multiple spectral bands may allow for more robust detection of objects, resulting in improved collision avoidance with animals, debris, and/or other relatively smaller objects on driving surfaces, while YAN an advantage of this technology is that it can facilitate efficient and accurate discovery of videos and key frames within videos using natural language descriptions of visual relationships between objects depicted in the key frames, and may reduce a number of queries required to be entered by a user in order to find a particular video of interest. This in turn reduces the number of computer resources required to execute multiple queries until the appropriate video has been identified. Please see BURLINA et al. (US 12525025 B1), Col. 2, Line [3-26], and YAN (US 20230086735 A1), Paragraph [0014].
Regarding claim 20, BURLINA in view of YAN explicitly teach the method of claim 18, further comprising:
BURLINA fails to explicitly teach encoding the first object as a first node in the scene graph; encoding the second object as a second node in the scene graph; and encoding the relationship as an edge between the first node and the second node.
However, YAN explicitly teaches encoding the first object as a first node in the scene graph (Fig. 2A. Paragraph [0051]-YAN discloses each scene graph 202 can define a set of objects that are represented by respective nodes 204, e.g., where a first object is represented by a first node from the set of nodes, and a second object is represented by a second node from the set of nodes. The first node and the second node can be connected by an edge representing a relationship feature that is defining of a relationship between the two objects.);
encoding the second object as a second node in the scene graph (Fig. 2A. Paragraph [0051]-YAN discloses each scene graph 202 can define a set of objects that are represented by respective nodes 204, e.g., where a first object is represented by a first node from the set of nodes, and a second object is represented by a second node from the set of nodes. The first node and the second node can be connected by an edge representing a relationship feature that is defining of a relationship between the two objects.); and
encoding the relationship as an edge between the first node and the second node (Fig. 2A. Paragraph [0051]-YAN discloses each scene graph 202 can define a set of objects that are represented by respective nodes 204, e.g., where a first object is represented by a first node from the set of nodes, and a second object is represented by a second node from the set of nodes. The first node and the second node can be connected by an edge representing a relationship feature that is defining of a relationship between the two objects.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of BURLINA in view of YAN of a method for data collection, comprising: detecting a set of objects in an environment based on an obtained set of multimodal data from a plurality of sensors; generating a scene graph based on the set of objects; with the teachings of YAN of encoding the first object as a first node in the scene graph; encoding the second object as a second node in the scene graph; and encoding the relationship as an edge between the first node and the second node.
Wherein having BURLINA’s apparatus for receiving and processing data from a vehicle having encoding the first object as a first node in the scene graph; encoding the second object as a second node in the scene graph; and encoding the relationship as an edge between the first node and the second node.
The motivation behind the modification would have been to obtain an apparatus/method for receiving and processing data from a vehicle that enhances efficiently and accurately detects objects surrounding a vehicle and planning navigation for the vehicle. Since both BURLINA and YAN relate to detecting objects and generating scene graphs of the detected objects, wherein BURLINA using data from multiple spectral bands may allow for more robust detection of objects, resulting in improved collision avoidance with animals, debris, and/or other relatively smaller objects on driving surfaces, while YAN an advantage of this technology is that it can facilitate efficient and accurate discovery of videos and key frames within videos using natural language descriptions of visual relationships between objects depicted in the key frames, and may reduce a number of queries required to be entered by a user in order to find a particular video of interest. This in turn reduces the number of computer resources required to execute multiple queries until the appropriate video has been identified. Please see BURLINA et al. (US 12525025 B1), Col. 2, Line [3-26], and YAN (US 20230086735 A1), Paragraph [0014].
Claims 3-4 and 16-17 are rejected under 35 U.S.C. 103 as being unpatentable over BURLINA et al. (US 12525025 B1), hereinafter referenced as BURLINA, in view of YAN (US 20230086735 A1), hereinafter referenced as YAN, and further in view of TAY et al. (US 20190311546 A1), hereinafter referenced as TAY.
Regarding claim 3, BURLINA in view of YAN explicitly teach the apparatus of claim 2,
BURLINA further explicitly teaches wherein the at least one processor is configured to (Fig. 5, #520 called processor and #522 called memory. Col. 24, Lines [56-62]-BURLINA discloses the vehicle computing device(s) 504 may include processor(s) 520 and memory 522 communicatively coupled with the processor(s) 520. The computing device(s) 518 may also include processor(s) 524, and/or memory 526. The processor(s) 520 and/or 524 may be any suitable processor capable of executing instructions to process data and perform operations as described herein.):
determine a driving context of the apparatus (Fig. 5. Col. 26, Line [43-49]-BURLINA discloses the planning component 532 can determine a route to travel from a first location (e.g., a current location) to a second location (e.g., a target location). For the purpose of this discussion, a route can be a sequence of waypoints for traveling between two locations. As non-limiting examples, waypoints include streets, intersections, global positioning system (GPS) coordinates, etc. Further in Col. 27, Line [8-11]-BURLINA discloses the planning component 532 can determine a route to travel from a first location (e.g., a current location) to a second location (e.g., a target location) to avoid objects in an environment.); and
BURLINA in view of YAN fail to explicitly teach receive a second description of a second scenario of interest; determine to output the second description of the second scenario of interest instead of the first description of the first scenario of interest based on the driving context of the apparatus.
However, TAY explicitly teaches receive a second description of a second scenario of interest (Fig. 4. Paragraph [0116]-TAY discloses if the autonomous vehicle determines that the billboard is an advertisement for a local business (e.g., a coffee shop, a retail store) based on iconography detected on the billboard, the autonomous vehicle can render a prompt—to reroute the autonomous vehicle to this local business—over the billboard depicted in the augmented 3D on the interior display; accordingly, the autonomous vehicle can update its navigation path and reroute to a known location of this local business responsive to the user selecting this prompt on the interior display (wherein the second description of a second scenario of interest is rerouting to the local business).);
determine to output the second description of the second scenario of interest instead of the first description of the first scenario of interest based on the driving context of the apparatus (Fig. 4. Paragraph [0116]-TAY discloses if the autonomous vehicle determines that the billboard is an advertisement for a local business (e.g., a coffee shop, a retail store) based on iconography detected on the billboard, the autonomous vehicle can render a prompt—to reroute the autonomous vehicle to this local business—over the billboard depicted in the augmented 3D on the interior display; accordingly, the autonomous vehicle can update its navigation path and reroute to a known location of this local business responsive to the user selecting this prompt on the interior display (wherein rerouting is outputting the second description, the initial navigation path is the first description, and the billboards and user selecting prompt are driving contexts).).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of BURLINA in view of YAN of an apparatus for data collection, comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: detect a set of objects in an environment based on an obtained set of multimodal data from a plurality of sensors; generate a scene graph based on the set of objects; with the teachings of TAY of receive a first description of a first scenario of interest, wherein the first description comprises a textual description of the first scenario of interest; and output the textual description of the first scenario of interest.
Wherein having BURLINA’s apparatus for receiving and processing data from a vehicle having receive a first description of a first scenario of interest, wherein the first description comprises a textual description of the first scenario of interest; and output the textual description of the first scenario of interest.
The motivation behind the modification would have been to obtain an apparatus/method for receiving and processing data from a vehicle that enhances efficiently and accurately detects objects surrounding a vehicle and planning navigation for the vehicle. Since both BURLINA and TAY relate to detecting objects in an autonomous vehicle, wherein BURLINA using data from multiple spectral bands may allow for more robust detection of objects, resulting in improved collision avoidance with animals, debris, and/or other relatively smaller objects on driving surfaces, while TAY enabling the rider to build confidence in the autonomous vehicle's perception of its environment while limiting computational load necessary to generate and render this representation of the field.. Please see BURLINA et al. (US 12525025 B1), Col. 2, Line [3-26], and TAY et al. (US 20190311546 A1), Paragraph [0114].
Regarding claim 4, BURLINA in view of YAN and further in view of TAY explicitly teach the apparatus of claim 3,
BURLINA further explicitly teaches wherein the driving context is based on a location of the apparatus (Fig. 5. Col. 26, Line [43-49]-BURLINA discloses the planning component 532 can determine a route to travel from a first location (e.g., a current location) to a second location (e.g., a target location). For the purpose of this discussion, a route can be a sequence of waypoints for traveling between two locations. As non-limiting examples, waypoints include streets, intersections, global positioning system (GPS) coordinates, etc. Further in Col. 27, Line [8-11]-BURLINA discloses the planning component 532 can determine a route to travel from a first location (e.g., a current location) to a second location (e.g., a target location) to avoid objects in an environment.).
Regarding claim 16, BURLINA in view of YAN explicitly teach the method of claim 15, BURLINA further explicitly teaches further comprising:
determining a driving context of a vehicle (Fig. 5. Col. 26, Line [43-49]-BURLINA discloses the planning component 532 can determine a route to travel from a first location (e.g., a current location) to a second location (e.g., a target location). For the purpose of this discussion, a route can be a sequence of waypoints for traveling between two locations. As non-limiting examples, waypoints include streets, intersections, global positioning system (GPS) coordinates, etc. Further in Col. 27, Line [8-11]-BURLINA discloses the planning component 532 can determine a route to travel from a first location (e.g., a current location) to a second location (e.g., a target location) to avoid objects in an environment.); and
BURLINA in view of YAN fail to explicitly teach receiving a second description of a second scenario of interest; determining to output the second description of the second scenario of interest instead of the first description of the first scenario of interest based on the driving context of the vehicle.
However, TAY explicitly teaches receiving a second description of a second scenario of interest (Fig. 4. Paragraph [0116]-TAY discloses if the autonomous vehicle determines that the billboard is an advertisement for a local business (e.g., a coffee shop, a retail store) based on iconography detected on the billboard, the autonomous vehicle can render a prompt—to reroute the autonomous vehicle to this local business—over the billboard depicted in the augmented 3D on the interior display; accordingly, the autonomous vehicle can update its navigation path and reroute to a known location of this local business responsive to the user selecting this prompt on the interior display (wherein the second description of a second scenario of interest is rerouting to the local business).);
determining to output the second description of the second scenario of interest instead of the first description of the first scenario of interest based on the driving context of the vehicle (Fig. 4. Paragraph [0116]-TAY discloses if the autonomous vehicle determines that the billboard is an advertisement for a local business (e.g., a coffee shop, a retail store) based on iconography detected on the billboard, the autonomous vehicle can render a prompt—to reroute the autonomous vehicle to this local business—over the billboard depicted in the augmented 3D on the interior display; accordingly, the autonomous vehicle can update its navigation path and reroute to a known location of this local business responsive to the user selecting this prompt on the interior display (wherein rerouting is outputting the second description, the initial navigation path is the first description, and the billboards and user selecting prompt are driving contexts).).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of BURLINA in view of YAN of a method for data collection, comprising: detecting a set of objects in an environment based on an obtained set of multimodal data from a plurality of sensors; generating a scene graph based on the set of objects; with the teachings of TAY of receiving a second description of a second scenario of interest; determining to output the second description of the second scenario of interest instead of the first description of the first scenario of interest based on the driving context of the vehicle.
Wherein having BURLINA’s apparatus for receiving and processing data from a vehicle having receiving a second description of a second scenario of interest; determining to output the second description of the second scenario of interest instead of the first description of the first scenario of interest based on the driving context of the vehicle.
The motivation behind the modification would have been to obtain an apparatus/method for receiving and processing data from a vehicle that enhances efficiently and accurately detects objects surrounding a vehicle and planning navigation for the vehicle. Since both BURLINA and TAY relate to detecting objects in an autonomous vehicle, wherein BURLINA using data from multiple spectral bands may allow for more robust detection of objects, resulting in improved collision avoidance with animals, debris, and/or other relatively smaller objects on driving surfaces, while TAY enabling the rider to build confidence in the autonomous vehicle's perception of its environment while limiting computational load necessary to generate and render this representation of the field.. Please see BURLINA et al. (US 12525025 B1), Col. 2, Line [3-26], and TAY et al. (US 20190311546 A1), Paragraph [0114].
Regarding claim 17, BURLINA in view of YAN and further in view of TAY explicitly teach the method of claim 16,
BURLINA further explicitly teaches wherein the driving context is based on a location of the vehicle (Fig. 5. Col. 26, Line [43-49]-BURLINA discloses the planning component 532 can determine a route to travel from a first location (e.g., a current location) to a second location (e.g., a target location). For the purpose of this discussion, a route can be a sequence of waypoints for traveling between two locations. As non-limiting examples, waypoints include streets, intersections, global positioning system (GPS) coordinates, etc. Further in Col. 27, Line [8-11]-BURLINA discloses the planning component 532 can determine a route to travel from a first location (e.g., a current location) to a second location (e.g., a target location) to avoid objects in an environment.).
Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over BURLINA et al. (US 12525025 B1), hereinafter referenced as BURLINA, in view of YAN (US 20230086735 A1), hereinafter referenced as YAN, and further in view of HOVIS et al. (US 20180341821 A1), hereinafter referenced as HOVIS.
Regarding claim 9, BURLINA in view of YAN explicitly teach the apparatus of claim 1,
BURLINA further explicitly teaches wherein the at least one processor is configured to (Fig. 5, #520 called processor and #522 called memory. Col. 24, Lines [56-62]-BURLINA discloses the vehicle computing device(s) 504 may include processor(s) 520 and memory 522 communicatively coupled with the processor(s) 520. The computing device(s) 518 may also include processor(s) 524, and/or memory 526. The processor(s) 520 and/or 524 may be any suitable processor capable of executing instructions to process data and perform operations as described herein.):
BURLINA in view of YAN fail to explicitly teach detect a second set of objects in the environment based on an obtained second set of multimodal data; and update the scene graph based on the second set of objects.
However, HOVIS explicitly teaches detect a second set of objects in the environment (Fig. 5, #506 called compare location and classification of newly detected objects with existing objects on PSG. Paragraph [0064]-HOVIS discloses the newly detected objects are compared with existing objects in the PSG 112 that were previously detected (wherein the newly detected objects are a second set of objects).) based on an obtained second set of multimodal data (Fig. 5. Paragraph [0064]-HOVIS discloses in this step 506, information gathered by the external sensors 206 and information received by the V2X receivers 208 may be fused to increase the confidence factors of the objects detected together with the range and direction of the objects relative to the motor vehicle 300. The newly detected objects are compared with existing objects in the PSG 112 that were previously detected. Further in paragraph [0049]-HOVIS discloses the external sensors 206 include, but are not limited to, radar, laser, scanning laser, camera, sonar, ultra-sonic devices, LIDAR, and the like.) ; and
update the scene graph based on the second set of objects (Fig. 5. [0067]-HOVIS discloses in step 512, the PSG 112 is generated, published, and becomes accessible by various vehicle systems that require information about the surroundings of the motor vehicle. The PSG 112 contains information on a set of localized objects, categories of each object, and relationship between each object and the motor vehicle 300. The PSG 112 is continuously updated and historical events of the PSG 112 may be recorded.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of BURLINA in view of YAN of an apparatus for data collection, comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: detect a set of objects in an environment based on an obtained set of multimodal data from a plurality of sensors; generate a scene graph based on the set of objects; with the teachings of HOVIS of detect a second set of objects in the environment based on an obtained second set of multimodal data; and update the scene graph based on the second set of objects.
Wherein having BURLINA’s apparatus for receiving and processing data from a vehicle having detect a second set of objects in the environment based on an obtained second set of multimodal data; and update the scene graph based on the second set of objects.
The motivation behind the modification would have been to obtain an apparatus/method for receiving and processing data from a vehicle that enhances efficiently and accurately detects objects surrounding a vehicle and planning navigation for the vehicle. Since both BURLINA and HOVIS relate to detecting objects and generating scene graphs of the detected objects, wherein BURLINA using data from multiple spectral bands may allow for more robust detection of objects, resulting in improved collision avoidance with animals, debris, and/or other relatively smaller objects on driving surfaces, while HOVIS while current ADAS having vehicle controllers adequate to process information from predetermined specific types of exterior sensors to achieve their intended purpose, there is a need for a new and improved system and method for a perception system to accommodate new sensor types without the need of upgrading the processors and/or routines of the vehicle controllers. Please see BURLINA et al. (US 12525025 B1), Col. 2, Line [3-26], and HOVIS et al. (US 20180341821 A1), Paragraph [0006].
Claims 10-11 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over YAN (US 20230086735 A1), hereinafter referenced as YAN, in view of GRAF et al. (US 20180307967 A1), hereinafter referenced as GRAF.
Regarding claim 10, YAN explicitly teaches an apparatus for data collection (Fig. 1. Paragraph [0033]-YAN discloses FIG. 1 depicts an example operating environment 100 of a visual relationship system 102. Visual relationship system 102 can be hosted on a local device, e.g., user device 104, one or more local servers, a cloud-based service, or a combination thereof (wherein a user device is an apparatus for data collection).), comprising:
at least one memory (Fig. 7, #704 called secondary storage, #706 called ROM, and #08 called RAM. Paragraph [0110]-YAN discloses memory devices including secondary storage 704, and memory, such as ROM 706 and RAM 708); and
at least one processor (Fig. 7, #702 called processor. Paragraph [0109]) coupled to the at least one memory and configured to (Fig. 7. Paragraph [0109-0110]-YAN discloses the general-purpose network component or computer system includes a processor 702 (which may be referred to as a central processor unit or CPU) that is in communication with memory devices including secondary storage 704, and memory, such as ROM 706 and RAM 708, input/output (I/O) devices 710, and a network.):
receive a description of a scenario of interest (Fig. 3, #302 called parsed query. Paragraph [0070]-YAN discloses a query 302 including terms descriptive of a visual relationship can be provided to the visual relationship system 102. In some implementations, query 302 is a textual query that is generated by the speech-to-text converter 106 from a query 120 received by the visual relationship system 102 from a user on a user device 104. Further in paragraph [0071]-YAN discloses a query 302 is “I want a boy holding a ball” where the object-terms are determined as “boy” and “ball” and relationship feature-terms are determined as “holding.”),
wherein the description comprises a textual description of the scenario of interest (Fig. 3, #302 called parsed query. Paragraph [0070]-YAN discloses a query 302 including terms descriptive of a visual relationship can be provided to the visual relationship system 102. In some implementations, query 302 is a textual query that is generated by the speech-to-text converter 106 from a query 120 received by the visual relationship system 102 from a user on a user device 104. Further in paragraph [0071]-YAN discloses a query 302 is “I want a boy holding a ball” where the object-terms are determined as “boy” and “ball” and relationship feature-terms are determined as “holding.”);
parse the description of the scenario of interest (Fig. 3, #302 called parsed query. Paragraph [0071]-YAN discloses visual relationship system 102 can receive the query 302 as input and perform feature/object extraction 304 on the query 302 to determine terms of the query 302 defining objects 306 and relationship features 308. Visual relationship system 102 can extract objects 306 and relationship features 308 from the input query 302, for example, by using natural language processing to parse the terms of the query and identify objects/relationship features.) to generate a query scene graph based on the description of the scenario of interest (Fig. 3, #312 called query graph. Paragraph [0072]-YAN discloses the visual relationship system 102 can utilize the extracted objects 306 and relationship features 308 that are defined in the terms of the query 302 to perform query graph generation 310. A query graph 312 can be generated where objects 306 and relationship features 308 extracted from the terms of the query 302 are utilized as nodes 314 and edges 316 between nodes, respectively. Continuing the example provided above, a query graph 312 can include a first node “boy” and a second node “ball” with an edge “holding” connecting the first and second nodes 314.);
YAN fails to explicitly teach output the description of the scenario of interest for transmission to a vehicle; and output the query scene graph for transmission to the vehicle.
However, GRAF explicitly teaches output the description of the scenario of interest for transmission to a vehicle (Fig. 6. Paragraph [0036]-GRAF discloses the user interface control panel 150 can display a plurality of different data and/or information to the driver of the vehicle 102. For example, a speed 180 of vehicle A can be displayed relative to a speed 182 of vehicle B and a speed 184 of the vehicle C. Of course, one skilled in the art can contemplate displaying a plurality of other information to the user (e.g., position information related to each vehicle A, B, C, D, etc.) (wherein the description is the displayed speed and position information of vehicles).); and
output the query scene graph (Fig. 1-3. Paragraph [0023]-GRAF discloses the set of detections that form observations of the driving or traffic scene are treated as a graph. The graph can be defined as a set of nodes (e.g., vehicles) and edges (e.g., relationship between the vehicles). One or more nodes can also represent persons. Further in paragraph [0024-0025]-GRAF discloses relationships are derived between vehicles A and B, between vehicles A and C, between vehicles A, C, B, and between vehicles A, B, C. Thus, the paths 16 (FIG. 1) enable the extraction of data and/or information between nodes of a graph. Referring to FIG. 3, a neural network 30 is constructed for each extracted substructure (e.g., path). The neural network 30 can be constructed by a construction module. These networks are constructed by using the same neural network building blocks at each step through the graph.) for transmission to the vehicle (Fig. 6. Paragraph [0038]-GRAF discloses the vehicle 102 can further receive data and/or information from a plurality of networks. For example, the vehicle 102 can receive data from a first network 130 (e.g., Internet) and a second network 140 (e.g., a deep convolutional neural network). One skilled in the art can contemplate a plurality of other networks for communicating with the vehicle 102 (wherein the neural network facilitates transmission to the vehicle as the graph is input into the neural network). Further see paragraph [0023-0026].).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of YAN of an apparatus for data collection, comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: receive a description of a scenario of interest, wherein the description comprises a textual description of the scenario of interest; parse the description of the scenario of interest to generate a query scene graph based on the description of the scenario of interest; with the teachings of GRAF of output the description of the scenario of interest for transmission to a vehicle; and output the query scene graph for transmission to the vehicle.
Wherein having YAN’s apparatus for receiving and processing image data having output the description of the scenario of interest for transmission to a vehicle; and output the query scene graph for transmission to the vehicle.
The motivation behind the modification would have been to obtain an apparatus/method for receiving and processing image data that enhances efficiently and accurately detects objects surrounding a vehicle and planning navigation for the vehicle. Since both YAN and GRAF relate to detecting objects and generating scene graphs of the detected objects, wherein YAN an advantage of this technology is that it can facilitate efficient and accurate discovery of videos and key frames within videos using natural language descriptions of visual relationships between objects depicted in the key frames, and may reduce a number of queries required to be entered by a user in order to find a particular video of interest. This in turn reduces the number of computer resources required to execute multiple queries until the appropriate video has been identified, while GRAF to improve accuracy of the positioning by using lane detection or structure from motion. Please see YAN (US 20230086735 A1), Paragraph [0014], and GRAF et al (US 20180307967 A1), Paragraph [0021].
Regarding claim 11, YAN in view of GRAF explicitly teach the apparatus of claim 10,
YAN further explicitly teaches wherein the at least one processor is further configured to (Fig. 7. Paragraph [0109-0110]-YAN discloses the general-purpose network component or computer system includes a processor 702 (which may be referred to as a central processor unit or CPU) that is in communication with memory devices including secondary storage 704, and memory, such as ROM 706 and RAM 708, input/output (I/O) devices 710, and a network.):
store the scene graph in a dataset (Fig. 1, #118 called scene graph database. Paragraph [0066]-YAN discloses the scene graph 202 for the key frame 115 is stored in scene graph database 118.).
Although YAN explicitly teaches receive a scene graph matching the query scene graph. YAN fails to explicitly teach receive a scene graph matching the query scene graph from the vehicle.
However, GRAF explicitly teaches receive a scene graph matching the query scene graph from the vehicle (Fig. 1, illustrates a query scene graph. Paragraph [0020]-GRAF discloses a driving danger prediction system is realized by continuously matching a current TS to a codebook of (or predetermined or predefined or pre-established) TSs that have been identified as leading to dangerous situations. When a match occurs, a warning can be transmitted to the driver. The exemplary invention describes how to fit an end-to-end scene parsing neural network learning approach (e.g., a graph network) to the challenge of matching TSs (wherein a traffic scene (TS) is a scene graph and wherein a warning is transmitted when the system receives a scene graph matching the query scene graph).); and
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of YAN of an apparatus for data collection, comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: receive a description of a scenario of interest, wherein the description comprises a textual description of the scenario of interest; parse the description of the scenario of interest to generate a query scene graph based on the description of the scenario of interest; with the teachings of GRAF of receive a scene graph matching the query scene graph from the vehicle.
Wherein having YAN’s apparatus for receiving and processing image data having receive a scene graph matching the query scene graph from the vehicle.
The motivation behind the modification would have been to obtain an apparatus/method for receiving and processing image data that enhances efficiently and accurately detects objects surrounding a vehicle and planning navigation for the vehicle. Since both YAN and GRAF relate to detecting objects and generating scene graphs of the detected objects, wherein YAN an advantage of this technology is that it can facilitate efficient and accurate discovery of videos and key frames within videos using natural language descriptions of visual relationships between objects depicted in the key frames, and may reduce a number of queries required to be entered by a user in order to find a particular video of interest. This in turn reduces the number of computer resources required to execute multiple queries until the appropriate video has been identified, while GRAF to improve accuracy of the positioning by using lane detection or structure from motion. Please see YAN (US 20230086735 A1), Paragraph [0014], and GRAF et al (US 20180307967 A1), Paragraph [0021].
Regarding claim 13, YAN in view of GRAF explicitly teach the apparatus of claim 10,
YAN further explicitly teaches wherein the description of the scenario includes a first object, a second object, and a relationship between the first object and second object (Fig. 3. Paragraph [0071]-YAN discloses a query 302 is “I want a boy holding a ball” where the object-terms are determined as “boy” and “ball” and relationship feature-terms are determined as “holding” (wherein boy is a first object, ball is a second object, and holding is a relationship between the objects).), and
wherein the at least one processor is further configured to (Fig. 7. Paragraph [0109-0110]-YAN discloses the general-purpose network component or computer system includes a processor 702 (which may be referred to as a central processor unit or CPU) that is in communication with memory devices including secondary storage 704, and memory, such as ROM 706 and RAM 708, input/output (I/O) devices 710, and a network.):
encode the first object as a first node in the query scene graph (Fig. 3, #312 called query graph. Paragraph [0072]-YAN discloses a query graph 312 can include a first node “boy” and a second node “ball” with an edge “holding” connecting the first and second nodes 314.);
encode the second object as a second node in the query scene graph (Fig. 3, #312 called query graph. Paragraph [0072]-YAN discloses a query graph 312 can include a first node “boy” and a second node “ball” with an edge “holding” connecting the first and second nodes 314.); and
encode the relationship as an edge between the first node and the second node (Fig. 3, #312 called query graph. Paragraph [0072]-YAN discloses a query graph 312 can include a first node “boy” and a second node “ball” with an edge “holding” connecting the first and second nodes 314 (wherein holding is the relationship).).
Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over YAN (US 20230086735 A1), hereinafter referenced as YAN, in view of GRAF et al. (US 20180307967 A1), hereinafter referenced as GRAF, and further in view of KABZAN et al (US 11550851 B1), hereinafter referenced as KABZAN.
Regarding claim 12, YAN in view of GRAF explicitly teach the apparatus of claim 11,
YAN in view of GRAF fail to explicitly teach wherein the description of a scenario of interest is generated based on scenarios which are underrepresented in the dataset.
However, KABZAN explicitly teaches wherein the description of a scenario of interest is generated based on scenarios which are underrepresented in the dataset (Fig. 5. Col. 18, Line [7-13]-KABZAN discloses the scenario database 530 may be updated by the vehicle log data store 535 with new scenarios. The scenario mining controller 590 may be configured to add metadata to vehicle scenarios from the vehicle log data store 535 to update the scenario database 530 based on new environments the vehicle has been in or simulations in which the machine learning model 570 has been trained (wherein a new scenario is an underrepresented scenario). Further in Col. 17, Line [39-49]-KABZAN discloses the scenario mining controller 590 may search through the scenario database 530 to identify a vehicle scenario having metadata that includes a large truck as an agent vehicle, a pedestrian detected in the crosswalk, a stroller detected in the crosswalk, and the ego vehicle located 30 feet away from the crosswalk (wherein generating a description of a scenario is adding metadata to the scenario in the database).).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of YAN in view of GRAF of an apparatus for data collection, comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: receive a description of a scenario of interest, wherein the description comprises a textual description of the scenario of interest; parse the description of the scenario of interest to generate a query scene graph based on the description of the scenario of interest; with the teachings of KABZAN of wherein the description of a scenario of interest is generated based on scenarios which are underrepresented in the dataset.
Wherein having YAN’s apparatus for receiving and processing image data having wherein the description of a scenario of interest is generated based on scenarios which are underrepresented in the dataset.
The motivation behind the modification would have been to obtain an apparatus/method for receiving and processing image data that enhances efficiently and accurately detects objects surrounding a vehicle and planning navigation for the vehicle. Since both YAN and KABZAN relate to detecting objects and storing data on the detected objects, wherein YAN an advantage of this technology is that it can facilitate efficient and accurate discovery of videos and key frames within videos using natural language descriptions of visual relationships between objects depicted in the key frames, and may reduce a number of queries required to be entered by a user in order to find a particular video of interest. This in turn reduces the number of computer resources required to execute multiple queries until the appropriate video has been identified, while KABZAN as a technical improvement, the scenario database described herein mines complex scenarios using SQL queries using attributes of these complex scenarios.. Please see YAN (US 20230086735 A1), Paragraph [0014], and KABZAN et al (US 11550851 B1), Col. 4, Line [14-35].
Conclusion
Listed below are the prior arts made of record and not relied upon but are considered pertinent to applicant’s disclosure.
TANG et al. (US 20240086586 A1) – A computer-implemented method for simulating vehicle data and improving driving scenario detection is provided. The method includes retrieving, from vehicle sensors, key parameters from real data of validation scenarios to generate corresponding scenario configurations and descriptions, transferring target scenario descriptions and validation scenario descriptions to target scenario scripts and validation scenario scripts, respectively, to create first raw simulation data pertaining to target scenario descriptions and second raw simulation data pertaining to validation scenario descriptions, training, by an adjuster network, a deep neural network model to minimize differences between the first raw simulation data and the second raw simulation data, refining the first and second raw simulation data of rare driving scenarios to generate rare driving scenario training data, and outputting the rare driving scenario training data to a display screen of a computing device to enable a user to train a scenario detector for an autonomic driving assistant system…Abstract, Fig. 2.
CHEN et al. (US 20230047160 A1) – Systems, methods, computer-readable media, techniques, and methodologies are disclosed for performing end-to-end, learning-based keypoint detection and association. A scene graph of a signalized intersection is constructed from an input image of the intersection. The scene graph includes detected keypoints and linkages identified between the keypoints. The scene graph can be used along with a vehicle's localization information to identify which keypoint that represents a traffic signal is associated with the vehicle's current travel lane. An appropriate vehicle action may then be determined based on a transition state of the traffic signal keypoint and trajectory information for the vehicle. A control signal indicative of this vehicle action may then be output to cause an autonomous vehicle, for example, to implement the appropriate vehicle action…Abstract, Fig. 4.
PANDYA et al. (US 20220172606 A1) – An example method for extracting traffic scenarios from vehicle sensor data is disclosed. The example method includes acquiring vehicle data generated by one or more sensors coupled to a vehicle. The vehicle data is at least partially indicative of the surroundings of the vehicle during a particular time frame. The vehicle data is analyzed to identify objects in the surroundings of the vehicle and to determine the motion of the vehicle relative to the surroundings during the particular time frame. A plurality of events are defined, each indicative of a relationship between the vehicle and the objects. A scenario is defined as a particular combination of the events. Portions of the vehicle data in which the combination of elements occurs during a time interval are identified, and at least some of the identified data is extracted to a predefined data structure to create an extracted scenario…Abstract, Fig. 16A-16D.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ETHAN N WOLFSON whose telephone number is (571)272-1898. The examiner can normally be reached Monday - Friday 8:00 am - 5:00 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chineyere Wills-Burns can be reached at (571) 272-9752. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ETHAN N WOLFSON/Examiner, Art Unit 2673
/CHINEYERE WILLS-BURNS/Supervisory Patent Examiner, Art Unit 2673