DETAILED ACTION
Claims 1-20 are pending.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statements (IDS) submitted on October 3, 2024 and April 23, 2026 have been considered by the examiner.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1, 7, 8, 9, 10, 15 are rejected under 35 U.S.C. 103 as being unpatentable over Wang et al. (NPL titled "DriveMLM: Aligning multi-modal large language models with behavioral planning states for autonomous driving”) (“Wang”) in view of Wouhaybi et al. (U.S. Patent No. US 12348860 B2) (“Wouhaybi”).
Regarding claim 1, Wang discloses provide a prompt to a multi-modal language model (Abstract; Fig.3 and Section 3.1 paragraph 1; Section 3.3 paragraph 1; Tables B and C in “Supplementary Material” include examples of prompts) to generate, based at least on sensor data generated using a plurality of sensors of an ego-machine (Abstract; Section 3.1 paragraph 1; pg. 5 Fig. 3, wherein multi-modal language model (DriveMLM) receives sensor inputs (i.e. sensor data generated) from the vehicle (i.e. ego-machine) including multi-view images and LiDAR (i.e. from a plurality of sensors)), one or more initial responses (Table 1; Section 3.2(3), wherein an initial response relevant to a designated task from the given prompt is generated);
However, as indicated by the double strikethroughs above, Wang fails to teach one or more initial responses representative of a selected sensor of the plurality of sensors that is relevant to a designated task, specifically (emphasis added). Wouhaybi, on the other hand, teaches an ego-machine with multiple sensors using a neural network to select a camera based on relevancy to a designated task. More specifically and as it relates to the applicant’s claims, Wouhaybi discloses one or more initial responses representative of a selected sensor of the plurality of sensors that is relevant to a designated task (Column 8 lines 8-15 and 26-42, wherein sensor data is received from various external sensors of a robot and a neural network is configured to select (i.e. initial response) one or more sensors for data delivery based on usefulness/relevance of sensor data or task requirements (i.e. representative of a selected sensor of the plurality of sensors that is relevant to a designated task).
Wouhaybi is combinable with Wang because they are from the same art of image processing.
The suggestion/motivation for doing so would have been to increase power and computational efficiency (Wouhaybi, Column 5 lines 9-10).
Wang additionally discloses provide a prompt to the multi-modal language model to generate one or more subsequent responses evaluating the designated task based at least on one or more frames of the sensor data generated (Table 1; Section 3.2(3), wherein a subsequent response applied to the MLLM evaluates the designated task from the subsequent prompt; Section 4.1 paragraph 2, wherein images from the multi-view cameras are collected for each frame (i.e. based at least on one or more frames of the sensor data generated)); and
control one or more operations of the ego-machine based at least on the one or more subsequent responses (Section 1 paragraph 3, wherein DriveMLM performs autonomous driving of the vehicle (i.e. ego-machine) by converting decisions (i.e. subsequent responses) into vehicle control signals; Section 3.1 paragraph 2, wherein decision state output is input into a motion planning and control module, computing the final trajectory for vehicle control).
However, as indicated by the double strikethroughs above, Wang fails to teach generate one or more subsequent responses evaluating the designated task based at least on one or more frames of the sensor data generated using the selected sensor, specifically (emphasis added). Wouhaybi, on the other hand, teaches calculation of a confidence factor based on the selected sensor information. If the confidence factor is within a predetermined range, the robot continues to receive sensor data from the selected sensor. If the confidence factor is outside of the predetermined range, then a different sensor will be triggered. More specifically and as it relates to the applicant’s claims, Wouhaybi discloses generate one or more subsequent responses evaluating the designated task based at least on one or more frames of the sensor data generated using the selected sensor (Column 9 lines 33-53, wherein a confidence factor based on sensor information is calculated (i.e. evaluated designated task based on sensor data) and if the confidence factor is within a predetermined range, sensor data from the selected sensor will continually be received, or if the confidence factor is outside of the predetermined range, then a different sensor will be selected (i.e. subsequent response using the selected sensor)).
Wouhaybi is combinable with Wang because they are from the same art of image processing.
The suggestion/motivation for doing so would have been to only allow sufficient/relevant sensor data to be received for increased power and computational efficiency (Wouhaybi, Column 5 lines 9-10; column 9 lines 45-48).
Wang additionally fails to teach one or more processors comprising processing circuitry to perform the instructions above. Wouhaybi, on the other hand, discloses one or more processors comprising processing circuitry (Column 2 lines 45-54).
Wouhaybi is combinable with Wang because they are from the same art of image processing. It is well-known in the art that a processor comprising processing circuitry is required to execute the desired instructions.
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate one or more initial responses representative of a selected sensor of the plurality of sensors that is relevant to a designated task, generate one or more subsequent responses evaluating the designated task based at least on one or more frames of the sensor data generated using the selected sensor, and one or more processors comprising processing circuitry, as taught by Wouhaybi, into the instructions, as taught by Wang, to obtain the invention as specified in claim 1.
Claim 10 has limitations that are substantially similar to claim 1. Therefore, the rejection applied to claim 1, please see above, also applies equally to claim 10. Furthermore, Wang discloses a method (Section 3 “Proposed Method”; Section 4.2 paragraph 3; Tables 3 and 4, wherein results of different methods, including DriveMLM, are compared).
Regarding claim 7, Wang and Wouhaybi disclose the one or more processors of claim 1.
Wang and Wouhaybi teach one or more initial responses representative of a selected camera of the plurality of cameras that is relevant to the designated task. Wang additionally discloses wherein the sensor data comprises image data generated using a plurality of cameras of the ego-machine (Abstract; Section 3.1 paragraph 1; pg. 5 Fig. 3, wherein DriveMLM receives sensor inputs (i.e. sensor data) from the vehicle (i.e. ego-machine) including multi-view images and LiDAR (i.e. image data generated using a plurality of cameras)), the multi-modal language model comprises a vision language model (VLM) (Section 3.1 paragraph 1; Fig. 3; Section 3.3 – “MLLM Decoder”; Table 1 – “System Message”, wherein DriveMLM (i.e. the multi-modal language model) takes in inputs including multi-view images and LiDAR, which are visual inputs and therefore, DriveMLM comprises a VLM; Section 4.2 paragraph 1, wherein ViT-g/14 from EVA-CLIP is used as the visual encoder in DriveMLM), and the one or more processors are further to provide a prompt to the VLM to generate the one or more initial responses representative of a selected camera of the plurality of cameras that is relevant to the designated task (Table 1; Section 3.2(3), wherein an initial response relevant to a designated task from the given prompt is generated).
However, as indicated by the double strikethroughs above, Wang fails to teach one or more initial responses representative of a selected sensor of the plurality of sensors that is relevant to a designated task, specifically (emphasis added). Wouhaybi, on the other hand, teaches an ego-machine with multiple sensors using a neural network to select a camera based on relevancy to a designated task. More specifically and as it relates to the applicant’s claims, Wouhaybi discloses one or more initial responses representative of a selected sensor of the plurality of sensors that is relevant to a designated task (Column 8 lines 8-15 and 26-42, wherein sensor data is received from various external sensors of a robot and a neural network is configured to select (i.e. initial response) one or more sensors for data delivery based on usefulness/relevance of sensor data or task requirements (i.e. representative of a selected sensor of the plurality of sensors that is relevant to a designated task).
Wouhaybi is combinable with Wang because they are from the same art of image processing.
The suggestion/motivation for doing so would have been to increase power and computational efficiency (Wouhaybi, Column 5 lines 9-10).
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate one or more initial responses representative of a selected sensor of the plurality of sensors that is relevant to a designated task, as taught by Wouhaybi, into the instructions, as taught by Wang, to obtain the invention as specified in claim 7.
Regarding claim 8, Wang and Wouhaybi disclose the one or more processors of claim 1.
Wang additionally discloses wherein the one or more subsequent responses by the multi-modal language model indicate one or more results of performing one or more computer vision tasks (Section 4.6, wherein control process involves analyzing the road conditions (i.e. computer vision task, of processing images to extract information) to make decision choices and provide explanatory statements (i.e. subsequent responses)) using the multi-modal language model, at least one computer vision task of the one or more computer vision tasks including at least one of:
driver drowsiness detection,
driver distraction detection,
driver out-of-position detection,
driver or occupant identification,
seatbelt usage detection,
occupant presence detection,
occupant classification,
child presence detection,
gesture recognition,
sign recognition,
context-aware question answering (Section 4.4 wherein DriveMLM outputs explanation of decisions in the form of question and answer; Fig. 4 and 5, wherein directional positions of obstacles and road conditions are identified (i.e. context-aware)),
recognition of one or more objects left behind, or
suspicious activity monitoring.
Regarding claim 9, Wang and Wouhaybi disclose the one or more processors of claim 1.
Wang additionally discloses wherein the one or more processors are comprised in at least one of:
a control system for an autonomous or semi-autonomous machine (Abstract; Section 3.1 paragraph 1, wherein DriveMLM’s decision-stage output is sent to a motion planning and control module (i.e. control system) that computes the trajectory of the vehicle (i.e. autonomous machine) in an autonomous driving system);
a perception system for an autonomous or semi-autonomous machine (Abstract; Section 3.2; Table 2, wherein perception of DriveMLM is compared; Table 1, wherein Q1/A1 is a scene-perception function performed as part of the autonomous driving system of the vehicle (i.e. perception system for an autonomous machine));
a system for performing simulation operations;
a system for performing digital twin operations;
a system for performing light transport simulation;
a system for performing collaborative content creation for 3D assets;
a system for performing deep learning operations (Abstract; section 3.3, wherein a multi-modal tokenizer that converts raw data into image token embeddings and the MLLM decoder translates the tokenized inputs into decision states and explanations; section 4.2 paragraph 1, wherein DriveMLM is trained with instruction following data);
a system for performing remote operations;
a system for performing real-time streaming;
a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content;
a system implemented using an edge device;
a system implemented using a robot;
a system for performing conversational AI operations;
a system implementing one or more language models (Abstract; Sections 3.1-3.3, wherin DriveMLM is described as a multi-modal language model);
a system implementing one or more large language models (LLMs) (Abstract, section 4.2 paragraph 1, wherein LLaMA-7B is used as the LLM in DriveMLM);
a system implementing one or more vision language models (VLMs) (Section 3.1 paragraph 1; Fig. 3; Section 3.3 – “MLLM Decoder”; Table 1 – “System Message”, wherein DriveMLM takes in inputs including multi-view images and LiDAR, which are visual inputs and therefore, DriveMLM comprises a VLM; Section 4.2 paragraph 1, wherein ViT-g/14 from EVA-CLIP is used as the visual encoder in DriveMLM);
a system implementing one or more multi-modal language models (Abstract; Sections 3.1-3.3, wherin DriveMLM is described as a multi-modal language model);
a system for generating synthetic data;
a system for generating synthetic data using AI;
a system for performing one or more generative AI operations;
a system incorporating one or more virtual machines (VMs);
a system implemented at least partially in a data center; or
a system implemented at least partially using cloud computing resources.
Claim 15 has limitations that are substantially similar to claim 9. Therefore, the rejection applied to claim 9, please see above, also applies equally to claim 15.
Claims 2, 6, and 11 are rejected under 35 U.S.C. 103 as being unpatentable over Wang et al. (NPL titled "DriveMLM: Aligning multi-modal large language models with behavioral planning states for autonomous driving”) (“Wang”) in view of Wouhaybi et al. (U.S. Patent No. US 12348860 B2) (“Wouhaybi”) and further in view of Ding et al. (NPL titled “HiLM-D: Towards high-resolution understanding in multimodal large language models for autonomous driving”) (“Ding”).
Regarding claim 2, Wang and Wouhaybi disclose the one or more processors of claim 1.
Although Wang and Wouhaybi teach the multi-modal language model generating the one or more initial responses representative of the selected sensor, Wang and Wouhaybi fail to teach wherein the processing circuitry is further to increase an input resolution of at least one frame of the one or more frames of the sensor data applied to the multi-modal language model based at least on the multi-modal language model generating the one or more initial responses representative of the selected sensor. Ding, on the other hand, teaches increasing the resolution of an input frame for HiLM-D. More specifically and as it relates to the applicant’s claims, Ding discloses wherein the processing circuitry is further to increase an input resolution of at least one frame of the one or more frames of the sensor data applied to the multi-modal language model (Fig. 5; pg. 6 section "Performance comparison of the baseline and ours across different resolution inputs", wherein the resolution of inputs is increased for the high-resolution perception branch of HiLM-D (i.e. multi-modal language model)) based at least on the multi-modal language model generating the one or more initial responses representative of the selected sensor.
Ding is combinable with Wang and Wouhaybi because they are from the same art of image processing.
The suggestion/motivation for doing so would have been to enhance perception ability, especially for small objects (Ding, pg. 5 section "Comparison with the State-of-the-Art-Methods" paragraph 2; pg. 6 section "Performance comparison of the baseline and ours across different resolution inputs").
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate wherein the processing circuitry is further to increase an input resolution of at least one frame of the one or more frames of the sensor data applied to the multi-modal language model based at least on the multi-modal language model generating the one or more initial responses representative of the selected sensor, as taught by Ding, into the one or more processors, as taught by Wang and Wouhaybi, to obtain the invention as specified in claim 2.
Claim 11 has a limitation that is substantially similar to claim 2. Therefore, the rejection applied to claim 2, please see above, also applies equally to claim 11.
Regarding claim 6, Wang and Wouhaybi disclose the one or more processors of claim 1.
Wang and Wouhaybi fail to teach wherein the processing circuitry is further to prompt the multi-modal language model to identify, within the one or more frames of the sensor data generated using the selected sensor, one or more regions of interest associated with the designated task. Ding, on the other hand, teaches identifying bounding boxes representing risk objects. More specifically and as it relates to the applicant’s claims, Ding discloses wherein the processing circuitry is further to prompt the multi-modal language model (pg. 4 section "Enumeration module", wherein high-risk objects are ones that occupy a small part of the entire image and a prompt of "Where are the vehicles, traffic lights/cones, and people?" is given as an example prompt) to identify, within the one or more frames of the sensor data generated using the selected sensor, one or more regions of interest associated with the designated task (Fig. 4, pg. 4 section "Query detection head", wherein bounding boxes (i.e. regions of interest) representing risk objects (i.e. designated task, from the example prompt) are identified).
Ding is combinable with Wang and Wouhaybi because they are from the same art of image processing.
The suggestion/motivation for doing so would have been to generate more precise bounding boxes to accurately identify the highest risk object (Ding, pg. 4 section "Query detection head"; pg. 5 Fig. 4 caption).
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate wherein the processing circuitry is further to prompt the multi-modal language model to identify, within the one or more frames of the sensor data generated using the selected sensor, one or more regions of interest associated with the designated task, as taught by Ding, into the one or more processors, as taught by Wang and Wouhaybi, to obtain the invention as specified in claim 2.
Claims 3-4 and 12-13 are rejected under 35 U.S.C. 103 as being unpatentable over Wang et al. (NPL titled "DriveMLM: Aligning multi-modal large language models with behavioral planning states for autonomous driving”) (“Wang”) in view of Wouhaybi et al. (U.S. Patent No. US 12348860 B2) (“Wouhaybi”) and further in view of Kosugi et al. (U.S. Publication No. US 2021/0096632 A1) (“Kosugi”).
Regarding claim 3, Wang and Wouhaybi disclose the one or more processors of claim 1.
Wang and Wouhaybi teach wherein the processing circuitry is further to prompt the multi-modal language model to generate the one or more initial responses
However, as indicated by the double strikethroughs above, Wang and Wouhaybi fail to teach wherein the processing circuitry is further to prompt the multi-modal language model to generate the one or more initial responses based at least on at least one first frame of the one or more frames of the sensor data, and prompt the multi-modal language model to generate the one or more subsequent responses based at least on at least one second frame of the one or more frames of the sensor data, the at least one first frame having a first resolution and the at least one second frame having a second resolution that is higher than the first resolution, specifically (emphasis added). Kosugi, on the other hand, teaches identifying an object based off a first frame and then identifying a human face based on a second frame with a resolution higher than the resolution of the first frame. More specifically and as it relates to the applicant’s claims, Kosugi discloses wherein the processing circuitry is further to prompt the multi-modal language model to generate the one or more initial responses based at least on at least one first frame of the one or more frames of the sensor data, and prompt the multi-modal language model to generate the one or more subsequent responses based at least on at least one second frame of the one or more frames of the sensor data, the at least one first frame having a first resolution and the at least one second frame having a second resolution that is higher than the first resolution ([0134], wherein an object is detected (i.e. initial response) in the first image (i.e. first frame) having a first resolution and then a second image (i.e. second frame) at a second resolution higher than the first resolution is used to detect a human face, subsequently (i.e. subsequent response)).
Kosugi is combinable with Wang and Wouhaybi because they are from the same art of image processing.
The suggestion/motivation for doing so would have been to reduce false detection while reducing consumed power and amount of processing (Kosugi, [0130]).
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate wherein the processing circuitry is further to prompt the multi-modal language model to generate the one or more initial responses based at least on at least one first frame of the one or more frames of the sensor data, and prompt the multi-modal language model to generate the one or more subsequent responses based at least on at least one second frame of the one or more frames of the sensor data, the at least one first frame having a first resolution and the at least one second frame having a second resolution that is higher than the first resolution, as taught by Kosugi, into the one or more processors, as taught by Wang and Wouhaybi, to obtain the invention as specified in claim 3.
Regarding claim 4, Wang and Wouhaybi disclose the one or more processors of claim 1.
Wang and Wouhaybi teach wherein the processing circuitry is further to increase a rate of providing prompts to the multi-modal language model to generate additional subsequent responses based at least on the multi-modal language model identifying the selected sensor, specifically (emphasis added). Kosugi, on the other hand, teaches increasing the frame rate of the image sensor to detect a human face and whether a user is present. More specifically and as it relates to the applicant’s claims, Kosugi discloses wherein the processing circuitry is further to increase a rate of providing prompts to the multi-modal language model to generate additional subsequent responses ([0134], wherein the image processor receives (i.e. prompted) a first image (i.e. first frame) at a first frame rate to detect an object, and then receives a second image (i.e. second frame) detected at a second frame rate higher than the first frame rate is used to detect a human face (i.e. additional subsequent responses) based at least on the multi-modal language model identifying the selected sensor.
Kosugi is combinable with Wang and Wouhaybi because they are from the same art of image processing.
The suggestion/motivation for doing so would have been reduce false detection while reducing consumed power and amount of processing (Kosugi, [0130]).
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate wherein the processing circuitry is further to increase a rate of providing prompts to the multi-modal language model to generate additional subsequent responses based at least on the multi-modal language model identifying the selected sensor., as taught by Kosugi, into the one or more processors, as taught by Wang and Wouhaybi, to obtain the invention as specified in claim 4.
Claims 12 and 13 have limitations that are substantially similar to claims 3 and 4, respectively. Therefore, the rejection applied to claims 3 and 4, please see above, also applies equally to claim 12 and 13, respectively.
Claims 5 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Wang et al. (NPL titled "DriveMLM: Aligning multi-modal large language models with behavioral planning states for autonomous driving”) (“Wang”) in view of Wouhaybi et al. (U.S. Patent No. US 12348860 B2) (“Wouhaybi”) and further in view of Appaya Dhanabalan et al. (U.S. Publication No. US 2025/0138546 A1) (“Appaya Dhanabalan”).
Regarding claim 5, Wang and Wouhaybi discloses the one or more processors of claim 1.
Wang and Wouhaybi teach wherein the processing circuitry is further to provide a prompt to the multi-modal language model to generate the one or more initial responses representative of the selected sensor
However, as indicated by the double strikethroughs above, Wang and Wouhaybi fail to teach wherein the processing circuitry is further to provide a prompt to the multi-modal language model to generate the one or more initial responses representative of the selected sensor based at least on applying at least one of:
a tiled representation of a plurality of frames of sensor data generated using the plurality of sensors; or
one or more compressed representations of a plurality of frames of sensor data generated using the plurality of sensors, specifically (emphasis added). Appaya Dhanabalan, on the other hand, teaches representing the inputs from the multiple sensors in a single grid. More specifically and as it relates to the applicant’s claims, Appaya Dhanabalan discloses wherein the processing circuitry is further to provide a prompt to the multi-modal language model to generate the one or more initial responses representative of the selected sensor based at least on applying at least one of:
a tiled representation of a plurality of frames of sensor data generated using the plurality of sensors (Fig. 7 and 8; [0021], [0023], [0052], and [0054], wherein detections from multiple sensor frames are combined onto a single occupancy grid (i.e. tiled representation of a plurality of frames of sensor data generated using the plurality of sensors); or
one or more compressed representations of a plurality of frames of sensor data generated using the plurality of sensors.
Appaya Dhanabalan is combinable with Wang and Wouhaybi because they are from the same art of image processing.
The suggestion/motivation for doing so would have been to increase effectiveness of perception modules on a vehicle while reducing processing latency (Appaya Dhanabalan, [0024]).
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate wherein the processing circuitry is further to provide a prompt to the multi-modal language model to generate the one or more initial responses representative of the selected sensor based at least on applying at least one of: a tiled representation of a plurality of frames of sensor data generated using the plurality of sensors; or one or more compressed representations of a plurality of frames of sensor data generated using the plurality of sensor, as taught by Appaya Dhanabalan, into the one or more processors, as taught by Wang and Wouhaybi, to obtain the invention as specified in claim 5.
Claim 14 has limitations that are substantially similar to claim 5. Therefore, the rejection applied to claim 5, please see above, also applies equally to claim 14.
Claims 16 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Isonni et al. (U.S. Publication No. US 2023/0131458 A1) (“Isonni”) in view of Wang et al. (NPL titled "DriveMLM: Aligning multi-modal large language models with behavioral planning states for autonomous driving”) (“Wang”) and further in view of Wouhaybi et al. (U.S. Patent No. US 12348860 B2) (“Wouhaybi”).
Regarding claim 16, Isonni discloses a system ([0042], [0044], and [0049]) comprising one or more processors to control ([0178] and [0180-0181]), within a simulation of an environment that is rendered using one or more light transport simulation algorithms ([0060] and [0121], wherein the environment is rendered using ray tracing (i.e. light transport simulation algorithm)), one or more operations of a simulated ego-machine in the simulated environment (Abstract; [0045-0046] and [0082], wherein robot corresponds to ego-machine)
However, as shown with the double strikethroughs above, Isonni fails to teach a system comprising one or more processors to control, within a simulation of an environment that is rendered using one or more light transport simulation algorithms, one or more operations of a simulated ego-machine in the simulated environment based at least on one or more outputs of one or more multi-modal language models, specifically (emphasis added). Wang, on the other hand, teaches controlling a vehicle based on the outputs of DriveMLM. More specifically and as it relates to the applicant’s claims, Isonni in view of Wang discloses a system comprising one or more processors to control, within a simulation of an environment that is rendered using one or more light transport simulation algorithms, one or more operations of a simulated ego-machine in the simulated environment based at least on one or more outputs of one or more multi-modal language models (Section 1 paragraph 3, wherein DriveMLM performs autonomous driving of the vehicle (i.e. ego-machine) by converting decisions (i.e. outputs) of DriveMLM (i.e. the multi-modal language model) into vehicle control signals; Section 3.1 paragraph 2, wherein decision state output is input into a motion planning and control module, computing the final trajectory for vehicle control).
Wang is combinable with Isonni because they are from the same art of image processing.
The suggestion/motivation for doing so would have been to ensure a comprehensive dataset encompassing decision states, decision explanations, and user commands (Wang, section 3.1 paragraph 1).
Although Isonni in view of Wang teaches the one or more outputs generated based at least on the multi-modal language model and the simulated sensor data generated using the plurality of simulated sensors of the simuilated ego-machine (Isonni, [0045], [0060], and [0082]), Isonni and Wang fail to teach the one or more outputs generated based at least on the multi-modal language model evaluating whether one or more conditions associated with a designated task are present in one or more frames of simulated sensor data generated using a selected simulated sensor, the selected simulated sensor being selected from a plurality of simulated sensors based at least on the multi-modal language model evaluating a plurality of frames of simulated sensor data generated using the plurality of simulated sensors of the simulated ego-machine, specifically.
Wouhaybi, on the other hand, teaches an ego-machine with multiple sensors using a neural network to select a camera based on relevancy to a designated task and the calculation of a confidence factor based on the selected sensor information relative to the designated task. If the confidence factor is within a predetermined range, the robot continues to receive sensor data from the selected sensor. If the confidence factor is outside of the predetermined range, then a different sensor will be triggered. More specifically and as it relates to the applicant’s claims, Wouhaybi discloses the one or more outputs generated based at least on the multi-modal language model evaluating whether one or more conditions associated with a designated task are present in one or more frames of simulated sensor data generated using a selected simulated sensor (Column 9 lines 33-53, wherein a confidence factor based on sensor information is calculated for the frame (i.e. evaluating whether one or more conditions associated with a designated task are present) and if the confidence factor is within a predetermined range, sensor data from the selected sensor will continually be received, or if the confidence factor is outside of the predetermined range, then a different sensor will be selected (i.e. subsequent response using the selected sensor)), the selected simulated sensor being selected from a plurality of simulated sensors based at least on the multi-modal language model evaluating a plurality of frames of simulated sensor data generated using the plurality of simulated sensors of the simulated ego-machine (Column 8 lines 8-15 and 26-42, wherein sensor data is received from various external sensors of a robot and a neural network is configured to select one or more sensors for data delivery based on usefulness/relevance of sensor data or task requirements).
Wouhaybi is combinable with Isonni and Wang because they are from the same art of image processing.
The suggestion/motivation for doing so would have been to increase power and computational efficiency (Wouhaybi, Column 5 lines 9-10).
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate controlling one or more operations of a simulated ego-machine in the simulated environment based at least on one or more outputs of one or more multi-modal language models, as taught by Wang, and the one or more outputs generated based at least on the multi-modal language model evaluating whether one or more conditions associated with a designated task are present in one or more frames of simulated sensor data generated using a selected simulated sensor, the selected simulated sensor being selected from a plurality of simulated sensors based at least on the multi-modal language model evaluating a plurality of frames of simulated sensor data generated using the plurality of simulated sensors of the simulated ego-machine, as taught by Wouhaybi, into the system, as taught by Isonni, to obtain the invention as specified in claim 16.
Regarding claim 19, Isonni, Wang, and Wouhaybi disclose the system of claim 16.
Isonni, Wang, and Wouhaybi teach wherein the evaluating whether one or more conditions associated with a designated task are present in one or more frames of simulated sensor data comprises
However, as indicated by the double strikethroughs above, Isonni, Wang, and Wouhaybi fail to teach wherein the evaluating whether one or more conditions associated with a designated task are present in one or more frames of simulated sensor data comprises providing one or more text prompts to the multi-modal language model to generate one or more responses to evaluate a presence of the one or more conditions associated with the designated task based at least on one or more frames of generated simulated sensor data corresponding to the selected simulated sensor, specifically (emphasis added). Wang, on the other hand, teaches providing one or more text prompts to the multi-modal language model to generate one or more responses. More specifically and as it relates to the applicant’s claims, Wang discloses providing one or more text prompts to the multi-modal language model to generate one or more responses (Table 1; Section 4.4 paragraph 1; Tables B-D in “Supplementary Material”)
The suggestion/motivation for doing so would have been to efficiently receive desired answers based on description, driving decisions, or explanations (Wang, Section A paragraph 2).
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate providing one or more text prompts to the multi-modal language model to generate one or more responses, as taught by Wang, into the system, as taught by Isonni, Wang, and Wouhaybi, to obtain the invention as specified in claim 19.
Claims 17-18 are rejected under 35 U.S.C. 103 as being unpatentable over Isonni et al. (U.S. Publication No. US 2023/0131458 A1) (“Isonni”) in view of Wang et al. (NPL titled "DriveMLM: Aligning multi-modal large language models with behavioral planning states for autonomous driving”) (“Wang”) and Wouhaybi et al. (U.S. Patent No. US 12348860 B2) (“Wouhaybi”), and further in view of Tran et al. (NPL titled “Immersion into 3D Biomedical Data Via Holographic AR Interfaces Based on the Universal Scene Description (USD) Standard”) (“Tran”).
Regarding claim 17, Isonni, Wang, and Wouhaybi disclose the system of claim 16.
However, Isonni, Wang, and Wouhaybi fail to teach wherein the simulation is generated, at least in part, using a three-dimensional (3D) content collaboration platform for 3D assets. Tran, on the other hand, teaches using Omniverse to create simulation solutions for perception AI applications. More specifically and as it relates to the applicant’s claims, Tran discloses wherein the simulation is generated, at least in part, using a three-dimensional (3D) content collaboration platform for 3D assets (Abstract; Section I paragraphs 2-3; Section IV paragraph 2, wherein Omniverse is a content collaboration platform for 3D assets).
Tran is combinable with Isonni, Wang, and Wouhaybi because they are from the same art of image processing.
The suggestion/motivation for doing so would have been to robustly and scalably interchange and augment 3D scenes among different tools (Tran, Section I paragraph 2).
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate wherein the simulation is generated, at least in part, using a three-dimensional (3D) content collaboration platform for 3D assets, as taught by Tran, into the system, as taught by Isonni, Wang, and Wouhaybi, to obtain the invention as specified in claim 17.
Regarding claim 18, Isonni, Wang, and Wouhaybi disclose the system of claim 17.
However, Isonni, Wang, and Wouhaybi fail to teach wherein one or more files of the 3D content collaboration platform for 3D assets uses an OpenUSD format. Tran, on the other hand, teaches using Omniverse, which is built on OpenUSD. More specifically and as it relates to the applicant’s claims, Tran discloses wherein one or more files of the 3D content collaboration platform for 3D assets uses an OpenUSD format (Abstract; Section I paragraphs 2-3, wherein Omniverse is built on USD, which is also referred to as OpenUSD).
Tran is combinable with Isonni, Wang, and Wouhaybi because they are from the same art of image processing.
The suggestion/motivation for doing so would have been to provide rich and varied ways to combine assets into larger assemblies and enable collaborative workflows (Tran, Section I paragraph 2).
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate wherein one or more files of the 3D content collaboration platform for 3D assets uses an OpenUSD format, as taught by Tran, into the system, as taught by Isonni, Wang, and Wouhaybi, to obtain the invention as specified in claim 17.
Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Isonni et al. (U.S. Publication No. US 2023/0131458 A1) (“Isonni”) in view of Wang et al. (NPL titled "DriveMLM: Aligning multi-modal large language models with behavioral planning states for autonomous driving”) (“Wang”) and Wouhaybi et al. (U.S. Patent No. US 12348860 B2) (“Wouhaybi”), and further in view of Cain et al. (U.S. Publication No. US 2025/0247213 A1) (“Cain”).
Regarding claim 20, Isonni, Wang, and Wouhaybi discloses the system of claim 16.
Although Isonni, Wang, and Wouhaybi teach a multi0modal language model, Isonni, Wang, and Wouhaybi fail to teach wherein at least one multi-modal language model of the one or more multi-modal language models is implemented in at least one processing node of a plurality of processing nodes of a data center and accessible to one or more remote clients via at least one of an application programming interface (API), or an application plug-in. Cain, on the other hand, teaches a processor, using a large language model, located in a data center and is located in a remote server to be accessed via an API. More specifically and as it relates to the applicant’s claims, Cain discloses wherein at least one multi-modal language model of the one or more multi-modal language models is implemented in at least one processing node of a plurality of processing nodes of a data center and accessible to one or more remote clients via at least one of an application programming interface (API), or an application plug-in ([0016], [0050], [0057], [0095], and [0101]).
Cain is combinable with Isonni, Wang, and Wouhaybi because they are from the same art of data processing.
The suggestion/motivation for doing so would have been to allow convenient, secure, and cost-effective sharing of similar application features between multiple sets of users (Cain, [0021]).
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate wherein at least one multi-modal language model of the one or more multi-modal language models is implemented in at least one processing node of a plurality of processing nodes of a data center and accessible to one or more remote clients via at least one of an application programming interface (API), or an application plug-in, as taught by Cain, into the system, as taught by Isonni, Wang, and Wouhaybi, to obtain the invention as specified in claim 20.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Hermann et al. (U.S. Publication No. US 2021/0110115 A1) teaches controlling an autonomous vehicle by perceiving the current state of the environment through sensors by inputting a text prompt regarding a task to the language model to produce an action selection output that selects the best possible action to take. Example tasks may be navigating to a certain location or finding a specific object.
Contact
Any inquiry concerning this communication or earlier communications from the examiner should be directed to RACHEL Y DANG whose telephone number is (571)438-9519. The examiner can normally be reached Monday - Thursday: 7am - 4:30pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, John Villecco can be reached at (571) 272-7319. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/RACHEL Y DANG/Examiner, Art Unit 2661
/AARON W CARTER/Primary Examiner, Art Unit 2661