Detailed Action
1. Claims 1-20 are pending in this Application.
Notice of Pre-AIA or AIA Status
2. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Election /Restriction
3 Applicant’s election without traverse, claims 1-20 based on amendment of claim 12, in the reply filed on 08/10/2026 is acknowledged.
4. No claims are withdrawn from further consideration because the applicant has amended claim 12 to match claim 19 and Group I. Prior to the amendment, claim 12 corresponded to Group II.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
5. Claims 1-8 and 10-20 are rejected under 35 U.S.C. 103 as being unpatentable over
Sehrish Malik et al., (hereafter Sehrish), “Hybrid Inference Based Scheduling Mechanism for Efficient Real Time Task and Resource Management in Smart Cars for Safe Driving”, Electronics 2019, pub. 21 March 2019, in view of Yu Huang et al., (hereafter Yu), “Applications of Large Scale Foundation Models for Autonomous Diving”, arXiv pub., Nov. 20, 2023.
As to claim 1, Sehrish teaches One or more processors comprising processing circuitry (Figure 2, page 4, Section 3. Figure 2 illustrates autonomous smart cars that includes an Intelligence framework model for hybrid task scheduler and inference engine including Scheduling unit. Scheduling unit implements a scheduling policy as per which the arrived tasks are to be executed at the CPU) to:
queue one or more inference requests representing one or more detection tasks associated with an ego-machine (Page 3 section 3.2.1,3.2.2 and 3.2.3, Intelligence framework model for hybrid task scheduler and inference engine of the autonomous smart cars includes Fair Emergency First Scheduling Policy, Priority Based Scheduling Policy and Earliest Deadline First Scheduling Policy. Earliest deadline first (EDF) or least time to go is a dynamic priority scheduling algorithm used in real-time operating systems to place processes in a priority queue. Whenever a scheduling event occurs (task finishes, new task released, etc.) the queue will be searched for the process closest to its deadline);
prompt one or more (Figs.4-5 Section 3.3Smart car systems use an inference engine to analyze contextual scenarios and generate automated responses. For instance, when a driver is detected, the system relies on input data from sensors like a camera and pressure sensor to trigger a task like automatic seat adjustment. Similarly, when a pedestrian is detected or environmental changes occur, the system fires specific rules and processes exceptions to ensure safe driving choices); and
control one or more operations of the ego-machine based at least on the one or more responses (Section 4.3, 5th par., The inference engine classifies the incoming data sensing values, matches the rules and follows the rules when the goal is met. The followed rules contain the response control tasks to be executed for the safe driving of a smart car. The fired control tasks are lined at control unit and communicated back to scheduler via hybrid agent and executed at control unit when scheduled to execute).
However, it is noted that Sehrish does not specifically teach vision-language models (VLMs)
On the other hand the methods of LLM/ Vision-Language Models (VLMs) -based
autonomous driving models disclosed by Yu teaches Vision-Language Models (VLMs) (page 5 right col. section III, Vision-Language Models (VLMs) bridge the capabilities of Natural Language Processing (NLP) and Computer Vision (CV), breaking down the boundaries between text and visual information to connect multimodal data)
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Vision-Language Models (VLMs) taught by YU into Sehrish.
The suggestion/motivation for doing so would have been to allow user of Sehrish to
to simultaneously process, reason about, and bridge both visual and textual information
As to claim 2, Yu teaches the one or more VLMs comprise a first VLM to produce an output corresponding to different types of detection tasks associated with the ego- machine(Fig. 29, TABEL III, page 12 left col., 1 section IV abstract, VI, In Autonomous driving solutions, the LLMs and VLMs provide a contextual understanding of driving environment, such as detection, tracking and segmentation of obstacles, road signs/marking and free space drivable areas).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Vision-Language Models (VLMs) taught by YU into Sehrish.
The suggestion/motivation for doing so would have been to allow user of Sehrish to
to simultaneously process, reason about, and bridge both visual and textual information
As to claim 3, Yu teaches the one or more VLMs comprise a first VLM to produce an output corresponding to one or more interior detection tasks ( as discussed above in claim 2, the LLMs and VLMs provide a contextual understanding of driving environment, such as detection, tracking and segmentation of obstacles, road signs/marking and free space drivable areas)and a second VLM to produce an output corresponding to one or more exterior detection tasks associated with the ego-machine(Table , the motion planning for trajectory prediction , decision making for action command, make decision and planning…)
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Vision-Language Models (VLMs) taught by YU into Sehrish.
The suggestion/motivation for doing so would have been to allow user of Sehrish to
to simultaneously process, reason about, and bridge both visual and textual information.
As to claim 4, Sehrish teaches the processing circuitry is further to queue a plurality of inference requests submitted by a plurality of detection applications of the ego- machine ( page 6-7, section 3.2.2, 3.2.3. Earliest Deadline First Scheduling Policy Earliest deadline first (EDF) or least time to go is a dynamic priority scheduling algorithm used in real-time operating systems to place processes in a priority queue. Whenever a scheduling event occurs (task finishes, new task released, etc.) the queue will be searched for the process closest to its deadline. This process is the next one to be scheduled for execution).
As to claim 5, Sehrish teaches the processing circuitry is further to prioritize scheduling the one or more inference requests based at least on an assessed importance of the one or more detection tasks to safety ( page 13, section 4.3, the hybrid agent then passes them onto the inference engine based on the priorities and levels of urgency. The Figure 11 above shows the simulation execution flow from the input sensor readings to the scheduler. At the scheduler, the sensing tasks are executed based on priorities and sent to the hybrid agent for further processing).
As to claim 6, Sehrish teaches the processing circuitry is further to prioritize scheduling one or more first requests for the one or more page 6 3.2.1 . The tasks in fair emergency first scheduling policy are classified into four categories: high urgency event driven tasks, normal event driven tasks, high priority periodic tasks and normal periodic tasks. A default high priority is given to the event driven tasks over the periodic tasks. Urgent event driven tasks have priority over normal event driven tasks and priority periodic tasks have priority over normal periodic tasks.)
However, it is noted that Sehrish does not specifically teach vision-language models (VLMs).
On the other hand the methods of LLM/ Vision-Language Models (VLMs) -based
autonomous driving models disclosed by Yu teaches Vision-Language Models (VLMs) (page 5 right col. section III, Vision-Language Models (VLMs) bridge the capabilities of Natural Language Processing (NLP) and Computer Vision (CV), breaking down the boundaries between text and visual information to connect multimodal data).
As to claim 7, Sehrish teaches the processing circuitry is further to prioritize scheduling one or more first requests for the one or more Abstract, page 6 3.2.1, Our proposed hybrid inference engine and task scheduler
mechanism provides an efficient way of controlling smart cars in different scenarios such as heavy rainfall, obstacle detection, driver’s focus diversion etc., The tasks in fair emergency first scheduling policy are classified into four categories: high urgency event driven tasks, normal event driven tasks, high priority periodic tasks and normal periodic tasks. A default high priority is given to the event driven tasks over the periodic tasks. Urgent driven tasks have priority over normal event driven tasks and priority periodic tasks have priority over normal periodic tasks.).
However, it is noted that Sehrish does not specifically teach vision-language models (VLMs)
On the other hand the methods of LLM/ Vision-Language Models (VLMs) -based
autonomous driving models disclosed by Yu teaches Vision-Language Models (VLMs) (page 5 right col. section III, Vision-Language Models (VLMs) bridge the capabilities of Natural Language Processing (NLP) and Computer Vision (CV), breaking down the boundaries between text and visual information to connect multimodal data. Further The local planner generates an obstacle avoiding trajectory to reach the locally generated waypoints and executes it by employing a low-level controller).
As to claim 8, Sehrish teaches the one or more detection tasks comprise at least one of driver drowsiness detection, driver distraction detection, suspicious activity monitoring ( page 4, section 3, page 9, The driver gets alerts/ notification through a user interface. The art includes notifying the driver different scenarios such as heavy rainfall, obstacle detection, driver’s focus diversion etc.,(see abstract) . Further the smart car sensing data unit collects the input data from the smart car sensing environment. This data can be various kinds, e.g., environmental data, road state data, traffic state data, driver health state data, smart car physical self-state data etc., ).
As to claim 10, Sehrish teaches the one or more operations of the ego- machine comprise at least one of issuing an audible or visual alert, adjusting one or more in-vehicle infotainment settings, activating one or more safety systems, or executing a navigational maneuver(page 3 1st par., page 9 , driver gets alerts through a user interface. The vehicular sensing data has car tire, car brake, and car speed and car distance sensing data. The sensing data sensors provide alert signal that indicate , safe warning and high alert).
As to claim 11, the combination of Sehrish and YU teaches the one or more processors are comprised in at least one of:
a control system for an autonomous or semi-autonomous machine (Sehrish:Fig.1 control unit );
a perception system for an autonomous or semi-autonomous machine (Sehrish: Table 2 below explains the environment sensing data scenario classifications, Table 3 explains the human sensing data scenario classifications and Table 4 explains the sensing data scenario);
a system for performing simulation operations (Sehrish: page 11 section 4, Simulation of Inference based Scheduling Mechanism for Safe Driving in Smart Cars);
a system for performing digital twin operations;
a system for performing light transport simulation (Sehrish: Fig.9) ;
a system for performing collaborative content creation for 3D assets ( YU: page 9 left col., 3rd par., A method called Point-E in [138] for 3D object generation is proposed.
First it generates a single synthetic view using a text to-image diffusion model, and then produces a 3D point cloud using a second diffusion model which conditions on the
generated image);
a system for performing deep learning operations (YU: page 7, right col. 2nd par., Robotics Transformer 1 (RT-1) [85] can absorb large amounts of data, effectively
generalize, and output actions at real-time rates for practical robotic control);
a system for performing remote operations (Sehrish: Fig.9: page 2 3rd par: The necessity of safe driving has resulted in the potential need for substitute technologies in automobiles such as smart cars. Internet of things (IoT) and the power of computer vision are the primary actors that enabled the existence of smart cars),
a system for performing real-time streaming (YU: page 14 left col. last par.,- right col. 1st par., Fig. 34, proposed by Wayve (an UK startup), is a generative
world model that leverages video, text, and action inputs to build realistic driving scenarios while a fine-grained control over ego-vehicle behavior and scene features is given).
a system implemented using a robot (YU: page 18 right col. 1st par., ConBaT uses a causal transformer, derived from the Perception-Action Causal Transformer (PACT), that learns to predict safe robot actions autoregressively using a critic that requires minimal safety data labeling. During deployment, a lightweight online optimization is employ);
a system for performing conversational Al operations (YU: Abstract););
a system implementing one or more language models (YU: Abstract);
a system implementing one or more large language models (LLMs) (YU: Abstract);
a system implementing one or more vision language models (VLMs) (YU: Abstract);
a system for generating synthetic data (YU: page 8 right col, 2nd par., Experimental
evidence also showed that training with synthetic data generated by diffusion models can improve task performance on tasks );
a system for generating synthetic data using Al (YU: Abstract, page 8 right col, 2nd par);
a system for performing one or more generative Al operations (YU: Abstract);
a system incorporating one or more virtual machines (VMs) (YU: page 8 left col., 2nd par. Habitat is a simulation platform for training virtual robots in interactive 3D environments and complex physics-enabled scenarios). ; or
a system implemented at least partially using cloud computing resources (page 9 left col., 3rd par., A method called Point-E in [138] for 3D object generation is proposed.
First it generates a single synthetic view using a text to-image diffusion model, and then produces a 3D point cloud using a second diffusion model which conditions on the
generated image.).
Claim 12 is rejected the same as claim 1 except claim 12 is directed to a system claim. All the limitations of claim 12 are addressed in claim 1. Thus, argument analogous to that presented above for claim 1 is applicable to claim 12.
Claim 13 is rejected the same as claim 2 except claim 13 is directed to a system claim. All the limitations of claim 13 are addressed in claim 2. Thus, argument analogous to that presented above for claim 2 is applicable to claim 13.
Claim 14 is rejected the same as claim 3 except claim 14 is directed to a system claim. All the limitations of claim 14 are addressed in claim 3. Thus, argument analogous to that presented above for claim 3 is applicable to claim 14.
Claim 15 is rejected the same as claim 4 except claim 15 is directed to a system claim. All the limitations of claim 15 are addressed in claim 4. Thus, argument analogous to that presented above for claim 4 is applicable to claim 15.
Claim 16 is rejected the same as claim 5 except claim 16 is directed to a system claim. All the limitations of claim 16 are addressed in claim 5. Thus, argument analogous to that presented above for claim 5 is applicable to claim 16.
Claim 17 is rejected the same as claim 6 except claim 17 is directed to a system claim. All the limitations of claim 17 are addressed in claim 6. Thus, argument analogous to that presented above for claim 6 is applicable to claim 17.
Claim 18 is rejected the same as claim 11 except claim 18 is directed to a system claim. All the limitations of claim 18 are addressed in claim 11. Thus, argument analogous to that presented above for claim 11 is applicable to claim 18.
Claim 19 is rejected the same as claim 1 except claim 19 is directed to a method claim. All the limitations of claim 19 are addressed in claim 1. Thus, argument analogous to that presented above for claim 1 is applicable to claim 19.
Claim 20 is rejected the same as claim 11 except claim 20 is directed to a method claim. All the limitations of claim 20 are addressed in claim 11. Thus, argument analogous to that presented above for claim 11 is applicable to claim 20.
5. Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Sehrish in view of Yu further in view of QASEMI et al., ( hereafter QASEMI), US 20240029422 A1, pub 01/25/2024.
Regarding, claim 9, while modified Sehrish teaches claim 1, but fails to teach claim 9
On the other hands the vision language model disclosed by QASEMI teaches he one or more responses indicate one or more results of malicious intent detection performed using the one or more VLMs based at least on the one or more frames of image data(claims 1, 10 and 12
the deep learning model is a natural language model or vision language model which is applied to detect one or more fine-grained classification includes an indication of suspicious or criminal activity based upon movement characteristics of the one or more objects detected by the camera).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to incorporate a fine-grained classification method configured to detect suspicious or criminal activity, based upon movement characteristics of the one or more objects detected by the cameras taught by QASEMI, into the modified Sehrish system.
The motivation for doing so would have been to allow users of Sehrish to develop a robust method of identifying suspicious or criminal activity.
Prior arts are not used in rejections but pertinent to the claims or disclosure.
a.. “DriveMLM: Aligning Multi-Modal Large Language Models with Behavioral Planning States for Autonomous Driving by Erfei Cui et al., disclosed:
In this work, we have presented DriveMLM, a novel framework that leverages large language models (LLMs) for autonomous driving (AD). DriveMLM can perform close-loop AD in realistic simulators by using a multimodalLLM (MLLM) to model the behavior planning module of a modular AD system. DriveMLM can also generate natural language explanations for its driving decisions, which can enhance the transparency and trustworthiness of the AD system. We have shown that DriveMLM can outperform the Apollo baseline on the CARLA Town05 Long benchmark. We believe that our work can inspire more research on the integration of LLMs and AD( see section 5 page 11)
PNG
media_image1.png
388
350
media_image1.png
Greyscale
b.. “A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models”
by Jindong Gu et al., disclosed:
This survey paper on prompt engineering of pre-trained vision language models has provided valuable insights into the current state of research in this field. The main findings and trends identified through the analysis shed light on the effective utilization of prompts in adapting large pre-trained models for vision-language
tasks.
One key finding is the versatility and applicability of prompt engineering across different types of vision-language models, including multimodal-to-text generation models, image-text matching models, and text-to-image generation models. The survey explored each model type from their respective characteristics, highlighting various prompting methods on them.
The implications of these findings are significant for both academia and industry. By leveraging prompt engineering techniques, researchers can achieve remarkable performance gains in vision-language models without the need for extensive labeled data. This has the potential to reduce the burden of data annotation and accelerate the deployment of vision-language models in real-world applications (see page 16 section 8)
Contact Information
Any inquiry concerning this communication or earlier communication from the examiner should be directed to Mekonen Bekele whose telephone number is (469) 295-9077.The examiner can normally be reached on Monday -Friday from 9:00AM to 6:50 PM Eastern Time.
If attempt to reach the examiner by telephone are unsuccessful, the examiner’s supervisor Eng, George can be reached on (571) 272-7495.The fax phone number for the organization where the application or proceeding is assigned is 571-237-8300. Information regarding the status of an application may be obtained from the patent Application Information Retrieval (PAIR) system. Status information for published application may be obtained from either Private PAIR or Public PAIR.
Status information for unpublished application is available through Privet PAIR only.
For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have question on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866.217-919 (tool-free)
/MEKONEN T BEKELE/ Primary Examiner, Art Unit 2699