Prosecution Insights
Last updated: October 02, 2026
Application No. 19/040,312

OBJECT IDENTIFICATION USING AUDIBLE CUES FOR AUTONOMOUS AND SEMI-AUTONOMOUS SYSTEMS AND APPLICATIONS

Final Rejection §103
Filed
Jan 29, 2025
Priority
Nov 18, 2020 — continuation of 11/816,987 +1 more
Examiner
ALAM, MIRZA F
Art Unit
2688
Tech Center
2600 — Communications
Assignee
NVIDIA Corporation
OA Round
2 (Final)
74%
Grant Probability
Favorable
3-4
OA Rounds
8m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 74% — above average
74%
Career Allowance Rate
767 granted / 1032 resolved
+12.3% vs TC avg
Strong +34% interview lift
Without
With
+33.9%
Interview Lift
resolved cases with interview
Typical timeline
2y 4m
Avg Prosecution
32 currently pending
Career history
1055
Total Applications
across all art units

Statute-Specific Performance

§101
6.0%
-34.0% vs TC avg
§103
67.1%
+27.1% vs TC avg
§102
3.2%
-36.8% vs TC avg
§112
12.4%
-27.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1032 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status 1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION 2. Applicant’s amendment filed August 24, 2012 amends claims 1-20. Claims 1-20 have been presented for examination. Applicant’s amendment has been fully considered and entered. Claim Rejections - 35 USC § 103 3. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 4. The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103(a) are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. 5. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. 6. Claims 1-3, 5, 7-13 and 15-18 and 20-21 are rejected under 35 U.S.C. 103(a) as being unpatentable over Nister (US 20190243371 A1) (hereinafter Nister) in view of Lockwood (US Patent 10268191 B1) (hereinafter Lockwood) and further in view of Salekin (US 20210005067 A1) (hereinafter Salekin). Regarding claim 1, Nister discloses an autonomous or semi-autonomous machine (FIG. 1, f autonomous vehicle system, autonomous vehicle 102 ) comprising: one or more central processing units (CPUs); one or more graphics processing units (GPUs) ( para 236, vehicle 102 include CPU(s) 1106, GPU(s) 1108, processor(s) 1110, accelerator(s) 1); one or more hardware accelerators (para 219, operate steering system 1154 via actuators 1156 to operate system 1150 via one or more accelerators 1152); and one or more sensors that are at least one of external or internal to the autonomous or semi-autonomous machine (para 59, autonomous vehicle 102 include a sensor manager 108 manage or abstract sensor data from sensors of the vehicle 102, GNSS) sensor(s) 1158, RADAR sensor(s) 1160, ultrasonic sensor(s) 1162, LIDAR sensor(s) 1164, inertial measurement unit (IMU) sensor(s) 1166, microphone(s) 1196, stereo camera(s) 1168, wide-view camera(s) 1170, infrared camera(s) 1172, surround camera(s) 1174, long range and/or mid-range camera(s) 1198, and/or other sensor types (i.e. external or internal sensors), para 220, controller(s) 1136 controlling components of vehicle 102 in response to sensor data received from sensors). Nister specifically fails to disclose the autonomous or semi-autonomous machine is to: determine, based at least on one or more neural networks processing first audio data obtained using the one or more sensors, at least a temporal interference state associated with the one or more neural networks; determine, based at least on one or more neural networks processing the temporal inference state and second audio data obtained using the one or more sensors, classification information associated with one or more sounds as represented by at least the second audio data, the first audio data associated with a first time window that is offset from and at least partially overlaps with a second time window associated with the second audio data; and perform one or more control operations based at least on the classification information. In analogous art, Lockwood discloses the autonomous or semi-autonomous machine is to: determine, based at least on one or more neural networks processing first audio data obtained using one or more sensors (Abstract, train or use a model to determine from vehicle audio data (i.e., audio data) , col. 25, lines 51-67, model 508 trained includes vehicle data 514 (e.g., sensor data, audio data, etc.) and model 508 implemented as a convolutional neural network, col. 18, lines 57-65, audio captured by external microphone, second audio captured by microphone, present three-dimensional representation data, col. 40, lines 60-67, model 508 to determine between first and second audio data from first sensor data and second sensor data)). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify teaching of audio analysis for the detection of one or more emergency by autonomous vehicles disclosed by Nister and Salekin to determine detection result indicating whether multi-tone time patterns at corresponding frequencies present in acoustic signal as taught by Lockwood to obtain assistance from a driverless vehicle and to obtain data in response to request and use train a model and use a model to determine from vehicle data option for presentation. [Lockwood, Abstract]. Nister and Lockwood fails to disclose at least a temporal interference state associated with the one or more neural networks; determine, based at least on one or more neural networks processing the temporal inference state and second audio data obtained using the one or more sensors, classification information associated with one or more sounds as represented by at least the second audio data, the first audio data associated with a first time window that is offset from and at least partially overlaps with a second time window associated with the second audio data; and perform one or more control operations based at least on the classification information. In analogous art, Salekin discloses at least a temporal interference state associated with the one or more neural networks (para 07, determining target audio event is present in the audio clip using a first neural network based on plurality of audio features, and determining position in time of the target audio event within the audio clip using a second neural network, para 34, each window segment S.sub.i has predetermined amount or percentage of temporal overlap with adjacent window segments and predetermined length of audio clip 102, predetermined length of each window segment, and predetermined amount or percentage of temporal overlap with adjacent window segments); determine, based at least on one or more neural networks processing the temporal inference state and second audio data obtained using the one or more sensors (FIG. 3 neural networks audio tagging model of audio event detection program, para 28, audio event detection model utilize neural networks, para 08, determine audio event is present in the audio clip using a neural networks based on the plurality of audio features; determine, in response to determining that audio event is present in audio clip, plurality of audio features and audio event; and determine position in time of the target audio event within audio clip using a neural networks based on the plurality of vectors, para 35, each window segment S.sub.i and has a predetermined amount or percentage of temporal overlap with adjacent sub-segments), classification information associated with one or more sounds as represented by at least one of the first audio data or the second audio data (FIG. 6-7, block 330, 340, classifier model of audio event detection program, para 27, audio event detection model, comprises audio feature extractor 32, convolution neural network (DCNN) audio tagging model 34, classifier model 38 is configured to identify boundaries or positions in time of detected target audio event in audio clip, para 41, processor 22 execute program instructions to DCNN audio tagging model 34 to determine a classification), the first audio data associated with a first time window that is offset from and at least partially overlaps with a second time window associated with the second audio data (para 27, audio extractor 32 to segment individual audio clip into a plurality of overlapping windows and represent state of audio clip in each window. DCNN audio model 34 detect presence of a target audio event and model 36 and generate vector representation of each window of audio, para 34, total number of window segments N (e.g., 148) is a function of predetermined length (e.g., 30 seconds) of audio clip 102, first predetermined length (e.g., 500 ms) of each window segment, and predetermined amount or percentage of overlap with adjacent window segments (e.g., 300 ms or 60% partially overlap)); and perform one or more control operations based at least on the classification information (Abstract, use classifier model and determine audio event based on the audio vector representations, para 29, audio event detection program 30 and/or the audio event detection model, statements that software method step performs some process/function or is configured to perform some process/function means that a processor or controller (e.g., the processor 22) executes program instructions, para 69, plurality of LSTM cells 138 perform functions to a set parameters, learned and optimized during a training process and execute program instructions corresponding to classifier model 38, para 75, processor 22 determine a classification indicating whether audio event is present using a DCNN and perform a sequence of dilated convolution operation, para 32, training dataset generated for each target audio event based on modest number of audio samples). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify teaching of audio analysis for the detection of one or more emergency by autonomous vehicles disclosed by Nister and Lockwood to determine detection result indicating whether multi-tone time patterns at corresponding frequencies present in acoustic signal as taught by Lockwood to use classifier to model long term dependencies and determine the boundaries in time of the target audio event within the audio clip based on the audio vector representations for improved audio surveillance. [Salekin, Abstract]. Regarding claim 2, Nister discloses the autonomous or semi-autonomous machine of claim 1, wherein the classification information is determined, at least, by: determining, based at least on the one or more neural networks processing the first audio data, a first portion of the classification information that includes one or more first probabilities associated with the one or more sounds (para 73, temporal filtering may be used to reduce oscillations experienced from delays that are not modeled (e.g., when no temporal element is used), para 260, neural network that outputs measure of confidence for each object detection and interpreted as a probability, or as providing a relative “weight” of each detection compared to other detections, DLA run a neural network for regressing the confidence value. The neural network take as its input for subset of parameters, such as inertial measurement unit (IMU) sensor 1166 output, 3D location estimates from the neural network and/or other sensors (e.g., LIDAR sensor(s) 1164 or RADAR sensor(s) 1160), among others, para 278, CNN for classifying environmental and urban sounds, para 334, receive data from GPU(s) 1208, CPU(s) 1206 and output data (e.g., sound, etc.)); and Nister fails to discloses determining, based at least on the one or more neural network processing the temporal inference state and the second audio data, a second portion of the classification information that includes one or more second probabilities associated with the one or more sounds. In analogous art, Lockwood discloses determining, based at least on the one or more neural network processing the temporal inference state and the second audio data, a second portion of the classification information that includes one or more second probabilities associated with the one or more sounds (col. 8, lines 27-37, predict multiple object trajectories based on probabilities determinations or multi-modal distributions of predicted positions, trajectories, velocities associated with an object, col. 11, lines 7020, calculating a Bayesian probabilities, similarly to the confidence level and determination of the criticality or priority determined by a teleoperator device after receiving of sensor data). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify teaching of audio analysis for the detection of one or more emergency by autonomous vehicles disclosed by Nister to determine detection result indicating whether multi-tone time patterns at corresponding frequencies present in acoustic signal as taught by Lockwood to determine a position in time of the target audio event within the audio clip using a second neural network based on the plurality of vectors to audio surveillance [Lockwood, para 08]. Regarding claim 3, Nister discloses the autonomous or semi-autonomous machine of claim 1, wherein the autonomous or semi-autonomous machine is further to: determine, based at least on the classification information, a sound type of one or more sounds, wherein the one or more control operations are performed based at least on the sound type (para 28, FIG. 11B, system used for depth-based object detection, especially for objects for which a neural network, object detection and classification, as well as object tracking, para 278, CNN for classifying environmental and urban sounds). Regarding claim 5, Nister discloses the autonomous or semi-autonomous machine of claim 1, wherein the autonomous or semi-autonomous machine is further to: determine, based at least on the classification information, a type of object associated with the one or more sounds, wherein the one or more control operations are performed based at least on the type of object (para 28, FIG. 11B, system used for depth-based object detection, especially for objects for which a neural network, object detection and classification, as well as basic object tracking, para 278, CNN for classifying environmental and urban sounds). Regarding claim 7, Nister fails to discloses the autonomous or semi-autonomous machine of claim 1, wherein the autonomous or semi-autonomous machine is further to: determine, based at least on at least one of the first audio data or the second audio data, a location of an object associated with the one or more sounds, wherein the one or more control operations are further performed based at least on the location of the object. In analogous art, Lockwood discloses the autonomous or semi-autonomous machine of claim 1, wherein the autonomous or semi-autonomous machine is further to: determine, based at least on at least one of the first audio data or the second audio data, a location of an object associated with the one or more sounds, wherein the one or more control operations are further performed based on the location of the object (col. 7, lines 19-25, sensors 204 include sensors configured to identify a location. The sensors 204 include light detection and ranging sensors (LIDAR), cameras (e.g. RGB-cameras, intensity (grey scale) cameras, infrared cameras, depth cameras, stereo cameras, and the like), radio detection and ranging sensors (RADAR), sound navigation and ranging (SONAR) sensors, microphones for sensing sounds). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify teaching of audio analysis for the detection of one or more emergency by autonomous vehicles disclosed by Nister to determine detection result indicating whether multi-tone time patterns at corresponding frequencies present in acoustic signal as taught by Lockwood for determining audio event using neural network based on the plurality of audio features for determining in response to determining that the target audio event is present in the audio clip for correlation between audio features in the plurality of audio and target audio event; [Lockwood, para 07]. Regarding claim 8, Nister and Lockwood fails to discloses the autonomous or semi-autonomous machine of claim 1, wherein the autonomous or semi-autonomous machine is further to: determine, based on at least one of the first audio data or the second audio data, a direction of travel of an object associated with the one or more sounds, wherein the one or more control operations are further performed based at least on the direction of travel of the object. In analogous art, Salekin discloses the autonomous or semi-autonomous machine of claim 1, wherein the autonomous or semi-autonomous machine is further to: determine, based at least on one of the first audio data or the second audio data, a direction of travel of an object associated with the one or more sounds, wherein the one or more control operations (para 18, microphones record omni directional audio, whereas video cameras generally have angular field of view, para 12, FIG. 3 operations of convolutional neural network audio tagging model of audio event detection program). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify teaching of audio analysis for the detection of one or more emergency by autonomous vehicles disclosed by Nister and Lockwood to determine detection result indicating whether multi-tone time patterns at corresponding frequencies present in acoustic signal as taught by Lockwood to use classifier to model long term dependencies and determine the boundaries in time of the target audio event within the audio clip based on the audio vector representations for improved audio surveillance. [Salekin, Abstract]. Regarding claim 9, Nister and Lockwood fails to discloses the autonomous or semi-autonomous machine of claim 1, wherein the autonomous or semi-autonomous machine is further to: detect, based at least on one or more second neural networks processing sensor data obtained using one or more perception sensors of the autonomous or semi-autonomous machine, one or more objects; and associate the one or more sounds with the one or more objects based at least on the detection of the one or more objects and the classification information associated with the one or more sounds. In analogous art, Salekin discloses wherein the autonomous or semi-autonomous machine is further to: detect, based at least on one or more second neural networks processing sensor data obtained using one or more perception sensors of the autonomous or semi-autonomous machine, one or more objects; and associate the one or more sounds with the one or more objects based at least on the detection of the one or more objects and the classification information associated with the one or more sounds (para 18, microphones can record omni directional audio, whereas video cameras generally have a limited angular field of view, para 12, FIG. 3 operations of convolutional neural network audio tagging model of audio event detection program). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify teaching of audio analysis for the detection of one or more emergency by autonomous vehicles disclosed by Nister and Lockwood to determine detection result indicating whether multi-tone time patterns at corresponding frequencies present in acoustic signal as taught by Salekin for detection of tonal components are determined and pattern model is applied to determine result indicating whether multi-tone is present in the acoustic signal. [Salekin, para 0052]. Regarding claim 10, Nister discloses a system (FIG. 1 is a block diagram of autonomous vehicle system, autonomous vehicle 102 ) comprising: one or more central processing units (CPUs); one or more graphics processing units (para 236, vehicle 102 include CPU(s) 1106, GPU(s) 1108, processor(s) 1110, accelerator(s) 1); one or more hardware accelerators (para 219, operate the steering system 1154 via one or more steering actuators 1156, to operate the propulsion system 1150 via one or more accelerators 1152); one or more perception sensors having one or more fields of view or one or more sensor fields external to a machine including the system (col. 7, lines 12-40, FIG. 2, vehicle system 202 includes a plurality of sensors 204, include, for example, inertial measurement units (IMUs), accelerometers , and gyroscopes); and one or more audio sensors at least one of internal to the machine or external to the machine (para 59, autonomous vehicle 102 include a sensor manager 108 manage or abstract sensor data from sensors of the vehicle 102, GNSS) sensor(s) 1158, RADAR sensor(s) 1160, ultrasonic sensor(s) 1162, LIDAR sensor(s) 1164, inertial measurement unit (IMU) sensor(s) 1166, microphone(s) 1196, stereo camera(s) 1168, wide-view camera(s) 1170, infrared camera(s) 1172, surround camera(s) 1174, long range and/or mid-range camera(s) 1198, and/or other sensor types (i.e. external or internal sensors), para 220, controller(s) 1136 controlling components of vehicle 102 in response to sensor data received from sensors). Nister specifically fails to discloses the system is to: determine, based at least on one or more neural networks processing first audio data obtained using the one or more audio sensors, information associated with one or more objects in an environment of the machine; based at least on the information associated with the one or more objects being determined, verify, based at least one the one or more neural networks processing second audio data obtained using the one or more audio sensors, the information associated with the one or more objects, the first audio data associated with a first time window that is offset from and at least partially overlaps with a second time window associated with the second audio data; and cause the machine to perform one or more operations based at least on the information. In analogous art, Lockwood discloses the system is to: determine, based at least on one or more neural networks processing first audio data obtained using the one or more audio sensors (Abstract, train or use a model to determine from vehicle data (i.e., audio data) , col. 25, lines 51-67, model 508 trained includes vehicle data 514 (e.g., sensor data, audio data, etc.) and model 508 implemented as a convolutional neural network, col. 18, lines 57-65, first audio captured by an external microphone, and second audio captured by an internal microphone, and present three-dimensional representation data, col. 40, lines 60-67, model 508 to determine between first and second data from first sensor data and second sensor data)), Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify teaching of audio analysis for the detection of one or more emergency by autonomous vehicles disclosed by Nister and Salekin to determine detection result indicating whether multi-tone time patterns at corresponding frequencies present in acoustic signal as taught by Lockwood to obtain assistance from a driverless vehicle and to obtain data in response to request and use train a model and use a model to determine from vehicle data option for presentation. [Lockwood, Abstract]. Nister and Lockwood specifically fails to disclose information associated with one or more objects in an environment of the machine; based at least on the information associated with the one or more objects being determined, verify, based at least one the one or more neural networks processing second audio data obtained using the one or more audio sensors, the information associated with the one or more objects, the first audio data associated with a first time window that is offset from and at least partially overlaps with a second time window associated with the second audio data; and cause the machine to perform one or more operations based at least on the information. In analogous art, Salekin discloses information associated with one or more objects in an environment of the machine (para 07, determining target audio event is present in the audio clip using a first neural network based on plurality of audio features, and determining position in time of the target audio event within the audio clip using a second neural network, para 34, each window segment S.sub.i has predetermined amount or percentage of temporal overlap with adjacent window segments and predetermined length of audio clip 102, predetermined length of each window segment, and predetermined amount or percentage of temporal overlap with adjacent window segments); based at least on the information associated with the one or more objects being determined, verify, based at least one the one or more neural networks processing second audio data obtained using the one or more audio sensors (FIG. 3 neural networks audio tagging model of audio event detection program, para 28, audio event detection model utilize neural networks, para 08, determine audio event is present in the audio clip using a neural networks based on the plurality of audio features; determine, in response to determining that audio event is present in audio clip, plurality of audio features and audio event; and determine position in time of the target audio event within audio clip using a neural networks based on the plurality of vectors, para 35, each window segment S.sub.i and has a predetermined amount or percentage of temporal overlap with adjacent sub-segments), information associated with one or more objects in an environment of the machine, the first audio data associated with a first time window that is offset from and at least partially overlaps with a second time window associated with the second audio data (FIG. 6-7, block 330, 340, classifier model of the audio event detection program, para 27, audio event detection model, which comprises audio feature extractor 32, convolution neural network (DCNN) audio tagging model 34, classifier model 38 is configured to identify the boundaries and/or positions in time of the detected target audio event in the audio clip, para 41, processor 22 is configured to execute program instructions corresponding to the DCNN audio tagging model 34 to determine a classification, para 34, total number of window segments N (e.g., 148) is a function of predetermined length (e.g., 30 seconds) of audio clip 102, first predetermined length (e.g., 500 ms) of each window segment, and predetermined amount or percentage of overlap with adjacent window segments (e.g., 300 ms or 60% overlap), para 56, HLD features distorted with distance, environmental noise (i.e offset) and reverberation); and cause the machine to perform one or more operations based at least on the information (Abstract, use classifier model and determine audio event based on the audio vector representations, para 29, audio event detection program 30 and/or the audio event detection model, statements that a software component or method step performs some process/function or is configured to perform some process/function means that a processor or controller (e.g., the processor 22) executes corresponding program instructions, para 69, plurality of LSTM cells 138 perform functions to a set parameters, which are learned and optimized during a training process and execute program instructions corresponding to classifier model 38, para 75, processor 22 determine a classification indicating whether audio event is present using a DCNN and configured to perform a sequence of dilated convolution operation, para 32, training dataset can be generated for each target audio event based on a modest number of isolated audio samples). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify teaching of audio analysis for the detection of one or more emergency by autonomous vehicles disclosed by Nister and Lockwood to determine detection result indicating whether multi-tone time patterns at corresponding frequencies present in acoustic signal as taught by Lockwood to use classifier to model long term dependencies and determine the boundaries in time of the target audio event within the audio clip based on the audio vector representations for improved audio surveillance. [Salekin, Abstract]. Regarding claim 11, Nister and Lockwood fails to discloses the system of claim 10, wherein the: system is further to: determine, based at least on the one or more neural network processing the first audio data, a temporal inference state associated with the first audio data, wherein the information associated with the one or more objects is verified based at least on the one or more neural networks processing the temporal inference state and the second audio data. In analogous art, Salekin discloses the system of claim 10, wherein the system is further to: determine, based at least on the one or more neural network processing the first audio data, a temporal inference state associated with the first audio data (para 07, determining target audio event is present in the audio clip using a first neural network based on plurality of audio features, and determining position in time of the target audio event within the audio clip using a second neural network, para 34, each window segment S.sub.i has predetermined amount or percentage of temporal overlap with adjacent window segments and predetermined length of audio clip 102, predetermined length of each window segment, and predetermined amount or percentage of temporal overlap with adjacent window segments); based at least on the information associated with the one or more objects being determined, verify, based at least one the one or more neural networks processing second audio data obtained using the one or more audio sensors (FIG. 3 neural networks audio tagging model of audio event detection program, para 28, audio event detection model utilize neural networks, para 08, determine audio event is present in the audio clip using a neural networks based on the plurality of audio features; determine, in response to determining that audio event is present in audio clip, plurality of audio features and audio event; and determine position in time of the target audio event within audio clip using a neural networks based on the plurality of vectors). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify teaching of audio analysis for the detection of one or more emergency by autonomous vehicles disclosed by Nister and Lockwood to determine detection result indicating whether multi-tone time patterns at corresponding frequencies present in acoustic signal as taught by Lockwood to use classifier to model long term dependencies and determine the boundaries in time of the target audio event within the audio clip based on the audio vector representations for improved audio surveillance. [Salekin, Abstract]. Regarding claim 12, Nister discloses the system of claim 10, wherein: the information is associated one or more sound types corresponding to the one or more objects; the system is further to determine, based at least on the one or more sound types, a type of object; and the machine is caused to perform the one or more operations based at least on the type of object (para 28, FIG. 11B, system used for depth-based object detection, especially for objects for which a neural network, object detection and classification, as well as basic object tracking, para 278, CNN for classifying environmental and urban sounds). Regarding claim 13, Nister fails to discloses the system of claim 10, wherein: the information is associated with one or more types of objects corresponding to the one or more objects; the system is further to determine, based at least on the information, a type of object of the one or more types of objects; and the machine is caused to perform the one or more operations based at least on the type of object. In analogous art, Lockwood discloses the system of claim 10, wherein: the information is associated with one or more types of objects corresponding to the one or more objects; the system is further to determine, based at least on the information, a type of object of the one or more types of objects; and the machine is caused to perform the one or more operations based at least on the type of object (Abstract, train or use a model to determine from vehicle data (i.e., audio data) , col. 25, lines 51-67, model 508 trained includes vehicle data 514 (e.g., sensor data, audio data) and model 508 implemented as convolutional neural network, col. 18, lines 57-65, first audio captured by external microphone, second audio captured by microphone). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify teaching of audio analysis for the detection of one or more emergency by autonomous vehicles disclosed by Nister to determine detection result indicating whether multi-tone time patterns at corresponding frequencies present in acoustic signal as taught by Lockwood to include surveillance system to receive audio surveillance signals from the audio input devices and to process the audio surveillance signals to detect target audio events for improved surveillance . [Lockwood, para 26]. Regarding claim 15, Nister discloses the system of claim 10, wherein the information further includes perception information corresponding to the one or more objects, the perception information determined based at least on one or more second neural networks processing sensor data obtained using the one or more perception sensors (para 63, obstacle perceiver 110 perform obstacle perception based on vehicle 102 allowed to drive, para 88, control parameters perception state based on sensor data generated by sensors of vehicle 102), para 311, perception block use information identifying objects, secondary computer use neural network trained). Regarding claim 16, Lockwood discloses the system of claim 10, wherein the system is comprised in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing deep learning operations; a system on chip (SoC); a system including a programmable vision accelerator (PVA); a system including a vison processing unit; a system implemented using an edge device; a system implemented using a robot; a system incorporating one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources (para 249, accelerator(s) 1114 (e.g., the hardware acceleration cluster) may include a programmable vision accelerator(s) (PVA), include augmented reality (AR) and/or virtual reality (VR) applications); a system including a vison processing unit; a system implemented using an edge device; a system implemented using a robot; a system incorporating one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources (para 213, visualization techniques may be used in computer games, robotics, virtual reality, physical simulations, or other technology areas, para 281, wireless connectivity over the Internet with the cloud (e.g., with server(s) 104 and/or other network devices). Regarding claim 17, Nister discloses at least one system-on-a-chip (SoC) comprising: one or more central processing units (CPUs) (para 219, Controller(s) 1136, which include one or more system on chips (SoCs)); one or more graphics processing units (GPUs) (para 219, PU(s), provide signals (e.g., representative of commands) to one or more components and/or systems of the vehicle 102); one or more hardware accelerators (para 219, operate the steering system 1154 via actuators 1156, to operate propulsion system 1150 via one or more accelerators 1152); wherein the at least one SoC is to cause a machine to perform one or more operations based at least on the one or more sounds being associated with the object (para 220, controller(s) 1136 controlling components and systems of vehicle 102 in response to sensor data received from one or more sensors (e.g., sensor inputs, para 234, each SoC 1104, each controller 1136, and/or each computer within the vehicle have access to the same input data (e.g., inputs from sensors of vehicle 102)), , Nister specifically fails to disclose detect data obtained using one or more perception sensors of a machine, first information associated with an object; determine, based at least on one or more second neural networks processing audio data obtained using one or more audio sensors of the machine, second information associated with one or more sounds; associate, based at least on the first information and the second information, the one or more sounds with the object. In analogous art, Lockwood discloses detect data obtained using one or more perception sensors of a machine, first information associated with an object (Abstract, train or use a model to determine from vehicle data (i.e., audio data) , col. 25, lines 51-67, model 508 trained includes vehicle data 514 (e.g., sensor data, audio data, etc.) and model 508 implemented as a convolutional neural network, col. 18, lines 57-65, first audio captured by an external microphone, and second audio captured by an internal microphone, and present three-dimensional representation data, col. 40, lines 60-67, model 508 to determine between first and second data from first sensor data and second sensor data)). Th Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify teaching of audio analysis for the detection of one or more emergency by autonomous vehicles disclosed by Nister to determine detection result indicating whether multi-tone time patterns at corresponding frequencies present in acoustic signal as taught by Lockwood to include surveillance system to receive audio surveillance signals from the audio input devices and to process the audio surveillance signals to detect target audio events for improved surveillance . [Lockwood, para 26]. Nister and Lockwood specifically fails to disclose determine, based at least on one or more second neural networks processing audio data obtained using one or more audio sensors of the machine, second information associated with one or more sounds; associate, based at least on the first information and the second information, the one or more sounds with the object. In analogous art, Salekin discloses determine, based at least on one or more second neural networks processing audio data obtained using one or more audio sensors of the machine, second information associated with one or more sounds (FIG. 3 neural networks audio tagging model of audio event detection program, para 28, audio event detection model utilize neural networks, para 08, determine audio event is present in the audio clip using a neural networks based on the plurality of audio features; determine, in response to determining that audio event is present in audio clip, plurality of audio features and audio event; and determine position in time of the target audio event within audio clip using a neural networks based vectors, para 35, window segment S.sub.i and has a predetermined amount or percentage of temporal overlap with adjacent sub-segments); associate, based at least on the first information and the second information, the one or more sounds with the object (FIG. 6-7, block 330, 340, classifier model of the audio event detection program, para 27, audio event detection model, which comprises audio feature extractor 32, convolution neural network (DCNN) audio tagging model 34, classifier model 38 is configured to identify boundaries and/or positions in time of the detected target audio event in the audio clip, para 27, audio extractor 32 to segment individual audio clip into a plurality of overlapping windows and represent state of audio clip in each window. DCNN audio model 34 detect presence of a target audio event and model 36 and generate vector representation of each window of the audio), Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify teaching of audio analysis for detection of emergency by autonomous vehicles disclosed by Nister and Lockwood to determine detection result indicating whether multi-tone time patterns at corresponding frequencies present in acoustic signal as taught by Salekin for detection of tonal components determined and pattern model applied to determine detection result indicating whether multi-tone is present in the acoustic signal. [Salekin, para 0052]. Regarding claim 18, Nister discloses the at least one SoC of claim 17, wherein: the first information includes a first location associated with the object within an environment (para 08, system may determine a state (e.g., location, velocity, orientation, yaw rate, etc.)); the second information includes a second location associated with the one or more sounds within the environment (para 67, localization manager 120 based on a particular location of the vehicle 102, and the localized mapping outputs may be used by the world model manager 122 to generate and/or update the world model, para 89, sensor data received from one or more sensors (e.g., of the vehicle 102) to determine locations, orientations, and velocities of objects 106 (or other actors) in the environment); and the one or more sounds are associated with the object based at least on the first location and the second location (para 272, SoC(s) 1104 include a broad range of peripheral interfaces to enable communication with peripherals, audio codecs, power management, and/or other devices, para 84, determine state of the objects 106 in the environment using any combination of the stereo, the microphone(s) 1196, claim 3, determine locations , orientations, and velocities of objects in the environment; and based at least in part on the control parameters and the locations). Regarding claim 20, Nister discloses the at least one SoC of claim 17, wherein the SoC is comprised in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing deep learning operations; a system on chip (SoC); a system including a programmable vision accelerator (PVA) (para 249, accelerator(s) 1114 (e.g., the hardware acceleration cluster) may include a programmable vision accelerator(s) (PVA), include augmented reality (AR) and/or virtual reality (VR) applications); a system including a vison processing unit; a system implemented using an edge device; a system implemented using a robot; a system incorporating one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources (para 213, visualization techniques may be used in computer games, robotics, virtual reality, physical simulations, or other technology areas, para 281, wireless connectivity over the Internet with the cloud (e.g., with server(s) 104 and/or other network devices)). Regarding claim 21, Nister discloses the at least one SoC of claim 17, wherein: the first information includes a location associated with the object within an environment (para 08, system may determine a state (e.g., location, velocity, orientation, yaw rate, etc.)); the second information includes a direction associated with the one or more sounds within the environment (para 67, localization manager 120 based on a particular location of the vehicle 102, and the localized mapping outputs may be used by the world model manager 122 to generate and/or update the world model, para 89, sensor data received from one or more sensors (e.g., of the vehicle 102) to determine locations, orientations, and velocities of objects 106 (or other actors) in the environment)); and the one or more sounds are associated with the object based at least on the location and the direction (para 84, determine state of the objects 106 in the environment using any combination of the stereo, the microphone(s) 1196, claim 3, determine locations , orientations, and velocities of objects in the environment; and based at least in part on the control parameters and the locations). 7. Claims 4-6 and 14 are rejected under 35 U.S.C. 103(a) as being unpatentable Nister (US 20190243371 A1) (hereinafter Nister) in view of Lockwood (US Patent 10268191 B1) (hereinafter Lockwood) and further in view of Salekin (US 20210005067 A1) (hereinafter Salekin) and further in view of Watkins (US 20220122620 A1) (hereinafter Watkins). Regarding claim 4, Nister and Lockwood and Salekin fails to discloses the autonomous or semi-autonomous machine of claim 1, wherein the autonomous or semi-autonomous machine is further to: determine, based at least on a first portion of the classification information that is associated with the first audio data, a sound type of the one or more sounds; and verify, based at least on a second portion of the classification information that is associated with the second audio data, the sound type, wherein the one or more control operations are performed based at least on the sound type being verified. In analogous art, Watkins discloses the autonomous or semi-autonomous machine of claim 1, wherein the autonomous or semi-autonomous machine is further to: determine, based at least on a first portion of the classification information that is associated with the first audio data, a sound type of the one or more sounds; and verify, based at least on a second portion of the classification information that is associated with the second audio data, the sound type, wherein the one or more control operations are performed based at least on the sound type being verified (para 05, using a first audio recording device coupled to an autonomous vehicle; separating, using a computing device coupled to the autonomous vehicle, the audio segment into one or more audio clips, para 06, generating the spectrogram includes performing a transformation on each of the audio clips to a lower dimensional feature representation, para 44, perform a localization analysis on the audio dada (for example, using audio data captured from a first and a second audio recording device 106) using the CNN in order to localize the vehicle generating the detected emergency siren). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify teaching of audio analysis for the detection of one or more emergency by autonomous vehicles disclosed by Nister and Lockwood and Salekin to determine detection result indicating whether multi-tone time patterns at corresponding frequencies present in signal as taught by Watkins for sing audio segment into audio clips for generating a spectrogram of the one or more audio clips, and inputting each spectrogram into a Convolutional Neural Network (CNN) for determining a course of action of autonomous vehicle. [Watkins, Abstract]. Regarding claim 6, Nister and Lockwood and Salekin fails to discloses the autonomous or semi-autonomous machine of claim 1, wherein the autonomous or semi-autonomous machine is further to: determine, based at least on a first portion of the classification information that is associated with the first audio data, a type of object associated with the one or more sounds; and verify, based at least on a second portion of the classification information that is associated with the second audio data, the type of object, wherein the one or more control operations are performed based at least on the type of object being verified. In analogous art, Watkins discloses the autonomous or semi-autonomous machine of claim 1, wherein the autonomous or semi-autonomous machine is further to: determine, based at least on a first portion of the classification information that is associated with the first audio data, a type of object associated with the one or more sounds; and verify, based at least on a second portion of the classification information that is associated with the second audio data, the type of object, wherein the one or more control operations are performed based at least on the type of object being verified (para 56, plurality of audio clips correlates to the same record time, and the audio clips may include an emergency siren, para 05, using a first audio recording device coupled to an autonomous vehicle; separating, using a computing device coupled to the autonomous vehicle, the audio segment into one or more audio clips, para 06, generating the spectrogram includes performing a transformation on each of the audio clips to a lower dimensional feature representation, para 44, perform a localization analysis on the audio dada (for example, using audio data captured from a first and a second audio recording device 106) using the CNN in order to localize the vehicle generating the detected emergency siren). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify teaching of audio analysis for the detection of one or more emergency by autonomous vehicles disclosed by Nister and Lockwood and Salekin to determine detection result indicating whether multi-tone time patterns at corresponding frequencies present in signal as taught by Watkins for sing audio segment into audio clips for generating a spectrogram of the one or more audio clips, and inputting each spectrogram into a Convolutional Neural Network (CNN) for determining a course of action of autonomous vehicle. [Watkins, Abstract]. Regarding claim 14, Nister and Lockwood and Salekin fails to discloses the system of claim 10, wherein the second audio data includes a portion of the first audio data. In analogous art, Watkins discloses the system of claim 10, wherein the second audio data includes a portion of the first audio data (para 05, using a first audio recording device coupled to an autonomous vehicle into one or more audio clips, para 06, generating the spectrogram includes performing a transformation on each of the audio clips to a lower dimensional feature representation, para 44, perform a localization analysis on the audio dada (for example, using audio data captured from a first and a second audio recording device 106) using the CNN in order to localize the vehicle generating the detected emergency siren). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify teaching of audio analysis for the detection of one or more emergency by autonomous vehicles disclosed by Nister and Lockwood and Salekin to determine detection result indicating whether multi-tone time patterns at corresponding frequencies present in signal as taught by Watkins for sing audio segment into audio clips for generating a spectrogram of the one or more audio clips, and inputting each spectrogram into a Convolutional Neural Network (CNN) for determining a course of action of autonomous vehicle. [Watkins, para 05]. Response to Arguments 8. Applicant's arguments filed July 24, 2026 have been fully considered but they are not persuasive. On page 16, lines 1-9, and page 17, lines 3-18, and lines , and page 18, lines 9-18, and page 19, lines 1-7, the applicant argues that the reference(s) do not teach or even suggest each and every limitations as claimed. The examiner respectfully disagrees and points out that the Nister teaches as in FIG. 1, f autonomous vehicle system, autonomous vehicle 102 , and autonomous vehicle 102 include a sensor manager 108 manage or abstract sensor data from sensors of the vehicle 102, GNSS) sensor(s) 1158 [059], and, controller(s) 1136 controlling components of vehicle 102 in response to sensor data received from sensors [220]. Lockwood teaches train or use a model to determine from vehicle audio data (i.e., audio data) [Abstract] and , model 508 trained includes vehicle data 514 (e.g., sensor data, audio data, etc.) and model 508 implemented as a convolutional neural network [col. 25, lines 51-67] and, audio captured by external microphone, second audio captured by microphone, present three-dimensional representation data [col. 18, lines 57-65] and, model 508 to determine between first and second audio data from first sensor data and second sensor data [col. 40, lines 60-67], and Salekin teaches system determining target audio event is present in the audio clip using a first neural network based on plurality of audio features, and determining position in time of the target audio event within the audio clip using a second neural network [07] and, each window segment S.sub.i has predetermined amount or percentage of temporal overlap with adjacent window segments and predetermined length of audio clip 102, predetermined length of each window segment, and predetermined amount or percentage of temporal overlap with adjacent window segments [034] and as in FIG. 3 neural networks audio tagging model of audio event detection program and audio event detection model utilize neural networks [028] and, determine audio event is present in the audio clip using a neural networks based on the plurality of audio features; determine, in response to determining that audio event is present in audio clip, plurality of audio features and audio event; and determine position in time of the target audio event within audio clip using a neural networks based on the plurality of vectors [028[ and each window segment S.sub.i and has a predetermined amount or percentage of temporal overlap with adjacent sub-segments [035] and as in FIG. 6-7, block 330, 340, classifier model of audio event detection program, and audio event detection model, comprises audio feature extractor 32, convolution neural network (DCNN) audio tagging model 34, classifier model 38 is configured to identify boundaries or positions in time of detected target audio event in audio clip [027] and, processor 22 execute program instructions to DCNN audio tagging model 34 to determine a classification [041] and, total number of window segments N (e.g., 148) is a function of predetermined length (e.g., 30 seconds) of audio clip 102, first predetermined length (e.g., 500 ms) of each window segment, and predetermined amount or percentage of overlap with adjacent window segments (e.g., 300 ms or 60% partially overlap) [034] and processor 22 determine a classification indicating whether audio event is present using a DCNN and perform a sequence of dilated convolution operation [075] and, training dataset generated for each target audio event based on modest number of audio samples [032]. On page 23, lines 1-9, the applicant argues that the reference(s) do not teach or even suggest the applicant argues that the reference(s) do not teach or even suggest each and every limitations as claimed in claims 4-6 and 14. The examiner respectfully disagrees and points out that the Watkins teaches system using a first audio recording device coupled to an autonomous vehicle; separating, using a computing device coupled to the autonomous vehicle, the audio segment into one or more audio clips and , and generating the spectrogram includes performing a transformation on each of the audio clips to a lower dimensional feature representation [05-06] and , system perform a localization analysis on the audio dada (for example, using audio data captured from a first and a second audio recording device 106) using the CNN in order to localize the vehicle generating the detected emergency siren [044]. Thus, Nister (US 20190243371 A1) and Lockwood (US Patent 10268191 B1) and Salekin (US 20210005067 A1) and Watkins (US 20220122620 A1) disclose the applicant’s whole invention. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to MIRZA ALAM whose telephone number is (469) 295-9286. The examiner can normally be reached on 8:00AM-5:00PM. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Steven Lim can be reached on 571-270-1210. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000./ /MIRZA F ALAM/ Primary Examiner, Art Unit 2688
Read full office action

Prosecution Timeline

Jan 29, 2025
Application Filed
May 14, 2026
Non-Final Rejection mailed — §103
Jul 23, 2026
Examiner Interview Summary
Jul 23, 2026
Applicant Interview (Telephonic)
Jul 24, 2026
Response Filed
Sep 11, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12736442
BENCHMARK TEST POINT WITH INTEGRATED RFID TECHNOLOGY FOR A PARTICLE DETECTOR
2y 2m to grant Granted Sep 15, 2026
Patent 12730180
RANGING METHOD AND APPARATUS
2y 9m to grant Granted Sep 08, 2026
Patent 12725512
SYSTEMS, METHODS, AND PROCESSES OF PROVIDING FIRST RESPONDERS WITH LIVE CONTEXTUAL INFORMATION OF BUILDING DESTINATIONS FOR ENHANCED PUBLIC SAFETY OPERATIONS
2y 1m to grant Granted Sep 01, 2026
Patent 12726807
Systems and Methods for Users to Avoid Active Danger and Get to a Safety Zone
1y 8m to grant Granted Sep 01, 2026
Patent 12720467
PRIORITIZATION AND PERFORMANCE OF OVERLAPPING POSITIONING METHOD REQUESTS
2y 7m to grant Granted Aug 25, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
74%
Grant Probability
99%
With Interview (+33.9%)
2y 4m (~8m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 1032 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month