DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 8, 11, 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Narayan et al. (US 20230368544 A1).
Regarding claims 1 and 12, Narayan et al. disclose an apparatus configured to perform a perception task the apparatus comprising: a memory; and processing circuitry connected to the memory (perception software stack advantageously performs the method, [0045], The computer device 104 includes the perception software stack 108 or communicates with a computer readable medium having the perception software stack 108 stored thereon. The computer device includes a processor 110 and a memory 112. The perception software stack 108 is preferably implemented as computer-readable instructions stored on the memory 112 and executable by the processor 110, [0102]), the processing circuitry configured to; and method for performing a perception task (perception software stack advantageously performs the method, [0045]), the method comprising: generate sensor features from data from one or more sensors (a first layer configured to pre-process an original image received from a camera or sensor, [0045]); process the sensor features with a time-continuous recurrent neural network (RNN) to produce time-continuous features (The neural network in the third layer has a continuous time recurrent neural network (CTRNN) architecture that is preferably built to be low resolution, in the form of a low resolution recurrent active vision neural network (LRRAVNN), [0047], The neural network 400 has a continuous time recurrent neural network (CTRNN) architecture. The CTRNN architecture has an input layer 202, a hidden recurrent layer 204 and an output layer 206, [0122]); and perform the perception task using the time-continuous features (The first layer and second layer, which prepare the original image for the neural network, and the fourth layer, which post-processes the output value of the neural network, help to make the perception software stack more accurate and reliable at identifying a feature in the environment. As with the method above, the neural network and the perception software stack as a whole can be modified to perform different tasks, such as image classification, image segmentation and object detection, [0047], perception software stack 108, [0101], [0122]).
Narayan et al. do not use the language “to produce time-continuous features”. It would have been obvious at the time of filing to one of ordinary skill in the art the CTRNN would produce time-continuous features as this uses a continuous time recurrent neural network ([0047]) and the output is object detection ([0047]) thereby making it obvious the output produced would be time-continuous features.
Regarding claim 8, Narayan et al. disclose the apparatus of claim 1. Narayan et al. further indicate the perception task includes one or more of semantic segmentation, semantic occupancy prediction, lane tracking, or 3D object detection (Preferably, the neural network stored in the memory is configured to perform one or more specific tasks including image classification, object detection and road segmentation, [0011], [0023], [0047], image captured by a camera or LIDAR sensor, The neural network may classify the image, segment the image, to identify road for example, or detect objects such as cars and pedestrians in the image, [0063]).
Regarding claim 11, Narayan et al. disclose the apparatus of claim 1. Narayan et al. further indicate the processing circuitry is part of an advanced driver assistance system (ADAS), and wherein the ADAS is configured to control a vehicle at least in part based on an output of the perception task (“This invention relates to a device and system for autonomous vehicle control, and in particular, a device for identifying a feature in the environment of a vehicle and a system to control the vehicle based on the identified features of the environment. A key component of an autonomous vehicle/assistance system is sensory perception. Sensory perception is the term that describes the capability to process input data from sensors and hardware, to obtain meaningful and useful results that can then be used to inform control of a certain system. Autonomous vehicles conventionally combine a variety of sensors to perceive their surroundings, including radar, Lidar, sonar, Global Positioning System (GPS), cameras and inertial measurement units. The data from these sensors is fed into advanced control systems to identify appropriate navigation paths, obstacles, hazards and relevant signage. The advanced control systems may process several functions at once, and need to do so in a fast and efficient manner in order to deal with real-world problems whilst the autonomous vehicle is in transit. For example, the control systems must be able to extract which areas of an image plane, obtained from a camera, correspond to road (freespace), non-road (non-drivable areas), or obstacles such as cars, bicycles and pedestrians. This function must be performed continuously and accurately if the autonomous vehicle is to perform in an effective and safe manner”, [0001]-[0003], According to a second aspect of the invention, a vehicle control system for fitting in or on a vehicle is provided. The system comprises a sensor or camera; a control computer; and the computer device of the first aspect of the invention, wherein the computer device is configured to receive sensor data or an original image from the sensor or camera, and output an output value to the control computer based on the sensor data or original image; the control computer being configured to control one or more components of a vehicle based on the output value received from the computer device. Preferably the control computer is configured to autonomously control the vehicle. [0014]-[0015], The method may further include controlling the speed and/or direction of the vehicle based on the identified feature, [0043], The specific tasks are thus computer vision tasks, the results of which are used to inform control of an autonomous vehicle, [0101]).
Claim(s) 5, 6, 16 and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Narayan et al. (US 20230368544 A1) as applied to claim 1 above, further in view of Karasev et al. (US 20240404256 A1).
Regarding claims 5 and 16, Narayan et al. disclose the apparatus of claim 1. Narayan et al. do not explicitly disclose the sensor features are birds-eye-view (BEV) sensor features, and wherein to generate the sensor features from the data from the one or more sensors, the processing circuitry is configured to: generate respective sensor features from the one or more sensors; and generate, using the respective sensor features, a BEV representation having the BEV sensor features.
Karasev et al. teach the sensor features are birds-eye-view (BEV) sensor features, and wherein to generate the sensor features from the data from the one or more sensors, the processing circuitry is configured to: generate respective sensor features from the one or more sensors; and generate, using the respective sensor features, a BEV representation having the BEV sensor features (FIG. 5 depicts a block diagram of an example computing device capable of implementing camera-radar data fusion to generate a fused bird's-eye view (BEV) grid for efficient object detection, [0009], Camera images and radar data, however, are usually acquired in different frames of reference (e.g., the perspective view for camera images and spherical or cylindrical system of coordinates for radar data). As disclosed herein, efficient processing of combined camera and radar data (and lidar data, in L4 systems) can be achieved by using the top-down view, also known as the bird's-eye view (BEV), in which objects are represented on a convenient manifold, e.g., a plane viewed from above and characterized by a simple set of Cartesian coordinates, [0020], Disclosed herein is an end-to-end perception model (EEPM), which can include a neural network architecture, can be trained to make predictions related to AV perception, [0024], deploys a recurrent neural network (RNN) to smooth and interpolate locations and velocities over time, [0046], FIG. 2A is a diagram illustrating example network architecture of an end-to-end perception model (EEPM) 132 that can be deployed as part of a perception system of a vehicle, in accordance with some implementations of the present disclosure. Input data 201 can include data obtained by various components of the sensing system 110 (as depicted in FIG. 1), e.g., lidar(s) 112, radar(s) 114, optical (e.g., visible) range camera(s) 118, IR sensors(s) 119. For example, as shown, the input data 201 can include camera data 210 and radar data 220. Although not shown, the input data 201 can further include, e.g., lidar data., [0051]).
Narayan et al. and Karasev et al. are in the same art of RNNs (Narayan et al., [0047]; Karasev et al., [0046]). The combination of Karasev et al. with Narayan et al. will enable creating a BEV representation. It would have been obvious at the time of filing to one of ordinary skill in the art to combine the BEV of Karasev et al. with the invention of Narayan et al. as this was known at the time of filing, the combination would have predictable results, and Karasev et al. indicate “Camera images and radar data, however, are usually acquired in different frames of reference (e.g., the perspective view for camera images and spherical or cylindrical system of coordinates for radar data). As disclosed herein, efficient processing of combined camera and radar data (and lidar data, in L4 systems) can be achieved by using the top-down view, also known as the bird's-eye view (BEV), in which objects are represented on a convenient manifold, e.g., a plane viewed from above and characterized by a simple set of Cartesian coordinates. Object identification and tracking can subsequently be performed directly within the BEV representation. Success of such techniques depends on accurate mapping of the objects to the BEV representation” ([0020]) providing an accuracy benefit to combining inventions that will also improve safety in the autonomous driving applications of Narayan et al..
Regarding claims 6 and 17, Narayan et al. and Karasev et al. disclose the apparatus of claim 5. Karasev et al. further indicate to process the sensor features with the time-continuous RNN to produce time-continuous features, the processing circuitry is configured to: receive current BEV sensor features at a current time; receive previous BEV sensor features from a previous time; warp the previous BEV sensor features to a pose of the current BEV sensor features to create warped BEV sensor features; combine the warped BEV sensor features and the current BEV sensor features to form combined BEV sensor features; and process the combined BEV sensor features with the time-continuous RNN to form the time-continuous features (The output of EEPM 132 can be used for tracking of detected objects. In some implementations, tracking can be reactive and can include history of poses (positions and orientations) and velocities of the tracked objects, deploys a recurrent neural network (RNN) to smooth and interpolate locations and velocities over time, [0046], The set of BEV grids defining the multi-BEV space can be recurrent, e.g., some proportion of the features obtained at time t.sub.1 can be warped (using a differentiable warp such as a spatial transformer) and aggregated into new grids at time t.sub.2 obtained together with the new features from time step t.sub.2, e.g., using the smooth pose delta (i.e., pose change between time t.sub.1 and time t.sub.2). The multi-scale BEV space can be in a smooth pose consistent frame. The multi-scale BEV space can be spatially consistent for a period of time used for the aggregation in detection. In some implementations, a process for clearing distant portions of the grid and shifting values over as the AV moves through the world. Various priors in the global frame (e.g., elevation tiles, road graph) may undergo an accurate global-to-smooth transform. Dynamic objects may be represented using a flow field in combination with an occupancy map to perform additional recurrent aggregation, [0064]).
Claim(s) 7 and 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Narayan et al. (US 20230368544 A1) and Karasev et al. (US 20240404256 A1) as applied to claim 6 above, further in view of Malloch (US 12579802 B1).
Regarding claims 7 and 18, Narayan et al. and Karasev et al. disclose the apparatus of claim 6. Narayan et al. and Karasev et al. do not explicitly disclose to perform the perception task using the time-continuous features, the processing circuitry is configured to: process the time-continuous features and the current BEV features using a transformer decoder.
Malloch teaches process the time-continuous features and the current BEV features using a transformer decoder (Another example of an adaptive model for the camera attention module 304 to learn when to use images from each camera includes a temporal aggregation model in which heuristic rules such as those discussed above are mixed with features from previous camera timestamps, col. 8, lines 55-60, A multi-layer perception network can transform the 3D coordinates to 3D position embedded data, “At step 418, a birds eye view image of the multi-camera image data is generated. In particular, the birds eye view image is based on the 3D position-aware features F.sup.3d. In some examples, the 3D position-aware features F.sup.3d can be input to a decoder 314, which can be a transformer decoder, and the 3D position-aware features F.sup.3d can be flattened to generate the birds eye view image. In some examples, objects in the images can be detected based on the 3D position-aware features F.sup.3d and object queries. A decoder can predict 3D bounding boxes for detected objects, and the decoder can determine object classes for the detected objects”, col. 11, lines 20-37).
Narayan et al. and Malloch are in the same art of perception networks carried out by neural networks (Narayan et al., [0045], [0047]; Malloch, col. 1, lines 50-55, col. 5, lines 50-55). The combination of Malloch with Narayan et al. and Karasev et al. will enable processing the time-continuous features and the current BEV features using a transformer decoder. It would have been obvious at the time of filing to one of ordinary skill in the art to combine the transformer decoder of Malloch with the invention of Narayan et al. and Karasev et al. as this was known at the time of filing, the combination would have predictable results, and Malloch indicate, “In general, autonomous vehicles have a limited amount of computational power and resources. In some examples, the computer vision stack of an autonomous vehicle can use more than half of the computational power of the vehicle. Thus, decreasing the amount of processing at the computer vision stack can result in significant saving of resources such as computational power. For example, the vehicle may consume less power from doing fewer computations. In some examples, the vehicle cost can be lowered since less hardware is provisioned. To save resources and increase vehicle efficiency, a subset of the vehicle cameras can be turned off during certain situations to save computational resources. However, this can result in a blind spot in a selected area. Thus, systems and methods are described herein for downscaling the computational resources dedicated to selected cameras in various situations to save on compute resources while continuing imaging of the vehicle surroundings and preventing blind spots” (col. 2, line 60 – col. 3, line 15) demonstrating a computational benefit to combining inventions.
Claim(s) 10 and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Narayan et al. (US 20230368544 A1) as applied to claim 1 above, further in view of Wang et al. (US 20250278624 A1).
Regarding claims 10 and 20, Narayan et al. disclose the apparatus of claim 1, wherein the one or more sensors include one or more camera sensors, one or more sonar sensors, one or more radar sensors, or one or more LiDAR sensors ([0063], [0104]). Narayan et al. do not explicitly disclose to generate the sensor features from the data from the one or more sensors, the processing circuitry is configured to: receive the data from the one or more sensors at asynchronous observation times; and generate the sensor features from the data from the one or more sensors at each of the asynchronous observation times.
Wang et al. teach the one or more sensors include one or more camera sensors, one or more sonar sensors, one or more radar sensors, or one or more LiDAR sensors (For example, for effective control of autonomous vehicles, image data from different sensors such as LiDARs and RGB cameras is utilized to identify objects in the field of view of the vehicle, [0042]), and wherein to generate the sensor features from the data from the one or more sensors, the processing circuitry is configured to: receive the data from the one or more sensors at asynchronous observation times; and generate the sensor features from the data from the one or more sensors at each of the asynchronous observation times (The shared continuous-time axis forces the neural decoder of each such autoencoder to generate virtual latent dynamic states at the same time instances even though the time series data input to each such autoencoder may be asynchronous with the time series data input to other autoencoders, [0012], Some embodiments also provide a fusion framework utilizing such autoencoders with neural ODEs for generating time synchronized latent dynamic states of multiple time asynchronous input streams and a post-ODE latent fusion block for combining the generated latent states for further processing, [0016], The coordinate estimation loss is backpropagated through the whole network, from the shared coordinate decoder back to the ODE encoders for all the input sample streams to achieve the end-to-end training. According to some embodiments, the loss function may be derived based on the evidence lower bound (ELBO) principle, which is a weighted sum of waveform reconstruction losses in asynchronous time instances, coordinate estimation errors in these synchronized time instances, and Kullback-Leibler (KL) divergence loss term that regularizes the distribution, [0017], Accordingly, some embodiments provide a fusion framework utilizing such autoencoders with neural ODEs for generating time synchronized latent dynamic states of multiple time asynchronous input streams and a post-ODE latent fusion block for combining the generated latent states for further processing tasks such as tracking the state trajectory of a device., [0044], FIG. 1A illustrates a block diagram of an artificial intelligence (AI) system 100 for tracking a state of a device with continuous-time latent dynamics, according to some embodiments., Each neural ODE subnetwork 101 comprises an encoder 107 along with a latent subnetwork 109 and may be implemented as a recurrent neural network (RNN) architecture. The neural ODE subnetwork 101 transforms each stream of the input data 102 from its input state space to a latent space. The latent states for each input stream are time synchronized with the latent states of other input streams in the latent space., [0045], The input stream 130 has non-periodic data captures at time instances t, 2t, 3.75t, and 4.95t while the input stream 135 has non-periodic data captures at time instances 0.75t′ and 2.25t′ where t′ is greater than t. The two streams 130 and 135 differ in terms of frame rates as well as timing of capture and are therefore asynchronous in time., [0052]).
Narayan et al. and Wang et al. are in the same art of RNNs (Narayan et al., [0047]; Wang et al., [0045]). The combination of Wang et al. with Narayan et al. will enable receiving the data from the one or more sensors at asynchronous observation times. It would have been obvious at the time of filing to one of ordinary skill in the art to combine the receiving of Wang et al. with the invention of Narayan et al. as this was known at the time of filing, the combination would have predictable results, and Wang et al. indicate, “Such a transformation is advantageous in many technical fields including location tracking, anomaly detection, smart grid applications, and data completeness applications to name a few” ([0002]) “Accordingly, improved techniques and processing architectures are required for achieving dynamic, efficient and robust state space transformation for time-series data of a dynamic system” ([0004]) and “Particularly, for robustness and better accuracy, several real-world applications require fusion of data from a plurality of sources. However, meaningful fusion of such data is constrained due to timing disparity between data of two types or from two different sources” providing a real world benefit by allowing data from disparate sources to be optimally combined.
Allowable Subject Matter
Claims 2-4, 9, 13-15, and 19 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The following art is cited as relating to one of the objected to claims but not sufficient alone or in combination to disclose the limitations when incorporated into the base claims:
US 8411141 B2: “That is, the video processing apparatus obtains information describing a feature amount of the image at a steady state, obtains information describing a feature amount of a video frame, which is imaged by the imaging apparatus, obtains information describing an amount of displacement therebetween based on the information describing the feature amount of the video frame and the feature amount of the image at the steady state, delimits the video duration and determines the length of the video duration based on changes in the amount of displacement of the video frames in chronological order, obtains the information describing the average of the amounts of displacement of the video frames included in each video duration as information describing the amount of displacement of the video duration, and records the information describing the amount of displacement of the obtained video duration and images of the video duration correspondingly in recording means. Thus, since the video duration is delimited based on the change in amount of displacement of a video frame in chronological order, the video duration can be delimited for each set of video frames having closer amounts of displacement. Furthermore, the information describing the amount of displacement of each video duration and images of the video duration are recorded correspondingly, which is useful for processing of playing or searching based on the amount of displacement. In this case, the information describing the feature amount of an image at a steady state may be an average value among multiple video frames”, col. 6, lines 1-35.
US 20070213786 A1 An autonomous dynamical system is described by ordinary differential equations {dot over (x)}=f(x,d) where the vector x.epsilon.R.sup.m defines the dynamical variables and d is a scalar parameter available for an external adjustment, such as the desired STLmax. We envision a scalar variable y(t)=g(x(t)) that is a function of dynamical variables x(t) that can be measured as a system output. Assuming that at d=d.sub.0 the system has an unstable fixed point x* that satisfies f(x*, d.sub.0)=0, if the steady state value y*=g(x*) of the observable corresponding to the fixed point were known, we could stabilize the system by using a standard proportional feedback control, i.e., adjusting the control parameter by the law d=d.sub.0-k(y-y*). However, supposing that the reference value y* is unknown, our aim is to construct a reference-free feedback perturbation that automatically locates and stabilizes the fixed point. Such a perturbation should vanish when the system settles on the fixed point. Recurrent neural networks (RNN) and time lagged feedforward neural networks (TLFNs) are capable of learning and reproducing chaotic dynamics in a variety of realistic and synthetic nonlinear systems.
US 20210034949 A1 Accordingly, RNN/GRU 304 may predict a failure of an information handling resource before it actually occurs. As explained in greater detail below, RNN/GRU 304 may be unable to handle any uneven time gaps in the sample or the time series of its training data, thus imputing missing data from the training data in order to perform training and prediction.
US 20200265307 A1 The target task may correspond to a task that is to be performed by the neural network apparatus using a learned neural network. For example, learned neural network may be considered a multi-task trained neural network (also referred to as a neural network for a plurality of tasks), where the target task may be one of a plurality of tasks the neural network had been trained to implement at different discontinuous times and/or sequentially in time, as a non-limiting example.
“CARRNN: A Continuous Autoregressive Recurrent Neural Network for Deep Representation Learning From Sporadic Temporal Data” Learning temporal patterns from multivariate longitudinal data is challenging especially in cases when data is sporadic, as often seen in, e.g., healthcare applications where the data can suffer from irregularity and asynchronicity as the time between consecutive data points can vary across features and samples, hindering the application of existing deep learning models that are constructed for complete, evenly spaced data with fixed sequence lengths. In this article, a novel deep learning-based model is developed for modeling multiple temporal features in sporadic data using an integrated deep learning architecture based on a recurrent neural network (RNN) unit and a continuous-time autoregressive (CAR) model. The proposed model, called CARRNN, uses a generalized discrete-time autoregressive (AR) model that is trainable end-to-end using neural networks modulated by time lags to describe the changes caused by the irregularity and asynchronicity. It is applied to time-series regression and classification tasks for Alzheimer’s disease progression modeling, intensive care unit (ICU) mortality rate prediction, human activity recognition, and event-based digit recognition, where the proposed model based on a gated recurrent unit (GRU) in all cases achieves significantly better predictive performance than the state-of-the-art methods using RNNs, GRUs, and long short-term memory (LSTM) networks.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHELLE M ENTEZARI HAUSMANN whose telephone number is (571)270-5084. The examiner can normally be reached 10-7 M-F.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vincent M Rudolph can be reached at (571) 272-8243. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MICHELLE M ENTEZARI HAUSMANN/Primary Examiner, Art Unit 2671