DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Summary
The Amendment filed on 5 May 2026 has been acknowledged.
Claims 1, 9, 16 – 17 and 19 are amended.
Currently, claims 1 – 20 are pending and considered as set forth.
Response to Arguments
Applicant's arguments filed on 5 May 2026 have been fully considered but they are not persuasive.
The Applicant first argues, “Applicant respectfully submits that Shashua does not teach or suggest, at least, "a plurality of sensors of two or more sensor modalities; ... "apply[ing], to an end-to-end (E2E) neural network, sensor data obtained using the plurality of sensors; [and] directly comput[ing], by the E2E neural network and based at least on the E2E neural network processing the sensor data, data representative of one or more trajectory points in three-dimensional (3D) world space," as amended claim 1 recites. In the rejection of previously presented independent claim 1, the Office cites Shashua as allegedly teaching, "a plurality of sensors of two or more sensor modalities; ... "apply, to an end- to-end (E2E) neural network, sensor data obtained using the plurality of sensors; [and] directly compute, based at least on the E2E neural network processing the sensor data, data representative of one or more trajectory points in three-dimensional (3D) world space." Office Action, pp. 4-7. … As shown, Shashua describes a bottom-up lane detection module that detects lane marks within images and an end-to-end deep neural network that detects paths in images. Id., paras. [0423] and [0569]. However, Shashau does not teach or suggest that sensor data obtained using "sensors of two or more sensor modalities" are applied to either the bottom-up lane detection module or the end-to-end deep neural network. For instance, Shashua merely describes both the bottom-up lane detection module and the end-to-end deep neural network generate information for images, a single sensor modality.”
The Examiner respectfully traverses that as, considering the reference as a whole, paragraph 576 teaches how the camera use radar or laser while detecting landmark which is directly involving end-to-end deep network as updated rejection below. Therefore, Shashua does teaches two or more of sensor modalities applied to either the bottom-up lane detection module or the end-to-end deep neural network.
The Applicant further argues, “Shashua does not teach or suggest that either the bottom-up lane detection module or the end-to-end deep neural network directly generate "data representative of one or more trajectory points in three-dimensional (3D) world space." Rather, Shashua describes that both the bottom-up lane detection module and the end-to-end deep neural network detect the road model in an image coordinate frame which is then transformed to a three-dimensional space. As such, Shashua does not teach or suggest that either the bottom-up lane detection module or the end-to-end deep neural network "directly" detects the road model in the three-dimensional space using end-to-end processing. … As shown, Shashua describes that a sparse map may represent a plurality of trajectories for guiding autonomous driving or navigation along a road segment. Id., paras. [0380] and [0485]. Shashua then describes that a reconstructed trajectory may be modified based on sensed data. Id. However, Applicant initially notes that these portions of Shashua do not describe and/or relate to paragraphs [0423] and [0569] cited above. For instance, these portion of Shashua do not teach or suggest using the bottom-up lane detection module or the end-to-end deep neural network to determine the trajectory and/or reconstruct the trajectory.”
While the Examiner respectfully traverses that while paragraph 380 and 485 focuses on trajectory aspect, considering the reference as a whole, paragraph 421 – 423 teaches using bottom up lane detection module and end-to-end deep network to determine trajectory or correct short range path.
The Applicant further argues, “hese portions of Shashua do not teach or suggest applying sensor data from a "plurality of sensors of two or more sensor modalities" to a "neural network" to determine the trajectory and/or reconstruct the trajectory, let alone an "end-to-end (E2E) neural network." For similar reasons, these portions of Shashua do not teach or suggest that a "neural network" directly computes "data representative of one or more trajectory points in three-dimensional (3D) world space." For instance, these portions of Shashua do not teach or suggest any type of neural network processing. In other words, Shashua does not teach or suggest applying sensor data obtained using "a plurality of sensors of two or more sensor modalities" to a "neural network" and/or that a "neural network" directly computes "data representative of one or more trajectory points in three- dimensional (3D) world space," let alone an "end-to-end (E2E) neural network." Additionally, while Shashua does describe other types of sensor data, such as RADAR data, Shashua merely describes processing the RADAR data with images to determine features within an environment. However, Shashua does not teach or suggest processing the RADAR data with the images using a model that generates "data representative of one or more trajectory points in three-dimensional (3D) world space." For instance, Shashua does not teach or suggest a trajectory model that processes the RADAR data with the images. Consequently, Shashua does not teach or suggest "a plurality of sensors of two or more sensor modalities; ... "apply[ing], to an end-to-end (E2E) neural network, sensor data obtained using the plurality of sensors; [and] directly comput[ing], by the E2E neural network and based at least on the E2E neural network processing the sensor data, data representative of one or more trajectory points in three-dimensional (3D) world space," as amended claim 1 recites.”
The Examiner respectfully traverses that similar argument has been responded above and reflected in the updated rejection below.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1 – 20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Shashua et al. (Hereinafter Shashua) (US 2017/0010618 A1).
As per claim 1, Shashua teaches the limitations of:
an autonomous or semi-autonomous machine (See at least paragraph 6; systems and methods for autonomous vehicle navigation) comprising:
a plurality of sensors of two or more sensor modalities (See at least paragraph 511; navigation system 1700 may include one or more sensors, such as camera 122, GPS unit 1710, road profile sensor 1730, speed sensor 1720, and accelerometer 1725. Vehicle 1205 may include other sensors, such as radar sensors. The sensors included in vehicle 1205 may collect data related to road segment 1200 as vehicle 1205 travels along road segment 1200.);
one or more controllers (See at least paragraph 115; the control system may include at least one of a steering control, an acceleration control, and a braking control.);
one or more actuation components (See at least paragraph 138; FIG. 2F, vehicle 200 may include throttling system 220, braking system 230, and steering system 240.); and
one or more processors, the one or more processors comprising processing circuitry (See at least abstract and paragraph 266; navigation system for a vehicle may include at least one processor. The at least one processor may be programmed to determine a navigational maneuver for the vehicle based, at least in part, on a comparison of a motion of the vehicle with respect to a predetermined model representative of a road segment. The at least one processor may be further programmed to receive, from a camera, at least one image representative of an environment of the vehicle. The at least one processor may be further programmed to determine, based on analysis of the at least one image, an existence in the environment of the vehicle of an navigational adjustment condition, cause the vehicle to adjust the navigational maneuver based on the existence of the navigational adjustment condition, and store information relating to the navigational adjustment condition. … Both applications processor 180 and image processor 190 may include various types of processing devices. For example, either or both of applications processor 180 and image processor 190 may include a microprocessor, preprocessors (such as an image preprocessor), graphics processors, a central processing unit (CPU), support circuits, digital signal processors, integrated circuits, memory, or any other types of devices suitable for running applications and for image processing and analysis) to:
apply, to an end-to-end (E2E) neural network, sensor data obtained using the plurality of sensors (See at least paragraph 423, 506, 569 and 576; This module may look for edges in the image and assembles them together to form the lane marks. A second module may be used together with the bottom-up lane detection module. The second module is an end-to-end deep neural network, which may be trained to predict the correct short range path from an input image. … a database augmenting that of the first scheme with precise vehicle position, vehicle orientation and image pixel depth using Light Detection And Ranging (LIDAR) measurements may be used to accurately match world positions in different drives. … The system may take into account both the local shape of the lamppost and the arrangement of the lamppost in the scene: lampposts are typically at the side of the road (or on the divider), lampposts often appear more than once in a single image and at different sizes, Lampposts on highways may have fixed spacing based on country standards (e.g., around 25 m to 50 m spacing). The disclosed systems may use a convolutional neural network algorithm to classify a constant strip from the image (e.g., 136×72 pixels) that may be sufficient to catch almost all the street poles. The network may not contain any affine layers, and may only be composed of convolution layers, Max Pooling vertical layers and ReLu layers. The network's output dimension may be 3 times of the strip width, these three channels may have 3 degrees of freedom for each column in the strip. The first degree of freedom may indicate whether there is a street pole in this column, the second degree of freedom may indicate this pole's top, and the third degree of freedom may indicate its bottom. With the network's output results, the system may take all the local maximums that are above a threshold, and built rectangles bounding the poles. … the camera for detecting landmarks may be augmented any distance measurement apparatus, such as laser or radar);
directly compute, by the E2E neural network and based on at least on the E2E neural network processing the sensor data, data representative of one or more trajectory points in three-dimensional (3D) world space (See at least paragraph 380, 421 – 423, 485 and 576; sparse map 800 may include representations of a plurality of target trajectories 810 for guiding autonomous driving or navigation along a road segment. Such target trajectories may be stored as three-dimensional splines. The target trajectories stored in sparse map 800 may be determined based on two or more reconstructed trajectories of prior traversals of vehicles along a particular road segment. A road segment may be associated with a single target trajectory or multiple target trajectories. For example, on a two lane road, a first target trajectory may be stored to represent an intended path of travel along the road in a first direction, and a second target trajectory may be stored to represent an intended path of travel along the road in another direction (e.g., opposite to the first direction). Additional target trajectories may be stored with respect to a particular road segment. For example, on a multi-lane road one or more target trajectories may be stored representing intended paths of travel for vehicles in one or more lanes associated with the multi-lane road. In some embodiments, each lane of a multi-lane road may be associated with its own target trajectory. In other embodiments, there may be fewer target trajectories stored than lanes present on a multi-lane road. In such cases, a vehicle navigating the multi-lane road may use any of the stored target trajectories to guides its navigation by taking into account an amount of lane offset from a lane for which a target trajectory is stored (e.g., if a vehicle is traveling in the left most lane of a three lane highway, and a target trajectory is stored only for the middle lane of the highway, the vehicle may navigate using the target trajectory of the middle lane by accounting for the amount of lane offset between the middle lane and the left-most lane when generating navigational instructions). … The disclosed systems and methods may enable autonomous vehicle navigation (e.g., steering control) with low footprint models, which may be collected by the autonomous vehicles themselves without the aid of expensive surveying equipment. To support the autonomous navigation (e.g., steering applications), the road model may include the geometry of the road, its lane structure, and landmarks that may be used to determine the location or position of vehicles along a trajectory included in the model. Generation of the road model may be performed by a remote server that communicates with vehicles travelling on the road and that receives data from the vehicles. The data may include sensed data, trajectories reconstructed based on the sensed data, and/or recommended trajectories that may represent modified reconstructed trajectories. The server may transmit the model back to the vehicles or other vehicles that later travel on the road to aid in autonomous navigation. The geometry of a reconstructed trajectory (and also a target trajectory) along a road segment may be represented by a curve in three dimensional space, which may be a spline connecting three dimensional polynomials. The reconstructed trajectory curve may be determined from analysis of a video stream or a plurality of images captured by a camera installed on the vehicle. In some embodiments, a location is identified in each frame or image that is a few meters ahead of the current position of the vehicle. This location is where the vehicle is expected to travel to in a predetermined time period. This operation may be repeated frame by frame, and at the same time, the vehicle may compute the camera's ego motion (rotation and translation). At each frame or image, a short range model for the desired path is generated by the vehicle in a reference frame that is attached to the camera. The short range models may be stitched together to obtain a three dimensional model of the road in some coordinate frame, which may be an arbitrary or predetermined coordinate frame. The three dimensional model of the road may then be fitted by a spline, which may include or connect one or more polynomials of suitable orders. To conclude the short range road model at each frame, one or more detection modules may be used. For example, a bottom-up lane detection module may be used. The bottom-up lane detection module may be useful when lane marks are drawn on the road. This module may look for edges in the image and assembles them together to form the lane marks. A second module may be used together with the bottom-up lane detection module. The second module is an end-to-end deep neural network, which may be trained to predict the correct short range path from an input image. In both modules, the road model may be detected in the image coordinate frame and transformed to a three dimensional space that may be virtually attached to the camera. … The principle underlying the maps generation is the integration of ego motion. The vehicles sense the motion of the camera in space (3D translation and 3D rotation). The vehicles or the server may reconstruct the trajectory of the vehicle by integration of ego motion over time, and this integrated path may be used as a model for the road geometry. This process may be combined with sensing of close range lane marks, and then the reconstructed route may reflect the path that a vehicle should follow, and not the particular path that it did follow. In other words, the reconstructed route or trajectory may be modified based on the sensed data relating to close range lane marks, and the modified reconstructed trajectory may be used as a recommended trajectory or target trajectory, which may be saved in the road model or sparse map for use by other vehicles navigating the same road segment.);
determine, using the one or more controllers, one or more controls to control the autonomous or semi-autonomous machine according to the one or more trajectory points (See at least paragraph 439 and 443; As vehicles 1205-1225 travel on road segment 1200, navigation information collected (e.g., detected, sensed, or measured) by vehicles 1205-1225 may be transmitted to server 1230. In some embodiments, the navigation information may be associated with the common road segment 1200. The navigation information may include a trajectory associated with each of the vehicles 1205-1225 as each vehicle travels over road segment 1200. In some embodiments, the trajectory may be reconstructed based on data sensed by various sensors and devices provided on vehicle 1205. For example, the trajectory may be reconstructed based on at least one of accelerometer data, speed data, landmarks data, road geometry or profile data, vehicle positioning data, and ego motion data. In some embodiments, the trajectory may be reconstructed based on data from inertial sensors, such as accelerometer, and the velocity of vehicle 1205 sensed by a speed sensor. In addition, in some embodiments, the trajectory may be determined (e.g., by a processor onboard each of vehicles 1205-1225) based on sensed ego motion of the camera, which may indicate three dimensional translation and/or three dimensional rotations (or rotational motions). The ego motion of the camera (and hence the vehicle body) may be determined from analysis of one or more images captured by the camera. …The autonomous vehicle road navigation model may use map data included in sparse map 800 for determining target trajectories along road segment 1200 for guiding autonomous navigation of autonomous vehicles 1205-1225 or other vehicles that later travel along road segment 1200. For example, when the autonomous vehicle road navigation model is executed by a processor included in a navigation system of vehicle 1205, the model may cause the processor to compare the trajectories determined based on the navigation information received from vehicle 1205 with predetermined trajectories included in sparse map 800 to validate and/or correct the current traveling course of vehicle 1205. ); and
send one or more control signals corresponding to the one or more controls to the one or more actuation components to cause the autonomous or semi-autonomous machine to navigate according to the one or more trajectory points (See at least paragraph 39 – 40; he recognized landmark may include at least one of a traffic sign, an arrow marking, a lane marking, a dashed lane marking, a traffic light, a stop line, a directional sign, a reflector, a landmark beacon, a lamppost, a change is spacing of lines on the road, or a sign for a business. The predetermined road model trajectory may include a three-dimensional polynomial representation of a target trajectory along the road segment. Navigation between recognized landmarks may include integration of vehicle velocity to determine a location of the vehicle along the predetermined road model trajectory. The processor may be further programmed to adjust the steering system of the vehicle based on the autonomous steering action to navigate the vehicle. The processor may be further programmed to: determine a distance of the vehicle from the at least one recognized landmark; and determine whether the vehicle is positioned on the predetermined road model trajectory associated with the road segment based on the distance. The processor may be further programmed to adjust the steering system of the vehicle to move the vehicle from a current position of the vehicle to a position on the predetermined road model trajectory when the vehicle is not positioned on the predetermined road model trajectory. … A method of navigating a vehicle may include receiving, from an image capture device associated with the vehicle, at least one image representative of an environment of the vehicle; analyzing, using a processor associated with the vehicle, the at least one image to identify at least one recognized landmark; determining a current position of the vehicle relative to a predetermined road model trajectory associated with the road segment based, at least in part, on a predetermined location of the recognized landmark; determining an autonomous steering action for the vehicle based on a direction of the predetermined road model trajectory at the determined current location of the vehicle relative to the predetermined road model trajectory; and adjusting a steering system of the vehicle based on the autonomous steering action to navigate the vehicle.).
As per claim 2, Shashua teaches the limitation of:
wherein the one or more processors include at least one of: one or more graphics processing units (GPUs), one or more central processing units (CPUs), or one or more hardware accelerators (See at least paragraph 266).
As per claim 3, Shashua teaches the limitation of:
wherein the autonomous or semi-autonomous machine further comprises one or more systems-on-a-chip (SOCs), and the one or more processors are included in the one or more SOCs (See at least paragraph 266 - 267).
As per claim 4, Shashua teaches the limitation of:
wherein the one or more trajectory points correspond to a turn or a lane change (See at least paragraph 517).
As per claim 5, Shashua teaches the limitation of:
wherein map data is further applied to the E2E neural network, and the data representative of the one or more trajectory points in 3D world space are directly computed further based at least on the E2E neural network processing the map data (See at least paragraph 12 and 540).
As per claim 6, Shashua teaches the limitation of:
wherein vehicle state data is further applied to the E2E neural network, and the data representative of the one or more trajectory points in 3D world space are directly computed further based at least on the E2E neural network processing the vehicle state data (See at least paragraph 422 – 423).
As per claim 7, Shashua teaches the limitation of:
wherein the plurality of sensors include at least two of: a LiDAR sensor; an image sensor; a SONAR sensor; a depth sensor; a microphone sensor; a RADAR sensor; or an ultrasonic sensor (See at least abstract, paragraph 330 and 506).
As per claim 8, Shashua teaches the limitation of:
wherein the autonomous or semi-autonomous machine is a passenger vehicle, a truck, a bus, a robot, a warehouse vehicle, a flying vessel, or a boat (See at least paragraph 6 and 88).
As per claim 10, Shashua teaches the limitation of:
wherein the autonomous or semi-autonomous machine further comprises one or more internal sensors having fields of view or sensory fields internal to the autonomous or semi-autonomous machine, and wherein the computing system or another computing system of the autonomous or semi-autonomous machine perform in-cabin monitoring of one or more passengers using second sensor data obtained using the one or more internal sensors (See at least paragraph 6 and 418).
Regarding claims 9 and 11 – 20:
Claims 9 and 11 – 20 are rejected using the same rationale, mutatis mutandis, applied to claims 1 – 8 above, respectively.
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to IG T AN whose telephone number is (571)270-5110. The examiner can normally be reached M - F: 10:00AM- 4:00PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Aniss Chad can be reached at (571) 270-3832. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
IG T AN
Primary Examiner
Art Unit 3662
/IG T AN/Primary Examiner, Art Unit 3662