DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Applicant’s response to the Non-final Office Action dated 03/12/2026, filed with the office on 04/20/2026, has been entered and made of record.
Status of Claims
Claims 1-20 are pending. Claims 1-4, 6, 8-10, 12, 14-17 and 19 are amended.
Response to Arguments
Applicant's arguments filed on April 20, 2026 with respect to rejection of claims under 35 U.S.C. 103 has been fully considered; but they are not found persuasive. Specifically, in page 10 of its reply, Applicant argues in first paragraph that Chen does not teach temporal processing on forward and backward recurrence. In response to applicant's arguments against the references individually, one cannot show nonobviousness by attacking references individually where the rejections are based on combinations of references. See In re Keller, 642 F.2d 413, 208 USPQ 871 (CCPA 1981); In re Merck & Co., 800 F.2d 1 091, 231 USPQ 375 (Fed. Cir. 1986). While Chen teaches temporal processing of features— ¶0055: “multi-head neural network 110 is configured to extract the spatio-temporal features from the BEV maps”. And Lee discloses calculating the flow of features in both forward and backward directions—
¶0112: “flows in both the forward and the backward directions from the two bird's-eye view image embeddings can be calculated. In this example, flow net 541 calculates the flow of the feature pillars from image one to image two while flow net 542 calculates the flow of the feature pillars from image two to image one” wherein the two images represent the environment across a time sequence— ¶0080: “two birds-eye view images 332, 334 having 3-D embeddings (i.e., pillar features), one representing the first point cloud (e.g., the point cloud at time t−1) and one representing the second point cloud (e.g., the point cloud at time t)”). Therefore, Applicant’s argument is not found persuasive.
Applicant further argues in page 10, second paragraph, that Chen does not disclose each BEV feature images representing a distinct time step. Examiner respectfully disagrees. Chen discloses outputting BEV maps for each time step— ¶0075: “output of each of the classification head, the motion prediction head, and the current motion state estimation head is utilized to generate extended BEV maps for each time step”. And in addition, Lee uses BEV images from distinct time steps—¶0105: “two birds-eye view images 531, 532 having BeV embeddings (e.g., birds-eye view images 332, 334), one representing the first point cloud (e.g., the point cloud at time t−1) and one representing the second point cloud (e.g., the point cloud at time t)”). Therefore, Applicant’s arguments are not found persuasive.
Applicant adds in page 10, third paragraph, that Chen only outputs a single BEV feature map, not the claimed temporally aggregated BEV feature images that are maintained on a per-time-step basis. Examiner respectfully disagrees. Chen discloses multiple BEV maps that are maintained on a per-time-step basis— Chen, ¶0009: “a temporal sequence of BEV maps is generated, where the deep model executes joint reasoning about the category and motion information for each cell in an end-to-end manner”.
Applicant continues to argue in page 11, first paragraph, that Lee does not operate on a sequence level recurrent setting. Examiner respectfully disagrees. Lee discloses in ¶0073: “The second point cloud 312 represents the scene surrounding the vehicle at a time t subsequent to the time t−1 of the scene represented by the first point cloud 314”, therefore, the BEV images used for feature flow estimation are in a sequence. Accordingly, Applicant’s arguments are not found persuasive.
Applicant further argues in page 11, second paragraph, that Lee does not disclose BEV feature extractor aggregating features to generate temporally enriched BEV feature images for downstream heads. Examiner respectfully disagrees. Lee discloses BEV feature extractor in ¶0073: “pillar feature network 320 receives a point cloud from the vehicle LIDAR unit and operates on the data to extract a two-dimensional Birds-eye View pseudo-image from the point cloud”, and aggregating features across two or more images to generate temporally aggregated features— ¶0109: “features from the two (or more) BeV images are aggregated. In the example of FIG. 5, aggregator 552 receives all of the pillar features from the various points of all BeV images” and then the aggregated images are used for further downstream efficient processing— ¶0107: “Aggregator 552 may be configured to group similar features (in the form of pillars) together and represent them as a single feature for more efficient processing. This may allow the system to approximate the original problem with fewer-states in the form of an aggregated problem”. Therefore, Applicant’s arguments are not found persuasive.
Applicant further adds in page 11, third paragraph, that Lee’s optical flow directionality is fundamentally different and does not teach feature aggregation across a time sequence. Examiner respectfully disagrees. Lee discloses aggregating features across a time sequence— ¶0073: “The second point cloud 312 represents the scene surrounding the vehicle at a time t subsequent to the time t−1 of the scene represented by the first point cloud 314”. Therefore, Applicant’s arguments are not found persuasive.
Lastly, Applicant argues in page 11, fourth paragraph, that the combination of Chen and Lee does not teach or suggest the claimed BEV map with feature-level temporal aggregation. Examiner respectfully disagrees. The broadest reasonable interpretation of the claimed ‘feature-level temporal aggregation’ includes combining feature data across time steps, which means, the aggregated features from multiple time steps are represented by a single BEV image. Chen discloses representing information from multiple frames (motion trajectory) into a single frame— ¶0044: “the motion trajectory can be represented as one or combination of a sequence of Cartesian coordinates with a time associated to each coordinate, a sequence of positions and velocities of the vehicle 116, and a sequence of headings of the vehicle 116. To generate the motion trajectory, the motion planner 112 utilizes the output of the multi-head neural network 110. For instance, the motion planner 112 produces the motion trajectory of the vehicle 116 based on the extended BEV image”, additionally, Lee discloses combining features from multiple images and representing the combined feature as a single feature— ¶0107: “group similar features (in the form of pillars) together and represent them as a single feature”. Therefore, Applicant’s arguments are not found persuasive.
Consequently, THIS ACTION IS MADE FINAL.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1, 3-5, 7, 8, 10, 11, 13, 14, 16-18 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Chen et al. (US 2021/0302992 A1), in view of Lee et al. (US 2021/0358296 A1).
Regarding claim 1, Chen teaches, A system for (Chen, ¶0019: “a control system for controlling a motion of a vehicle”) generating perception data, (Chen, ¶0003: “the perception of environment… through two-dimensional (2D) object detection based on camera data”) the system comprising: a processor (Chen, ¶0039: “the control system 100 includes an image processor 106”) operating in an offline processing environment; (Chen, ¶0037: “control system 100 for controlling a motion of a vehicle 116”) and a memory storing machine-readable instructions that, (Chen, ¶0042: “control system 100 includes a memory 108 that stores instructions executable by a controller 104”) when executed by the processor, cause the processor to: (Chen, ¶0042: “controller 104 may be configured to execute the stored instructions in order to control operations”) extract first features (Chen, ¶0094: “multi-head neural network 110 executes feature extraction operation”) from a time sequence of perceptual sensor data to (Chen, ¶0051: “a sequence of video frames including the one or more objects in the environment captured by the camera”) generate a first set of bird's-eye-view (BEV) feature images; (Chen, ¶0007: “environmental state is detected based on a bird's eye view (BEV) map”) extract second features from the first set of BEV feature images (Chen, ¶0055: “extract the spatio-temporal features from the BEV maps”) using a BEV feature extractor that performs feature-level temporal aggregation including (Chen, ¶0055: “multi-head neural network 110 is configured to extract the spatio-temporal features from the BEV maps”). However, Chen does not explicitly teach, both forward recurrence and backward recurrence to generate a second set of BEV feature images, wherein each BEV feature image in the second set of BEV feature images corresponds to a distinct time step in the time sequence of perceptual sensor data and incorporates information from all time steps in the time sequence of perceptual sensor data; and consume the second set of BEV feature images using one or more neural-network heads to perform one of: generating automatically labeled perception data to train one or more of an online perception model, an online prediction model, -or an online planning model used to control an autonomous robot; -or validating performance of an online autonomous stack used to control an autonomous robot.
In an analogous field of endeavor, Lee teaches, both forward recurrence and backward recurrence (Lee, ¶0112: “flows in both the forward and the backward directions from the two bird's-eye view image embeddings can be calculated”) to generate a second set of BEV feature images, wherein each BEV feature image in the second set of BEV feature images corresponds to (Lee, ¶0109: “aggregator 552 receives all of the pillar features from the various points of all BeV images”) a distinct time step in the time sequence of perceptual sensor data and incorporates information from all time steps in the time sequence of perceptual sensor data; (Lee, ¶0105: “two birds-eye view images 531, 532… one representing the first point cloud (e.g., the point cloud at time t−1) and one representing the second point cloud (e.g., the point cloud at time t”) and consume the second set of BEV feature images using (Lee, ¶0107: “two birds-eye view images 531, 532 are aggregated to train classifiers for the features”) one or more neural-network heads (Lee, ¶0108: “classifiers may include, for example, Classification And Regression Tree (CART), K-nearest neighbor, neural network and mixture models”) to perform one of: generating automatically labeled perception data to train one or more of an online perception model, (Lee, ¶0102: “training data for the model may be autonomously labelled”) an online prediction model, or an online planning model (Lee, ¶0008: “Online versions typically model this state…. with a recurrent network trained by self-supervised labeling to predict future states”) used to control an autonomous robot; (Lee, ¶0040: “The systems and methods disclosed herein may be implemented for use in scene flow estimation for robotics, autonomous vehicles and other automated technologies”) or validating performance of an online autonomous stack used to control an autonomous robot. (Lee, ¶0005: “develop a new evaluation by looking at the 3D reconstruction quality of dynamic models”).
Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Chen using the teachings of Lee to introduce generating automatic labels from forward and backward recurrences to train a classification model. A person skilled in the art would be motivated to combine the known elements as described above and achieve the predictable result of detecting the surrounding of an autonomous robot for automatic operation. Therefore, it would have been obvious to combine the analogous arts Chen and Lee to obtain the invention in claim 1.
Regarding claim 3, Chen in view Lee teaches, The system of claim 1, wherein the time sequence of perceptual sensor data includes one or more of camera images, Light Detection and Ranging (LIDAR) data, radar data, sonar data, map data, -or audio data. (Chen, ¶0050: “a plurality of sensors on the vehicle 116 such as a light detection and ranging (LiDAR) sensor, a radio detection and ranging (RADAR) sensor, a camera, and the like”).
Regarding claim 4, Chen in view Lee teaches, The system of claim 1, wherein the one or more neural-network heads include (Chen, ¶0040: “The multi-head neural network 110 includes”) one or more of a three-dimensional (3D) detection head, (Chen, ¶0062: “the classification head may execute… 3D object detection based on LiDAR data, or fusion-based detection”) a 3D semantic-occupancy head, an occupancy-flow head, (Chen, ¶0066: “motion state estimation head outputs occupancy information indicating a state of each of the one or more objects”) a map-elements head, (Chen, ¶0058: “cell classification head executes BEV map segmentation”) an instance-segmentation head, a panoptic-segmentation head, (Chen, ¶0059: “the semantic segmentation corresponds to classification of every pixel of the BEV maps into a corresponding class”) a drivable- surface-estimation head, (Chen, ¶0004: “OGM pipeline can be utilized to specify a future drivable space and thereby provide support for motion planning”) -or an elevation-estimation head. (Chen, ¶0054: “converts the 3D voxel lattice into a 2D pseudo-image with a height dimension”).
Regarding claim 5, Chen in view Lee teaches, The system of claim 1, wherein the autonomous robot is an autonomous vehicle. (Chen, ¶0038: “The vehicle 116 may be an autonomous vehicle or a semi-autonomous vehicle”).
Regarding claim 7, Chen in view Lee teaches, The system of claim 1, wherein the online autonomous stack includes perception, prediction, and planning models. (Chen, ¶0002: “autonomous systems such as autonomous vehicles… facilitates motion planning… (1) perception, which identifies the foreground objects from the background; and (2) motion prediction”).
Regarding claim 8, it recites a computer-readable medium including instructions corresponding to the elements of the system recited in claim 1. Therefore, the recited instructions of the computer-readable medium of claim 8 are mapped to the proposed combination in the same manner as the corresponding elements of the system claim 1. Additionally, the rationale and motivation to combine Chen and Lee presented in rejection of claim 1, apply to this claim. In addition, Chen teaches, A non-transitory computer-readable medium for generating perception data and storing instructions that, when executed by a processor, cause the processor to: (Chen, ¶0109: “the program code or code segments to perform the necessary tasks may be stored in a machine readable medium. A processor(s) may perform the necessary tasks”).
Regarding claim 10, it recites a computer-readable medium including instructions corresponding to the elements of the system recited in claim 4. Therefore, the recited instructions of the computer-readable medium of claim 10 are mapped to the proposed combination in the same manner as the corresponding elements of the system claim 4. Additionally, the rationale and motivation to combine Chen and Lee presented in rejection of claim 1, apply to this claim.
Regarding claim 11, it recites a computer-readable medium including instructions corresponding to the elements of the system recited in claim 5. Therefore, the recited instructions of the computer-readable medium of claim 11 are mapped to the proposed combination in the same manner as the corresponding elements of the system claim 5. Additionally, the rationale and motivation to combine Chen and Lee presented in rejection of claim 1, apply to this claim.
Regarding claim 13, it recites a computer-readable medium including instructions corresponding to the elements of the system recited in claim 7. Therefore, the recited instructions of the computer-readable medium of claim 13 are mapped to the proposed combination in the same manner as the corresponding elements of the system claim 7. Additionally, the rationale and motivation to combine Chen and Lee presented in rejection of claim 1, apply to this claim.
Regarding claim 14, it recites a method with steps corresponding to the elements of the system recited in claim 1. Therefore, the recited steps of the method claim 14 are mapped to the proposed combination in the same manner as the corresponding elements of the system claim 1. Additionally, the rationale and motivation to combine Chen and Lee presented in rejection of claim 1, apply to this claim. In addition, Chen teaches, A method (Chen, ¶00111: “Embodiments of the present disclosure may be embodied as a method”).
Regarding claim 16, it recites a method with steps corresponding to the elements of the system recited in claim 3. Therefore, the recited steps of the method claim 16 are mapped to the proposed combination in the same manner as the corresponding elements of the system claim 3. Additionally, the rationale and motivation to combine Chen and Lee presented in rejection of claim 1, apply to this claim.
Regarding claim 17, it recites a method with steps corresponding to the elements of the system recited in claim 4. Therefore, the recited steps of the method claim 17 are mapped to the proposed combination in the same manner as the corresponding elements of the system claim 4. Additionally, the rationale and motivation to combine Chen and Lee presented in rejection of claim 1, apply to this claim.
Regarding claim 18, it recites a method with steps corresponding to the elements of the system recited in claim 5. Therefore, the recited steps of the method claim 18 are mapped to the proposed combination in the same manner as the corresponding elements of the system claim 5. Additionally, the rationale and motivation to combine Chen and Lee presented in rejection of claim 5, apply to this claim.
Regarding claim 20, it recites a method with steps corresponding to the elements of the system recited in claim 7. Therefore, the recited steps of the method claim 20 are mapped to the proposed combination in the same manner as the corresponding elements of the system claim 7. Additionally, the rationale and motivation to combine Chen and Lee presented in rejection of claim 1, apply to this claim.
Claims 2, 6, 9, 12, 15 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Chen et al. (US 2021/0302992 A1), in view of Lee et al. (US 2021/0358296 A1) and in further view of Beaver et al. (US 2023/0274541 A1).
Regarding claim 2, Chen in view Lee teaches, The system of claim 1. However, the combination of Chen and Lee does not explicitly teach, wherein the BEV feature extractor includes one of a plurality of Gated Recurrent Units (GRUs), a plurality of Long Short-Term Memory (LSTM) networks, -or a plurality of transformer networks to perform the feature-level temporal aggregation including both forward recurrence and backward recurrence.
In an analogous field of endeavor, Beaver teaches, wherein the BEV feature extractor includes one of a plurality of Gated Recurrent Units (GRUs), a plurality of Long Short-Term Memory (LSTM) networks, and a plurality of transformer networks to perform the feature-level temporal aggregation including both forward recurrence and backward recurrence. (Beaver, ¶0019: “a sequence-to-sequence model such a recurrent neural network (RNN), long short-term memory (LSTM) network, gated recurrent unit (GRU) network, and/or a Bidirectional Encoder Representations from Transformers (BERT) transformer network, may be used to process individual pixels or groups of pixels as a sequence”).
Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Chen in view of Lee using the teachings of Beaver to introduce a plurality of extractor models. A person skilled in the art would be motivated to combine the known elements as described above and achieve the predictable result of detecting the surrounding of an autonomous robot using extracted features. Therefore, it would have been obvious to combine the analogous arts Chen, Lee and Beaver to obtain the invention in claim 2.
Regarding claim 6, Chen in view Lee teaches, The system of claim 1. However, the combination of Chen and Lee does not explicitly teach, wherein the autonomous robot is one of a search and rescue robot, a delivery robot, an aerial drone, -or an indoor robot.
In an analogous field of endeavor, Beaver teaches, wherein the autonomous robot is one of a search and rescue robot, a delivery robot, an aerial drone, and an indoor robot. (Beaver, ¶0030: “An individual robot 108.sub.1-M may take various forms, such as an unmanned aerial vehicle 108-1”).
Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Chen in view of Lee using the teachings of Beaver to introduce an unmanned aerial vehicle. A person skilled in the art would be motivated to combine the known elements as described above and achieve the predictable result of autonomously operating an aerial drone using the perception data. Therefore, it would have been obvious to combine the analogous arts Chen, Lee and Beaver to obtain the invention in claim 6.
Regarding claim 9, it recites a computer-readable medium including instructions corresponding to the elements of the system recited in claim 2. Therefore, the recited instructions of the computer-readable medium of claim 9 are mapped to the proposed combination in the same manner as the corresponding elements of the system claim 2. Additionally, the rationale and motivation to combine Chen, Lee and Beaver presented in rejection of claim 2, apply to this claim.
Regarding claim 12, it recites a computer-readable medium including instructions corresponding to the elements of the system recited in claim 6. Therefore, the recited instructions of the computer-readable medium of claim 12 are mapped to the proposed combination in the same manner as the corresponding elements of the system claim 6. Additionally, the rationale and motivation to combine Chen, Lee and Beaver presented in rejection of claim 6, apply to this claim.
Regarding claim 15, it recites a method with steps corresponding to the elements of the system recited in claim 2. Therefore, the recited steps of the method claim 15 are mapped to the proposed combination in the same manner as the corresponding elements of the system claim 2. Additionally, the rationale and motivation to combine Chen, Lee and Beaver presented in rejection of claim 2, apply to this claim.
Regarding claim 19, it recites a method with steps corresponding to the elements of the system recited in claim 6. Therefore, the recited steps of the method claim 19 are mapped to the proposed combination in the same manner as the corresponding elements of the system claim 6. Additionally, the rationale and motivation to combine Chen, Lee and Beaver presented in rejection of claim 6`, apply to this claim.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MEHRAZUL ISLAM whose telephone number is (571)270-0489. The examiner can normally be reached Monday-Friday: 8am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Saini Amandeep can be reached on (571) 272-3382. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MEHRAZUL ISLAM/Examiner, Art Unit 2662
/AMANDEEP SAINI/Supervisory Patent Examiner, Art Unit 2662