DETAILED ACTION
This office action is responsive to applicant’s amendments and arguments filed 07/06/2026.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Application is acknowledged as a National Stage application of PCT/CN2022/085560. Priority to PCT/CN2022/085560 with a priority date of 04/07/2022 is acknowledged under 35 USC 119(e) and 37 CFR 1.78.
Response to Arguments
Applicant's arguments filed regarding the rejection of independent claims 1, 7, 13, 19, and 25 have been fully considered but they are not persuasive.
Applicant argues that the combination of Huang and Roddick fails to teach or suggest the amended limitations: “use the neural network to generate, from the input information, a unified 3D representation including features corresponding to the one or more objects and the 3D environment”, and “use the neural network to jointly infer, based, at least in part, on the unified 3D representation”, both positions of objects and a segmented map. Applicant states that the unified 3D representation is not taught for two main reasons:
1) Huang teaches determining object positions. It “uses a framework/pipeline inspired by, and in common with, other tools, such as tools to generate segmented maps”, but does not explicitly teach the segmentation aspect and therefore does not teach a unified 3D representation.
2) Roddick teaches segmentation from monocular images without being incorporated into a unified 3D representation.
However, Huang teaches a pipeline that is not only “inspired by, and in common with” segmentation tools, but is designed to be used in conjunction with them. Huang uses a birds-eye-view (BEV) representation specifically because this is the representation used by segmentation tools such as the invention of Roddick. The BEV representation is proposed as a potential unified 3D representation for both segmentation and object detection, as explained in the introduction of Huang (pg. 1-2):
“For example, in the nuScenes [1] benchmark, image-view-based methods like FCOS3D [49] and PGD [50] have leading performances in the multi-camera 3D object detection track, while the BEV semantic segmentation track is dominated by the BEV-based methods like PON [39], Lift-Splat-Shoot [33], and VPN [31]. Which view space is more reasonable for perception in autonomous driving, and can we handle these tasks in a unified framework?… With BEVDet, we explore the advantages of detecting 3D objects in BEV, expecting a superior performance compared to the latest image-view-based methods and a consistent paradigm with BEV semantic segmentation. In this way, we can further verify the feasibility of multi-task learning, which is meaningful for time-efficient inference.”,
as well as the conclusion of Huang (pg. 15):
“Future works will focus on (1) improving the performance of BEVDet, particularly on targets’ attribute prediction. (2) studying multi-task learning based on BEVDet.”
Huang does not explicitly teach the implementation details of the segmentation aspect of the unified 3D representation, since it is left to the compatible inventions, but it makes multiple references to the concept of the unified 3D representation. Upon reading Huang, one of ordinary skill in the art would have been directed toward attempting to implement the combination of Huang with a compatible BEV segmentation tool, such as Roddick, in order to test Huang’s proposal of a unified system. Huang is very clear about its intent to be combined with other tools:
Pg. 2 “The proposed BEVDet, as illustrated in Fig. 1, shares a similar framework with the up-to-date BEV semantic segmentation algorithms [33,39,54]. It is modularly designed with an image-view encoder for encoding features in image view, a view transformer for transforming the feature from image view into BEV, a BEV encoder for further encoding features in the BEV perspective, and a task specific head for performing 3D object detection in the BEV space. Benefiting from this modular design, we can reuse a mass of existing works which have been proved effective in other areas...”
Each of the modules mentioned in this section are also present in the invention of Roddick, as explained in the “Response to Arguments” section of the previous office action, supporting the compatibility of the inventions of Huang and Roddick.
Applicant is correct that Roddick alone does not teach a unified 3D representation for both segmentation and object detection. However, Roddick also teaches transforming image features into a BEV representation (fig. 2 “(3) A stack of dense transformer layers map the image-based features into the birds-eye-view.”). In the context of Huang, the BEV representation taught by Roddick is the unified 3D representation, since Huang proposes sharing it between segmentation and object detection tasks. The combination of Huang in view of Roddick, when considered together, teaches the amended limitations.
Therefore, the rejection of claim 1 and the other independent claims under 35 U.S.C. 103 is maintained.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-3, 7-9, 13-15, and 25-27 is/are rejected under 35 U.S.C. 103 as being unpatentable over Huang et al. (“BEVDet: High-performance Multi-camera 3D Object Detection in Bird-Eye-View”. arXiv preprint (31 Mar 2022). https://arxiv.org/abs/2112.11790v2; hereinafter "Huang") in view of Roddick et al. (“Predicting Semantic Map Representations from Images using Pyramid Occupancy Networks”. arXiv preprint (30 Mar 2020). https://arxiv.org/abs/2003.13402v1; hereinafter "Roddick").
Regarding claim 1, Huang teaches: A processor, comprising: one or more circuits (pg. 9 section 4.1 “Experimental Settings” subsection “Training Parameters”: “Models are trained with AdamW [28] optimizer, in which gradient clip is exploited with learning rate 2e-4, a total batch size of 64 on 8 NVIDIA GeForce RTX 3090 GPUs.”) to:
input information to a neural network based on two dimensional (2D) images of one or more objects in a three-dimensional (3D) environment (fig. 1, first step “Image-view Encoder” has multiple 2D images as input; section 1 “Introduction” discusses 2D visual perception; pg. 2 states the purpose is “performing 3D object detection”), wherein respective images of the 2D images capture different points of view of the 3D environment (section 4.1 “Experimental Settings” subsection “Dataset”: “We conduct comprehensive experiments on the large-scale benchmark nuScenes [1]. The nuScenes benchmark includes 1000 scenes with images from 6 cameras.”; the associated reference (see References Cited) shows an example where each of the 6 cameras faces a different direction);
use the neural network to generate, from the input information, a unified 3D representation including features corresponding to the one or more objects and the 3D environment (fig. 1 first two modules generate 3D bird’s-eye-view (BEV) representation from image features: “Image-view encoder, including a backbone and a neck, is applied at first for image feature extraction. View transformer transforms the feature from the image view to BEV.”, see pg. 5-6 section 3.1 “Image-view Encoder” and “View Transformer” subsection for more detail;
pg. 1-2 Introduction section proposes the use of the BEV representation as a unified 3D representation which can handle both segmentation and object detection: “…image-view-based methods like FCOS3D [49] and PGD [50] have leading performances in the multi-camera 3D object detection track, while the BEV semantic segmentation track is dominated by the BEV-based methods like PON [39], Lift-Splat-Shoot [33], and VPN [31]. Which view space is more reasonable for perception in autonomous driving, and can we handle these tasks in a unified framework?… With BEVDet, we explore the advantages of detecting 3D objects in BEV, expecting a superior performance compared to the latest image-view-based methods and a consistent paradigm with BEV semantic segmentation. In this way, we can further verify the feasibility of multi-task learning, which is meaningful for time-efficient inference.”); and
use the neural network to infer, based, at least in part, on the unified 3D representation, one or more positions of the one or more objects within the 3D environment (fig. 1 “Head”; pg. 6 section 3.1 “Network Structure” subsection “Head” teaches a network head for 3D object detection) from a viewpoint distinct from the different points of view of the 2D images (fig. 1 “View Transformer”; pg. 6 section 3.1 “Network Structure” subsection “View Transformer” teaches transforming features from the initial 2D image view space to a shared 3D birds-eye-view space).
Huang also teaches the possibility of adapting its invention to use the neural network to jointly infer, based, at least in part, on the unified 3D representation, one or more positions of the one or more objects within the 3D environment and a segmented map of the 3D environment from a viewpoint distinct from the different points of view of the 2D images (pg. 1-2 section 1 “Introduction”: “However, with respect to the scene of vision-based autonomous driving where both accuracy and time efficiency are desired, major tasks like 3D object detection and map restoration (i.e., Bird-Eye-View (BEV) semantic segmentation) are still conducted by different paradigms in the up-to-date benchmarks… Which view space is more reasonable for perception in autonomous driving, and can we handle these tasks in a unified framework? Aiming at these questions, we propose BEVDet in this paper. With BEVDet, we explore the advantages of detecting 3D objects in BEV, expecting a superior performance compared to the latest image-view-based methods and a consistent paradigm with BEV semantic segmentation. In this way, we can further verify the feasibility of multi-task learning, which is meaningful for time-efficient inference.”; it is suggested that “multi-task learning” refers to training a neural network to perform both 3D object detection and map segmentation;
pg. 15 section 5 “Conclusion”: “Future works will focus on studying multi-task learning based on BEVDet.”).
However, the invention of Huang does not itself perform the segmentation task. Thus, Huang does not explicitly teach: use the neural network to jointly infer, based, at least in part, on the unified 3D representation, a segmented map of the 3D environment from a viewpoint distinct from the different points of view of the 2D images.
Roddick teaches: A processor, comprising: one or more circuits (pg. 2 col. 1 “The method is fast enough to be used in real time applications, processing 23.2 frames per second on a single GeForce RTX 2080 Ti graphics card.”) to:
input information to a neural network based on two dimensional (2D) images of one or more objects in a three-dimensional (3D) environment (fig. 2 shows image input to neural network), wherein respective images of the 2D images capture different points of view of the 3D environment (fig. 1 upper part shows six 2D input images with different viewpoints); and
use the neural network to infer, based on the input information, a segmented map of the 3D environment from a viewpoint distinct from the different points of view of the 2D images (fig. 1 lower part, “Given a set of surround-view images, we predict a full 360◦ birds-eye-view semantic map, which captures both static elements like road and sidewalk as well as dynamic actors such as cars and pedestrians.”, where none of the input images were from a bird’s-eye perspective; fig. 2 shows the neural network architecture).
Huang and Roddick are both analogous to the claimed invention because they are in the same field of 2D visual perception for autonomous vehicle navigation. Furthermore, Huang explicitly teaches that it is designed using similar architecture to Roddick (along with several other birds-eye-view segmentation methods) in order to test the feasibility of combining the object detection functionality of Huang with the semantic segmentation functionality of Roddick and others (Huang pg. 1-2 section “Introduction”; citation 39 is Roddick; also see “Response to Arguments” section). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified the invention of Huang with the teachings of Roddick to establish a framework in which a neural network can simultaneously perform both object detection and semantic segmentation of 2D input images for autonomous vehicle navigation. The motivation would have been to improve the speed and efficiency of the navigation system (mentioned in Huang “Introduction”, pg. 2), which is an important consideration when operating in real time.
Claims 7, 13, and 25 are rejected with the same rationales, references, and motivations to combine as claim 1.
Regarding claim 2, the combination of Huang in view of Roddick teaches the processor of claim 1, wherein the one or more circuits are further to cause an encoder of the neural network to extract the features, corresponding to the one or more objects and the 3D environment, into a 2D view space (Huang fig. 1 “Image-view Encoder”; pg. 5 section 3.1 “Network Structure” subsection “Image-view Encoder” teaches encoding features from the input images into 2D “image-view space”).
Claims 8, 14, and 26 are rejected with the same rationales, references, and motivations to combine as claim 2.
Regarding claim 3, the combination of Huang in view of Roddick teaches the processor of claim 2, wherein the one or more circuits are further to transform the features from the 2D view space into the unified 3D representation (Huang fig. 1 “View Transformer” and “BEV Encoder”; pg. 6 section 3.1 “Network Structure” subsection “View Transformer” teaches transforming features from multiple images, each with its own 2D image view space, to a shared 3D birds-eye-view space; subsection “BEV Encoder” teaches further encoding features in the 3D BEV space).
Claims 9, 15, and 27 are rejected with the same rationales, references, and motivations to combine as claim 3.
Claim(s) 4-6, 10-12, 16-18, 19-24, and 28-30 is/are rejected under 35 U.S.C. 103 as being unpatentable over Huang ("BEVDet: High-performance Multi-camera 3D Object Detection in Bird-Eye-View") in view of Roddick ("Predicting Semantic Map Representations from Images using Pyramid Occupancy Networks") as applied to claims 1, 7, 13, 19, and 25 above, and further in view of Vora et al. (US 20220371606 A1, hereinafter "Vora").
Claims 19, 20, and 21 are rejected with the same rationales, references, and motivations to combine as claims 1, 2, and 3 respectively, with the additional limitation of a machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to… (Vora [0071], [0213]).
Vora and the combination of Huang in view of Roddick are analogous to the claimed invention because they are in the same field of 2D visual perception for autonomous vehicle navigation. It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified the invention of Huang in view of Roddick with the teachings of Vora to store program instructions using a non-transitory storage medium in order for the system to be easily saved, modified, and run on demand.
Regarding claim 4, the combination of Huang in view of Roddick teaches the processor of claim 1, wherein the one or more circuits are further to utilize respective network heads of the neural network to determine the one or more positions of the one or more objects (Huang fig. 1 “Head”; pg. 6 section 3.1 “Network Structure” subsection “Head” teaches a network head for 3D object detection for “position, scale, orientation, and speed of movable objects like pedestrians, vehicles, barriers, and so on.”).
The combination of Huang in view of Roddick does not explicitly teach wherein the one or more circuits are further to utilize respective network heads of the neural network to generate the segmented map of the 3D environment.
Vora teaches wherein the one or more circuits are further to utilize respective network heads of the neural network to determine the one or more positions of the one or more objects (fig. 10 bounding box regression head 1018, [0155] “In examples, the network 100 includes…a classification head 1016 and bounding box regression head 1018 for 3D object detection”), and to generate the segmented map of the 3D environment (fig. 10 segmentation head 1014, [0105] “…at block 604, a segmentation head of the network determines a class associated with each pixel.”, [0155] “The network 1000 achieves panoptic segmentation by combining the outputs of the semantic segmentation and object detection.”).
Vora and the combination of Huang in view of Roddick are analogous to the claimed invention because they are in the same field of 2D visual perception for autonomous vehicle navigation. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified the invention of Huang in view of Roddick with the teachings of Vora to use network heads to perform both object detection and semantic segmentation. The motivation would have been to improve the speed and efficiency of the navigation system, which is an important consideration when operating in real time.
Claims 10, 16, 22, and 28 are rejected with the same rationales, references, and motivations to combine as claim 4. Claim 22 has the additional limitation of a machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to… (Vora [0071], [0213]), with the same motivation to combine as claims 19-21.
Regarding claim 5, the combination of Huang in view of Roddick and further in view of Vora teaches the processor of claim 4, wherein the one or more circuits are further to utilize one or more task-specific network heads to generate one or more additional inferences based, at least in part, upon the features in the unified 3D representation (Huang fig. 1 “Head”; pg. 6 section 3.1 “Network Structure” subsection “Head”: “The task-specific head is constructed upon the BEV feature. In common sense [1], 3D object detection in automatic pilot aims at the position, scale, orientation, and speed of movable objects like pedestrians, vehicles, barriers, and so on.” – task-specific head may perform a task other than determining object position).
Claims 11, 17, 23, and 29 are rejected with the same rationales, references, and motivations to combine as claim 5. Claim 23 has the additional limitation of a machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to… (Vora [0071], [0213]), with the same motivation to combine as claims 19-21.
Regarding claim 6, the combination of Huang in view of Roddick teaches the processor of claim 1, but does not explicitly teach wherein the one or more circuits are further to provide the one or more positions and the segmented map to a navigation system to generate navigation instructions for an automated device.
Vora teaches wherein the one or more circuits are further to provide the one or more positions and the segmented map to a navigation system to generate navigation instructions for an automated device ([0004] “The subject matter described in this specification is directed to a computer system and techniques for detecting objects in an environment surrounding an autonomous vehicle. Generally, the computer system is configured to receive input from one or more sensors of the vehicle, detect one or more objects in the environment surrounding the vehicle based on the received input, and operate the vehicle based upon the detection of the objects.”;
[0078] “In some embodiments, planning system 404 receives data associated with a destination and generates data associated with at least one route (e.g., routes 106) along which a vehicle (e.g., vehicles 102) can travel along toward a destination. In some embodiments, planning system 404 periodically or continuously receives data from perception system 402 (e.g., data associated with the classification of physical objects, described above) and planning system 404 updates the at least one trajectory or generates at least one different trajectory based on the data generated by perception system 402.”;
[0103] “In some embodiments, perception system 402 receives data associated with at least one physical object (e.g., data that is used by perception system 402 to detect the at least one physical object) in an environment and classifies the at least one physical object. A planning system (e.g., planning system 404) generates or updates at least one trajectory based on the data generated by the perception system.”;
[0105] “In some embodiments, the detected and segmented objects as output by the segmentation head 604, classification head 606, and bounding box head 608 are generated by or provided to the perception system 402, planning system 404, localization system 406, and/or control system 408 (FIG. 4).”).
Vora and the combination of Huang in view of Roddick are analogous to the claimed invention because they are in the same field of 2D visual perception for autonomous vehicle navigation. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified the invention of Huang in view of Roddick with the teachings of Vora to include a planning system which uses the object detection and semantic segmentation output of Huang in view of Roddick to plan a route for autonomous vehicle. The motivation would have been to apply the experimental teachings of Huang and Roddick towards their intended practical usage for autonomous vehicle navigation.
Claims 12, 18, 24, and 30 are rejected with the same rationales, references, and motivations to combine as claim 6. Claim 24 has the additional limitation of a machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to… (Vora [0071], [0213]), with the same motivation to combine as claims 19-21.
References Cited
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Philion et al. (“Lift, Splat, Shoot: Encoding Images From Arbitrary Camera Rigs by Implicitly Unprojecting to 3D”. arXiv preprint (13 Aug 2020). https://arxiv.org/abs/2008.05711v1) teaches a system for semantic segmentation of multiple 2D image inputs for autonomous vehicle navigation. Similarly to the invention of Roddick, the invention of Philion involves first encoding features in a 2D image-view space, then performing a transformation to a 3D birds-eye-view space, and finally performing semantic segmentation within the 3D BEV space.
Yin et al. (“Center-based 3D Object Detection and Tracking”. arXiv preprint (6 Jan 2021). https://arxiv.org/abs/2006.11275v2) teaches a neural network with task-specific heads to detect an object’s location, size/orientation, and velocity. The task-specific heads of Yin are stated to be incorporated into the invention of Huang.
Caesar et al. ("nuScenes: A multimodal dataset for autonomous driving". arXiv preprint (5 May 2020). https://arxiv.org/abs/1903.11027v5) is cited by Huang as the dataset used to test its invention. It features 6 camera images per scene, each captured from a different angle, as mentioned in the rejection of claim 1.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to BENJAMIN STATZ whose telephone number is (571)272-6654. The examiner can normally be reached Mon-Fri 8am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Tammy Goddard can be reached at (571)272-7773. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/BENJAMIN TOM STATZ/Examiner, Art Unit 2611
/DAVID T WELCH/Primary Examiner, Art Unit 2613