DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 02/25/2025 was filed and is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claim(s) 1-8, 10, 11, 13-20 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Gosala (“SkyEye: Self-Supervised Bird’s-Eye-View Semantic Mapping Using Monocular Frontal View Images”, 2023, as cited in IDS filed 02/25/2025).
Regarding claims 1 and 15, Gosala teaches A method, comprising:
receiving, by a device, video data (Gosala, Abstract: “FV semantic annotations of video sequences”) that includes video frames (Gosala, Abstract: “FV semantic annotations of video sequences”) depicting monocular frontal views (Gosala, Abstract: “FV semantic annotations of video sequences”. FV stands for Front View);
selecting, by the device, a reference video frame (Gosala, pg 1, column 2, “FV semantic ground truth labels”) and a target video frame (Gosala, pg 2, column 1, first full paragraph, reproduced below:
PNG
media_image1.png
684
742
media_image1.png
Greyscale
. “FV semantic predictions” show a target video frame is involved) from the video data (Gosala, pg 1, column 2, “overcomes the need for BEV ground truths by leveraging FV semantic ground truth labels along with the spatial and temporal consistency offered by video sequences”);
processing, by the device, the reference video frame (Gosala, pg 1, column 2, “FV semantic ground truth labels”), with a bird’s eye view (BEV) model (Gosala, see nearest image above, “BEV semantic pseudolabels” is being interpreted as involving BEV model), to generate a rendered semantic segmentation (Gosala, see nearest image above, “FV semantic prediction”, which is being interpreted as involving segmentation) and a BEV prediction (Gosala, see nearest image above, “BEV semantic pseudolabels” is being interpreted as involving BEV prediction);
sampling, by the device, class probability values (Gosala, pg 5, column 2, Section 4.3, ¶1: “Tab. 1 presents the results of this evaluation using the class-wise Intersection-over-Union (IoU) and overall mean IoU (mIoU) metrics for the KITTI-360 dataset”. The metrics shows that class probability values occur, this is how machine learning models work, they have probability values of a class when doing predictions, as PHOSITA would know) from the BEV prediction (Gosala, see nearest image above, “BEV semantic pseudolabels” is being interpreted as involving BEV prediction);
processing, by the device, the target video frame (Gosala, see nearest image above, “FV semantic predictions” show a target video frame is involved), with a geometry model (Gosala, see nearest image below, “BEV Map” and “shape of dynamic objects” is being interpreted as involving “a geometry model”), to generate densities (Gosala, pg 4, column 2, After Fig. 2 is mentioned, reproduced below:
PNG
media_image2.png
148
486
media_image2.png
Greyscale
. “Densify” is being interpreted as involving generating densities);
generating, by the device, a target semantic segmentation (Gosala, see pg 1 image above, “FV semantic predictions”) based on the class probability values (Gosala, pg 5, column 2, Section 4.3, ¶1: “Tab. 1 presents the results of this evaluation using the class-wise Intersection-over-Union (IoU) and overall mean IoU (mIoU) metrics for the KITTI-360 dataset”. The metrics shows that class probability values occur, this is how machine learning models work, they have probability values of a class when doing predictions, as PHOSITA would know) and the densities (Gosala, see nearest image above, “Densify” is being interpreted as involving generating densities);
calculating, by the device, a cross-entropy loss based on the rendered semantic segmentation and the target semantic segmentation (Gosala, pg 4, column 1, paragraph before equation 2, selection reproduced below:
PNG
media_image3.png
184
736
media_image3.png
Greyscale
); and
training (Gosala, see nearest image above, loss shows training occurs), by the device, the BEV model (Gosala, Figure 1, “BEV prediction” shows a BEV model is used), with the cross-entropy loss (Gosala, see nearest image above, “cross entropy loss”), in order to generate a trained BEV model (Gosala, Figure 1’s “BEV prediction” and nearest image above with “cross entropy loss” shows a trained BEV model is generated).
Regarding claim 2, Gosala teaches The method of claim 1, further comprising:
receiving additional video data that includes video frames (Gosala, pg 5, Section 4.1, “We evaluate SkyEye on the KITTI-360 dataset”. “Evaluation” is being interpreted as involving “receiving additional video data that includes video frames”) depicting monocular frontal views (Gosala, Figure 1, which shows “FV Image”; Abstract: “automated driving pipelines”, which are being interpreted as involving “monocular frontal views”);
processing the additional video data (Gosala, pg 5, Section 4.1, “We evaluate SkyEye on the KITTI-360 dataset”. “Evaluation” is being interpreted as involving “receiving additional video data that includes video frames”), with the trained BEV model (Gosala, pg 5, Section 4.1, “We evaluate SkyEye on the KITTI-360 dataset”. “SkyEye”, after training, is being interpreted as “trained BEV model”), to generate a new BEV prediction (Gosala, Abstract: “In this work, we address this limitation by proposing the first self-supervised approach for generating a BEV semantic map using a single monocular image from the frontal view (FV).” “Generating a BEV semantic map” is being interpreted as “BEV prediction”); and
providing the new BEV prediction (Gosala, Abstract: “In this work, we address this limitation by proposing the first self-supervised approach for generating a BEV semantic map using a single monocular image from the frontal view (FV).” “Generating a BEV semantic map” is being interpreted as “BEV prediction”; Figure 1, which shows “BEV Prediction”; Figure 5, column “SkyEye (Ours)”, is being interpreted a non-exhaustive example of BEV prediction).
Regarding claim 3, Gosala teaches The method of claim 1, wherein sampling the class probability values from the BEV prediction comprises:
collecting the class probability values associated with each pixel or segment within the BEV prediction to generate a probabilistic map (Gosala, Figure 1, which shows BEV Prediction; Section 3, Equation 1, which shows loss functions. PHOSITA would know for semantic segmentation that probabilities for each pixel or segment would have a probability attached to it. The loss function may be involved with these probabilities during training. Therefore, the BEV prediction is being interpreted as involving probabilistic map, the semantic segmentation is a non-exhaustive example).
Regarding claim 4, Gosala teaches The method of claim 1, wherein processing the target video frame, with the geometry model, to generate the densities comprises:
utilizing a point cloud processor (Gosala, pg 4, Section 3.3, ¶1, “a depth prediction pipeline to lift FV semantic annotations into BEV yielding a semantic point cloud”. Which shows a point cloud processor is involved to create a point cloud) to analyze the target video frame (Gosala, pg 4, Section 3.3, ¶1: “a semantic point cloud”, which shows an analysis on the target video frame to create a semantic point cloud) and convert image pixels into a three-dimensional point cloud representation (Gosala, pg 4, Section 3.3, ¶1: “point cloud”, which is being interpreted as having three-dimensions) to extract the densities from the target video frame (Gosala, pg 4, Section 3.3, ¶1: “a densification module to generate dense segmentation masks from sparse depth predictions for static classes”. Which is being interpreted as involving extracting densities from the target video frame).
Regarding claim 5, Gosala teaches The method of claim 1, wherein generating the target semantic segmentation based on the class probability values and the densities comprises:
performing a volumetric rendering (Gosala, pg 4, Section 3.2, ¶1: “We hypothesize that this formulation would help the network generate a spatially consistent volumetric representation of the scene from a single FV image”. “Volumetric representation” is being interpreted as “volumetric rendering”) of a semantic perspective view (Gosala, pg 4, Section 3.2, ¶1: “BEV semantic maps”) for the target video frame (Gosala, pg 4, Section 3.2, ¶1: “FV image”) using the class probability values (Gosala, pg 4, Section 3.2, ¶2: “Subsequently, we compute the cross entropy loss between the FV semantic predictions and their corresponding FV semantic ground truths to generate the implicit supervision signal for training the model”. “FV semantic predictions” are being interpreted as involving “class probability values”) and the densities (Gosala, pg 4, Section 3.2, ¶1, “volumetric representation” is being interpreted as involving densities, as PHOSITA would know that volumes have density), wherein the semantic perspective view (Gosala, pg 4, Section 3.2, ¶2: “FV semantic predictions“) corresponds to the target semantic segmentation (Gosala, pg 4, Section 3.2, ¶2 : “their corresponding FV semantic ground truths”).
Regarding claim 6, Gosala teaches The method of claim 1, further comprising:
generating semantic segmentation labels (Gosala, pg 5, Figure 4 image and text, “Overview of our pseudolabel generation pipeline”) based on the rendered semantic segmentation (Gosala, pg 5, Figure 4 image and text, “We lift semantic annotations in FV into the 3D world”. “Into the 3D world” is being interpreted as involving “rendered semantic segmentation”) and the target semantic segmentation (Gosala, pg 5, Figure 4 image and text, “We lift semantic annotations in FV”)and
training the BEV model (Gosala, Section 3.3, title: “Explicit Supervision”, is being interpreted as involving training the BEV model), with the semantic segmentation labels (Gosala, pg 5, Figure 4 image and text, “Overview of our pseudolabel generation pipeline”), to generate the trained BEV model (Gosala, Section 3.3, title: “Explicit Supervision”, is being interpreted as involving training the BEV model, which generates the trained BEV model).
Regarding claim 7, Gosala teaches The method of claim 1, further comprising:
receiving camera calibration and pose information (Gosala, pg 4, column 2, Subsection “Pseudolabel Generation”: “lift FV semantic ground truths into BEV using the known camera intrinsics and poses.” “Camera intrinsics” is being interpreted as involving “camera calibration”) associated with the video data (Gosala, pg 4, column 2, Subsection “Pseudolabel Generation”: “lift FV semantic ground truths”, which is being interpreted as involving associated video data); and
utilizing the camera calibration and pose information with the geometry model (Gosala, pg 4, column 2, Subsection “Pseudolabel Generation”: “lift FV semantic ground truths into BEV using the known camera intrinsics and poses.” “lift FV semantic” is being interpreted as involving the geometry model; as supported by Figure 4, “We lift semantic annotations in FV into the 3D world”, which is being interpreted as involving geometry) to generate the densities (Gosala, pg 4, column 2, Subsection “Pseudolabel Generation”: “We then densify the sparse BEV map”).
Regarding claim 11, Gosala teaches The device of claim 8, wherein the one or more processors, to train the BEV model, with the cross-entropy loss, in order to generate the trained BEV model, are configured to:
Backpropagate (Gosala, pg 4, Section 3.3, line 2: “learnable parameters”, when combined with cross-entropy loss, are being interpreted as involving backpropagation, otherwise the machine learning model would not learn) the cross-entropy loss (Gosala, pg 4, column 1, paragraph before equation 2, selection reproduced below:
PNG
media_image3.png
184
736
media_image3.png
Greyscale
); through the BEV model (Gosala, Figure 1 text: “SkyEye: The first self-supervised framework for semantic BEV mapping”. “SkyEye” is being interpreted as involving BEV model) to generate the trained BEV model (Gosala, Figure 1 text: “SkyEye: The first self-supervised framework for semantic BEV mapping”. SkyEye is being interpreted as involving training, after training, a trained BEV model is generated).
Regarding claim 13, Gosala teaches The device of claim 8, wherein the one or more processors, to train the BEV model, with the cross-entropy loss, in order to generate the trained BEV model, are configured to:
adjust parameters of the BEV model (Gosala, pg 4, Section 3.3, line 2: “learnable parameters”, when combined with cross-entropy loss, are being interpreted as involving backpropagation, otherwise the machine learning model would not learn) based on the cross-entropy loss (Gosala, pg 4, column 1, paragraph before equation 2, selection reproduced below:
PNG
media_image3.png
184
736
media_image3.png
Greyscale
) and to generate the trained BEV model (Gosala, Figure 1 text: “SkyEye: The first self-supervised framework for semantic BEV mapping”. SkyEye is being interpreted as involving training, after training, a trained BEV model is generated).
Regarding claim 14, Gosala teaches The device of claim 8, wherein the BEV model is a BEV semantic segmentation network model (Gosala, Figure 1 text: “Gosala, Figure 1 text: “SkyEye: The first self-supervised framework for semantic BEV mapping”. “Semantic BEV mapping” is being interpreted as involving BEV semantic segmentation network model).
Regarding claim 16, Gosala teaches The non-transitory computer-readable medium of claim 15, wherein the one or more instructions further cause the device to:
implement the trained BEV model in a camera that captured the video data (Gosala, Section 4.1, “We evaluate SkyEye on the KITTI-360 [20] dataset”. Evaluating using the dataset is being interpreted as using the trained BEV model for evaluation. The dataset serves as the camera that captured the video data).
Claim 8 is rejected using the same rationale as applied to claim 1 and 2 discussed above.
Claim 10 is rejected using the same rationale as applied to claim 1 discussed above.
Claim 17 is rejected using the same rationale as applied to claim 5 discussed above.
Claim 18 is rejected using the same rationale as applied to claim 6 discussed above.
Claim 19 is rejected using the same rationale as applied to claim 7 discussed above.
Claim 20 is rejected using the same rationale as applied to claim 10 discussed above.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 9 is/are rejected under 35 U.S.C. 103 as being unpatentable over Gosala, in view of Wimbauer (“Behind the Scenes: Density Fields for Single View Reconstruction”, 2023, as cited in IDS filed 02/25/2025).
Regarding claim 9, Gosala teaches The device of claim 8,
However, Gosala does not appear to explicitly teach wherein the geometry model is a pretrained neural field.
Pertaining to the same field of endeavor, Wimbauer teaches
wherein the geometry model is a pretrained (Wimbauer, pg 12, column 2, Section Networks: “For encoder, we use a ResNet-50 [17] backbone pretrained
on ImageNet”) neural field (Wimbauer, “Density Field”, which is being interpreted as involving geometry. Supported by Figure 2, which uses the pretrained encoder).
Gosala and Wimbauer are considered to be analogous art because they are directed to automated driving that involves images. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method and system for automated driving that involves images with densities and geometry (as taught by Gosala) to include wherein the geometry model is a pretrained neural field (as taught by Wimbauer) because the combination provides an improvement to geometry predictions related to volume rendering (Wimbauer, Abstract).
Claim(s) 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Gosala, in view of Yuan (US 2022/0004808 A1, 2019).
Regarding claim 12, Gosala teaches The device of claim 8, wherein the one or more processors, to train the BEV model, with the cross-entropy loss, in order to generate the trained BEV, are configured to:
the cross-entropy loss (Gosala, pg 4, column 1, paragraph before equation 2, selection reproduced below:
PNG
media_image3.png
184
736
media_image3.png
Greyscale
) across multiple video frames (Gosala, Figure 2, which shows bottom left having multiple video frames) to update parameters (Gosala, pg 4, Section 3.3, line 2: “learnable parameters”, which is being interpreted as involving “update parameters”) of the BEV model (Gosala, Figure 1, “SkyEye: The first self-supervised framework for semantic BEV mapping”. Which is being interpreted as involving the BEV model ) and to generate the trained BEV model (Gosala, Figure 1, “SkyEye: The first self-supervised framework for semantic BEV mapping”. Which is being interpreted as involving generating the trained BEV model after training.).
However, Gosala does not appear to explicitly teach average the cross-entropy loss.
Pertaining to the same field of endeavor, Yuan teaches
Average… (Yuan, [0407]: “The average of the cross-entropy loss function of all the pixels in the image may be used as the loss function during the network training”)
Gosala and Yuan are considered to be analogous art because they are directed to image segmentation. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method and system for semantic segmentation (as taught by Gosala) to include average the cross-entropy loss (as taught by Yuan) because the combination provides an improvement to semantic segmentation (Yuan, [0001-0004]).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
Gosala et al (US 2025/0209724 A1, Provisional Priority of 2023) discloses BEV semantic segmentation using monocular frontal view video frames.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JOHNNY B DUONG whose telephone number is (571)272-1358. The examiner can normally be reached Monday - Thursday 10a-9p (ET).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Matthew Bella can be reached at (571)272-7778. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/J.B.D./Examiner, Art Unit 2667 /MATTHEW C BELLA/Supervisory Patent Examiner, Art Unit 2667