Prosecution Insights
Last updated: October 02, 2026
Application No. 19/060,279

USING GAUSSIAN PRIMITIVES WHEN GENERATING BIRD’S EYE VIEW REPRESENTATIONS

Non-Final OA §103§112
Filed
Feb 21, 2025
Examiner
PROTAZI, BRIGITER DIVULALE
Art Unit
2612
Tech Center
2600 — Communications
Assignee
Qualcomm Incorporated
OA Round
1 (Non-Final)
0%
Grant Probability
At Risk
1-2
OA Rounds
7m
Est. Remaining
0%
With Interview

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 1 resolved
-62.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Fast prosecutor
2y 2m
Avg Prosecution
16 currently pending
Career history
23
Total Applications
across all art units

Statute-Specific Performance

§101
6.3%
-33.7% vs TC avg
§103
66.1%
+26.1% vs TC avg
§102
15.2%
-24.8% vs TC avg
§112
12.5%
-27.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1 resolved cases

Office Action

§103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement (IDS) submitted on 06/05/2025 is being considered by the examiner. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 19 and rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claims 19 and 20 invoke 35 U.S.C. 112(f). Claim limitation “means for...” invokes 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. However, the written description fails to disclose the corresponding structure, material, or acts for performing the entire claimed function and to clearly link the structure, material, or acts to the function. Claim 19 recites “ A device for processing media data,...”, “means for receiving, from a multi-camera system of a vehicle, one or more images of a three-dimensional space around the vehicle;” “means for receiving, from a depth sensing unit of the vehicle, point cloud data representing the three-dimensional space;” “means for forming an initial bird's eye view (BEV) representation of the three- dimensional space from the one or more images and the point cloud data;” “means for generating Gaussian primitives using the initial BEV representation, the Gaussian primitives representing features of objects in the three-dimensional space;” and “means for combining the Gaussian primitives into the initial BEV representation using Gaussian splatting to form an enhanced BEV representation of the three- dimensional space.” The use of “device for” and “means for” does not recite sufficient structure for preforming the respective functions. For a means plus function limitation involving a specialized function, the specification must disclose corresponding structure sufficient to perform the claimed function which requires a disclosure of an algorithm by which the processor performs the function. Merely identifying a generic processor or device without disclosing how the processor performs the claimed function does not provide sufficient corresponding structure. The specification merely identifies processing circuitry, one or more processors, and memory as performing the recited functions without disclosing adequate algorithm for the recited claim limitations. The spec fails to disclose sufficient corresponding structure for the respective means plus function limitations, thus rendering claim 19 indefinite. Claim 20 depends on claim 19 and recites “means for generating the Gaussian primitives comprises:” “means for determining input features from at least one of the initial BEV representation or the one or more images and the point cloud data;” “means for extracting high-level feature maps from the input features;” “means for progressively downsampling the high-level feature maps to form a set of downsampled feature maps;” “means for encoding the set of downsampled feature maps to form an encoded feature map including encoded features of the objects in the three-dimensional space;” and “means for decoding the encoded feature map, including means for progressively upsampling the encoded feature map to an original resolution of the three-dimensional space and means for predicting, during the decoding, the Gaussian primitives.” Each limitation employs the term “means for” followed with functional language without reciting sufficient structure for preforming the respective functions. For a means plus function limitation involving a specialized function, the specification must disclose corresponding structure sufficient to perform the claimed function which requires a disclosure of an algorithm by which the processor performs the function. Merely identifying a generic processor or device without disclosing how the processor performs the claimed function does not provide sufficient corresponding structure. The specification merely identifies processing circuitry, one or more processors, and memory as performing the recited functions without disclosing adequate algorithm for the recited claim limitations. The spec fails to disclose sufficient corresponding structure for the respective means plus function limitations, thus rendering claim 20, which depends on claim 19, as indefinite. Therefore, the claim is indefinite and is rejected under 35 U.S.C. 112(b) or pre-AIA 35 U.S.C. 112, second paragraph. Applicant may: (a) Amend the claim so that the claim limitation will no longer be interpreted as a limitation under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph; (b) Amend the written description of the specification such that it expressly recites what structure, material, or acts perform the entire claimed function, without introducing any new matter (35 U.S.C. 132(a)); or (c) Amend the written description of the specification such that it clearly links the structure, material, or acts disclosed therein to the function recited in the claim, without introducing any new matter (35 U.S.C. 132(a)). If applicant is of the opinion that the written description of the specification already implicitly or inherently discloses the corresponding structure, material, or acts and clearly links them to the function so that one of ordinary skill in the art would recognize what structure, material, or acts perform the claimed function, applicant should clarify the record by either: (a) Amending the written description of the specification such that it expressly recites the corresponding structure, material, or acts for performing the claimed function and clearly links or associates the structure, material, or acts to the claimed function, without introducing any new matter (35 U.S.C. 132(a)); or (b) Stating on the record what the corresponding structure, material, or acts, which are implicitly or inherently set forth in the written description of the specification, perform the claimed function. For more information, see 37 CFR 1.75(d) and MPEP §§ 608.01(o) and 2181. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-3, 8, 10-12, 17 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over LIANG (No. US-20200025931-A1 “Liang”) in view of ZHOU (Zhou, X., Lin, Z., Shan, X., Wang, Y., Sun, D., & Yang, M. H. (2024, June). Drivinggaussian: Composite gaussian splatting for surrounding dynamic autonomous driving scenes. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 21634-21643). IEEE. (Year: 2024) “Zhou”) and in further view of CHABOT (Chabot, F., Granger, N., & Lapouge, G. (2024, December). Gaussianbev: 3d gaussian representation meets perception models for bev segmentation. In 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) (pp. 2250-2259). IEEE. (Year: 2024) “Chabot”). Regarding claim 1, Liang teaches “A method of processing image and depth data, the method comprising:” (methods and processes; Para 0056); “receiving, from a multi-camera system of a vehicle, one or more images of a three-dimensional space around the vehicle;” (a camera system configured to capture image data associated with an environment surrounding an autonomous vehicle; Para 0008); (image data obtained from one or more cameras; Para 0038); “receiving, from a depth sensing unit of the vehicle, point cloud data representing the three-dimensional space;” (a LIDAR system configured to capture LIDAR point cloud data associated with the environment surrounding the autonomous vehicle; Para 0008); “forming an initial bird's eye view (BEV) representation of the three-dimensional space from the one or more images and the point cloud data;” (the LIDAR point cloud data to a bird's eye view representation of the LIDAR point cloud data; Para 0006); (to generate a feature map comprising the fused image features and the LIDAR features; Para 0011); (fuse image features from image data (e.g., image data captured by a camera system within an autonomous vehicle) with LIDAR features from LIDAR point cloud data (e.g., LIDAR point cloud data captured by a LIDAR system within an autonomous vehicle); Para 0045); (implementing data transformations and object detection using a bird's eye view (BEV) representation of the data (e.g., sensor data, map data, etc.). .... BEV representations of sensor data can advantageously retain the metric space associated with raw sensor data (e.g., image data captured by cameras and/or LIDAR point cloud data captured by a LIDAR system); Para 0064); Liang discloses a multi sensor system that contains cameras that capture the environment around a vehicle, both images and depth data of the vehicle. Then use the fused image and point cloud data into a BEV representation through a fused feature map. However, while Liang fails to teach “generating Gaussian primitives using the initial BEV representation, the Gaussian primitives representing features of objects in the three-dimensional space; and combining the Gaussian primitives with the initial BEV representation using Gaussian splatting to form an enhanced BEV representation of the three-dimensional space”. Zhou teaches “generating Gaussian primitives using the initial BEV representation, the Gaussian primitives representing features of objects in the three-dimensional space; and” (we first sequentially and progressively model the static background of the entire scene with incremental static 3D Gaussians. We then leverage a composite dynamic Gaussian graph to handle multiple moving objects, individually reconstructing each object and restoring their accurate positions and occlusion relationships within the scene; Abstract); Zhou discloses Gaussian primitives that can be generated to model the vehicle environment. The Gaussian representation then disclose dynamic objects that include the positions and relationships to the driving environment. Thus, teaching the Gaussian primitives represent object features in a 3D space. However, while Liang and Zhou fail to teach “combining the Gaussian primitives with the initial BEV representation using Gaussian splatting to form an enhanced BEV representation of the three-dimensional space”. Chabot teaches “combining the Gaussian primitives with the initial BEV representation using Gaussian splatting to form an enhanced BEV representation of the three-dimensional space.” (This representation is then splattered in BeV using a rasterizer module; Pg.2 Col 1, Para 3); (The BeV rasterizer module is used to obtain the BeV feature map B ∈ RHB×WB×C from the set of gaussians G predicted by the 3D gaussian generator. To this end, the differentiable rasterization process proposed in gaussian splatting; Pg.5, Col 1, Para 3.3); (enables fine 3D modeling of scenes. .... This representation is then rendered into a BeV feature map; Pg.2, Col 1, Para 2); Chabot discloses 3D Gaussian representation, Gaussian splatting and BEV feature map. Chabot identifies BEV details as the problem and uses Gaussian scene representations/splatting to get a more represented BEV. It would be obvious to one skilled in the art to modify the camera system of Liang with the teaching of Zhou and Chabot to generate a Gaussian representation of the scene features and splat the Gaussian representation into the BEV representation. Liang, Zhou and Chabot are analogous art as they are related to autonomous vehicles and 3D environment. The motivation for the above is to have an accurate BEV for a vehicle for better autonomous vehicle usage. Therefore, it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Liang by generating Gaussian primitives using the initial BEV representation, the Gaussian primitives representing features of objects in the three-dimensional space as taught by Zhou and by combining the Gaussian primitives with the initial BEV representation using Gaussian splatting to form an enhanced BEV representation of the three-dimensional space as taught by Chabot. Regarding claim 2, while Liang fails to teach the limitation, Zhou teaches “The method of claim 1, wherein generating the Gaussian primitives comprises, for each of the Gaussian primitives, calculating a mean value and covariance data.” (µ is the mean of the LiDAR points; Σ ∈ R3×3 is an anisotropic covariance matrix; and.... parameters of the Gaussian model, including position P(x,y,z), covariance matrix Σ; Pg.4, Col 3.1); Zhou discloses a mean value in the LIDAR and a covariance matrix. The motivation for the above is to have accurate calculations of Gaussian primitives for better generation of BEV. Therefore, it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Liang by generating the Gaussian primitives comprises, for each of the Gaussian primitives, calculating a mean value and covariance data as taught by Zhou. Regarding claim 3, Liang teaches “The method of claim 2, wherein forming the initial BEV representation comprises performing feature extraction using the sensor input to extract features representing positions of objects in the three-dimensional space, and” (cameras to perform very accurate localization of objects within three-dimensional space relative to an autonomous vehicle; Para 0036); (sensor data associated with the autonomous vehicle's surrounding environment as well as the position; Para 0038); (fuse image features from image data ... with LIDAR features from LIDAR point cloud data; Para 0045); Liang discloses camera sensor inputs that lead to feature extraction based on the inputs and the fused features lead to the location/position of the objects in the 3D space. However, Liang fails to teach “wherein generating the Gaussian primitives comprises, for each of the features, calculating the mean value as a central position of the feature and the covariance data as a representation of uncertainty around the mean value”. Zhou teaches “wherein generating the Gaussian primitives comprises, for each of the features, calculating the mean value as a central position of the feature and the covariance data as a representation of uncertainty around the mean value.” (µ is the mean of the LiDAR points; Σ ∈ R3×3 is an anisotropic covariance matrix; and.... parameters of the Gaussian model, including position P(x,y,z), covariance matrix Σ; Pg.4, Col 3.1); Zhou discloses calculating a mean and anisotropic covariance for the corresponding Gaussian. It would be obvious to a person skilled in the art to understand that the mean value would be the central position of the feature as it is obvious the mean is spatially the average, thus the position would be the average central position of the features. And it would be obvious that the covariance, by its regular statistically meaning, would be an uncertainty around the mean. The motivation for the above is to have accurate feature extraction and calculations for accurate generation of Gaussian primitives. Therefore, it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Liang by generating the Gaussian primitives comprises, for each of the features, calculating the mean value as a central position of the feature and the covariance data as a representation of uncertainty around the mean value as taught by Zhou. Regarding claim 8, Liang further teaches “The method of claim 1, further comprising at least partially autonomously controlling the vehicle using the enhanced BEV representation.” (controlling motion of an autonomous vehicle (e.g., vehicle 102 of FIG. 1) based at least in part on the motion plan determined; Para 0159); (Additional technical effects and benefits can be achieved by implementing data transformations and object detection using a bird's eye view (BEV) representation of the data (e.g., sensor data, map data, etc.); Para 0064); Claim 10 is directed at a device and its limitations are similar in scope and functions performed by the effect processing method of claim 1. Therefore, claim 10 limitations are also rejected with the same rationale as regarding claim 1. Claim 11 is directed at a device and its limitations are similar in scope and functions performed by the effect processing method of claim 2. Therefore, claim 11 limitations are also rejected with the same rationale as regarding claim 2. Claim 12 is directed to a device and its limitations are similar in scope and functions performed by the effect processing method of claim 3. Therefore, claim 12 limitations are also rejected with the same rationale as regarding claim 3. Claim 17 is directed to a device and its limitations are similar in scope and functions performed by the effect processing method of claim 8. Therefore, claim 17 limitations are also rejected with the same rationale as regarding claim 8. Claim 19 is directed at a device and its limitations are similar in scope and functions performed by the effect processing method of claim 1. Therefore, claim 19 limitations are also rejected with the same rationale as regarding claim 1. Claim(s) 4 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over LIANG in view of ZHOU and in further view of CHABOT and in further view of CHENG (Cheng, K., Long, X., Yang, K., Yao, Y., Yin, W., Ma, Y., ... & Chen, X. (2024, July). Gaussianpro: 3d gaussian splatting with progressive propagation. In Forty-first international conference on machine learning. (Year: 2024) “Cheng”). Regarding claim 4, while Liang, Zhou and Chabot fail to teach claim 4, Cheng teaches “The method of claim 2, wherein generating the Gaussian primitives comprises, for each region of a set of regions of the three-dimensional space: determining a local complexity of the region based on one or more characteristics associated with at least one of the region, the multi-camera system, the depth sensing unit, or a context of the vehicle;” (When tackling large-scale scenes that unavoidably contain texture-less surfaces, SfM techniques fail to produce enough points in these surfaces and cannot provide good initialization for 3DGS... the existing reconstructed geometries of the scene and utilizes patch matching to produce new Gaussians; Pg.1 Abstract); “determining a subset of the features in the region; and” (filter out the unreliable propagated depths and normals using geometric consistency, yielding filtered depths and filtered normals; Pg.4 Fig 2 Desc); “for each of the subset of the features in the region, adapting the covariance data for the features according to the local complexity of the region.” (The covariance matrix Σi of a 3D Gaussian can be considered as representing an ellipsoid, where the eigenvectors of Ri correspond to the three axes of the ellipsoid and the scale factors of Si refer to the axis lengths; Pg.4 Col 2 Para 3); Cheng discloses local scene characteristics determining when and how additional Gaussian representations are needed. As well as filtering and selection of useful features for Gaussian refinement which teaches the claimed subject matter of a subset of features. In addition, Cheng discloses altering the Gaussian scale/refinement which alters the Gaussian covariance representation. It would be obvious to a person skilled in the art to implement Zhou’s Gaussian representation with Cheng’s adaptive geometric selection and refinement technique to the features obtained from Liang’s vehicle sensors. It would be obvious to adapt the covariance of selected Gaussian features according to the geometric complexity of the region and allocate Gaussian representations to the regions requiring greater detail while also avoiding unnecessary complexity in other regions. Liang, Zhou, Chabot and Cheng are analogous art as they are related to autonomous vehicles and 3D environment. The motivation for the above is to have accurate generation of Gaussian primitives for better determination of region in the 3D space. Therefore, it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Liang, Zhou and Chabot by determining a local complexity of the region based on one or more characteristics associated with at least one of the region, the multi-camera system, the depth sensing unit, or a context of the vehicle, determining a subset of the features in the region and for each of the subset of the features in the region, adapting the covariance data for the features according to the local complexity of the region as taught by Cheng. Claim 13 is directed to a device and its limitations are similar in scope and functions performed by the effect processing method of claim 4. Therefore, claim 13 limitations are also rejected with the same rationale as regarding claim 4. Claim(s) 5, 14 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over LIANG in view of ZHOU and in further view of CHABOT and in further view of LU (Lu, L., Gao, H., Dai, T., Zha, Y., Hou, Z., Wu, J., & Xia, S. T. (2024, October). Large point-to-gaussian model for image-to-3d generation. In Proceedings of the 32nd ACM International Conference on Multimedia (pp. 10843-10852). (Year: 2024) “Lu”). Regarding claim 5, Liang further teaches “The method of claim 1, wherein generating the Gaussian primitives comprises: determining input features from at least one of the initial BEV representation or the one or more images and the point cloud data;” (to generate a feature map comprising the fused image features and the LIDAR features; Para 0011); (fuse image features from image data (e.g., image data captured by a camera system within an autonomous vehicle) with LIDAR features from LIDAR point cloud data (e.g., LIDAR point cloud data captured by a LIDAR system within an autonomous vehicle); Para 0045); While Liang and Zhou fail to teach “extracting high-level feature maps from the input features”. Chabot teaches “extracting high-level feature maps from the input features;” (extracts image features using an image backbone and a neck to obtain feature maps; Pg.3, Col 3.1); While Liang, Zhou and Chabot fail to teach “progressively downsampling the high-level feature maps to form a set of downsampled feature maps; encoding the set of downsampled feature maps to form an encoded feature map including encoded features of the objects in the three-dimensional space; and decoding the encoded feature map, including progressively upsampling the encoded feature map to an original resolution of the three-dimensional space and, during the decoding, predicting the Gaussian primitives”. Lu teaches “progressively downsampling the high-level feature maps to form a set of downsampled feature maps;” (During downsampling, the number of point clouds progressively decreases, and the current layer’s point cloud is derived via farthest point sampling (FPS) from the preceding shallower layer, thereby generating multi-scale point cloud features and expanding the receptive field; Pg.10847, Col 3.2.2); “encoding the set of downsampled feature maps to form an encoded feature map including encoded features of the objects in the three-dimensional space; and” (The Gaussian decoder’s architecture adopts a U-Net structure, akin to that described in ... During downsampling.... generating multi-scale point cloud features; Pg.10847, Col 3.2.2); “decoding the encoded feature map, including progressively upsampling the encoded feature map to an original resolution of the three-dimensional space and, during the decoding, predicting the Gaussian primitives.” (Multi-Scale Gaussian Decoder. The Gaussian decoder’s architecture adopts a U-Net structure; Pg.10847, Col 3.2.2); (After obtaining the enhanced features, we introduce the multi linear heads, which utilizes multiple decoder heads for different attributes in the Gaussian; Pg.10847, Col 3.2.2); Chabot discloses feature extraction into feature maps. While Lu discloses the encoder progressively down sampling the 3D representation and producing features. The U-Net encoder progressively produces encode features representing the 3D point cloud data. Lu also discloses a decoder in an encoder decoder architecture. The U-Net architecture reverses the point cloud down sampling through the decoder and Lu discloses the architecture as an encoder decoder in which the decoder reconstructs Gaussian attributes at different resolutions. It would be obvious to a person skilled in the art to implement the Gaussian generator of Liang, Zhou and Chabot by using the encoder decoder architecture of Lu because Lu teaches progressively down sampling 3D point cloud features to obtain multi scale features and using the decoder to predict Gaussian attributes. Liang, Zhou, Chabot and Cheng are analogous art as they are related to autonomous vehicles and 3D environment. The motivation for the above is to have high level information regarding BEV representation and feature maps for better generation of Gaussian primitives. Therefore, it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Liang and Zhou by extracting high-level feature maps from the input features as taught by Chabot and by progressively downsampling the high-level feature maps to form a set of downsampled feature maps, encoding the set of downsampled feature maps to form an encoded feature map including encoded features of the objects in the three-dimensional space and decoding the encoded feature map, including progressively upsampling the encoded feature map to an original resolution of the three-dimensional space and, during the decoding, predicting the Gaussian primitives as taught by Lu. Claim 14 is directed to a device and its limitations are similar in scope and functions performed by the effect processing method of claim 5. Therefore, claim 14 limitations are also rejected with the same rationale as regarding claim 5. Claim 20 is directed to a device and its limitations are similar in scope and functions performed by the effect processing method of claim 5. Therefore, claim 20 limitations are also rejected with the same rationale as regarding claim 5. Claim(s) 6 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over LIANG in view of ZHOU and in further view of CHABOT and in further view of LU in view of JIANG (Jiang, H., Liu, L., Cheng, T., Wang, X., Lin, T., Su, Z., ... & Wang, X. (2024). Gausstr: Foundation model-aligned gaussian transformer for self-supervised 3d spatial understanding. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 11960-11970). IEEE. (Year: 2024) “Jiang”). Regarding claim 6, while Liang, Zhou, Chabot and Lu fail to teach claim 6, Jiang teaches “The method of claim 5, wherein decoding the encoded feature map comprises decoding, by one or more self-attention layers of a decoder, the method further comprising using previously determined Gaussian primitives as context in the self- attention layers of the decoder.” (subsequent Transformer layers, GaussTR re fines Gaussian queries... GaussTR then applies self-attention [45] across Gaussian queries for 3D modeling and capturing contextual semantics across the scene, where 3D positional encoding PE of Gaussian means; Pg.4 Col 1); Jiang discloses self-attention within the Gaussian transformer which corresponds to a decoder. Each Gaussian query attends other Gaussian queries so that Gaussian representations have contextual information for refinement. The Gaussians determent from the earlier layer are carried forward and refined by subsequent transformer layers. It would be obvious to a person skilled in the art to implement the Gaussian decoder taught in claim 5 by Lu with Jiang’s transformer architecture so to represent the Gaussian primitives as Gaussian queries and apply self-attention across those queries during decoder layers. Liang, Zhou, Chabot, Lu and Jiang are analogous art as they are related to autonomous vehicles and 3D environment. The motivation for the above is to have accurate decoding of feature maps for better determination of Gaussian primitives. Therefore, it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Liang, Zhou, Chabot and Lu by decoding the encoded feature map comprises decoding, by one or more self-attention layers of a decoder, the method further comprising using previously determined Gaussian primitives as context in the self- attention layers of the decoder as taught by Jiang. Claim 15 is directed to a device and its limitations are similar in scope and functions performed by the effect processing method of claim 6. Therefore, claim 15 limitations are also rejected with the same rationale as regarding claim 6. Claim(s) 7 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over LIANG in view of ZHOU and in further view of CHABOT and in further view of LU in view of JIANG and further view of ZHANG (Zhang, Z., Song, T., Lee, Y., Yang, L., Peng, C., Chellappa, R., & Fan, D. (2024). Lp-3dgs: Learning to prune 3d gaussian splatting. Advances in Neural Information Processing Systems, 37, 122434-122457. (Year: 2024) “Zhang”). Regarding claim 7, while Liang, Zhou, Chabot and Lu fail to teach the limitation, Jiang teaches “modulating, by the one or more self-attention layers, the input features using the binary mask.” (followed by global self-attention across sparse Gaussian queries for effective 3D modeling; Pg.2 Col 1); However, while Liang, Zhou, Chabot, Lu, Jiang fail to teach “calculating a binary mask from the Gaussian primitives based on a threshold for the Gaussian primitives”. Zhang teaches “The method of claim 6, further comprising: calculating a binary mask from the Gaussian primitives based on a threshold for the Gaussian primitives; and” (LP-3DGS learns a trainable mask upon a previously defined importance metric to compress the number of Gaussians .... Specifically, to learn a binary mask, we first initialize a real-value mask mi for each point i, and then adopt the Gumbel-sigmoid technique to binarize the mask value; Pg.5 Col 3.3); (The binarization operation for real-value mask in pruning typically involves a hard threshold function, determining the binary mask should be 0 or 1; Pg.5 Col 3.3); Zhang discloses a mask associated with Gaussian primitives and determines the Gaussian information that is retained. As well as binary masks associated with the Gaussian points. Zhang also connects a threshold with the Gaussian associated mask into binary values. Jiang discloses self-attention across Gaussian representations. Binary masking of attention is already a conventional transformer mechanism. Zhang applies it teaching of binary Gaussian mask to determine which Gaussian information is processed. The masked attention used binary masks to determine which features partake in attention. Liang, Zhou, Chabot, Lu, Jiang and Zhang are analogous art as they are related to autonomous vehicles and 3D environment. The motivation for the above is to have accurate calculation for better generation or Gaussian primitives. Therefore, it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Liang, Zhou, Chabot, and Lu by modulating, by the one or more self-attention layers, the input features using the binary mask as taught by Jiang and by calculating a binary mask from the Gaussian primitives based on a threshold for the Gaussian primitives as taught by Zhang. Claim 16 is directed at a device and its limitations are similar in scope and functions performed by the effect processing method of claim 7. Therefore, claim 16 limitations are also rejected with the same rationale as regarding claim 7. Claim(s) 9 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over LIANG in view of ZHOU and in further view of CHABOT and in further view of CAI (Cai, H., Zhang, Z., Zhou, Z., Li, Z., Ding, W., & Zhao, J. (2023). Bevfusion4d: Learning lidar-camera fusion under bird's-eye-view via cross-modality guidance and temporal aggregation. arXiv preprint arXiv:2303.17099. (Year: 2023) “Cai”) in view of ZUO (Zuo, S., Zheng, W., Huang, Y., Zhou, J., & Lu, J. (2024, Dec). Gaussianworld: Gaussian world model for streaming 3d occupancy prediction. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 6772-6781). IEEE. (Year: 2024) “Zuo”). Regarding claim 9, while Liang, Zhou and Chabot fail to teach claim 9, Cai teaches “The method of claim 1, further comprising updating the enhanced BEV representation over time based on updated images from the multi-camera system,” (we extend our framework into the temporal domain with our proposed Temporal Deformable Alignment (TDA) module, which aims to aggregate BEV features from multiple historical frames; Pg 1 Abstract); (LiDAR BEV for extracting image features across multiple camera views; Pg 1 Abstract); “updated point cloud data from the depth sensing unit, and” (Integrating LiDAR and Camera information into Bird’s Eye-View (BEV) ... point clouds provide more accurate localization and geometry information; Pg 1 Abstract); However, while Liang, Zhou, Chabot and Cai fail to teach “updated Gaussian primitives from the updated images and the updated point cloud data”. Zuo teaches “updated Gaussian primitives from the updated images and the updated point cloud data.” (reformulate 3D occupancy prediction .... decompose the scene evolution into three factors: 1) ego motion alignment of static scenes; 2) local movements of dynamic objects; and 3) completion of newly observed scenes.... infer the scene evolution in the 3D Gaussian space considering the current RGB observation; Pg.1 Abstract); Cai discloses a BEV pipeline with image data and point-cloud data being fused temporally. Zuo discloses treating Gaussian scene as temporal rather than static, which teaches updating the scene representation to a current version. Zuo discloses the Gaussian scene changing to reflect vehicle motion, moving objects, etc., thus the Gaussian representation temporally evolved and is updating. In combination with Cui teaching of processing both images and point cloud data information as well as in combination with claim 1 Gaussian generation showcases images and point cloud data to be reflected in an updated Gaussian. It would be obvious for a person skilled in the art to modify the BEV system of claim 1 with the temporal teachings of Cui and Zuo so that the enhanced BEV and Gaussian primitives are updated as new image and point cloud data are received. Liang, Zhou, Chabot, Cai and Zuo are analogous art as they are related to autonomous vehicles and 3D environment. The motivation for the above is to have accurate updated of data received for better processing and representation of 3D environment of vehicle. Therefore, it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Liang, Zhou and Chabot by updating the enhanced BEV representation over time based on updated images from the multi-camera system, updated point cloud data from the depth sensing unit as taught by and by updated Gaussian primitives from the updated images and the updated point cloud data as taught by Cai. Claim 18 is directed to a device and its limitations are similar in scope and functions performed by the effect processing method of claim 9. Therefore, claim 18 limitations are also rejected with the same rationale as regarding claim 9. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. US-20250054286-A1 (Huang) – Discloses using images obtained from one or more sensors onboard a vehicle, a 2-dimensional (2D) feature extraction; performing, a 3-dimensional (3D) feature extraction on the images; detecting objects in the images by fusing detection results from the 2D feature extraction and the 3D feature extraction. US-20240312218-A1 (Krehl) – Discloses vehicle computing system can receive raw sensor data in both a traditional sensor data processing module and a learned sensor data processing module. Each module can reproject sensor data in BEV space and can optionally perform sensor fusion when multiple sensor data types are processed. CN-119360166-A (Tong) – Discloses a multi-sensor BEV fusion method based on image corrosion improvement and belongs to the technical field of multi-sensor fusion. The method comprises the steps of imagery carrying out feature extraction and BEV conversion on data of two sensors of an image and a radar and then processing the BEV features of the image by utilizing an image corrosion technology to remove noise and refine the outline of an object. Sathyam, R., & Li, Y. (2025). Foundation models for autonomous driving perception: A survey through core capabilities. IEEE Open Journal of Vehicular Technology. (Year: 2025) – Discloses how these models address critical challenges in autonomous perception, including limitations in generalization, scalability, and robustness to distributional shifts. And introduces a novel taxonomy structured around four essential capabilities for robust performance in dynamic driving environments: generalized knowledge, spatial understanding, multi-sensor robustness, and temporal reasoning. Kerbl, B., Kopanas, G., Leimkühler, T., & Drettakis, G. (2023). 3d gaussian splatting for real-time radiance field rendering. ACM Trans. Graph., 42(4), 139-1. (Year: 2023) – Discloses Gaussian Splatting for Real-Time Radiance Field Rendering. Unger, D., Gosala, N., Kumar, V. R., Borse, S., Valada, A., & Yogamani, S. (2024). Multi-camera bird’s eye view perception for autonomous driving. Computer Vision: Challenges, Trends, and Opportunities, 279. (Year: 2024) – Discloses he contemporary trends of multi-camera-based DNN (deep neural network) models outputting object representations directly in the BEV space. And transform camera images into BEV space using geometric constraints implicitly or explicitly in the network. Any inquiry concerning this communication or earlier communications from the examiner should be directed to BRIGITER D PROTAZI whose telephone number is (571)272-7995. The examiner can normally be reached Monday - Friday 7:30-5. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Said A Broome can be reached at 5712722931. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /B.D.P./Examiner, Art Unit 2612 /Said Broome/Supervisory Patent Examiner, Art Unit 2612
Read full office action

Prosecution Timeline

Feb 21, 2025
Application Filed
Sep 03, 2026
Non-Final Rejection mailed — §103, §112 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
0%
Grant Probability
0%
With Interview (+0.0%)
2y 2m (~7m remaining)
Median Time to Grant
Low
PTA Risk
Based on 1 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month