Prosecution Insights
Last updated: October 02, 2026
Application No. 18/606,470

POSITIONAL ENCODING FOR POINT CLOUD COMPRESSION

Non-Final OA §103
Filed
Mar 15, 2024
Examiner
SARMA, ABHISHEK
Art Unit
2621
Tech Center
2600 — Communications
Assignee
Qualcomm Incorporated
OA Round
1 (Non-Final)
84%
Grant Probability
Favorable
1-2
OA Rounds
0m
Est. Remaining
84%
With Interview

Examiner Intelligence

Grants 84% — above average
84%
Career Allowance Rate
492 granted / 589 resolved
+21.5% vs TC avg
Minimal +0% lift
Without
With
+0.3%
Interview Lift
resolved cases with interview
Fast prosecutor
2y 1m
Avg Prosecution
24 currently pending
Career history
615
Total Applications
across all art units

Statute-Specific Performance

§101
4.3%
-35.7% vs TC avg
§103
76.5%
+36.5% vs TC avg
§102
8.1%
-31.9% vs TC avg
§112
5.3%
-34.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 589 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . In the response to this Office Action, the Examiner respectfully requests that support be shown for language added to any original claims on amendment and any new claims. That is, indicate support for newly added claim language by specifically pointing to page(s) and line numbers in the specification and/or drawing figure(s). This will assist the Examiner in prosecuting this application. Election/Restrictions Applicant's election with traverse of Group II: in the reply filed on 06/16/2026 is acknowledged and is found persuasive. Therefore, the restriction requirement of 04/29/2026 is withdrawn. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-15, 17-33, 35-51, 53-69, and 71-72 are rejected under 35 U.S.C. 103 as being unpatentable over U.S. Patent Application Publication 2025/0139834 A1 to Pang et al. (hereinafter "Pang") in view of Non-Patent Publication "PCGFormer: Lossy Point Cloud Geometry Compression via Local Self-Attention" to Liu et al. (hereinafter "Liu", included in IDS provided by Applicant). Regarding Claims 1 and 19, Pang teaches an apparatus and a method configured to encode a point cloud, the apparatus comprising: one or more memories configured to store point cloud data; and processing circuitry in communication with the one or more memories, the processing circuitry configured to: receive a frame of point cloud data; encode coordinates of a geometry of the frame of point cloud data using a deep learning network (Figs. 1-5; Claims 16, 42; Para. 30-84 of Pang; base point cloud encoder (320)… feature encoder (370)… For an input point cloud PC0 to be compressed, it is first quantized using a quantizer (310) with a step size s (s>1). In one embodiment, for every point, say, A in PC0 with 3D coordinates (x, y, z), the quantizer divides the coordinates by s then converts them to integers, leading to ([x/s], [y/s], [z/s]), where the function [⋅] can be the floor, ceiling or rounding function. The quantizer then removes the duplicate points with the same 3D coordinates, that is, if there exist several points having the same coordinates, we only keep one of them and remove the rest, then we get PC1. The obtained quantized point cloud, PC1, is then compressed with a base point cloud encoder (320), which outputs a bitstream BS0), wherein the deep learning network includes one or more layers configured to generate first features for the coordinates of the geometry, and wherein the deep learning network further includes at least one positional encoding layer (Figs. 1-5; Claims 16, 42; Para. 30-84 of Pang; base point cloud encoder (320)… feature encoder (370)… For an input point cloud PC0 to be compressed, it is first quantized using a quantizer (310) with a step size s (s>1). In one embodiment, for every point, say, A in PC0 with 3D coordinates (x, y, z), the quantizer divides the coordinates by s then converts them to integers, leading to ([x/s], [y/s], [z/s]), where the function [⋅] can be the floor, ceiling or rounding function. The quantizer then removes the duplicate points with the same 3D coordinates, that is, if there exist several points having the same coordinates, we only keep one of them and remove the rest, then we get PC1. The obtained quantized point cloud, PC1, is then compressed with a base point cloud encoder (320), which outputs a bitstream BS0… feature codec can be implemented based on a deep neural network). Pang does not explicitly disclose at least one positional encoding layer configured to generate second features for the coordinates of the geometry and combine the first features with the second features to generate higher-dimensional features; and output an output tensor comprising encoded coordinates and the higher-dimensional features. However, Liu teaches at least one positional encoding layer configured to generate second features for coordinates of geometry and combine first features with the second features to generate higher-dimensional features; and output an output tensor comprising encoded coordinates and the higher-dimensional features (Figs. 1-2; Abstract; Section II-III of Liu; For the encoding of PCGFormer, input PCGs are gradually downscaled with a factor of 2 at each axis in a Cartesian coordinate system. Here three layers of spatial downscaling are exemplified to gradually analyze and embed neighborhood information. An Inception-Residual Network (IRN) layer (see Fig. 2a) and a Transformer layer using local self-attention upon k nearest neighbors (a.k.a., kNN self-attention) are stacked as modular components at each scale for feature extraction and aggregation)… proposed PCGFormer improves the ConvMST… like rules-based G-PCC and popular learning-based methods, our PCGFormer deals with voxelized PCGs… apply hierarchical binary cross-entropy (BCE) loss to guide end-to-end learning… at the beginning of each Trans-former module, k points of the current point including their coordinates and features are collected and combined as F′′X of N×k×d for the next step). Therefore, at the time when the invention was filed, it would have been obvious to a person of ordinary skill in the art to include at least one positional encoding layer configured to generate second features for the coordinates of the geometry and combine the first features with the second features to generate higher-dimensional features; and output an output tensor comprising encoded coordinates and the higher-dimensional features IC using the teachings of Liu in order to modify the device taught by Pang. The motivation to combine these analogous arts would have been to provide a framework of multiscale sparse tensor representation, with which spatial neighbors can be effectively aggregated and embedded for better compression performance (Section 5 of Liu). Regarding Claims 2 and 20, the combination of Pang and Liu teaches that the deep learning network includes one or more layers configured to downscale the coordinates of the geometry to generate downscaled coordinates, and wherein the output tensor comprises the downscaled coordinates and the higher-dimensional features (Figs. 1-2; Abstract; Section II-III of Liu; For the encoding of PCGFormer, input PCGs are gradually downscaled with a factor of 2 at each axis in a Cartesian coordinate system. Here three layers of spatial downscaling are exemplified to gradually analyze and embed neighborhood information. An Inception-Residual Network (IRN) layer (see Fig. 2a) and a Transformer layer using local self-attention upon k nearest neighbors (a.k.a., kNN self-attention) are stacked as modular components at each scale for feature extraction and aggregation). Regarding Claims 3 and 21, the combination of Pang and Liu teaches that the at least one position encoding layer of the deep learning network includes a first positional encoding layer that operates on an input of the deep learning network and a second positional encoding layer that operates on the downscaled coordinates (Figs. 1-2; Abstract; Section II-III of Liu; For the encoding of PCGFormer, input PCGs are gradually downscaled with a factor of 2 at each axis in a Cartesian coordinate system. Here three layers of spatial downscaling are exemplified to gradually analyze and embed neighborhood information. An Inception-Residual Network (IRN) layer (see Fig. 2a) and a Transformer layer using local self-attention upon k nearest neighbors (a.k.a., kNN self-attention) are stacked as modular components at each scale for feature extraction and aggregation). Regarding Claims 4 and 22, the combination of Pang and Liu teaches that the one or more layers of the deep learning network includes three layers configured to progressively downscale the coordinates of the geometry to generate first downscaled coordinates, second downscaled coordinates, and third downscaled coordinates, and wherein the output tensor comprises the third downscaled coordinates that are three-times downscaled and the higher-dimensional features (Figs. 1-2; Abstract; Section II-III of Liu; For the encoding of PCGFormer, input PCGs are gradually downscaled with a factor of 2 at each axis in a Cartesian coordinate system. Here three layers of spatial downscaling are exemplified to gradually analyze and embed neighborhood information. An Inception-Residual Network (IRN) layer (see Fig. 2a) and a Transformer layer using local self-attention upon k nearest neighbors (a.k.a., kNN self-attention) are stacked as modular components at each scale for feature extraction and aggregation). Regarding Claims 5 and 23, the combination of Pang and Liu teaches that the at least one position encoding layer of the deep learning network includes a first positional encoding layer that operates on an input of the deep learning network, a second positional encoding layer that operates on the first downscaled coordinates, a third positional encoding layer that operates on the second downscaled coordinates, and a fourth positional encoding layer that operates on the third downscaled coordinates (Figs. 1-2; Abstract; Section II-III of Liu; For the encoding of PCGFormer, input PCGs are gradually downscaled with a factor of 2 at each axis in a Cartesian coordinate system. Here three layers of spatial downscaling are exemplified to gradually analyze and embed neighborhood information. An Inception-Residual Network (IRN) layer (see Fig. 2a) and a Transformer layer using local self-attention upon k nearest neighbors (a.k.a., kNN self-attention) are stacked as modular components at each scale for feature extraction and aggregation). Regarding Claims 6 and 24, the combination of Pang and Liu teaches that at least one of the second positional encoding layer, the third positional encoding layer, or the fourth positional encoding layer is positioned between a layer configured to downscale the coordinates of the geometry and an inception-residual block (Figs. 1-2; Section III of Liu; IRN is the inception-residual network). Regarding Claims 7 and 25, the combination of Pang and Liu teaches that the at least one positional encoding layer comprises a learned positional encoding layer (Figs. 1-2; Section III of Liu; like rules-based G-PCC and popular learning-based methods, our PCGFormer deals with voxelized PCGs). Regarding Claims 8 and 26, the combination of Pang and Liu teaches that the learned positional encoding layer is configured as a convolutional layer, a fully connected layer, or a multilayer perceptron (Figs. 1-2; Section III of Liu; encoded using G- PCC Octree codec). Regarding Claims 9 and 27, the combination of Pang and Liu teaches that the deep learning network includes a plurality of positional encoding layers, and wherein each of the positional encoding layers is configured to use different sets of weights (Figs. 1-2; Section III of Liu; For every point, an attention map will be calculated. Through applying the attention map to k local neighbor features, we attain the attention weighted feature at a size of 1 × d of each point). Regarding Claims 10 and 28, the combination of Pang and Liu teaches that the deep learning network includes a plurality of positional encoding layers, and wherein each of the positional encoding layers is configured to use a same set of weights (Figs. 1-2; Section II-III of Liu; fixed-weightkernels). Regarding Claims 11 and 29, the combination of Pang and Liu teaches that the at least one positional encoding layer comprises one of a sine and cosine positional encoding layer, a linear positional encoding layer, a radial basis function positional encoding layer, a Fermi-Dirac positional encoding layer, a hybrid positional encoding layer, a relative positional encoding layer, or an image grid positional encoding layer (Figs. 1-2; Section III of Liu; encoded using G- PCC Octree codec). Regarding Claims 12 and 30, the combination of Pang and Liu teaches that the deep learning network includes a plurality of positional encoding layers, and wherein the plurality of positional encoding layers include at least two types of positional encoding (Figs. 3-5; Para. 62-84 of Pang; base point cloud encoder (320)… feature encoder (370)… feature codec can be implemented based on a deep neural network… Figs. 1-2; Section II-III of Liu; proposed PCGFormer improves the ConvMST… like rules-based G-PCC and popular learning-based methods, our PCGFormer deals with voxelized PCGs… apply hierarchical binary cross-entropy (BCE) loss to guide end-to-end learning… at the beginning of each Trans-former module, k points of the current point including their coordinates and features are collected and combined as F′′X of N×k×d for the next step). Regarding Claims 13 and 31, the combination of Pang and Liu teaches that to combine the first features with the second features to generate higher-dimensional features, the positional encoding layer is configured to concatenate the first features with the second features (Figs. 1-2; Section II-III of Liu; proposed PCGFormer improves the ConvMST… like rules-based G-PCC and popular learning-based methods, our PCGFormer deals with voxelized PCGs… apply hierarchical binary cross-entropy (BCE) loss to guide end-to-end learning… at the beginning of each Trans-former module, k points of the current point including their coordinates and features are collected and combined as F′′X of N×k×d for the next step). Regarding Claims 14 and 32, the combination of Pang and Liu teaches that the first features have a feature size of 64 and wherein the higher-dimensional features have a feature size of 192 (Figs. 1-2; Section II-III of Liu; proposed PCGFormer improves the ConvMST… like rules-based G-PCC and popular learning-based methods, our PCGFormer deals with voxelized PCGs… apply hierarchical binary cross-entropy (BCE) loss to guide end-to-end learning… at the beginning of each Trans-former module, k points of the current point including their coordinates and features are collected and combined as F′′X of N×k×d for the next step… Selecting a first feature size of 64 and a higher-dimensional feature size of 192 would only require routine skill for a person of ordinary skill in the art based on the teachings of Liu and can be accomplished without any undue experimentation and it has been held that where the general conditions of a claim are disclosed in the prior art, discovering the optimum or workable ranges involves only routine skill in the art. In re Aller, 105 USPQ 233). Regarding Claims 15 and 33, the combination of Pang and Liu teaches that the processing circuitry is further configured to: encode attributes of the frame of point cloud data using the deep learning network, wherein the deep learning network includes one or more layers configured to generate third features for the attributes, and wherein the deep learning network further includes at least one positional encoding layer configured to generate fourth features from the attributes and combine the third features with the fourth features to generate second higher-dimensional features (Figs. 1-2; Section II-III of Liu; proposed PCGFormer improves the ConvMST… like rules-based G-PCC and popular learning-based methods, our PCGFormer deals with voxelized PCGs… apply hierarchical binary cross-entropy (BCE) loss to guide end-to-end learning… at the beginning of each Trans-former module, k points of the current point including their coordinates and features are collected and combined as F′′X of N×k×d for the next step). Regarding Claims 17 and 35, the combination of Pang and Liu teaches the processing circuitry is further configured to: perform octree encoding on the coordinates of the output tensor to generate the encoded coordinates; quantize the higher-dimensional features to generate quantized higher-dimensional features; arithmetically encode the quantized higher-dimensional features of the output tensor using an entropy model to generate encoded higher-dimensional features; and output the encoded coordinates and the encoded higher-dimensional features in an encoded bitstream (Figs. 1-2; Section III of Liu; encoded using G- PCC Octree codec). Regarding Claims 18 and 36, the combination of Pang and Liu teaches a sensor configured to capture the frame of point cloud data (Para. 2, 45 of Pang; sensors like LiDARs produce (dynamic) point clouds that are used by the perception engine). Claims 37-53 and 55-71 are directed towards an apparatus and method configured to decode encoded point cloud data and are taught by the combination of Pang and Liu, which teach the apparatus and method configured to encode a point cloud as shown above. Regarding Claims 54 and 72, the combination of Pang and Liu teaches a display configured to display the encoded point cloud (Fig. 1; Para. 41 of Pang; system 100 may provide an output signal to various output devices, including a display 165). Allowable Subject Matter Claims 16, 34, 52, and 70 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. None of the references, either singularly or in combination, teach or fairly suggest the apparatus of claim 1, wherein the processing circuitry is further configured to: encode one or more syntax elements, the one or more syntax elements including: a first syntax element indicating whether positional encoding is enabled, a second syntax element indicating a type of positional encoding for the at least one positional encoding layer, a third syntax element indicating a feature size for the at least one positional encoding layer, a fourth syntax element indicating a location of the at least one positional encoding layer, or a fifth syntax element indicating an input type for the at least one positional encoding layer. None of the references, either singularly or in combination, teach or fairly suggest the apparatus of claim 37, wherein the processing circuitry is further configured to: decode one or more syntax elements, the one or more syntax elements including: a first syntax element indicating whether positional encoding is enabled, a second syntax element indicating a type of positional encoding for the at least one positional encoding layer, a third syntax element indicating a feature size for the at least one positional encoding layer, a fourth syntax element indicating a location of the at least one positional encoding layer, or a fifth syntax element indicating an input type for the at least one positional encoding layer. Any comments considered necessary by applicant must be submitted no later than the payment of the issue fee and, to avoid processing delays, should preferably accompany the issue fee. Such submissions should be clearly labeled “Comments on Statement of Reasons for Allowance.” Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to ABHISHEK SARMA whose telephone number is (571)272-9887. The examiner can normally be reached on Mon - Fri 8:00-5:00. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Amr Awad can be reached on 571-272-7764. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ABHISHEK SARMA/ Primary Examiner, Art Unit 2621
Read full office action

Prosecution Timeline

Mar 15, 2024
Application Filed
Aug 11, 2026
Examiner Interview (Telephonic)
Aug 31, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749200
METHODS AND SYSTEM FOR OBJECT PATH DETECTION IN A WORKPLACE
2y 1m to grant Granted Sep 29, 2026
Patent 12743152
SYSTEMS, APPARATUS, ARTICLES OF MANUFACTURE, AND METHODS FOR EYE GAZE CORRECTION IN CAMERA IMAGE STREAMS
4y 2m to grant Granted Sep 22, 2026
Patent 12739354
Synchronizing Image Signal Processing Across Multiple Image Sensors
2y 6m to grant Granted Sep 15, 2026
Patent 12737443
FINGERPRINT AUTHENTICATION-RELATED INDICATORS FOR CONTROLLING DEVICE ACCESS AND/OR FUNCTIONALITY
1y 4m to grant Granted Sep 15, 2026
Patent 12725299
THREE-DIMENSIONAL SCANNING SYSTEM AND THREE-DIMENSIONAL SCANNING METHOD
2y 11m to grant Granted Sep 01, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
84%
Grant Probability
84%
With Interview (+0.3%)
2y 1m (~0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 589 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month