DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
In the response to this Office Action, the Examiner respectfully requests that support be shown for language added to any original claims on amendment and any new claims. That is, indicate support for newly added claim language by specifically pointing to page(s) and line numbers in the specification and/or drawing figure(s). This will assist the Examiner in prosecuting this application.
Election/Restrictions
Applicant's election with traverse of Group II: in the reply filed on 06/16/2026 is acknowledged and is found persuasive. Therefore, the restriction requirement of 04/29/2026 is withdrawn.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-15, 17-33, 35-51, 53-69, and 71-72 are rejected under 35 U.S.C. 103 as being unpatentable over U.S. Patent Application Publication 2025/0139834 A1 to Pang et al. (hereinafter "Pang") in view of Non-Patent Publication "PCGFormer: Lossy Point Cloud Geometry Compression via Local Self-Attention" to Liu et al. (hereinafter "Liu", included in IDS provided by Applicant).
Regarding Claims 1 and 19, Pang teaches an apparatus and a method configured to encode a point cloud, the apparatus comprising: one or more memories configured to store point cloud data; and processing circuitry in communication with the one or more memories, the processing circuitry configured to: receive a frame of point cloud data; encode coordinates of a geometry of the frame of point cloud data using a deep learning network (Figs. 1-5; Claims 16, 42; Para. 30-84 of Pang; base point cloud encoder (320)… feature encoder (370)… For an input point cloud PC0 to be compressed, it is first quantized using a quantizer (310) with a step size s (s>1). In one embodiment, for every point, say, A in PC0 with 3D coordinates (x, y, z), the quantizer divides the coordinates by s then converts them to integers, leading to ([x/s], [y/s], [z/s]), where the function [⋅] can be the floor, ceiling or rounding function. The quantizer then removes the duplicate points with the same 3D coordinates, that is, if there exist several points having the same coordinates, we only keep one of them and remove the rest, then we get PC1. The obtained quantized point cloud, PC1, is then compressed with a base point cloud encoder (320), which outputs a bitstream BS0), wherein the deep learning network includes one or more layers configured to generate first features for the coordinates of the geometry, and wherein the deep learning network further includes at least one positional encoding layer (Figs. 1-5; Claims 16, 42; Para. 30-84 of Pang; base point cloud encoder (320)… feature encoder (370)… For an input point cloud PC0 to be compressed, it is first quantized using a quantizer (310) with a step size s (s>1). In one embodiment, for every point, say, A in PC0 with 3D coordinates (x, y, z), the quantizer divides the coordinates by s then converts them to integers, leading to ([x/s], [y/s], [z/s]), where the function [⋅] can be the floor, ceiling or rounding function. The quantizer then removes the duplicate points with the same 3D coordinates, that is, if there exist several points having the same coordinates, we only keep one of them and remove the rest, then we get PC1. The obtained quantized point cloud, PC1, is then compressed with a base point cloud encoder (320), which outputs a bitstream BS0… feature codec can be implemented based on a deep neural network).
Pang does not explicitly disclose at least one positional encoding layer configured to generate second features for the coordinates of the geometry and combine the first features with the second features to generate higher-dimensional features; and output an output tensor comprising encoded coordinates and the higher-dimensional features.
However, Liu teaches at least one positional encoding layer configured to generate second features for coordinates of geometry and combine first features with the second features to generate higher-dimensional features; and output an output tensor comprising encoded coordinates and the higher-dimensional features (Figs. 1-2; Abstract; Section II-III of Liu; For the encoding of PCGFormer, input PCGs are gradually downscaled with a factor of 2 at each axis in a Cartesian coordinate system. Here three layers of spatial downscaling are exemplified to gradually analyze and embed neighborhood information. An Inception-Residual Network (IRN) layer (see Fig. 2a) and a Transformer layer using local self-attention upon k nearest neighbors (a.k.a., kNN self-attention) are stacked as modular components at each scale for feature extraction and aggregation)… proposed PCGFormer improves the ConvMST… like rules-based G-PCC and popular learning-based methods, our PCGFormer deals with voxelized PCGs… apply hierarchical binary cross-entropy (BCE) loss to guide end-to-end learning… at the beginning of each Trans-former module, k points of the current point including their coordinates and features are collected and combined as F′′X of N×k×d for the next step).
Therefore, at the time when the invention was filed, it would have been obvious to a person of ordinary skill in the art to include at least one positional encoding layer configured to generate second features for the coordinates of the geometry and combine the first features with the second features to generate higher-dimensional features; and output an output tensor comprising encoded coordinates and the higher-dimensional features IC using the teachings of Liu in order to modify the device taught by Pang. The motivation to combine these analogous arts would have been to provide a framework of multiscale sparse tensor representation, with which spatial neighbors can be effectively aggregated and embedded for better compression performance (Section 5 of Liu).
Regarding Claims 2 and 20, the combination of Pang and Liu teaches that the deep learning network includes one or more layers configured to downscale the coordinates of the geometry to generate downscaled coordinates, and wherein the output tensor comprises the downscaled coordinates and the higher-dimensional features (Figs. 1-2; Abstract; Section II-III of Liu; For the encoding of PCGFormer, input PCGs are gradually downscaled with a factor of 2 at each axis in a Cartesian coordinate system. Here three layers of spatial downscaling are exemplified to gradually analyze and embed neighborhood information. An Inception-Residual Network (IRN) layer (see Fig. 2a) and a Transformer layer using local self-attention upon k nearest neighbors (a.k.a., kNN self-attention) are stacked as modular components at each scale for feature extraction and aggregation).
Regarding Claims 3 and 21, the combination of Pang and Liu teaches that the at least one position encoding layer of the deep learning network includes a first positional encoding layer that operates on an input of the deep learning network and a second positional encoding layer that operates on the downscaled coordinates (Figs. 1-2; Abstract; Section II-III of Liu; For the encoding of PCGFormer, input PCGs are gradually downscaled with a factor of 2 at each axis in a Cartesian coordinate system. Here three layers of spatial downscaling are exemplified to gradually analyze and embed neighborhood information. An Inception-Residual Network (IRN) layer (see Fig. 2a) and a Transformer layer using local self-attention upon k nearest neighbors (a.k.a., kNN self-attention) are stacked as modular components at each scale for feature extraction and aggregation).
Regarding Claims 4 and 22, the combination of Pang and Liu teaches that the one or more layers of the deep learning network includes three layers configured to progressively downscale the coordinates of the geometry to generate first downscaled coordinates, second downscaled coordinates, and third downscaled coordinates, and wherein the output tensor comprises the third downscaled coordinates that are three-times downscaled and the higher-dimensional features (Figs. 1-2; Abstract; Section II-III of Liu; For the encoding of PCGFormer, input PCGs are gradually downscaled with a factor of 2 at each axis in a Cartesian coordinate system. Here three layers of spatial downscaling are exemplified to gradually analyze and embed neighborhood information. An Inception-Residual Network (IRN) layer (see Fig. 2a) and a Transformer layer using local self-attention upon k nearest neighbors (a.k.a., kNN self-attention) are stacked as modular components at each scale for feature extraction and aggregation).
Regarding Claims 5 and 23, the combination of Pang and Liu teaches that the at least one position encoding layer of the deep learning network includes a first positional encoding layer that operates on an input of the deep learning network, a second positional encoding layer that operates on the first downscaled coordinates, a third positional encoding layer that operates on the second downscaled coordinates, and a fourth positional encoding layer that operates on the third downscaled coordinates (Figs. 1-2; Abstract; Section II-III of Liu; For the encoding of PCGFormer, input PCGs are gradually downscaled with a factor of 2 at each axis in a Cartesian coordinate system. Here three layers of spatial downscaling are exemplified to gradually analyze and embed neighborhood information. An Inception-Residual Network (IRN) layer (see Fig. 2a) and a Transformer layer using local self-attention upon k nearest neighbors (a.k.a., kNN self-attention) are stacked as modular components at each scale for feature extraction and aggregation).
Regarding Claims 6 and 24, the combination of Pang and Liu teaches that at least one of the second positional encoding layer, the third positional encoding layer, or the fourth positional encoding layer is positioned between a layer configured to downscale the coordinates of the geometry and an inception-residual block (Figs. 1-2; Section III of Liu; IRN is the inception-residual network).
Regarding Claims 7 and 25, the combination of Pang and Liu teaches that the at least one positional encoding layer comprises a learned positional encoding layer (Figs. 1-2; Section III of Liu; like rules-based G-PCC and popular learning-based methods, our PCGFormer deals with voxelized PCGs).
Regarding Claims 8 and 26, the combination of Pang and Liu teaches that the learned positional encoding layer is configured as a convolutional layer, a fully connected layer, or a multilayer perceptron (Figs. 1-2; Section III of Liu; encoded using G- PCC Octree codec).
Regarding Claims 9 and 27, the combination of Pang and Liu teaches that the deep learning network includes a plurality of positional encoding layers, and wherein each of the positional encoding layers is configured to use different sets of weights (Figs. 1-2; Section III of Liu; For every point, an attention map will be calculated. Through applying the attention map to k local neighbor features, we attain the attention weighted feature at a size of 1 × d of each point).
Regarding Claims 10 and 28, the combination of Pang and Liu teaches that the deep learning network includes a plurality of positional encoding layers, and wherein each of the positional encoding layers is configured to use a same set of weights (Figs. 1-2; Section II-III of Liu; fixed-weightkernels).
Regarding Claims 11 and 29, the combination of Pang and Liu teaches that the at least one positional encoding layer comprises one of a sine and cosine positional encoding layer, a linear positional encoding layer, a radial basis function positional encoding layer, a Fermi-Dirac positional encoding layer, a hybrid positional encoding layer, a relative positional encoding layer, or an image grid positional encoding layer (Figs. 1-2; Section III of Liu; encoded using G- PCC Octree codec).
Regarding Claims 12 and 30, the combination of Pang and Liu teaches that the deep learning network includes a plurality of positional encoding layers, and wherein the plurality of positional encoding layers include at least two types of positional encoding (Figs. 3-5; Para. 62-84 of Pang; base point cloud encoder (320)… feature encoder (370)… feature codec can be implemented based on a deep neural network… Figs. 1-2; Section II-III of Liu; proposed PCGFormer improves the ConvMST… like rules-based G-PCC and popular learning-based methods, our PCGFormer deals with voxelized PCGs… apply hierarchical binary cross-entropy (BCE) loss to guide end-to-end learning… at the beginning of each Trans-former module, k points of the current point including their coordinates and features are collected and combined as F′′X of N×k×d for the next step).
Regarding Claims 13 and 31, the combination of Pang and Liu teaches that to combine the first features with the second features to generate higher-dimensional features, the positional encoding layer is configured to concatenate the first features with the second features (Figs. 1-2; Section II-III of Liu; proposed PCGFormer improves the ConvMST… like rules-based G-PCC and popular learning-based methods, our PCGFormer deals with voxelized PCGs… apply hierarchical binary cross-entropy (BCE) loss to guide end-to-end learning… at the beginning of each Trans-former module, k points of the current point including their coordinates and features are collected and combined as F′′X of N×k×d for the next step).
Regarding Claims 14 and 32, the combination of Pang and Liu teaches that the first features have a feature size of 64 and wherein the higher-dimensional features have a feature size of 192 (Figs. 1-2; Section II-III of Liu; proposed PCGFormer improves the ConvMST… like rules-based G-PCC and popular learning-based methods, our PCGFormer deals with voxelized PCGs… apply hierarchical binary cross-entropy (BCE) loss to guide end-to-end learning… at the beginning of each Trans-former module, k points of the current point including their coordinates and features are collected and combined as F′′X of N×k×d for the next step… Selecting a first feature size of 64 and a higher-dimensional feature size of 192 would only require routine skill for a person of ordinary skill in the art based on the teachings of Liu and can be accomplished without any undue experimentation and it has been held that where the general conditions of a claim are disclosed in the prior art, discovering the optimum or workable ranges involves only routine skill in the art. In re Aller, 105 USPQ 233).
Regarding Claims 15 and 33, the combination of Pang and Liu teaches that the processing circuitry is further configured to: encode attributes of the frame of point cloud data using the deep learning network, wherein the deep learning network includes one or more layers configured to generate third features for the attributes, and wherein the deep learning network further includes at least one positional encoding layer configured to generate fourth features from the attributes and combine the third features with the fourth features to generate second higher-dimensional features (Figs. 1-2; Section II-III of Liu; proposed PCGFormer improves the ConvMST… like rules-based G-PCC and popular learning-based methods, our PCGFormer deals with voxelized PCGs… apply hierarchical binary cross-entropy (BCE) loss to guide end-to-end learning… at the beginning of each Trans-former module, k points of the current point including their coordinates and features are collected and combined as F′′X of N×k×d for the next step).
Regarding Claims 17 and 35, the combination of Pang and Liu teaches the processing circuitry is further configured to: perform octree encoding on the coordinates of the output tensor to generate the encoded coordinates; quantize the higher-dimensional features to generate quantized higher-dimensional features; arithmetically encode the quantized higher-dimensional features of the output tensor using an entropy model to generate encoded higher-dimensional features; and output the encoded coordinates and the encoded higher-dimensional features in an encoded bitstream (Figs. 1-2; Section III of Liu; encoded using G- PCC Octree codec).
Regarding Claims 18 and 36, the combination of Pang and Liu teaches a sensor configured to capture the frame of point cloud data (Para. 2, 45 of Pang; sensors like LiDARs produce (dynamic) point clouds that are used by the perception engine).
Claims 37-53 and 55-71 are directed towards an apparatus and method configured to decode encoded point cloud data and are taught by the combination of Pang and Liu, which teach the apparatus and method configured to encode a point cloud as shown above.
Regarding Claims 54 and 72, the combination of Pang and Liu teaches a display configured to display the encoded point cloud (Fig. 1; Para. 41 of Pang; system 100 may provide an output signal to various output devices, including a display 165).
Allowable Subject Matter
Claims 16, 34, 52, and 70 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
None of the references, either singularly or in combination, teach or fairly suggest the apparatus of claim 1, wherein the processing circuitry is further configured to: encode one or more syntax elements, the one or more syntax elements including: a first syntax element indicating whether positional encoding is enabled, a second syntax element indicating a type of positional encoding for the at least one positional encoding layer, a third syntax element indicating a feature size for the at least one positional encoding layer, a fourth syntax element indicating a location of the at least one positional encoding layer, or a fifth syntax element indicating an input type for the at least one positional encoding layer.
None of the references, either singularly or in combination, teach or fairly suggest the apparatus of claim 37, wherein the processing circuitry is further configured to: decode one or more syntax elements, the one or more syntax elements including: a first syntax element indicating whether positional encoding is enabled, a second syntax element indicating a type of positional encoding for the at least one positional encoding layer, a third syntax element indicating a feature size for the at least one positional encoding layer, a fourth syntax element indicating a location of the at least one positional encoding layer, or a fifth syntax element indicating an input type for the at least one positional encoding layer.
Any comments considered necessary by applicant must be submitted no later than the payment of the issue fee and, to avoid processing delays, should preferably accompany the issue fee. Such submissions should be clearly labeled “Comments on Statement of Reasons for Allowance.”
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ABHISHEK SARMA whose telephone number is (571)272-9887. The examiner can normally be reached on Mon - Fri 8:00-5:00.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Amr Awad can be reached on 571-272-7764. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ABHISHEK SARMA/
Primary Examiner, Art Unit 2621