Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Objections
Claims 4 and 20 are objected to because of the following informalities:
Claim 20, which is dependent on claim 18, further comprises that the fifth feature is “fifth feature is obtained by combining the second feature and the sixth feature using the second neural network module”. Claim 18 comprises “obtaining a fifth feature, using a fourth neural network module, based on the second feature”. Since claim 20 depends on claim 18, this suggests that both manners of generating the fifth feature are required, which creates uncertainty on how the fifth feature is created. Paragraphs 8-9 suggests that the generation of the fifth feature in claims 18 and 20 are separate/alternative examples and therefore Examiner will interpret for claim 20 to be dependent on claim 17. The claim should be amended to clarify the intended relationship between these limitations.
Claim 4 recites similar limitations to claim 20 and is objected to for the same reasonings as used above.
Appropriate correction is required.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-2, 4, 7-11, 14, 17, 18, 20, 23-27, and 30 is/are rejected under 35 U.S.C. 103 as being unpatentable over Huang et al (US 11676310 B2), Chen et al (“Point Cloud Compression with Sibling Context and Surface Priors”), and Wu (“PointPWC-Net: Cost Volume on Point Clouds for (Self-)Supervised Scene Flow Estimation”), hereinafter Huang, Chen, and Wu respectively.
Regarding claim 17, Huang teaches a device comprising: a processor; and a transceiver operatively coupled to the processor;
“The data encoding system can receive the point cloud data from the LIDAR sensor. Once the point cloud data has been received, the data encoding system can parse the point cloud data to generate a tree-based data structure for representing and/or storing the point cloud data. In some examples, the tree-based data structure may be an octree wherein each node can have up to eight child nodes with which it is associated." - Col 5, Lines 5-12
NOTE: Huang discloses a system for taking input LIDAR point cloud data using an encoding system to generate a tree-based structure such as an octree to represent the point cloud. A processor would naturally be required to execute the instructions for this point cloud encoder device and a communication interface for receiving/transmitting the point cloud bit-stream would be needed to receive the point cloud data.
the processor is configured to access a second feature, wherein the second feature is a deep feature vector from a previous LoD;
“Then, starting with the feature h.sub.i.sup.(0) for each node, we perform K aggregations between the current node feature and the feature of its parent. At iteration k, the aggregation can also be modeled as an MLP: h.sub.i.sup.(k)=MLP.sup.(k)([h.sub.i.sup.k−1],[h.sub.pa(i).sup.k−1]) where h.sub.pa(i).sup.(k−1) is the hidden feature of node i's parent node." - Col 6 Lines 65-66 and Col 7, Lines 1-4
NOTE: Huang discloses generating hidden features by aggregating the current node feature with a parent-node hidden feature. The already generated parent hidden features are accessed and used as contextual information to generate the current node hidden feature. Since the parent node is naturally located at a coarser octree level compared to the current level, the parent-node hidden feature corresponds to the second feature which is a deep feature vector from a previous level of detail.
the processor is configured to combine the first feature and the second feature using a second neural network module upon the whole point cloud frame at the current LoD to obtain a third feature,
“Then, starting with the feature h.sub.i.sup.(0) for each node, we perform K aggregations between the current node feature and the feature of its parent. At iteration k, the aggregation can also be modeled as an MLP: h.sub.i.sup.(k)=MLP.sup.(k)([h.sub.i.sup.k−1],[h.sub.pa(i).sup.k−1]) where h.sub.pa(i).sup.(k−1) is the hidden feature of node i's parent node." - Col 6 Lines 65-66 and Col 7, Lines 1-4
NOTE: Huang discloses using a multi-layer perceptron (MLP) network to perform the combination of feature h.sub.i.sup.k−1 (current node which can be seen as the first feature) and h.sub.pa(i).sup.k−1 (parent node which can be seen as the second feature) to generate a better representation of the current node that uses information from the parent node feature h.sub.i.sup.(k) (which can be seen as the third feature). This corresponds to combining a first and the second feature using a second neural network upon the whole point cloud frame at the current LoD to obtain a third feature.
Huang does not teach wherein the processor and the transceiver are configured to access a set of occupancy bits of neighboring voxels at a current level of detail (LoD) that are already encoded or decoded, using a tree-based point cloud decoder; the processor is configured to compute a first feature, using a first neural network module upon a whole point cloud frame at the current LoD, based on the accessed set of occupancy bits of neighboring voxels at the current LoD, wherein the first neural network module is a convolution based neural network module; the processor is configured to access a second feature, wherein the second feature is a deep feature vector from a previous LoD; the processor is configured to combine the first feature and the second feature using a second neural network module upon the whole point cloud frame at the current LoD to obtain a third feature, wherein the second neural network module is a convolution based neural network module (NOTE: Huang teaches combining a first and second feature to obtain a third feature, but not the specific first feature as defined in the claims or that the second neural network is a convolutional based neural network); the processor is configured to concatenate the third feature along with one or more features of the current voxels to be encoded or decoded to compose a fourth feature, wherein the fourth feature is a comprehensive feature; and the processor is configured to predict a probability distribution of occupancy of voxels at the current LoD based on the fourth feature using a third neural network module. However, Chen teaches wherein the processor and the transceiver are configured to access a set of occupancy bits of neighboring voxels at a current level of detail (LoD) that are already encoded or decoded,
"The octree at depth l (l ∈ [1,L]) can be considered as a discretization of the 3D space at the resolution of 2l · 2l · 2l. Inspired by [24], we construct the neighbor context of an octant ni by locally forming a K × K × K binary voxel block Vi centered at ni and K is empirically set to 9 in experiments. Each binary value indicates the existence of its corresponding neighbors. The neighbor voxel Vi is transformed by a feature extraction function fneigh: h_i^neigh = f_neigh(V_i), where h_neigh, is the output feature vector that represents the neighbor contextual information." - Pg 6, Section 3.3.1, Par 1.
NOTE: Chen discloses constructing an octree depth and binary voxel block centered at a current octant. The binary values in the binary voxel block indicate the occupancy of the neighboring voxels. This functionally corresponds to an accessed set of occupancy bits of neighboring voxels at a current LoD. Since occupancy information is used as input for predicting the current octant, a decoder would have accessed the occupancy information that has already been encoded/decoded.
using a tree-based point cloud decoder;
"In the decoding stage, we first reconstruct an octree with L levels from the compressed bitstream." - Pg 4, Section 3.1, Par 2
NOTE: The tree-based point cloud decoder would naturally be required for Chen's decoding stage to reconstruct an octree.
the processor is configured to compute a first feature, using a first neural network module upon a whole point cloud frame at the current LoD, based on the accessed set of occupancy bits of neighboring voxels at the current LoD, wherein the first neural network module is a convolution based neural network module;
"The octree at depth l (l ∈ [1,L]) can be considered as a discretization of the 3D space at the resolution of 2l · 2l · 2l. Inspired by [24], we construct the neighbor context of an octant ni by locally forming a K × K × K binary voxel block Vi centered at ni and K is empirically set to 9 in experiments. Each binary value indicates the existence of its corresponding neighbors. The neighbor voxel Vi is transformed by a feature extraction function fneigh: h_i^neigh = f_neigh(V_i), where h_neigh, is the output feature vector that represents the neighbor contextual information." - Pg 6, Section 3.3.1, Par 1.
NOTE: Chen discloses constructing an octree depth and binary voxel block centered at a current octant. The binary values in the binary voxel block indicate the occupancy of the neighboring voxels. This functionally corresponds to an accessed set of occupancy bits of neighboring voxels at a current LoD. Chen also teaches a feature extraction function that outputs a feature vector representing the neighbor contextual information. Chen also discloses in Fig. 2 that computing a first feature based on the accessed set of occupancy bits of neighboring voxels at the current LoD is done using a 3D Convolutional Layer (first neural network).
the processor is configured to combine the first feature and the second feature using a second neural network module upon the whole point cloud frame at the current LoD to obtain a third feature,
“Then, starting with the feature h.sub.i.sup.(0) for each node, we perform K aggregations between the current node feature and the feature of its parent. At iteration k, the aggregation can also be modeled as an MLP: h.sub.i.sup.(k)=MLP.sup.(k)([h.sub.i.sup.k−1],[h.sub.pa(i).sup.k−1]) where h.sub.pa(i).sup.(k−1) is the hidden feature of node i's parent node." – Huang Col 6 Lines 65-66 and Col 7, Lines 1-4
NOTE: Chen discloses a feature extraction function that outputs a feature vector representing the neighbor contextual information which corresponds to obtaining the first feature based on the accessed set of occupancy bits of neighboring voxels at the current level. This combination of features is done using a multi-layer perceptron (second neural network). After the combination, the first feature obtained by the process of Chen and the second feature obtained by accessing a parent-node feature which corresponds to the second feature being a deep feature vector from a previous LoD, can be used by Huang’s process for combining a first and second feature to obtain a third feature.
the processor is configured to concatenate the third feature along with one or more features of the current voxels to be encoded or decoded to compose a fourth feature, wherein the fourth feature is a comprehensive feature;
"we first concatenate the hneigh i with the octant information ci and then apply an MLP denoted as φ to extract more condensed features hchild i processed features hchild i . The will then be passed to its children as their ancestral dependence. A following linear projection layer γ is applied on hchild i hneigh i to scale back the features to the original dimension. Then the feature is added with through skip connection to recover the original neighbor feature ˆ hneigh i the current octant: hchild i =φ(Concat(hneigh i , ci)), ˆ hneigh i =γ(hchild i ) +hneigh i." - Pg 6, Section 3.3.2, Par 2
NOTE: Chen teaches the formula: hchild i =φ(Concat(hneigh i , ci)), ˆ hneigh i =γ(hchild i )+hneigh which demonstrates the concatenation of the neighbor-context feature information and the context associated with the current octant. This functionally corresponds to concatenating a feature along with one or more features of the current voxels to be encoded or decoded to compose a fourth feature. The resultant feature after the concatenation will naturally be a comprehensive feature since information from the previous and current level was used.
and the processor is configured to predict a probability distribution of occupancy of voxels at the current LoD based on the fourth feature using a third neural network module.
"Given an occupancy symbol stream s = [s1,s2,...,sn] with the number of non leaf octants as its length n, the goal of our entropy model is to minimize the bitstream length. According to information theory, this goal can be reached by minimizing cross-entropy loss Es∼p[−logq(s)], where q(s) is the estimated probability distribution of occupancy symbol s and p is the actual distribution." - Chen Pg 5, Section 3.3, Par 1
[AltContent: rect]NOTE: Chen shows in Fig. 2 that the probability distribution representing the occupancy possibilities for the current node as the output and shows that the 3D Convolution layer, FC layer, and softmax layer are a part of the prediction calculation. This functionally corresponds to predicting a probability distribution of occupancy of voxels at the current LoD using a third neural network module.
PNG
media_image1.png
528
1195
media_image1.png
Greyscale
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Huang by incorporating the teachings of Chen to access a set of occupancy bits of neighboring voxels at a current level of detail using a tree-based point cloud decoder to compute a first feature using a first convolutional based neural network, concatenating a third feature along one or more features to compose a fourth feature, and predicting a probability distribution of occupancy of voxels at a current LoD based on the fourth feature using a third neural network. One would be motivated to make this combination to improve the prediction process of voxel occupancy in the point cloud by providing access to neighboring voxel information to obtain feature information which results in generating a representation of the current level which is then used to generate a more accurate occupancy probability distribution.
Huang in view of Chen still does not teach the processor is configured to access a second feature, wherein the second feature is a deep feature vector from a previous LoD; the processor is configured to combine the first feature and the second feature using a second neural network module upon the whole point cloud frame at the current LoD to obtain a third feature, wherein the second neural network module is a convolution based neural network module (NOTE: Huang in view of Chen does teach combining the specific first and second feature as defined in the claims, but still does not teach the combination being performed with a convolutional based neural network). However, Wu teaches the processor is configured to access a second feature, wherein the second feature is a deep feature vector from a previous LoD; the processor is configured to combine the first feature and the second feature using a second neural network module upon the whole point cloud frame at the current LoD to obtain a third feature, wherein the second neural network module is a convolution based neural network module;
"We generate an L-level pyramid of feature representations, with the top level being the input point clouds, i.e., l0 = P/Q. For each level l, we use furthest point sampling [54] to downsample the points by a factor of 4 from previous level l − 1, and use PointConv [80] to perform convolution on the features from level l −1. As a result, we can generate a feature pyramid with L levels for each input point cloud. After this, we enlarge the receptive field at level l of the pyramid by upsampling the feature in level l + 1 and concatenate it to the feature at level l." - Pg 7, Section 3.2, Par 2, Lines 12-19
NOTE: Wu discloses constructing a hierarchical feature pyramid for a 3D point cloud and using PointConv to perform convolution on features from the preceding levels. It then upsamples the feature level at the previous level l-1 and concatenates it with the feature level at the current level l. This functionally corresponds to using a convolution based neural network module to combine a first and second feature. After the combination, the PointConv architecture for combining a previous and current feature as taught by Wu can modify Huang's system of combining the feature of a previous level and the feature of a current level using a multi-layer perceptron. This modification would allow the second feature, wherein the second feature is a deep feature vector from a previous LoD as taught by Huang and the first feature based on the accessed set of occupancy bits of neighboring voxels at the current LoD as taught by Chen to instead be used as input to Wu's convolutional network to combine the first and second features to create a third feature.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Huang by incorporating the teachings of Wu to have the combination of a first and second feature to produce a third feature be performed by a convolutional based neural network. One would be motivated to make this combination since convolutional-based neural networks are a standard neural network architecture. Implementing the CNN of Wu in Huang’s device will lead to the predicted result of retaining and incorporating information learned at a previous level to improve calculating predictions of voxel occupancy at a current level.
Regarding claim 1, the claim recites similar limitations to claim 17. Therefore, method claim 1 corresponds to the device disclosed in claim 17 and is rejected for the same reasons of obviousness as used above.
Regarding claim 18, Huang in view of Chen and Wu teaches the device of claim 17. Huang as modified further teaches wherein the processor is further configured to obtain a fifth feature, using a fourth neural network module, based on the second feature, wherein the fourth neural network module is a convolution based neural network module,
"We generate an L-level pyramid of feature representations, with the top level being the input point clouds, i.e., l0 = P/Q. For each level l, we use furthest point sampling [54] to downsample the points by a factor of 4 from previous level l − 1, and use PointConv [80] to perform convolution on the features from level l −1. As a result, we can generate a feature pyramid with L levels for each input point cloud. After this, we enlarge the receptive field at level l of the pyramid by upsampling the feature in level l + 1 and concatenate it to the feature at level l." – Wu, Pg 7, Section 3.2, Par 2, Lines 12-19
NOTE: Wu teaches that a feature from a previous level is processed by the PointConv convolution based neural network model to generate a feature at another level. After the combination, the PointConv operation of Wu can be applied to the obtained second feature as taught by Huang, see rejection of claim 17, to then obtain the fifth feature. The PointConv model provides the fourth neural network in which the second feature from Huang is used as input to obtain the fifth feature.
and wherein the third feature is obtained by combining the first feature and the fifth feature.
“Then, starting with the feature h.sub.i.sup.(0) for each node, we perform K aggregations between the current node feature and the feature of its parent. At iteration k, the aggregation can also be modeled as an MLP: h.sub.i.sup.(k)=MLP.sup.(k)([h.sub.i.sup.k−1],[h.sub.pa(i).sup.k−1]) where h.sub.pa(i).sup.(k−1) is the hidden feature of node i's parent node." – Huang, Col 6 Lines 65-66 and Col 7, Lines 1-4
NOTE: After obtaining the fifth feature using the Wu’s process, the third feature can be obtained by combining the first feature and the fifth feature (instead of the first and second) using the process as taught by Huang.
Regarding claim 2, the claim recites similar limitations to claim 18. Therefore, method claim 2 corresponds to the device disclosed in claim 18 and is rejected for the same reasons of obviousness as used above.
Regarding claim 20, Huang in view of Chen and Wu teaches the device of claim 18. Huang as modified teaches wherein the processor and the transceiver are further configured to access a sixth feature, wherein the sixth feature is based on bits from already encoded or decoded voxels at the current LoD,
"The octree at depth l (l ∈ [1,L]) can be considered as a discretization of the 3D space at the resolution of 2l · 2l · 2l. Inspired by [24], we construct the neighbor context of an octant ni by locally forming a K × K × K binary voxel block Vi centered at ni and K is empirically set to 9 in experiments. Each binary value indicates the existence of its corresponding neighbors. The neighbor voxel Vi is transformed by a feature extraction function fneigh: h_i^neigh = f_neigh(V_i), where h_neigh, is the output feature vector that represents the neighbor contextual information." – Chen Pg 6, Section 3.3.1, Par 1.
NOTE: Chen discloses constructing an octree depth and binary voxel block centered at a current octant. The binary values in the binary voxel block indicate the occupancy of the neighboring voxels. The neighboring voxels are transformed by a feature extraction function which outputs a feature vector representing the neighbor voxels information which may be understood as the sixth feature. Chen further discloses: “Since our entropy model encodes the occupancy symbols of nj’s children sequentially, the previous siblings’ occupancy symbols are available while predicting the probability of the current octant ni. The available occupancy symbols are filled in V sib i .” – Pg 7, Section 3.3.3, Par 2, Lines 4-7. This implies that since the occupancy symbols are encoded sequentially, the sixth feature would naturally be based on bits from already encoded voxels at the current level.
wherein the fifth feature is obtained by combining the second feature and the sixth feature using the second neural network module.
“Then, starting with the feature h.sub.i.sup.(0) for each node, we perform K aggregations between the current node feature and the feature of its parent. At iteration k, the aggregation can also be modeled as an MLP: h.sub.i.sup.(k)=MLP.sup.(k)([h.sub.i.sup.k−1],[h.sub.pa(i).sup.k−1]) where h.sub.pa(i).sup.(k−1) is the hidden feature of node i's parent node." – Huang Col 6 Lines 65-66 and Col 7, Lines 1-4
NOTE: Huang as modified by Chen and Wang discloses using the convolutional based neural network to perform Huang’s process of combining a first and second feature to obtain a third feature, see claim 17. After the combination, the sixth feature obtained by Chen may then be substituted as input so that the second and sixth feature can be combined to obtain the fifth feature.
Regarding claim 4, the claim recites similar limitations to claim 20. Therefore, method claim 4 corresponds to the device disclosed in claim 20 and is rejected for the same reasons of obviousness as used above.
Regarding claim 23, Huang in view of Chen and Wu teaches the device of claim 17. Huang as modified teaches wherein the fourth feature is composed further based on the second feature combined with the third feature.
“Then, starting with the feature h.sub.i.sup.(0) for each node, we perform K aggregations between the current node feature and the feature of its parent. At iteration k, the aggregation can also be modeled as an MLP: h.sub.i.sup.(k)=MLP.sup.(k)([h.sub.i.sup.k−1],[h.sub.pa(i).sup.k−1]) where h.sub.pa(i).sup.(k−1) is the hidden feature of node i's parent node." – Huang Col 6 Lines 65-66 and Col 7, Lines 1-4
NOTE: Huang as modified teaches the combination of a first and second feature to create a third feature, see rejection of claim 17. The same process can be performed to obtain the fourth feature by combining the second feature with the newly obtained third feature. This would correspond to obtaining a fourth feature composed based on the combination of the second and third feature.
Regarding claim 7, the claim recites similar limitations to claim 23. Therefore, method claim 7 corresponds to the device disclosed in claim 23 and is rejected for the same reasons of obviousness as used above.
Regarding claim 24, Huang in view of Chen and Wu teaches the device of claim 17. Huang as modified teaches wherein the third neural network module is part of a multi-layer perceptron (MLP)-based module.
“The contextual information can include, but is not limited to, the octant value of the respective symbol, the location of the respective symbol, the level of the respective symbol in the tree-based data structure, and the occupancy value of the parent node associated with the respective symbol. This contextual information can then be used as input into a machine-learned model (a multi-layer perceptron) that can produce, as output, a statistical distribution for the expected occupancy for the respective node.” – Huang Col 2 Lines 66-67, Col 3, Lines 1-8
NOTE: Huang discloses a multi-layer perceptron system that outputs a statistical distribution for the expected occupancy for the respective node. This functionally corresponds to the third neural network which is configured to generate a probability distribution for occupancy of voxels at a current LoD as defined in claim 17.
Regarding claim 8, the claim recites similar limitations to claim 24. Therefore, method claim 8 corresponds to the device disclosed in claim 24 and is rejected for the same reasons of obviousness as used above.
Regarding claim 25, Huang in view of Chen and Wu teaches the device of claim 17. Huang as modified teaches wherein the prediction is further based on contextual information about the current voxels to be encoded or decoded.
“The contextual information can include, but is not limited to, the octant value of the respective symbol, the location of the respective symbol, the level of the respective symbol in the tree-based data structure, and the occupancy value of the parent node associated with the respective symbol. This contextual information can then be used as input into a machine-learned model (a multi-layer perceptron) that can produce, as output, a statistical distribution for the expected occupancy for the respective node.” – Huang Col 2 Lines 66-67, Col 3, Lines 1-8
NOTE: Huang discloses determining contextual information for a node ang generating, using the contextual information as input to a neural network, a statistical distribution associated with the node. This functionally corresponds to the prediction based on contextual information about the current voxels to be encoded or decoded.
Regarding claim 9, the claim recites similar limitations to claim 25. Therefore, method claim 9 corresponds to the device disclosed in claim 25 and is rejected for the same reasons of obviousness as used above.
Regarding claim 26, Huang in view of Chen and Wu teaches the device of claim 17. Huang as modified teaches wherein the second feature is based on a first point cloud of the previous LoD.
“Then, starting with the feature h.sub.i.sup.(0) for each node, we perform K aggregations between the current node feature and the feature of its parent. At iteration k, the aggregation can also be modeled as an MLP: h.sub.i.sup.(k)=MLP.sup.(k)([h.sub.i.sup.k−1],[h.sub.pa(i).sup.k−1]) where h.sub.pa(i).sup.(k−1) is the hidden feature of node i's parent node." – Huang Col 6 Lines 65-66 and Col 7, Lines 1-4
NOTE: Huang discloses generating hidden features by aggregating the current node feature with a parent-node hidden feature. The already generated parent hidden features are accessed and used as contextual information to generate the current node hidden feature, see claim 17. Huang further discloses that the disclosure relates to receiving point cloud data and generating a tree-based data structure from the input point cloud which comprises a plurality of nodes in a hierarchical structure to represent the point cloud, see Col 1, Lines 41-49. This shows that the feature extracted from the parent-node which relates to the first point cloud and naturally would exist on a previous level compared to a current node.
Regarding claim 10, the claim recites similar limitations to claim 26. Therefore, method claim 10 corresponds to the device disclosed in claim 26 and is rejected for the same reasons of obviousness as used above.
Regarding claim 27, Huang in view of Chen and Wu teaches the device of claim 17. Huang as modified teaches wherein the occupancy bits of the neighboring voxels at the current LoD are received in a bitstream.
“Our entropy model estimates occupancy probability distribution for each non-empty octant, conditioning on ancestral dependence, neighbor dependence, sibling dependence, and surface priors. Finally, the octree symbols are encoded into a more compact bitstream with arithmetic coding.” – Chen Fig. 1 Caption
NOTE: Chen discloses that neighbor occupancy information is used as part of the octree occupancy-symbol bitstream. The decoding process: “Our entropy model can accurately predict symbol probabilities with these strong priors and further compress the symbols into a more compact bitstream by entropy coding. In the decoding stage, we first reconstruct an octree with L levels from the compressed bitstream.”, see Pg 4, Par 1 and 2, establishes reconstructing the point cloud based on the compressed bitstream. This functionally corresponds to the occupancy bits of the neighboring voxels at the current LoD are received in a bitstream.
Regarding claim 11, the claim recites similar limitations to claim 27. Therefore, method claim 11 corresponds to the device disclosed in claim 27 and is rejected for the same reasons of obviousness as used above.
Regarding claim 30, Huang in view of Chen and Wu teaches the device of claim 17. Huang as modified teaches wherein the third neural network module is a fully connected (FC) module.
“Specifically, the entropy model can include multiple MLP layers. The first MLP can be a 5-layer MLP with 128-dimensional hidden features. All subsequent MLPs can be 3-layer MLPs (with residual layers) with 128-dimensional hidden features. A final linear layer followed by a soft-max is used to make a 256-way prediction. The 256-way prediction can represent the likelihood, for a respective node of each potential eight-bit value, that the symbol associated with the respective node has that potential eight-bit value.” – Huang Col 6, Lines 23-31
NOTE: Huang discloses an entropy model with multiple MLP layers for processing the features and generating the occupancy probability distribution (which is what the third neural network is configured for). Specifically, Huang’s multi-layer perceptron system comprises multiple MLPs with varying number of layers and then goes through a final layer followed by a softmax to generate a prediction. This architecture is a conventional description of a fully connected (FC) module.
Regarding claim 14, the claim recites similar limitations to claim 30. Therefore, method claim 14 corresponds to the device disclosed in claim 30 and is rejected for the same reasons of obviousness as used above.
Claim(s) 12 and 28 is/are rejected under 35 U.S.C. 103 as being unpatentable over Huang, Chen, Wu, and Sinharoy et al (US 20190197739 A1), hereinafter Sinharoy.
Regarding claim 28, Huang in view of Chen and Wu teaches the device of claim 17
wherein the processor is further configured to reconstruct an octree structure based on the probability distribution of the occupancy of the voxels at the current LoD;
“As illustrated in Fig. 4, considering a reconstructed octree with L levels, we denote the set of leaf octants as NL. We first apply our entropy model at level L+1 to estimate the occupancy symbol probabilities for each leaf octant in NL. We take the predicted occupancy symbols SL+1 with top-1 accuracy to form a new set of octants NL+1 located at the predicted level L + 1. We denote the represented 3D coordinate of NL+1 as xp ∈ Rn× 3. In the second step, we apply the refinement module at level L + 1 [24] that takes neighbor context Vi and octant information ci of leaf as an input to predict offsets for the coordinates.” – Chen Pg 9, Section 3.5, Par 2
NOTE: Huang discloses reconstructing an octree: “In the decoding stage, we first reconstruct an octree with L levels from the compressed bitstream. We then propose a two-step heuristic strategy to produce a point cloud with better quality”, see Pg 4, Par 2, Lines 1-3. The entropy model as described estimates the occupancy symbol probabilities for child octants and uses the predicted occupancy symbols to form a new set of octants. This functionally corresponds to reconstructing an octree structure based on the probability distribution of the occupancy of voxels at a current LoD.
wherein the processor is further configured to reconstruct a second point cloud based on leaf nodes of the reconstructed octree structure;
“Each leaf octant at level L represents a subspace in the large 3D space. Because of the quantization, we can only recover a single point from each leaf octant by taking the center coordinate of its corresponding subspace… we approach this problem by first retrieving missing points with the help of our trained entropy model at level L + 1 and then further refine the aggregated points. By this strategy, we can reconstruct point clouds with the precision near to level L + 1 while keeping the bitrate at level L.” – Chen, Pg 9, Section 3.5, Par 1
NOTE: Chen discloses that each leaf octant represents a 3D subspace and that a point is recovered from each leaf octant (leaf node) using the center coordinates of its corresponding subspace. Chen then retrieves additional points at the next level using the leaf octants and refines their coordinates to produce a reconstructed point cloud (second point cloud). This functionally corresponds to reconstructing a second point cloud based on leaf nodes of the reconstructed octree structure.
Huang as modifies still does not teach wherein the processor and the transceiver are further configured to transmit bits of the second point cloud. However, Sinharoy teaches wherein the processor and the transceiver are further configured to transmit bits of the second point cloud.
“identify duplicate points in at least one of the two or more geometry frames based on the identified depth values of the corresponding pixels in the two or more geometry frames; and remove or ignore the identified duplicate points while reconstructing the 3D point cloud data; and a communication interface configured to transmit the encoded bit stream comprising the 3D point cloud data.” – Claim 8
NOTE: Sinharoy discloses reconstructing 3D point cloud data as part of its encoding process to then transmit an encoded bitstream comprising the 3D point cloud data using a communication interface. After the combination, the point cloud encoding and transmission process as taught by Sinharoy can be applied to the second reconstructed point cloud as obtained using the process of Chen. This modification is then added to Huang’s device so that the reconstructed second point cloud is encoded into a bitstream and then the resulting bits can be transmitted
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Huang by incorporating the teachings of Sinharoy to have the processor and transceiver transmit the bits of the second point cloud. One would be motivated to apply this well-known encoding and transmission technique to the reconstructed second point cloud because it would lead to the predicted result of communicating the reconstructed point cloud data to another device.
Regarding claim 12, the claim recites similar limitations to claim 28. Therefore, method claim 12 corresponds to the device disclosed in claim 28 and is rejected for the same reasons of obviousness as used above.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DAVID V. NGUYEN whose telephone number is (571)272-6111. The examiner can normally be reached M-F 9:00-5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, King Y Poon can be reached at 571-270-0728. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/DAVID VAN NGUYEN/Examiner, Art Unit 2617 /KING Y POON/Supervisory Patent Examiner, Art Unit 2617