Prosecution Insights
Last updated: September 17, 2026
Application No. 18/931,102

NEURAL NETWORK POINT CLOUD DATA ANALYZING METHOD AND COMPUTER PROGRAM PRODUCT

Non-Final OA §101§103
Filed
Oct 30, 2024
Priority
Nov 03, 2023 — TW 112142509
Examiner
YANG, WEI WEN
Art Unit
Tech Center
Assignee
National Changhua University Of Education
OA Round
1 (Non-Final)
82%
Grant Probability
Favorable
1-2
OA Rounds
7m
Est. Remaining
93%
With Interview

Examiner Intelligence

Grants 82% — above average
82%
Career Allowance Rate
559 granted / 682 resolved
+22.0% vs TC avg
Moderate +11% lift
Without
With
+11.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 5m
Avg Prosecution
31 currently pending
Career history
705
Total Applications
across all art units

Statute-Specific Performance

§101
7.9%
-32.1% vs TC avg
§103
74.9%
+34.9% vs TC avg
§102
9.2%
-30.8% vs TC avg
§112
7.9%
-32.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 682 resolved cases

Office Action

§101 §103
DETAILED ACTION Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 7-10 are rejected under 35 U.S.C. 101 as not falling within one of the four statutory categories of invention because: claims 7-10 recite “computer program product….”, which does not have a physical or tangible form, and may include signal such as carrier waves {subject matter that does not fall within a statutory category}, the claims as a whole were not to a statutory category and thus failed the first criterion for eligibility. See MPEP 2106(I). A claim directed toward “a non-transitory computer-readable storage medium ….” would establish a sufficient functional relationship between the instructions/program and a computer/processor. MPEP 2111.05(III). Hence, amending the limitation to “a non-transitory computer-readable storage medium ….….”, would resolve this issue for claims 7-10; Appropriate correction is required. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103(a) which forms the basis for all obviousness rejections set forth in this Office action: (a) A patent may not be obtained though the invention is not identically disclosed or described as set forth in section 102 of this title, if the differences between the subject matter sought to be patented and the prior art are such that the subject matter as a whole would have been obvious at the time the invention was made to a person having ordinary skill in the art to which said subject matter pertains. Patentability shall not be negatived by the manner in which the invention was made. Claims 1-10 are rejected under 35 U.S.C. 103(a) as being unpatentable over Maturana et al. (VoxNet: A 3D Convolutional Neural Network for Real-Time Object Recognition, 2015 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Sept 28 - Oct 2, 2015; pages 922-928), and in view of SEDAGHAT (Orientation-boosted Voxel Nets for 3D Object Recognition, BMVC 2017, 19 OCT. 2017, pages 1-18), and further in view of IANCU (US 20220114807 A1), and in view of ZHANG (US 20230363610 A1). Re Claim 1, Maturana et al. discloses a neural network point cloud data analyzing method (see Maturana: e.g., -- VoxNet, an architecture to tackle this problem by integrating a volumetric Occupancy Grid representation with a supervised 3D Convolutional Neural Network (3D CNN). We evaluate our approach on publicly available benchmarks using LiDAR, RGBD, and CAD data.--, in abstract), comprising: a data input step, wherein a processor receives a point cloud data (see Maturana: e.g., Fig. 1, and, --The input to our algorithm is a point cloud segment, which can originate from segmentation methods such as [12], [29], or a “sliding box” if performing detection. The segment is usually given by the intersection of a point cloud with a bounding box and may include background clutter. Our task is to predict an object class label for the segment.--; under III. APPROACH, left col., page 923); and an analyzing step, wherein the processor analyzes the point cloud data based on an VoxNet 3D model (see Maturana: e.g., -- VoxNet, an architecture to tackle this problem by integrating a volumetric Occupancy Grid representation with a supervised 3D Convolutional Neural Network (3D CNN). We evaluate our approach on publicly available benchmarks using LiDAR, RGBD, and CAD data.--, in abstract; and, Fig. 2, --A. Volumetric Occupancy Grid Occupancy grids ([30], [31]) represent the state of the environment as a 3D lattice of random variables (each corresponding to a voxel) and maintain a probabilistic estimate of their occupancy as a function of incoming sensor data and prior knowledge. There are two main reasons we use occupancy grids. First, they allow us to efficiently estimate free, occupied and unknown space from range measurements, even for measurements coming from different viewpoints and time instants. This representation is richer than those which only consider occupied space versus free space such as point clouds, as the distinction between free and unknown space can potentially be a valuable shape cue. Second, they can be stored and manipulated with simple and efficient data structures. In this work, we use dense arrays to perform all our CNN processing, as we use small volumes (323 voxels) and GPUs work best with dense data. To keep larger spatial extents in memory we use hierarchical data structures and copy specific segments to dense arrays as needed. Theoretically this allows us to store a potentially unbounded volume while using small occupancy grids for CNN processing.--, in page 923); Maturana however does not explicitly disclose enhanced VoxNet 3D model; Sedaghat discloses an enhanced VoxNet 3D model (see Sedaghat: e.g., Fig. 2, and, -- 3 Method The core network architecture is based on VoxNet [21] and is illustrated in Fig. 2. It takes a 3D voxel grid as input and contains two convolutional layers with 3D filters followed by two fully connected layers. Although this choice may not be optimal, we keep it to be able to directly compare our modifications to VoxNet. In addition, we experimented with a slightly deeper network that has four convolutional layers. Point clouds and CAD models are converted to voxel grids (occupancy grids). For the NYUv2 dataset we used the provided tools for the conversion; for the other datasets we implemented our own version. We tried both binary-valued and continuous-valued occupancy grids. In the end, the difference in the results was negligible and thus we only report the results of the former one. Multi-task learning We modify the baseline architecture by adding orientation estimation as an auxiliary parallel task. We call the resulting architecture the ORIentation-boosted vOxel Net – ORION. Without loss of generality, we only consider rotation around the z-axis (azimuth) as the most varying component of orientation in practical applications. Throughout this paper we use the term ’orientation’ to refer to this component.--, in page 4); Maturana and SEDAGHAT are combinable as they are in the same field of endeavor: using VOXNET 3D model for object recognition and classification from point cloud data. Therefore it would have been obvious to one of ordinary skill in the art at the time the invention was made to further modify Maturana’s method using SEDAGHAT’s teachings by including an enhanced VoxNet 3D model to Maturana’s VoxNet 3D model in order to achieved better classification results (see SEDAGHAT: e.g., at abstract, and in page 4); Maturana as modified by SEDAGHAT further disclose the enhanced VoxNet 3D model comprises: an input layer for inputting the point cloud data (see Maturana: e.g., Fig. 1, and, --The input to our algorithm is a point cloud segment, which can originate from segmentation methods such as [12], [29], or a “sliding box” if performing detection. The segment is usually given by the intersection of a point cloud with a bounding box and may include background clutter. Our task is to predict an object class label for the segment.--; under III. APPROACH, left col., page 923also see Sedaghat: e.g., Fig. 2, and, -- 3 Method The core network architecture is based on VoxNet [21] and is illustrated in Fig. 2. It takes a 3D voxel grid as input and contains two convolutional layers with 3D filters followed by two fully connected layers. Although this choice may not be optimal, we keep it to be able to directly compare our modifications to VoxNet. In addition, we experimented with a slightly deeper network that has four convolutional layers. Point clouds and CAD models are converted to voxel grids (occupancy grids). For the NYUv2 dataset we used the provided tools for the conversion; for the other datasets we implemented our own version. We tried both binary-valued and continuous-valued occupancy grids. In the end, the difference in the results was negligible and thus we only report the results of the former one. Multi-task learning We modify the baseline architecture by adding orientation estimation as an auxiliary parallel task. We call the resulting architecture the ORIentation-boosted vOxel Net – ORION. Without loss of generality, we only consider rotation around the z-axis (azimuth) as the most varying component of orientation in practical applications. Throughout this paper we use the term ’orientation’ to refer to this component.--, in page 4); Maturana as modified by SEDAGHAT however still do not explicitly disclose a first hidden unit signally connected to the input layer and comprising a first convolution layer, a first batch-normalization layer and a first activation layer in order; IANCU discloses a first hidden unit signally connected to the input layer (see IANCU: e.g., -- [0011] A neural network may include multiple layers of nodes. The layers may include an input layer, an output layer, and hidden layers in-between. The calculations of the neural network are propagated from the input layer through the hidden layers to the output layer. Each layer may include nodes associated with node values calculated from a prior layer through edges connecting nodes between the present layer and the prior layer. Edges may connect the nodes in a layer to nodes in an adjacent layer. Each edge may be associated with a weight value. Therefore, the node values associated with nodes of the present layer can be a weighed summation of the node values of the prior layer. [0012] One type of the neural networks is the convolutional neural networks (CNNs) where the calculation performed at the hidden layers can be convolutions of node values associated with the prior layer and weight values associated with edges. For example, a processing device may apply convolution operations to the input layer and generate the node values for the first hidden layer connected to the input layer through edges, and apply convolution operations to the first hidden layer to generate node values for the second hidden layer, and so on until the calculation reaches the output layer. The processing device may apply a soft combination operation to the output data and generate a detection result. The detection result may include the identities of the detected objects and their locations.--, in [0011]-[0012]); and Maturana (as modified by SEDAGHAT) and IANCU are combinable as they are in the same field of endeavor: object detection, recognition and classification from image/video data, Therefore it would have been obvious to one of ordinary skill in the art at the time the invention was made to further modify Maturana (as modified by SEDAGHAT)’s method using IANCU’s teachings by including a first hidden unit signally connected to the input layer to Maturana (as modified by SEDAGHAT)’s convolutional neural network (CNN) in order to apply convolution operations to the input layer and generate the node values for the first hidden layer connected to the input layer through edges (see IANCU: e.g., in [0011]-[0012]); Maturana as modified by SEDAGHAT and IANCU however still do not explicitly disclose a first hidden unit comprising a first convolution layer, a first batch-normalization layer and a first activation layer in order; Zhang discloses a first hidden unit comprising a first convolution layer, a first batch-normalization layer and a first activation layer in order (see Zhang: e.g., -- [0123] A convolutional neural network is a type of neural network. The fundamental difference between a densely connected layer and a convolution layer is this: Dense layers learn global patterns in their input feature space, whereas convolution layers learn local patters: in the case of images, patterns found in small 2D windows of the inputs. This key characteristic gives convolutional neural networks two interesting properties: (1) the patterns they learn are translation invariant and (2) they can learn spatial hierarchies of patterns. [0124] Regarding the first, after learning a certain pattern in the lower-right corner of a picture, a convolution layer can recognize it anywhere: for example, in the upper-left corner. A densely connected network would have to learn the pattern anew if it appeared at a new location. This makes convolutional neural networks data efficient because they need fewer training samples to learn representations, and they have generalization power. [0125] Regarding the second, a first convolution layer can learn small local patterns such as edges, a second convolution layer will learn larger patterns made of the features of the first layers, and so on. This allows convolutional neural networks to efficiently learn increasingly complex and abstract visual concepts. [0126] A convolutional neural network learns highly non-linear mappings by interconnecting layers of artificial neurons arranged in many different layers with activation functions that make the layers dependent. It includes one or more convolutional layers, interspersed with one or more sub-sampling layers and non-linear layers, which are typically followed by one or more fully connected layers. Each element of the convolutional neural network receives inputs from a set of features in the previous layer. The convolutional neural network learns concurrently because the neurons in the same feature map have identical weights. These local shared weights reduce the complexity of the network such that when multi-dimensional input data enters the network, the convolutional neural network avoids the complexity of data reconstruction in feature extraction and regression or classification process.--, in [0123]-[0126]; and, -- [0139] The algorithm includes computing the activation of all neurons in the network, yielding an output for the forward pass. The activation of neuron m in the hidden layers …[0140] This is done for all the hidden layers to get the activation--, in [0139]-[0140], and, -- Batch Normalization [0187] Batch normalization is a method for accelerating deep network training by making data standardization an integral part of the network architecture. Batch normalization can adaptively normalize data even as the mean and variance change over time during training. It works by internally maintaining an exponential moving average of the batch-wise mean and variance of the data seen during training. The main effect of batch normalization is that it helps with gradient propagation—much like residual connections—and thus allows for deep networks. Some very deep networks can only be trained if they include multiple Batch Normalization layers. [0188] Batch normalization can be seen as yet another layer that can be inserted into the model architecture, just like the fully connected or convolutional layer. The BatchNormalization layer is typically used after a convolutional or densely connected layer. It can also be used before a convolutional or densely connected layer. Both implementations can be used by the technology disclosed and are shown in FIG. 17. The BatchNormalization layer takes an axis argument, which specifies the feature axis that should be normalized. This argument defaults to −1, the last axis in the input tensor. This is the correct value when using Dense layers, Conv1D layers, RNN layers, and Conv2D layers with data_format set to “channels_last”. But in the niche use case of Conv2D layers with data_format set to “channels_first”, the features axis is axis 1; the axis argument in BatchNormalization can be set to 1. [0189] Batch normalization provides a definition for feed-forwarding the input and computing the gradients with respect to the parameters and its own input via a backward pass. In practice, batch normalization layers are inserted after a convolutional or fully connected layer, but before the outputs are fed into an activation function. For convolutional layers, the different elements of the same feature map—i.e., the activations—at different locations are normalized in the same way in order to obey the convolutional property. Thus, all activations in a mini-batch are normalized over all locations, rather than per activation.--, in [0187]-[0189]); Maturana (as modified by SEDAGHAT and IANCU) and ZHANG are combinable as they are in the same field of endeavor: object detection, recognition and classification from image/video data, Therefore it would have been obvious to one of ordinary skill in the art at the time the invention was made to further modify Maturana (as modified by SEDAGHAT and IANCU)’s method using ZHANG’s teachings by including a first hidden unit comprising a first convolution layer, a first batch-normalization layer and a first activation layer in order to Maturana (as modified by SEDAGHAT and IANCU)’s convolutional neural network (CNN) {such as in Fig. 1 of Maturana) in order to interconnect layers of artificial neurons arranged in many different layers with activation functions that make the layers dependent, and accelerat deep network training by making data standardization an integral part of the network architecture (see ZHANG: e.g., in [0123]-[0126], [0139]-[0140], and [0187]-[0189]); Maturana as modified by SEDAGHAT and IANCU and ZHANG further disclose a first pooling layer receiving an output from the first activation layer (see Maturana: e.g., --Pooling Layers (Pm). These layers downsample the input volume by a factor of by m along the spatial dimensions by replacing each mxmxm non-overlapping block of voxels with their maximum.--, in right col., page 924, also see Zhang: e.g., -- [0164] FIG. 10 depicts one implementation of a two-layer convolution of the convolution layers. In FIG. 10, an input of size 2048 dimensions is convolved. At convolution 1, the input is convolved by a convolutional layer comprising of two channels of sixteen kernels of size 3×3. The resulting sixteen feature maps are then rectified by means of the ReLU activation function at ReLU1 and then pooled in Pool 1 by means of average pooling using a sixteen channel pooling layer with kernels of size 3×3.--, in [0164]); a 1st to a 3rd second hidden units sequentially connected after the first pooling layer, each of the 1st to the 3rd second hidden units comprising a second convolution layer, a second batch-normalization layer, a second activation layer, a third convolution layer, a third batch-normalization layer, an adding layer and a third activation layer in order, wherein the adding layer receives and summarizes an output from the second batch-normalization layer and an output from the third batch-normalization layer (see Sedaghat: e.g., Fig. 2, and, -- 3 Method The core network architecture is based on VoxNet [21] and is illustrated in Fig. 2. It takes a 3D voxel grid as input and contains two convolutional layers with 3D filters followed by two fully connected layers. Although this choice may not be optimal, we keep it to be able to directly compare our modifications to VoxNet. In addition, we experimented with a slightly deeper network that has four convolutional layers. Point clouds and CAD models are converted to voxel grids (occupancy grids). For the NYUv2 dataset we used the provided tools for the conversion; for the other datasets we implemented our own version. We tried both binary-valued and continuous-valued occupancy grids. In the end, the difference in the results was negligible and thus we only report the results of the former one. Multi-task learning We modify the baseline architecture by adding orientation estimation as an auxiliary parallel task. We call the resulting architecture the ORIentation-boosted vOxel Net – ORION. Without loss of generality, we only consider rotation around the z-axis (azimuth) as the most varying component of orientation in practical applications. Throughout this paper we use the term ’orientation’ to refer to this component.--, in page 4; also see Zhang: e.g., -- [0123] A convolutional neural network is a type of neural network. The fundamental difference between a densely connected layer and a convolution layer is this: Dense layers learn global patterns in their input feature space, whereas convolution layers learn local patters: in the case of images, patterns found in small 2D windows of the inputs. This key characteristic gives convolutional neural networks two interesting properties: (1) the patterns they learn are translation invariant and (2) they can learn spatial hierarchies of patterns. [0124] Regarding the first, after learning a certain pattern in the lower-right corner of a picture, a convolution layer can recognize it anywhere: for example, in the upper-left corner. A densely connected network would have to learn the pattern anew if it appeared at a new location. This makes convolutional neural networks data efficient because they need fewer training samples to learn representations, and they have generalization power. [0125] Regarding the second, a first convolution layer can learn small local patterns such as edges, a second convolution layer will learn larger patterns made of the features of the first layers, and so on. This allows convolutional neural networks to efficiently learn increasingly complex and abstract visual concepts. [0126] A convolutional neural network learns highly non-linear mappings by interconnecting layers of artificial neurons arranged in many different layers with activation functions that make the layers dependent. It includes one or more convolutional layers, interspersed with one or more sub-sampling layers and non-linear layers, which are typically followed by one or more fully connected layers. Each element of the convolutional neural network receives inputs from a set of features in the previous layer. The convolutional neural network learns concurrently because the neurons in the same feature map have identical weights. These local shared weights reduce the complexity of the network such that when multi-dimensional input data enters the network, the convolutional neural network avoids the complexity of data reconstruction in feature extraction and regression or classification process.--, in [0123]-[0126]; and, -- [0139] The algorithm includes computing the activation of all neurons in the network, yielding an output for the forward pass. The activation of neuron m in the hidden layers …[0140] This is done for all the hidden layers to get the activation--, in [0139]-[0140], and, -- Batch Normalization [0187] Batch normalization is a method for accelerating deep network training by making data standardization an integral part of the network architecture. Batch normalization can adaptively normalize data even as the mean and variance change over time during training. It works by internally maintaining an exponential moving average of the batch-wise mean and variance of the data seen during training. The main effect of batch normalization is that it helps with gradient propagation—much like residual connections—and thus allows for deep networks. Some very deep networks can only be trained if they include multiple Batch Normalization layers. [0188] Batch normalization can be seen as yet another layer that can be inserted into the model architecture, just like the fully connected or convolutional layer. The BatchNormalization layer is typically used after a convolutional or densely connected layer. It can also be used before a convolutional or densely connected layer. Both implementations can be used by the technology disclosed and are shown in FIG. 17. The BatchNormalization layer takes an axis argument, which specifies the feature axis that should be normalized. This argument defaults to −1, the last axis in the input tensor. This is the correct value when using Dense layers, Conv1D layers, RNN layers, and Conv2D layers with data_format set to “channels_last”. But in the niche use case of Conv2D layers with data_format set to “channels_first”, the features axis is axis 1; the axis argument in BatchNormalization can be set to 1. [0189] Batch normalization provides a definition for feed-forwarding the input and computing the gradients with respect to the parameters and its own input via a backward pass. In practice, batch normalization layers are inserted after a convolutional or fully connected layer, but before the outputs are fed into an activation function. For convolutional layers, the different elements of the same feature map—i.e., the activations—at different locations are normalized in the same way in order to obey the convolutional property. Thus, all activations in a mini-batch are normalized over all locations, rather than per activation.--, in [0187]-[0189]); a second pooling layer connected between the third activation layer of the 1st second hidden unit and the second convolution layer of a 2nd second hidden unit of the 1st to the 3rd second hidden units (see Maturana: e.g., --Pooling Layers (Pm). These layers downsample the input volume by a factor of by m along the spatial dimensions by replacing each mxmxm non-overlapping block of voxels with their maximum.--, in right col., page 924, also see Zhang: e.g., -- [0164] FIG. 10 depicts one implementation of a two-layer convolution of the convolution layers. In FIG. 10, an input of size 2048 dimensions is convolved. At convolution 1, the input is convolved by a convolutional layer comprising of two channels of sixteen kernels of size 3×3. The resulting sixteen feature maps are then rectified by means of the ReLU activation function at ReLU1 and then pooled in Pool 1 by means of average pooling using a sixteen channel pooling layer with kernels of size 3×3. At convolution 2, the output of Pool 1 is then convolved by another convolutional layer comprising of sixteen channels of thirty kernels with a size of 3×3. This is followed by yet another ReLU2 and average pooling in Pool 2 with a kernel size of 2×2. The convolution layers use varying number of strides and padding, for example, zero, one, two and three. The resulting feature vector is five hundred and twelve (512) dimensions, according to one implementation.--, in [0164]); a third pooling layer connected between the third activation layer of the 2nd second hidden unit and the second convolution layer of the 3rd second hidden unit (see Maturana: e.g., --Pooling Layers (Pm). These layers downsample the input volume by a factor of by m along the spatial dimensions by replacing each mxmxm non-overlapping block of voxels with their maximum.--, in right col., page 924, also see Zhang: e.g., -- [0158] FIG. 9 is one implementation of sub-sampling layers in accordance with one implementation of the technology disclosed. Sub-sampling layers reduce the resolution of the features extracted by the convolution layers to make the extracted features or feature maps—robust against noise and distortion. In one implementation, sub-sampling layers employ two types of pooling operations, average pooling and max pooling. The pooling operations divide the input into non-overlapping two-dimensional spaces. For average pooling, the average of the four values in the region is calculated. For max pooling, the maximum value of the four values is selected. [0159] In one implementation, the sub-sampling layers include pooling operations on a set of neurons in the previous layer by mapping its output to only one of the inputs in max pooling and by mapping its output to the average of the input in average pooling. In max pooling, the output of the pooling neuron is the maximum value that resides within the input… [0163] In FIG. 9, the input is of size 4×4. For 2×2 sub-sampling, a 4×4 image is divided into four non-overlapping matrices of size 2×2. For average pooling, the average of the four values is the whole-integer output. For max pooling, the maximum value of the four values in the 2×2 matrix is the whole-integer output.--, in [0158]-[0163]; and, -- [0164] FIG. 10 depicts one implementation of a two-layer convolution of the convolution layers. In FIG. 10, an input of size 2048 dimensions is convolved. At convolution 1, the input is convolved by a convolutional layer comprising of two channels of sixteen kernels of size 3×3. The resulting sixteen feature maps are then rectified by means of the ReLU activation function at ReLU1 and then pooled in Pool 1 by means of average pooling using a sixteen channel pooling layer with kernels of size 3×3. At convolution 2, the output of Pool 1 is then convolved by another convolutional layer comprising of sixteen channels of thirty kernels with a size of 3×3. This is followed by yet another ReLU2 and average pooling in Pool 2 with a kernel size of 2×2. The convolution layers use varying number of strides and padding, for example, zero, one, two and three. The resulting feature vector is five hundred and twelve (512) dimensions, according to one implementation.--, in [0164]; and, --Global Average Pooling [0198] FIG. 19 illustrates how global average pooling (GAP) works. Global average pooling can be use used to replace fully connected (FC) layers for classification, by taking the spatial average of features in the last layer for scoring. The reduces the training load and bypasses overfitting issues. Global average pooling applies a structural prior to the model and it is equivalent to linear transformation with predefined weights. Global average pooling reduces the number of parameters and eliminates the fully connected layer. Fully connected layers are typically the most parameter and connection intensive layers, and global average pooling provides much lower-cost approach to achieve similar results. The main idea of global average pooling is to generate the average value from each last layer feature map as the confidence factor for scoring, feeding directly into the softmax layer. [0199] Global average pooling have three benefits: (1) there are no extra parameters in global average pooling layers thus overfitting is avoided at global average pooling layers; (2) since the output of global average pooling is the average of the whole feature map, global average pooling will be more robust to spatial translations; and (3) because of the huge number of parameters in fully connected layers which usually take over 50% in all the parameters of the whole network, replacing them by global average pooling layers can significantly reduce the size of the model, and this makes global average pooling very useful in model compression. [0200] Global average pooling makes sense, since stronger features in the last layer are expected to have a higher average value. In some implementations, global average pooling can be used as a proxy for the classification score. The feature maps under global average pooling can be interpreted as confidence maps, and force correspondence between the feature maps and the categories. Global average pooling can be particularly effective if the last layer features are at a sufficient abstraction for direct classification; however, global average pooling alone is not enough if multilevel features should be combined into groups like parts models, which is best performed by adding a simple fully connected layer or other classifier after the global average pooling.--, in [0198]-[0200]); and a fourth pooling layer connected after the third activation layer of the 3rd second hidden unit (see SEDAGHAT: e.g., Table 10, Pool4 after Conv4, in page 18; also see Zhang: e.g., -- [0164] FIG. 10 depicts one implementation of a two-layer convolution of the convolution layers. In FIG. 10, an input of size 2048 dimensions is convolved. At convolution 1, the input is convolved by a convolutional layer comprising of two channels of sixteen kernels of size 3×3. The resulting sixteen feature maps are then rectified by means of the ReLU activation function at ReLU1 and then pooled in Pool 1 by means of average pooling using a sixteen channel pooling layer with kernels of size 3×3. At convolution 2, the output of Pool 1 is then convolved by another convolutional layer comprising of sixteen channels of thirty kernels with a size of 3×3. This is followed by yet another ReLU2 and average pooling in Pool 2 with a kernel size of 2×2. The convolution layers use varying number of strides and padding, for example, zero, one, two and three. The resulting feature vector is five hundred and twelve (512) dimensions, according to one implementation.--, in [0164]; and, --Global Average Pooling [0198] FIG. 19 illustrates how global average pooling (GAP) works. Global average pooling can be use used to replace fully connected (FC) layers for classification, by taking the spatial average of features in the last layer for scoring. The reduces the training load and bypasses overfitting issues. Global average pooling applies a structural prior to the model and it is equivalent to linear transformation with predefined weights. Global average pooling reduces the number of parameters and eliminates the fully connected layer. Fully connected layers are typically the most parameter and connection intensive layers, and global average pooling provides much lower-cost approach to achieve similar results. The main idea of global average pooling is to generate the average value from each last layer feature map as the confidence factor for scoring, feeding directly into the softmax layer. [0199] Global average pooling have three benefits: (1) there are no extra parameters in global average pooling layers thus overfitting is avoided at global average pooling layers; (2) since the output of global average pooling is the average of the whole feature map, global average pooling will be more robust to spatial translations; and (3) because of the huge number of parameters in fully connected layers which usually take over 50% in all the parameters of the whole network, replacing them by global average pooling layers can significantly reduce the size of the model, and this makes global average pooling very useful in model compression. [0200] Global average pooling makes sense, since stronger features in the last layer are expected to have a higher average value. In some implementations, global average pooling can be used as a proxy for the classification score. The feature maps under global average pooling can be interpreted as confidence maps, and force correspondence between the feature maps and the categories. Global average pooling can be particularly effective if the last layer features are at a sufficient abstraction for direct classification; however, global average pooling alone is not enough if multilevel features should be combined into groups like parts models, which is best performed by adding a simple fully connected layer or other classifier after the global average pooling.--, in [0198]-[0200]). Re Claims 2, Maturana as modified by SEDAGHAT, IANCU and ZHANG further disclose wherein the enhanced VoxNet 3D model further comprises a first fully connected layer, a fourth batch-normalization layer, a fourth activation layer, and a dropout layer connected after the fourth pooling layer in order (see SEDAGHAT: e.g., Fig. 2, and, -- 3 Method The core network architecture is based on VoxNet [21] and is illustrated in Fig. 2. It takes a 3D voxel grid as input and contains two convolutional layers with 3D filters followed by two fully connected layers. Although this choice may not be optimal, we keep it to be able to directly compare our modifications to VoxNet. In addition, we experimented with a slightly deeper network that has four convolutional layers. Point clouds and CAD models are converted to voxel grids (occupancy grids). For the NYUv2 dataset we used the provided tools for the conversion; for the other datasets we implemented our own version. We tried both binary-valued and continuous-valued occupancy grids. In the end, the difference in the results was negligible and thus we only report the results of the former one. Multi-task learning We modify the baseline architecture by adding orientation estimation as an auxiliary parallel task. We call the resulting architecture the ORIentation-boosted vOxel Net – ORION. Without loss of generality, we only consider rotation around the z-axis (azimuth) as the most varying component of orientation in practical applications. Throughout this paper we use the term ’orientation’ to refer to this component.--, in page 4; and, Table 10, Pool4 after Conv4, in page 18; also see Zhang: e.g., -- [0164] FIG. 10 depicts one implementation of a two-layer convolution of the convolution layers. In FIG. 10, an input of size 2048 dimensions is convolved. At convolution 1, the input is convolved by a convolutional layer comprising of two channels of sixteen kernels of size 3×3. The resulting sixteen feature maps are then rectified by means of the ReLU activation function at ReLU1 and then pooled in Pool 1 by means of average pooling using a sixteen channel pooling layer with kernels of size 3×3. At convolution 2, the output of Pool 1 is then convolved by another convolutional layer comprising of sixteen channels of thirty kernels with a size of 3×3. This is followed by yet another ReLU2 and average pooling in Pool 2 with a kernel size of 2×2. The convolution layers use varying number of strides and padding, for example, zero, one, two and three. The resulting feature vector is five hundred and twelve (512) dimensions, according to one implementation.--, in [0164]; and, --Global Average Pooling [0198] FIG. 19 illustrates how global average pooling (GAP) works. Global average pooling can be use used to replace fully connected (FC) layers for classification, by taking the spatial average of features in the last layer for scoring. The reduces the training load and bypasses overfitting issues. Global average pooling applies a structural prior to the model and it is equivalent to linear transformation with predefined weights. Global average pooling reduces the number of parameters and eliminates the fully connected layer. Fully connected layers are typically the most parameter and connection intensive layers, and global average pooling provides much lower-cost approach to achieve similar results. The main idea of global average pooling is to generate the average value from each last layer feature map as the confidence factor for scoring, feeding directly into the softmax layer. [0199] Global average pooling have three benefits: (1) there are no extra parameters in global average pooling layers thus overfitting is avoided at global average pooling layers; (2) since the output of global average pooling is the average of the whole feature map, global average pooling will be more robust to spatial translations; and (3) because of the huge number of parameters in fully connected layers which usually take over 50% in all the parameters of the whole network, replacing them by global average pooling layers can significantly reduce the size of the model, and this makes global average pooling very useful in model compression. [0200] Global average pooling makes sense, since stronger features in the last layer are expected to have a higher average value. In some implementations, global average pooling can be used as a proxy for the classification score. The feature maps under global average pooling can be interpreted as confidence maps, and force correspondence between the feature maps and the categories. Global average pooling can be particularly effective if the last layer features are at a sufficient abstraction for direct classification; however, global average pooling alone is not enough if multilevel features should be combined into groups like parts models, which is best performed by adding a simple fully connected layer or other classifier after the global average pooling.--, in [0198]-[0200]). Re Claims 3, Maturana as modified by SEDAGHAT, IANCU and ZHANG further disclose wherein the enhanced VoxNet 3D model further comprises a second fully connected layer, a softmax layer and a CrossEntropyLoss layer connected after the dropout layer in order (see -- Fully Connected Layer FC(n). Fully connected layers have n output neurons. The output of each neuron is a learned linear combination of all the outputs from the previous layer, passed through a nonlinearity. We use ReLUs save for the final output layer, where the number of outputs corresponds to the number of class labels and a softmax nonlinearity is used to provide a probabilistic output.--, in page 924, and, -- we implemented a multiresolution VoxNet, inspired by the “foveal” architecture of [24] for video analysis. In this model we use two networks with an identical VoxNet architectures, each receiving occupancy grids at different resolutions: (0:1 m)3 and (0:2 m)3. Both inputs are centered on the same location, but the coarser network covers a larger area at low resolution while the finer network covers a smaller area at high resolution. To fuse the information from both networks, we concatenate the outputs of their respective FC(128) layers and connect them to a softmax output layer.--, in page 925; and, see SEDAGHAT: e.g., --We choose multinomial cross-entropy losses [25] for both tasks, so we can combine them by summing them up: L = (1􀀀g)LC +gLO (2) where LC and LO indicate losses for object classification and orientation estimation tasks respectively. We used equal loss weights (g = 0:5) and found in our classification experiments that the results do not depend on the exact choice of the weight g around this value. However, in one of the detection experiments, where the orientation estimation is not an auxiliary task anymore, we used a higher weight for the orientation output to improve its accuracy--, in page 5 Re Claims 4, Maturana as modified by SEDAGHAT, IANCU and ZHANG further disclose wherein an output size of the second pooling layer is smaller than an output size of the first pooling layer, an output size of the third pooling layer is smaller than the output size of the second pooling layer, and an output size of the fourth pooling layer is smaller than the output size of the third pooling layer and is equal to 1 (see SEDAGHAT: e.g., Fig. 2, and, -- 3 Method The core network architecture is based on VoxNet [21] and is illustrated in Fig. 2. It takes a 3D voxel grid as input and contains two convolutional layers with 3D filters followed by two fully connected layers. Although this choice may not be optimal, we keep it to be able to directly compare our modifications to VoxNet. In addition, we experimented with a slightly deeper network that has four convolutional layers. Point clouds and CAD models are converted to voxel grids (occupancy grids). For the NYUv2 dataset we used the provided tools for the conversion; for the other datasets we implemented our own version. We tried both binary-valued and continuous-valued occupancy grids. In the end, the difference in the results was negligible and thus we only report the results of the former one. Multi-task learning We modify the baseline architecture by adding orientation estimation as an auxiliary parallel task. We call the resulting architecture the ORIentation-boosted vOxel Net – ORION. Without loss of generality, we only consider rotation around the z-axis (azimuth) as the most varying component of orientation in practical applications. Throughout this paper we use the term ’orientation’ to refer to this component.--, in page 4; and, Table 10, Pool4 after Conv4, in page 18; also see Zhang: e.g., -- [0164] FIG. 10 depicts one implementation of a two-layer convolution of the convolution layers. In FIG. 10, an input of size 2048 dimensions is convolved. At convolution 1, the input is convolved by a convolutional layer comprising of two channels of sixteen kernels of size 3×3. The resulting sixteen feature maps are then rectified by means of the ReLU activation function at ReLU1 and then pooled in Pool 1 by means of average pooling using a sixteen channel pooling layer with kernels of size 3×3. At convolution 2, the output of Pool 1 is then convolved by another convolutional layer comprising of sixteen channels of thirty kernels with a size of 3×3. This is followed by yet another ReLU2 and average pooling in Pool 2 with a kernel size of 2×2. The convolution layers use varying number of strides and padding, for example, zero, one, two and three. The resulting feature vector is five hundred and twelve (512) dimensions, according to one implementation.--, in [0164]; and, --Global Average Pooling [0198] FIG. 19 illustrates how global average pooling (GAP) works. Global average pooling can be use used to replace fully connected (FC) layers for classification, by taking the spatial average of features in the last layer for scoring. The reduces the training load and bypasses overfitting issues. Global average pooling applies a structural prior to the model and it is equivalent to linear transformation with predefined weights. Global average pooling reduces the number of parameters and eliminates the fully connected layer. Fully connected layers are typically the most parameter and connection intensive layers, and global average pooling provides much lower-cost approach to achieve similar results. The main idea of global average pooling is to generate the average value from each last layer feature map as the confidence factor for scoring, feeding directly into the softmax layer. [0199] Global average pooling have three benefits: (1) there are no extra parameters in global average pooling layers thus overfitting is avoided at global average pooling layers; (2) since the output of global average pooling is the average of the whole feature map, global average pooling will be more robust to spatial translations; and (3) because of the huge number of parameters in fully connected layers which usually take over 50% in all the parameters of the whole network, replacing them by global average pooling layers can significantly reduce the size of the model, and this makes global average pooling very useful in model compression. [0200] Global average pooling makes sense, since stronger features in the last layer are expected to have a higher average value. In some implementations, global average pooling can be used as a proxy for the classification score. The feature maps under global average pooling can be interpreted as confidence maps, and force correspondence between the feature maps and the categories. Global average pooling can be particularly effective if the last layer features are at a sufficient abstraction for direct classification; however, global average pooling alone is not enough if multilevel features should be combined into groups like parts models, which is best performed by adding a simple fully connected layer or other classifier after the global average pooling.--, in [0198]-[0200]). Re Claims 5, Maturana as modified by SEDAGHAT, IANCU and ZHANG further disclose wherein each of the first activation layer, the second activation layer and the third activation layer uses a ReLU activation function, and a ratio thereof is set to 0.01 (see Sedaghat: e.g., Fig. 2, and, -- 3 Method The core network architecture is based on VoxNet [21] and is illustrated in Fig. 2. It takes a 3D voxel grid as input and contains two convolutional layers with 3D filters followed by two fully connected layers. Although this choice may not be optimal, we keep it to be able to directly compare our modifications to VoxNet. In addition, we experimented with a slightly deeper network that has four convolutional layers. Point clouds and CAD models are converted to voxel grids (occupancy grids). For the NYUv2 dataset we used the provided tools for the conversion; for the other datasets we implemented our own version. We tried both binary-valued and continuous-valued occupancy grids. In the end, the difference in the results was negligible and thus we only report the results of the former one. Multi-task learning We modify the baseline architecture by adding orientation estimation as an auxiliary parallel task. We call the resulting architecture the ORIentation-boosted vOxel Net – ORION. Without loss of generality, we only consider rotation around the z-axis (azimuth) as the most varying component of orientation in practical applications. Throughout this paper we use the term ’orientation’ to refer to this component.--, in page 4; see Zhang: e.g., -- [0123] A convolutional neural network is a type of neural network. The fundamental difference between a densely connected layer and a convolution layer is this: Dense layers learn global patterns in their input feature space, whereas convolution layers learn local patters: in the case of images, patterns found in small 2D windows of the inputs. This key characteristic gives convolutional neural networks two interesting properties: (1) the patterns they learn are translation invariant and (2) they can learn spatial hierarchies of patterns. [0124] Regarding the first, after learning a certain pattern in the lower-right corner of a picture, a convolution layer can recognize it anywhere: for example, in the upper-left corner. A densely connected network would have to learn the pattern anew if it appeared at a new location. This makes convolutional neural networks data efficient because they need fewer training samples to learn representations, and they have generalization power. [0125] Regarding the second, a first convolution layer can learn small local patterns such as edges, a second convolution layer will learn larger patterns made of the features of the first layers, and so on. This allows convolutional neural networks to efficiently learn increasingly complex and abstract visual concepts. [0126] A convolutional neural network learns highly non-linear mappings by interconnecting layers of artificial neurons arranged in many different layers with activation functions that make the layers dependent. It includes one or more convolutional layers, interspersed with one or more sub-sampling layers and non-linear layers, which are typically followed by one or more fully connected layers. Each element of the convolutional neural network receives inputs from a set of features in the previous layer. The convolutional neural network learns concurrently because the neurons in the same feature map have identical weights. These local shared weights reduce the complexity of the network such that when multi-dimensional input data enters the network, the convolutional neural network avoids the complexity of data reconstruction in feature extraction and regression or classification process.--, in [0123]-[0126]; and, -- [0139] The algorithm includes computing the activation of all neurons in the network, yielding an output for the forward pass. The activation of neuron m in the hidden layers …[0140] This is done for all the hidden layers to get the activation--, in [0139]-[0140], and, -- Batch Normalization [0187] Batch normalization is a method for accelerating deep network training by making data standardization an integral part of the network architecture. Batch normalization can adaptively normalize data even as the mean and variance change over time during training. It works by internally maintaining an exponential moving average of the batch-wise mean and variance of the data seen during training. The main effect of batch normalization is that it helps with gradient propagation—much like residual connections—and thus allows for deep networks. Some very deep networks can only be trained if they include multiple Batch Normalization layers. [0188] Batch normalization can be seen as yet another layer that can be inserted into the model architecture, just like the fully connected or convolutional layer. The BatchNormalization layer is typically used after a convolutional or densely connected layer. It can also be used before a convolutional or densely connected layer. Both implementations can be used by the technology disclosed and are shown in FIG. 17. The BatchNormalization layer takes an axis argument, which specifies the feature axis that should be normalized. This argument defaults to −1, the last axis in the input tensor. This is the correct value when using Dense layers, Conv1D layers, RNN layers, and Conv2D layers with data_format set to “channels_last”. But in the niche use case of Conv2D layers with data_format set to “channels_first”, the features axis is axis 1; the axis argument in BatchNormalization can be set to 1. [0189] Batch normalization provides a definition for feed-forwarding the input and computing the gradients with respect to the parameters and its own input via a backward pass. In practice, batch normalization layers are inserted after a convolutional or fully connected layer, but before the outputs are fed into an activation function. For convolutional layers, the different elements of the same feature map—i.e., the activations—at different locations are normalized in the same way in order to obey the convolutional property. Thus, all activations in a mini-batch are normalized over all locations, rather than per activation.--, in [0187]-[0189]). Re Claims 6, Maturana as modified by SEDAGHAT, IANCU and ZHANG further disclose wherein a number of convolution kernels of the first convolution layer, a number of convolution kernels of the second convolution layer of the 1st second hidden unit, and a number of convolution kernels of the third convolution layer of the 1st second hidden unit are all 32, wherein a number of convolution kernels of the second convolution layer of the 2nd second hidden unit and a number of convolution kernels of the third convolution layer of the 2nd second hidden unit are all 64, wherein a number of convolution kernels of the second convolution layer of the 3rd second hidden unit and a number of convolution kernels of the third convolution layer of the 3rd second hidden unit are all 128 (see Sedaghat: e.g., Fig. 2, and, -- 3 Method The core network architecture is based on VoxNet [21] and is illustrated in Fig. 2. It takes a 3D voxel grid as input and contains two convolutional layers with 3D filters followed by two fully connected layers. Although this choice may not be optimal, we keep it to be able to directly compare our modifications to VoxNet. In addition, we experimented with a slightly deeper network that has four convolutional layers. Point clouds and CAD models are converted to voxel grids (occupancy grids). For the NYUv2 dataset we used the provided tools for the conversion; for the other datasets we implemented our own version. We tried both binary-valued and continuous-valued occupancy grids. In the end, the difference in the results was negligible and thus we only report the results of the former one. Multi-task learning We modify the baseline architecture by adding orientation estimation as an auxiliary parallel task. We call the resulting architecture the ORIentation-boosted vOxel Net – ORION. Without loss of generality, we only consider rotation around the z-axis (azimuth) as the most varying component of orientation in practical applications. Throughout this paper we use the term ’orientation’ to refer to this component.--, in page 4; see Zhang: e.g., -- [0123] A convolutional neural network is a type of neural network. The fundamental difference between a densely connected layer and a convolution layer is this: Dense layers learn global patterns in their input feature space, whereas convolution layers learn local patters: in the case of images, patterns found in small 2D windows of the inputs. This key characteristic gives convolutional neural networks two interesting properties: (1) the patterns they learn are translation invariant and (2) they can learn spatial hierarchies of patterns. [0124] Regarding the first, after learning a certain pattern in the lower-right corner of a picture, a convolution layer can recognize it anywhere: for example, in the upper-left corner. A densely connected network would have to learn the pattern anew if it appeared at a new location. This makes convolutional neural networks data efficient because they need fewer training samples to learn representations, and they have generalization power. [0125] Regarding the second, a first convolution layer can learn small local patterns such as edges, a second convolution layer will learn larger patterns made of the features of the first layers, and so on. This allows convolutional neural networks to efficiently learn increasingly complex and abstract visual concepts. [0126] A convolutional neural network learns highly non-linear mappings by interconnecting layers of artificial neurons arranged in many different layers with activation functions that make the layers dependent. It includes one or more convolutional layers, interspersed with one or more sub-sampling layers and non-linear layers, which are typically followed by one or more fully connected layers. Each element of the convolutional neural network receives inputs from a set of features in the previous layer. The convolutional neural network learns concurrently because the neurons in the same feature map have identical weights. These local shared weights reduce the complexity of the network such that when multi-dimensional input data enters the network, the convolutional neural network avoids the complexity of data reconstruction in feature extraction and regression or classification process.--, in [0123]-[0126]; and, -- [0139] The algorithm includes computing the activation of all neurons in the network, yielding an output for the forward pass. The activation of neuron m in the hidden layers …[0140] This is done for all the hidden layers to get the activation--, in [0139]-[0140], and, -- Batch Normalization [0187] Batch normalization is a method for accelerating deep network training by making data standardization an integral part of the network architecture. Batch normalization can adaptively normalize data even as the mean and variance change over time during training. It works by internally maintaining an exponential moving average of the batch-wise mean and variance of the data seen during training. The main effect of batch normalization is that it helps with gradient propagation—much like residual connections—and thus allows for deep networks. Some very deep networks can only be trained if they include multiple Batch Normalization layers. [0188] Batch normalization can be seen as yet another layer that can be inserted into the model architecture, just like the fully connected or convolutional layer. The BatchNormalization layer is typically used after a convolutional or densely connected layer. It can also be used before a convolutional or densely connected layer. Both implementations can be used by the technology disclosed and are shown in FIG. 17. The BatchNormalization layer takes an axis argument, which specifies the feature axis that should be normalized. This argument defaults to −1, the last axis in the input tensor. This is the correct value when using Dense layers, Conv1D layers, RNN layers, and Conv2D layers with data_format set to “channels_last”. But in the niche use case of Conv2D layers with data_format set to “channels_first”, the features axis is axis 1; the axis argument in BatchNormalization can be set to 1. [0189] Batch normalization provides a definition for feed-forwarding the input and computing the gradients with respect to the parameters and its own input via a backward pass. In practice, batch normalization layers are inserted after a convolutional or fully connected layer, but before the outputs are fed into an activation function. For convolutional layers, the different elements of the same feature map—i.e., the activations—at different locations are normalized in the same way in order to obey the convolutional property. Thus, all activations in a mini-batch are normalized over all locations, rather than per activation.--, in [0187]-[0189]). Re Claims 7-8, and 10, claims 7-8, and 10 are corresponding computer program product claim to claims 1-3, and 5 respectively. Claims 7-8, and 10 thus are rejected for the similar reasons for claims 1-3, and 5. See above discussions with regard to claims 1-3, and 5 respectively. Furthermore, Maturana as modified by SEDAGHAT, IANCU and ZHANG further disclose a computer program product, being applied for a processor to conduct (see ZHANG: e.g., --[0561] Storage subsystem 2910 stores programming and data constructs that provide the functionality of some or all of the modules and methods described herein. These software modules are generally executed by deep learning processors 2978. [0562] Deep learning processors 2978 can be graphics processing units (GPUs) or field-programmable gate arrays (FPGAs). Deep learning processors 2978 can be hosted by a deep learning cloud platform--, in [0561]-[0562]). Re Claims 9, Maturana as modified by SEDAGHAT, IANCU and ZHANG further disclose wherein each of the first pooling layer, the second pooling layer, the third pooling layer and the fourth pooling layer is a maximum pooling layer (see Maturana: e.g., --Pooling Layers (Pm). These layers downsample the input volume by a factor of by m along the spatial dimensions by replacing each mxmxm non-overlapping block of voxels with their maximum.--, in right col., page 924,. also see Zhang: e.g., -- [0158] FIG. 9 is one implementation of sub-sampling layers in accordance with one implementation of the technology disclosed. Sub-sampling layers reduce the resolution of the features extracted by the convolution layers to make the extracted features or feature maps—robust against noise and distortion. In one implementation, sub-sampling layers employ two types of pooling operations, average pooling and max pooling. The pooling operations divide the input into non-overlapping two-dimensional spaces. For average pooling, the average of the four values in the region is calculated. For max pooling, the maximum value of the four values is selected. [0159] In one implementation, the sub-sampling layers include pooling operations on a set of neurons in the previous layer by mapping its output to only one of the inputs in max pooling and by mapping its output to the average of the input in average pooling. In max pooling, the output of the pooling neuron is the maximum value that resides within the input… [0163] In FIG. 9, the input is of size 4×4. For 2×2 sub-sampling, a 4×4 image is divided into four non-overlapping matrices of size 2×2. For average pooling, the average of the four values is the whole-integer output. For max pooling, the maximum value of the four values in the 2×2 matrix is the whole-integer output.--, in [0158]-[0163]). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to WEIWEN YANG whose telephone number is (571)270-5670. The examiner can normally be reached on Monday-Friday 8:30am-4:30pm east. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Amandeep Saini can be reached on 571-272-3382. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /WEI WEN YANG/ Primary Examiner, Art Unit 2662
Read full office action

Prosecution Timeline

Oct 30, 2024
Application Filed
Aug 11, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737879
Machine Learning for Detection of Diseases from External Anterior Eye Images
3y 9m to grant Granted Sep 15, 2026
Patent 12737880
SYSTEMS AND METHODS OF ANALYZING MICROBIOMES USING ARTIFICIAL INTELLIGENCE
3y 4m to grant Granted Sep 15, 2026
Patent 12738057
CUT-PASTE TRAINING AUGMENTATION FOR MACHINE LEARNING MODELS
2y 7m to grant Granted Sep 15, 2026
Patent 12737851
ENHANCED QUALITY BOREHOLE IMAGE GENERATION AND METHOD
2y 3m to grant Granted Sep 15, 2026
Patent 12729363
SYSTEM AND METHOD FOR SELECTING COLONIES
4y 1m to grant Granted Sep 08, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
82%
Grant Probability
93%
With Interview (+11.3%)
2y 5m (~7m remaining)
Median Time to Grant
Low
PTA Risk
Based on 682 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month