Prosecution Insights
Last updated: August 17, 2026
Application No. 18/176,216

DEEP NEURAL NETWORK MODEL COMPRESSION

Non-Final OA §101§102§103§112
Filed
Feb 28, 2023
Examiner
DETERDING, GWYNEVERE AMELIA
Art Unit
2125
Tech Center
2100 — Computer Architecture & Software
Assignee
NXP Semiconductors N.V.
OA Round
3 (Non-Final)
83%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
83%
With Interview

Examiner Intelligence

Grants 83% — above average
83%
Career Allowance Rate
5 granted / 6 resolved
+28.3% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
3y 4m
Avg Prosecution
20 currently pending
Career history
26
Total Applications
across all art units

Statute-Specific Performance

§101
31.1%
-8.9% vs TC avg
§103
34.1%
-5.9% vs TC avg
§102
14.4%
-25.6% vs TC avg
§112
15.2%
-24.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 6 resolved cases

Office Action

§101 §102 §103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claims 1-24 are presented for examination. Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 6/11/2026 has been entered. Response to Amendment The objection to the claims, the objection to the specification, the 112(b) rejection, and the “software per se” 101 rejection set forth in the previous Office Action have been obviated by the amendments. Thus, these objections/rejections are withdrawn. Claim Rejections - 35 USC § 112 Claims 1-24 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. The term “relatively low” in claims 1 and 13 is a relative term which renders the claim indefinite. The term “relatively low” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. The limitation of “relatively low accumulated alpha values” has been rendered indefinite by the use of the term “relatively low.” For examination purposes, Examiner will interpret “relatively low accumulated alpha values among accumulated alpha values” to mean a lowest predetermined number of the accumulated alpha values. Claims 2-12 and 14-24 are rejected due to dependency on a rejected base claim. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-24 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The analysis of the claims will follow the 2019 Revised Patent Subject Matter Eligibility Guidance (“2019 PEG”). Claim 1 Step 1: The claim is directed to a non-transitory computer readable medium, and is therefore directed to the statutory category of articles of manufacture. Step 2A Prong 1: The claim recites: -calculating, during training of the machine learning model, alpha values for different parts of the machine learning model based on gradients used in training the machine learning model, wherein the alpha values are an importance metric; this limitation recites the mathematical concept of calculating alpha values based on gradients used in training a machine learning model -accumulating the calculated alpha values across multiple training iterations during a single training process; this limitation encompasses mentally accumulating the calculated alpha values across multiple training iterations during a single training process Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites “training the machine learning model using training input data,” however this limitation amounts to generally linking the use of a judicial exception to the technological environment of model training (MPEP 2106.05(h)). The claim additionally recites “pruning the machine learning model based upon the accumulated alpha values by removing one or more channels or filters from one or more convolutional layers associated with relatively low accumulated alpha values among accumulated alpha values calculated for channels or filters of the one or more convolutional layers, to generate a compressed machine learning model having, relative to the machine learning model before pruning, at least one of fewer parameters, fewer activations, or fewer floating-point operations,” however this limitation amounts to generally linking the use of a judicial exception to the field of use of model pruning (MPEP 2106.05(h)), is it merely recites to prune the structures that were determined by the judicial exception. The claim also recites “A data processing system comprising instructions embodied in a non-transitory computer readable medium, the instructions for pruning a machine learning model in a processor, the instructions, comprising… [the method],” however this limitation amounts to mere instructions to apply the judicial exception using a generic computer programmed with a generic class of computer algorithms (MPEP 2106.05(f)). Step 2B: The claim does not contain significantly more than the judicial exception. The analysis at this step mirrors that of Step 2A Prong 2 above. As an ordered whole, the claim is directed to the abstract idea of calculating alpha values and accumulating the calculated alpha values across training iterations to determine which parts of a machine learning model to prune. Nothing in the claim provides significantly more than this. As such, the claim is not patent eligible. Claim 2 Step 1: An article of manufacture, as above. Step 2A Prong 1: The claim recites: -sorting the accumulated calculated alpha values; this limitation encompasses mentally sorting the accumulated calculated alpha values -selecting a lowest predetermined number of the sorted values; this limitation encompasses mentally selecting a lowest predetermined number of sorted values Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites “pruning the machine learning model based upon the selected values,” however this limitation amounts to generally linking the use of a judicial exception to the field of use of model pruning (MPEP 2106.05(h)). Step 2B: The claim does not contain significantly more than the judicial exception. The analysis at this step mirrors that of Step 2A Prong 2 above. Claim 3 Step 1: An article of manufacture, as above. Step 2A Prong 1: The claim recites: -summing the gradients for the different parts of the machine learning model over the different parts of the machine learning model; this limitation recites the mathematical concept of summing gradients for different parts of a machine learning model over the different parts of the model Step 2A Prong 2: This judicial exception is not integrated into a practical application, see analysis of claim 1. Step 2B: The claim does not contain significantly more than the judicial exception, see analysis of claim 1. Claim 4 Step 1: An article of manufacture, as above. Step 2A Prong 1: The claim recites: -assigning an importance score to filters in the machine learning model at a class level; this limitation encompasses mentally determining an importance score to assign to filters in the model at a class level Step 2A Prong 2: This judicial exception is not integrated into a practical application, see analysis of claim 1. Step 2B: The claim does not contain significantly more than the judicial exception, see analysis of claim 1. Claim 5 Step 1: An article of manufacture, as above. Step 2A Prong 1: The claim recites: -weighing the gradients before summing; this limitation encompasses mentally determining a weight for each of the gradients before summing Step 2A Prong 2: This judicial exception is not integrated into a practical application, see analysis of claim 1. Step 2B: The claim does not contain significantly more than the judicial exception, see analysis of claim 1. Claim 6 Step 1: An article of manufacture, as above. Step 2A Prong 1: The claim recites a mathematical calculation (not repeated here for formatting purposes). Step 2A Prong 2: This judicial exception is not integrated into a practical application, see analysis of claim 1. Step 2B: The claim does not contain significantly more than the judicial exception, see analysis of claim 1. Claim 7 Step 1: An article of manufacture, as above. Step 2A Prong 1: The claim recites: -summing a i k c over a training set; this limitation recites the mathematical concept of summing the calculated alpha values over a training set Step 2A Prong 2: This judicial exception is not integrated into a practical application, see analysis of claim 1. Step 2B: The claim does not contain significantly more than the judicial exception, see analysis of claim 1. Claim 8 Step 1: An article of manufacture, as above. Step 2A Prong 1: The claim recites the mathematical calculation of a summation (not repeated here for formatting purposes). Step 2A Prong 2: This judicial exception is not integrated into a practical application, see analysis of claim 1. Step 2B: The claim does not contain significantly more than the judicial exception, see analysis of claim 1. Claim 9 Step 1: An article of manufacture, as above. Step 2A Prong 1: The claim recites: -sorting the a i k values; this limitation encompasses mentally sorting the values -selecting a lowest predetermined number of the sorted values; this limitation encompasses mentally selecting a lowest predetermined number of the sorted values Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites “pruning the machine learning model based upon the selected values,” however this limitation amounts to merely generally linking the judicial exception to the field of use of model pruning (MPEP 2106.05(h)). Step 2B: The claim does not contain significantly more than the judicial exception. The analysis at this step mirrors that of Step 2A Prong 2 above. Claim 10 Step 1: An article of manufacture, as above. Step 2A Prong 1: The claim recites the same judicial exception as claim 9. Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites “pruning one or more of weights, kernels, features, layers, units, or neurons associated with the removed one or more channels or filters.” However, this limitation amounts to merely generically linking the judicial exception to the field of use of model pruning (MPEP 2106.05(h)). Step 2B: The claim does not contain significantly more than the judicial exception. The analysis at this step mirrors that of Step 2A Prong 2 above. Claim 11 Step 1: An article of manufacture, as above. Step 2A Prong 1: The claim recites the same judicial exception as claim 1. Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites “wherein the machine learning model is one of a deep-learning neural network and a convolutional neural network.” However, this limitation amounts to mere instructions to apply the judicial exception on a generic computer programmed with a generic class of computer algorithms (MPEP 2106.05(f)), because it is merely further limiting the generic machine learning model that is being pruned. Step 2B: The claim does not contain significantly more than the judicial exception. The analysis at this step mirrors that of Step 2A Prong 2 above. Claim 12 Step 1: An article of manufacture, as above. Step 2A Prong 1: The claim recites: -estimating a gradient update based upon an output of the machine learning model; this limitation encompasses mentally estimating a gradient update based upon an output of the machine learning model -updating weights of the machine learning model based upon the gradient update using backpropagation; this limitation recites the mathematical concept of updating weights using backpropagation Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites “initializing the machine learning model” and “inputting a plurality of training input data tensors into the machine learning model in a plurality of iterations,” however these limitations amount to insignificant extra-solution activity because they are merely necessary pre-solution steps for training a machine learning model in order to perform model pruning and do not meaningfully limit the claim. Step 2B: The claim does not contain significantly more than the judicial exception. The initializing the model limitation, in addition to being insignificant extra-solution activity, is also well-understood, routine, and conventional (Instant Application Specification, [0065]: Any known method of initializing the specific type of model being trained and pruned may be used). The inputting data into the model limitation, in addition to being insignificant extra-solution activity, is also well-understood, routine, and conventional (US20230306257, Sun et al., [0003]: Referring now to FIG. 1A, a method of training a neural network according to the conventional art is shown. The method can include inputting a training data set to a neural network model to generate an output in a forward pass). Claims 13-22, 24 Step 1: The claims recite a method and are therefore directed to the statutory category of processes. Step 2A Prong 1: Claims 13-22 and 24 recite the same judicial exceptions as claims 1-10 and 12 respectively. Step 2A Prong 2: The claims do not integrate the judicial exception into a practical application. The analysis at this step mirrors that of claims 1-10 and 12 respectively, except insofar as claims 13-22 and 24 are method claims and do not recite the “data processing system” limitation. Step 2B: The claims do not contain significantly more than the judicial exception. The analysis at this step mirrors that of claims 1-10 and 12 respectively, except insofar as claims 13-22 and 24 are method claims and do not recite the “data processing system” limitation. Claim 23 Step 1: A method, as above. Step 2A Prong 1: The claim recites the same judicial exceptions as claim 10 above. Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites “wherein the machine learning model is one of a deep-learning neural network and a convolutional neural network.” However, this limitation is mere instructions to apply the judicial exception on a generic computer programmed with a generic class of computer algorithms (MPEP 2106.05(f)), because it is merely further limiting the generic machine learning model that is being pruned. Step 2B: The claim does not contain significantly more than the judicial exception. The further limitation of the claim amounts to mere instructions to apply the judicial exception for the same reasons given above. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claims 1, 11, and 13 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Molchanov et al. (NPL: “Importance Estimation for Neural Network Pruning”) (“Molchanov”). Regarding claim 1, Molchanov discloses “A data processing system comprising instructions embodied in a non-transitory computer readable medium having stored thereon instructions for pruning a machine learning model in a processor (Molchanov, 4.2.2 Pruning and fine-tuning: “We use the following settings: 4 GPUs”), the instructions, comprising: training the machine learning model using training input data (Molchanov, 3.1 Pruning algorithm: “Our pruning method takes a trained network as input and prunes it during an iterative fine-tuning process with a small learning rate” and 4.2.2 Pruning and fine-tuning: “We use the following settings: 4 GPUs and a batch size of 256 examples”; Examiner notes that “iterative fine-tuning process with a small learning rate” corresponds to “training”); calculating, during training of the machine learning model, alpha values for different parts of the machine learning model based on gradients used in training the machine learning model, wherein the alpha values are an importance metric (Molchanov, 3.1 Pruning algorithm: “During each epoch, the following steps are repeated: 1. For each minibatch, we compute parameter gradients and update network weights by gradient descent. We also compute the importance of each neuron (or filter) using the gradient averaged over the minibatch, as described in (7) or (8)”; Examiner notes that “importance of each neuron using the gradient averaged over the minibatch” corresponds to “alpha values for different parts of the machine learning model based on gradients used in training the machine learning model”); accumulating the calculated alpha values across multiple training iterations during a single training process (Molchanov, 3.1. Pruning algorithm: “During each epoch, the following steps are repeated… 2. After a predefined number of minibatches, we average the importance score of each neuron (or filter) over the of minibatches” and Molchanov, 3.2. Implementation details: “Importance score accumulation. During training or fine-tuning with minibatches, observed gradients are combined to compute a single importance score Î = 𝔼 I " ; Examiner notes that each minibatch being processed corresponds to a “training iteration” since the network weights are updated for each minibatch, and each epoch corresponds to a single training process since an epoch is one complete pass through the entire training dataset, therefore averaging the importance score of each neuron over the minibatches during an epoch corresponds to “accumulating the calculated alpha values across multiple training iterations during a single training process”); and pruning the machine learning model based upon the accumulated alpha values by removing one or more channels or filters from one or more convolutional layers associated with relatively low accumulated alpha values among accumulated alpha values calculated for channels or filters of the one or more convolutional layers (Molchanov, 3.1. Pruning algorithm: “During each epoch, the following steps are repeated… 2. After a predefined number of minibatches, we average the importance score of each neuron (or filter) over the of minibatches, and remove the N neurons with the smallest importance scores” and 3.2. Implementation details: “Number of neurons pruned per iteration needs to be chosen based on how correlated the neurons are to each other. We observed that a filter’s contribution changes during pruning and we usually prune around 2% of initial filters per iteration” and 4.1.1 LeNet3, All layers pruning: “We consider both a direct application to convolutional filter weights (“on weight”) and the use of gates following each convolutional layer (“on gate”).] We treat linear layers as 1 × 1 convolutions. In all cases, pruning removes the entire filter and its corresponding bias”), to generate a compressed machine learning model having, relative to the machine learning model before pruning, at least one of fewer parameters, fewer activations, or fewer floating-point operations” (Molchanov, page 8, Table 3: see ResNet-101: No pruning: 7.80 GFLOPs 4.47 Params(107) compared to Taylor-FO-BN-75% (Ours): 4.70 GFLOPs 3.12 Params(107)). Regarding claim 11, the rejection of claim 1 is incorporated. Molchanov further discloses “wherein the machine learning model is one of a deep-learning neural network and a convolutional neural network” (Molchanov, 4.1.1 LeNet3: “We start with a simple network, LeNet3, trained on the CIFAR-10 dataset to achieve 73% test accuracy. The architecture of LeNet consists of 2 convolutional and 3 linear layers”). Regarding claim 13, Molchanov discloses “A method of pruning a machine learning model, comprising: training the machine learning model using training input data (Molchanov, 3.1 Pruning algorithm: “Our pruning method takes a trained network as input and prunes it during an iterative fine-tuning process with a small learning rate” and 4.2.2 Pruning and fine-tuning: “We use the following settings: 4 GPUs and a batch size of 256 examples”; Examiner notes that “iterative fine-tuning process with a small learning rate” corresponds to “training”); calculating, during training of the machine learning model, alpha values for different parts of the machine learning model based on gradients used in training the machine learning model, wherein the alpha values are an importance metric (Molchanov, 3.1 Pruning algorithm: “During each epoch, the following steps are repeated: 1. For each minibatch, we compute parameter gradients and update network weights by gradient descent. We also compute the importance of each neuron (or filter) using the gradient averaged over the minibatch, as described in (7) or (8)”; Examiner notes that “importance of each neuron using the gradient averaged over the minibatch” corresponds to “alpha values for different parts of the machine learning model based on gradients used in training the machine learning model”); accumulating the calculated alpha values across multiple training iterations during a single training process (Molchanov, 3.1. Pruning algorithm: “During each epoch, the following steps are repeated… 2. After a predefined number of minibatches, we average the importance score of each neuron (or filter) over the of minibatches” and Molchanov, 3.2. Implementation details: “Importance score accumulation. During training or fine-tuning with minibatches, observed gradients are combined to compute a single importance score Î = 𝔼 I " ; Examiner notes that each minibatch being processed corresponds to a “training iteration” since the network weights are updated for each minibatch, and each epoch corresponds to a single training process since an epoch is one complete pass through the entire training dataset, therefore averaging the importance score of each neuron over the minibatches during an epoch corresponds to “accumulating the calculated alpha values across multiple training iterations during a single training process”); and pruning the machine learning model based upon the accumulated alpha values by removing one or more channels or filters from one or more convolutional layers associated with relatively low accumulated alpha values among accumulated alpha values calculated for channels or filters of the one or more convolutional layers (Molchanov, 3.1. Pruning algorithm: “During each epoch, the following steps are repeated… 2. After a predefined number of minibatches, we average the importance score of each neuron (or filter) over the of minibatches, and remove the N neurons with the smallest importance scores” and 3.2. Implementation details: “Number of neurons pruned per iteration needs to be chosen based on how correlated the neurons are to each other. We observed that a filter’s contribution changes during pruning and we usually prune around 2% of initial filters per iteration” and 4.1.1 LeNet3, All layers pruning: “We consider both a direct application to convolutional filter weights (“on weight”) and the use of gates following each convolutional layer (“on gate”).] We treat linear layers as 1 × 1 convolutions. In all cases, pruning removes the entire filter and its corresponding bias”), to generate a compressed machine learning model having, relative to the machine learning model before pruning, at least one of fewer parameters, fewer activations, or fewer floating-point operations” (Molchanov, page 8, Table 3: see ResNet-101: No pruning: 7.80 GFLOPs 4.47 Params(107) compared to Taylor-FO-BN-75% (Ours): 4.70 GFLOPs 3.12 Params(107)). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 2-4 and 14-16 are rejected under 35 U.S.C. 103 as being unpatentable over Molchanov in view of Ni et al. (NPL: “Interpretable Analysis and Pruning of Modulation Recognition Network Based on Deep Learning”) (“Ni”). Regarding claim 2, the rejection of claim 1 is incorporated. Molchanov does not appear to explicitly disclose the further limitations of the claim. However, Ni discloses, “sorting the accumulated calculated alpha values (Ni, 3.4: “All the data in the training set are input into the network in turn, and the gradient values of each convolution filter are accumulated to obtain their overall contribution to the task. Figure 11 shows the ranking results of the contributions of the VGG16 obtained by using the method this paper proposed”; the examiner notes that ranking the contribution values corresponds to sorting the accumulated calculated alpha values); selecting a lowest predetermined number of the sorted values (Ni, 3.4: The 13 convolution layers of the VGG16 are pruned one by one according to the order of the contribution values of the convolution filters. Figure 12 shows the accuracy of VGG16 after pruning each convolution layer without retraining, which shows that the deletion of unimportant convolution filters in some convolution layers has little effect on the performance of the model. Figure 13 shows the accuracy of VGG16 with retraining after pruning each convolution layer. For VGG16, the accuracy on the test set increases from 0.951 to 0.954 after removing 70% of the convolution filters, and the accuracy decreases within 1% after removing 80% of the convolution filters; the examiner notes that the percentage of filters to be removed (e.g. 70%) corresponds to a predetermined number); and pruning the machine learning model based upon the selected values” (Ni, 3.4: The 13 convolution layers of the VGG16 are pruned one by one according to the order of the contribution values of the convolution filters). Ni and the instant application both relate to pruning neural networks and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified Molchanov such that pruning the machine learning model includes “sorting the accumulated calculated alpha values; selecting a lowest predetermined number of the sorted values; and pruning the machine learning model based upon the selected values,” as disclosed by Ni, and one would have been motivated to do so, as doing so would reduce the model size, memory, and computing time of the model (see Ni, 3.4). Claim 14 is a method claim corresponding to system claim 2 and is rejected for the same reasons as claim 2 above. Regarding claim 3, the rejection of claim 1 is incorporated. Molchanov does not appear to explicitly disclose the further limitations of the claim. However, Ni discloses “wherein calculating alpha values for different parts of the machine learning model based on gradients… includes summing the gradients for the different parts of the machine learning model over the different parts of the machine learning model” (Ni, 2.5: “the size of the feature map corresponding to each convolution filter is u × v… Define the activation value of the kth feature map as Ak and the output score of category C as yc (before softmax). Then the weight value a k c of the kth feature map for the category C can be obtained by formula (4): PNG media_image1.png 54 372 media_image1.png Greyscale where A ij k represents the value of the pixel point with the coordinate of (i, j) in the kth feature map”; the examiner notes that the filters correspond to “different parts of the machine learning model,” the partial derivative represents gradients for the kth feature map (corresponding to a filter) of the machine learning model which are being summed together, and dividing the sum by the size of the filter (u x v) reads on summing the gradients for the filter over the filter). Ni and the instant application both relate to pruning neural networks and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified Molchanov such that calculating alpha values for different parts of the machine learning model based on gradients used in training the machine learning model includes “summing the gradients for the different parts of the machine learning model over the different parts of the machine learning model” as disclosed by Ni, and one would have been motivated to do so, as doing so would allow for identifying convolution filters that do not play an active role in the task of identifying a particular class, thus improving the interpretability of pruning decisions and allowing for compression of the model without influencing the accuracy of the network (see Ni, 2.6 Model Pruning Based on Grad-CAM, and Introduction, paragraphs 7-8). Claim 15 is a method claim corresponding to system claim 3 and is rejected for the same reasons as claim 3 above. Regarding claim 4, the rejection of claim 3 is incorporated. Molchanov as modified by Ni further discloses, “wherein summing the gradients for the different parts of the machine learning model over the different parts includes assigning an importance score to filters in the machine learning model at a class level” (Ni, 3.4: a k c also represents the contribution value of each convolution filter to the recognition and classification; the examiner notes that a k c corresponds to the contribution of the kth feature map (corresponding to a filter) towards category C (Ni, 2.5: the weight value a k c of the kth feature map for the category C can be obtained by formula (4)) and therefore corresponds to an importance score of filters at a class level). Ni and the instant application both relate to pruning neural networks and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified Molchanov to include “wherein summing the gradients for the different parts of the machine learning model over the different parts includes assigning an importance score to filters in the machine learning model at a class level” as disclosed by Ni, and one would have been motivated to do so, as doing so would allow for identifying convolution filters that do not play an active role in the task of identifying a particular class, thus improving the interpretability of pruning decisions and allowing for compression of the model without influencing the accuracy of the network (see Ni, 2.6 Model Pruning Based on Grad-CAM, and Introduction, paragraphs 7-8). Claim 16 is a method claim corresponding to system claim 4 and is rejected for the same reasons as claim 4 above. Claims 5 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Molchanov in view of Ni and Chattopadhyay et al. (NPL: Grad-CAM++: Improved Visual Explanations for Deep Convolutional Networks) (“Chattopadhyay”). Regarding claim 5, the rejection of claim 3 is incorporated. Neither Molchanov nor Ni appears to explicitly disclose the further limitations of the claim. However, Chattopadhyay discloses “wherein summing the gradients for the different parts of the machine learning model… includes weighing the gradients before summing” (Chattopadhyay, 3.1: This problem can be fixed by taking a weighted average of the pixel-wise gradients. In particular, we reformulate Eqn 3 by explicitly coding the structure of the weights w k c as: PNG media_image2.png 46 308 media_image2.png Greyscale where relu is the Rectified Linear Unit activation function. Here the a i j k c ’s are weighting co-efficients for the pixel-wise gradients for class c and convolutional feature map Ak). Chattopadhyay and the instant application both relate to neural networks and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the step of summing the gradients for the different parts of the machine learning model over the different parts disclosed by the combination of Ni and Molchanov to include “weighing the gradients before summing” as disclosed by Chattopadhyay, and one would have been motivated to do so, as doing so would provide a measure of importance of each pixel in a feature map towards the overall decision of the CNN, so that the alpha values are accurate even when there are multiple occurrences of the same object in an image (see Chattopadhyay, Introduction, paragraph 4). Claim 17 is a method claim corresponding to system claim 5 and is rejected for the same reasons as claim 5 above. Claims 12 and 24 are rejected under 35 U.S.C. 103 as being unpatentable over Molchanov in view of Fairhart (US20200327410). Regarding claim 12, the rejection of claim 1 is incorporated. Molchanov further discloses “wherein training the machine learning model using training input data includes initializing the machine learning model” (Molchanov, 3.1 Pruning algorithm: “Our pruning method takes a trained network as input and prunes it during an iterative fine-tuning process with a small learning rate”; Examiner notes that the pre-training of the model before the fine-tuning process corresponds to “initializing the machine learning model”), but does not appear to explicitly disclose the further limitations of the claim. However, Fairhart discloses, “wherein training the machine learning model using training input data includes… inputting a plurality of training input data tensors into the machine learning model in a plurality of iterations (Fairhart, [0052]: In training mode 802, the invention depicted in this embodiment is provided a series of sets of input data tensors 806, each input data tensor 806 comprising multiplexed visual data 808 and distance data 810 representations of an image which may be of the target individual or some other object; the examiner notes that “sets of input data tensors 806” correspond to “a plurality of training input data tensors” and “a series of…” corresponds to “a plurality of iterations”); estimating a gradient update based upon an output of the machine learning model (Fairhart, [0045]: A loss function is applied to compute a loss value based upon the difference between the value of the target and the value of the output; the examiner notes that “loss value” corresponds to “gradient update” because it is used during the gradient descent process to update the model (Fairhart, [0047]: Advantageously, a form of gradient descent is applied to the layers of the neural network in an effort to optimize performance by reducing the loss value)); and updating weights of the machine learning model based upon the gradient update using backpropagation” (Fairhart, [0046-0047]: After applying a loss function to derive the loss value, the neural network performs a back propagation. In back propagation, levels are traversed in reverse order beginning at the highest level. A combination of program and data structures are employed to determine which weights on each level contribute most to the loss value. This is equivalent to computing the loss gradient dL/dW over the current weighting values W on each layer... Application of gradient descent operations results in values for a parameter update that changes weighting in a direction opposite to that of the loss gradient, as illustrated symbolically in FIG. 7; the examiner notes that “a parameter update” corresponds to “updating machine learning model weights” and the parameter update is “based upon the gradient update” and done “using backpropagation” because during backpropagation, the loss value (corresponding to “gradient update”) is used to compute the loss gradient, which is then used to update the weights of the model in a direction opposite to the loss gradient). Fairhart and the instant application both relate to neural networks and are analogous. It would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the training step disclosed by Molchanov to include “inputting a plurality of training input data tensors into the machine learning model in a plurality of iterations; estimating a gradient update based upon an output of the machine learning model; and updating weights of the machine learning model based upon the gradient update using backpropagation,” as disclosed by Fairhart, and one would have been motivated to do so, as doing so would minimize the error rate of the machine learning model (see Fairhart, [0012]). Claim 24 is a method claim corresponding to system claim 12 and is rejected for the same reasons as claim 12 above. Response to Arguments 35 USC § 101 Applicant's arguments regarding the rejections under 35 U.S.C. 101 filed 6/11/2026 have been fully considered, but they are not persuasive (Remarks, pages 10-13, B. 1-5). Regarding (1), Applicant argues that the claims recite a specific improvement in neural-network compression technology, citing paragraphs [0039] and [0040] of the instant application as evidence. However, these paragraphs refer to model compression and structured pruning as common techniques, thus it is unclear how these paragraphs support an improvement to technology provided by the claimed invention. Applicant further states that the alleged mathematical concepts are “applied in a particular technological process that structurally modifies the neural network and improves its suitability for resource-constrained inference.” Examiner respectfully disagrees that this amounts to a specific improvement in neural-network compression technology, as it merely generally links the mathematical concepts to the field of use of structured pruning (MPEP 2106.05(h)). Regarding (2), Applicant argues that Table 1 of the specification provides concrete evidence of this technical improvement, as the results in the table demonstrate “that the claimed gradient-derived pruning criterion improves the selection of structures to remove rather than merely applying generic pruning.” However, the selection of structures to remove based on the gradient-derived pruning criterion comes from the claimed judicial exception of calculating gradient-based alpha values and accumulating the alpha values across training iterations, thus this argument amounts to an assertion that the improvement is provided by the abstract idea of performing mathematical calculations to determine which structures to remove, rather than any particular technological pruning process. The judicial exception alone cannot provide the improvement (MPEP § 2106.05(a)). Regarding (3), Applicant argues that the claims are analogous to USPTO Example 48 Claim 2. Examiner respectfully disagrees. In Example 48 Claim 2, the claim recites additional elements beyond the judicial exception that directly reflect an improvement over existing speech-separation methods, while the instant application claims merely generally link the judicial exception to the field of use of model pruning. Regarding (4), Applicant’s argument that the training limitation is not insignificant extra-solution activity is moot, as the new ground of rejection characterizes the training limitation as merely generally linking the use of the judicial exception to a technological environment (MPEP 2106.05(h)). The model training merely acts as an environment in which the judicial exception of calculating and accumulating alpha values is performed, and thus does not integrate the judicial exception into a practical application. Regarding (5), Applicant argues that the amended claims remedy Examiner’s concerns that the claim itself does not reflect the asserted improvement. However, the amendment of “removing one or more channels or filters from one or more convolutional layers associated with relatively low accumulated alpha values” amounts to generally linking the judicial exception of calculating and accumulating alpha values to the field of use of model pruning and “to generate a compressed machine learning model” merely amounts to an intended result of the pruning, thus Examiner still submits that the pruning limitation does not reflect an improvement to technology. 35 U.S.C. § 103 Applicant's arguments regarding the rejections under 35 U.S.C. 103 filed 6/11/2026 have been fully considered but they, except insofar as rendered moot by the new grounds of rejection, are not persuasive (Remarks, pages 14-16). Applicant’s argument that Ni does not teach or suggest calculating the claimed alpha values during training of the machine learning model or accumulating those alpha values across multiple training iteration during a single training process is moot, as the new grounds of rejection relies on Molchanov to teach these limitations. Applicant’s argument that Molchanov does not teach calculating the claimed alpha values is unpersuasive. Applicant conflates the claimed alpha values with Ni’s Grad-CAM/class-activation-map contribution values, however the claim only requires that the alpha values be “based on gradients used in training the machine learning model” and “an importance metric,” both of which are disclosed by Molchanov, as Molchanov discloses importance scores for each neuron calculated using the gradient averaged over the minibatch (see Molchanov, 3.1 Pruning algorithm). Thus, Molchanov teaches calculating the claimed alpha values. Applicant’s argument that the combination of Ni and Molchanov lacks a sufficient articulated rationale is moot, as the new grounds of the rejection does not rely on the combination of Ni and Molchanov to teach the contested limitations. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to GWYNEVERE A DETERDING whose telephone number is (571)272-7657. The examiner can normally be reached Mon-Fri. 9am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kamran Afshar can be reached at (571) 272-7796. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /G.A.D./Examiner, Art Unit 2125 /KAMRAN AFSHAR/Supervisory Patent Examiner, Art Unit 2125
Read full office action

Prosecution Timeline

Feb 28, 2023
Application Filed
Jan 06, 2026
Non-Final Rejection mailed — §101, §102, §103
Mar 05, 2026
Response Filed
Apr 20, 2026
Final Rejection mailed — §101, §102, §103
Jun 11, 2026
Response after Non-Final Action
Jun 30, 2026
Request for Continued Examination
Jul 01, 2026
Response after Non-Final Action
Jul 29, 2026
Non-Final Rejection mailed — §101, §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705494
MACHINE LEARNING NETWORKS, ARCHITECTURES AND TECHNIQUES FOR DETERMINING OR PREDICTING DEMAND METRICS IN ONE OR MORE CHANNELS
1y 1m to grant Granted Aug 11, 2026
Patent 12682041
METHOD, DEVICE AND COMPUTER PROGRAM PRODUCT FOR GENERATING NEURAL NETWORK MODEL
3y 5m to grant Granted Jul 14, 2026
Patent 12675736
MACHINE LEARNING METHOD
3y 5m to grant Granted Jul 07, 2026
Patent 12651176
CONTINUOUS MAINTENANCE OF MODEL EXPLAINABILITY
3y 4m to grant Granted Jun 09, 2026
Study what changed to get past this examiner. Based on 4 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
83%
Grant Probability
83%
With Interview (+0.0%)
3y 4m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 6 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month