Prosecution Insights
Last updated: October 02, 2026
Application No. 18/411,542

IMPORTANCE-AWARE MODEL PRUNING AND RE-TRAINING FOR EFFICIENT CONVOLUTIONAL NEURAL NETWORKS

Final Rejection §103
Filed
Jan 12, 2024
Priority
Jun 30, 2016 — nonprovisional of PCTCN2016087859 +1 more
Examiner
VARNDELL, ROSS E
Art Unit
2674
Tech Center
2600 — Communications
Assignee
Intel Corporation
OA Round
2 (Final)
85%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
98%
With Interview

Examiner Intelligence

Grants 85% — above average
85%
Career Allowance Rate
535 granted / 632 resolved
+22.7% vs TC avg
Moderate +13% lift
Without
With
+13.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 3m
Avg Prosecution
37 currently pending
Career history
668
Total Applications
across all art units

Statute-Specific Performance

§101
6.9%
-33.1% vs TC avg
§103
67.0%
+27.0% vs TC avg
§102
6.2%
-33.8% vs TC avg
§112
12.1%
-27.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 632 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This final action is responsive to the amendment and remarks filed 07 July 2026. Claims 26-45 are pending. Claims 26, 29, 30, 33, 36, 37, 40, 43 and 44 are amended. Claims 26, 33 and 40 are independent. Response to Arguments The rejection under 35 U.S.C. § 101 is withdrawn. Applicant's Desjardins argument and specification ¶ 24 provide a sufficient basis for eligibility. Applicant's argument that the cited references are silent regarding the determining step have been considered but are moot in view of new ground(s) of rejection (Hassibi and Stork, “Second Order Derivatives for Network Pruning: Optimal Brain Surgeon,” (hereinafter “OBS”).) because of the amendments. OBS supplies the covariance term: “We shall show that the Hessian can be reduced to the sample covariance matrix associated with certain gradient vectors” (OBS, p. 166). OBS also teaches calculation using “a single sequential pass through the training data” (OBS, p. 167), and this covariance over the training data is the claimed influence of the input data. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 26-31, 33- 38, and 40-45 is/are rejected under 35 U.S.C. 103 as being unpatentable over Le Cun et al., “Optimal Brain Damage” (hereinafter “OBD”) in view of Hassibi and Stork, “Second Order Derivatives for Network Pruning: Optimal Brain Surgeon,” hereinafter “OBS”). Claims 26, 33, and 40. OBD discloses one or more non-transitory computer-readable media storing instructions executable to perform operations (ODB describes a computational algorithm executed on a trained network: “training has converged” and the method proceeds by “computing the second derivates … hkk” via backpropagation, then iterating “to step 2.” The entire ODB procedure (forward pass, Hessian diagonal computation, saliency ranking, and weight deletion) is an algorithmic sequence that can only be performed by a processor operating on weights stored in memory. The paper reports empirical results on a digit recognition system, confirming actual computer implementation. The examiner takes official notice that implementing such an algorithm on a processor and storing it on a non-transitory computer-readable medium was well-known and the only possible means of execution.), the operations comprising: providing input data to a neural network, the neural network comprising one or more layers with weights, the input data processed in the one or more layers (OBD: "The method was validated using our handwritten digit recognition network trained with backpropagation ... The network state is computed using the standard formulae PNG media_image1.png 49 167 media_image1.png Greyscale ... where xi is the state of unit i, ai its total input (weighted sum) ... and Wij is the connection going from unit j to unit i" (Section 2.1). This teaches providing input data to a multi-layer neural network whose layers have weight parameters Wij.); computing a loss of the neural network based on the input data and the weights (OBD: "We assume the objective function is the usual mean-squared error (MSE); generalization to other additive error measures is straightforward ... We approximate the objective function E by a Taylor series. A perturbation δU of the parameter vector will change the objective function by PNG media_image2.png 51 438 media_image2.png Greyscale " (Section 2). This teaches computing a loss function E (the MSE objective) as a function of the input data and the network weights.); determining importance scores for the weights based on the loss and covariance matrix values of the weights (see the teaching of OBS set forth below), an importance score of a weight indicating a measurement of a change in the loss by removing the weight (OBD: “it is more than reasonable to define the saliency of a parameter to be the change in the objective function caused by deleting that parameter ... 4. Compute the saliencies for each parameter: sk = hkk uk2 / 2” (Section 2 and Section 2.2, The Recipe, Step 4). The saliency sk derived from the diagonal second derivative hkk of the loss E and the weight value uk, and directly approximate ΔE – the change in the loss – caused by deleting (setting to zero) weight parameter k. This teaches determining an importance score (saliency) for each weight that indicates a measurement of the change in the loss by removing that weight.), the covariance matrix values indicating influence of the input data (see the teaching of OBS set forth below); selecting one or more weights based on the importance scores of the weights (OBD: "Sort the parameters by saliency and delete some low-saliency parameters" (Section 2.2, The Recipe, Step 5). This teaches selecting one or more weights – specifically those with the lowest importance scores (saliencies) – for removal.); and generating a pruned neural network by changing the one or more selected weights to one or more zeros (OBD: “Deleting a parameter is defined as setting it to 0 and freezing it there” (Section 2.2). This teaches changing the selected low-saliency weights to zero.). OBS states that “We shall show that the Hessian can be reduced to the sample covariance matrix associated with certain gradient vectors” (p. 166). OBS further states that “the covariance form of the Hessian yields a recursive formula for computing the inverse” (p. 166). OBS reports that “Equations 12 and 14 show that H is the sample covariance matrix associated with the gradient variable X” and that “Equation 17 permits the calculation of H-1 using a single sequential pass through the training data” (p. 167). The covariance matrix values are a function of the training inputs. The gradient variable is input dependent. Per Eq. 11 the gradient variable X[k] is the derivative of F(w, in[k]) with respect to w, so X[k] is a function of the k-th training input in[k], and per Eq. 12 the Hessian is the average of the outer products of those gradient variables over the P training patterns (OBS, Eqs. 11 and 12, p. 167). Because H is the sample covariance matrix associated with that input-dependent gradient variable, the covariance matrix values indicate influence of the input data. One of ordinary skill would have replaced OBD’s diagonal saliency with OBS's covariance-based Hessian because OBS identifies OBD’s diagonal assumption as causing pruning of the wrong weights. OBS states, “For computational simplicity, OBD assumes that the Hessian matrix is diagonal: in fact. however. Hessians for every problem we have considered are strongly non-diagonal, and this leads OBD to eliminate the wrong weights” (OBS p. 165). OBS further states that “The superiority of the method described here - Optimal Brain Surgeon - lies in great part to the fact that it makes no restrictive assumptions about the form of the network’s Hessian, and thereby eliminates the correct weights” (OBS p. 165). A skilled artisan would therefore have used OBS’s covariance-based Hessian in OBD’s saliency determination to avoid eliminating the wrong weights. Expectation of success was reasonable because OBS’s saliency definition is “more general than Le Cun et al.'s, and which includes theirs in the special case of diagonal H” (OBS p. 165). Claims 27, 34, and 41. ODB and OBS discloses the one or more non-transitory computer-readable media of The one or more non-transitory computer-readable media of wherein selecting the one or more weights based on the importance scores of the weights comprises: comparing an importance score of a first weight with an importance score of a second weight; and selecting the first weight over the second weight based on the importance score of the first weight being smaller than the importance score of the second weight (OBD: "Sort the parameters by saliency and delete some low-saliency parameters" (Section 2.2, Step 5); "It is clear that deleting parameters by order of saliency causes a significantly smaller increase of the objective function than deleting them according to their magnitude" (Section 2.3). Sorting by saliency rank and deleting those with the lowest (smallest) saliency scores directly teaches comparing the importance scores of individual weights and selecting a weight with a smaller importance score over a weight with a larger importance score for deletion.). Claims 28, 35, and 42. ODB and OBS discloses the one or more non-transitory computer-readable media of The one or more non-transitory computer-readable media of wherein the input data is training data used to train the neural network (OBD: "It was trained on a database of segmented handwritten zip code digits and printed digits containing approximately 9300 training examples" (Section 2.3). This teaches that the input data provided to the neural network is training data used in training the network.). Claims 29, 36, and 43. ODB and OBS discloses the one or more non-transitory computer-readable media of The one or more non-transitory computer-readable media of wherein the neural network has been trained, and the operations further comprise: maintaining one or more values of one or more unselected weights; and after changing the one or more selected weights to the one or more zeros and maintaining the one or more values of the one or more unselected weights, further training the neural network (OBD: "Train the network until a reasonable solution is obtained" (Section 2.2, Step 2) – the neural network has been trained prior to pruning. "Deleting a parameter is defined as setting it to 0 and freezing it there" (Section 2.2) – only the selected low-saliency weights are frozen at zero; the unselected weights retain their current values (are maintained). "Iterate to step 2" (Section 2.2, Step 6) – the network undergoes further training (retraining) after the selected weights are set to zero. This teaches maintaining the values of unselected weights while further training the network after changing selected weights to zeros.). Claims 30, 37, and 44. ODB and OBS discloses the one or more non-transitory computer-readable media wherein further training the pruned neural network comprises: maintaining the one or more zeros; and modifying the one or more values of the one or more unselected weights (OBD: "Deleting a parameter is defined as setting it to 0 and freezing it there" (Section 2.2). The phrase "freezing it there" expressly teaches that the zeroed weights are maintained (held at zero) throughout further training. The remaining unselected weights are then updated through continued backpropagation in Step 2, thereby modifying their values while the zeros are maintained. This teaches maintaining zeros of selected weights and modifying values of unselected weights during further training.). Claims 31, 38, and 45. ODB and OBS discloses the one or more non-transitory computer-readable media of The one or more non-transitory computer-readable media of wherein the operations further comprise: selecting an additional weight from the one or more unselected weights based on one or more importance scores of the one or more unselected weights; and changing the additional weight to a zero (OBD: "Iterate to step 2" (Section 2.2, Step 6). In each subsequent iteration, OBD recomputes the second-derivative saliency scores (Steps 3-4) for the remaining non-zero (previously unselected) weights, then "sort[s] the parameters by saliency and delete[s] some low-saliency parameters" (Step 5) – i.e., selects one or more additional weights from the unselected pool based on their updated importance scores and sets them to zero. This teaches iteratively selecting additional weights from the unselected weights based on importance scores and changing them to zero. See also OBD Section 2.3, Figure 2, showing the performance benefit of iterative pruning with retraining.). Claims 32 and 39 is/are rejected under 35 U.S.C. 103 as being unpatentable over OBD and OBS as applied to claims 26 and 33 above, in view of Han et al., “Learning both Weights and Connections for Efficient Neural Networks,” (hereinafter “Han”). Claims 32 and 39. The one or more non-transitory computer-readable media of The one or more non-transitory computer-readable media of wherein the one or more layers comprises one or more convolutional layers. OBD does not explicitly recite that the one or more layers comprise one or more convolutional layers, as OBD describes its network as a shared-weight architecture. However, Han, in the same field of neural network pruning for computational efficiency, explicitly teaches that the weight selection and zeroing methodology is applicable to one or more convolutional layers (Han: "Both CONV and FC layers can be pruned, but with different sensitivity" (Section 5); "pruning reduces the number of weights by 12x and computation by 6x ... [for layers] conv1, conv2" (Table 3, LeNet-5 results); "We further examine the performance of pruning on the lmageNet ... dataset ... VGG-16 has far more convolutional layers ... We aggressively pruned both convolutional and fully-connected layers" (Section 4.3). This teaches that a weight-pruning-and-zeroing methodology like that of OBD applies to and is particularly beneficial for one or more convolutional layers in a CNN.) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply OBD's importance-score-based weight pruning to convolutional layers, as taught by Han. The motivation for this combination would have been to reduce the storage and computational burden of convolutional layers; which, as Han demonstrates, account for the majority of parameters and arithmetic operations in deep CNNs used for computer vision tasks. Thereby enabling deployment of accurate neural network models on resource constrained mobile and embedded devices without loss of predictive accuracy. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Ross Varndell whose telephone number is (571)270-1922. The examiner can normally be reached M-F, 9-5 EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, O’Neal Mistry can be reached at (313)446-4912. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see https://ppair-my.uspto.gov/pair/PrivatePair. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Ross Varndell/Primary Examiner, Art Unit 2674
Read full office action

Prosecution Timeline

Jan 12, 2024
Application Filed
Apr 07, 2026
Non-Final Rejection mailed — §103
Jul 07, 2026
Response Filed
Sep 21, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749288
EXPLOITING HIERARCHICAL STRUCTURE LEARNING WITH HYPERBOLIC DISTANCE TO ENHANCE OPEN WORLD OBJECT DETECTION
2y 12m to grant Granted Sep 29, 2026
Patent 12749290
STORAGE MEDIUM, DATA GENERATION METHOD, AND INFORMATION PROCESSING DEVICE
2y 11m to grant Granted Sep 29, 2026
Patent 12731395
PROCESSING METHOD, AND PROCESSING SYSTEM
2y 8m to grant Granted Sep 08, 2026
Patent 12731248
SYSTEM AND METHOD FOR DEFECT DETECTION USING A CONDITIONAL MASKED AUTOENCODER
1y 8m to grant Granted Sep 08, 2026
Patent 12664610
MACHINE LEARNING TECHNIQUES TO CREATE HIGHER RESOLUTION COMPRESSED DATA STRUCTURES REPRESENTING TEXTURES FROM LOWER RESOLUTION COMPRESSED DATA STRUCTURES AND TRAINING THEREFOR
4y 9m to grant Granted Jun 23, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
85%
Grant Probability
98%
With Interview (+13.3%)
2y 3m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 632 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month