Prosecution Insights
Last updated: August 17, 2026
Application No. 17/398,673

PERFORMANCE-AWARE SIZE REDUCTION FOR NEURAL NETWORKS

Non-Final OA §103§112
Filed
Aug 10, 2021
Examiner
BOSTWICK, SIDNEY VINCENT
Art Unit
2124
Tech Center
2100 — Computer Architecture & Software
Assignee
NVIDIA Corporation
OA Round
5 (Non-Final)
52%
Grant Probability
Moderate
5-6
OA Rounds
0m
Est. Remaining
89%
With Interview

Examiner Intelligence

Grants 52% of resolved cases
52%
Career Allowance Rate
76 granted / 147 resolved
-3.3% vs TC avg
Strong +37% interview lift
Without
With
+36.9%
Interview Lift
resolved cases with interview
Typical timeline
4y 5m
Avg Prosecution
41 currently pending
Career history
214
Total Applications
across all art units

Statute-Specific Performance

§101
25.3%
-14.7% vs TC avg
§103
45.2%
+5.2% vs TC avg
§102
4.9%
-35.1% vs TC avg
§112
24.3%
-15.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 147 resolved cases

Office Action

§103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 6/4/2026 has been entered. Remarks This Office Action is responsive to Applicants' Amendment filed on June 4, 2026, in which claims 1-18 and 25-30 are currently amended. Claims 1-18 and 25-30 are currently pending. Information Disclosure Statement The information disclosure statement (IDS) submitted on June 4, 2026 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Response to Arguments The rejections to claims 2, 3, 6, 8, 9, 10, 12, 14, 15, 16, 18, 26, 27, 28, and 30 under 35 U.S.C. § 112(b) are hereby withdrawn, as necessitated by applicant's amendments and remarks made to the rejections. Applicant’s arguments with respect to rejection of claims 1-18 and 25-30 under 35 U.S.C. 103 based on amendment have been considered and are persuasive. The argument is moot in view of a new ground of rejection set forth below. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 1-18 and 25-30 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Regarding claims 1, 7, 13, and 25, "different latency step sizes for different layers of one or more neural networks" is indefinite. It's unclear if the relationship is one-to-one or one-to-many. For example, the claim limitation may be interpreted as every layer having a step size that is completely unique to that layer where no step sizes are repeated, or there being two or more step sizes among two or more layers (Layer A with Step Size A, Layer B with Step Size B, Layer C with Step Size A), or something else altogether. Since these interpretations are contradictory the scope of the claim cannot be determined. In the interest of further examination the claim is interpreted as having two or more step sizes among two or more layers. The remaining claims are rejected with respect to their dependence on the rejected claims. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1, 3-7, 9-13, 15-18, 25, and 27-30 are rejected under U.S.C. §103 as being unpatentable over the combination of Gamanayake (“Cluster Pruning: An Efficient Filter Pruning Method for Edge AI Vision Applications”, 2020) and Hoang (US20210110235A1). Regarding claim 1, Gamanayake teaches A processor, comprising: one or more circuits to: ([p. 7] "To test the inference accuracy of the test datasets using GPU and CPU") access pre-generated data including different latency step sizes for different layers of one or more neural networks,([p. 4] "we prune the less significant filters in a layer, then profile the accuracy and latency response for a given hardware architecture […] we start to prune them in ascending order of the rank and profiled the accuracy and latency of the network for each pruning instance as shown in Fig. 4, Fig. 5, and Fig. 8." [p. 4] "in Fig. 5 and Fig. 8, we can identify periodic bottoms with significant drops and periodic peaks with significant rises in the latency and accuracy graphs" [p. 4] "The optimum cluster size Pl can be calculated as the LCM(pacc l , plat l), where LCM(.,.) represent the calculation of least common multiple of the two given periodic lengths" See Algorithm 1 where layer specific optimum cluster size is an input, in other words it is explicitly pre-generated data. The optimum cluster size being a direct function of the step size/periodic length.) the different latency step sizes corresponding to size granularities of the different layers at which inference time on a specified hardware target changes according to a step function([p. 4] "If the hardware architecture is susceptible to workload imbalance, the influence will be reflected in the performance graphs when we analyse the single layer pruning results. If there exist a particular pattern of inference time drops in latency graphs with respect to the number of filters left in a layer after pruning, networks forward inference time is influenced") order individual neurons of the different layers by individual neuron accuracy importance scores for the individual neurons of the different layers;([p. 4] "evaluate the importance of each filter in a neural network […] Using the Eq. 3, we ranked the filters according to their increasing order of significant" [p. 4] "we iterate through each layer of the network and identify the importance of individual filter in corresponding layer according to the minimum weight criteria using Eq. 3." Filter interpreted as synonymous with neuron) group the ordered individual neurons into groups having respective size granularities determined from the pre-generated data; ([p. 4 Algorithm 1] "Optimum Cluster Sizes per layer: P […] Select a cluster of filters" [p. 4] "For each layer, filter clusters are formed according to the optimum cluster size") determine a respective score for each group based on a combination of the individual neuron accuracy importance scores for the ordered individual neurons of each group;([P. 4] "The importance of the filter groups are calculated by taking the average of ΘMW values of the filters inside the corresponding group. After that, all the groups in the network are ranked and pruned according to their increasing order of significance") and cause individual groups of ordered individual neurons to be removed at a group level based, at least in part, on the respective score for each group ([P. 4] "The importance of the filter groups are calculated by taking the average of ΘMW values of the filters inside the corresponding group. After that, all the groups in the network are ranked and pruned according to their increasing order of significance") and a specified performance target such that different numbers of neurons are removed from different layers([P. 4 Algorithm 1] "GR = Rank G according to MWG Until the pruning objective is reached, prune filter groups in GR consecutively" [P. 4] "The importance of the filter groups are calculated by taking the average of ΘMW values of the filters inside the corresponding group. After that, all the groups in the network are ranked and pruned according to their increasing order of significance"). However, Gamanayake does not explicitly teach the relationship between filters and neurons. Hoang, in the same field of endeavor, teaches that filters are neurons ([¶0046] "Each neuron in a neural network computes an output value by applying a specific function to the input values coming from the receptive field in the previous layer. The function that is applied to the input values is determined by a vector of weights and a bias. Learning, in a neural network, progresses by making iterative adjustments to these biases and weights. The vector of weights and the bias are called filters and represent particular features of the input (e.g., a particular shape)"). Gamanayake as well as Hoang are directed towards pruning convolutional neural networks. Therefore, Gamanayake as well as Hoang are analogous art in the same field of endeavor. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of Gamanayake with the teachings of Hoang by treating filters as neurons. Hoang provides as additional motivation for combination ([¶0064] “In practice, the equivalent convolution is normally implemented by statically identical copies of the neuron to different input regions.”). This motivation for combination also applies to the remaining claims which depend on this combination. Regarding claim 3, the combination of Gamanayake, and Hoang teaches The processor of claim 1, wherein respective individual neuron accuracy importance scores are calculated for the individual neurons after pre-training of the one or more neural networks.(Gamanayake [p. 4 Algorithm 1] "Input: Pretrained Network with Filter Set: F […] Set for ΘMW(.) values in current layer: MWl = {} for each filter in current layer k = 1,...,Kl do MWl∪{ΘMW(Fk l)} [...] FR = Rank Fl according to values in MWl" [p. 5] "we use two models from this network, which are pre-trained on above mentioned two datasets"). Regarding claim 4, the combination of Gamanayake, and Hoang teaches The processor of claim 3, wherein the one or more circuits are further to group sets of neurons based at least in part upon a similarity of the individual neuron accuracy importance scores(Gamanayake [p. 4 Algorithm 1] "Input: Pretrained Network with Filter Set: F […] Set for ΘMW(.) values in current layer: MWl = {} for each filter in current layer k = 1,...,Kl do MWl∪{ΘMW(Fk l)} [...] FR = Rank Fl according to values in MWl" [p. 5] "we use two models from this network, which are pre-trained on above mentioned two datasets"). Regarding claim 5, the combination of Gamanayake, and Hoang teaches The processor of claim 1, wherein the specified hardware target is a target type of hardware to be used to perform inferencing using the one or more neural networks. (Gamanayake [p. 2] "We formulate an optimization problem to measure the hardware response towards the performance (accuracy and inference latency) of the network" See FIG. 5). Regarding claim 6, the combination of Gamanayake, and Hoang teaches The processor of claim 1, wherein the one or more circuits are further to utilize an optimization solver to determine the individual groups to be removed to optimize accuracy for the specified performance target(Gamanayake [p. 2] "We formulate an optimization problem to measure the hardware response towards the performance (accuracy and inference latency) of the network" [p. 4 Algorithm 1] "Optimum Cluster Sizes per layer: P […] Select a cluster of filters" [p. 4] "For each layer, filter clusters are formed according to the optimum cluster size" See also Algorithm 1). Regarding claims 7 and 9-12, 7 claims 7 and 9-12 are substantially similar to claims 1 and 3-6. Therefore, the rejections applied to claims 1 and 3-6 also apply to claims 7 and 9-12. Regarding claims 13 and 15, claims 13 and 15 are directed towards the method performed by the processor of claims 1 and 3. Therefore, the rejections applied to claims 1 and 3 also apply to claims 13 and 15. Regarding claim 16, the combination of Gamanayake, and Hoang teaches The method of claim 15, further comprising: grouping sets of neurons based at least in part upon a similarity of the individual neuron accuracy importance scores and the latency step sizes (Gamanayake [p. 4 Algorithm 1] "Input: Pretrained Network with Filter Set: F […] Set for ΘMW(.) values in current layer: MWl = {} for each filter in current layer k = 1,...,Kl do MWl∪{ΘMW(Fk l)} [...] FR = Rank Fl according to values in MWl" [p. 5] "we use two models from this network, which are pre-trained on above mentioned two datasets" See Algorithm 1 where layer specific optimum cluster size is an input, in other words it is explicitly pre-generated data. The optimum cluster size being a direct function of the latency step size/periodic length.). Regarding claims 17 and 18, claims 17 and 18 are directed towards the method performed by the processor of claims 5 and 6. Therefore, the rejections applied to claims 5 and 6 also apply to claims 17 and 18. Regarding claim 25, claim 25 is substantially similar to claim 1. Therefore, the rejection applied to claim 1 also applies to claim 25. Claim 25 also recites additional elements memory for storing network parameters for the one or more neural networks. (Hoang ([¶Abstract] “Techniques are presented for accelerating in-memory matrix multiplication operations for a convolution neural network (CNN) inference in which the weights of a filter are stored in the memory of a storage class memory device”). Similarly, regarding claims 27-30, claims 27-30 are substantially similar to claims 3, 16, 5, and 6, respectively. Therefore, the rejections applied to claims 3, 16, 5, and 6 also apply to claims 27-30. Claims 2, 8, 14, and 26 are rejected under U.S.C. §103 as being unpatentable over the combination of Gamanayake and Hoang and in further view of Elkerdawy ("To Filter Prune, or to Layer Prune, That Is The Question", 2019). Regarding claim 2, the combination of Gamanayake, and Hoang teaches the processor of claim 1. However, the combination of Gamanayake, and Hoang doesn't explicitly teach, wherein to access the pre-generated data the one or more circuits are further to access a look-up table of pre-measured performance values for the specified hardware target.. Elkerdawy, in the same field of endeavor, teaches the processor of claim 1, wherein to access the pre-generated data the one or more circuits are further to access a look-up table of pre-measured performance values for the specified hardware target.([p. 5 §2] "A lookup table is built for latency prediction and then multiple candidates are generated at each pruning iteration by pruning a ratio of filters from each layer independently. The candidate with the highest accuracy is then selected"). The combination of Gamanayake and Hoang as well as Elkerdawy are directed towards pruning neural networks. Therefore, the combination of Gamanayake and Hoang as well as Elkerdawy are analogous art in the same field of endeavor. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of the combination of Gamanayake and Hoang with the teachings of Elkerdawy by performing the accuracy importance score based pruning of Elkerdawy in addition to the step-wise pruning of The combination of Gamanayake and Hoang (for example in the iterative hardware and accuracy aware pruning step in The combination of Gamanayake and Hoang). Elkerdawy provides as additional motivation for combination of LayerPrune ([p. 13 §4.3] "LayerPrune outperforms SSS on the same latency budget even when SSS supports block pruning for ResNet50, which shows the effectiveness of accuracy approximation as layer importance"). This motivation for combination also applies to the remaining claims which depend on this combination. Regarding claim 8, claim 8 is substantially similar to claim 2. Therefore, the rejection applied to claim 2 also applies to claim 8. Regarding claim 14, the combination of Gamanayake, and Hoang teaches The method of claim 13. However, the combination of Gamanayake, and Hoang doesn't explicitly teach further comprising: calculating, according to a look-up table of pre-measured performance values for the specified hardware target, an impact on performance for each of the one or more groups before determining the one or more groups to be removed.. Elkerdawy, in the same field of endeavor, teaches The method of claim 13, further comprising: calculating, according to a look-up table of pre-measured performance values for the specified hardware target, an impact on performance for each of the one or more groups before determining the one or more groups to be removed.([p. 5 §2] "A lookup table is built for latency prediction and then multiple candidates are generated at each pruning iteration by pruning a ratio of filters from each layer independently. The candidate with the highest accuracy is then selected"). The combination of Gamanayake and Hoang as well as Elkerdawy are directed towards pruning neural networks. Therefore, the combination of Gamanayake and Hoang as well as Elkerdawy are analogous art in the same field of endeavor. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of the combination of Gamanayake and Hoang with the teachings of Elkerdawy by performing the accuracy importance score based pruning of Elkerdawy in addition to the step-wise pruning of The combination of Gamanayake and Hoang (for example in the iterative hardware and accuracy aware pruning step in The combination of Gamanayake and Hoang). Elkerdawy provides as additional motivation for combination of LayerPrune ([p. 13 §4.3] "LayerPrune outperforms SSS on the same latency budget even when SSS supports block pruning for ResNet50, which shows the effectiveness of accuracy approximation as layer importance"). This motivation for combination also applies to the remaining claims which depend on this combination. Regarding claim 26, claim 26 is substantially similar to claim 2. Therefore, the rejection applied to claim 2 also applies to claim 26. Claim 26 also recites additional elements memory for storing network parameters for the one or more neural networks. (Hoang ([¶Abstract] “Techniques are presented for accelerating in-memory matrix multiplication operations for a convolution neural network (CNN) inference in which the weights of a filter are stored in the memory of a storage class memory device”). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to SIDNEY VINCENT BOSTWICK whose telephone number is (571)272-4720. The examiner can normally be reached M-F 7:30am-5:00pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Miranda Huang can be reached on (571)270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SIDNEY VINCENT BOSTWICK/Examiner, Art Unit 2124
Read full office action

Prosecution Timeline

Show 4 earlier events
Sep 15, 2025
Request for Continued Examination
Oct 05, 2025
Response after Non-Final Action
Oct 29, 2025
Non-Final Rejection mailed — §103, §112
Jan 29, 2026
Response Filed
Mar 04, 2026
Final Rejection mailed — §103, §112
Jun 04, 2026
Request for Continued Examination
Jun 06, 2026
Response after Non-Final Action
Jul 21, 2026
Non-Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12699874
Leveraging Redundancy in Attention with Reuse Transformers
3y 10m to grant Granted Aug 04, 2026
Patent 12675673
NEURAL NETWORK PROCESSING DEVICE, METHOD, AND COMPUTER-READABLE RECORDING MEDIUM
3y 7m to grant Granted Jul 07, 2026
Patent 12645914
INSTRUCTION PRUNING FOR NEURAL NETWORKS
3y 6m to grant Granted Jun 02, 2026
Patent 12626139
SECRET SOFTMAX FUNCTION CALCULATION SYSTEM, SECRET SOFTMAX FUNCTION CALCULATION APPARATUS, SECRET SOFTMAX FUNCTION CALCULATION METHOD, SECRET NEURAL NETWORK CALCULATION SYSTEM, SECRET NEURAL NETWORK LEARNING SYSTEM, AND PROGRAM
4y 3m to grant Granted May 12, 2026
Patent 12619815
Magnitude Invariant Multimodal Agent for Efficient Image-Text Interface Automation
1y 6m to grant Granted May 05, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
52%
Grant Probability
89%
With Interview (+36.9%)
4y 5m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 147 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month