Prosecution Insights
Last updated: October 01, 2026
Application No. 18/358,629

SYSTEMS AND METHODS FOR SIMULTANEOUS NETWORK PRUNING AND PARAMETER OPTIMIZATION

Final Rejection §101§103
Filed
Jul 25, 2023
Priority
Aug 12, 2022 — provisional 63/371,299
Examiner
CADY, MATTHEW ALAN
Art Unit
2145
Tech Center
2100 — Computer Architecture & Software
Assignee
JPMorgan Chase Bank, N.A.
OA Round
2 (Final)
0%
Grant Probability
At Risk
3-4
OA Rounds
2m
Est. Remaining
0%
With Interview

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 1 resolved
-55.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
3y 4m
Avg Prosecution
22 currently pending
Career history
19
Total Applications
across all art units

Statute-Specific Performance

§101
10.4%
-29.6% vs TC avg
§103
68.7%
+28.7% vs TC avg
§102
11.3%
-28.7% vs TC avg
§112
9.6%
-30.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-4, 6-14, 16-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yi Guo (Hereinafter Guo) (“GDP: Stabilized Neural Network Pruning via Gates with Differentiable Polarization”, 09/08/2021), in view of Andreas Veit (hereinafter Veit) (“Convolutional Networks with Adaptive Inference Graphs”, 05/08/2020). Regarding claim 1, Guo teaches; receiving, by a network optimization computer program, a network to optimize, the network comprising a plurality of layers; [pg. 4] PNG media_image1.png 164 381 media_image1.png Greyscale NOTE: Teaches receiving a network to optimize (original network with learnable weights) with a plurality of layers (L is the total number of layers, which indicates that there can be multiple, i.e., a plurality). [pg. 4] PNG media_image2.png 187 373 media_image2.png Greyscale NOTE: Guo discloses that network weights and gate parameters are optimized by gradient descent and proximal-SGD in its pruning algorithm, which are computer-implemented (via a computer program) optimization procedures for a neural network. Thus, Guo teaches a network optimization computer program for receiving and optimizing a network with a plurality of layers. training, by the network optimization computer program, parameters for the network; [pg. 4] PNG media_image1.png 164 381 media_image1.png Greyscale PNG media_image3.png 208 375 media_image3.png Greyscale NOTE: Teaches training, by the network optimization computer program, parameters for the network (discloses optimizing the weights of the network via gradient descent, which is considered training the parameters of the network). optimizing, by the network optimization computer program, a loss function for the network, [pg. 4] PNG media_image3.png 208 375 media_image3.png Greyscale NOTE: Teaches optimizing the loss function of the network. wherein the network optimization computer program uses a polarization regularizer to reach a consensus static sub-network; ([Abstract] During the training process, the polarization effect will drive a subset of gates to smoothly decrease to exact zero, while other gates gradually stay away from zero by a large margin.) NOTE: The core idea of Guo is to encourage gate values to update smoothly “towards polarization”. [pg. 4] PNG media_image4.png 956 736 media_image4.png Greyscale PNG media_image5.png 130 552 media_image5.png Greyscale NOTE: In the above function gε(x), as ε gets smaller, the gate value becomes polarized (i.e. pushed towards one of two poles, 0 or 1). PNG media_image6.png 523 554 media_image6.png Greyscale NOTE: Each learnable gate parameter α defines a gate value through gε(α), where gε becomes polarized as ε decreases. Those gate parameters are then used to define the resource regularization term R(α), which is added to the task loss. Thus, the objective function uses a regularizer built on a polarizing gate function, i.e. a polarization regularizer under broadest reasonable interpretation (BRI). [pg. 7] PNG media_image7.png 162 457 media_image7.png Greyscale NOTE: When training terminates, there is a final sub-network, which is a fixed/static final sub-network rather than a dynamically changing sub-network. This final sub-network is a compact, stable, and essential sub-component of the super-model, and includes layers and channels that are consistently identified as important across the multiple training runs and can thus be considered a consensus sub-network under BRI. The training/optimization process uses a polarization regularizer, as previously explained. Guo therefore teaches the network optimization computer program using a polarization regularizer to reach a consensus static sub-network. and updating, by the network optimization computer program, parameters for each gating module in the network consistent with the consensus static sub-network. [pg. 4] PNG media_image8.png 267 367 media_image8.png Greyscale NOTE: The gates inserted before each convolution layer can be considered a gating module under BRI. PNG media_image9.png 202 385 media_image9.png Greyscale NOTE: Teaches updating, by the network optimization computer program, parameters for each gating module (gates with learnable parameters PNG media_image10.png 22 28 media_image10.png Greyscale updated during the optimization process) in the network. ([pg. 7] We do not need to sample sub-net or calculate expectation when training. Figure 11 shows the training stability of GDP. When training terminates, there is only one candidate sub-net left. Different value of lambda in Function (7) result in sub-net with different FLOPs and performance. We want the performance of the super-net to reflect the performance of the final sub-net as much as possible.) NOTE: Discloses that the result of the training process is the aforementioned consensus static-subnetwork (the final sub-net). Thus, the aforementioned updates for the parameters for each gating module in the network performed during the training/optimization process are consistent with the resulting consensus static sub-network. selecting, by the network optimization computer program, layer pruning and/or channel pruning for the layers within the network; [pg. 3] PNG media_image11.png 139 368 media_image11.png Greyscale NOTE: Teaches selecting by the network optimization computer program, layer pruning and/or channel pruning for the layers within the network (gates before layers to control the on-and-off of each channel or the whole block/layer, which is considered pruning via gates). providing, by the network optimization computer program, a gating module at layers within the network, [pg. 3] PNG media_image11.png 139 368 media_image11.png Greyscale NOTE: Teaches providing, by the network optimization computer program, a gating module at layers within the network (plugging trainable gates before layers of the network) Guo fails to teach but Veit teaches; receiving, by the network optimization computer program a sparsity hyperparameter (Note: Because the target rate controls the proportion of gated layers maintained in the executed/open state, t constitutes a sparsity hyperparameter under BRI … ([pg. 3] a gate … decides whether to execute the next layer. The gate chooses between two discrete states: 0 for ‘off’ and 1 for ‘on’) … [pg. 5] constrain how often each layer is allowed to be used … we use ... an additional loss term … target rate … Each layer has a target rate t [pg. 10] we set the target rate for the early layers of ConvNet-AIG 101 to 1 so that they are always executed); wherein each gating module opens or closes a gate in the gating module based on an output of a binary head ([pg. 3] gl(xl−1) is a gate that … decides whether to execute the next layer. The gate chooses between two discrete states: 0 for ‘off’ and 1 for ‘on’ … [pg. 4] Each gate comprises two parts. The first part estimates the relevance of the layer to be executed… The goal of the second component is to make a discrete decision based on the relevance scores); extracting, by the network optimization computer program, gate open/close features from the gating modules ([pg. 4] Each gate comprises two parts. The first part estimates the relevance of the layer to be executed … relevance score for the layer … is a vector β containing two unnormalized scores for the actions of (a) computing and (b) skipping the following layer) modifying by the network optimization computer program a number of open gating modules ([pg. 3] a gate … decides whether to execute the next layer. The gate chooses between two discrete states: 0 for ‘off’ and 1 for ‘on’) in the [network] based on the sparsity hyperparameter ([pg. 5] constrain how often each layer is allowed to be used … we use ... an additional loss term that encourages each layer to be executed at a certain target rate [pg. 10] we set the target rate for the early layers of ConvNet-AIG 101 to 1 so that they are always executed); OBVIOUSNESS TO COMBINE VEIT: Veit is analogous art to the present disclosure as it pertains to a binary gating architecture for neural networks. It would have been obvious to one of ordinary skill in the art, before the effective filing date, to modify Guo’s gate-based pruning method to incorporate Veit’s target rate-controlled gating technique and associated binary gating architecture to provide a user controllable degree of network sparsity while reducing unnecessary computation. Veit expressly teaches that its target rate loss constrains how often each layer is permitted to execute and that “[pg. 4] The target rate provides an easy instrument to adjust computation time.” Veit further teaches that its lightweight gating mechanism introduces minimal computational overhead while permitting a substantial number of layers to be skipped (“[pg. 4] the gating function adds only a computational overhead of 0.04%, but allows to skip 38% of its layers on average”). A person of ordinary skill therefore would have been motivated to employ Veit’s target rate control in Guo’s pruning framework so that the desired degree of gate retention pruning could be specified through a sparsity hyperparameter and reflected in the number of gates retained in Guo’s consensus status sub-network, thereby predictably providing a greater control over the tradeoff between network computation and retained model capacity. Regarding claim 2, Guo teaches; wherein the gating modules comprise a channel pruning gating module and/or a layer pruning gating module. [pg. 3] PNG media_image11.png 139 368 media_image11.png Greyscale NOTE: Teaches the gating modules (trainable gates before the layers used for pruning components of the network) comprising a channel pruning gating module and/or a layer pruning gating module (the gates of the aforementioned gating modules control the on-and-off of each channel or the entire layer block, which is considered channel/layer pruning). Regarding claim 3, Guo teaches; wherein the network comprises a residual network. [Abstract] PNG media_image12.png 93 369 media_image12.png Greyscale NOTE: Teaches the network comprising a residual network (Guo applies GDP to the stated ResNet model, where ResNet is a residual network). Regarding claim 4, Guo fails to explicitly teach but Veit teaches; wherein the network comprises a sequential network ([pg. 1] In this work, we propose convolutional networks with adaptive inference graphs (ConvNet-AIG)). NOTE: From the applicant’s spec; “[0030] Examples of sequential networks include convolutional neural networks” OBVIOUSNESS: Using the same reasoning from claim 1. Regarding claim 6, Guo teaches; wherein the gating modules are added at a beginning of each layer and each channel within the network. ([Abstract] In view of the research gaps, we present a new module named Gates with Differentiable Polarization (GDP), inspired by principled optimization ideas. GDP can be plugged before convolutional layers without bells and whistles, to control the on-and-off of each channel or whole layer block.) NOTE: Teaches the gating modules (GDP is a gating module) added at the beginning of each layer (GDP can be plugged before convolutional layers) and each channel (each channel of the layer) within the network. Regarding claim 7, Guo fails to teach but Veit teaches; wherein each gating module comprises a fully connected layer ([pg. 4] The goal of the gate’s first component is to estimate the associated layer’s relevance given the input features … we add a simple non-linear function of two fully-connected layers) and a binary head ([pg. 3] gl(xl−1) is a gate that … decides whether to execute the next layer. The gate chooses between two discrete states: 0 for ‘off’ and 1 for ‘on’ … [pg. 4] Each gate comprises two parts … The goal of the second component is to make a discrete decision), wherein the fully connected layer receives a one-dimensional vector ([pg. 4] The input to the gate is the output of the previous layer … we only consider channel-wise means … This compresses the input features into a 1 × 1 × C channel descriptor zc … To capture the dependencies between channels, we add a simple non-linear function of two fully-connected layers), multiplies the one-dimensional vector by a weight matrix ([pg. 4] W2σ(W1z)), and outputs an output vector ([pg. 4] The output of this operation is the relevance score for the layer … vector β … β =W2σ(W1z)), and the binary head receives the output vector and returns a binary value indicating whether the layer will be computed ([pg. 3] gl(xl−1) is a gate that … decides whether to execute the next layer. The gate chooses between two discrete states: 0 for ‘off’ and 1 for ‘on’ … [pg. 4] Each gate comprises two parts … The goal of the second component is to make a discrete decision based on the relevance scores). OBVIOUSNESS: Using the same reasoning from claim 1. Regarding claim 8, Guo fails to teach but Veit teaches; wherein the binary head comprises a straight-through estimator ([pg. 4] the second component … For this, … we utilize the Gumbel-Max trick [9] and its recent continuous relaxation [19,25] … [pg. 5] One option to employ the Gumbel-softmax estimator is to use the continuous version … An alternative is the straight through version [19] of the Gumbel-softmax estimator). OBVIOUSNESS: Using the same reasoning from claim 1. Additionally, it would further have been obvious to employ Veit’s straight through Gumbel-softmax implementation because Veit states that “[pg. 5] the straight-through estimator performs better.” Regarding claim 9, Guo fails to teach but Veit teaches; wherein a gradient of the straight-through estimator updates the parameters of the gating modules through back propagation. ([pg. 2] train … the discrete gates … end-to-end [pg. 5] An alternative is the straight through version [19] of the Gumbel-softmax estimator. There, during training, … during the backwards pass we compute the gradient of the softmax relaxation in Equation 9… We illustrate the two different paths during the forward and backward pass in Figure 3) PNG media_image13.png 323 969 media_image13.png Greyscale NOTE: Veit teaches a straight through Gumbel-softmax implementation, where, during training, the gradient of the softmax relaxation is propagated backward through the gating unit during the backward pass. Because the gating unit comprises trainable parameters, including the aforementioned fully-connected layer weight matrices, the backward propagated gradient is thereby used to train / update the parameters of the gating unit through backpropagation. OBVIOUSNESS: Using the same reasoning from claim 1. Additionally, it would further have been obvious to employ Veit’s straight through Gumbel-softmax implementation because Veit states that “[pg. 5] the straight-through estimator performs better.” Regarding claim 10, Guo fails to teach but Veit teaches; wherein the parameters comprise weight matrices. [pg. 4] PNG media_image14.png 68 478 media_image14.png Greyscale NOTE: Veit teaches weight matrices W1 W2 as parameters for the gating unit. OBVIOUSNESS: Using the same reasoning from claim 1. Regarding claims 11-14, 16-20 Claims 11-14, 16-20 are computer readable medium claims directly corresponding to claims 1-4, 6-10, respectively and are therefore rejected using the same reasoning. Response to Arguments The 35 USC 112(b) rejections of claims 7 and 17 have been withdrawn considering the amendments. Claims 1-20 were not rejected under 35 USC 101 in the previous office action. Accordingly, claims 1-4, 6-14, and 16-20 are also not rejected under 35 USC 101 in the present office action. Applicant's arguments filed 07/08/2026 regarding the rejections under 35 USC 103 have been fully considered but they are not persuasive. Starting on page 4, the applicant states that “Claims 1-6 and 11-16 stand rejected under 35 U.S.C. § 103 as allegedly rendered obvious by the Non-Patent Literature "GDP: Stabilized Neural Network Pruning via Gates with Differentiable Polarization" by Guo and the Non-Patent Literature "Dynamic Channel and Layer Gating in Convolutional Neural Networks" by Bejnordi et al. ("Bejnordi") … Applicant respectfully disagrees, as the Office Action has failed to establish a prima facie case of obviousness.” Examiner respectfully disagrees. In the present office action, claims 1-4, 6-14, 16-20 stand rejected under Guo in view of Veit, as necessitated by the amendments to independent claims 1 and 11. Guo continues to teach the gate-polarization/pruning process and resulting consensus static sub-network, while Veit teaches controlling gated layer utilization according to a target rate hyperparameter. As reflected by the present office action, it would have been obvious to employ Veit’s target rate control in Guo’s gate-based pruning process to permit control over the desired degree of gate activation/pruning and corresponding network computation. The applicant further states that “Without conceding that the proposed combination of Guo and Bejnordi is proper, Applicant respectfully submits that the proposed combination fails to disclose all elements of independent claim 1 … claim 1 has been amended to include the steps of "receiving, by the network optimization computer program, a sparsity hyperparameter" and "modifying, by the network optimization computer program, a number of open gating modules in the consensus static sub- network based on the sparsity hyperparameter." The proposed combination does not disclose at least these elements. Instead, Bejnordi actually teaches away from the use of a sparsity hypermeter: “Unlike ConvNet-AIG[], we propose to remove the target rate. This would allow different layers/channels to take varying dynamic execution rates. This way, the network may automatically learn to use more units for a specific layer and less for another, without us having to determine a target rate in advance. Besides that, we give weight to the sparsity loss by the coefficients k and y which control the pressure on the sparsity loss for layer gating and channel gating, respectively.” Bejnordi, page 8. Thus, Bejnordi does not disclose the use of a sparsity hyperparameter as claimed, and teaches away from the ConvNet's use of sparsity parameters. Guo does not cure these deficiencies. While of a different scope, independent claim 11 recites similar elements and is allowable for at least these reasons set forth above. Therefore, for at least these reasons, Applicant respectfully requests that the rejections of independent claims 1 and 11, and of all claims dependent thereon, be withdrawn. Claims 7-10 and 17-20 stand rejected under 35 U.S.C. § 103 … Applicant notes that these claims are dependent on independent claims 1 or 11 and are allowable for at least the reasons discussed above.” Examiner respectfully disagrees. In the present office action, claims 1-4, 6-14, 16-20 stand rejected under Guo in view of Veit, as necessitated by the amendments to independent claims 1 and 11. Thus, the Applicant’s argument that Bejnordi teaches away from the target rate approach does not overcome the present rejection, because the present does not rely of Bejnordi for the newly added limitations. The present rejection relies on Veit for the newly added sparsity hyperparameter limitations, while Guo supplies the consensus static sub-network into which known control is incorporated for the reasons set forth above. As reflected by the present office action, Veit expressly employs a sparsity hyperparameter (target-rate t) to constrain how often gated layers are executed and teaches that the target rate provides an instrument for controlling computation. Thus, Veit affirmatively teaches the newly included limitations of amended claims 1 and 11. Accordingly, the 35 USC 103 rejections of the present office action for claims 1-4, 6-14, 16-20 stand. CONCLUSION Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Matthew Alan Cady whose telephone number is (571) 272-7229. The examiner can normally be reached Monday - Friday, 7:30 am - 5:00 pm ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Cesar Paula can be reached on (571)272-4128. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MATTHEW ALAN CADY/ Examiner, Art Unit 2145 /CESAR B PAULA/ Supervisory Patent Examiner, Art Unit 2145
Read full office action

Prosecution Timeline

Jul 25, 2023
Application Filed
Apr 09, 2026
Non-Final Rejection mailed — §101, §103
Jul 08, 2026
Response Filed
Sep 14, 2026
Final Rejection mailed — §101, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
0%
Grant Probability
0%
With Interview (+0.0%)
3y 4m (~2m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 1 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month