Prosecution Insights
Last updated: October 04, 2026
Application No. 17/855,955

SYSTEM AND METHOD FOR EVALUATING WEIGHT INITIALIZATION FOR NEURAL NETWORK MODELS

Final Rejection §103§112
Filed
Jul 01, 2022
Priority
Sep 17, 2021 — provisional 63/245,281
Examiner
BOSTWICK, SIDNEY VINCENT
Art Unit
2124
Tech Center
2100 — Computer Architecture & Software
Assignee
Cognizant Technology Solutions US Corp.
OA Round
4 (Final)
51%
Grant Probability
Moderate
5-6
OA Rounds
1m
Est. Remaining
86%
With Interview

Examiner Intelligence

Grants 51% of resolved cases
51%
Career Allowance Rate
78 granted / 152 resolved
-3.7% vs TC avg
Strong +35% interview lift
Without
With
+35.1%
Interview Lift
resolved cases with interview
Typical timeline
4y 5m
Avg Prosecution
43 currently pending
Career history
216
Total Applications
across all art units

Statute-Specific Performance

§101
24.6%
-15.4% vs TC avg
§103
46.5%
+6.5% vs TC avg
§102
4.6%
-35.4% vs TC avg
§112
24.0%
-16.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 152 resolved cases

Office Action

§103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Remarks This Office Action is responsive to Applicants' Amendment filed on June 3, 2026, in which claims 1-9, 11, 13, 15-23, and 25 are currently amended. Claims 1-25 are currently pending. Response to Arguments The previous rejections to claims 1-25 under 35 U.S.C. § 112(b) are hereby withdrawn, as necessitated by applicant's amendments and remarks made to the rejections. Applicant’s arguments with respect to rejection of claims 1-25 under 35 U.S.C. 101 based on amendment have been considered and are persuasive. The previous rejections to claims 1-25 under 35 U.S.C. § 101 are hereby withdrawn, as necessitated by applicant's amendments and remarks made to the rejections. Applicant’s arguments with respect to rejection of claims 1-25 under 35 U.S.C. 103 based on amendment have been considered and are persuasive. The argument is moot in view of a new ground of rejection set forth below. Claim Objections Claims 1, 13, and 25 are objected to because of the following informalities: "the initial value of mean of" should read "an initial value of a mean of". Similarly, “the initial value of variance” should read “an initial value of a variance”. Appropriate correction is required. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 1-25 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Regarding claims 1, 13, and 25, "wherein if it is determined that a mean-variance mapping function (g-layer) associated with a corresponding layer is not pre-defined, then deriving the mean variance mapping function (g-layer) for that layer using data analytics based on user inputs and storing it in the mean-variance mapping table, or the mean and variance are assumed to remain unchanged after propagation through the corresponding layer" is grammatically indefinite. It is unclear what the word "or" is connecting. The language can be read to mean that when a mean-variance function is not predefined the processor has two options: either derive and store a new mapping function, or assume that the mean and variance remain unchanged. Alternatively, it can also be read to mean that the assumption that the mean and variance remain unchanged is a separate alternative to the entire preceding step. Because the sentence does not clearly identify which two actions are alternatives one of ordinary skill in the art could not reasonably determine the scope of the claim. In the interest of further examination the claim is interpreted as the mean and variance remaining unchanged is a separate alternative to the entire preceding step. Regarding claims 1, 13, and 25, “or the mean and variance” lacks antecedent basis. “Or a mean and variance” is recommended. Regarding claims 1, 13, and 25, “the initial value of mean” and “the initial value of variance” lack antecedent basis. Claims 1, 13, and 25 recite “setting an initial value of respective weight parameters”, however, there is no clear antecedent for an initial value of a mean or variance. “An initial value of a mean” and “an initial value of a variance” are recommended. Regarding claims 6, 14, and 19, “the neural network” lacks antecedent basis. Claim 1 from which claim 6 depends recites “one or more neural network models” (simply “neural network models” in claim 13 from which claims 14 and 19 depend) and further specifies “an untrained neural network” such that it would be unclear which neural network was “the neural network”. In the interest of further examination “the neural network” is interpreted as “the untrained neural network”. The remaining claims are rejected with respect to their dependence on the rejected claims. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1-6, 8-13, 15-19, and 21-25 are rejected under U.S.C. §103 as being unpatentable over the combination of Shekhovtsov (“Normalization of Neural Networks using Analytic Variance Propagation”, 2018) and Schilling (“The Effect of Batch Normalization on Deep Convolutional Neural Networks”, 2016). Regarding claim 1, Shekhovtsov teaches A method for evaluating weight initialization technique for individual layers of neural network models to analytically preserve mean and variance across layers, ([p. 1] "The proposed estimation uses an analytic propagation of mean and variance of the training set through the network" This is essentially Shechovtsov's stated purpose. It analytically determines neural network activation moments through the network and expressly uses those statistics for initialization) wherein the method is implemented by a processor executing program instructions stored in a memory, the method comprising:([p. 5] "init=BN: BN is introduced with s random and b = 0 (pytorch default) […] Our implementation in pytorch will be made publicly available") building, by the processor, a mean-variance mapping table ([p. 2] "A basic neural network composes linear transforms and coordinate-wise non-linearities" [p. 3] " for many cases of practical interest it is tractable" The "table" can reasonably be read as an organized, machine-usable association between layer types and their corresponding moment-transfer functions in view of the instant specification which does not explicitly limit the table structure.) comprising a plurality of mean- variance mapping functions (g-layer), ([p. 2] "The statistics of a linear transform Y = wTX are given by [Eqn. 3a and 3b]" Shekhovtsov supplies multiple mathematical functions taking incoming moments, and for weighted layers w, to outgoing moments. These are functionally the claimed g-layers) each of the mean-variance mapping functions (g-layer) of the plurality of mean-variance mapping functions (g-layer) corresponds to a respective layer of a plurality of layers used in one or more neural network models, ([p. 3] "the approximation can be applied layer-by-layer, propagating mean and variance. Let us consider a deep NN") wherein each of the plurality of mean-variance mapping functions (g-layer) are pre-defined in the mean-variance mapping table([p. 3] "for many cases of practical interest it is tractable. In particular it is the case for sigmoid f(x) = 1/(1 + e−x) and ReLU(x) = max(0,x) non-linearities and for max(x1,x2), which can be used to implement max pooling and max-out" These mappings exist as general analytic rules before they are applied to a particular layer instance. Network-specific mean and variance are arguments to the rule, not newly learned definitions of the rule. In the disclosed Pytorch implementation those predetermined operation-dependent rules supply the functional contents of the claimed table.) receiving, by the processor, an untrained neural network comprising a plurality of layers ([p. 2] "Starting from some initialization of network weights, such as random or orthogonal" [p. 5] "Initialize weights randomly") and an input dataset ([p. 6] "Datasets We used MNIST and CIFAR10") deriving, by the processor, a mean-variance mapping function (g-layer) for each of the respective layers of the untrained neural network from the mean-variance mapping table, ([p. 3] "the approximation can be applied layer-by-layer, propagating mean and variance. Let us consider a deep NN […] Propagate the moments until the first normalization layer" The implementation identifies the operation being processed and invokes/instantiates its predefined analytic moment rule. This is functionally selecting the corresponding g-layer from the organized collection of available mappings) and wherein if it is determined that a mean-variance mapping function (g-layer) associated with a corresponding layer is not pre-defined, then deriving the mean variance mapping function (g-layer_ for that layer using data analytics based on user inputs and storing it in the mean-variance mapping table, or the mean and variance are assumed to remain unchanged after propagation through the corresponding layer;([p. 3] "for many cases of practical interest it is tractable. In particular it is the case for sigmoid f(x) = 1/(1 + e−x) and ReLU(x) = max(0,x) non-linearities and for max(x1,x2), which can be used to implement max pooling and max-out" The mapping unavailable condition never occurs.) determining, by the processor, association of a weight parameter (θ) with each of the plurality of layers of the untrained neural network; ([p. 5] "Initialize weights randomly" [p. 4] "our normalization depends as well on the current weights w" [p. 5] "Normalization is introduced after every linear (conv or fully connected) layer") and evaluating, by the processor, a weight initialization technique for setting an initial value of respective weight parameters (θ) associated with one or more layers, out of the plurality of layers which are determined to have associated weight parameters (θ)([p. 5] "Initialize weights randomly" [p. 4] "our normalization depends as well on the current weights w" [p. 5] "Normalization is introduced after every linear (conv or fully connected) layer") by applying the derived mean-variance mapping function (g-layer) corresponding to the one or more layers, ([p. 4] "In the projecting initialization, approximate expectations E[Xi] and Var[Xi]") such that the initial value of mean (μOut) of output signals of the one or more layers is set as zero ([p. 2] "This choice indeed matches the expectations of X when Z has zero mean and unit variance") and the initial value of variance (vout) of output signals of the one or more layers is set as one, ([p. 2] "This choice indeed matches the expectations of X when Z has zero mean and unit variance") thereby eliminating the problem of exploding or vanishing output signals,([p. 2] "this projecting initialization approximately compensates accumulated biases and scaling of linear and non-linear transforms and efficiently reinitializes the network to a point where non-linearities are not saturated on average. It makes the network training in variant to a scale-bias preprocessing transformation of the input images and the initial scale of the random weights" Invariant to scale bias preprocessing interpreted as synonymous with eliminating the problem of exploding or vanishing signals) wherein the method of evaluating the weight initialization technique is employed to improve deep learning of neural network models and adapt to a plurality of different and unique neural network architectures([p. 7] "the proposed technique improves robustness to initialization points and achieves a lower training objective in many cases"). However, Shekhovtsov does not explicitly teach from one or more client devices, one or more input/output devices and one or more external resources via an integration interface configured with one or more Application Programming Interfaces (APIs) or a Graphical User Interface (GUI) accessible via a user module. Schilling, in the same field of endeavor, teaches from one or more client devices, one or more input/output devices and one or more external resources via an integration interface configured with one or more Application Programming Interfaces (APIs) or a Graphical User Interface (GUI) accessible via a user module ([p. 49] "The datasets used for the following experiments are arguably the most popular ones used for image classification tasks in the literature. Ordered by increasing difficulty to generalize, those are MNIST [39], SVHN [43], CIFAR10, and CIFAR100 [32]" [p. 89] "The software framework used for conducting experiments is Torch [6], a scientific computing platform with wide support for machine learning algorithms that is used and maintained by various technology companies such as Google DeepMind and Facebook AI Research. At its core, Torch supplies a flexible n-dimensional array with many useful routines such as indexing, slicing and transposing. Moreover, Torch features an extremely fast scripting language and is based on LuaJIT, a powerful just-in-time compiler. LuaJIT is implemented in the C programming language and thus provides a clean interface to the GPU using NVIDIA’s CUDA libraries. CUDA is a parallel computing platform and application programming interface that enables general purpose GPU processing. In this research, we make heavy use of both the CUDA Toolkit 7.5 and cuDNN v4, a deep neural network library that extends the toolkit with useful operations such as highly optimized convolutions. An exciting piece of trivia is that NVIDIA advertises cuDNN v4 with the following slogan: “Train neural networks up to 14x faster with batch normalization” [...] Torch’s API bears significant resemblance to the pseudocode used in the back propagation algorithm (algorithm 1) since each module provides a forward and backward function that implements the mathematical notation introduced in sections 2.9.1, 2.9.2, 2.9.3, and 3" [p. 90] "The machine used for the experiments features a 3.4GHz Intel Core i7-2600K with 16GB of DDR3 RAM, and a NVIDIA Tesla K40c GPU with 2880 CUDA cores clocked at 745MHz and 12GB of RAM. The system is running Ubuntu 14.04.1 LTS with GNU/Linux kernel 3.19.0-51" The machine in the experiment interpreted as a client device comprising multiple input/output devices (CPU/GPU/RAM) and obtaining neural network and training dataset using external resources (MNIST, SVHN, CIFAR10, CIFAR100) via integration interface (CUDA, Torch, cDNN, etc.) with one or more API and/or a GUI). Shekhovtsov as well as Schilling are directed towards mean, variance, and batch normalization in convolutional neural networks. Therefore, Shekhovtsov as well as Schilling are reasonably pertinent analogous art. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of Shekhovtsov with the teachings of Schilling by performing the method of Shekhovtsov on the same generic computing system as Schilling as applying the same insights regarding batch normalization. Schilling provides as additional motivation for combination ([Abstract] “Batch normalization is a recently popularized method for accelerating the training of deep feed-forward neural networks. Apart from speed improvements, the technique reportedly enables the use of higher learning rates, less careful parameter initialization, and saturating nonlinearities.”). This motivation for combination also applies to the remaining claims which depend on this combination. Regarding claim 2, the combination of Shekhovtsov and Schilling teaches The method as claimed in claim 1, wherein the derived mean-variance mapping functions (g-layer) mapped to the respective layers of the untrained neural network are stored in the mean-variance mapping table.(Shekhovtsov [p. 3] "the approximation can be applied layer-by-layer, propagating mean and variance. Let us consider a deep NN […] Propagate the moments until the first normalization layer" The implementation identifies the operation being processed and invokes/instantiates its predefined analytic moment rule. This is functionally selecting the corresponding g-layer from the organized collection of available mappings). Regarding claim 3, the combination of Shekhovtsov and Schilling teaches The method as claimed in claim 1, wherein the mean-variance mapping function (g-layer) associated with each of the respective layers of the untrained neural network is derived using data analytics based on any one of the following: a weight parameter associated with the respective layer, a type of said respective layer, an activation function associated with said respective layer or any combination thereof.(Shekhovtsov [p. 1] "The proposed estimation uses an analytic propagation of mean and variance of the training set through the network" [p. 5] "Initialize weights randomly" [p. 4] "our normalization depends as well on the current weights w" [p. 5] "Normalization is introduced after every linear (conv or fully connected) layer" This is essentially Shechovtsov's stated purpose. It analytically determines neural network activation moments through the network and expressly uses those statistics for initialization). Regarding claim 4, the combination of Shekhovtsov and Schilling teaches The method as claimed in claim 1, wherein the mean-variance mapping function (g-layer) corresponding to any layer of the untrained neural network having no associated weight parameter (θ) is representative of a function derived to map a mean and a variance of input signal of the any layer with a mean and a variance of output signal after propagation through said any layer.(Schilling [pp. 24-25 §2.9.3] "It is common practice to periodically insert a pooling layer in between successive convolutional layers [...] Since the pooling layer does not have any learnable parameters, the backward pass is merely an upsampling operation of the upstream derivatives. In case of the max-pooling operation, it is common practice to keep track of the index of the maximum activation so that the gradient can be routed towards its origin during backpropagation" Schilling explicitly teaches that the pooling layers have no associated weight parameters.). Regarding claim 5, the combination of Shekhovtsov and Schilling teaches The method as claimed in claim 1, wherein the mean-variance mapping function (g-layer) corresponding to the one or more layer of the untrained neural network determined to have an associated weight parameter (θ) is representative of a function derived to map a mean and a variance of an input signal of the one or more layers and the weight parameter (θ) associated with the one or more layers with a mean and a variance of an output signal after propagation through the one or more layers.(Shekhovtsov [p. 1] "The proposed estimation uses an analytic propagation of mean and variance of the training set through the network" [p. 5] "Initialize weights randomly" [p. 4] "our normalization depends as well on the current weights w" [p. 5] "Normalization is introduced after every linear (conv or fully connected) layer" [p. 3] "the approximation can be applied layer-by-layer, propagating mean and variance. Let us consider a deep NN" This is essentially Shechovtsov's stated purpose. It analytically determines neural network activation moments through the network and expressly uses those statistics for initialization). Regarding claim 6, the combination of Shekhovtsov and Schilling teaches The method as claimed in claim 1, wherein each layer of the untrained neural network is connected with its next layer or any subsequent layer such that an output of any layer (L) of the neural network is an input of its next layer (L+1) or a subsequent layer connected directly to said any layer (L), (Shekhovtsov [p. 7] "Noisy-CIFAR-10 We tested with a CNN network with ReLU and the following conv layers: ksize =[3, 3, 3, 3, 3, 3, 3, 1, 1 ] stride=[1, 1, 2, 1, 1, 2, 1, 1, 1 ] depth =[96, 96, 96, 192, 192, 192, 192, 192, 10]") and a mean (μOut) and a variance (vout)of an output signal of said any layer (L) is a mean (μin) and a variance (vin) of an input signal of the next layer (L+1) or the subsequent layer.(Shekhovtsov [p. 1] "The proposed estimation uses an analytic propagation of mean and variance of the training set through the network" [p. 5] "Initialize weights randomly" [p. 4] "our normalization depends as well on the current weights w" [p. 5] "Normalization is introduced after every linear (conv or fully connected) layer" [p. 3] "the approximation can be applied layer-by-layer, propagating mean and variance. Let us consider a deep NN" This is essentially Shechovtsov's stated purpose. It analytically determines neural network activation moments through the network and expressly uses those statistics for initialization). Regarding claim 8, the combination of Shekhovtsov and Schilling teaches The method as claimed in claim 1, wherein the evaluation of the weight initialization technique for setting the initial value of the respective weight parameter (θ) associated with the one or more layers determined to have associated weight parameter (θ) comprises: a. computing and incorporating a mean (μin) and a variance (vin) of an input signal of a layer (L) out of the each layer determined to have associated weight parameter (θ) in the derived mean-variance mapping function (g-layer) corresponding to the layer (L), wherein the derived mean-variance mapping function maps the mean (μin) and the variance (vin) and the weight parameter (θ) associated with said layer (L) with a mean (μOut) and a variance (vout) of an output signal after propagation through said layer (L);(Shekhovtsov [p. 7] "Noisy-CIFAR-10 We tested with a CNN network with ReLU and the following conv layers: ksize =[3, 3, 3, 3, 3, 3, 3, 1, 1 ] stride=[1, 1, 2, 1, 1, 2, 1, 1, 1 ] depth =[96, 96, 96, 192, 192, 192, 192, 192, 10]" [p. 1] "The proposed estimation uses an analytic propagation of mean and variance of the training set through the network" Convolution layers have weight parameters each feature map is both an output signal of one layer and an input signal to a subsequent layer.) b. evaluating a weight distribution for the weight parameter (θ) associated with the layer (L) and ascertaining a sampling range for said weight parameter (θ); c. selecting the initial value of the weight parameter (θ) from the ascertained sampling range such that the mean (μOut) and variance (vout) of the output signal of said layer (L) is zero and one, respectively on incorporating the selected initial value in said derived mean-variance mapping function (g-layer); and d. repeating a-c for the each layer determined to have associated weight parameter (θ).(Schilling [p. 49 §3.5] "Batch normalization alleviates some of the hopelessness by assuming that if x is drawn from a unit normal distribution, all the downstream layers will be normally distributed because the intermediate transformations are linear" [p. 52] "The goal of batch normalization is to achieve a stable distribution of activation values throughout training" [p. 56 §5.4] "Weight Initialization [...] To assess how the scale of the weight initialization affects the training behavior, we train both the vanilla and batch normalized network with initial values drawn from distributions with different variances"). Regarding claim 9, the combination of Shekhovtsov and Schilling teaches The method as claimed in claim 8, wherein the mean (μin) and the variance (vin) of the input signal of the layer (L) is same as a mean (μOut) and a variance (vout) of an output signal of any preceding layer of the neural network directly providing input to said layer (L).(Shekhovtsov [p. 7] "Noisy-CIFAR-10 We tested with a CNN network with ReLU and the following conv layers: ksize =[3, 3, 3, 3, 3, 3, 3, 1, 1 ] stride=[1, 1, 2, 1, 1, 2, 1, 1, 1 ] depth =[96, 96, 96, 192, 192, 192, 192, 192, 10]" [p. 1] "The proposed estimation uses an analytic propagation of mean and variance of the training set through the network" Convolution layers have weight parameters each feature map is both an output signal of one layer and an input signal to a subsequent layer.). Regarding claim 10, the combination of Shekhovtsov and Schilling teaches The method as claimed in claim 9, wherein the mean (μOut) and the variance (vout) of the output signal of the any preceding layer is computed using a mean-variance mapping function (g-layer) corresponding to said any preceding layer if no weight parameter (θ) is associated with said any preceding layer; or the mean (μOut) and the variance (vout) of the output signal of the any preceding layer is zero and one respectively, if said any preceding layer has an associated weight parameter (θ).(Shekhovtsov [p. 2] "This choice indeed matches the expectations of X when Z has zero mean and unit variance"). Regarding claim 11, the combination of Shekhovtsov and Schilling teaches The method as claimed in claim 1, wherein a mean and a variance of an output signal of any layer of the untrained neural network having no associated weight parameter (θ) is computed using the derived mean-variance mapping function (g-layer) corresponding to said any layer by: computing a mean and a variance of the input signal of said any layer, wherein the mean and the variance of the input signal of said any layer is same as a mean and a variance of an output signal of any preceding layer of the untrained neural network directly providing input to said any layer; and(Schilling [p. 78 §7.2] "in the convolutional case, we normalize each feature map over the current mini-batch and learn the scale and shift parameters per feature map, rather than per activation" [p. 38 §3] "we normalize the distribution of each input feature in each layer across each mini-batch to have zero mean and a standard deviation of one" See also Table 6.3 which shows that each convolutional layer is followed by a pooling layer which has no associated weight parameter ([pp. 24-25 §2.9.3] "the pooling layer does not have any learnable parameters").) incorporating the computed mean and the variance of the input signal of said any layer in the derived mean-variance mapping function (g-layer) to compute the mean and variance of the output signal after propagating through said any layer.(Schilling See also Table 6.3 which shows that each convolutional layer is followed by a pooling layer which has no associated weight parameter. Convolutional layer having output feature map with zero mean and unit variance subsequent to max pooling interpreted as synonymous with computing the mean and variance of the output signal after propagating through said any layer.). Regarding claim 12, the combination of Shekhovtsov and Schilling teaches The method as claimed in claim 8, wherein the mean and the variance of the input signal of the layer (L) is computed by aggregation of input data if the layer (L) is an input layer.(Shekhovtsov [p. 1] "The proposed estimation uses an analytic propagation of mean and variance of the training set through the network" [p. 4] "to estimate quantities such as mean activations and even norms of gradients" [p. 6] "Datasets We used MNIST and CIFAR10"). Regarding claims 13, 15-19, and 21-24, claims 13, 15-19, and 21-24 are directed towards a system for performing the method of claims 1-6, 8, and 10-12, respectively. Therefore, the rejections applied to claims 1-6, 8, and 10-12 also apply to claims 13, 15-19, and 21-24. Regarding claim 25, claim 25 is directed towards a computer program product for performing the method of claim 1. Therefore, the rejection applied to claim 1 also applies to claim 25. Claims 7 and 20 are rejected under U.S.C. §103 as being unpatentable over the combination of Shekhovtsov and Schilling and in further view of Amiri (US 20230022401 A1). Regarding claim 7, the combination of Shekhovtsov and Schilling teaches The method as claimed in claim 1, wherein the step of determining association of the weight parameter (θ) with each of the plurality of layers of the untrained neural network comprises: identifying a type of the layer based on analysis of the layer; (Schilling [pp. 24-25 §2.9.3] "It is common practice to periodically insert a pooling layer in between successive convolutional layers [...] Since the pooling layer does not have any learnable parameters, the backward pass is merely an upsampling operation of the upstream derivatives. In case of the max-pooling operation, it is common practice to keep track of the index of the maximum activation so that the gradient can be routed towards its origin during backpropagation" FIG. 27 shows a neural network architecture with determined types having determined and known weight parameter associations. §2.9.1 and 2.9.2 describe layers having weight parameters. 2.9.3 describes layers not having weight parameters.). However, the combination of Shekhovtsov and Schilling doesn't explicitly teach and determining association of weight parameter (θ) with the layer based on the identified type of the layer by accessing a predefined database, said predefined database comprising information associated with types of layers having weights and not having weights. Amiri, in the same field of endeavor, teaches and determining association of weight parameter (θ) with the layer based on the identified type of the layer by accessing a predefined database, said predefined database comprising information associated with types of layers having weights and not having weights.([¶0040] "corresponding parameters and weights are stored in the model database 214 to be used in the prediction phase. When the best forecasters (denoted by ft in FIG. 2) are determined, time series data and the forecaster are provided to TS classifier 210 for training the TS classifier 210. The best forecaster need not be a different forecaster and may be the same forecaster with different parameters or hyper-parameters e.g., parameters may be the trainable weights of a DNN, or trainable parameters of a linear regression model, while the hyper-parameters may be describing the number of hidden layers, type of activation functions"). The combination of Shekhovtsov and Schilling as well as Amiri are directed towards neural network processing. Therefore, the combination of Shekhovtsov and Schilling as well as Amiri are reasonably pertinent analogous art. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of the combination of Shekhovtsov and Schilling with the teachings of Amiri by using a database comprising information associated with types of layers having weights and not having weights. Amiri teaches that the database is used to determine a best forecaster for the particular application and further provides at motivation for combination ([¶0020] “The forecaster of various embodiments of the disclosure offer, in some cases, over 40% improvement over conventional forecasters because the best forecaster for the time series is selected based on the output of a time series classifier.”). Regarding claim 20, claim 20 is directed towards a system for performing the method of claim 7. Therefore, the rejection applied to claim 7 also applies to claim 20. Claim 14 is rejected under U.S.C. §103 as being unpatentable over the combination of Shekhovtsov and Schilling and Lidman (US20200379923A1). Regarding claim 14, the combination of Shekhovtsov and Schilling teaches The system as claimed in claim 13. However, the combination of Shekhovtsov and Schilling doesn't explicitly teach, wherein the weight initialization engine comprises an interface unit executed by the processor, said interface unit configured to facilitate user interaction, and receive the neural network model. Lidman, in the same field of endeavor, teaches The system as claimed in claim 13, wherein the weight initialization engine comprises an interface unit executed by the processor, said interface unit configured to facilitate user interaction, and receive the neural network model.([¶0024] "the user device 110 may receive the neural network models 102 (including updated neural network models) from the deep learning environment 101 at any time."). The combination of Shekhovtsov and Schilling as well as Lidman are directed towards neural network processing. Therefore, the combination of Shekhovtsov and Schilling as well as Lidman are reasonably pertinent analogous art. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of the combination of Shekhovtsov and Schilling with the teachings of Lidman by receiving the neural network model through a user interface. Lidman provides as additional motivation for combination ([¶0021] “the user device 110 may use the neural network models 102 to generate inferences about a user and/or content the user is viewing or listening to”). Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SIDNEY VINCENT BOSTWICK whose telephone number is (571)272-4720. The examiner can normally be reached M-F 7:30am-5:00pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Miranda Huang can be reached on (571)270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SIDNEY VINCENT BOSTWICK/Examiner, Art Unit 2124
Read full office action

Prosecution Timeline

Show 1 earlier event
Jul 10, 2025
Non-Final Rejection mailed — §103, §112
Oct 06, 2025
Response Filed
Oct 28, 2025
Final Rejection mailed — §103, §112
Dec 30, 2025
Request for Continued Examination
Jan 16, 2026
Response after Non-Final Action
Mar 03, 2026
Non-Final Rejection mailed — §103, §112
Jun 03, 2026
Response Filed
Aug 19, 2026
Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12699874
Leveraging Redundancy in Attention with Reuse Transformers
3y 10m to grant Granted Aug 04, 2026
Patent 12675673
NEURAL NETWORK PROCESSING DEVICE, METHOD, AND COMPUTER-READABLE RECORDING MEDIUM
3y 7m to grant Granted Jul 07, 2026
Patent 12645914
INSTRUCTION PRUNING FOR NEURAL NETWORKS
3y 6m to grant Granted Jun 02, 2026
Patent 12626139
SECRET SOFTMAX FUNCTION CALCULATION SYSTEM, SECRET SOFTMAX FUNCTION CALCULATION APPARATUS, SECRET SOFTMAX FUNCTION CALCULATION METHOD, SECRET NEURAL NETWORK CALCULATION SYSTEM, SECRET NEURAL NETWORK LEARNING SYSTEM, AND PROGRAM
4y 3m to grant Granted May 12, 2026
Patent 12619815
Magnitude Invariant Multimodal Agent for Efficient Image-Text Interface Automation
1y 6m to grant Granted May 05, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
51%
Grant Probability
86%
With Interview (+35.1%)
4y 5m (~1m remaining)
Median Time to Grant
High
PTA Risk
Based on 152 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month