Prosecution Insights
Last updated: October 02, 2026
Application No. 18/295,791

NEURAL NETWORK LAYER FOR NON-LINEAR NORMALIZATION

Final Rejection §101§103
Filed
Apr 04, 2023
Priority
May 13, 2022 — EU 22 17 3331.4
Examiner
RAMESH, TIRUMALE K
Art Unit
2121
Tech Center
2100 — Computer Architecture & Software
Assignee
Robert Bosch GmbH
OA Round
2 (Final)
26%
Grant Probability
At Risk
3-4
OA Rounds
1y 2m
Est. Remaining
50%
With Interview

Examiner Intelligence

Grants only 26% of cases
26%
Career Allowance Rate
13 granted / 49 resolved
-28.5% vs TC avg
Strong +24% interview lift
Without
With
+23.7%
Interview Lift
resolved cases with interview
Typical timeline
4y 9m
Avg Prosecution
19 currently pending
Career history
85
Total Applications
across all art units

Statute-Specific Performance

§101
26.8%
-13.2% vs TC avg
§103
63.9%
+23.9% vs TC avg
§102
4.2%
-35.8% vs TC avg
§112
4.6%
-35.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 49 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . 1234 Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefore, subject to the conditions and requirements of this title. Regarding 101 Rejections Claims 1-8 were rejected for “software per se” as the claim did not recite any computer hardware. The applicant has added the limitations “processor” to the claim 1. Although, the memory or storage is not explicitly specified, claim 1 satisfy the machine category. As a result, the examiner REMOVES the 101 rejections for claims 1-8. The claim 13 is a new dependent claim on claim 1. Claim Interpretation This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation is: “ A training system configured to train a machine learning system including a plurality of layers, the training system configured to: provide an output signal based on an input signal by forwarding the input signal through the plurality of layers of the machine learning system, at least one of the layers of the plurality of layers being configured to receive a layer input, which is based on the input signal, and provide a layer output based on which the output signal is determined, the layer being configured to determine the layer output using a non-linear normalization of the layer input” in claim 11. Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Regarding 103 Rejections The applicants argue on Pages 7-9, the following: Page 7: Argument #1-Sadri Reference Reference “Sadri “does not a take a group of values and compute the empirical percentile and map these to the target distribution. The limitations call for group transformed based on where it falls in the empirical distribution. Examiner’s Response The examiner “disagrees” with the argument. In probability and machine learning, extracting a random variable p from a noise probability distribution means sampling values from a predefined probability distribution (often Gaussian or uniform) to represent “noise” in a system. This is a common step in modeling real-world data, where noise is treated as a random fluctuation. As recited in Sadri [0035]” lower bound of a support set of the noise probability distribution may be equal to p l and an upper bound of the support set may be equal to p u   where 0≤ p l ≤ p u ≤1, ( p l + p u )/2 ≤ μ and μ is a mean of the noise probability distribution”, the noise in the lower bound of a support set” or “noise in the upper bound” means that the actual measured or estimated bounds p u and p l   of the support may differ slightly from the theoretical or ideal bounds due to measurement error, sampling variability, or other sources of uncertainty. Reference “Sadri” recites in [0005]” In an exemplary embodiment, the first iterative process may include extracting an i t h   denoised training image from an output of the FCN, generating a plurality of updated weights, and replacing the plurality of initial weights with the plurality of updated weights”. Perhaps known to a POSITA that in denoising tasks, the model learns the statistical structure of clean data so it can predict and restore the original image when given a noisy version where the clean image is the “ground truth” version with no noise. Further, reference “Sadri” recites in [0007] “ in an exemplary embodiment, the ( i + 1 ) t h plurality of normalized training feature maps may be generated by applying a batch normalization process on the (l+1).sup.th plurality of filtered training feature maps”. Perhaps it is known to POSITA that a batch normalization process can provide non-linear transformation during training due to how it computes the normalization parameters. As mapped in “Sadri” [0035]” when μ is close to upper bound p u a ratio of number of images that are corrupted by higher levels of impulsive noise to a number of plurality of training images 202 may increase” does define a group of values with lower bound x and upper bound y, compute empirical percentiles from the noise distribution, and map them to the target distribution’s quantiles. This preserves the percentile structure while transforming the data to the target distribution’s support. Further, reference “Sadri” recites in [0042]” generating the plurality of initial weights may include generating a plurality of random variables from a predetermined probability distribution. In an exemplary embodiment, the predetermined probability distribution may be determined by a designer of FCN 200 according to a required range of each of the plurality of initial weights. In an exemplary embodiment, the predetermined probability distribution may be selected from Gaussian or uniform probability distributions”. Perhaps known to a POSITA that a uniform probability distribution whether discrete or continuous assigns equal probability to all outcomes within a defined range or set. In the continuous case, this means the probability density is constant over the interval [a, b], so every value in that range is equally likely. In the discrete case, each possible outcome has the same probability mass. Reference “Sadri” recites in [0048] “ in an exemplary embodiment, all elements of the set of ( i + 1 ) t h   plurality of filtered training feature maps 218 may be scaled and shifted by a scale and a shift variable, which may be learned during training process. Therefore, in an exemplary embodiment, all elements of ( i + 1 ) t h   plurality of normalized training feature maps 224 may follow a normal distribution, which may considerably reduce a required time for training FCN 200” and further recites in [0049]” In an exemplary embodiment, implementing l.sup.th non-linear activation function 228 may include implementing one of a rectified linear unit (ReLU) function or an exponential linear unit (ELU) function. In an exemplary embodiment, implementing l.sup.th non-linear activation function 228 may include implementing other types of non-linear activation functions such as leaky ReLU, scaled ELU, parametric ReLU, etc”. Perhaps as known to a POSITA that non-linear activation functions that include ReLU, Scaled ELU, Parametric ReLU, and others does represent a non-linear transformation and that the normalization followed by a non-linear transformation does provide a non-linear normalization based on the statutory and procedural framework for prior art and obviousness. The combined process of obtaining non-linear normalization, the data maybe rescaled to a standard range, but the mapping from original to final values is not linear (end to end), but it is non-linear . This can be useful for introducing non-linearity, improving model performance, or matching perceptual scales. Normalization [0048] -> Non-Linear Activation + RELU [0048] -> Provided Non-Linear Normalization Perhaps it is known to POSITA that a target distribution in statistics or machine learning refers to the desired probability distribution that a model or process should follow. This could be a theoretical distribution (e.g., normal, uniform) or an empirical distribution derived from data. Page 8: Argument #2- Sadri Reference The reference “Sadri’s” probability distribution operates on training image noise levels and is a data augmentation mechanism. It does not operate on any group of values of the layer input. Examiner’s Response The examiner “disagrees” with the argument over noise. Perhaps it is known to a POSITA that in probability and machine learning, extracting a random variable p from a noise probability distribution means sampling values from a predefined probability distribution (often Gaussian or uniform) to represent “noise” in a system. This is a common step in modeling real-world data, where noise is treated as a random fluctuation. Extracting a random variable p from a noise distribution means drawing noise samples from a defined probability distribution. Using these samples as inputs (or as part of the data) and applying a model (e.g., a generator) produces a training array of data points used to train a model, often to simulate realistic data or improve generalization. Reference “Sadri” also recites in [0006]” In an exemplary embodiment, a support set of the uniform probability distribution may be equal to an intensity range of the n.sup.th original image”. Page 9: Argument #3-Lin Reference The reference “Lin” does not disclose “where the layer is configured to determine the layer output using a non-linear normalization of the layer input”. Lin teaches a non-linear activation function that operates on a single node’s output and takes the weighted sum of inputs at each individual node. The non-linear activation and non-linear normalization are not the same and are different concepts. Examiner’s Response The examiner “disagrees” with the argument over noise. Reference Lin recites in [0042] “ A normalization operation can bring the numerical data to a common or balanced scale without distorting the data” and recites in [0007] “ By determining the architecture of the normalization-activation layer as described in this specification, the system can identify architectures that, particularly on image understanding or generation tasks, outperform conventional human-designed architectures, e.g., batch norm + ReLU. In other words, the inclusion of the identified architecture in a neural network can result in the neural network being able to be trained to convergence quickly by improving the stability of the training process and results in the trained neural network having improved performance on the target ask relative to conventional human-designed or searched architectures” and further recites in [0055] The machine-learning model processes the input to generate an output. An artificial neural network includes an input layer that consists of values in a data point. The next layer is called a hidden layer, and nodes at the hidden layer each receive one or more of the input values. Each node contains parameters (e.g., weights) to apply to the input values. Each node therefore essentially inputs the input values into a multivariate function (e.g., a non-linear mathematical transformation) to produce an output value” Based on the same context as detailed for Sadri reference above, the normalization can bring numerical data to a common or balanced scale, and in some cases, it can be combined with non-linear transformations to produce a non-linear normalization effect. Using the Batch Norm+ReLu as taught in [0007], the Batch Normalization layer followed by a ReLU layer does represent a form of non-linear normalization. Lin further supports in [0023] “ As discussed above, typical CNNs consist of convolutional (Conv) layer, pooling layer, non-linear layer (e.g. ReLU), normalization layer (e.g. local response normalization (LRN)) and fully connected (FC) layer, etc. The convolutional layer generally includes a set of trainable kernels, which extract local features from a small spatial region but cross-depth volume of the input tensor”. In CONCLUSION, the references “Sadri” and “Lin” supports the rejections of claims 1, and 9-12. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 4-13 are rejected under 35 U.S.C. 103 as being unpatentable over Chun-Fu CHEN et al. (hereinafter Chen) US 2019/0122113 A1, in view of Jiu-Che Lin et.al. (hereinafter Lin) US 2023/0342016 A1, in view of Seyedeh Sahar Sadrizadeh et.al. (hereinafter Sadri) US 2021/0209735 A1. Regarding claim 1: (Currently Amended) Chen discloses: - A computer-implemented machine learning system, the machine learning system comprising: a plurality of layers, the machine learning system being configured to provide an output signal based on an input signal by forwarding the input signal through the plurality of layers of the machine learning system, wherein at least one of the layers of the plurality of layers is configured to receive a layer input, which is based on the input signal, and to provide a layer output based on which the output signal is determined, [0021]: FIG. 2 illustrates a convolutional neural network, according to one embodiment described herein. PNG media_image1.png 416 765 media_image1.png Greyscale As shown, the CNN 200 includes an input layer 210, a convolutional layer 215, a subsampling layer 220, a convolutional layer 225, subsampling layer 230, fully connected layers 235 and 240, and an output layer 245. The input layer 210 in the depicted embodiment is configured to accept a 32×32 pixel image. The convolutional layer 215 generates 6 28×28 feature maps from the input layer, and so on. While a particular CNN 200 is depicted, more generally a CNN is composed of one or more convolutional layers, frequently with a subsampling step, and then one or more fully connected layers. [0037]: Generally, performing ISBP from 3-way tensor to 3-way tensor tends to be more complicated than the above two cases, as the operations of the forward propagation between the input and output [0027]: Accordingly, DCNN optimization component 140 can perform Importance Score Back Propagation (ISBP) for optimizing a CNN. [0041]: If B P c o n v f n i ,   j ≠ 1 , this can indicate that the i.sup.th position in the output layer comes from a convolution operation involving the j t h   position in the input layer. ( BRI: The layer output based on which the final output signal of a neural network is determined is the output layer, which is the final, topmost layer in the network architecture) [0023]: As discussed above, typical CNNs consist of convolutional (Conv) layer, pooling layer, non-linear layer (e.g. ReLU), normalization layer (e.g. local response normalization (LRN)) and fully connected (FC) layer, etc. The convolutional layer generally includes a set of trainable kernels, which extract local features from a small spatial region but cross-depth volume of the input tensor. Each kernel can be trained as a feature extractor for some specific visual features, such as an edge or a color in the first layer (BRI: normalization provide an output that is a nonlinear transformation of their input) Chen does not explicitly disclose: - wherein the layer is configured to determine the layer output using a non-linear normalization of the layer input. - wherein for determining the layer output, the layer is configured to normalize at least one group of values of the layer input, wherein the group includes all values of the layer input or a subset of the values of the layer input. However, Lin discloses: - wherein the layer is configured to determine the layer output using a non-linear normalization of the layer input. [0055]: The machine-learning model processes the input to generate an output. An artificial neural network includes an input layer that consists of values in a data point. The next layer is called a hidden layer, and nodes at the hidden layer each receive one or more of the input values. Each node contains parameters (e.g., weights) to apply to the input values. Each node therefore essentially inputs the input values into a multivariate function (e.g., a non-linear mathematical transformation) to produce an output value. A next layer can be another hidden layer or an output layer. In either case, the nodes at the next layer receive the output values from the nodes at the previous layer, and each node applies weights to those values and then generates its own output value. wherein for determining the layer output, the layer is configured to normalize at least one group of values of the layer input, wherein the group includes all values of the layer input or a subset of the values of the layer input; [0088]: At operation 412, processing logic performs one or more preprocessing operations on the input data. In some embodiments, the preprocessing operations can include a smoothing operation, a normalization operation, a dimensions reduction operations, a sort features operation, or any other operation configured to prepare data for training a machine-learning model. In [0055]: The machine-learning model processes the input to generate an output. An artificial neural network includes an input layer that consists of values in a data point. The next layer is called a hidden layer, and nodes at the hidden layer each receive one or more of the input values. Each node contains parameters (e.g., weights) to apply to the input values. Each node therefore essentially inputs the input values into a multivariate function (e.g., a non-linear mathematical transformation) to produce an output value It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Chen and Lin. Chen a machine learning system to receive layer inputs and determine the layer output and teaches non-linear normalization of the layer input within the context of the CNN consisting of a “normalization layer”. Lin teaches a non-linear normalization of the layer input and determining layer output using sorting, smoothing within the context of a probability distribution. One ordinary skill would be motivated to combine Chen and Lin that can provide optimized model parameters and improvement to the model in terms of its accuracy (Lin [0059]). Chen and Lin do not explicitly disclose: - wherein the non-linear normalization includes mapping empirical percentiles of values from the group to percentiles of a predefined probability distribution. However, Sadri discloses: - wherein the non-linear normalization includes mapping empirical percentiles of values from the group to percentiles of a predefined probability distribution. [0035]: For further detail with respect to step 112, in an exemplary embodiment, a lower bound of a support set of the noise probability distribution may be equal to p l and an upper bound of the support set may be equal to p u   where 0≤ p l ≤ p u ≤1, PNG media_image2.png 68 432 media_image2.png Greyscale where μ is the mean of the noise probability distribution an exemplary embodiment, when μ is close to upper bound   p u   , a ratio of number of images that are corrupted by higher levels of impulsive noise to a number of plurality of training images 202 may increase It would be obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Chen, Lin and Sadri. Chen a machine learning system to receive layer inputs and determine the layer output using a non-linear normalization of the layer input. Lin teaches determining layer output using sorting, smoothing within the context of a probability distribution. Sadri teaches mapping empirical percentiles of values from the group to percentiles of a predefined probability distribution. One of ordinary skill would be motivated to combine Chen, Lin and Sadri that can provide minimized loss function using gradient descent (Sadri ([0052]) Regarding claim 4: (Currently Amended) Chen and Lin do not explicitly disclose: - wherein the predefined probability distribution is a standard normal distribution. However, Sadri discloses: - wherein the predefined probability distribution is a standard normal distribution. [0049]: In an exemplary embodiment, step 140 may include generating (l+1).sup.th plurality of training feature maps 216. [0049]: In an exemplary embodiment, implementing l.sup.th non-linear activation function 228 may include implementing one of a rectified linear unit (ReLU) function or an exponential linear unit (ELU) function. In an exemplary embodiment, implementing l.sup.th non-linear activation function 228 may include implementing other types of non-linear activation functions such as leaky ReLU, scaled ELU, parametric ReLU, etc. In [0035]: In an exemplary embodiment, random variable p may be generated from a truncated Gaussian probability distribution defined within ( p l , p u ) 202. As a result, in an exemplary embodiment, a minimum level of impulsive noise in plurality of training images 202 may be equal to   p   l . In contrast, in an exemplary embodiment, a maximum level of impulsive noise in plurality of training images 202 may be equal to p u . It would be obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Chen, Lin and Sadri. Chen a machine learning system to receive layer inputs and determine the layer output using a non-linear normalization of the layer input. Lin teaches determining layer output using sorting, smoothing within the context of a probability distribution. Sadri teaches mapping empirical percentiles of values from the group to percentiles of a predefined probability distribution. One of ordinary skill would be motivated to combine Chen, Lin and Sadri that can provide minimized loss function using gradient descent (Sadri ([0052]) Regarding claim 5: (Currently Amended) Chen does not explicitly disclose: - wherein to determine the layer output, the layer is configured to: receive a group of values of the layer input; - sort the received values; - compute percentile values for each position of the sorted values; - compute interpolation targets using a quantile function of the predefined probability distribution; - determine a function characterizing a linear interpolation of the sorted values and the interpolation targets. However, Lin discloses: - wherein to determine the layer output, the layer is configured to: receive a group of values of the layer input; [0050]: Deep learning is a class of machine learning algorithms that use a cascade of multiple layers of nonlinear processing units for feature extraction and transformation. Each successive layer uses the output from the previous layer as input. Deep neural networks can learn in a supervised (e.g., classification) and/or unsupervised (e.g., pattern analysis) manner. [0020]: generate a virtual knob for each feature that is used to train the machine learning model. Each virtual knob can be capable of adjusting the representative value of the corresponding feature, thus adjusting the output generated by the machine-learning model. - sort the received values; [0041]: Data interaction tool 154 can perform one or more preprocessing operations on the received input data. In some embodiments, the preprocessing operations can include a smoothing operation, a normalization operation, a dimensions reduction operations, a sort features operation, or any other operation configured to prepare data for training a machine-learning model. In [0041]: the sort features operation can include ranking the selected feature based on an order of importance. - compute percentile values for each position of the sorted values; [0091]: determine confidence interval for each virtual knob. A confidence interval displays the probability that a parameter (e.g., virtual knob value) will fall between a pair of values around a mean [0091]: in one example, processing logic can determine the confidence level of the intercept value using Formula 4, expressed below. PNG media_image3.png 30 442 media_image3.png Greyscale [0092] where: C I 0.95 is the confidence interval with a 95% confidence level [0093]: β 0   ^ is the estimator of a parameter (e.g., a virtual knob value) In [0094]: t 0.95 is the t-statistic (e.g., ratio of the departure of the estimated value of a parameter from its hypothesized value to its standard error) [0095] SE is the standard error of the estimator In [0096] : k i ,       0.95 is the i-th knob value with 95% confidence level In [0097]: w i   is the i-th knob weighting. - compute interpolation targets using a quantile function of the predefined probability distribution; [0091]: In some embodiments, processing logic can determine confidence interval for each virtual knob. A confidence interval displays the probability that a parameter (e.g., virtual knob value) will fall between a pair of values around a mean. Confidence intervals measure the degree of uncertainty or certainty in a sampling method. In one embodiment, to determine a confidence interval for a virtual knob, processing logic can first determine a confidence interval of the intercept value (e.g., Po) - determine a function characterizing a linear interpolation of the sorted values and the interpolation targets; [0090]: Each virtual knob can be capable of adjusting the representative value of the corresponding feature, thus adjusting the machine-learning model. In some embodiments, to generate a virtual knob, processing logic can use a transform function (or any other applicable function) to modify one or more values representing each feature in the machine-learning model. [0042]: The machine-learning model can include a representative function for each selected feature. In an illustrative example, machine-learning model 190 can be trained using linear regression, and expressed as seen in Formula 1 below, where the x value(s) represent each selected feature, the c value(s) represent corresponding coefficients, and the P 0 value represents the intercept value, and y represents the label value: PNG media_image4.png 22 462 media_image4.png Greyscale - determine the layer output by processing the received values with the function. [0042]: The machine-learning model can include a representative function for each selected feature. In an illustrative example, machine-learning model 190 can be trained using linear regression, and expressed as seen in Formula 1 below, where the x value(s) represent each selected feature, the c value(s) represent corresponding coefficients, and the P 0 value represents the intercept value, and y represents the label value: PNG media_image4.png 22 462 media_image4.png Greyscale It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Chen and Lin. Chen a machine learning system to receive layer inputs and determine the layer output using a non-linear normalization of the layer input. Lin teaches determining layer output using sorting, smoothing within the context of a probability distribution. One ordinary skill would be motivated to combine Chen and Lin that can provide optimized model parameters and improvement to the model in terms of its accuracy (Lin [0059]). Regarding claim 6: (Original) Chen does not explicitly disclose: - wherein to determine the layer output, before determining the function, the layer is configured to smooth the sorted values using a smoothing operation. However, Lin discloses: - wherein to determine the layer output, before determining the function, the layer is configured to smooth the sorted values using a smoothing operation. [0088]: At operation 412, processing logic performs one or more preprocessing operations on the input data. In some embodiments, the preprocessing operations can include a smoothing operation, a normalization operation, a dimensions reduction operations, a sort features operation, or any other operation configured to prepare data for training a machine-learning model. Regarding claim 7: (Original) Chen does not explicitly disclose: - wherein, to determine the layer output, the layer is further configured to scale and/or shift the values obtained after processing the received values with the function. However, Lin discloses: - wherein, to determine the layer output, the layer is further configured to scale and/or shift the values obtained after processing the received values with the function. [0041]: A normalization operation can bring the numerical data to a common or balanced scale without distorting the data. In [0045]: The scaling constant can be configured to increase or decrease the adjustment factor of the virtual knob. This allows for calibrating the virtual knobs without needing to retrain the machine-learning model. The scaling constant can be generated manually (e.g., user input), or automatically (e.g., based on a predefined value associated with a particular feature) using, for example, a data table, optimization tool 160, etc. It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Chen and Lin. Chen a machine learning system to receive layer inputs and determine the layer output using a non-linear normalization of the layer input. Lin teaches determining layer output using sorting, smoothing within the context of a probability distribution. One of ordinary skill would be motivated to combine Chen and Lin that can provide optimized model parameters and improvement to the model in terms of its accuracy (Lin [0059]). Regarding claim 8: (Original) Chen does not explicitly disclose: - wherein the input signal characterizes a signal obtained from a sensor. However, Lin discloses: - wherein the input signal characterizes a signal obtained from a sensor. [0061]: running trained machine-learning model 190 on the current sensor data input to obtain one or more outputs. [0057]: After one or more rounds of training, processing logic can determine whether a stopping criterion has been met. A stopping criterion can be a target level of accuracy, a target number of processed images from the training dataset It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Chen and Lin. Chen a machine learning system to receive layer inputs and determine the layer output using a non-linear normalization of the layer input. Lin teaches determining layer output using sorting, smoothing within the context of a probability distribution. One ordinary skill would be motivated to combine Chen and Lin that can provide optimized model parameters and improvement to the model in terms of its accuracy (Lin [0059]). Regarding claim 9: (Currently Amended) Chen discloses: - A computer-implemented method for training a machine learning system, the machine learning system including a plurality of layers, the method comprising the following: providing an output signal based on an input signal by forwarding the input signal through the plurality of layers of the machine learning system, at least one of the layers of the plurality of layers receiving a layer input, which is based on the input signal, and providing a layer output based on which the output signal is determined, [0021]: FIG. 2 illustrates a convolutional neural network, according to one embodiment described herein. PNG media_image1.png 416 765 media_image1.png Greyscale As shown, the CNN 200 includes an input layer 210, a convolutional layer 215, a subsampling layer 220, a convolutional layer 225, subsampling layer 230, fully connected layers 235 and 240, and an output layer 245. The input layer 210 in the depicted embodiment is configured to accept a 32×32 pixel image. The convolutional layer 215 generates 6 28×28 feature maps from the input layer, and so on. While a particular CNN 200 is depicted, more generally a CNN is composed of one or more convolutional layers, frequently with a subsampling step, and then one or more fully connected layers. [0037]: Generally, performing ISBP from 3-way tensor to 3-way tensor tends to be more complicated than the above two cases, as the operations of the forward propagation between the input and output [0027]: Accordingly, DCNN optimization component 140 can perform Importance Score Back Propagation (ISBP) for optimizing a CNN. [0041]: If B P c o n v f n i ,   j ≠ 1 , this can indicate that the i.sup.th position in the output layer comes from a convolution operation involving the j t h   position in the input layer. [0023]: As discussed above, typical CNNs consist of convolutional (Conv) layer, pooling layer, non-linear layer (e.g. ReLU), normalization layer (e.g. local response normalization (LRN)) and fully connected (FC) layer, etc. The convolutional layer generally includes a set of trainable kernels, which extract local features from a small spatial region but cross-depth volume of the input tensor. Each kernel can be trained as a feature extractor for some specific visual features, such as an edge or a color in the first layer [0020] : The deep convolutional neural network (DCNN) optimization component 140 is generally configured to optimize the structure of the trained DCNN model 138. In order to achieve a balance between the predictive power and model redundancy of CNNs, the DCNN optimization component 140 can learn the importance of convolutional kernels and neurons in FC layers from feature selection perspective. The DCNN optimization component 140 can optimize a CNN by pruning less important kernels and neurons based on their importance scores. The DCNN optimization component 140 can further fine-tune the remaining kernels and neurons to achieve a minimum loss of accuracy in the optimized DCNN. Chen does not explicitly disclose: - the layer determining the layer output using a non-linear normalization of the layer input. wherein for determining the layer output, the layer is configured to normalize at least one group of values of the layer input, wherein the group includes all values of the layer input or a subset of the values of the layer input; However, Lin discloses: - wherein the layer is configured to determine the layer output using a non-linear normalization of the layer input. [0055]: The machine-learning model processes the input to generate an output. An artificial neural network includes an input layer that consists of values in a data point. The next layer is called a hidden layer, and nodes at the hidden layer each receive one or more of the input values. Each node contains parameters (e.g., weights) to apply to the input values. Each node therefore essentially inputs the input values into a multivariate function (e.g., a non-linear mathematical transformation) to produce an output value. A next layer can be another hidden layer or an output layer. In either case, the nodes at the next layer receive the output values from the nodes at the previous layer, and each node applies weights to those values and then generates its own output value. (BRI: each node (or neuron) in a hidden or output layer contains learnable parameters—specifically weights and a bias—that are applied to input values to produce an output) wherein for determining the layer output, the layer is configured to normalize at least one group of values of the layer input, wherein the group includes all values of the layer input or a subset of the values of the layer input; [0088]: At operation 412, processing logic performs one or more preprocessing operations on the input data. In some embodiments, the preprocessing operations can include a smoothing operation, a normalization operation, a dimensions reduction operations, a sort features operation, or any other operation configured to prepare data for training a machine-learning model. [0055]: The machine-learning model processes the input to generate an output. An artificial neural network includes an input layer that consists of values in a data point. The next layer is called a hidden layer, and nodes at the hidden layer each receive one or more of the input values. Each node contains parameters (e.g., weights) to apply to the input values. Each node therefore essentially inputs the input values into a multivariate function (e.g., a non-linear mathematical transformation) to produce an output value It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Chen and Lin. Chen a machine learning system to receive layer inputs and determine the layer output and teaches non-linear normalization of the layer input within the context of the CNN consisting of a “normalization layer”. Lin teaches a non-linear normalization of the layer input and determining layer output using sorting, smoothing within the context of a probability distribution. One ordinary skill would be motivated to combine Chen and Lin that can provide optimized model parameters and improvement to the model in terms of its accuracy (Lin [0059]). Chen and Lin do not explicitly disclose: - wherein the non-linear normalization includes mapping empirical percentiles of values from the group to percentiles of a predefined probability distribution. However, Sadri discloses: - wherein the non-linear normalization includes mapping empirical percentiles of values from the group to percentiles of a predefined probability distribution. [0035]: For further detail with respect to step 112, in an exemplary embodiment, a lower bound of a support set of the noise probability distribution may be equal to p l and an upper bound of the support set may be equal to p u   where 0≤ p l ≤ p u ≤1, PNG media_image2.png 68 432 media_image2.png Greyscale where μ is the mean of the noise probability distribution an exemplary embodiment, when μ is close to upper bound   p u   , a ratio of number of images that are corrupted by higher levels of impulsive noise to a number of plurality of training images 202 may increase It would be obvious to one of ordinary skills in the art before the effective filing date of the present application to combine Chen, Lin and Sadri. Chen a machine learning system to receive layer inputs and determine the layer output using a non-linear normalization of the layer input. Lin teaches determining layer output using sorting, smoothing within the context of a probability distribution. Sadri teaches mapping empirical percentiles of values from the group to percentiles of a predefined probability distribution. One of ordinary skill would be motivated to combine Chen, Lin and Sadri that can provide minimized loss function using gradient descent (Sadri ([0052]) Regarding claim 10: (Currently Amended) Chen discloses: - A computer-implemented method for determining an output signal based on an input signal, the method comprising: determining the output signal by providing the input signal to a machine learning system, the machine learning system including a plurality of layers, the machine learning system providing the output signal based on an input signal by forwarding the input signal through the plurality of layers of the machine learning system, at least one of the layers of the plurality of layers receiving a layer input, which is based on the input signal, and providing a layer output based on which the output signal is determined, [0021]: FIG. 2 illustrates a convolutional neural network, according to one embodiment described herein. PNG media_image1.png 416 765 media_image1.png Greyscale As shown, the CNN 200 includes an input layer 210, a convolutional layer 215, a subsampling layer 220, a convolutional layer 225, subsampling layer 230, fully connected layers 235 and 240, and an output layer 245. The input layer 210 in the depicted embodiment is configured to accept a 32×32 pixel image. The convolutional layer 215 generates 6 28×28 feature maps from the input layer, and so on. While a particular CNN 200 is depicted, more generally a CNN is composed of one or more convolutional layers, frequently with a subsampling step, and then one or more fully connected layers. [0037]: Generally, performing ISBP from 3-way tensor to 3-way tensor tends to be more complicated than the above two cases, as the operations of the forward propagation between the input and output [0027]: Accordingly, DCNN optimization component 140 can perform Importance Score Back Propagation (ISBP) for optimizing a CNN. [0041]: If B P c o n v f n i ,   j ≠ 1 , this can indicate that the i.sup.th position in the output layer comes from a convolution operation involving the j t h   position in the input layer. [0023]: As discussed above, typical CNNs consist of convolutional (Conv) layer, pooling layer, non-linear layer (e.g. ReLU), normalization layer (e.g. local response normalization (LRN)) and fully connected (FC) layer, etc. The convolutional layer generally includes a set of trainable kernels, which extract local features from a small spatial region but cross-depth volume of the input tensor. Each kernel can be trained as a feature extractor for some specific visual features, such as an edge or a color in the first layer (BRI: normalization provides an output that is a nonlinear transformation of their input] [0020] : The deep convolutional neural network (DCNN) optimization component 140 is generally configured to optimize the structure of the trained DCNN model 138. In order to achieve a balance between the predictive power and model redundancy of CNNs, the DCNN optimization component 140 can learn the importance of convolutional kernels and neurons in FC layers from feature selection perspective. The DCNN optimization component 140 can optimize a CNN by pruning less important kernels and neurons based on their importance scores. The DCNN optimization component 140 can further fine-tune the remaining kernels and neurons to achieve a minimum loss of accuracy in the optimized DCNN. Chen does not explicitly disclose: - the layer determining the layer output using a non-linear normalization of the layer input. wherein for determining the layer output, the layer is configured to normalize at least one group of values of the layer input, wherein the group includes all values of the layer input or a subset of the values of the layer input; However, Lin discloses: - the layer determining the layer output using a non-linear normalization of the layer input. [0055]: The machine-learning model processes the input to generate an output. An artificial neural network includes an input layer that consists of values in a data point. The next layer is called a hidden layer, and nodes at the hidden layer each receive one or more of the input values. Each node contains parameters (e.g., weights) to apply to the input values. Each node therefore essentially inputs the input values into a multivariate function (e.g., a non-linear mathematical transformation) to produce an output value. A next layer can be another hidden layer or an output layer. In either case, the nodes at the next layer receive the output values from the nodes at the previous layer, and each node applies weights to those values and then generates its own output value. wherein for determining the layer output, the layer is configured to normalize at least one group of values of the layer input, wherein the group includes all values of the layer input or a subset of the values of the layer input; [0088]: At operation 412, processing logic performs one or more preprocessing operations on the input data. In some embodiments, the preprocessing operations can include a smoothing operation, a normalization operation, a dimensions reduction operations, a sort features operation, or any other operation configured to prepare data for training a machine-learning model. [0055]: The machine-learning model processes the input to generate an output. An artificial neural network includes an input layer that consists of values in a data point. The next layer is called a hidden layer, and nodes at the hidden layer each receive one or more of the input values. Each node contains parameters (e.g., weights) to apply to the input values. Each node therefore essentially inputs the input values into a multivariate function (e.g., a non-linear mathematical transformation) to produce an output value It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Chen and Lin. Chen a machine learning system to receive layer inputs and determine the layer output and teaches non-linear normalization of the layer input within the context of the CNN consisting of a “normalization layer”. Lin teaches a non-linear normalization of the layer input and determining layer output using sorting, smoothing within the context of a probability distribution. One ordinary skill would be motivated to combine Chen and Lin that can provide optimized model parameters and improvement to the model in terms of its accuracy (Lin [0059]). Chen and Lin do not explicitly disclose: - wherein the non-linear normalization includes mapping empirical percentiles of values from the group to percentiles of a predefined probability distribution. However, Sadri discloses: - wherein the non-linear normalization includes mapping empirical percentiles of values from the group to percentiles of a predefined probability distribution. [0035]: For further detail with respect to step 112, in an exemplary embodiment, a lower bound of a support set of the noise probability distribution may be equal to p l and an upper bound of the support set may be equal to p u   where 0≤ p l ≤ p u ≤1, PNG media_image2.png 68 432 media_image2.png Greyscale where μ is the mean of the noise probability distribution an exemplary embodiment, when μ is close to upper bound   p u   , a ratio of number of images that are corrupted by higher levels of impulsive noise to a number of plurality of training images 202 may increase It would be obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Chen, Lin and Sadri. Chen a machine learning system to receive layer inputs and determine the layer output using a non-linear normalization of the layer input. Lin teaches determining layer output using sorting, smoothing within the context of a probability distribution. Sadri teaches mapping empirical percentiles of values from the group to percentiles of a predefined probability distribution One of ordinary skill would be motivated to combine Chen, Lin and Sadri that can provide minimized loss function using gradient descent (Sadri ([0052]) Regarding claim 11: (Currently Amended) Chen discloses: - A training system configured to train a machine learning system including a plurality of layers, the training system configured to: provide an output signal based on an input signal by forwarding the input signal through the plurality of layers of the machine learning system, at least one of the layers of the plurality of layers being configured to receive a layer input, which is based on the input signal, and provide a layer output based on which the output signal is determined, the layer being configured to determine the layer output using a non-linear normalization of the layer input. [0021]: FIG. 2 illustrates a convolutional neural network, according to one embodiment described herein. PNG media_image1.png 416 765 media_image1.png Greyscale As shown, the CNN 200 includes an input layer 210, a convolutional layer 215, a subsampling layer 220, a convolutional layer 225, subsampling layer 230, fully connected layers 235 and 240, and an output layer 245. The input layer 210 in the depicted embodiment is configured to accept a 32×32 pixel image. The convolutional layer 215 generates 6 28×28 feature maps from the input layer, and so on. While a particular CNN 200 is depicted, more generally a CNN is composed of one or more convolutional layers, frequently with a subsampling step, and then one or more fully connected layers. [0037]: Generally, performing ISBP from 3-way tensor to 3-way tensor tends to be more complicated than the above two cases, as the operations of the forward propagation between the input and output [0027]: Accordingly, DCNN optimization component 140 can perform Importance Score Back Propagation (ISBP) for optimizing a CNN. [0041]: If B P c o n v f n i ,   j ≠ 1 , this can indicate that the i.sup.th position in the output layer comes from a convolution operation involving the j t h   position in the input layer. [0023]: As discussed above, typical CNNs consist of convolutional (Conv) layer, pooling layer, non-linear layer (e.g. ReLU), normalization layer (e.g. local response normalization (LRN)) and fully connected (FC) layer, etc. The convolutional layer generally includes a set of trainable kernels, which extract local features from a small spatial region but cross-depth volume of the input tensor. Each kernel can be trained as a feature extractor for some specific visual features, such as an edge or a color in the first layer [0020] : The deep convolutional neural network (DCNN) optimization component 140 is generally configured to optimize the structure of the trained DCNN model 138. In order to achieve a balance between the predictive power and model redundancy of CNNs, the DCNN optimization component 140 can learn the importance of convolutional kernels and neurons in FC layers from feature selection perspective. The DCNN optimization component 140 can optimize a CNN by pruning less important kernels and neurons based on their importance scores. The DCNN optimization component 140 can further fine-tune the remaining kernels and neurons to achieve a minimum loss of accuracy in the optimized DCNN. Chen does not explicitly disclose: - the layer determining the layer output using a non-linear normalization of the layer input. wherein for determining the layer output, the layer is configured to normalize at least one group of values of the layer input, wherein the group includes all values of the layer input or a subset of the values of the layer input; However, Lin discloses: - wherein the layer is configured to determine the layer output using a non-linear normalization of the layer input. [0055]: The machine-learning model processes the input to generate an output. An artificial neural network includes an input layer that consists of values in a data point. The next layer is called a hidden layer, and nodes at the hidden layer each receive one or more of the input values. Each node contains parameters (e.g., weights) to apply to the input values. Each node therefore essentially inputs the input values into a multivariate function (e.g., a non-linear mathematical transformation) to produce an output value. A next layer can be another hidden layer or an output layer. In either case, the nodes at the next layer receive the output values from the nodes at the previous layer, and each node applies weights to those values and then generates its own output value. wherein for determining the layer output, the layer is configured to normalize at least one group of values of the layer input, wherein the group includes all values of the layer input or a subset of the values of the layer input; [0088]: At operation 412, processing logic performs one or more preprocessing operations on the input data. In some embodiments, the preprocessing operations can include a smoothing operation, a normalization operation, a dimensions reduction operations, a sort features operation, or any other operation configured to prepare data for training a machine-learning model. [0055]: The machine-learning model processes the input to generate an output. An artificial neural network includes an input layer that consists of values in a data point. The next layer is called a hidden layer, and nodes at the hidden layer each receive one or more of the input values. Each node contains parameters (e.g., weights) to apply to the input values. Each node therefore essentially inputs the input values into a multivariate function (e.g., a non-linear mathematical transformation) to produce an output value It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Chen and Lin. Chen a machine learning system to receive layer inputs and determine the layer output and teaches non-linear normalization of the layer input within the context of the CNN consisting of a “normalization layer”. Lin teaches a non-linear normalization of the layer input and determining layer output using sorting, smoothing within the context of a probability distribution. One ordinary skill would be motivated to combine Chen and Lin that can provide optimized model parameters and improvement to the model in terms of its accuracy (Lin [0059]). Chen and Lin do not explicitly disclose: - wherein the non-linear normalization includes mapping empirical percentiles of values from the group to percentiles of a predefined probability distribution. However, Sadri discloses: - wherein the non-linear normalization includes mapping empirical percentiles of values from the group to percentiles of a predefined probability distribution. [0035]: For further detail with respect to step 112, in an exemplary embodiment, a lower bound of a support set of the noise probability distribution may be equal to p l and an upper bound of the support set may be equal to p u   where 0≤ p l ≤ p u ≤1, PNG media_image2.png 68 432 media_image2.png Greyscale where μ is the mean of the noise probability distribution an exemplary embodiment, when μ is close to upper bound   p u   , a ratio of number of images that are corrupted by higher levels of impulsive noise to a number of plurality of training images 202 may increase It would be obvious to one of ordinary skills in the art before the effective filing date of the present application to combine Chen, Lin and Sadri. Chen a machine learning system to receive layer inputs and determine the layer output using a non-linear normalization of the layer input. Lin teaches determining layer output using sorting, smoothing within the context of a probability distribution. Sadri teaches mapping empirical percentiles of values from the group to percentiles of a predefined probability distribution One of ordinary skill would be motivated to combine Chen, Lin and Sadri that can provide minimized loss function using gradient descent (Sadri ([0052]) Regarding claim 12: (Currently Amended) Chen discloses: - A non-transitory machine-readable storage medium on which is stored a computer program [0050]: The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device - for determining an output signal based on an input signal, the computer program, when executed by a computer, causing the computer to perform: determining the output signal by providing the input signal to a machine learning system, the machine learning system including a plurality of layers, the machine learning system providing the output signal based on an input signal by forwarding the input signal through the plurality of layers of the machine learning system, at least one of the layers of the plurality of layers receiving a layer input, which is based on the input signal, and providing a layer output based on which the output signal is determined, [0050]: A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire. [0021]: FIG. 2 illustrates a convolutional neural network, according to one embodiment described herein. PNG media_image1.png 416 765 media_image1.png Greyscale As shown, the CNN 200 includes an input layer 210, a convolutional layer 215, a subsampling layer 220, a convolutional layer 225, subsampling layer 230, fully connected layers 235 and 240, and an output layer 245. The input layer 210 in the depicted embodiment is configured to accept a 32×32 pixel image. The convolutional layer 215 generates 6 28×28 feature maps from the input layer, and so on. While a particular CNN 200 is depicted, more generally a CNN is composed of one or more convolutional layers, frequently with a subsampling step, and then one or more fully connected layers. [0037]: Generally, performing ISBP from 3-way tensor to 3-way tensor tends to be more complicated than the above two cases, as the operations of the forward propagation between the input and output [0027]: Accordingly, DCNN optimization component 140 can perform Importance Score Back Propagation (ISBP) for optimizing a CNN. [0041]: If B P c o n v f n i ,   j ≠ 1 , this can indicate that the i.sup.th position in the output layer comes from a convolution operation involving the j t h   position in the input layer. [0023]: As discussed above, typical CNNs consist of convolutional (Conv) layer, pooling layer, non-linear layer (e.g. ReLU), normalization layer (e.g. local response normalization (LRN)) and fully connected (FC) layer, etc. The convolutional layer generally includes a set of trainable kernels, which extract local features from a small spatial region but cross-depth volume of the input tensor. Each kernel can be trained as a feature extractor for some specific visual features, such as an edge or a color in the first layer [0020] : The deep convolutional neural network (DCNN) optimization component 140 is generally configured to optimize the structure of the trained DCNN model 138. In order to achieve a balance between the predictive power and model redundancy of CNNs, the DCNN optimization component 140 can learn the importance of convolutional kernels and neurons in FC layers from feature selection perspective. The DCNN optimization component 140 can optimize a CNN by pruning less important kernels and neurons based on their importance scores. The DCNN optimization component 140 can further fine-tune the remaining kernels and neurons to achieve a minimum loss of accuracy in the optimized DCNN. Chen does not explicitly disclose: - the layer determining the layer output using a non-linear normalization of the layer input. wherein for determining the layer output, the layer is configured to normalize at least one group of values of the layer input, wherein the group includes all values of the layer input or a subset of the values of the layer input; However, Lin discloses: - wherein the layer is configured to determine the layer output using a non-linear normalization of the layer input. [0055]: The machine-learning model processes the input to generate an output. An artificial neural network includes an input layer that consists of values in a data point. The next layer is called a hidden layer, and nodes at the hidden layer each receive one or more of the input values. Each node contains parameters (e.g., weights) to apply to the input values. Each node therefore essentially inputs the input values into a multivariate function (e.g., a non-linear mathematical transformation) to produce an output value. A next layer can be another hidden layer or an output layer. In either case, the nodes at the next layer receive the output values from the nodes at the previous layer, and each node applies weights to those values and then generates its own output value. wherein for determining the layer output, the layer is configured to normalize at least one group of values of the layer input, wherein the group includes all values of the layer input or a subset of the values of the layer input; [0088]: At operation 412, processing logic performs one or more preprocessing operations on the input data. In some embodiments, the preprocessing operations can include a smoothing operation, a normalization operation, a dimensions reduction operations, a sort features operation, or any other operation configured to prepare data for training a machine-learning model. In [0055]: The machine-learning model processes the input to generate an output. An artificial neural network includes an input layer that consists of values in a data point. The next layer is called a hidden layer, and nodes at the hidden layer each receive one or more of the input values. Each node contains parameters (e.g., weights) to apply to the input values. Each node therefore essentially inputs the input values into a multivariate function (e.g., a non-linear mathematical transformation) to produce an output value It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Chen and Lin. Chen a machine learning system to receive layer inputs and determine the layer output and teaches non-linear normalization of the layer input within the context of the CNN consisting of a “normalization layer”. Lin teaches a non-linear normalization of the layer input and determining layer output using sorting, smoothing within the context of a probability distribution. One of ordinary skill would be motivated to combine Chen and Lin that can provide optimized model parameters and improvement to the model in terms of its accuracy (Lin [0059]). Chen and Lin do not explicitly disclose: wherein the non-linear normalization includes mapping empirical percentiles of values from the group to percentiles of a predefined probability distribution. However, Sadri discloses: wherein the non-linear normalization includes mapping empirical percentiles of values from the group to percentiles of a predefined probability distribution. [0035]: For further detail with respect to step 112, in an exemplary embodiment, a lower bound of a support set of the noise probability distribution may be equal to p l and an upper bound of the support set may be equal to p u   where 0≤ p l ≤ p u ≤1, PNG media_image2.png 68 432 media_image2.png Greyscale where μ is the mean of the noise probability distribution an exemplary embodiment, when μ is close to upper bound   p u   , a ratio of number of images that are corrupted by higher levels of impulsive noise to a number of plurality of training images 202 may increase It would be obvious to one of ordinary skills in the art before the effective filing date of the present application to combine Chen, Lin and Sadri. Chen a machine learning system to receive layer inputs and determine the layer output using a non-linear normalization of the layer input. Lin teaches determining layer output using sorting, smoothing within the context of a probability distribution. Sadri teaches mapping empirical percentiles of values from the group to percentiles of a predefined probability distribution One of ordinary skill would be motivated to combine Chen, Lin and Sadri that can provide minimized loss function using gradient descent (Sadri ([0052]) Claim 13 is rejected under 35 U.S.C. 103 as being unpatentable over Chun-Fu CHEN et al. (hereinafter Chen) US 2019/0122113 A1, in view of Jiu-Che Lin et.al. (hereinafter Lin) US 2023/0342016 A1, in view of Seyedeh Sahar Sadrizadeh et.al. (hereinafter Sadri) US 2021/0209735 A1. further in view of Hanxiao Liu et.al. (hereinafter Liu) US 2023/0121404 A1. Regarding claim 13: (New) Chen, Lin and Sadri do not explicitly disclose: - wherein the at least one of the layers of the plurality of layers is configured to accept a layer input of an arbitrary shape and to normalize each element of the layer input to determine a layer output of the same shape as the layer input. However, Liu discloses: - wherein the at least one of the layers of the plurality of layers is configured to accept a layer input of an arbitrary shape and to normalize each element of the layer input to determine a layer output of the same shape as the layer input. [Abstract]: an architecture for an activation-normalization layer to be included in a neural network to replace a set of layers that receive a layer input comprising a plurality of values, apply one or more normalization operations to the values in the layer input to generate a normalized layer input, and apply an element-wise activation function to the normalized layer input to generate a layer output. [0005]: an architecture for a normalization - activation neural network layer (also referred to as an “NA layer”), [0039]: As a particular example, each possible candidate architecture in the search space can be represented as a computation graph that transforms one input tensor into an output tensor of the same shape. [0033]: Examples of normalization operations include those performed by batch normalization, group normalization, and layer normalization layers in order to normalize received inputs. Examples of element-wise activation functions include rectified linear unit (ReLU), inverse tangent, and sigmoid functions. [0029]: At a high level, the system 100 determines the architecture for the NA layer by generating a plurality of candidate architectures for the normalization - activation layer, evaluating how each of the generated candidate architectures perform when included in multiple different neural network architectures, and then selecting an architecture based on the results of the evaluation. It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Chen, Lin, Sadri and Liu. Chen a machine learning system to receive layer inputs and determine the layer output and teaches non-linear normalization of the layer input within the context of the CNN consisting of a “normalization layer”. Lin teaches a non-linear normalization of the layer input and determining layer output using sorting, smoothing within the context of a probability distribution. Liu teaches the output shape same as input shape. Sadri teaches mapping empirical percentiles of values from the group to percentiles of a predefined probability distribution. One of ordinary skills would be to combine Chen, Lin, Sadri and Liu that provide an improved stability of the training process (Liu [0007]) Conclusion THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to TIRUMALE KRISHNASWAMY RAMESH whose telephone number is (571)272-4605. The examiner can normally be reached by phone. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Li B Zhen can be reached on phone (571-272-3768). The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /TIRUMALE K RAMESH/Examiner, Art Unit 2121 /Li B. Zhen/Supervisory Patent Examiner, Art Unit 2121
Read full office action

Prosecution Timeline

Apr 04, 2023
Application Filed
Nov 01, 2023
Response after Non-Final Action
Mar 05, 2026
Non-Final Rejection mailed — §101, §103
May 27, 2026
Response Filed
Aug 26, 2026
Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12699873
NEURAL NETWORK PROCESSING USING MIXED-PRECISION DATA REPRESENTATION
6y 6m to grant Granted Aug 04, 2026
Patent 12688395
Neural Network Processor with On-Chip Convolution Kernel Storage
8y 5m to grant Granted Jul 21, 2026
Patent 12518153
TRAINING MACHINE LEARNING SYSTEMS
5y 12m to grant Granted Jan 06, 2026
Patent 12293284
META COOPERATIVE TRAINING PARADIGMS
4y 4m to grant Granted May 06, 2025
Patent 12229651
BLOCK-BASED INFERENCE METHOD FOR MEMORY-EFFICIENT CONVOLUTIONAL NEURAL NETWORK IMPLEMENTATION AND SYSTEM THEREOF
4y 4m to grant Granted Feb 18, 2025
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
26%
Grant Probability
50%
With Interview (+23.7%)
4y 9m (~1y 2m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 49 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month