DETAILED ACTION
This action is in response to the amendments and remarks filed 06/22/2026. Claims 1-11 are pending and have been examined.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Acknowledgment is made of applicant's claim for foreign priority based on an application filed in Japan on 05/12/2022. It is noted, however, that applicant has not filed a certified copy of the JP2022-078954 application as required by 37 CFR 1.55.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 04/01/2026 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 5-7 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter without significantly more.
Claim 5
Step 1: The claim’s ancestor, claim 1, recites “An information processing apparatus”. Therefore, claim 5 is therefore directed to the statutory category of machine
Step 2A Prong 1: The claim recites the following judicial exception(s)
wherein the loss is calculated by a loss function including a regularization item that is large when the size of the output exceeds the quantization parameter: This recites a mathematical concept in the form of an equation with a term that scales directly with size of the output exceeding the quantization parameter.
Step 2A Prong 2: The judicial exception(s) are not integrated into a practical application through the following additional element(s)
Ancestral claim 1:
at least one processor; and at least one memory having stored thereon instructions which, when executed by the at least one processor, cause the information processing apparatus to: This is mere instruction to execute the recited judicial exceptions with generic computer hardware (MPEP 2106.05(f)).
obtain information indicating a size of an output as a result of a first operation in a neural network that performs the first operation using a weight coefficient for input data and a second operation of quantizing a result of the first operation, in order to obtain data of an intermediate layer: This is directed to mere reception of data and is insignificant extra-solution activity (MPEP 2106.05(g)).
control the first operation by controlling the weight coefficient used in the first operation to adjust the size of the output based on the information and a quantization parameter used for the quantization: This amounts to updating data. In other words, mere reception of data updates and is insignificant extra-solution activity (MPEP 2106.05(g)).
Ancestral claim 4:
wherein executing the stored instructions by the at least one processor further causes the information processing apparatus to: This is mere instruction to execute the recited judicial exceptions with generic computer hardware (MPEP 2106.05(f)).
control the first operation by controlling the weight coefficient of the neural network, through learning in which the size of the output exceeding the quantization parameter results in a large loss: This amounts to updating data. In other words, mere reception of data updates and is insignificant extra-solution activity (MPEP 2106.05(g)).
Step 2B: The following additional element(s) of the claim, taken alone or in combination, do not amount to significantly more than the recited judicial exception(s)
Ancestral claim 1:
at least one processor; and at least one memory having stored thereon instructions which, when executed by the at least one processor, cause the information processing apparatus to: This is mere instruction to execute the recited judicial exceptions with generic computer hardware (MPEP 2106.05(f)).
obtain information indicating a size of an output as a result of a first operation in a neural network that performs the first operation using a weight coefficient for input data and a second operation of quantizing a result of the first operation, in order to obtain data of an intermediate layer: This is an instance of retrieving information from memory, a limitation known to be well-understood, routine, and conventional (MPEP 2106.05(d) II. iv.).
control the first operation by controlling the weight coefficient used in the first operation to adjust the size of the output based on the information and a quantization parameter used for the quantization: This is an instance of updating a value in memory. In other words, of retrieving information from memory, a well-known, understood, and routine limitation (MPEP 2106.05(d) II. iv.)
Ancestral claim 4:
wherein executing the stored instructions by the at least one processor further causes the information processing apparatus to: This is mere instruction to execute the recited judicial exceptions with generic computer hardware (MPEP 2106.05(f)).
control the first operation by controlling the weight coefficient of the neural network, through learning in which the size of the output exceeding the quantization parameter results in a large loss: This is an instance of updating a value in memory. In other words, of retrieving information from memory, a well-known, understood, and routine limitation (MPEP 2106.05(d) II. iv.)
Claim 6
Step 1: The claim recites a machine, as in claim 5
Step 2A Prong 1: The claim recites the following further judicial exception(s)
evaluate a recognition accuracy of the neural network for a detection target: This can be performed as a mental process. One can merely gauge the accuracy of the network’s results.
evaluate the recognition accuracy of the neural network for the detection target after the weight coefficient has been quantized: This can be performed as a mental process. One can merely gauge the accuracy of the network’s results after quantization.
Step 2A Prong 2: The judicial exception(s) are not integrated into a practical application through the further additional element(s)
wherein executing the stored instructions by the at least one processor further causes the information processing apparatus to: This is mere instruction to execute the recited judicial exceptions with generic computer hardware (MPEP 2106.05(f)).
quantize the weight coefficient of the neural network: This is an instance of conventional quantization and amounts to insignificant extra-solution activity (MPEP 2106.05(g)).
correct the regularization item included in the loss function, based on the recognition accuracy evaluated and the recognition accuracy evaluated: This is mere instruction to apply the judicial exceptions to the loss in a generic manner (MPEP 2106.05(f)).
Step 2B: The further additional element(s) of the claim, taken alone or in combination, do not amount to significantly more than the recited judicial exception(s)
wherein executing the stored instructions by the at least one processor further causes the information processing apparatus to: This is mere instruction to execute the recited judicial exceptions with generic computer hardware (MPEP 2106.05(f)).
quantize the weight coefficient of the neural network: This is an instance of using quantization on weights of a neural network, a standard technique in quantized neural networks, as noted by Choi ( METHOD AND DEVICE FOR DETERMINING SATURATION RATIO-BASED QUANTIZATION RANGE FOR QUANTIZATION OF NEURAL NETWORK, filed 7/22/2022, US 20240320464 A1): “Here, quantization means mapping tensor values from a dimension with a wide data representation range to a dimension with a narrow data representation range. In other words, quantization means that a processor that processes neural network operations maps high-precision tensors to low-precision values. In artificial neural networks, quantization can be applied to tensors including activations, weights, and biases of a layer.” (Choi, [0006])
correct the regularization item included in the loss function, based on the recognition accuracy evaluated and the recognition accuracy evaluated: This is mere instruction to apply the judicial exceptions to the loss in a generic manner (MPEP 2106.05(f)).
Claim 7
Step 1: The claim recites a machine, as in claim 6.
Step 2A Prong 1: The claim recites the following further judicial exception(s)
evaluate a deterioration degree of the recognition accuracy of the neural network for the detection target due to the quantization of the weight coefficient using the recognition accuracy evaluated and the recognition accuracy evaluated: This can be performed as a mental process. One can merely subtract the second accuracy from the first to measure a difference caused by quantization.
Step 2A Prong 2: The judicial exception(s) are not integrated into a practical application through the following further additional element(s)
wherein executing the stored instructions by the at least one processor further causes the information processing apparatus to: This is mere instruction to execute the recited judicial exceptions with generic computer hardware (MPEP 2106.05(f)).
correct the regularization item using the deterioration degree: This is mere instruction to correct a regularization item in a generic manner based on a judicial exception (MPEP 2106.05(f)).
Step 2B: The following additional element(s) of the claim, taken alone or in combination, do not amount to significantly more than the recited judicial exception(s)
wherein executing the stored instructions by the at least one processor further causes the information processing apparatus to: This is mere instruction to execute the recited judicial exceptions with generic computer hardware (MPEP 2106.05(f)).
correct the regularization item using the deterioration degree: This is mere instruction to correct a regularization item in a generic manner based on a judicial exception (MPEP 2106.05(f)).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-9 and 11 are rejected under 35 U.S.C. 103 as being unpatentable over Banner (ACIQ: Analytical Clipping for Integer Quantization of neural networks, published 9/27/2018, ICLR 2019 Conference Blind Submission, retrieved from https://openreview.net/forum?id=B1x33sC9KQ) in view of Diril (LAYER-LEVEL QUANTIZATION IN NEURAL NETWORKS, published 6/6/2019, US 2019/0171927 A1) and Sasagawa (Neural Network Derivation Method, published 8/12/2021, US 2021/0248463 A1).
Regarding claim 1, Banner discloses [a]n information processing apparatus … to:
obtain information indicating a size of an output as a result of a first operation in a neural network that performs the first operation using a weight coefficient for input data and a second operation of quantizing a result of the first operation, in order to obtain data of an intermediate layer:
“In Section 4, we provide a rigorous formulation to optimize the quantization effect of activation tensors (output[s]) using clipping by analyzing both the Gaussian and the Laplace priors. This formulation is henceforth refered [sic] to as Analytical Clipping for Integer Quantization (ACIQ).” (Banner, page 2, paragraph 2)
“Commonly (e.g., in GEMMLOWP), integer tensors are uniformly quantized in the range
[
-
α
,
α
]
, where
α
is determined by the tensor maximal absolute value. In the following we show that the this [sic] choice of
α
is suboptimal, and suggest a model where the tensor values (output[s]) are clipped to reduce quantization noise. For any
x
∈
R
(size of an output), we define the clipping function
c
l
i
p
(
x
,
α
)
as follows
PNG
media_image1.png
84
482
media_image1.png
Greyscale
” (Banner, page 4, paragraph 4). To clip an activation tensor, the size of it must be measured and compared against alpha.
Examiner’s note: As one of ordinary skill in the art would know, each activation is calculated as a weighted sum (first operation) of weights (weight coefficient[s]) and inputs from the previous layer.
“For each method, we select N layers (tensors) to be quantized (second operation) to 4 bits using our optimized clipping method and compare it against the standard GEMMLOWP approach. In figure 3 we present this accuracy-quantization tradeoff.” (Banner, page 7, paragraph 3). The size of each activation must be measured for activations in N layers to be input into the clipping function.
PNG
media_image2.png
200
400
media_image2.png
Greyscale
(Banner, page 8, Figure 3). For N > 2, N - 2 intermediate layer activations are being clipped and quantized.
control the first operation by controlling the weight coefficient used in the first operation to adjust the size of the output based on the information and a quantization parameter used for the quantization:
“Commonly (e.g., in GEMMLOWP), integer tensors are uniformly quantized in the range
[
-
α
,
α
]
, where
α
(quantization parameter) is determined by the tensor maximal absolute value. In the following we show that the this [sic] choice of
α
is suboptimal, and suggest a model where the tensor values (output[s]) are clipped to reduce quantization noise. For any
x
∈
R
(the information), we define the clipping function
c
l
i
p
(
x
,
α
)
as follows
PNG
media_image1.png
84
482
media_image1.png
Greyscale
” (Banner, page 4, paragraph 4).
“the expected mean-square-error between X and its quantized version Q(X) can be written as follows:
PNG
media_image3.png
332
551
media_image3.png
Greyscale
” (Banner, page 4, paragraph 6). As can be seen in the formula above, the mean square error depends partly on (x -
α
). In other words, the value of the loss increases as the difference between the activation tensor size and quantization threshold increases.
Banner relates to quantizing neural networks and is analogous to the claimed invention.
While Banner fails to disclose the further limitations of the claim, Diril discloses [a]n information processing apparatus comprising: at least one processor; and at least one memory having stored thereon instructions which, when executed by the at least one processor, cause the information processing apparatus to: “Processor 814 generally represents any type or form of physical processing unit (e.g., a hardware-implemented central processing unit) capable of processing data or interpreting and executing instructions. In certain embodiments, processor 814 may receive instructions from a software application or module. These instructions may cause processor 814 to perform the functions of one or more of the example embodiments described and/or illustrated herein.” (Diril, [0057]); “System memory 816 generally represents any type or form of volatile or non-volatile storage device or medium capable of storing data and/or other computer-readable instructions” (Diril, [0058])
Diril relates to quantizing neural networks and is analogous to the claimed invention. Banner teaches a method of quantizing neural networks. The claimed invention improves upon this method by storing it in the form of instructions on computer hardware. Diril teaches standard computer hardware, applicable to Banner. A person of ordinary skill in the art would have recognized that storing Banner’s method as computer instructions on Diril’s hardware would lead to the predictable result of the method being executable by a computing system, and would improve the known device by allowing it to be performed with real data (MPEP 2143 I. (D) Applying a known technique to a known device (method, or product) ready for improvement to yield predictable results). While Diril fails to disclose the further limitations of the claim, Sasagawa discloses instructions to control the first operation by controlling the weight coefficient used in the first operation to adjust the size of the output based on the information and a quantization parameter used for the quantization: “in the present disclosure, when an inference model is generated, training is performed using a loss function for optimization to which a regularization term is added. This regularization term prevents a weight parameter used by a neural network from becoming a weight parameter likely to change the accuracy of an inferred value. For example, in the present disclosure, when a neural network is trained, a weight parameter is updated so that a value of “loss function+regularization term (loss)” becomes smaller, and an inference model is generated. Accordingly, even when a weight parameter is quantized at the time of mounting, it is possible to reduce a significant decrease in accuracy of an inferred value.” (Sasagawa, [0037])
Sasagawa relates to quantized neural networks and is analogous to the claimed invention. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the existing combination to adjust the weights based on loss during learning, as disclosed by Sasagawa. Doing so would avoid reductions in accuracy, even for quantized weights. See Sasagawa, [0037].
Regarding claim 2, the rejection of claim 1 is incorporated. Diril further discloses an apparatus, wherein the information indicating the size of the output is information calculated based on a distribution of values of the output: “This second limit value may correspond to a maximum value for the activation layer, such as an absolute maximum weight or filter value (e.g., the highest value of an activation layer (distribution of values of the output), which may be identified by passing output values through a min-max unit) or an estimated maximum weight (information indicating the size of the output) or filter value (e.g., an approximate maximum that discards outliers, a maximum within a predetermined standard deviation of values for a particular layer, etc.).” (Diril, [0045]).
Diril relates to quantization of neural networks and is analogous to the claimed invention. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the existing combination to calculate approximate maximum activation values, as disclosed by Diril. Such maximum values can be used to determine an upper bound for quantization that preserves as much accuracy as possible while quantizing values. See Diril, [0046].
Regarding claim 3, the rejection of claim 2 is incorporated. Diril further discloses an apparatus, wherein the information indicating the size of the output is information indicating an upper limit excluding an outlier of the output: “This second limit value may correspond to a maximum value for the activation layer, such as an absolute maximum weight or filter value (e.g., the highest value of an activation layer, which may be identified by passing output values through a min-max unit) or an estimated maximum weight (information indicating the size of the output) or filter value (e.g., an approximate maximum that discards outliers, a maximum within a predetermined standard deviation of values for a particular layer, etc.).” (Diril, [0045]).
Diril relates to quantization of neural networks and is analogous to the claimed invention. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the existing combination to calculate approximate maximum activation values, as disclosed by Diril. Such maximum values can be used to determine an upper bound for quantization that preserves as much accuracy as possible while quantizing values. See Diril, [0046].
Regarding claim 4, the rejection of claim 1 is incorporated. Banner, in combination with Diril, discloses an information processing apparatus, wherein executing the stored instructions by the at least one processor further causes the information processing apparatus to control the first operation by controlling the weight coefficient of the neural network, through learning in which the size of the output exceeding the quantization parameter results in a large loss:
(Banner) “Commonly (e.g., in GEMMLOWP), integer tensors are uniformly quantized in the range
[
-
α
,
α
]
, where
α
(quantization parameter) is determined by the tensor maximal absolute value. In the following we show that the this [sic] choice of
α
is suboptimal, and suggest a model where the tensor values (output[s]) are clipped to reduce quantization noise. For any
x
∈
R
(size of the output), we define the clipping function
c
l
i
p
(
x
,
α
)
as follows
PNG
media_image1.png
84
482
media_image1.png
Greyscale
” (Banner, page 4, paragraph 4)
(Banner) “the expected mean-square-error between X and its quantized version Q(X) can be written as follows:
PNG
media_image3.png
332
551
media_image3.png
Greyscale
” (Banner, page 4, paragraph 6). As can be seen in the formula above, the mean square error depends partly on (x -
α
). In other words, the value of the loss increases as the difference between the activation tensor size and quantization threshold increases.
Examiner’s note: As noted in the rejection of ancestral claim 1, Diril discloses processing hardware that can execute stored program instructions.
While Banner and Diril fail to disclose the further limitations of the claim, Sasagawa discloses instructions to: control the first operation by controlling the weight coefficient of the neural network, through learning in which the size of the output exceeding the quantization parameter results in a large loss: “in the present disclosure, when an inference model is generated, training is performed using a loss function for optimization to which a regularization term is added. This regularization term prevents a weight parameter used by a neural network from becoming a weight parameter likely to change the accuracy of an inferred value. For example, in the present disclosure, when a neural network is trained, a weight parameter is updated so that a value of “loss function+regularization term (loss)” becomes smaller, and an inference model is generated. Accordingly, even when a weight parameter is quantized at the time of mounting, it is possible to reduce a significant decrease in accuracy of an inferred value.” (Sasagawa, [0037])
Sasagawa relates to quantized neural networks and is analogous to the claimed invention. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the existing combination to adjust the weights based on loss during learning, as disclosed by Sasagawa. Doing so would avoid reductions in accuracy, even for quantized weights. See Sasagawa, [0037]
Regarding claim 5, the rejection of claim 4 is incorporated. Banner further discloses an apparatus, wherein the loss is calculated by a loss function including a regularization item that is large when the size of the output exceeds the quantization parameter: “the expected mean-square-error between X and its quantized version Q(X) can be written as follows:
PNG
media_image3.png
332
551
media_image3.png
Greyscale
” (Banner, page 4, paragraph 6).
Regarding claim 6, the rejection of claim 5 is incorporated. Sasagawa further discloses instructions to:
evaluate a recognition accuracy of the neural network for a detection target:
“Discrimination training model 40 is a model for training discriminator 41 that determines the accuracy of an inferred value” (Sasagawa, [0048])
“Pre-quantization weight parameter w+Δw and quantized weight parameter w.sup.q are inputted to discriminator 41. Discriminator 41 outputs inferred value D(w+Δw) in response to weight parameter w+Δw, and outputs inferred value D(w.sup.q) in response to weight parameter w.sup.q.” (Sasagawa, [0049])
“Inferred value (first inferred value) x of pre-quantization model 20 (the neural network) and inferred value (second inferred value) G(z) of quantized model 30 are inputted to discrimination training model 40. Discrimination training model 40 contrasts inputted inferred value x and inferred value G(z) with above-described inferred values D(w+Δw) and D(w.sup.q), and trains discriminator 41 by performing backpropagation.” (Sasagawa, [0050])
quantize the weight coefficient of the neural network: “Quantized model 30 includes a second neural network having weight parameter (second parameter) w.sup.q. Weight parameter w.sup.q is obtained by converting weight parameter w of pre-quantization model 20 into a second numeric representation different from the above-described first numeric representation … Specifically, weight parameter w.sup.q is obtained by quantizing weight parameter w+Δw obtained by adding Δw to weight parameter w.” (Sasagawa, [0047])
evaluate the recognition accuracy of the neural network for the detection target after the weight coefficient has been quantized: “Inferred value (first inferred value) x of pre-quantization model 20 and inferred value (second inferred value) G(z) of quantized model 30 (neural network for the detection target after the weight coefficient has been quantized) are inputted to discrimination training model 40. Discrimination training model 40 contrasts inputted inferred value x and inferred value G(z) with above-described inferred values D(w+Δw) and D(w.sup.q), and trains discriminator 41 by performing backpropagation.” (Sasagawa, [0050])
correct the regularization item included in the loss function, based on the recognition accuracy evaluated and the recognition accuracy evaluated:
“The regularization term in above “loss function+regularization term” is determined to be larger when the accuracy of an inferred value decreases, and is determined to be smaller when the accuracy of an inferred value increases.” (Sasagawa, [0038])
“FIG. 2 is a schematic diagram illustrating a function of a discriminator that determines the accuracy of an inferred value. FIG. 2 shows a state in which weight parameters that decrease the accuracy of an inferred value when the weight parameters are quantized (the lower region of FIG. 2) and weight parameters that are less likely to decrease the accuracy of an inferred value even when the weight parameters are quantized (the upper region of FIG. 2) are classified by the function of the discriminator. When such a discriminator can be generated, it is possible to determine whether an unknown weight parameter is to decrease the accuracy of an inferred value, and it is possible to decide whether to increase or decrease a regularization term based on the determination result.” (Sasagawa, [0039])
Sasagawa relates to quantizing neural networks and is analogous to the claimed invention. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the existing combination to add a regularization term proportional to accuracy discrepancies in quantized weights to the loss function, as disclosed by Sasagawa. Doing so would reduce loss of accuracy due to quantized weights. See Sasagawa, [0037].
Regarding claim 7, the rejection of claim 6 is incorporated. Sasagawa further discloses instructions to:
evaluate a deterioration degree of the recognition accuracy of the neural network for the detection target due to the quantization of the weight coefficient using the recognition accuracy evaluated and the recognition accuracy evaluated: “FIG. 2 is a schematic diagram illustrating a function of a discriminator that determines the accuracy of an inferred value. FIG. 2 shows a state in which weight parameters that decrease the accuracy of an inferred value when the weight parameters are quantized (the lower region of FIG. 2) and weight parameters that are less likely to decrease the accuracy of an inferred value even when the weight parameters are quantized (the upper region of FIG. 2) are classified by the function of the discriminator. When such a discriminator can be generated, it is possible to determine whether an unknown weight parameter is to decrease the accuracy of an inferred value, and it is possible to decide whether to increase or decrease a regularization term based on the determination result.” (Sasagawa, [0039])
correct the regularization item using the deterioration degree: “The regularization term in above “loss function+regularization term” is determined to be larger when the accuracy of an inferred value decreases, and is determined to be smaller when the accuracy of an inferred value increases.” (Sasagawa, [0038])
Sasagawa relates to quantizing neural networks and is analogous to the claimed invention. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the existing combination to add a regularization term proportional to accuracy discrepancies in quantized weights to the loss function, as disclosed by Sasagawa. Doing so would reduce loss of accuracy due to quantized weights. See Sasagawa, [0037].
Regarding claim 8, the rejection of claim 1 is incorporated. While Banner fails to disclose the further limitations of the claim, Sasagawa discloses an apparatus, wherein the at least one processor adjusts the size of the output by correcting the weight coefficient of the intermediate layer:
“These models each have a multi-layer structure and include an input layer, an intermediate layer, and an output layer, etc. Each of the layers includes nodes (not shown) corresponding to neurons. The strength of a connection between neurons is represented by a weight parameter. Although a neural network has weight parameters, in order to facilitate understanding, a weight parameter will be described below as an example of weight parameters.” (Sasagawa, [0045]). The weight parameters of a network are described as a single ‘weight parameter’.
“in the present disclosure, when a neural network is trained, a weight parameter is updated so that a value of “loss function+regularization term” becomes smaller, and an inference model is generated.” (Sasagawa, [0037]). Weights are adjusted to minimize the regularization term.
“The regularization term in above “loss function+regularization term” is determined to be larger when the accuracy of an inferred value decreases, and is determined to be smaller when the accuracy of an inferred value increases. The following describes how to determine whether the accuracy of an inferred value is to increase or decrease.” (Sasagawa, [0038]). Weight changes cause changes to the output, and consequently the accuracy of a model.
Sasagawa relates to quantized neural networks and is analogous to the claimed invention. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the existing combination to adjust the weights based on loss during learning, as disclosed by Sasagawa. Doing so would avoid reductions in accuracy, even for quantized weights. See Sasagawa, [0037].
Regarding claim 9, the rejection of claim 8 is incorporated. Banner further discloses an apparatus, wherein the at least one processor controls the first operation in the neural network when the size of the output exceeds a predetermined value: “Commonly (e.g., in GEMMLOWP), integer tensors are uniformly quantized in the range
[
-
α
,
α
]
, where
α
is determined by the tensor maximal absolute value. In the following we show that the this [sic] choice of
α
(predetermined value) is suboptimal, and suggest a model where the tensor values are clipped to reduce quantization noise. For any
x
∈
R
(size of the output), we define the clipping function
c
l
i
p
(
x
,
α
)
as follows
PNG
media_image1.png
84
482
media_image1.png
Greyscale
” (Banner, page 4, paragraph 4). As noted in the rejection for claim 1, the first operation is the weighted sum comprising the activation. The result is clipped depending on its size relative to alpha.
Regarding claim 11, Banner discloses a program that causes the computer to:
obtain information indicating a size of an output as a result of a first operation in a neural network that performs the first operation using a weight coefficient for input data and a second operation of quantizing a result of the first operation, in order to obtain data of an intermediate layer:
“In Section 4, we provide a rigorous formulation to optimize the quantization effect of activation tensors (output[s]) using clipping by analyzing both the Gaussian and the Laplace priors. This formulation is henceforth refered [sic] to as Analytical Clipping for Integer Quantization (ACIQ).” (Banner, page 2, paragraph 2)
“Commonly (e.g., in GEMMLOWP), integer tensors are uniformly quantized in the range
[
-
α
,
α
]
, where
α
is determined by the tensor maximal absolute value. In the following we show that the this [sic] choice of
α
is suboptimal, and suggest a model where the tensor values (output[s]) are clipped to reduce quantization noise. For any
x
∈
R
(size of an output), we define the clipping function
c
l
i
p
(
x
,
α
)
as follows
PNG
media_image1.png
84
482
media_image1.png
Greyscale
” (Banner, page 4, paragraph 4). To clip an activation tensor, the size of it must be measured and compared against alpha.
Examiner’s note: As one of ordinary skill in the art would know, each activation is calculated as a weighted sum (first operation) of weights (weight coefficient[s]) and inputs from the previous layer.
“For each method, we select N layers (tensors) to be quantized (second operation) to 4 bits using our optimized clipping method and compare it against the standard GEMMLOWP approach. In figure 3 we present this accuracy-quantization tradeoff.” (Banner, page 7, paragraph 3). The size of each activation must be measured for activations in N layers to be input into the clipping function.
PNG
media_image2.png
200
400
media_image2.png
Greyscale
(Banner, page 8, Figure 3). For N > 2, N - 2 intermediate layer activations are being clipped and quantized.
control the first operation by controlling the weight coefficient used in the first operation to adjust the size of the output based on the information and a quantization parameter used for the quantization:
“Commonly (e.g., in GEMMLOWP), integer tensors are uniformly quantized in the range
[
-
α
,
α
]
, where
α
(quantization parameter) is determined by the tensor maximal absolute value. In the following we show that the this [sic] choice of
α
is suboptimal, and suggest a model where the tensor values (output[s]) are clipped to reduce quantization noise. For any
x
∈
R
(the information), we define the clipping function
c
l
i
p
(
x
,
α
)
as follows
PNG
media_image1.png
84
482
media_image1.png
Greyscale
” (Banner, page 4, paragraph 4).
“the expected mean-square-error between X and its quantized version Q(X) can be written as follows:
PNG
media_image3.png
332
551
media_image3.png
Greyscale
” (Banner, page 4, paragraph 6). As can be seen in the formula above, the mean square error depends partly on (x -
α
). In other words, the value of the loss increases as the difference between the activation tensor size and quantization threshold increases.
Banner relates to quantizing neural networks and is analogous to the claimed invention.
While Banner fails to disclose the further limitations of the claim, Diril discloses [a] non-transitory computer-readable storage medium storing a program which, when executed by a computer comprising a processor and memory, causes the computer to: “The computer-readable medium containing the computer program may be loaded into computing system 810. All or a portion of the computer program stored on the computer-readable medium may then be stored in system memory 816 and/or various portions of storage devices 832 and 833. When executed by processor 814, a computer program loaded into computing system 810 may cause processor 814 to perform and/or be a means for performing the functions of one or more of the example embodiments described and/or illustrated herein.” (Diril, [0073]); “Examples of computer-readable media include, without limitation, transmission-type media, such as carrier waves, and non-transitory-type media, such as magnetic storage media ( e.g., hard disk drives, tape drives, and floppy disks), optical-storage media (e.g., Compact Disks (CDs), Digital Video Disks (DVDs), and BLU-RAY disks), electronic- storage media (e.g., solid-state drives and flash media), and other distribution systems.” (Diril, [0080])
Diril relates to quantizing neural networks and is analogous to the claimed invention. Banner teaches a method of quantizing neural networks. The claimed invention improves upon this method by storing it in the form of instructions on computer hardware. Diril teaches standard computer hardware, applicable to Banner. A person of ordinary skill in the art would have recognized that storing Banner’s method as computer instructions on Diril’s hardware would lead to the predictable result of the method being executable by a computing system, and would improve the known device by allowing it to be performed with real data (MPEP 2143 I. (D) Applying a known technique to a known device (method, or product) ready for improvement to yield predictable results).
While Diril fails to disclose the further limitations of the claim, Sasagawa discloses a method of controlling the first operation by controlling the weight coefficient used in the first operation to adjust the size of the output based on the information and a quantization parameter used for the quantization: “in the present disclosure, when an inference model is generated, training is performed using a loss function for optimization to which a regularization term is added. This regularization term prevents a weight parameter used by a neural network from becoming a weight parameter likely to change the accuracy of an inferred value. For example, in the present disclosure, when a neural network is trained, a weight parameter is updated so that a value of ‘loss function+regularization term (loss)’ becomes smaller, and an inference model is generated. Accordingly, even when a weight parameter is quantized at the time of mounting, it is possible to reduce a significant decrease in accuracy of an inferred value.” (Sasagawa, [0037])
Sasagawa relates to quantized neural networks and is analogous to the claimed invention. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the existing combination to adjust the weights based on loss during learning, as disclosed by Sasagawa. Doing so would avoid reductions in accuracy, even for quantized weights. See Sasagawa, [0037].
Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Banner (ACIQ: Analytical Clipping for Integer Quantization of neural networks, published 9/27/2018, ICLR 2019 Conference Blind Submission, retrieved from https://openreview.net/forum?id=B1x33sC9KQ) in view of Sasagawa (Neural Network Derivation Method, published 8/12/2021, US 2021/0248463 A1).
Regarding claim 10, Banner discloses [a]n information processing method comprising:
obtaining information indicating a size of an output as a result of a first operation in a neural network that performs the first operation using a weight coefficient for input data and a second operation of quantizing a result of the first operation, in order to obtain data of an intermediate layer:
“In Section 4, we provide a rigorous formulation to optimize the quantization effect of activation tensors (output[s]) using clipping by analyzing both the Gaussian and the Laplace priors. This formulation is henceforth refered [sic] to as Analytical Clipping for Integer Quantization (ACIQ).” (Banner, page 2, paragraph 2)
“Commonly (e.g., in GEMMLOWP), integer tensors are uniformly quantized in the range
[
-
α
,
α
]
, where
α
is determined by the tensor maximal absolute value. In the following we show that the this [sic] choice of
α
is suboptimal, and suggest a model where the tensor values (output[s]) are clipped to reduce quantization noise. For any
x
∈
R
(size of an output), we define the clipping function
c
l
i
p
(
x
,
α
)
as follows
PNG
media_image1.png
84
482
media_image1.png
Greyscale
” (Banner, page 4, paragraph 4). To clip an activation tensor, the size of it must be measured and compared against alpha.
Examiner’s note: As one of ordinary skill in the art would know, each activation is calculated as a weighted sum (first operation) of weights (weight coefficient[s]) and inputs from the previous layer.
“For each method, we select N layers (tensors) to be quantized (second operation) to 4 bits using our optimized clipping method and compare it against the standard GEMMLOWP approach. In figure 3 we present this accuracy-quantization tradeoff.” (Banner, page 7, paragraph 3). The size of each activation must be measured for activations in N layers to be input into the clipping function.
PNG
media_image2.png
200
400
media_image2.png
Greyscale
(Banner, page 8, Figure 3). For N > 2, N - 2 intermediate layer activations are being clipped and quantized.
controlling the first operation by controlling the weight coefficient used in the first operation to adjust the size of the output based on the information and a quantization parameter used for the quantization:
“Commonly (e.g., in GEMMLOWP), integer tensors are uniformly quantized in the range
[
-
α
,
α
]
, where
α
(quantization parameter) is determined by the tensor maximal absolute value. In the following we show that the this [sic] choice of
α
is suboptimal, and suggest a model where the tensor values (output[s]) are clipped to reduce quantization noise. For any
x
∈
R
(the information), we define the clipping function
c
l
i
p
(
x
,
α
)
as follows
PNG
media_image1.png
84
482
media_image1.png
Greyscale
” (Banner, page 4, paragraph 4).
“the expected mean-square-error between X and its quantized version Q(X) can be written as follows:
PNG
media_image3.png
332
551
media_image3.png
Greyscale
” (Banner, page 4, paragraph 6). As can be seen in the formula above, the mean square error depends partly on (x -
α
). In other words, the value of the loss increases as the difference between the activation tensor size and quantization threshold increases.
Banner relates to quantizing neural networks and is analogous to the claimed invention.
While Banner fails to disclose the further limitations of the claim, Sasagawa discloses a method of controlling the first operation by controlling the weight coefficient used in the first operation to adjust the size of the output based on the information and a quantization parameter used for the quantization: “in the present disclosure, when an inference model is generated, training is performed using a loss function for optimization to which a regularization term is added. This regularization term prevents a weight parameter used by a neural network from becoming a weight parameter likely to change the accuracy of an inferred value. For example, in the present disclosure, when a neural network is trained, a weight parameter is updated so that a value of “loss function+regularization term (loss)” becomes smaller, and an inference model is generated. Accordingly, even when a weight parameter is quantized at the time of mounting, it is possible to reduce a significant decrease in accuracy of an inferred value.” (Sasagawa, [0037])
Sasagawa relates to quantized neural networks and is analogous to the claimed invention. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the primary reference to adjust the weights based on loss during learning, as disclosed by Sasagawa. Doing so would avoid reductions in accuracy, even for quantized weights. See Sasagawa, [0037].
Response to Arguments
The following responses address arguments and remarks made in the instant remarks dated 06/22/2026.
Objections
Previous objections to the specification and claims have been withdrawn in light of the instant amendments.
Claim Interpretation
In light of the instant amendments, the claims are no longer interpreted under 35 U.S.C. 112(f).
112 Rejections
In light of the instant amendments, previous rejections under 35 U.S.C. 112(a) and 112(b) have been withdrawn.
101 Rejections
On page 9 of the instant remarks, the Applicant requests that rejections under 35 U.S.C. 101 be withdrawn in light of the instant amendments:
“Claims 5-7 are rejected under 35 U.S.C. § 101 as allegedly being directed to non-statutory
subject matter without significantly more. Specifically the Office Action states that claim 5 "recites
a mathematical concept in the form of an equation with a term that scales directly with size of the
output exceeding the quantization parameter," that claim 6, "can be performed as a mental
process," and that claim 7, "can be performed as a mental process." (Office Action, p. 10-16.)
Without conceding the propriety of the rejection, the claims are amended in a manner
believed to overcome the rejections under 35 U.S.C. § 101.
Therefore, Applicant respectfully requests that the rejections of the claims under 35 U.S.C.
§ 101 be withdrawn.”
The Examiner respectfully maintains that claims 5-7 are directed to judicial exceptions which aren’t integrated into a practical application, and said claims don’t amount to significantly more than these exceptions. For example, claim 5 recites “the loss is calculated by a loss function including a regularization item that is large when the size of the output exceeds the quantization parameter”, a mathematical concept, which is not practically integrated through the additional elements of the claims. Similar reasoning is applicable to dependent claims 6-7.
See the 101 rejections section for more detail. No rejections are withdrawn on this basis.
102 / 103 Rejections
On pages 9-11 of the instant remarks, the Applicant argues that the relied upon prior art fails to disclose the amended claims:
“However, nowhere does Banner teach or suggest the above-noted features
of amended independent claim 1. Banner merely discloses analytical clipping for integer
quantization of neural networks by clipping tensor values to reduce quantization noise and then
quantizing the clipped tensor values. Banner's configuration is limited to selecting a clipping value
based on tensor statistics and limiting the tensor values themselves when the tensor values exceed
the clipping value. Although Banner explains that such clipping may reduce quantization noise,
that disclosure concerns replacing an out-of-range tensor value with a threshold value before
quantization. Banner does not disclose controlling a first operation in a neural network by
controlling a weight coefficient used in the first operation to adjust the size of the output based on
information indicating the size of the output and a quantization parameter used for the quantization
Rather, Banner's clipping function is merely a post-operation limitation applied to tensor values
before quantization, and does not control the weight coefficient used to produce the output of the
first operation. Therefore, Banner fails to teach or suggest the above-noted features of amended
independent claim 1. Thus, amended independent claim 1 is patentable over Banner.
Diril and Sasagawa are not relied on to cure, nor do they cure, the deficiencies of Banner
with respect to the above-noted features of amended independent claim 1.”
Regarding the Applicant’s arguments above, the Examiner respectfully disagrees.
Regarding “control the first operation by controlling the weight coefficient used in the first operation to adjust the size of the output based on the information and a quantization parameter used for the quantization”, Banner discloses clipping activation tensors outside a quantization threshold, and adjusting the network via loss such that activation tensors outside this threshold are discouraged (Banner, page 4, paragraphs 4-6). While Banner fails to explicitly disclose adjusting network weights, Sasagawa discloses adjusting the weights in a neural network such that a regularization term in the loss formula is minimized (Sasagawa, [0037]). Combining Sasagawa and Banner would yield the benefit of decreasing reductions in accuracy as measured by the loss, even for quantized weights, as made clear by Sasagawa ([0037]).
Regarding “at least one processor; and at least one memory having stored thereon instructions which, when executed by the at least one processor, cause the information processing apparatus to”, Diril discloses computer processors that can execute program instructions stored in memory (Diril, [0057-0058]).
Similar arguments are applicable to the substantially similar other amended independent claims. See the 103 rejections section for more detail. No rejections are withdrawn on these grounds.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
Zhao et al. (Improving Neural Network Quantization without Retraining using Outlier Channel Splitting, published 2019, arXiv:1901.09504v1) discloses methods of quantizing neural networks using various value clipping techniques.
Jantscher et al. (ERROR COMPENSATION IN ANALOG NEURAL NETWORKS, filed 2020, US 20220309331 A1) discloses a method of constructing a neural network loss based on activation clipping thresholds.
Nagel et al. (A White Paper on Neural Network Quantization, published 2021, arXiv:2106.08295v1) discloses a myriad of methods for quantizing neural networks, including clipping errors.
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Aaron P Gormley whose telephone number is (571)272-1372. The examiner can normally be reached Monday - Friday 12:00 PM - 8:00 PM EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michelle T Bechtold can be reached at (571) 431-0762. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/AG/Examiner, Art Unit 2148 /Ryan Barrett/Primary Examiner, Art Unit 2148