DETAILED ACTION
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This action is responsive to the Amendment filed on 6/5/2026. Claims 1-11 are pending in the case. Claims 1, 8, and 11 are independent claims.
Response to Arguments
Applicant’s amendments regarding 35 U.S.C. § 101 rejections are persuasive. These rejections are respectfully withdrawn.
Applicant’s prior art arguments have been considered but are moot because the new grounds of rejection presented below do not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the arguments.
Claim Rejections - 35 U.S.C. § 112
The following is a quotation of 35 U.S.C. § 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
Claims 1-11 are rejected under 35 U.S.C. § 112(b) as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor regards as the invention. Each independent claim recites “low-bit,” which renders the claims indefinite. The term “low-bit” is not defined by the claims, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. For the purposes of prior art and subject matter eligibility analyses Examiner assumes any number of bits may be considered low. Dependent claims inherit the same issue from parent claims and do not resolve it.
Claim Rejections - 35 U.S.C. § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. §§ 102 and 103 (or as subject to pre-AIA 35 U.S.C. §§ 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. § 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1 and 8-11 are rejected under 35 U.S.C. § 102(a)(1) as being anticipated by Liu et al. (US 2021/0065011 A1, hereinafter Liu1).
As to independent claim 1, Liu1 discloses a training method for a low-bit quantized neural network model including at least one quantization layer, the method comprising:
quantizing, in a forward transfer process (“The forward propagation process may include the quantization process of the weight of arbitrary bits and the feature map,” paragraph 0057 lines 8-10), a network parameter represented by a continuous real value into a quantized value, and calculating a quantization error generated by the quantizing (“there is a quantization error
m
i
n
V
~
t
∈
F
Y
t
-
Ỹ
t
when the weight w of a high-accuracy floating-point type is quantized to wb of a low-accuracy fixed-point type (Wherein
Υ
t
+
1
=
V
t
+
1
V
t
η
l
+
1
η
l
, Ỹ is calculated in the same manner as Υt+1, except that Υt+1 is based on a fully refined network and Ỹ is based on a quantized network), which results in a difference between the gradient g of the weight w and the gradient ĝ of the weight wb,” paragraph 0043 lines 3-12);
determining, in a backward transfer process (“Conversely, if the difference value between the actual output result and the expected output result of the neural network model exceeds a predetermined threshold, the back propagation process needs to be continued, that is, based on the difference value between the actual output result and the expected output result, the operation is performed layer by layer in the neural network model from bottom to top, and the weights in the model are updated so that the performance of the network model with updated weights is closer to the expected performance,” paragraph 0057 lines 17-27), an original gradient of a weight in a low-bit quantized neural network model including at least one quantization layer (“the constraint threshold range of the gradient is set,” paragraph 0051 lines 16-17 – the range is set based on the original gradient);
generating, in the backward transfer process, a corrected gradient by correcting the original gradient of the weight using the quantization error calculated in the forward transfer process, wherein the correcting comprises correcting a magnitude of the original gradient using the quantization error and correcting a direction of the original gradient using the quantization error (“when the gradient of the low-accuracy weights is distorted due to quantization errors, the distorted gradient is constrained to be within the set constraint threshold range, the distortion of the gradient is corrected and thus the trained network model can achieve the expected performance,” paragraph 0051 lines 19-24); and
updating the low-bit quantized neural network model according to the corrected gradient (“the distortion of the gradient is corrected and thus the trained network model can achieve the expected performance,” paragraph 0051 lines 22-24).
As to independent claim 8, Liu1 discloses an apparatus for training a low-bit quantized neural network model including at least one quantization layer, the apparatus comprising:
one or more storage media (“The internal memory 31 includes a random access memory (RAM), a read-only memory (ROM), and the like. The RAM can be used as a main memory, a work area, and the like of the processor 30. The ROM may be used to store control program of the processor 30, and may also be used to store files or other data to be used when the control program is operated,” paragraph 0129 lines 3-9); and
one or more processors (“The processor 30 may be a CPU or a GPU, which is used to perform overall control on the training apparatus,” paragraph 0129 lines 1-2), wherein the one or more processors and the one or more storage media are configured to
quantize, in a forward transfer process (“The forward propagation process may include the quantization process of the weight of arbitrary bits and the feature map,” paragraph 0057 lines 8-10), a network parameter represented by a continuous real value into a quantized value, and calculating a quantization error generated by the quantizing (“there is a quantization error
m
i
n
V
~
t
∈
F
Y
t
-
Ỹ
t
when the weight w of a high-accuracy floating-point type is quantized to wb of a low-accuracy fixed-point type (Wherein
Υ
t
+
1
=
V
t
+
1
V
t
η
l
+
1
η
l
, Ỹ is calculated in the same manner as Υt+1, except that Υt+1 is based on a fully refined network and Ỹ is based on a quantized network), which results in a difference between the gradient g of the weight w and the gradient ĝ of the weight wb,” paragraph 0043 lines 3-12);
determine, in a backward transfer process (“Conversely, if the difference value between the actual output result and the expected output result of the neural network model exceeds a predetermined threshold, the back propagation process needs to be continued, that is, based on the difference value between the actual output result and the expected output result, the operation is performed layer by layer in the neural network model from bottom to top, and the weights in the model are updated so that the performance of the network model with updated weights is closer to the expected performance,” paragraph 0057 lines 17-27), an original gradient of a weight in a low-bit quantized neural network model including at least one quantization layer (“the constraint threshold range of the gradient is set,” paragraph 0051 lines 16-17 – the range is set based on the original gradient);
generate, in the backward transfer process, a corrected gradient by correcting the original gradient of the weight using the quantization error calculated in the forward transfer process, wherein the correcting comprises correcting a magnitude of the original gradient using the quantization error and correcting a direction of the original gradient using the quantization error (“when the gradient of the low-accuracy weights is distorted due to quantization errors, the distorted gradient is constrained to be within the set constraint threshold range, the distortion of the gradient is corrected and thus the trained network model can achieve the expected performance,” paragraph 0051 lines 19-24); and
update the low-bit quantized neural network model according to the corrected gradient (“the distortion of the gradient is corrected and thus the trained network model can achieve the expected performance,” paragraph 0051 lines 22-24).
As to dependent claim 9, Liu1 further discloses a method comprising:
receiving a data set corresponding to a requirement of a task that the low-bit quantized neural network model is capable of performing (“Taking a case where a network model trained according to the manner in the first exemplary embodiment is stored in a security camera as an example, it is assumed that the security camera is to execute a target detection application, after the security camera takes a picture as a data set, the taken picture is input into the network model,” paragraph 0130 lines 4-10);
performing operations on the data set in layers from top to bottom in the low-bit quantized neural network model (“that the picture is calculated in each layer from the top to the bottom in the network model,” paragraph 0130 lines 10-11); and
outputting a result (“the target detection result is output,” paragraph 0130 line 12).
As to dependent claim 10, Liu1 further discloses an apparatus wherein the one or more processors and the one or more storage media are further configured to:
receive a data set corresponding to a requirement of a task that the low-bit quantized neural network model is capable of performing (“Taking a case where a network model trained according to the manner in the first exemplary embodiment is stored in a security camera as an example, it is assumed that the security camera is to execute a target detection application, after the security camera takes a picture as a data set, the taken picture is input into the network model,” paragraph 0130 lines 4-10);
perform operations on the data set in layers from top to bottom in the low-bit quantized neural network model (“that the picture is calculated in each layer from the top to the bottom in the network model,” paragraph 0130 lines 10-11); and
output a result (“the target detection result is output,” paragraph 0130 line 12).
As to independent claim 11, Liu1 discloses a non-transitory computer-readable storage medium (“The internal memory 31 includes a random access memory (RAM), a read-only memory (ROM), and the like. The RAM can be used as a main memory, a work area, and the like of the processor 30. The ROM may be used to store control program of the processor 30, and may also be used to store files or other data to be used when the control program is operated,” paragraph 0129 lines 3-9) storing instructions that, when executed by a computer, cause the computer to perform operations comprising:
quantizing, in a forward transfer process (“The forward propagation process may include the quantization process of the weight of arbitrary bits and the feature map,” paragraph 0057 lines 8-10), a network parameter represented by a continuous real value into a quantized value, and calculating a quantization error generated by the quantizing (“there is a quantization error
m
i
n
V
~
t
∈
F
Y
t
-
Ỹ
t
when the weight w of a high-accuracy floating-point type is quantized to wb of a low-accuracy fixed-point type (Wherein
Υ
t
+
1
=
V
t
+
1
V
t
η
l
+
1
η
l
, Ỹ is calculated in the same manner as Υt+1, except that Υt+1 is based on a fully refined network and Ỹ is based on a quantized network), which results in a difference between the gradient g of the weight w and the gradient ĝ of the weight wb,” paragraph 0043 lines 3-12);
determining, in a backward transfer process (“Conversely, if the difference value between the actual output result and the expected output result of the neural network model exceeds a predetermined threshold, the back propagation process needs to be continued, that is, based on the difference value between the actual output result and the expected output result, the operation is performed layer by layer in the neural network model from bottom to top, and the weights in the model are updated so that the performance of the network model with updated weights is closer to the expected performance,” paragraph 0057 lines 17-27), an original gradient of a weight in a low-bit quantized neural network model including at least one quantization layer (“the constraint threshold range of the gradient is set,” paragraph 0051 lines 16-17 – the range is set based on the original gradient);
generating, in the backward transfer process, a corrected gradient by correcting the original gradient of the weight using the quantization error calculated in the forward transfer process, wherein the correcting comprises correcting a magnitude of the original gradient using the quantization error and correcting a direction of the original gradient using the quantization error (“when the gradient of the low-accuracy weights is distorted due to quantization errors, the distorted gradient is constrained to be within the set constraint threshold range, the distortion of the gradient is corrected and thus the trained network model can achieve the expected performance,” paragraph 0051 lines 19-24); and
updating the low-bit quantized neural network model according to the corrected gradient (“the distortion of the gradient is corrected and thus the trained network model can achieve the expected performance,” paragraph 0051 lines 22-24).
Claim Rejections - 35 U.S.C. § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. §§ 102 and 103 (or as subject to pre-AIA 35 U.S.C. §§ 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. § 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 C.F.R. § 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. § 102(b)(2)(C) for any potential 35 U.S.C. § 102(a)(2) prior art against the later invention.
Claims 2-7 are rejected under 35 U.S.C. § 103 as being unpatentable over Liu1 in view of Liu et al. (US 2020/0394523 A1, hereinafter Liu2).
As to dependent claim 2, the rejection of claim 1 is incorporated.
Liu1 does not appear to expressly teach a method wherein in the quantizing,
a quantization step is determined according to a quantization interval and a quantization bit width,
the continuous real value is mapped to a discrete quantization value, and
the discrete quantization value is limited in a range that is representable by the quantization bit width.
Liu2 teaches a method wherein in the quantizing,
a quantization step is determined according to a quantization interval (“The formula (1) shows that when the data to be quantized is quantized by using the quantization parameter corresponding to the first situation, a quantization interval is 2s and is marked as C,” paragraph 0073 lines 16-19) and a quantization bit width (“a maximum value A of a floating-point number may be represented by an n-bit fixed-point number as 2s(2n-1−1), then a maximum value in a number field of the data to be quantized may be represented by an n-bit fixed-point number as 2s(2n-1−1), and a minimum value in the number field of the data to be quantized may be represented by an n-bit fixed-point number as −2s(2n-1−1),” paragraph 0073 lines 9-16),
the continuous real value is mapped to a discrete quantization value (“the following formula (1) may be used to quantize the data to obtain quantized data Ix:
I
x
=
r
o
u
n
d
F
x
2
s
,” paragraph 0072 lines 2-5), and
the discrete quantization value is limited in a range that is representable by the quantization bit width (“a maximum value A of a floating-point number may be represented by an n-bit fixed-point number as 2s(2n-1−1), then a maximum value in a number field of the data to be quantized may be represented by an n-bit fixed-point number as 2s(2n-1−1), and a minimum value in the number field of the data to be quantized may be represented by an n-bit fixed-point number as −2s(2n-1−1),” paragraph 0073 lines 9-16).
Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the quantization of Liu1 to comprise the interval and width of Liu2. (1) The Examiner finds that the prior art included each claim element listed above, although not necessarily in a single prior art reference, with the only difference between the claimed invention and the prior art being the lack of actual combination of the elements in a single prior art reference. (2) The Examiner finds that one of ordinary skill in the art could have combined the elements as claimed by known development methods, and that in combination, each element merely performs the same function as it does separately. (3) The Examiner finds that one of ordinary skill in the art would have recognized that the results of the combination were predictable, namely quantization according to an interval and bit width (“a maximum value A of a floating-point number may be represented by an n-bit fixed-point number as 2s(2n-1−1), then a maximum value in a number field of the data to be quantized may be represented by an n-bit fixed-point number as 2s(2n-1−1), and a minimum value in the number field of the data to be quantized may be represented by an n-bit fixed-point number as −2s(2n-1−1),” Liu2 paragraph 0073 lines 9-16). Therefore, the rationale to support a conclusion that the claim would have been obvious is that the combining prior art elements according to known methods to yield predictable results to one of ordinary skill in the art. See MPEP § 2143(I)(A).
As to dependent claim 3, the rejection of claim 1 is incorporated.
Liu1 does not appear to expressly teach a method wherein in generating the corrected gradient,
an updated value corresponding to a discrete quantization value is calculated,
the direction of the original gradient is corrected according to the updated value of the discrete quantization value and the continuous real value, and
the magnitude of the original gradient is corrected according to the quantization error in the forward transfer process of the low-bit quantized neural network model.
Liu2 teaches a method wherein in generating the corrected gradient,
an updated value corresponding to a discrete quantization value is calculated (“The quantization error diffbit is determined according to the pre-quantized data and the corresponding quantized data, and the quantization error diffbit is compared with the threshold to obtain a comparison result,” paragraph 0143 lines 7-10),
the direction of the original gradient is corrected (“perform back propagation on a gradient, which means to propagate the gradient from the output layer to the hidden layer, and finally to the input layer, and sequentially adjust weights and biases of each layer in the neural network according to the gradient,” paragraph 0053 lines 5-9) according to the updated value of the discrete quantization value and the continuous real value (“The quantization error diffbit is determined according to the pre-quantized data and the corresponding quantized data, and the quantization error diffbit is compared with the threshold to obtain a comparison result,” paragraph 0143 lines 7-10), and
the magnitude of the original gradient is corrected (“perform back propagation on a gradient, which means to propagate the gradient from the output layer to the hidden layer, and finally to the input layer, and sequentially adjust weights and biases of each layer in the neural network according to the gradient,” paragraph 0053 lines 5-9) according to the quantization error (“The quantization error diffbit is determined according to the pre-quantized data and the corresponding quantized data, and the quantization error diffbit is compared with the threshold to obtain a comparison result,” paragraph 0143 lines 7-10) in the forward transfer process of the low-bit quantized neural network model (“perform a forward processing on a signal, which means to transmit the signal from the input layer to the output layer through the hidden layer,” paragraph 0053 lines 2-4).
Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the quantization of Liu1 to comprise the calculation of Liu2. (1) The Examiner finds that the prior art included each claim element listed above, although not necessarily in a single prior art reference, with the only difference between the claimed invention and the prior art being the lack of actual combination of the elements in a single prior art reference. (2) The Examiner finds that one of ordinary skill in the art could have combined the elements as claimed by known development methods, and that in combination, each element merely performs the same function as it does separately. (3) The Examiner finds that one of ordinary skill in the art would have recognized that the results of the combination were predictable, namely calculating the corrected quantization (“The quantization error diffbit is determined according to the pre-quantized data and the corresponding quantized data, and the quantization error diffbit is compared with the threshold to obtain a comparison result,” Liu2 paragraph 0143 lines 7-10). Therefore, the rationale to support a conclusion that the claim would have been obvious is that the combining prior art elements according to known methods to yield predictable results to one of ordinary skill in the art. See MPEP § 2143(I)(A).
As to dependent claim 4, the rejection of claim 3 is incorporated. Liu1/Liu2 further teaches a method wherein if a direction of the continuous real value pointing to the updated value of the discrete quantization value is consistent with the direction of the original gradient, then in the generating the corrected gradient, a direction of the corrected gradient is set to an opposite direction of the original gradient, otherwise the direction of the corrected gradient is set to the direction of the original gradient (“perform back propagation on a gradient, which means to propagate the gradient from the output layer to the hidden layer, and finally to the input layer, and sequentially adjust weights and biases of each layer in the neural network according to the gradient,” Liu2 paragraph 0053 lines 5-9).
As to dependent claim 5, the rejection of claim 3 is incorporated. Liu1/Liu2 further teaches a method wherein if the direction of the original gradient is positive and the continuous real value is less than the discrete quantization value while being greater than the updated value of the discrete quantization value, or if the direction of the original gradient is negative and the continuous real value is greater than the discrete quantization value while being less than the updated value of the discrete quantization value, then the magnitude of the original gradient is reduced, wherein the direction of the original gradient is positive when a value of the original gradient is positive (“perform back propagation on a gradient, which means to propagate the gradient from the output layer to the hidden layer, and finally to the input layer, and sequentially adjust weights and biases of each layer in the neural network according to the gradient,” Liu2 paragraph 0053 lines 5-9).
As to dependent claim 6, the rejection of claim 3 is incorporated. Liu1/Liu2 further teaches a method wherein if the direction of the original gradient is positive and the continuous real value is greater than the discrete quantization value while also being greater than the updated value of the discrete quantization value, or if the direction of the original gradient is negative and the continuous real value is less than the discrete quantization value while also being less than the updated value of the discrete quantization value, then the magnitude of the original gradient is increased (“perform back propagation on a gradient, which means to propagate the gradient from the output layer to the hidden layer, and finally to the input layer, and sequentially adjust weights and biases of each layer in the neural network according to the gradient,” Liu2 paragraph 0053 lines 5-9).
As to dependent claim 7, the rejection of claim 5 is incorporated. Liu1/Liu2 further teaches a method wherein in correcting the original gradient, the calculated quantization error is scaled, and the original gradient of the weight is corrected based on the scaled quantization error (“perform back propagation on a gradient, which means to propagate the gradient from the output layer to the hidden layer, and finally to the input layer, and sequentially adjust weights and biases of each layer in the neural network according to the gradient,” Liu2 paragraph 0053 lines 5-9).
Conclusion
The prior art made of record and not relied upon is considered pertinent to Applicant’s disclosure:
Kepp (“Quantized Back Propagation for Efficient Deep Learning,” 19 October 2018, https://www.kevinkepp.de/assets/docs/masters-thesis.pdf) disclosing gradient correction for low-bit quantized training
Karimireddy et al. (“Error Feedback Fixes SignSGD and other Gradient Compression Schemes,” 29 May 2019, https://arxiv.org/abs/1901.09847 https://proceedings.mlr.press/v97/karimireddy19a/karimireddy19a.pdf) disclosing gradient correction for low-bit quantized training
Wolfe (“Quantized Training with Deep Networks,” 30 August 2022, https://cameronrwolfe.substack.com/p/quantized-training-with-deep-networks-82ea7f516dc6) disclosing gradient correction for low-bit quantized training
Applicant is required under 37 C.F.R. § 1.111(c) to consider these references fully when responding to this action.
Applicant’s amendments necessitated the new grounds of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 C.F.R. § 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 C.F.R. § 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action.
In the interests of compact prosecution, Applicant is invited to contact the examiner via electronic media pursuant to USPTO policy outlined MPEP § 502.03. All electronic communication must be authorized in writing. Applicant may wish to file an Internet Communications Authorization Form PTO/SB/439. Applicant may wish to request an interview using the Interview Practice website: http://www.uspto.gov/patent/laws-and-regulations/interview-practice.
Applicant is reminded Internet e-mail may not be used for communication for matters under 35 U.S.C. § 132 or which otherwise require a signature. A reply to an Office action may NOT be communicated by Applicant to the USPTO via Internet e-mail. If such a reply is submitted by Applicant via Internet e-mail, a paper copy will be placed in the appropriate patent application file with an indication that the reply is NOT ENTERED. See MPEP § 502.03(II).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Ryan Barrett whose telephone number is 571 270 3311. The examiner can normally be reached 9:00am to 5:30pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, Applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor Michelle Bechtold can be reached at 571 431 0762. The fax phone number for the organization where this application or proceeding is assigned is 571 273 8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Ryan Barrett/
Primary Examiner, Art Unit 2148