DETAILED ACTION
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This action is responsive to the original application filed on 12/15/2023. Acknowledgment is made with respect to a claim of priority to Provisional Application 62/880,475 filed on 7/30/2019. This application is a Continuation Application of Issued Patent No. 11,847,568.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 23, 27, 28, 29, 33, 37, 38, and 39 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claims 23 and 33 recite the limitations “wherein modifying the plurality of parameters of the neural network comprises backpropagating the computed loss” and “wherein the set of instructions for modifying the plurality of parameters of the neural network comprises a set of instructions for backpropagating the computed loss” (emphasis added). There is insufficient antecedent basis for the claimed “the computed loss” in both claims. Further, it appears that claims 23 and 33 should depend on dependent claims 22 and 32, respectively, to provide for antecedent basis for the “modifying” and “the computed loss”. For examination purposes, claim 23 will be interpreted to mean “The method of claim [[21]] 22, wherein modifying the plurality of parameters of the neural network comprises backpropagating the computed loss” and claim 33 will be interpreted to mean “The method of claim [[31]] 32, wherein modifying the plurality of parameters of the neural network comprises backpropagating the computed loss” (emphasis added). Appropriate correction is required.
Claims 27 and 37 recite the limitation “wherein the quantized values are one of 4-bit values and 8-bit values” (emphasis added). Claims 21 and 31 use “quantized values” twice for two different things: (a) “the parameters are defined as quantized values in the trained neural network” (weights), and (b) “a particular range of quantized values defined by a neural network inference circuit” (activation range). Claims 27 and 37 both recite “the quantized values” without specifying which antecedent they refer back to. It is unclear whether the 4-bit/8-bit limitation applies to the trained weight parameters, the quantized activations, or both. Please explain. For examination purposes, the limitation will be interpreted to mean “wherein the quantized values of the parameters are one of 4-bit values and 8-bit values” (emphasis added). Appropriate correction is required.
Claims 28 and 38 recite the limitation “wherein the floating point values are stored using a variably-positioned binary point and the quantized values use a fixed binary point position” (emphasis added). Claims 21 and 31 use “quantized values” twice for two different things: (a) “the parameters are defined as quantized values in the trained neural network” (weights), and (b) “a particular range of quantized values defined by a neural network inference circuit” (activation range). Claims 28 and 38 both recite “the quantized values” without specifying which antecedent they refer back to. It is unclear whether the fixed-binary point limitation applies to the trained weight parameters, the quantized activations, or both. Please explain. For examination purposes, the limitation will be interpreted to mean “wherein the floating point values are stored using a variably-positioned binary point and the quantized values of the parameters use a fixed binary point position” (emphasis added). Appropriate correction is required.
Claims 29 and 39 recite the limitations “the plurality of parameters comprises a plurality of weights for computing dot products in convolutional layers of the neural network; and training the neural network comprises constraining the set of weights to ternary values” (emphasis added). There is insufficient antecedent basis for the claimed “the set of weights” in both claims. For examination purposes, claims 29 and 39 will be interpreted to mean “the plurality of parameters comprises a plurality of weights for computing dot products in convolutional layers of the neural network; and training the neural network comprises constraining the [[set]] plurality of weights to ternary values” (emphasis added). Appropriate correction is required.
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/process/file/efs/guidance/eTD-info-I.jsp.
Claim 21 is non-provisionally rejected on the ground of nonstatutory double patenting as being anticipated by claim 1 of U.S. Patent No. 11,847,568. That is, claim 1 of U.S. Patent No. 11,847,568 discloses a method that comprises receiving a definition of floating-point neural network, selecting a set of scaling and shift values to be applied to input values or activations, training a neural network using the selected values, and executing a neural network inference circuit.
Claim 22 is non-provisionally rejected on the ground of nonstatutory double patenting as being anticipated by claim 1 of U.S. Patent No. 11,847,568. That is, claim 1 of U.S. Patent No. 11,847,568 discloses propagating values through the network, computing a loss, and modifying parameters of the network using the loss.
Claim 23 is non-provisionally rejected on the ground of nonstatutory double patenting as being anticipated by claim 1 of U.S. Patent No. 11,847,568. That is, claim 1 of U.S. Patent No. 11,847,568 discloses backpropagating a loss.
Claim 24 is non-provisionally rejected on the ground of nonstatutory double patenting as being anticipated by claims 1 and 2 of U.S. Patent No. 11,847,568. That is, claims 1 and 2 of U.S. Patent No. 11,847,568 disclose determining constraints based on a type of computation.
Claim 25 is non-provisionally rejected on the ground of nonstatutory double patenting as being anticipated by claims 1, 2, and 4 of U.S. Patent No. 11,847,568. That is, claims 1, 2, and 4 of U.S. Patent No. 11,847,568 disclose the element-wise multiplication layer, selecting scaling and shift values, and that the first and second shift values are equal to zero.
Claim 26 is non-provisionally rejected on the ground of nonstatutory double patenting as being anticipated by claims 1, 2, and 3 of U.S. Patent No. 11,847,568. That is, claims 1, 2, and 3 of U.S. Patent No. 11,847,568 disclose the element-wise addition layer, selecting scaling and shift values, and that the first and second scaling values are equal.
Claim 27 is non-provisionally rejected on the ground of nonstatutory double patenting as being anticipated by claims 1, 6, and 7 of U.S. Patent No. 11,847,568. That is, claims 1, 6, and 7 of U.S. Patent No. 11,847,568 disclose that the quantized values are 8- or 4-bit values.
Claim 28 is non-provisionally rejected on the ground of nonstatutory double patenting as being anticipated by claims 1, 8, and 9 of U.S. Patent No. 11,847,568. That is, claims 1, 8, and 9 of U.S. Patent No. 11,847,568 disclose storing using variably positioned binary point and using a fixed binary point position.
Claim 29 is non-provisionally rejected on the ground of nonstatutory double patenting as being anticipated by claims 1 and 10 of U.S. Patent No. 11,847,568. That is, claims 1 and 10 of U.S. Patent No. 11,847,568 disclose the weights for computing dot products and constraining weights to ternary values.
Claim 30 is non-provisionally rejected on the ground of nonstatutory double patenting as being anticipated by claim 1 of U.S. Patent No. 11,847,568. That is, claim 1 of U.S. Patent No. 11,847,568 discloses selecting scaling and shift values to be applied to input or activation values.
Claim 31 is non-provisionally rejected on the ground of nonstatutory double patenting as being anticipated by claim 14 of U.S. Patent No. 11,847,568. That is, claim 14 of U.S. Patent No. 11,847,568 discloses a non-transitory computer-readable medium that comprises receiving a definition of floating-point neural network, selecting a set of scaling and shift values to be applied to input values or activations, training a neural network using the selected values, and executing a neural network inference circuit.
Claim 32 is non-provisionally rejected on the ground of nonstatutory double patenting as being anticipated by claim 14 of U.S. Patent No. 11,847,568. That is, claim 14 of U.S. Patent No. 11,847,568 discloses propagating values through the network, computing a loss, and modifying parameters of the network using the loss.
Claim 33 is non-provisionally rejected on the ground of nonstatutory double patenting as being anticipated by claim 14 of U.S. Patent No. 11,847,568. That is, claim 14 of U.S. Patent No. 11,847,568 discloses backpropagating a loss.
Claim 34 is non-provisionally rejected on the ground of nonstatutory double patenting as being anticipated by claims 14 and 15 of U.S. Patent No. 11,847,568. That is, claims 14 and 15 of U.S. Patent No. 11,847,568 disclose determining constraints based on a type of computation.
Claim 37 is non-provisionally rejected on the ground of nonstatutory double patenting as being anticipated by claims 14, 18, and 20 of U.S. Patent No. 11,847,568. That is, claims 14, 18, and 20 of U.S. Patent No. 11,847,568 disclose that the quantized values are 8- or 4-bit values.
Claim 40 is non-provisionally rejected on the ground of nonstatutory double patenting as being anticipated by claim 14 of U.S. Patent No. 11,847,568. That is, claim 14 of U.S. Patent No. 11,847,568 discloses selecting scaling and shift values to be applied to input or activation values.
This is a non-provisional double patenting rejection because the claims of the parent case (Issued Patent No. 11,847,568) of the present application have in fact been patented.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 21-40 are rejected under 35 U.S.C 101 because the claimed invention is directed to an abstract idea without significantly more. The analysis of the claims will follow the 2019 Revised Patent Subject Matter Eligibility Guidance, 84 Fed. Reg. 50 (“2019 PEG”).
When considering subject matter eligibility under 35 U.S.C. 101, it must be determined whether the claim is directed to one of the four statutory categories of invention, i.e., process, machine, manufacture, or composition of matter (Step 1). If the claim does fall within one of the statutory categories, the second step in the analysis is to determine whether the claim is directed to a judicial exception (Step 2A). The Step 2A analysis is broken into two prongs. In the first prong (Step 2A, Prong 1), it is determined whether or not the claims recite a judicial exception (e.g., mathematical concepts, mental processes, certain methods of organizing human activity). If it is determined in Step 2A, Prong 1 that the claims recite a judicial exception, the analysis proceeds to the second prong (Step 2A, Prong 2), where it is determined whether or not the claims integrate the judicial exception into a practical application. If it is determined at step 2A, Prong 2 that the claims do not integrate the judicial exception into a practical application, the analysis proceeds to determining whether the claim is a patent-eligible application of the exception (Step 2B). If an abstract idea is present in the claim, any element or combination of elements in the claim must be sufficient to ensure that the claim integrates the judicial exception into a practical application, or else amounts to significantly more than the abstract idea itself.
Claim 21
Step 1: The claim recites a method; therefore, it is directed to the statutory category of a process.
Step 2A Prong 1: The claim recites, inter alia:
based on a distribution of input activation values for a particular layer of the neural network, selecting a set of scaling and shift values to be applied to input activation values of the layer in order for the input activation values to match a particular range of quantized values defined by a neural network inference circuit for which the neural network is trained: Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mental process of selecting scaling and shift values to be applied to activation values to match a range of quantized values, which is an evaluation or observation that is practically capable of being performed in the human mind with the assistance of pen and paper. For example, one can practically and mentally select or determine values that match a range of other values that are then used for further downstream processing.
Step 2A Prong 2: The claim does not recite any additional limitations which integrate the abstract idea into a practical application. Specifically, the additional elements consist of “a neural network inference circuit”, “receiving a definition of a neural network comprising a plurality of parameters at a plurality of layers, the parameters defined as floating point values”, “training the neural network using the selected set of scaling and shift values, wherein the parameters are defined as quantized values in the trained neural network”, and “generating program instructions for the neural network inference circuit to execute the trained neural network with the quantized parameter values”.
The additional elements of “a neural network inference circuit” and “generating program instructions for the neural network inference circuit to execute the trained neural network with the quantized parameter values” amount to generic computer components used as a tool to perform an existing process. The additional element of “training the neural network using the selected set of scaling and shift values, wherein the parameters are defined as quantized values in the trained neural network” amounts to reciting only the idea of a solution or outcome i.e., the claim fails to recite details of how a solution to a problem is accomplished because it is not clear how the generic neural network is broadly trained using scaling and shift values. Thus, the additional elements amount to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
The additional element “receiving a definition of a neural network comprising a plurality of parameters at a plurality of layers, the parameters defined as floating point values” is insignificant extra-solution activity required for any uses of the abstract ideas (see MPEP § 2106.05(g)).
Thus, even when viewed individually and as an ordered combination, these additional elements do not integrate the abstract idea into a practical application, and the claim is thus directed to the abstract idea.
Step 2B: Finally, the claim taken as a whole does not contain an inventive concept which provides significantly more than the abstract idea.
The additional elements of “a neural network inference circuit” and “generating program instructions for the neural network inference circuit to execute the trained neural network with the quantized parameter values” amount to generic computer components used as a tool to perform an existing process. The additional element of “training the neural network using the selected set of scaling and shift values, wherein the parameters are defined as quantized values in the trained neural network” amounts to reciting only the idea of a solution or outcome i.e., the claim fails to recite details of how a solution to a problem is accomplished because it is not clear how the generic neural network is broadly trained using scaling and shift values. Thus, the additional elements amount to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
The additional element “receiving a definition of a neural network comprising a plurality of parameters at a plurality of layers, the parameters defined as floating point values” is insignificant extra-solution activity required for any uses of the abstract ideas (see MPEP § 2106.05(g), and is a well-understood, routine, conventional activity (see MPEP § 2106.05(d)(II)(i); “Receiving or transmitting data over a network”).
Taken alone or in combination, the additional elements of the claim do not provide an inventive concept and thus the claim is subject-matter ineligible.
Claim 22
Step 1: A process, as above.
Step 2A Prong 1: The claim recites, inter alia:
propagating sets of input values through the neural network to generate sets of output values, said propagation comprising applying the selected set of scaling and shift values to input activation values of the particular layer: Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mathematical concept of propagating values through a network or model, which is performed through mathematical computation as evidenced by paragraph [0088] of the originally filed specification (“The forward and backward propagation, in some embodiments, use the approximate quantization function and affine transformations to calculate the output and gradients as described above in relation to FIG. 3”).
computing a loss based on comparing the generated sets of output values to expected sets of output values for the sets of input values: Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mathematical concept of calculating a loss or gradient based on a comparison, which is performed through mathematical computation as evidenced by paragraph [0088] of the originally filed specification (“The forward and backward propagation, in some embodiments, use the approximate quantization function and affine transformations to calculate the output and gradients as described above in relation to FIG. 3”).
modifying the plurality of parameters of the neural network based on the computed loss: Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mathematical concept of modifying parameters or backwards propagating based on a loss or using gradients, which is performed through mathematical computation as evidenced by paragraph [0088] of the originally filed specification (“The forward and backward propagation, in some embodiments, use the approximate quantization function and affine transformations to calculate the output and gradients as described above in relation to FIG. 3”).
Step 2A Prong 2, Step 2B: The claim does not recite any additional elements that are sufficient to integrate the judicial exceptions into a practical application or amount to significantly more than the judicial exception. As such, the claim is ineligible.
Claim 23
Step 1: A process, as above.
Step 2A Prong 1: The claim recites, inter alia:
wherein modifying the plurality of parameters of the neural network comprises backpropagating the computed loss: Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mathematical concept of backpropagating a loss to modify parameters, which is performed through mathematical computation as evidenced by paragraph [0088] of the originally filed specification (“The forward and backward propagation, in some embodiments, use the approximate quantization function and affine transformations to calculate the output and gradients as described above in relation to FIG. 3”).
Step 2A Prong 2, Step 2B: The claim does not recite any additional elements that are sufficient to integrate the judicial exceptions into a practical application or amount to significantly more than the judicial exception. As such, the claim is ineligible.
Claim 24
Step 1: A process, as above.
Step 2A Prong 1: The claim recites, inter alia:
wherein selecting the set of scaling and shift values comprises determining a set of constraints on the set of scaling and shift values based on a type of computation performed by the particular layer: Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mental process of determining constraints on scaling and shift values, which is an evaluation or observation that is practically capable of being performed in the human mind with the assistance of pen and paper.
Step 2A Prong 2, Step 2B: The claim does not recite any additional elements that are sufficient to integrate the judicial exceptions into a practical application or amount to significantly more than the judicial exception. As such, the claim is ineligible.
Claim 25
Step 1: A process, as above.
Step 2A Prong 1: The claim recites, inter alia:
selecting the set of scaling and shift values comprises selecting (i) a first scaling value and first shift value to be applied to the input activation values from a first previous layer and (ii) a second scaling value and second shift value to be applied to the input activation values from the second previous layer: Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mental process of selecting scaling and shift values, which is an evaluation or observation that is practically capable of being performed in the human mind with the assistance of pen and paper.
Step 2A Prong 2, Step 2B: The additional elements of “the particular layer is an element-wise multiplication layer with input activation values from two previous layers of the neural network … the set of constraints comprises a constraint that the first and second shift values are equal to zero” amount to no more than generally linking the use of a judicial exception to a particular technological environment or field of use (see MPEP § 2106.05(h). Taken alone or in combination, the additional elements of the claim do not provide an inventive concept, integrate the abstract ideas into a practical application, or provide significantly more than the abstract ideas of the claim and thus the claim is subject-matter ineligible.
Claim 26
Step 1: A process, as above.
Step 2A Prong 1: The claim recites, inter alia:
selecting the set of scaling and shift values comprises selecting (i) a first scaling value and first shift value to be applied to the input activation values from a first previous layer and (ii) a second scaling value and second shift value to be applied to the input activation values from the second previous layer: Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mental process of selecting scaling and shift values, which is an evaluation or observation that is practically capable of being performed in the human mind with the assistance of pen and paper.
Step 2A Prong 2, Step 2B: The additional elements of “the particular layer is an element-wise addition layer with input activation values from two previous layers of the neural network… the set of constraints comprises a constraint that the first scaling value equals the second scaling value” amount to no more than generally linking the use of a judicial exception to a particular technological environment or field of use (see MPEP § 2106.05(h). Taken alone or in combination, the additional elements of the claim do not provide an inventive concept, integrate the abstract ideas into a practical application, or provide significantly more than the abstract ideas of the claim and thus the claim is subject-matter ineligible.
Claim 27
Step 1: A process, as above.
Step 2A Prong 1: The claim recites the abstract ideas of the preceding claims from which it depends.
Step 2A Prong 2, Step 2B: The additional element of “wherein the quantized values are one of 4-bit values and 8-bit values” amounts to no more than generally linking the use of a judicial exception to a particular technological environment or field of use (see MPEP § 2106.05(h). Taken alone or in combination, the additional elements of the claim do not provide an inventive concept, integrate the abstract ideas into a practical application, or provide significantly more than the abstract ideas of the claim and thus the claim is subject-matter ineligible.
Claim 28
Step 1: A process, as above.
Step 2A Prong 1: The claim recites the abstract ideas of the preceding claims from which it depends.
Step 2A Prong 2, Step 2B: The additional element of “wherein the floating point values are stored using a variably-positioned binary point and the quantized values use a fixed binary point position” amounts to no more than generally linking the use of a judicial exception to a particular technological environment or field of use (see MPEP § 2106.05(h). Taken alone or in combination, the additional elements of the claim do not provide an inventive concept, integrate the abstract ideas into a practical application, or provide significantly more than the abstract ideas of the claim and thus the claim is subject-matter ineligible.
Claim 29
Step 1: A process, as above.
Step 2A Prong 1: The claim recites, inter alia:
training the neural network comprises constraining the set of weights to ternary values: Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mental process of constraining weights to ternary values, such as -1, 0, and 1, which is an evaluation or observation that is practically capable of being performed in the human mind with the assistance of pen and paper.
Step 2A Prong 2, Step 2B: The additional element of “the plurality of parameters comprises a plurality of weights for computing dot products in convolutional layers of the neural network” amounts to no more than generally linking the use of a judicial exception to a particular technological environment or field of use (see MPEP § 2106.05(h). Taken alone or in combination, the additional elements of the claim do not provide an inventive concept, integrate the abstract ideas into a practical application, or provide significantly more than the abstract ideas of the claim and thus the claim is subject-matter ineligible.
Claim 30
Step 1: A process, as above.
Step 2A Prong 1: The claim recites, inter alia:
for each respective layer of the plurality of layers, selecting respective scaling and shift values to be applied to input activation values of the respective layer based on a respective distribution of the input activation values for the respective layer: Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mental process of selecting scaling and shift values based on a distribution, which is an evaluation or observation that is practically capable of being performed in the human mind with the assistance of pen and paper.
Step 2A Prong 2, Step 2B: The claim does not recite any additional elements that are sufficient to integrate the judicial exceptions into a practical application or amount to significantly more than the judicial exception. As such, the claim is ineligible.
Claims 31-40
Claims 31-40 recite a non-transitory machine-readable medium (step 1: a manufacture) using a processing unit and program to perform the steps of claims 21-30, respectively, which by MPEP 2106.05(f) (“apply it”) cannot integrate an abstract idea into a practical application or provide significantly more than the abstract idea by itself, and are thus rejected for the same reasons set forth in the rejection of claims 21-30, respectively.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 21-24, 27, 28, 30-34, 37, 38, and 40 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Jacob et al. (Jacob et al., “Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference”, Dec. 15, 2017, arXiv:1712.05877v1, pp. 1-14, hereinafter “Jacob”).
Regarding claim 21, Jacob discloses [a] method for quantizing a neural network, the method comprising: (Abstract; “We propose a quantization scheme that allows inference to be carried out using integer-only arithmetic, which can be implemented more efficiently than floating point inference on commonly available integer-only hardware”)
receiving a definition of a neural network comprising a plurality of parameters at a plurality of layers, the parameters defined as floating point values; (§3; “A common approach to training quantized networks is to train in floating point and then quantize the resulting weights (sometimes with additional post-quantization training for fine-tuning). We found that this approach works sufficiently well for large models with considerable representational capacity”, which discloses a training graph of a floating point model that is functionally equivalent to receiving a network definition whose weights or parameters at each layer are ordinary floating point values; and Algorithm 1, Step 1; “Create a training graph of the floating-point model”, which confirms that a floating point neural network is received and subsequently quantized)
based on a distribution of input activation values for a particular layer of the neural network, selecting a set of scaling and shift values to be applied to input activation values of the layer in order for the input activation values to match a particular range of quantized values defined by a neural network inference circuit for which the neural network is trained; (§3.1; “For activations, ranges depend on the inputs to the network. To estimate the ranges, we collect [a;b] ranges seen on activations during training and then aggregate them via exponential moving averages (EMA) with the smoothing parameter being close to 1 so that observed ranges are smoothed across thousands of training steps”, which discloses, for activations, the [a;b] value ranges are observed on a layer’s activations during training and then these values are aggregated using an exponential moving average or activation-value distribution; and Equation 13; the equation discloses a relationship between a distribution-derived range of activation values and the scale “S” and the zero point “Z”, which are the claimed scale and shift values, respectively; and §1; “We provide a quantized inference framework that is efficiently implementable on integer-arithmetic-only hard ware such as the Qualcomm Hexagon (sections 2.2, 2.3), and we describe an efficient, accurate implementation on ARMNEON(AppendixB)”, which discloses a neural network inference circuit for which the neural network is trained; and §2.1; and Equation 2)
training the neural network using the selected set of scaling and shift values, wherein the parameters are defined as quantized values in the trained neural network; (§3, Equation 12, and Algorithm 1, Steps 2-3; the section discloses inserting “fake quantization” operations that implement the point-wise quantization function of equation 12, and then trains “in simulated quantized mode until convergence” using the selected scale (scaling) and zero-point (shift) parameters, thereby producing a trained network whose weights are quantized values)
generating program instructions for the neural network inference circuit to execute the trained neural network with the quantized parameter values (Algorithm 1, Steps 4-5; the algorithm creates and optimizes “the inference graph for running in a low bit inference engine” and runs inference using that quantized inference graph; and §2.4; the section describes the corresponding fused-layer implementation compiled for “ARMand x86CPU architectures”).
Regarding claim 31, it is a non-transitory machine-readable medium claim corresponding to the steps of claim 21 and is rejected for the same reasons as claim 21.
Regarding claims 22 and 32, the rejection of claims 21 and 31 are incorporated and Jacob further discloses propagating sets of input values through the neural network to generate sets of output values, said propagation comprising applying the selected set of scaling and shift values to input activation values of the particular layer; computing a loss based on comparing the generated sets of output values to expected sets of output values for the sets of input values; and modifying the plurality of parameters of the neural network based on the computed loss (Algorithm 1, Step 3; “Train in simulated quantized mode until convergence”; and §3; “We propose an approach that simulates quantization effects in the forward pass of training. Backpropagation still happens as usual, and all weights and biases are stored in floating point so that they can be easily nudged by small amounts. The forward propagation pass however simulates quantized inference as it will happen in the inference engine, by implementing in floating-point arithmetic the rounding behavior of the quantization scheme that we introduced in section 2”).
Regarding claims 23 and 33, the rejection of claims 21 and 31 are incorporated and Jacob further discloses wherein modifying the plurality of parameters of the neural network comprises backpropagating the computed loss (Algorithm 1, Step 3; “Train in simulated quantized mode until convergence”; and §3; “We propose an approach that simulates quantization effects in the forward pass of training. Backpropagation still happens as usual, and all weights and biases are stored in floating point so that they can be easily nudged by small amounts”).
Regarding claims 24 and 34, the rejection of claims 21 and 31 are incorporated and Jacob further discloses wherein selecting the set of scaling and shift values comprises determining a set of constraints on the set of scaling and shift values based on a type of computation performed by the particular layer (§2.4, and Equation 11; “In order to have the quantized bias-addition be the addition of an int32 bias into this int32 accumulator, the bias-vector is quantized such that: it uses int32 as its quantized data type; it uses 0 as its quantization zero-point Zbias; and its quantization scale Sbias is the same as that of the accumulators, which is the product of the scales of the weights and of the input activations”; and §3; “Weights are quantized before they are convolved with the input. If batch normalization (see [17]) is used for the layer, the batch normalization parameters are “folded into” the weights before quantization, see section 3.2. • Activations are quantized at points where they would be during inference, e.g. after the activation function is ap plied to a convolutional or fully connected layer’s output, or after a bypass connection adds or concatenates the out puts of several layers together such as in ResNets.”).
Regarding claims 27 and 37, the rejection of claims 21 and 31 are incorporated and Jacob further discloses wherein the quantized values are one of 4-bit values and 8-bit values (§2.1; “For 8-bit quantization, q is quantized as an 8-bit integer (for B-bit quantization, q is quantized as an B-bit integer). Some arrays, typically bias vectors, are quantized as 32-bit integers”; and §1).
Regarding claims 28 and 38, the rejection of claims 21 and 31 are incorporated and Jacob further discloses wherein the floating point values are stored using a variably-positioned binary point and the quantized values use a fixed binary point position (§2.1; “For 8-bit quantization, q is quantized as an 8-bit integer (for B-bit quantization, q is quantized as an B-bit integer). Some arrays, typically bias vectors, are quantized as 32-bit integers … The constant S (for “scale”) is an arbitrary positive real number. It is typically represented in software as a floating point quantity, like the real values r. Section 2.2 describes methods for avoiding the representation of such floating point quantities in the inference workload”; and §1).
Regarding claims 30 and 40, the rejection of claims 21 and 31 are incorporated and Jacob further discloses for each respective layer of the plurality of layers, selecting respective scaling and shift values to be applied to input activation values of the respective layer based on a respective distribution of the input activation values for the respective layer (§2.1; “A basic requirement of our quantization scheme is that it permits efficient implementation of all arithmetic using only integer arithmetic operations on the quantized values (we eschew implementations requiring lookup tables because these tend to perform poorly compared to pure arithmetic on SIM hardware). This is equivalent to requiring that the quantization scheme be an affine mapping of integers q to real numbers r, i.e. of the form r =S(q−Z) (1) for some constants S and Z. Equation (1) is our quantization scheme and the constants S and Z are our quantization parameters. Our quantization scheme uses a single set of quantization parameters for all values within each activations array and within each weights array; separate arrays use separate quantization parameters.”; and §3.1).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 29 and 39 are rejected under 35 USC § 103 as being obvious over Jacob in view of Li et al. (Li et al., “Ternary weight networks”, Nov. 19, 2016, arXiv:1605.04711v2, pp. 1-5, hereinafter “Li”).
Regarding claims 29 and 39, the rejection of claims 21 and 31 are incorporated and Jacob discloses the plurality of parameters comprises a plurality of weights for computing dot products in convolutional layers of the neural network; and (§3; “We propose an approach that simulates quantization effects in the forward pass of training. Backpropagation still happens as usual, and all weights and biases are stored in floating point so that they can be easily nudged by small amounts; and §1).
Jacob fails to explicitly disclose but Li discloses training the neural network comprises constraining the set of weights to ternary values (§2; “We address the limited storage and limited computational resources issues by introducing ternary weight networks (TWNs), which constrain the weights to be ternary-valued: +1, 0 and-1. TWNs seek to make a balance between the full precision weight networks (FPWNs) counterparts and the binary precision weight networks (BPWNs) counterparts”; and Abstract).
Jacob and Li are analogous art because both are concerned with neural network compression techniques. Before the effective filing date of the claimed invention, it would have been obvious to one skilled in neural network compression to combine the ternary value constraining of Li and the quantization technique of Jacob to yield to the predictable result of training the neural network comprises constraining the set of weights to ternary values. The motivation for doing so would be to provide for higher accuracy and lower computational requirements for neural networks (Li; §4 Conclusion).
Conclusion
Claims 25, 26, 35, and 36 have been searched, but no prior art was uncovered.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Brent Hoover whose telephone number is (303)297-4403. The examiner can normally be reached Monday - Friday 9-5 MST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Abdullah Kawsar can be reached at 571-270-3169. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/BRENT JOHNSTON HOOVER/ Primary Examiner, Art Unit 2127