DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Objections
Claims 1-10 are objected to because of the following informalities: The claim(s) are narrative in form and replete with indefinite language. The structure must be organized and correlated in such a manner as to present a complete understandable claim. The claim(s) must be in one sentence form only. For example, “Current total bit operation counts are computed based on this single-layer sensitivity until reaching the preset maximum bit operation counts. Meanwhile, quantized layers and precision are recorded, determining the mixed-precision quantization strategy”, the claim is not clear where the middle of the limitation is ended and then starts with “meanwhile”. The dependent claims have similar issues, where the dependent claims do not clearly write what is further limiting. Appropriate corrections are required.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that use the word “means” or “step” but are nonetheless not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph because the claim limitation(s) do not recite(s) sufficient structure, materials, or acts to entirely perform the recited function. Such claim limitation(s) is/are:
“Module M1: Perform uniform bit-width high-precision quantization…” (claim 6).
“Module M2: Each layer in the neural network undergoes post-training quantization…” (claim 6).
“Module M3: Single-layer sensitivity is computed based on the baseline inference accuracy…” (claim 6).
“Module M4: Current total bit operation counts are computed…” (claim 6).
For an analysis of the structure, material, or acts corresponding to the claimed functions, see rejection under 35 USC § 112(b) infra.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant wishes to provide further explanation or dispute the examiner’s interpretation of the corresponding structure, applicant must identify the corresponding structure with reference to the specification by page and line number, and to the drawing, if any, by reference characters in response to this Office action.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 6-10 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor, or for pre-AIA the applicant regards as the invention.
This application includes one or more claim limitations that use the word “means” or “step” but are nonetheless not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph because the claim limitation(s) do not recite(s) sufficient structure, materials, or acts to entirely perform the recited function. Such claim limitation(s) is/are:
“Module M1: Perform uniform bit-width high-precision quantization…” (claim 6).
“Module M2: Each layer in the neural network undergoes post-training quantization…” (claim 6).
“Module M3: Single-layer sensitivity is computed based on the baseline inference accuracy…” (claim 6).
“Module M4: Current total bit operation counts are computed…” (claim 6).
However, the written description fails to disclose the corresponding structure, material, or acts for performing the entire claimed function and to clearly link the structure, material, or acts to the function.
The closest paragraph is ¶ 68, “alternatively, the modules for implementing various functions can be regarded as both software programs for implementing methods and structures within hardware components.”
Therefore, the claim is indefinite and is rejected under 35 U.S.C. 112(b) or pre-AIA 35 U.S.C. 112, second paragraph. For the purpose of examination, any computer capable of performing the claimed functions reads on the claims.
Applicant may:
(a) Amend the claim so that the claim limitation will no longer be interpreted as a limitation under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph;
(b) Amend the written description of the specification such that it expressly recites what structure, material, or acts perform the entire claimed function, without introducing any new matter (35 U.S.C. 132(a)); or
(c) Amend the written description of the specification such that it clearly links the structure, material, or acts disclosed therein to the function recited in the claim, without introducing any new matter (35 U.S.C. 132(a)).
If applicant is of the opinion that the written description of the specification already implicitly or inherently discloses the corresponding structure, material, or acts and clearly links them to the function so that one of ordinary skill in the art would recognize what structure, material, or acts perform the claimed function, applicant should clarify the record by either:
(a) Amending the written description of the specification such that it expressly recites the corresponding structure, material, or acts for performing the claimed function and clearly links or associates the structure, material, or acts to the claimed function, without introducing any new matter (35 U.S.C. 132(a)); or
(b) Stating on the record what the corresponding structure, material, or acts, which are implicitly or inherently set forth in the written description of the specification, perform the claimed function. For more information, see 37 CFR 1.75(d) and MPEP §§ 608.01(o) and 2181.
Claims 7-10 are rejected as they are being directly or indirectly dependent on rejected claim 6.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Tsuji et al. (“Greedy Search Algorithm for Mixed Precision in Post-Training Quantization of Convolutional Neural Network Inspired by Submodular Optimization”, Proceedings of Machine Learning Research 157, 2021) in view Wang et al. (“HAQ: Hardware-Aware Automated Quantization with Mixed Precision”, 2019 IEEE/CVF).
Regarding claim 1.
Tsuji teaches a hardware-aware mixed-precision quantization method based on the greedy search is characterized (see abstract, “we propose a quantization scheme that considers the effects by continuously updating the accuracy after each layer quantization. Additionally, for more data compression, we extend that scheme to mixed precision, which applies a layer-by-layer fitted bit-width…we derive practical solutions to the bit allocation problem in polynomial time O(N2) using a deterministic greedy search algorithm inspired by submodular optimization without any training.”) in that comprises:
Step S1: Perform uniform bit-width high-precision quantization on all layers of the neural network, conduct training-aware quantization, and acquire the training model, baseline inference accuracy, and total bit operation counts (see page 7 section 3.3 greedy search algorithm, “Figure 2 is simple: an algorithm that greedily selects a quantization layer of smaller bit-width (4-bit or 6-bit) at a time until quantization of all layers is completed. We set 4-bit or 6-bit in advance for the smaller bit-width used in the algorithm as step 0, according to quantization tolerance of the target model. The first step is to prepare a pre-trained FP32 model and quantize all its layers to INT8. The second step is to calculate (involve inference) the objective function in Equation 2 for all 8-bit quantized layers, select one layer with the largest value, and quantize it to smaller bit-width.”);
Step S2: Each layer in the neural network undergoes post-training quantization with low precision individually, and the corresponding inference accuracy and total bit operation counts for each layer are recorded (see page 8, algorithm 1,
PNG
media_image1.png
338
582
media_image1.png
Greyscale
);
Step S3: Single-layer sensitivity is computed based on the baseline inference accuracy, total bit operation counts, as well as the inference accuracy and total bit operation counts corresponding to each layer (see page 10, section 4.1, “the accuracy degradation because of quantization is calculated only once in advance as the sensitivity (no update) of each layer. The accuracy is maintained by quantizing the layers in order of descending value. We evaluated the quantization performance of the algorithm by extending this one-shot search method to our experimental conditions and by comparing it with our proposed method. Comparison between the flowcharts of the two algorithms is presented in Figure 3.”, also see page 7, section 3.3, “The second step is to calculate (involve inference) the objective function in Equation 2 for all 8-bit quantized layers, select one layer with the largest value, and quantize it to smaller bit-width. This second step is iterated until all layers have been quantized to smaller bit-width. In other words, the layers are selected in order of quantization efficiency as defined by the objective function in Equation 2. Details of this proposed algorithm are presented in Algorithm 1. By plotting the inference accuracy after quantizing the selected layer to the smaller bit-width at each step, we eventually obtain an improved tradeoff between quantization progress as the model size and accuracy.”);
Step S4: Current total bit operation counts are computed based on this single-layer sensitivity until reaching the preset maximum bit operation counts. Meanwhile, quantized layers and precision are recorded, determining the mixed-precision quantization strategy (see page 7, section 3.3, “Users can choose desirable quantization architectures of bit allocation among the models shown for each computational resource.”, see page 8, algorithm 1,
PNG
media_image1.png
338
582
media_image1.png
Greyscale
also see page 10, section 4.1, “the accuracy degradation because of quantization is calculated only once in advance as the sensitivity (no update) of each layer. The accuracy is maintained by quantizing the layers in order of descending value. We evaluated the quantization performance of the algorithm by extending this one-shot search method to our experimental conditions and by comparing it with our proposed method. Comparison between the flowcharts of the two algorithms is presented in Figure 3.”).
Tsuji do not specifically teach Each layer in the neural network undergoes post-training quantization with low precision individually, and the corresponding inference accuracy and total bit operation counts for each layer are recorded.
Wang teaches Each layer in the neural network undergoes post-training quantization with low precision individually, and the corresponding inference accuracy and total bit operation counts for each layer are recorded (see abstract, “Rather than relying on proxy signals such as FLOPs and model size, we employ a hardware simulator to generate direct feedback signals (latency and energy) to the RL agent.”, also see 8606, “Figure 2: An overview of our Hardware-Aware Automated Quantization (HAQ) framework. We leverage the reinforcement learning to automatically search over the huge quantization design space with hardware in the loop. The agent propose an optimal bitwidth allocation policy given the amount of computation resources (i.e., latency, power, and model size). Our RL agent integrates the hardware accelerator into the exploration loop so that it can obtain the direct feedback from the hardware, instead of relying on indirect proxy signals.”).
Both Tsuji and Wang pertain to the problem of Quantization with Mixed Precision, thus being analogous. It would have been obvious to one skilled in the art before the effective filing date of the claimed invention to combine Tsuji and Wang to teach the above limitations. The motivation for doing so would be “the Hardware-Aware Automated Quantization (HAQ) framework which leverages the reinforcement learning to automatically determine the quantization policy, and we take the hardware accelerator’s feedback in the design loop. Rather than relying on proxy signals such as FLOPs and model size, we employ a hardware simulator to generate direct feedback signals (latency and energy) to the RL agent. Compared with conventional methods, our framework is fully automated and can specialize the quantization policy for different neural network architectures and hardware architectures. Our framework effectively reduced the latency by 1.4-1.95× and the energy consumption by 1.9× with negligible loss of accuracy compared with the fixed bitwidth (8 bits) quantization. Our framework reveals that the optimal policies on different hardware architectures (i.e., edge and cloud architectures) under different resource constraints (i.e., latency, energy and model size) are drastically different.” (see Wang abstract).
Regarding claim 2.
Tsuji and Wang teaches the method of claim 1,
Tsuji further teaches the hardware-aware mixed-precision quantization method based on greedy search, as described in claim 1, is characterized by individually quantizing each neural network layer with single-layer low precision after training; The other layers remain unchanged during the single-layer low-precision quantization of the current layer (see page 10, section 4.1, “the accuracy degradation because of quantization is calculated only once in advance as the sensitivity (no update) of each layer. The accuracy is maintained by quantizing the layers in order of descending value. We evaluated the quantization performance of the algorithm by extending this one-shot search method to our experimental conditions and by comparing it with our proposed method. Comparison between the flowcharts of the two algorithms is presented in Figure 3.”, also see page 7, section 3.3, “The second step is to calculate (involve inference) the objective function in Equation 2 for all 8-bit quantized layers, select one layer with the largest value, and quantize it to smaller bit-width. This second step is iterated until all layers have been quantized to smaller bit-width. In other words, the layers are selected in order of quantization efficiency as defined by the objective function in Equation 2. Details of this proposed algorithm are presented in Algorithm 1. By plotting the inference accuracy after quantizing the selected layer to the smaller bit-width at each step, we eventually obtain an improved tradeoff between quantization progress as the model size and accuracy.”).
Regarding claim 4.
Tsuji and Wang teaches the method of claim 1,
Tsuji further teaches the hardware-aware mixed-precision quantization method based on greedy search, according to claim 1, is characterized in that Step S4 comprises: Sorting the sensitivity of each layer from high to low, sequentially conducting low-precision quantization on each layer according to the sorted results, and calculating the current total bit operation counts until the current total bit operation counts reach the preset maximum bit operation counts. Recording the currently quantized layers and their corresponding quantization precision to determine the optimal mixed-precision quantization strategy (see page 7, section 3.3, “The second step is to calculate (involve inference) the objective function in Equation 2 for all 8-bit quantized layers, select one layer with the largest value, and quantize it to smaller bit-width. This second step is iterated until all layers have been quantized to smaller bit-width. In other words, the layers are selected in order of quantization efficiency as defined by the objective function in Equation 2. Details of this proposed algorithm are presented in Algorithm 1. By plotting the inference accuracy after quantizing the selected layer to the smaller bit-width at each step, we eventually obtain an improved tradeoff between quantization progress as the model size and accuracy.”).
Wang also teaches sorting in page 8607, “The feedback is directly obtained from the hardware accelerator, which we will discuss in Section 3.3. If the current policy exceeds our resource budget (on latency, energy or model size), we will sequentially decrease the bitwidth of each layer until the constraint is finally satisfied”.
The motivation utilized in the combination of claim 1, super, applies equally as well to claim 4.
Regarding claim 5.
Tsuji and Wang teaches the method of claim 4,
Tsuji further teaches The hardware-aware mixed-precision quantization method, based on greedy search according to claim 4, is characterized in that the preset maximum total bit operation counts are set based on the actual maximum bit operation counts allowed by the hardware platform (see page 7, section 3.3, “Users can choose desirable quantization architectures of bit allocation among the models shown for each computational resource.”, see page 8, algorithm 1,
PNG
media_image1.png
338
582
media_image1.png
Greyscale
also see page 10, section 4.1, “the accuracy degradation because of quantization is calculated only once in advance as the sensitivity (no update) of each layer. The accuracy is maintained by quantizing the layers in order of descending value. We evaluated the quantization performance of the algorithm by extending this one-shot search method to our experimental conditions and by comparing it with our proposed method. Comparison between the flowcharts of the two algorithms is presented in Figure 3.”).
Claims 6-7 and 9-10 recites a system to perform the method recited in claims 1-2 and 4-5. Therefore the rejection of claims 1-2 and 4-5 above applies equally here.
Allowable Subject Matter
Claims 3 and 8 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims and overcoming 112b rejection above.
Related prior arts:
Bijalwan et al. (US 20230281423 A1) teaches generating a mixed precision quantization model for performing image processing. The method comprises receiving a validation dataset of images to train a neural network model. The method comprises for each image of the validation dataset, generating a union sensitivity list, selecting a group of layers, generating a mixed precision quantization model by quantizing the selected group of layers into a high precision format; computing accuracy of the mixed precision quantization model for comparison with a target accuracy; in response to determining the accuracy is less than the target accuracy, generating another mixed precision model by selecting a next group of layers and computing the accuracy. In response to determining the accuracy is greater than or equal to the target accuracy, storing the mixed precision quantization model as a final mixed precision quantization model for image processing.
VASYLTSOV et al. (US 20230297836 A1) teaches generate, based on a determination of sensitivity of layers in a model to be trained, sensitivity results, and train the model by applying quantization to a layer of the layers with a low sensitivity of the sensitivity results lower than a predetermined threshold. Quantization is one of the approaches to optimizing a predetermined DNN model to be executed on predetermined hardware, and mixed-precision quantization may be one of the most promising types of quantization for optimizing a DNN model.
LU et al. (US 20210089922 A1) teaches joint optimization of mixed-precision bit-width quantization and structured pruning. For ease of explanation mixed-precision bit-width quantization may also be referred to as mixed-precision quantization. a general metric may be specified to measure compression performance. In one configuration, a bit-operations (BOPs) count is specified to measure compression performance. Aspects of the present disclosure are not limited to the BOPs count and may be modified to optimize for specific hardware.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to IMAD M KASSIM whose telephone number is (571)272-2958. The examiner can normally be reached 10:30AM-5:30PM, M-F (E.S.T.).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michael J. Huntley can be reached at (303) 297 - 4307. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/IMAD KASSIM/Primary Examiner, Art Unit 2129