DETAILED ACTION
This is a non-final, first office action on the merits. Claims 1-16 are pending. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). The certified copy has been filed in parent Application No. KR10-2022-0033175, filed on 03/17/2022.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(d):
(d) REFERENCE IN DEPENDENT FORMS.—Subject to subsection (e), a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers.
The following is a quotation of pre-AIA 35 U.S.C. 112, fourth paragraph:
Subject to the following paragraph [i.e., the fifth paragraph of pre-AIA 35 U.S.C. 112], a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers.
Claim 16 is rejected under 35 U.S.C. 112(d) or pre-AIA 35 U.S.C. 112, 4th paragraph, as being of improper dependent form for failing to further limit the subject matter of the claim upon which it depends, or for failing to include all the limitations of the claim upon which it depends. Claim 16 fails to further limit claim 1, from which it depends and is therefore an improper dependent claim. For example, claim 16 can be infringed without infringing claim 1 since there is no requirement in claim 16 that the method of claim 1 be actually performed. Applicant may cancel the claim(s), amend the claim(s) to place the claim(s) in proper dependent form, rewrite the claim(s) in independent form, or present a sufficient showing that the dependent claim(s) complies with the statutory requirements.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-16 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. Specifically, claims 1-16 are directed to an abstract idea without additional elements amounting to significantly more than the abstract idea.
With respect to Step 2A Prong One of the framework, claims 1, 9, and 16 recite an abstract idea. Claims 1, 9, and 16 include “determining a zone and an operand that uses a candidate data format; obtaining a first parameter gradient through a first simulation on input data by applying an original data format to the operand in the zone; obtaining a second parameter gradient through a second simulation on the input data by applying the candidate data format to the operand in the zone; and determining a performance indicator according to the candidate data format based on the first parameter gradient and the second parameter gradient”.
The limitations above recite an abstract idea under Step 2A Prong One. More particularly, the elements above recite mental processes-concepts performed in the human mind (including an observation, evaluation, judgment, opinion) because the elements describe a process for determining a performance indicator. As a result, claims 1, 9, and 16 recite an abstract idea under Step 2A Prong One.
Claims 2-8 and 10-15 further describe the process for determining a performance indicator. As a result, claims 2-8 and 10-15 recite an abstract idea under Step 2A Prong One for the same reasons as stated above with respect to claims 1, 9, and 16.
With respect to Step 2A Prong Two of the framework, claims 1, 9, and 16 do not include additional elements that integrate the abstract idea into a practical application. Claims 1, 9, and 16 include additional elements that do not recite an abstract idea under Step 2A Prong One. The additional elements of claims 1, 9, and 16 include an artificial neural network, a processor, a memory, instruction, a processor, and a computer-readable non-transitory recording medium. When considered in view of the claim as a whole, the additional elements do not integrate the abstract idea into a practical application because the additional computing elements are generic computing elements that are merely used as a tool to perform the recited abstract idea. As a result, claims 1, 9, and 16 do not include additional elements that integrate the abstract idea into a practical application under Step 2A Prong Two.
Claims 2-4 and 10 do not include any additional elements beyond those recited with respect to claims 1, 9, and 16. As a result, claims 2-4 and 10 do not include additional elements that integrate the abstract idea into a practical application under Step 2A Prong Two for the same reasons as stated above with respect to claims 1, 9, and 16.
Claims 5-8 and 11-15 include additional elements that do not recite an abstract idea under Step 2A Prong One. The additional elements of claims 5-8 and 11-15 include an artificial neural network and a processor. When considered in view of the claims as a whole, the additional elements do not integrate the abstract idea into a practical application because the additional computing elements do no more than generally link the use of the recited abstract idea to a particular technological environment. As a result, claims 5-8 and 11-15 do not include additional elements that integrate the abstract idea into a practical application under Step 2A Prong Two.
With respect to Step 2B of the framework, claims 1, 9, and 16 do not include additional elements amounting to significantly more than the abstract idea. As noted above, claims 1, 9, and 16 include additional elements that do not recite an abstract idea under Step 2A Prong One. The additional elements of claims 1, 9, and 16 include an artificial neural network, a processor, a memory, instruction, a processor, and a computer-readable non-transitory recording medium. The additional elements do not amount to significantly more than the abstract idea because the additional computing elements are generic computing elements that are merely used as a tool to perform the recited abstract idea. Further, looking at the additional elements as an ordered combination adds nothing that is not already present when considering the additional elements individually. As a result, independent claims 1, 9, and 16 do not include additional elements that amount to significantly more than the abstract idea under Step 2B.
Claims 2-4 and 10 do not include any additional elements beyond those recited with respect to claims 1, 9, and 16. As a result, claims 2-4 and 10 do not include additional elements that amount to significantly more than the abstract idea under Step 2B for the same reasons as stated above with respect to claims 1, 9, and 16.
Claims 5-8 and 11-15 include additional elements that do not recite an abstract idea under Step 2A Prong One. The additional elements of claims 5-8 and 11-15 include an artificial neural network and a processor. The additional elements do not amount to significantly more than the abstract idea because the additional computing elements do no more than generally link the use of the recited abstract idea to a particular technological environment. Further, looking at the additional elements as an ordered combination adds nothing that is not already present when considering the additional elements individually. As a result, claims 5-8 and 11-15 do not include additional elements that amount to significantly more than the abstract idea under Step 2B.
Therefore, the claims are directed to an abstract idea without additional elements amounting to significantly more than the abstract idea. Accordingly, claims 1-16 are rejected under 35 U.S.C. 101 as being directed to non-statutory subject matter.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention.
Claim 1-16 rejected under 35 U.S.C. 102(a)(1) as being anticipated by Sriram et al. (US Pub No. 2022/0044114) (hereinafter Sriram et al.), hereinafter Sriram.
Regarding claims 1, 9, and 16, Sriram discloses an artificial neural network performance prediction method according to data format which is performed by an artificial neural network performance prediction device including a processor (see Sriram, para [0214], wherein artificial intelligence functionality, such as by executing redundant and/or different neural networks; and para [0002], wherein processors or computing systems used to train neural networks using low-precision quantization), the artificial neural network performance prediction method comprising:
a memory storing at least one instruction; and a processor (see Sriram, Fig. 11);
a computer-readable non-transitory recording medium (see Sriram, para [0111]);
determining a zone and an operand of an artificial neural network that uses a candidate data format (see Sriram, (a "zone" construed as a data from a specific layer or region "zone" as the input for those calculations) para [0058], wherein quantizing all weights and activations of the network except for layers that require finer granularity in representation than the 8-bit quantization can provide (e.g., regression layers); para [0319], wherein processor 1505 to set up operands and access computation; and para [0061], wherein perform quantization on parameters while training a neural network to generate a model to run inference at a lower precision);
obtaining a first parameter gradient through a first simulation of the artificial neural network on input data by applying an original data format to the operand in the zone (see Sriram, para [0131], wherein bias values, gradient information, momentum values, and/or other parameter or hyperparameter information; paras [0062]-[0063], wherein training of a neural network 106, first applies QAT 108 to generate a first trained model 110 and then applies PTQ 112 on the first trained model 110 to output a second trained model 114 that has parameters (e.g., weights and activations) represented by low-bit integers (e.g., INTS data type)…..A training graph may be modified to simulate the lower precision behavior in the forward pass of the training process, and thus introduces the quantization errors as part of the training loss, which the optimizer tries to minimize during the training……modeling quantization errors during training which helps in maintaining accuracy as compared to floating-point 16 (FP16) or floating point 32 (FP32). In at least one embodiment, GPU 116, using one or more processors, first applies QAT 108, during training 106, to output a first trained model 110. The first trained model 110 may also be referred to herein as an intermediate trained model, a QAT trained model, QAT quantized model, or a QAT model); and para [0058], wherein quantizing all weights and activations of the network except for layers that require finer granularity in representation than the 8-bit quantization can provide (e.g., regression layers);
obtaining a second parameter gradient through a second simulation of the artificial neural network on the input data by applying the candidate data format to the operand in the zone (see Sriram, paras [0062]-[0063], wherein using one or more processors, applies PTQ 112 on the first trained model 110 to output a second trained model 114 that has both QAT and PTQ applied. The second trained model 114 may comprise parameters (e.g., weights and activations) that are represented by 8-bit integers. The process and steps of applying PTQ 112 to the first trained model 110 is described in more detail below. In some embodiments, GPU 116 performs QAT on one or more pre trained models for re-training using framework 100; and para [0058], wherein quantizing all weights and activations of the network except for layers that require finer granularity in representation than the 8-bit quantization can provide (e.g., regression layers)); and
determining a performance indicator according to the candidate data format based on the first parameter gradient and the second parameter gradient (see Sriram, paras [0057]-[0059], wherein transform (e.g., quantize) a model to have weights represented by lower-precision values (e.g., low-bit integers) instead of using higher-precision values (e.g., values with full floating point-precision) to conserve memory usage and reduces computation when the trained model is deployed; para [0131], wherein bias values, gradient information, momentum values, and/or other parameter or hyperparameter information; and para [0112], wherein using fewer bits to represent the parameters, a smaller die-area for dedicated hardware and/or less computations per second may be used to perform the same processing tasks, and thus reducing cost and power requirements. In an embodiment, accuracies of the model is improved by applying PTQ to a QAT model at 8-bit/lower bit quantization. The quantization may be performed during the DNN training. In some instances, the range of weights is typically limited due to regularization, which penalizes growth of weight magnitudes and their statistics and is known in each training iteration. Therefore, the weights can be quantized by finding the absolute maximum value of the weights).
Regarding claims 2 and 10, Sriram discloses the artificial neural network performance prediction method of claim 1, wherein the candidate data format includes at least one data format that is lower in precision than the original data format (see Sriram, para [0057], wherein transform (e.g., quantize) a model to have weights represented by lower-precision values (e.g., low-bit integers) instead of using higher-precision values (e.g., values with full floating point-precision) to conserve memory usage and reduces computation when the trained model is deployed).
Regarding claim 3, Sriram discloses the artificial neural network performance prediction method of claim 1, wherein the operand includes at least one of an activation value, an error indicating an activation gradient, and a weight gradient (see Sriram, para [0057], wherein quantizing models such as rounding the weights and activations to lower-bit integers).
Regarding claims 4 and 11, Sriram discloses the artificial neural network performance prediction method of claim 1, wherein the determining of the performance indicator includes:
determining a magnitude between the first parameter gradient and the second parameter gradient (see Sriram, para [0112], wherein using fewer bits to represent the parameters, a smaller die-area for dedicated hardware and/or less computations per second may be used to perform the same processing tasks, and thus reducing cost and power requirements. In an embodiment, accuracies of the model is improved by applying PTQ to a QAT model at 8-bit/lower bit quantization. The quantization may be performed during the DNN training. In some instances, the range of weights is typically limited due to regularization, which penalizes growth of weight magnitudes and their statistics and is known in each training iteration. Therefore, the weights can be quantized by finding the absolute maximum value of the weights); and
determining a misalignment between the first parameter gradient and the second parameter gradient (see Sriram, para [0591], wherein (i.e., gradient misalignment occurs when some bias in the gradient) initial model 4004 may have previously fine-tuned parameters (e.g., weights and/or biases) that remain from prior training, so training or retraining 3614 may not take as long or require as much processing as training a model from scratch. In at least one embodiment, during model training 3614, by having reset or replaced output or loss layer(s) of initial model 4004, parameters may be updated and re-tuned for a new data set based on loss calculations associated with accuracy of output or loss layer(s) at generating predictions on new).
Regarding claims 5 and 12, Sriram discloses 5. The artificial neural network performance prediction method of claim 1, wherein the determining of the zone and the operand of the artificial neural network includes:
determining a first zone associated with a forward path of the artificial neural network (see Sriram, para [0058], wherein quantizing all weights and activations of the network except for layers that require finer granularity in representation than the 8-bit quantization can provide (e.g., regression layers); para [0070], wherein numerically expressed (e.g., activation values) where each node is given a number. This node is then passed through the neural network during training; and para [0634], wherein quantize the one or more parameters of the one or more layers of the neural network during a forward pass by using an absolute maximum value of the one or more weights); and
determining an activation value associated with forward propagation of the first zone as the operand (see Sriram, para [0605], wherein changing precision of one or more weights and one or more activation values of a portion of a neural network to generate a first trained model; para [0634], wherein quantize the one or more parameters of the one or more layers of the neural network during a forward pass; para [0571], wherein during model training 3614, by having reset or replaced output or loss layer(s) of initial model 4004, parameters may be updated and re-tuned for a new data set based on loss calculations associated with accuracy of output or loss layer(s) at generating predictions on new… calculating the output of a specific layer, or region ("zone") to use as the input (operand) for the next mathematical operation).
Regarding claims 6 and 13, Sriram discloses the artificial neural network performance prediction method of claim 1, wherein the determining of the zone and the operand of the artificial neural network includes:
determining a second zone associated with a backward path of the artificial neural network (see Sriram, para [0058], wherein quantizing all weights and activations of the network except for layers that require finer granularity in representation than the 8-bit quantization can provide (e.g., regression layers); para [0292], wherein store current register values to a designated region in memory (e.g., identified by a context pointer); and para [0611], wherein updating the portion of the neural network, during training, by using the calculated gradients in a backward propagation pass); and
determining at least one of an activation gradient and a weight gradient associated with backward propagation of the second zone as the operand (see Sriram, para [0057], wherein quantizing models such as rounding the weights and activations to lower-bit integers; para [0058], wherein quantizing all weights and activations of the network except for layers that require finer granularity in representation than the 8-bit quantization can provide (e.g., regression layers); and para [0611], wherein updating the portion of the neural network, during training, by using the calculated gradients in a backward propagation pass).
Regarding claims 7 and 14, Sriram discloses the artificial neural network performance prediction method of claim 1, wherein the determining of the zone and the operand of the artificial neural network includes:
determining a third zone associated with at least one layer of the artificial neural network; and determining at least one of an activation value, an activation gradient, and a weight gradient of the third zone as the operand (see Sriram, paras [0068] & [0071], wherein GPU 116 causes a model 110 to process said particular training image, calculates loss using a loss function, and determines a gradient for said particular training image based on calculated loss and said loss function……the first trained model's 110 weights are quantized again, and range/scale factors for activations are calibrated again using the PTQ 112 process. While the activations were quantized using the running-statistics of their distribution during QAT training, the activations may be re-quantized by calculating their statistics again against the calibration dataset which is expected during deployment. By applying PTQ 112 on a QAT quantized model 110, all layers of the neural network can be quantized. The resulting output is a trained model 114 that has both QAT and PTQ applied).
Regarding claims 8 and 15, Sriram discloses the artificial neural network performance prediction method of claim 1, wherein the candidate data format includes at least one candidate data format, and the artificial neural network performance prediction method further comprises determining an optimal data format for the zone among the at least one candidate data format based on the performance indicator (see Sriram, paras [0278] & [0545], wherein software 3618 and/or services 3620 may be optimized for GPU processing with respect to deep learning, machine learning, and/or high-performance computing, as non-limiting examples. In at least one embodiment, at least some of computing environment of deployment system 3606 and/or training system 3604 may be executed in a datacenter one or more supercomputers or high-performance computing systems, with GPU optimized software (e.g., hardware and software combination of NVIDIA's DGX system). In at least one embodiment, datacenters may be compliant with provisions of HIPAA, such that receipt, processing, and transmission of imaging data and/or other patient data is securely handled with respect to privacy of patient data. In at least one embodiment, hardware 3622 may include any number of GPU s that may be called upon to perform processing of data in parallel, as described herein. In at least one embodiment, cloud platform may further include GPU processing for GPU-optimized execution of deep learning tasks, machine learning tasks, or other computing tasks…..; and paras [0057]-[0059], wherein transform (e.g., quantize) a model to have weights represented by lower-precision values (e.g., low-bit integers) instead of using higher-precision values (e.g., values with full floating point-precision) to conserve memory usage and reduces computation when the trained model is deployed; and para [0112], wherein using fewer bits to represent the parameters, a smaller die-area for dedicated hardware and/or less computations per second may be used to perform the same processing tasks, and thus reducing cost and power requirements. In an embodiment, accuracies of the model is improved by applying PTQ to a QAT model at 8-bit/lower bit quantization. The quantization may be performed during the DNN training. In some instances, the range of weights is typically limited due to regularization, which penalizes growth of weight magnitudes and their statistics and is known in each training iteration. Therefore, the weights can be quantized by finding the absolute maximum value of the weights).
Conclusion
The prior arts made of record and not relied upon is considered pertinent to applicant's disclosure. (US Pub No. 2020/0143231; US Pat No. 11,232,360; US Pub No. 2012/0079456; US Pub No. 2020/0193273; US Pat No. 8,972,940; US Pub No. 2018/0314940; US Pub No. 2023/0252224; and I Hubara, M Courbariaux, D Soudry, R El-Yaniv (Quantized neural networks: Training neural networks with low precision weights and activations), - journal of machine …, 2018 - jmlr.org.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to HAFIZ A KASSIM whose telephone number is (571)272-8534. The examiner can normally be reached 9:00 - 5:00 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Rutao Wu can be reached at 571-272-6045. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/HAFIZ A KASSIM/Primary Examiner, Art Unit 3623 09/09/2026