DETAILED ACTION
This action is in response to the submission filed 27 06 2023 for application 18/342,661. Currently claims 1-20 are pending and have been examined.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Acknowledgment is made of applicant's claim for foreign priority based on an application filed in Japan on 28 April, 2023. It is noted, however, that applicant has not filed a translation of the TW112116132 application as required by 37 CFR 1.55.
Information Disclosure Statement
Information disclosure statements (IDS) were submitted on 27 June 2023, 25 January 2024, 23 April 2024, and 06 August 2024. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1 - 20 are rejected under 35 U.S.C. 101 because the claimed invention is directed towards abstract ideas without significantly more.
Regarding claims 1-6:
According to the first step (Step 1) of the 101 analysis, claims 1-13 are directed to an optimizing method for a deep learning network (process) and falls within one of the four statutory categories (i.e., process, machine, manufacture, or composition of matter).
Regarding claim 1:
In the next step (Step 2A, prong 1) of the analysis, the limitations of:
quantizing the first data into a first format or a second format through a power of two quantization, wherein numbers of first values in the first format or the second format are different;
Under the broadest reasonable interpretation, the above limitation is a process step that recites mathematical relationships and calculations such as power of two quantization. If a claim, under its broadest reasonable interpretation covers mathematical concepts but for the recitation of generic computer components, then it falls within the “Mathematical concepts” grouping of abstract ideas.
In the next step (Step 2A, prong 2) of the analysis, the limitations:
using the first format or the second format as a target format;
and performing an operation related to a deep learning network by using the first data quantized based on the target format.
The above limitations are considered to be additional elements and it does not integrate the abstract idea into a practical application because the additional elements are recited so generically (no details whatsoever are provided other than that it is a method using the first format or the second format as a target format; and performing an operation related to a deep learning network by using the first data quantized based on the target format) that it represents no more than mere instructions to apply the judicial exception on a computer. As discussed in MPEP 2106.05(f), mere instructions to implement an abstract idea on a computer as a tool to perform an abstract idea is not indicative of integration into a practical application.
In the same step (Step 2A, prong 2) of the analysis, the limitation, obtaining a first data, is considered to be an additional element and as recited represent insignificant extra-solution activity because it is mere data gathering. See MPEP 2106.05(g), discussing limitations that the Federal Circuit has considered to be insignificant extra-solution activity.
In the last step (Step 2B) of the analysis, the additional element does not amount to significantly more than the judicial exceptions. As explained with respect to Step 2A Prong Two, the method using the first format or the second format as a target format; and performing an operation related to a deep learning network by using the first data quantized based on the target format, is at best the equivalent of merely adding the words “apply it” to the judicial exception. See MPEP 2106.05(f). Even when considered in combination, mere instructions to apply an exception cannot provide an inventive concept and does not amount to significantly more than the judicial exception.
In the last step (Step 2B) of the analysis, as discussed above the additional element of obtaining a first data, which is recited at a high level of generality and amounts to extra-solution activity of receiving data i.e. pre-solution activity of gathering data for use in the claimed process. The courts have found limitations directed to obtaining information electronically, recited at a high level of generality, to be well-understood, routine, and conventional (see MPEP 2106.05(d)(II), “receiving or transmitting data over a network”, "electronic record keeping," and "storing and retrieving information in memory"). These limitations therefore remain insignificant extra-solution activity even upon reconsideration, and do not amount to significantly more.
Accordingly, at Step 2B, after considering all claim elements individually and as an ordered combination, it is determined that the claims do not amount to significantly more than the judicial exception. The claim is not patent eligible.
Regarding claim 2:
In the next step (Step 2A, prong 1) of the analysis, the limitations of:
wherein using the first format or the second format as the target format comprises: determining one of the first format and the second format as the target format according to a quantization error, wherein the quantization error is an error between the first data quantized by the power of two into the first format or the second format and the first data not quantized by the power of two quantization.
Under the broadest reasonable interpretation, the above limitation is a process step that recites mathematical relationships and calculations such as determining based on quantization error calculations. If a claim, under its broadest reasonable interpretation covers mathematical concepts but for the recitation of generic computer components, then it falls within the “Mathematical concepts” grouping of abstract ideas.
In the next step (Step 2A, prong 2) of the analysis, it does not integrate into a practical application because it does not add any additional elements that integrate the abstract idea into practical application.
In the last step (Step 2B) of the analysis, it does not add any additional elements that amount to significantly more than the abstract idea and thus fails to add an inventive concept.
Accordingly, at Step 2B, after considering all claim elements individually and as an ordered combination, it is determined that the claims do not amount to significantly more than the judicial exception. The claim is not patent eligible.
Regarding claim 3:
In the next step (Step 2A, prong 1) of the analysis, the limitations of:
and quantizing the first data into the first format or the second format through the power of two quantization comprises: determining a scaling factor of the layer;
determining an upper limit of a quantization value and a lower limit of the quantization value of the layer according to the scaling factor;
and determining a data set in the power of two quantization for the layer according to the upper limit of the quantization value and the lower limit of the quantization value, wherein the data set is used to define quantization values in the first format and the second format.
Under the broadest reasonable interpretation, the above limitations are process steps that recite mathematical relationships and because quantizing the first data into the first format or the second format through the power of two quantization as explained above is reciting a mathematical concept where determining a scaling factor looking at the data and then determining the upper and lower limits according to that scaling factor and determining a data set in the power of two quantization for the layer according to the upper limit of the quantization value and the lower limit of the quantization value, wherein the data set is used to define quantization values in the first format and the second format are all done mathematically. Specification paragraph [0035] for example supports that this mathematical as well. If a claim, under its broadest reasonable interpretation covers mathematical concepts but for the recitation of generic computer components, then it falls within the “Mathematical concepts” grouping of abstract ideas.
In the next step (Step 2A, prong 2) of the analysis, the limitation:
wherein the first data belongs to one of a plurality of layers in a pre-training model based on the deep learning network.
is considered to be an additional element and it does not integrate the abstract idea into a practical application because the additional element is recited so generically (no details whatsoever are provided other than that it is a method wherein the first data belongs to one of a plurality of layers in a pre-training model based on the deep learning network) that it represents no more than mere instructions to apply the judicial exception on a computer. As discussed in MPEP 2106.05(f), mere instructions to implement an abstract idea on a computer as a tool to perform an abstract idea is not indicative of integration into a practical application.
In the last step (Step 2B) of the analysis, the additional element does not amount to significantly more than the judicial exceptions. As explained with respect to Step 2A Prong Two, the method wherein the first data belongs to one of a plurality of layers in a pre-training model based on the deep learning network, is at best the equivalent of merely adding the words “apply it” to the judicial exception. See MPEP 2106.05(f). Even when considered in combination, mere instructions to apply an exception cannot provide an inventive concept and does not amount to significantly more than the judicial exception.
Accordingly, at Step 2B, after considering all claim elements individually and as an ordered combination, it is determined that the claims do not amount to significantly more than the judicial exception. The claim is not patent eligible.
Regarding claim 4:
In the next step (Step 2A, prong 2) of the analysis, the limitation:
wherein the first format is one-hot encoding, the second format is two-hot encoding, and the first data is a weight of a pre-training model based on the deep learning network.
is considered to be an additional element and it does not integrate the abstract idea into a practical application because the additional element is recited so generically (no details whatsoever are provided other than that it is a method wherein the first format is one-hot encoding, the second format is two-hot encoding, and the first data is a weight of a pre-training model based on the deep learning network) that it represents no more than mere instructions to apply the judicial exception on a computer. As discussed in MPEP 2106.05(f), mere instructions to implement an abstract idea on a computer as a tool to perform an abstract idea is not indicative of integration into a practical application.
In the last step (Step 2B) of the analysis, the additional element does not amount to significantly more than the judicial exceptions. As explained with respect to Step 2A Prong Two, the method wherein the first format is one-hot encoding, the second format is two-hot encoding, and the first data is a weight of a pre-training model based on the deep learning network, is at best the equivalent of merely adding the words “apply it” to the judicial exception. See MPEP 2106.05(f). Even when considered in combination, mere instructions to apply an exception cannot provide an inventive concept and does not amount to significantly more than the judicial exception.
Accordingly, at Step 2B, after considering all claim elements individually and as an ordered combination, it is determined that the claims do not amount to significantly more than the judicial exception. The claim is not patent eligible.
Regarding claim 5:
In the next step (Step 2A, prong 1) of the analysis, the limitations of:
wherein the operation is a multiplication operation, and the target format is the one-hot encoding, and performing the operation related to the deep learning network using the first data quantized based on the target format comprises:
shifting a second data through a shifter according to a position of the first value in the target format, wherein the second data is a parameter for performing the operation with the first data in the deep learning network.
Under the broadest reasonable interpretation, the above limitation is a process step that clearly recites mathematical relationships and calculations such as multiplication operation and shifting. If a claim, under its broadest reasonable interpretation covers mathematical concepts but for the recitation of generic computer components, then it falls within the “Mathematical concepts” grouping of abstract ideas.
In the next step (Step 2A, prong 2) of the analysis, it does not integrate into a practical application because it does not add any additional elements that integrate the abstract idea into practical application.
In the last step (Step 2B) of the analysis, it does not add any additional elements that amount to significantly more than the abstract idea and thus fails to add an inventive concept.
Accordingly, at Step 2B, after considering all claim elements individually and as an ordered combination, it is determined that the claims do not amount to significantly more than the judicial exception. The claim is not patent eligible.
Regarding claim 6:
In the next step (Step 2A, prong 1) of the analysis, the limitations of:
wherein the operation is a multiplication operation, and the target format is the two-hot encoding, and performing the operation related to the deep learning network using the first data quantized based on the target format comprises:
shifting a second data through a shifter according to two positions of the first value in the target format, wherein the second data is a parameter for performing the operation with the first data in the deep learning network;
and adding the shifted second data by an adder.
Under the broadest reasonable interpretation, the above limitation is a process step that clearly recites mathematical relationships and calculations such as multiplication operation, shifting, and adding. If a claim, under its broadest reasonable interpretation covers mathematical concepts but for the recitation of generic computer components, then it falls within the “Mathematical concepts” grouping of abstract ideas.
In the next step (Step 2A, prong 2) of the analysis, it does not integrate into a practical application because it does not add any additional elements that integrate the abstract idea into practical application.
In the last step (Step 2B) of the analysis, it does not add any additional elements that amount to significantly more than the abstract idea and thus fails to add an inventive concept.
Accordingly, at Step 2B, after considering all claim elements individually and as an ordered combination, it is determined that the claims do not amount to significantly more than the judicial exception. The claim is not patent eligible.
Regarding claim 7:
In the next step (Step 2A, prong 1) of the analysis, the limitations of:
quantizing the first data through a power of two quantization, wherein the first data quantized through the power of two quantization is a first format or a second format, and numbers of first values in the first format or the second format are different;
quantizing the second data through a dynamic fixed-point quantization;
and performing an operation related to a deep learning network on the quantized first data after the power of two quantization and the quantized second data after the dynamic fixed- point quantization.
Under the broadest reasonable interpretation, the above limitation is a process step that recites mathematical relationships and calculations such as power of two quantization. If a claim, under its broadest reasonable interpretation covers mathematical concepts but for the recitation of generic computer components, then it falls within the “Mathematical concepts” grouping of abstract ideas.
In the next step (Step 2A, prong 2) of the analysis, the limitations:
obtaining a first data;
obtaining a second data;
is considered to be an additional element and as recited represent insignificant extra-solution activity because it is mere data gathering. See MPEP 2106.05(g), discussing limitations that the Federal Circuit has considered to be insignificant extra-solution activity.
In the last step (Step 2B) of the analysis, as discussed above the additional elements of obtaining a first and second data, which is recited at a high level of generality and amounts to extra-solution activity of receiving data i.e. pre-solution activity of gathering data for use in the claimed process. The courts have found limitations directed to obtaining information electronically, recited at a high level of generality, to be well-understood, routine, and conventional (see MPEP 2106.05(d)(II), “receiving or transmitting data over a network”, "electronic record keeping," and "storing and retrieving information in memory"). These limitations therefore remain insignificant extra-solution activity even upon reconsideration, and do not amount to significantly more.
Accordingly, at Step 2B, after considering all claim elements individually and as an ordered combination, it is determined that the claims do not amount to significantly more than the judicial exception. The claim is not patent eligible.
Regarding claim 8:
In the next step (Step 2A, prong 1) of the analysis, the limitations of:
wherein quantizing the first data through the power of two quantization comprises: quantizing the first data into the first format or the second format through the power of two quantization;
wherein performing the operation related to the deep learning network on the quantized first data after the power of two quantization and the quantized second data after the dynamic fixed-point quantization comprises: performing the operation using the first data quantized based on the target format.
Under the broadest reasonable interpretation, the above limitations are process steps that recite mathematical relationships and because quantizing the first data into the first format or the second format through the power of two quantization as explained above is reciting a mathematical concept and performing the operation related to the deep learning network on the quantized first data after the power of two quantization and the quantized second data after the dynamic fixed-point quantization comprises are all done mathematically also. If a claim, under its broadest reasonable interpretation covers mathematical concepts but for the recitation of generic computer components, then it falls within the “Mathematical concepts” grouping of abstract ideas.
In the next step (Step 2A, prong 2) of the analysis, the limitation:
and using the first format or the second format as a target format.
is considered to be an additional element and it does not integrate the abstract idea into a practical application because the additional element is recited so generically (no details whatsoever are provided other than that it uses the first format or the second format as a target format) that it represents no more than mere instructions to apply the judicial exception on a computer. As discussed in MPEP 2106.05(f), mere instructions to implement an abstract idea on a computer as a tool to perform an abstract idea is not indicative of integration into a practical application.
In the last step (Step 2B) of the analysis, the additional element does not amount to significantly more than the judicial exceptions. As explained with respect to Step 2A Prong Two, the method wherein using the first format or the second format as a target format, is at best the equivalent of merely adding the words “apply it” to the judicial exception. See MPEP 2106.05(f). Even when considered in combination, mere instructions to apply an exception cannot provide an inventive concept and does not amount to significantly more than the judicial exception.
Accordingly, at Step 2B, after considering all claim elements individually and as an ordered combination, it is determined that the claims do not amount to significantly more than the judicial exception. The claim is not patent eligible.
Regarding claim 9:
Claim 9 is substantially similar to claim 2 and therefore is rejected on similar grounds as claim 2.
Regarding claim 10:
Claim 10 is substantially similar to claim 3 and therefore is rejected on similar grounds as claim 3.
Regarding claim 11:
Claim 11 is substantially similar to claim 4 and therefore is rejected on similar grounds as claim 4.
Regarding claim 12:
Claim 12 is substantially similar to claim 5 and therefore is rejected on similar grounds as claim 5.
Regarding claim 13:
Claim 13 is substantially similar to claim 6 and therefore is rejected on similar grounds as claim 6.
Regarding claims 14-20:
According to the first step (Step 1) of the 101 analysis, claims 14-20 are directed to a system (manufacture) and falls within one of the four statutory categories (i.e., process, machine, manufacture, or composition of matter).
Regarding claim 14:
In step (Step 2A, prong 2) of the analysis, the limitations of:
A computing system for a deep learning network, comprising: a memory, which is configured to store program codes;
and a processor, which is coupled to the memory and configured to load the program codes to:
are considered to be additional elements and it does not integrate the abstract idea into a practical application because the additional element is recited so generically (no details whatsoever are provided other than that it is a computing system for a deep learning network, comprising: a memory, which is configured to store program codes; and a processor, which is coupled to the memory and configured to load the program codes to perform something) that it represents no more than mere instructions to apply the judicial exception on a computer. As discussed in MPEP 2106.05(f), mere instructions to implement an abstract idea on a computer as a tool to perform an abstract idea is not indicative of integration into a practical application.
The rest of the limitations of claim 14 are substantially similar to claim 1 and therefore is rejected on similar grounds as claim 1 as explained above.
Regarding claim 15:
Claim 15 is substantially similar to claim 2 and therefore is rejected on similar grounds as claim 2.
Regarding claim 16:
Claim 16 is substantially similar to claim 3 and therefore is rejected on similar grounds as claim 3.
Regarding claim 17:
Claim 17 is substantially similar to claim 4 and therefore is rejected on similar grounds as claim 4.
Regarding claim 18:
Claim 18 is substantially similar to claim 5 and therefore is rejected on similar grounds as claim 5.
Regarding claim 19:
Claim 19 is substantially similar to claim 6 and therefore is rejected on similar grounds as claim 6.
Regarding claim 20:
In the next step (Step 2A, prong 1) of the analysis, the limitations of:
quantize the second data through a dynamic fixed-point quantization;
and perform the operation on the quantized first data after the power of two quantization and the quantized second data after the dynamic fixed-point quantization.
Under the broadest reasonable interpretation, the above limitation is a process step that recites mathematical relationships and calculations such as calculating a dynamic fixed-point quantization of the second data and perform the operation on the quantized first data after the power of two quantization and the quantized second data after the dynamic fixed-point quantization. If a claim, under its broadest reasonable interpretation covers mathematical concepts but for the recitation of generic computer components, then it falls within the “Mathematical concepts” grouping of abstract ideas.
In the next step (Step 2A, prong 2) of the analysis, the limitations:
obtaining a second data;
is considered to be an additional element and as recited represent insignificant extra-solution activity because it is mere data gathering. See MPEP 2106.05(g), discussing limitations that the Federal Circuit has considered to be insignificant extra-solution activity.
In the last step (Step 2B) of the analysis, as discussed above the additional elements of obtaining second data, which is recited at a high level of generality and amounts to extra-solution activity of receiving data i.e. pre-solution activity of gathering data for use in the claimed process. The courts have found limitations directed to obtaining information electronically, recited at a high level of generality, to be well-understood, routine, and conventional (see MPEP 2106.05(d)(II), “receiving or transmitting data over a network”, "electronic record keeping," and "storing and retrieving information in memory"). These limitations therefore remain insignificant extra-solution activity even upon reconsideration, and do not amount to significantly more.
Accordingly, at Step 2B, after considering all claim elements individually and as an ordered combination, it is determined that the claims do not amount to significantly more than the judicial exception. The claim is not patent eligible.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1, 2, 14, and 15 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Huang et al (Structured Term Pruning for Computational Efficient Neural Networks Inference, 2021).
Regarding claim 1:
Huang teaches: An optimizing method for a deep learning network, comprising: obtaining a first data ([Abstract] To address this, we propose an algorithm-architecture codesign, named Structured Term Pruning (STP), to boost the computation efficiency of neural networks inference. [Page 5, Column 1] Fig. 2 [Page 8, Column 1, Last Paragraph] The DNNs have a large number of calculations. Note: Figure 2 shows obtaining 84 as the first data. DNN corresponds to Deep Neural Network);
quantizing the first data into a first format or a second format through a power of two quantization, wherein numbers of first values in the first format or the second format are different ([Page 5, Column 1, Fig. 2] Note: Figure 2 shows PoT Quanitzation in the middle which corresponds to power-of-two quantization. 1st row of the middle column corresponds to First format and 2nd row of the middle column corresponds to second format and the numbers of the first values in the first format or the second format are different);
using the first format or the second format as a target format ([Page 5, Column 2, Paragraph 1] The quantization process is straightforward, only keeping the largest terms per group. This approach is similar to the sum of the power of two quantization. Targeting a term budget for a group is apparently more robust than for a single value. From Figure 2 we can see that the term quantization causes less quantization error than PoT with the same number of terms. Note: Fig. 2 shows the 1st row of the middle column corresponding to First format and 2nd row of the middle column corresponding to second format and either of them could be the target format);
and performing an operation related to a deep learning network by using the first data quantized based on the target format ([Page 5, Column 2, Paragraph 2] However, TQ is a post-training quantization method. To further improve the sparsity while maintaining accuracy, retraining of the network is necessary. Note: TQ corresponds to an operation related to deep learning by using the first data quantized).
Regarding claim 2:
Huang teaches: The optimizing method for the deep learning network according to claim 1, wherein using the first format or the second format as the target format comprises: determining one of the first format and the second format as the target format according to a quantization error, wherein the quantization error is an error between the first data quantized by the power of two into the first format or the second format and the first data not quantized by the power of two quantization ([Page 5, Column 1, Fig. 2] Note: Figure 2 shows the quantization error. The 1st row of the 1st column in the figure corresponds to the first data not quantized by the power of two quantization).
Regarding claim 14:
Claim 14 is substantially similar to claim 1 and therefore is rejected on similar grounds as claim 1.
Regarding claim 15:
Claim 15 is substantially similar to claim 2 and therefore is rejected on similar grounds as claim 2.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 3, 7-10, 16, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Huang et al (Structured Term Pruning for Computational Efficient Neural Networks Inference, 2021) in view of Sabir et al (Weight Quantization Retraining for Sparse and Compressed Spatial Domain Correlation Filters, 2021).
Regarding claim 3:
Huang teaches: The optimizing method for the deep learning network according to claim 1 (as shown above).
Huang further teaches: wherein the first data belongs to one of a plurality of layers in a pre-training model based on the deep learning network, and quantizing the first data into the first format or the second format through the power of two quantization comprises: determining a scaling factor of the layer ([Page 2, Column 1, Paragraph 2] On the software algorithm side, we convert the pre-trained full precision model to the group structured bit-sparse model with quantization and structured term pruning. [Page 4, Column 2, Paragraph 1] Where Hl, Wl and MULsl are the corresponding parameters of the lth layer. And the average bit operations BOPAvg reflects the complexity of the multiplication operation after term reduction. [Page 4, Column 2, Paragraph 2] We fix the quantization step size (scaling factor) in the following training steps);
determining an upper limit of a quantization value and a lower limit of the quantization value of the layer according to the scaling factor ([Page 4, Column 2, Last Paragraph] For the avoidance of over-aggressive scaling, we add a hyperparameter to bound the range. Note: Bound the range corresponds to upper and lower limit);
wherein the data set is used to define quantization values in the first format and the second format ([Page 5, Column 1, Fig. 2] Note: Figure 2 shows 1st row of the middle column corresponds to First format and 2nd row of the middle column corresponds to second format).
However, Huang does not explicitly disclose: and determining a data set in the power of two quantization for the layer according to the upper limit of the quantization value and the lower limit of the quantization value.
Sabir teaches, in an analogous system: and determining a data set in the power of two quantization for the layer according to the upper limit of the quantization value and the lower limit of the quantization value ([Page 4, Table 1] m1 Exponential power of two used for the upper-bound. m2 Exponential power of two used for the lower-bound).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the optimizing method for a deep learning network of Huang to incorporate the teachings of Sabir to determine a data set in the power of two quantization for the layer according to the upper limit of the quantization value and the lower limit of the quantization value. One would have been motivated to do this modification because doing so would give the benefit of reducing the computation workload of inference due to large spatial-domain CPR filters as taught by Sabir [Abstract].
Regarding claim 7:
Huang teaches: An optimizing method for a deep learning network, comprising: obtaining a first data ([Abstract] To address this, we propose an algorithm-architecture codesign, named Structured Term Pruning (STP), to boost the computation efficiency of neural networks inference. [Page 5, Column 1] Fig. 2 [Page 8, Column 1, Last Paragraph] The DNNs have a large number of calculations. Note: Figure 2 shows obtaining 84 as the first data. DNN corresponds to Deep Neural Network);
quantizing the first data through a power of two quantization, wherein the first data quantized through the power of two quantization is a first format or a second format, and numbers of first values in the first format or the second format are different ([Page 5, Column 1, Fig. 2] Note: Figure 2 shows PoT Quanitzation in the middle which corresponds to power-of-two quantization. 1st row of the middle column corresponds to First format and 2nd row of the middle column corresponds to second format and the numbers of the first values in the first format or the second format are different).
However, Huang does not explicitly disclose: obtaining a second data; quantizing the second data through a dynamic fixed-point quantization; and performing an operation related to a deep learning network on the quantized first data after the power of two quantization and the quantized second data after the dynamic fixed- point quantization.
Sabir teaches, in an analogous system: obtaining a second data; quantizing the second data through a dynamic fixed-point quantization; and performing an operation related to a deep learning network on the quantized first data after the power of two quantization and the quantized second data after the dynamic fixed- point quantization ([Page 18, Figure 11] Note; Figure 11 shows power of two quantization and the quantized second data after the dynamic fixed-point quantization. It also shows re-quantization corresponding to quantized first data after the power of two quantization and the quantized second data after the dynamic fixed-point quantization).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the optimizing method for a deep learning network of Huang to incorporate the teachings of Sabir to obtain a second data; and quantize the second data through a dynamic fixed-point quantization; and performing an operation related to a deep learning network on the quantized first data after the power of two quantization and the quantized second data after the dynamic fixed- point quantization. One would have been motivated to do this modification because doing so would give the benefit of reducing the computation workload of inference due to large spatial-domain CPR filters as taught by Sabir [Abstract].
Regarding claim 8:
Claim 8 is substantially similar to claim 1 and therefore is rejected on similar grounds as claim 1.
Regarding claim 9:
Claim 9 is substantially similar to claim 2 and therefore is rejected on similar grounds as claim 2.
Regarding claim 10:
Claim 10 is substantially similar to claim 3 and therefore is rejected on similar grounds as claim 3.
Regarding claim 16:
Claim 16 is substantially similar to claim 3 and therefore is rejected on similar grounds as claim 3.
Regarding claim 20:
Huang teaches: The computing system for the deep learning network according to claim 14 (as shown above).
However, Huang does not explicitly disclose: wherein the processor is further configured to: obtain a second data; quantize the second data through a dynamic fixed-point quantization; and perform the operation on the quantized first data after the power of two quantization and the quantized second data after the dynamic fixed-point quantization.
Sabir teaches, in an analogous system: wherein the processor is further configured to: obtain a second data; quantize the second data through a dynamic fixed-point quantization; and perform the operation on the quantized first data after the power of two quantization and the quantized second data after the dynamic fixed-point quantization ([Page 18, Figure 11] Note; Figure 11 shows power of two quantization and the quantized second data after the dynamic fixed-point quantization. It also shows re-quantization corresponding to quantized first data after the power of two quantization and the quantized second data after the dynamic fixed-point quantization).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the optimizing method for a deep learning network of Huang to incorporate the teachings of Sabir wherein the processor is further configured to: obtain a second data; quantize the second data through a dynamic fixed-point quantization; and perform the operation on the quantized first data after the power of two quantization and the quantized second data after the dynamic fixed-point quantization. One would have been motivated to do this modification because doing so would give the benefit of reducing the computation workload of inference due to large spatial-domain CPR filters as taught by Sabir [Abstract].
Claims 4-6 and 17-19 are rejected under 35 U.S.C. 103 as being unpatentable over Huang et al (Structured Term Pruning for Computational Efficient Neural Networks Inference, 2021) in view of Kraus et al (US 20220300696 A1).
Regarding claim 4:
Huang teaches: The optimizing method for the deep learning network according to claim 1 (as shown above).
Huang further teaches: and the first data is a weight of a pre-training model based on the deep learning network ([Page 2, Column 1, Paragraph 2] On the software algorithm side, we convert the pre-trained full precision model to the group structured bit-sparse model with quantization and structured term pruning. To ensure the network accuracy, we fine-tune the weights so that they can be aware of the group structure quantization).
However, Huang does not explicitly disclose: wherein the first format is one-hot encoding, the second format is two-hot encoding.
Kraus teaches, in an analogous system: wherein the first format is one-hot encoding, the second format is two-hot encoding ([0036] In some embodiments, style encoder 166 uses one-hot encoding or two-hot encoding. For example, in some implementations that support a single text style for a particular text element, one-hot encoding is used to map each style option to its own dimension. In some implementations that support both primary and secondary text styles for a particular text element, two-hot encoding is used to map each style option to its own feature column, which includes two values, one for the primary text style and one for the secondary text style).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the optimizing method for a deep learning network of Huang to incorporate the teachings of Kraus wherein the first format is one-hot encoding, the second format is two-hot encoding. One would have been motivated to do this modification because doing so would give the benefit of supporting a single text style for a particular text element (one-hot encoding) or supporting both primary and secondary text styles for a particular text element (two-hot encoding) as taught by Kraus [0036].
Regarding claim 5:
The system of Huang and Kraus teaches: The optimizing method for the deep learning network according to claim 4 (as shown above).
Huang further teaches: wherein the operation is a multiplication operation, and performing the operation related to the deep learning network using the first data quantized based on the target format comprises ([Page 4, Column 1, Paragraph 3] The 32-bit floating-point (FP32) network is firstly quantized to 8-bit fixed-point (INT8) format. The main purpose of bitlevel pruning is to simplify the arithmetic of the INT8 multiplication instead of reducing the memory accesses):
shifting a second data through a shifter according to a position of the first value in the target format, wherein the second data is a parameter for performing the operation with the first data in the deep learning network ([Page 5, Column 2, Paragraph 1] The quantization process is straightforward, only keeping the largest terms per group. This approach is similar to the sum of the power of two quantization. Targeting a term budget for a group is apparently more robust than for a single value. From Figure 2 we can see that the term quantization causes less quantization error than PoT with the same number of terms. Note: Fig. 2 shows the 1st row of the middle column corresponding to First format and 2nd row of the middle column corresponding to second format and either of them could be the target format. [Page 7, Column 2, Pargaraph 3] As the bold content shows at the second shift operation term, four outputs require four different inputs without any data reuse. In reality, this can happen in any shift operation term because of the different term distribution of weight. [Page 8, Column 1, Paragraph 2] The PE completes one shift-add during every cycle. A group of activation is stored in the PE, and the blue logic elements select the current activation for calculation. The yellow logic elements complete the multiplication operation with a shifter and decide whether to negate the product or not, based on the sign of the term. The green logic elements perform the addition of the shift result to the partial sum).
However, Huang does not explicitly disclose: and the target format is the one-hot encoding.
Kraus further teaches, in an analogous system: and the target format is the one-hot encoding ([0036] In some embodiments, style encoder 166 uses one-hot encoding).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the optimizing method for a deep learning network of Huang to incorporate the teachings of Kraus wherein the target format is the one-hot encoding. One would have been motivated to do this modification because doing so would give the benefit of supporting a single text style for a particular text element using one-hot encoding as taught by Kraus [0036].
Regarding claim 6:
The system of Huang and Kraus teaches: The optimizing method for the deep learning network according to claim 4 (as shown above).
Huang further teaches: wherein the operation is a multiplication operation, and performing the operation related to the deep learning network using the first data quantized based on the target format comprises ([Page 4, Column 1, Paragraph 3] The 32-bit floating-point (FP32) network is firstly quantized to 8-bit fixed-point (INT8) format. The main purpose of bitlevel pruning is to simplify the arithmetic of the INT8 multiplication instead of reducing the memory accesses):
shifting a second data through a shifter according to two positions of the first value in the target format, wherein the second data is a parameter for performing the operation with the first data in the deep learning network ([Page 5, Column 2, Paragraph 1] The quantization process is straightforward, only keeping the largest terms per group. This approach is similar to the sum of the power of two quantization. Targeting a term budget for a group is apparently more robust than for a single value. From Figure 2 we can see that the term quantization causes less quantization error than PoT with the same number of terms. Note: Fig. 2 shows the 1st row of the middle column corresponding to First format and 2nd row of the middle column corresponding to second format and either of them could be the target format. [Page 7, Column 2, Paragraph 3] As the bold content shows at the second shift operation term, four outputs require four different inputs without any data reuse. In reality, this can happen in any shift operation term because of the different term distribution of weight. [Page 8, Column 1, Paragraph 2] The PE completes one shift-add during every cycle. A group of activation is stored in the PE, and the blue logic elements select the current activation for calculation. The yellow logic elements complete the multiplication operation with a shifter and decide whether to negate the product or not, based on the sign of the term. The green logic elements perform the addition of the shift result to the partial sum);
and adding the shifted second data by an adder ([Page 8, Column 1, Paragraph 2] The green logic elements perform the addition of the shift result to the partial sum. [Page 9, Column 1, Paragraph 2] The accumulator is composed of several adders and buffers as shown in Figure 3).
However, Huang does not explicitly disclose: and the target format is the two-hot encoding.
Kraus further teaches, in an analogous system: and the target format is the two-hot encoding ([0036] In some embodiments, style encoder 166 uses two-hot encoding).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the optimizing method for a deep learning network of Huang to incorporate the teachings of Kraus wherein the and the target format is the two-hot encoding. One would have been motivated to do this modification because doing so would give the benefit of supporting both primary and secondary text styles for a particular text element using two-hot encoding as taught by Kraus [0036].
Regarding claim 17:
Claim 17 is substantially similar to claim 4 and therefore is rejected on similar grounds as claim 4.
Regarding claim 18:
Claim 18 is substantially similar to claim 5 and therefore is rejected on similar grounds as claim 5.
Regarding claim 19:
Claim 19 is substantially similar to claim 6 and therefore is rejected on similar grounds as claim 6.
Claims 11-13 are rejected under 35 U.S.C. 103 as being unpatentable over Huang et al (Structured Term Pruning for Computational Efficient Neural Networks Inference, 2021) in view of Sabir et al (Weight Quantization Retraining for Sparse and Compressed Spatial Domain Correlation Filters, 2021) and further in view of Kraus et al (US 20220300696 A1).
Regarding claim 11:
The system of Huang and Sabir teaches: The optimizing method for the deep learning network according to claim 8 (as shown above).
The rest of the limitations of Claim 11 are substantially similar to claim 4 and therefore is rejected on similar grounds as claim 4.
Regarding claim 12:
The system of Huang and Sabir teaches: The optimizing method for the deep learning network according to claim 11 (as shown above).
The rest of the limitations of Claim 12 are substantially similar to claim 5 and therefore is rejected on similar grounds as claim 5.
Regarding claim 13:
The system of Huang and Sabir teaches: The optimizing method for the deep learning network according to claim 11 (as shown above).
The rest of the limitations of Claim 13 are substantially similar to claim 6 and therefore is rejected on similar grounds as claim 6.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Rakka et al (MIXED-PRECISION NEURAL NETWORKS: A SURVEY, 2022) discloses summarizing the quantization techniques used generally in literature. Then, we present a thorough survey of the mixed-precision frameworks, categorized according to their optimization techniques such as reinforcement learning and quantization techniques like deterministic rounding. Furthermore, the advantages and shortcomings of each framework are discussed, where we present a juxtaposition. We finally give guidelines for future mixed-precision frameworks.
Vogel (Design and Implementation of Number Representations for Efficient Multiplierless Acceleration of Convolutional Neural Networks, 2020) discloses in the first part of the thesis, previous fixed-point quantization methods are reviewed and a self-supervised extension is proposed which specifically enhances the quantization results on pre-trained neural networks. This novel method is the basis for further quantization procedures in the remainder of the thesis. In the second part, CNNs trained on small-scale classification tasks with binary or ternary valued parameters are evaluated. Furthermore, a hardware efficient method for stochastic rounding is introduced, and it is experimentally shown to enhance classification performance while avoiding the need for multiplications. The third and fourth part of the thesis are devoted to the logarithmic quantization of pre-trained CNNs for large-scale image classification and semantic scene segmentation.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to CHAITANYA RAMESH JAYAKUMAR whose telephone number is (571)272-3369. The examiner can normally be reached Mon-Fri 9am-1pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Omar Fernandez Rivas can be reached at (571)272-2589. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/C.R.J./ Examiner, Art Unit 2128
/OMAR F FERNANDEZ RIVAS/Supervisory Patent Examiner, Art Unit 2128