NON-FINAL REJECTION, FIRST DETAILED ACTION
Status of Prosecution
The present application 18/358,571, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
The application was filed in the Office on July 25, 2023. PCT/US23/83563 was subsequently filed on Dec. 12, 2023.
Claims 1-20 are pending and are all rejected in this rejection. Claims 1, 8 and 15 are independent claims.
Status of Claims
Claims 1-20 are rejected under 35 U.S.C. § 101 because the claimed invention is directed to an abstract idea without significantly more.
Claims 1-3, 5-6, 8-10, 12-13, 15-20 are rejected under 35 USC § 103 as being unpatentable over Jiang et al. (“Jiang”), United States Patent Application Publication 2022/0224901 published on July 14, 2022, in view of Lewis et al. (“Lewis”), United States Patent Application Publication 2022/0237465 published on July 28, 2022.
Claims 4, 11 and 18 are rejected under 35 USC § 103 as being unpatentable over Jiang et al. (“Jiang”), United States Patent Application Publication 2022/0224901 published on July 14, 2022, in view of Lewis et al. (“Lewis”), United States Patent Application Publication 2022/0237465 published on July 28, 2022 in further view of Zimmer et al. (“Zimmer”), United States Patent 11,676,068 published on June 13, 2023.
Claims 7 and 14 are rejected under 35 USC § 103 as being unpatentable over Jiang et al. (“Jiang”), United States Patent Application Publication 2022/0224901 published on July 14, 2022, in view of Lewis et al. (“Lewis”), United States Patent Application Publication 2022/0237465 published on July 28, 2022 in further view of non-patent literature Shallue et al. (“Shallue”), “Measuring the Effects of Data Parallelism on Neural Network Training,” published in 2019.
Claim Rejections – § 101 Subject Matter Eligibility
Claims 1-20 are rejected under 35 U.S.C. § 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding representative claim 1, at step 1, the claim recites method for neural network training, and therefore is a process, which is a statutory category of invention. See MPEP § 2106.03.
At step 2A, prong one, the claim recites a method for neural network training .
The following limitations are the abstract idea of a mathematical calculation. See MPEP § 2106.04(a)(2)(I)(C):
partitioning the weight tensor into subtensors, a dimension of a subtensor corresponding to a subset of the input channels;
selecting one or more subtensors from the subtensors based on one or more weights in the one or more subtensor;
modifying values of the one or more weights in the one or more subtensors to zero; and
further training the neural network by modifying values of one or more weights in one or more other subtensors of the subtensors.
Therefore, the claim recites at least one abstract idea per this part of the analysis.
At step 2A prong 2, the claim language is analyzed to determine whether it recites additional elements that integrate the judicial exception into a practical application. See MPEP § 2106.04(d).
The limitation: “training a neural network by generating a weight tensor for a layer of the neural network, the weight tensor having a dimension corresponding to input channels of the layer;” is a step that, under its broadest reasonable interpretation, is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use, specifically model training. See MPEP §§ 2106.04(d), 2106.05(h).
Accordingly, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. The claim is therefore directed to an abstract idea.
Next, at step 2B of the analysis, the claim is considered if it recites additional elements that amount to significantly more than the judicial exception. See MPEP § 2106.05.
As discussed above with respect to integration of the abstract idea into a practical application, the additional elements of training by generating the weight tensor amount to nothing more than linking the use of the judicial exception to a particular technological environment or field of use. See MPEP § 2106.05(h).
Therefore, claim 1 is ineligible.
As to dependent claim 2, the analysis of the parent claim is incorporated. In the step 2A, prong 2 analysis, the additional limitation of “the layer has a plurality of weight tensors that comprises the weight tensor, the plurality of weight tensors corresponds to different output channels of the layer, and values of weights in a plurality of selected subtensors of the plurality of weight tensors are modified to zero.” is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP § 2106.05(g).
The claim is also ineligible.
As to dependent claim 3, the analysis of the parent claim is incorporated. In the step 2A, prong 2 analysis, the additional limitation of “wherein each of the plurality of weight tensors has the same number of one or more selected subtensors” is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP § 2106.05(g).
The claim is also ineligible.
As to dependent claim 4, the analysis of the parent claim is incorporated. In the step 2A, prong 2 analysis, the additional limitation of “the layer is executed by processing elements arranged in a plurality of columns, a column comprising one or more processing elements, and weights in different weight tensors are processed by different columns” is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer. See MPEP § 2106.05(f)(1).
The claim is also ineligible.
As to dependent claim 5, the analysis of the parent claim is incorporated. In the step 2A, prong 2 analysis, the additional limitation of “wherein different subtensors have the same number of input channels,” is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP § 2106.05(g).
The claim is also ineligible.
As to dependent claim 6, the analysis of the parent claim is incorporated. In the step 2A, prong 1 analysis, the additional limitations of: “wherein selecting the one or more subtensors from the subtensors comprises: determining norms of the subtensors, a norm of a subtensor determined based on values of weights in the subtensor; and selecting the one or more subtensors based on the norms, wherein one or more norms of the one or more subtensors are lower than one or more norms of the one or more other subtensors,” are additional elements that generally are the abstract idea of a mathematical calculation. See MPEP § 2106.04(a)(2)(I)(C).
As to dependent claim 7, the analysis of the parent claim is incorporated. In the step 2A, prong 2 analysis, the additional limitation of “training the neural network comprises passing a first training dataset through at least part of the neural network a first number of times, further training the neural network comprises passing a second training dataset through at least part of the neural network a second number of times, and the first number is greater than the second number,” is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP § 2106.05(g).
The claim is also ineligible.
As to the remaining claims, they are rejected similarly.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. § 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
-A.
Claims 1-3, 5-6, 8-10, 12-13, 15-20 are rejected under 35 USC § 103 as being unpatentable over Jiang et al. (“Jiang”), United States Patent Application Publication 2022/0224901 published on July 14, 2022, in view of Lewis et al. (“Lewis”), United States Patent Application Publication 2022/0237465 published on July 28, 2022.
As to Claim 1, Jiang teaches: A method for neural network training, the method comprising:
training a neural network by generating a weight tensor for a layer of the neural network, the weight tensor having a dimension corresponding to input channels of the layer (Jiang: par. 0052, weight parameters are a general 5D tensor with size corresponding to parameter c1, corresponding to the number of input channels and c2 the number of output channels);
partitioning the weight tensor into subtensors (Jiang: pars. 0056-58, the 5D tensor is reshaped into a 3D tensor of size c’1, c’2, k, and the blocks are partitioned into blocks);
selecting one or more subtensors from the subtensors based on one or more weights in the one or more subtensor (Jiang: par. 0071, Fig. 6, the pruning loss is calculated as a norm of the weights in the block);
modifying values of the one or more weights in the one or more subtensors to zero (Jiang: par. 0071, the computer pruning mask module [610] ranks the micro-structured blocks based on the pruning loss and prunes these selected subtensors them by setting their weights to zero); and
further training the neural network by modifying values of one or more weights in one or more other subtensors of the subtensors (Jiang: par. 0072, through multiple iterations, until loss convergence, the training takes place via the back-propagation and weight update [540] module).
PNG
media_image1.png
738
1300
media_image1.png
Greyscale
Jiang may not explicitly teach: partitioning the weight tensor into subtensors, a dimension of a subtensor corresponding to a subset of the input channels.
While Jiang does teach that the partitioning does may involve the number of channels, it may not necessarily be corresponding to a “subset of the input channels.” (Jiang: par. 0058, blocks of size g…).
Lewis teaches in general concepts related to training a sparse neural network by pruning one or more subsets of weights based on their signal-to-noise ratio (Lewis: Abstract). Specifically, Lewis teaches that subvolumes of the weights or layer outputs of a neural network may be pruned together (Lewis: par. 0037). These may be contiguous weights within a weight tensor (Lewis: par. 0037).
It would have been obvious to a person having ordinary skill in the art at a time before the effective filing date of the application to have modified the Jiang disclosures and teachings by partitioning the tensor into subtensors in contiguous fashion that would be a subset of the input channels as taught and suggested by Lewis. Such a person would have been motivated to do so with a reasonable expectation of success to allow for efficiency and optimization by dealing with contiguous parts of the tensor together that may also correspond to channels.
As to Claim 2, Jiang and Lewis teach the elements of claim 1.
Jiang further teaches: the layer has a plurality of weight tensors that comprises the weight tensor (Jiang: par. 0052, weight parameters are a general 5D tensor with size corresponding to parameter c1, corresponding to the number of input channels and c2 the number of output channels; the corresponding layer is a 4D tensor A of size (h, w ,d, c0)),
the plurality of weight tensors corresponds to different output channels of the layer (Examiner notes that the tensors would encompass the different output channel layers), and
values of weights in a plurality of selected subtensors of the plurality of weight tensors are modified to zero (Jiang: par. 0071, the computer pruning mask module [610] ranks the micro-structured blocks based on the pruning loss and prunes these selected subtensors them by setting their weights to zero).
As to Claim 3, Jiang and Lewis teach the elements of claim 2.
Jiang further teaches: wherein each of the plurality of weight tensors has the same number of one or more selected subtensors (Examiner notes that nothing in Jiang or Lewis indicate a limitation on the number of tensors being restricted, including being equal to a selected subtensor).
As to Claim 5, Jiang and Lewis teach the elements of claim 3.
Jiang further teaches: wherein different subtensors have the same number of input channels (Examiner notes that nothing in Jiang or Lewis indicate a limitation on the number of subtensors being restricted, including being equal the number of input channels. It would have been obvious to do so to encompass all channels additionally).
As to Claim 6, Jiang and Lewis teach the elements of claim 1.
Jiang further teaches: wherein selecting the one or more subtensors from the subtensors comprises:
determining norms of the subtensors, a norm of a subtensor determined based on values of weights in the subtensor (Jiang: par. 0071, the pruning loss may be norm of the weights in the block); and
selecting the one or more subtensors based on the norms, wherein one or more norms of the one or more subtensors are lower than one or more norms of the one or more other subtensors (Jiang: par. 0071, the Compute Pruning Mask module [610] will prune based on the pruning loss parameter, which is based on the norm).
As to Claim 8, it is rejected for similar reasons as claim 1. Jiang further teaches a processor and memory (Jiang: par. 0101).
As to Claim 9, it is rejected for similar reasons as claim 2.
As to Claim 10, it is rejected for similar reasons as claim 3.
As to Claim 12, it is rejected for similar reasons as claim 5.
As to Claim 13, it is rejected for similar reasons as claim 6.
As to Claim 15, it is rejected for similar reasons as claims 1 and 8.
As to Claim 16, it is rejected for similar reasons as claim 2.
As to Claim 17, it is rejected for similar reasons as claim 3.
As to Claim 19, it is rejected for similar reasons as claim 5.
As to Claim 20, it is rejected for similar reasons as claim 6.
B.
Claims 4, 11 and 18 are rejected under 35 USC § 103 as being unpatentable over Jiang et al. (“Jiang”), United States Patent Application Publication 2022/0224901 published on July 14, 2022, in view of Lewis et al. (“Lewis”), United States Patent Application Publication 2022/0237465 published on July 28, 2022 in further view of Zimmer et al. (“Zimmer”), United States Patent 11,676,068 published on June 13, 2023.
As to Claim 4, Jiang and Lewis teach the elements of claim 3.
Jiang and Lewis may not explicitly teach: the layer is executed by processing elements arranged in a plurality of columns, a column comprising one or more processing elements, and
weights in different weight tensors are processed by different columns.
Zimmer teaches in general concepts related to removing sparse data on a pixel by pixel basis (Zimmer: Abstract). Specifically, Zimmer a systolic array receives weights to be processed (Zimmer: Fig. 1B, col. 5, lines 34 to 44, the systolic array [120] and weights [110w]). The systolic array has processing elements [121a-12Nn] which process each column of the weights (Zimmer: col. 6, lines 24 to 40).
PNG
media_image2.png
740
992
media_image2.png
Greyscale
It would have been obvious to a person having ordinary skill in the art at a time before the effective filing date of the claimed invention to have modified the Jiang-Lewis combination utilizing the systolic array as taught and suggested by Zimmer. Such a person would have done so with an expectation for success, to allow for the determination of valid parameters to allow for avoiding some of the operations to multiply and accumulate sparse data (Zimmer: Abstract).
As to Claim 11, it is rejected for similar reasons as claim 4.
As to Claim 18, it is rejected for similar reasons as claim 4.
C.
Claims 7 and 14 are rejected under 35 USC § 103 as being unpatentable over Jiang et al. (“Jiang”), United States Patent Application Publication 2022/0224901 published on July 14, 2022, in view of Lewis et al. (“Lewis”), United States Patent Application Publication 2022/0237465 published on July 28, 2022 in further view of non-patent literature Shallue et al. (“Shallue”), “Measuring the Effects of Data Parallelism on Neural Network Training,” published in 2019.
As to Claim 7, Jiang and Lewis teach the elements of claim 1.
Jiang further teaches: training the neural network comprises passing a first training dataset through at least part of the neural network a first number of times;
further training the neural network comprises passing a second training dataset through at least part of the neural network a second number of times (Jiang: par. 0072, multiple iterations may be taken).
Jiang and Lewis may not explicitly teach: the first number is greater than the second number.
Shallue is non-patent literature discussing the increase in scale of data parallelism available for neural network training (Shallue: Abstract). Specifically, Shallue teaches that various empirical studies have shown that the training set sizes will have differing number of training epochs to converge (Shallue: Sec. 3.1.2).
It would have been obvious to a person having ordinary skill in the art at a time before the effective filing date of the claimed invention to have modified the Jiang-Lewis combination by tailoring the number of epochs based on the size of the training datasets, which may have a first and second where the first is greater than the second as taught and suggested by Shallue. Such a person would have done so with an expectation for success, to allow for optimizing and being efficient with the number of training epochs.
As to Claim 14, it is rejected for similar reasons as claim 7.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JAMES T TSAI whose telephone number is (571)270-3916. The examiner can normally be reached M-F 8-5 Eastern.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Viker Lamardo can be reached on 571-270-5871. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000./JAMES T TSAI
/JAMES T TSAI/ Primary Examiner, Art Unit 2147