Notice of Pre-AIA or AIA Status
This Action is responsive to Claims filed on 04/05/2023.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55.
Information Disclosure Statement
The information disclosure statement(s) (IDS) submitted on 04/05/2023 and 03/09/2026 were filed before the mailing date of the first Action. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Drawings
Receipt is acknowledged of Drawings filed 04/05/2023. These Drawings are acceptable.
Status of the Claims
Claims 1-20 are currently pending.
Claim Objections
Claims objected to because of the following informalities:
Claim 1: “An processor…” should be “A processor…”
Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-2, 6-8, 11-13, and 17-19 rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception (i.e., a law of nature, a natural phenomenon, or an abstract idea) without significantly more; and because the claims as a whole, considering all claim elements both individually and in combination, do not amount to significantly more than the abstract idea, see Alice Corporation Pty. Ltd. v. CLS Bank International, et al, 573 U.S. (2014). In determining whether the claims are subject matter eligible, the Examiner applies the 2019 USPTO Patent Eligibility Guidelines. (2019 Revised Patent Subject Matter Eligibility Guidance, 84 Fed. Reg. 50, Jan. 7, 2019.)
Step 1: All Claims
Claims 1-11 recite an apparatus, which falls under the statutory category of a machine. Claims 12-20 recite a method, which falls under the statutory category of a process.
Step 2A – Prong 1:
Claim 1 recites an abstract idea, law of nature, or natural phenomenon. The limitations of “…transform input feature maps (IFMs) by performing a forward transform operation in a Winograd convolution (WinConv) domain;”, “…multiply the transformed IFMs by transformed kernels and perform a first inverse transform operation based on results of the multiplying…”, and “…generate output feature maps (OFMs) based on a result of the first inverse transform operation.” under the broadest reasonable interpretation, cover a mental process including an observation, evaluation, judgment or opinion that could be performed in the human mind or with the aid of pencil and paper.
The above limitations recite an algorithmic set of data manipulation or calculation steps practically performed within the human mind or with the aid of pen and paper.
Step 2A – Prong 2:
The additional elements of claim 1 do not integrate the abstract idea into a judicial exception. The claim recites the additional elements “An processor-implemented apparatus”, “a forward transform module”, “multiply and accumulate array (MAA) units… the MAA units comprising adder trees and multipliers;”, and “an inverse transform module” are recognized as generic computer components recited at a high level of generality. Although they have and execute instructions to perform the abstract idea itself, this also does not serve to integrate the abstract idea into a practical application as it merely amounts to instructions to "apply it." (See MPEP 2106.04(d)(2) indicating mere instructions to apply an abstract idea does not amount to integrating the abstract idea into a practical application).
The additional elements recited in the limitation “Winograd convolution” are recognized as non-generic computer components, but are recited at a high level of generality and are found to generally link the abstract idea to a particular technological environment or field of use (See MPEP 2106.05(h)).
Step 2B:
The only limitation on the performance of the described method is a limitation reciting “An processor-implemented apparatus”, “a forward transform module”, “multiply and accumulate array (MAA) units… the MAA units comprising adder trees and multipliers;”, and “an inverse transform module” These elements are insufficient to transform a judicial exception to a patentable invention because the recited elements are considered insignificant extra-solution activity (generic computer system, processing resources, links the judicial exception to a particular, respective, technological environment). The claim thus recites computing components only at a high-level of generality such that it amounts to no more than mere instructions to apply the exception using generic computer components; mere instructions to apply an exception using a generic computer component cannot provide an inventive concept (see MPEP 2106.05(f)).
The additional elements recited in the limitation “Winograd convolution” are recognized as non-generic computer components, but are recited at a high level of generality and are found to generally link the abstract idea to a particular technological environment or field of use (See MPEP 2106.05(h)).
Taken alone or in ordered combination, these additional elements do not amount to significantly more than the above-identified abstract idea. There is no indication that the combination of elements improves the functioning of a computer or improves any other technology. Their collective functions merely provide conventional computer implementation.
For the reasons above, claims 1 and 12 are rejected as being directed to non-patentable subject matter under §101.
Claim 12 recites similar limitation to claim 1 save for “A processor-implemented method,” (generic computer components). The “transforming…”, “multiplying…”, “performing…”, and “generating…” steps are recited even more highly generally than their corresponding limitations of Claim 1.
Dependent Claims:
Claim 2 (claim 13) recites generic computer components and abstract idea mental process steps “…perform the first inverse transform operation based on the results of the multiplying and an output transformation matrix that is transposed…” and “…generate the OFMs by performing a second inverse transform operation on the result of the first inverse transform operation and the output transformation matrix.”
Claim 3 (claim 14) is not rejected under 101 because it does not recite non-eligible subject matter; however, incorporation of the dependent claim into the independent claims will require the additional element to be analyzed under the 2-step process.
Claim 4 (claim 15) is not rejected under 101 because it does not recite non-eligible subject matter; however, incorporation of the dependent claim into the independent claims will require the additional element to be analyzed under the 2-step process.
Claim 5 (claim 16) is not rejected under 101 because it does not recite non-eligible subject matter; however, incorporation of the dependent claim into the independent claims will require the additional element to be analyzed under the 2-step process.
Claim 6 (claim 17) recites generic computer components and abstract idea mental process steps “perform the first inverse transform operation based on the results of the multiplying, using an addition operation in the adder trees; and generate a plurality of dot products as the result of the first inverse transform operation.”
Claim 7 (claim 18) recites generic computer components and abstract idea mental process steps “perform a second inverse transform operation on the result of the first inverse transform operation, using a WinConv inverse transform operation; and generate the OFMs based on a result of the second inverse transform operation.”
Claim 8 (claim 19) recites instructions to apply the abstract idea mental process steps on the generic computer components (See MPEP 2016.05(f)).
Claim 9 (claim 20) is not rejected under 101 because it does not recite non-eligible subject matter; however, incorporation of the dependent claim into the independent claims will require the additional element to be analyzed under the 2-step process.
Claim 10 is not rejected under 101 because it does not recite non-eligible subject matter; however, incorporation of the dependent claim into the independent claims will require the additional element to be analyzed under the 2-step process.
Claim 11 recites generic computer components and abstract idea mental process steps “select a transformation matrix and a transposed transformation matrix based on a size of a kernel and a position of an IFM window; and transform the IFMs into the WinConv domain based on the size of the kernel, the selected transformation matrix, and the selected transposed transformation matrix, to generate the transformed IFMs.”
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim(s) 1-20 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Gopinath et al. (IN 201941039259, published 04/02/2021), hereinafter Gopinath.
In regards to claim 1: The present invention claims: “An processor-implemented apparatus comprising: a forward transform module to transform input feature maps (IFMs) by performing a forward transform operation in a Winograd convolution (WinConv) domain;” Gopinath teaches “Winograd Forward Transform module receives 4×4×16 IFM block in parallel from S0 to S15, and produces 4×4×16 transformed IFM block, which is distributed among 16 MPUs as 1×1×16 microbatches.” ([0044, Page 12).
“multiply and accumulate array (MAA) units to multiply the transformed IFMs by transformed kernels and perform a first inverse transform operation based on results of the multiplying, the MAA units comprising adder trees and multipliers;” Gopinath teaches “The xy-first storage CNN accelerator architecture includes a Multiply Accumulate Array Set (MAA Set) (136). The MAA Set (136) consists of 16 MAAs (138-140) . 25 Each MAA has 16 Multiply-Accumulate Units (MAUs). In each MAA, an IFM tile of size 4×4 is multiplied with a single kernel weight as shown in FIG. 1C. The IFM is broadcasted to all 16 MAAs, and the IFM is multiplied with same kernel index from 16 different kernels, which contribute to 16 OFMs.” ([00430) and “Here, the IFM microbatches are multiplied with the corresponding 16 kernel microbatches to produce 16 partial OFM microbatches.” ([0044], Page 12).
“and an inverse transform module to generate output feature maps (OFMs) based on a result of the first inverse transform operation.” Gopinath teaches “Finally, the resultant 16 microbatches, which together correspond to 16 channels of 4×4 20 partial OFMs, are taken through a Winograd inverse transform module to get a 2×2×16 OFM block. The Winograd inverse transform is given as 𝒚=𝑨𝑻[(𝑮𝒈𝑮𝑻)⊙(𝑩𝑻𝒅𝑩)]𝑨” ([0044], Page 12).
In regards to claim 2: The present invention claims: “wherein the MAA units are configured to perform the first inverse transform operation based on the results of the multiplying and an output transformation matrix that is transposed, and the inverse transform module is configured to generate the OFMs by performing a second inverse transform operation on the result of the first inverse transform operation and the output transformation matrix.” Gopinath teaches “In each MAA, an IFM tile of size 4×4 is multiplied with a single kernel weight as shown in FIG. 1C. The IFM is broadcasted to all 16 MAAs, and the IFM is multiplied with same kernel index from 16 different kernels, which contribute to 16 OFMs.” ([0043]) and “Finally, the resultant 16 microbatches, which together correspond to 16 channels of 4×4 20 partial OFMs, are taken through a Winograd inverse transform module to get a 2×2×16 OFM block. The Winograd inverse transform is given as…” ([0044, Page 12)
In regards to claim 3: The present invention claims: “wherein the MAA units comprise a first set of MAA units and a second set of MAA units, and the first set of MAA units corresponds to first alternate MAA units, and the second set of MAA units corresponds to second alternate MAA units.” Gopinath teaches “The xy-first storage CNN accelerator architecture includes a Multiply Accumulate Array Set (MAA Set) (136). The MAA Set (136) consists of 16 MAAs (138-140).” ([0043]) and “The hybrid traversal may be introduced in the MAAs (418-420) by adding 7 extra accumulators for each multiplier in MAU and multiplexer network to select one of the 8 accumulators for updating the MAAs (418-420).” ([0051]).
In regards to claim 4: The present invention claims: “wherein the first set of MAA units comprises a first set of multipliers among the multipliers, the second set of MAA units comprises a second set of multipliers, other than the first set of multipliers, among the multipliers, and the second set of MAA units is configured to disable the second set of multipliers based on a zero gating at input terminals of the second set of multipliers, during the multiplying of the transformed IFMs and the transformed kernels in the first set of MAA units.” Gopinath teaches “There is another option, where only energy can be saved without any performance improvement by switching off a number of multipliers that have one of their operands equal to zero. Kernel pruning plays an important role in increasing the sparsity of kernels for improved acceleration factor and reduction in size of the trained model.” ([0005])
In regards to claim 5: The present invention claims: “wherein a first number of multipliers in the first set of multipliers is used by the first set of MAA units for the multiplying of the transformed IFMs and the transformed kernels, and a second number of multipliers, other than the first number of multipliers, in the first set of multipliers, are disabled during the multiplying of the transformed IFMs and the transformed kernels based on a zero gating at input terminals of the second number of multipliers.” Gopinath teaches “
Winograd Forward Transform module receives 4×4×16 IFM block in parallel from S0 to S15, and produces 4×4×16 transformed IFM block, which is distributed among 16 MPUs as 1×1×16 microbatches.” ([0044], Page 12) and “There is another option, where only energy can be saved without any performance improvement by switching off a number of multipliers that have one of their operands equal to zero. Kernel pruning plays an important role in increasing the sparsity of kernels for improved acceleration factor and reduction in size of the trained model.” ([0005])
In regards to claim 6: The present invention claims: “wherein the plurality of MAA units is configured to: perform the first inverse transform operation based on the results of the multiplying, using an addition operation in the adder trees;” Gopinath teaches “The outputs of MPUs are selectively added using an OFM adder tree. To support WgConv, a Winograd Forward 15 Transform (WFT) unit is introduced after the pixel memories. OFM adder tree is reconfigured to support Winograd inverse transform function for WgConv.” ([0046], Page 13)
“and generate a plurality of dot products as the result of the first inverse transform operation.” Gopinath teaches “As the independent element-wise operations are distributed to different MPUs, the same MPU data path realizing dot product operations used in DConv can be 15 applied here without modifications.” ([0044], Page 12) and “The dot product includes the kernel memory (122), a single ported IFM memory (402), a single ported OFM memory (404), an IFM cache (406), multipliers (410), an adder tree (412), an adder (414), and a multiple accumulator registers (416). The dot product uses optimizations of a scheme 2 and introduces newer optimizations like a kernel cache for kernel reuse, multiple accumulators in MPU (124) to compute multiple OFMs for each kernel value.” ([0048]).
In regards to claim 7: The present invention claims: “wherein the inverse transform module is configured to: perform a second inverse transform operation on the result of the first inverse transform operation, using a WinConv inverse transform operation; and generate the OFMs based on a result of the second inverse transform operation.” Gopinath teaches “The hybrid traversal for 3×3 convolution includes (a) The Hybrid traversal for DConv has OFM in an outer loop and the kernel in an inner loop similar to output stationary traversal. In addition, due to computation of partial OFM pixels of a single channel, two additional loops are introduced inside the kernel loop, making it partial weight stationary 10 under output stationary traversal. (b) The hybrid traversal for WgConv has three additional loops due to computation of partial OFM pixels of two channels.“ ([0057], [0056] dictates the algorithm and shows the multiple transformations).
In regards to claim 8: The present invention claims: “wherein the transformed kernels are transformed into the WinConv domain by the MAA units.” See Gopinath [0056] for the algorithm dictating the transformed kernels in the WGConv portion.
In regards to claim 9: The present invention claims: “a plurality of memory banks configured to store channels of coordinates of each of the IFMs as IFM blocks in a z-first data storage layout and transmit the IFM blocks to an IFM fetcher; and the IFM fetcher configured to fetch the IFM blocks.” Gopinath teaches “To facilitate this, the microbatches from individual x-y location under every 4×4 IFM block are stored in different banks of pixel memory (102). Hence, the use of 16 different banks of pixel memory. The same pattern is followed for all the 10 IFM channels.“ ([0045], Page 12) and “The scheme 2 uses IFM cache (406) to increase IFM reuse along with the register clock gating of the scheme 1. In an existing z-direction storage CNN accelerator architecture the IFM value is broadcasted across multiple kernels.“ ([0049]).
In regards to claim 10: The present invention claims: “a data staging unit configured to distribute the transformed IFMs into a plurality of IFM buffers and rearrange the transformed IFMs so that at least four pixels per channel are provided together at an input terminal of each of the plurality of MAA units.” Gopinath teaches “FIG. 8 illustrates a strided convolution using the hybrid traversal, according to embodiments as disclosed herein. In the strided convolution, stride factors are considered while reading the IFM to IFM buffer (620). The plurality of kernel values (602-618) are multiplied with the IFM read into the IFM buffer. For example, consider 3×3 stride 2, first reading alternate pixels of IFM enables multiplication of them with corner pixels of 3×3 kernel. The idea is to reuse the read pixel in IFM buffer to the maximal extent. Each kernel element is multiplied with 8 IFMs in sequence (similar to direct convolution). Similarly, IFMs corresponding to other kernel elements are read into IFM buffers and reused.“ ([0058]).
In regards to claim 11: The present invention claims: “wherein the forward transform module is configured to: select a transformation matrix and a transposed transformation matrix based on a size of a kernel and a position of an IFM window;” Gopinath teaches “For the DConv, a sequential access of IFM microbatches under every IFM window is needed for convolution. Therefore, IFM is divided into batches of 16 channels that are stored in separate pixel memory banks. As the WgConv mode produces 4×16 OFM block, the hybrid traversal for 25 DConv is designed such that OFM block of same dimensions appears at the output.” ([0047], Page 13).
“and transform the IFMs into the WinConv domain based on the size of the kernel, the selected transformation matrix, and the selected transposed transformation matrix, to generate the transformed IFMs.” See Gopinath [0056] for the algorithm dictating the transformed kernels in the WGConv portion.
In regards to claim 12-20: Claims 12-20 recites similar limitations to those found in Claims 1-9, save for the recitation of “A processor-implemented method, comprising…” of Claim 12; therefore, both sets of claims are similarly rejected.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
US 2021/0117755 A1 (US Application which claims the cited Gopinath reference as priority.
US 20210357734 A1 (Also relevant to overarching structural similarities)
Any inquiry concerning this communication or earlier communications from the examiner should be directed to GRIFFIN T BEAN whose telephone number is (703)756-1473. The examiner can normally be reached M - F 7:30 - 4:30.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Li Zhen can be reached at (571) 272-3768. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/GRIFFIN TANNER BEAN/ Examiner, Art Unit 2121
/Li B. Zhen/ Supervisory Patent Examiner, Art Unit 2121