DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant's arguments filed 05/26/2026 have been fully considered but they are not persuasive.
Regarding applicant’s argument for 35 U.S.C. 101, applicant argues in page 8-9 “In particular, the Examiner asserts that claims 1, 11, and 17 recite the judicial exception of an abstract idea without significantly more because the claim scope under the broadest reasonable interpretation includes mental processes and do not integrate that abstract idea into a practical application. Applicant respectfully disagrees. For example, independent claims 1 and 17 each clearly recites a data router configured "to perform tensor tiling of an input tensor," which is a practical application …,
Applicant respectfully submits that these newly amended in limitations meet the practical application standard of MPEP 2106.04(d). In particular, and as disclosed at least in paragraphs [0003], [0021], [0028], [0032], [0053], and [0121] of the originally filed Specification, the presence of overlapping shared edges of tensor tiles with duplicated rows or columns improves processing efficiency and reduces power consumption for changing tensor size to perform convolution of tensors having different sizes, which is a technical solution to a technical problem that is necessarily rooted in computing technology …,
Accordingly, for at least this additional reason, Applicant asserts that the rejection of claims 1, 11, and 17, and claims 2-10, 12-16, and 18-20, which depend respectively therefrom, recite statutory subject matter under 35 U.S.C. § 101, and therefore the present rejection should be withdrawn.” Applicants argues that the amended limitations provided a practical application by overlapping shared edges of tensors tiles and provides a technical solution. However, the independent claims process lack specificity of how the tiling affects the convolution process or accomplished given the tiling operations in the claim. The MPEP 2106.05(a) explains that the claims must purport(s) the improvement as explain the process is directed towards performing tiling operations with overlapped memory sections and performing convolutional operations on tiles. Further the applicant argues amended limitations that have not been examined thus the argument is moot and not convincing.
(Examiner Note: Applicant could overcome my providing details on independent claims on how the process work to perform a convolution as explain in the specification.)
Regarding applicant’s argument for 35 U.S.C. 101, applicant argues in page 9 “ Applicant respectfully submits that Yu, Taba, and the additionally cited references fail to teach or suggest each and every one of the features of amended claim 1. For example, alone or in combination, Yu and Taba fail to teach or suggest "wherein the overlapping shared edge includes a duplicate row or column of the shared edge and "change a size of the first tile and the second tile without re-accessing the duplicated row or column" as recited in amended claim 1.
Accordingly, for at least these reasons, Applicant submits that independent claim 1 is patentable over Yu and Taba. Furthermore, independent claims 11 and 17 are patentable over the Yu and Taba for at least similar reasons and further in view of their own features. Still further, claims 2-10 which depend from claim 1, claims 12-16, which depend from claim 11, and claims 18-20 which depend from claim 17 are patentable over Yu and Taba for at least the reasons provided for their respective independent claims and further in view of their own features. Accordingly, Applicant respectfully requests that this rejection of claims 1-20 be reconsidered and withdrawn." Applicant argues how the amended limitations overcome the prior art cited. Further applicant argues the prior art cited Taba does not teach the amended section. However the updated 35 U.S.C. section use Taba for the cited amended section of claim 1. Further applicant argues amended limitations and thus the argument is moot and not convincing.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception, abstract idea, without significantly more.
Step 1: This part of the eligibility analysis evaluates whether the claim(s) falls within any statutory
category. MPEP 2106.03:
According to the first part of the Alice analysis, in the instant case, the claims were determined
to be directed to one of the four statutory categories: an article of manufacture, a method/process (Claims 11-16), a machine/system/product (Claims 1-10, 17-20), and a composition of matter. Based on the claims being determined to be within of the four categories (i.e., process, machine, manufacture, or composition of matter), (Step 1), it must be determined if the claims are directed to a judicial exception (i.e., law of nature, natural phenomenon, and abstract idea).
Step 2A Prong One: This part of the eligibility analysis evaluates whether the claim(s) recites a
judicial exception.
Regarding independent claims 1, 11, 17, the claims recite a judicial exception (i.e., an abstract idea enumerated in the 2019 PEG) without significantly more (Step-2A: Prong One). The applicant's claim limitations under broadest reasonable interpretation covers activities classified under mental processes - concepts performed in the human mind (including an observation, evaluation, judgment, opinion) (see MPEP § 2106.04(a)(2), subsection Ill) and the 2019 PEG. As evaluated below:
Claims 1, 11, 17:
“the data router configured to: determine a split of the input tensor into a plurality of tiles based on the array of interconnected PEs and dimensions of the input tensor” (mental process of judgement)
If the identified limitation(s) falls within at least one of the groupings of abstract ideas, it is
reasonable to conclude that the claim(s) recites an abstract idea in Step 2A Prong One.
Step 2A Prong Two: This part of the eligibility analysis evaluates whether the claim(s) as a whole integrates the recited judicial exception into a practical application of the exception. As evaluated below:
“a systolic array comprising an array of interconnected processing elements (PEs), each PE associated with a PE data memory configured to store at least a portion of a tensor”
“split the input tensor into the plurality of tiles, including a first tile and a second tile overlapping a shared edge, by routing the input tensor data to the PE data memories that store the plurality of tiles”
“wherein the overlapping shared edge includes a duplicated row or column of the shared edge;”
“and change a size of the first tile and the second tile without re-accessing the duplicated row or column.”
These recitations are deemed insufficient to transform the judicial exception to a patentable invention because the recitation is directed to instructions for mere data gathering or data output, see MPEP 2106.05(g).
“a data router configured to perform tensor tiling of an input tensor”
These recitations are deemed insufficient to transform the judicial exception to a patentable invention because the recitation is directed to instructions merely indicating a field of use or technological environment in which to apply a judicial exception, see MPEP 2106.05(h).
Accordingly, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea when considered as an ordered combination and as a whole.
Step 2B: This part of the eligibility analysis evaluates whether the claim, as a whole, amounts to
significantly more than the recited exception, i.e., whether any additional element, or combination of
additional elements, adds an inventive concept to the claim. MPEP 2106.05.
First, the additional elements considered as part of the preamble and the additional elements
directed to the use of computer technology are deemed insufficient to transform the judicial exception
to a patentable invention to a patentable invention because they generally link the judicial exception to
the technology environment, see MPEP 2106.05(h).
Second, the additional elements directed to mere application of the abstract idea or mere instructions to implement an abstract idea on a computer are deemed insufficient to transform the judicial exception to a patentable invention to a patentable invention because the limitations generally apply the use of a generic computer and/or process with the judicial exception, see MPEP 2106.05(f).
Third, the claims are directed to instructions merely indicating a field of use or technological environment in which to apply a judicial exception. The courts have found these types of limitations insufficient to transform the judicial exception to a patentable invention, see MPEP 2106.05(g).
Lastly, the claims directed to data gathering activity as noted above, are deemed directed to an insignificant extra-solution activity. The courts have found these types of limitations insufficient to
qualify as "significantly more", see MPEP 2106.05(g).
Furthermore, when considering evidence in view of Berkheimer v. HP, Inc., 881 F.3d 1360, 1368, 125 USPQ2d 1649, 1654 (Fed. Cir. 2018), see USPTO Berkheimer Memorandum (April 2018). Examiner notes Berkheimer: Option 2 - A citation to one or more of the court decisions discussed in MPEP § 2106.05(d}(II} as noting the well understood, routine, conventional nature of the additional element (s) (e.g., limitations directed to mere data gathering):
The courts have recognized the following computer functions as well understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity, see MPEP 2106.05(d).
The additional limitations, as analyzed, failed to integrate a judicial exception into a practical application at Step 2A and provide an inventive concept in Step 2B, per the analysis above. Thus, considering the additional elements individually and in combination and the claims as a whole, the additional elements do not provide significantly more than the abstract idea. This claim is not patent eligible. Therefore, in examining elements as recited by the limitations individually and as an ordered combination, as a whole, claims 1, 11, 17 do not recite what the courts have identified as "significantly more".
Furthermore, regarding dependent claims 2-10, which depend from claim 1, claims 12-16, which depend from claim 11, claims 18-20, which depend from claim 17, the claims are directed to a judicial exception (i.e., an abstract idea enumerated in the 2019 PEG, a law of nature, or a natural phenomenon) without significantly more as highlighted below in the claim limitations by evaluating the claim limitations under the Step2A and 2B:
Claim 2:
Incorporates the rejection of claim 1.
“an input handler configured to provide an indication of the determined split to the data router”
These recitations are deemed insufficient to transform the judicial exception to a patentable invention because the recitation is directed to instructions for mere data gathering or data output, see MPEP 2106.05(g).
Limitations directed to instructions for mere data gathering or data output cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B.
Claim 3:
Incorporates the rejection of claim 1.
“wherein each PE is associated with a PE convolution engine configured to perform a convolution on a respective portion of a tile stored in the associated PE data memory”
These recitations are deemed insufficient to transform the judicial exception to a patentable invention because the recitation is directed to instructions merely indicating a field of use or technological environment in which to apply a judicial exception, see MPEP 2106.05(h).
Limitations directed to mere instructions indicating a field of use or technological environment in which to apply a judicial exception cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B.
Claim 4:
Incorporates the rejection of claim 3.
“a systolic controller configured to control each of the PE convolution engines to perform the convolution on the respective portion of the tile stored in the associated PE data memory based on the split”
These recitations are deemed insufficient to transform the judicial exception to a patentable invention because the recitation is directed to instructions merely indicating a field of use or technological environment in which to apply a judicial exception, see MPEP 2106.05(h).
Limitations directed to mere instructions indicating a field of use or technological environment in which to apply a judicial exception cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B.
Claims 5, 13, 19:
Incorporates the rejections of claims 3, 12, 18, respectively.
“wherein the PE convolution engine is configured to perform the convolution on respective portions of multiple tiles stored in the associated PE data memory by reusing data in the associated PE data memory that overlaps the shared edge”
These recitations are deemed insufficient to transform the judicial exception to a patentable invention because the recitation is directed to instructions merely indicating a field of use or technological environment in which to apply a judicial exception, see MPEP 2106.05(h).
Limitations directed to mere instructions indicating a field of use or technological environment in which to apply a judicial exception cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B.
Claims 6, 14, 20:
Incorporates the rejections of claims 1, 11, 17, respectively.
“wherein the duplicated row or column of the shared edge includes shared edge duplicated in a first PE data memory associated with a first PE and the first tile and in a second PE data memory associated with a second PE and the second tile”
These recitations are deemed insufficient to transform the judicial exception to a patentable invention because the recitation is directed to instructions for mere data gathering or data output, see MPEP 2106.05(g).
Limitations directed to instructions for mere data gathering or data output cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B.
Claims 7, 15:
Incorporates the rejections of claims 1, 11, respectively.
“wherein the data router is further configured to transpose the plurality of tiles in the PE data memories by storing tile rows as columns in the PE data memories”
These recitations are deemed insufficient to transform the judicial exception to a patentable invention because the recitation is directed to instructions for mere data gathering or data output, see MPEP 2106.05(g).
Limitations directed to instructions for mere data gathering or data output cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B.
Claims 8, 16:
Incorporates the rejections of claims 1, 11, respectively.
“wherein each PE is further associated with a PE weight memory and wherein the data router is further configured to route weights to the PE weight memories based on the routing of the input tensor data to store the plurality of tiles in the PE data memories”
These recitations are deemed insufficient to transform the judicial exception to a patentable invention because the recitation is directed to instructions for mere data gathering or data output, see MPEP 2106.05(g).
Limitations directed to instructions for mere data gathering or data output cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B.
Claim 9:
Incorporates the rejection of claim 1.
“wherein the data router comprises a hardware-implemented algorithm”
These recitations are deemed insufficient to transform the judicial exception to a patentable invention because the recitation is directed to instructions merely indicating a field of use or technological environment in which to apply a judicial exception, see MPEP 2106.05(h).
Limitations directed to mere instructions indicating a field of use or technological environment in which to apply a judicial exception cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B.
Claim 10:
Incorporates the rejection of claim 1.
“wherein the systolic array comprises a scalable array of interconnected PEs”
These recitations are deemed insufficient to transform the judicial exception to a patentable invention because the recitation is directed to instructions merely indicating a field of use or technological environment in which to apply a judicial exception, see MPEP 2106.05(h).
Limitations directed to mere instructions indicating a field of use or technological environment in which to apply a judicial exception cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B.
Claim 12:
Incorporates the rejection of claim 11.
“further comprising: performing a convolution on the input tensor by performing, by a PE convolution engine associated with each PE, a convolution on respective portions of the input tiles stored in the associated PE data memory”
These recitations are deemed insufficient to transform the judicial exception to a patentable invention because the recitation is directed to instructions merely indicating a field of use or technological environment in which to apply a judicial exception, see MPEP 2106.05(h).
Limitations directed to mere instructions indicating a field of use or technological environment in which to apply a judicial exception cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B.
Claim 18:
Incorporates the rejection of claim 17.
“wherein the data router is further configured to route weights to ”
These recitations are deemed insufficient to transform the judicial exception to a patentable invention because the recitation is directed to instructions for mere data gathering or data output, see MPEP 2106.05(g).
“wherein each PE is associated with a PE convolution engine configured to perform a convolution on the input tensor by performing a convolution on respective portions of the plurality of tiles stored in the associated PE data memory with ”
These recitations are deemed insufficient to transform the judicial exception to a patentable invention because the recitation is directed to instructions merely indicating a field of use or technological environment in which to apply a judicial exception, see MPEP 2106.05(h).
Limitations directed to instructions for mere data gathering or data output or directed to instructions merely indicating a field of use or technological environment in which to apply a judicial exception cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B.
The dependent claims as analyzed above, do not recite limitations that integrated the judicial exception into a practical application. In addition, the claim limitations do not include additional elements that are sufficient to amount to significantly more than the judicial exception (Step-2B). Therefore, the claims do not recite any limitations, when considered individually or as a whole, that recite what have the courts have identified as "significantly more", see MPEP 2106.05; and therefore, as a whole the claims are not patent eligible. As shown above, the dependent claims do not provide any additional elements that when considered individually or as an ordered combination, amount to significantly more than the abstract idea identified. Therefore, as a whole, the dependent claims do not recite what have the courts have identified as "significantly more" than the recited judicial exception. Therefore, claims 2-10, 12-16, 18-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception and does not recite, when claim elements are examined individually and as a whole, elements that the courts have identified as "significantly more" than the recited judicial exception.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Yu et al. (U.S. Patent No. 11494321, hereinafter ‘Yu'), in view of Nama et al. (U.S. Patent No. 12321843B2, hereinafter 'Nama').
Regarding claim 1 and analogous claims 11, 17, Yu teaches A computing system, comprising:
a systolic array comprising an array of interconnected processing elements (PEs), each PE associated with a PE data memory configured to store at least a portion of a tensor (Yu Col. 18, Lines 33-50, FIG. 5 illustrates an example of a computing system 500 according to certain embodiments. In the illustrated example, computing system 500 includes a DMA engine 550, system memory 520, and one or more accelerators 502-1 to 502-m. Computing system 500 may include other components not specifically shown, such as a host processor. Accelerators 502-1 may be a neural network accelerator (e.g., a neural network processor or tensor processing unit), and may include a [a systolic array comprising an array of interconnected processing elements (PEs)] processing element array 510-1 (e.g., a systolic array), a state buffer 504-1, and a result buffer 512-1 as described above with respect to FIG. 3. Processing element array 510-1 may include an array of processing elements arranged in rows and columns. [each PE associated with a PE data memory configured to store at least a portion of a tensor] Each processing element is capable of performing a multiply-and-add operation. State buffer 504-1 may be used to store input data such as feature map values and weight values for processing element array 510-1, and/or may be used to store intermediate outputs that may be used in subsequent layers.); and
a data router configured to perform tensor tiling of an input tensor, the data router configured to:
determine a split of the input tensor into a plurality of tiles based on the array of interconnected PEs and dimensions of the input tensor (Yu Col. 20, Lines 45-57, The [a data router configured to perform tensor tiling of an input tensor] tensor copy instructions described above may be performed by a processing engine of the computing system, such as pooling engine 318 or activation engine 316 of accelerator 302. Pooling engine 318 may read from a first region of the state buffer (e.g., memory subsystem 304) and write the read data to a second region in the state buffer. The first region and the second region may have the same two-dimensional shape (e.g., the same number of partitions and the same number of bytes in each partition), or may have different two-dimensional shapes, such as different numbers of partitions and different numbers of elements (e.g., bytes) in each partition, but may have the same total size (number of partitions×number of elements in each partition).;
[Col. 23, Line 65-Col. 24, Line 11] According to techniques disclosed herein, before tensor 714 is written or loaded into state buffer 710 during period P1, a processing engine 730 (e.g., pooling engine 318) may execute a tensor copy instruction to read tensor 712 from state buffer 710, reshape tensor 712 into a tensor 720, and save tensor 720 to state buffer 710. determine a split of the input tensor into a plurality of tiles based on the array of interconnected PEs and dimensions of the input tensor Tensor 720 may have the same total size (and the same content) as tensor 712, but may have a shape different from a shape of tensor 712. For example, tensor 720 may include more partitions than tensor 712 but may have few elements in each partition compared with tensor 712. In one example, tensor 712 may have a shape of 32 partitions×256 elements/partition, whereas tensor 720 may have a shape of 128 partitions×64 elements/partition.); and
Yu fails to teach split the input tensor into the plurality of tiles, including a first tile and a second tile overlapping a shared edge, by routing the input tensor data to the PE data memories that store the plurality of tiles; wherein the overlapping shared edge includes a duplicated row or column of the shared edge; and change a size of the first tile and the second tile without re-accessing the duplicated row or column.
However Name teaches split the input tensor into the plurality of tiles, including a first tile and a second tile overlapping a shared edge, by routing the input tensor data to the PE data memories that store the plurality of tiles; wherein the overlapping shared edge includes a duplicated row or column of the shared edge; and change a size of the first tile and the second tile without re-accessing the duplicated row or column (Fig. 8B,
PNG
media_image1.png
405
1038
media_image1.png
Greyscale
[and split the input tensor into the plurality of tiles, including a first tile and a second tile overlapping a shared edge,]
Col 20 line 14-32,
For example, referring to the left section of FIG. 8B, the tensor 810 is a 36x36 tensor having four 20x20 overlapping tiles 830a, 830b, 830c, 830d. Thus, any two neighboring tiles in the tensor 810 have an overlap region, such as the overlap region 831 between the tiles 834a and 834b. A size of the overlap region 831 is 20x4. Thus, the overlap region 831 between the tiles 830a and 830b has a width of 4. Similarly, an overlap region between the tiles 830a and 830c has a height of 4. For purposes of ease of discussion, the tiles 830 of the tensor 810 are assumed to have an overlap of 4x4 (i.e., at least a height or a width of an overlap region is 4).
Now referring to the middle section of FIG. 8B, illustrated is a manner in which the tensor 810 is stored in the memory 140. For example, unlike the tiles 834 of the tensor 820 of FIG. SA, in FIG. 8B individual tiles 830 of the tensor 810 are materialized and stored in the memory 140 in an overlapping manner. Thus, for example, the overlap region 831 between the tiles 830a and 830b is stored merely once in the memory 140 [wherein the overlapping shared edge includes a duplicated row or column of the shared edge; and change a size of the first tile and the second tile without re-accessing the duplicated row or column].
Col 41 line 17-30,
FIG. 14A is a simplified diagram of a tile and an array level network usable in the configuration of FIG. 13, where the configurable units in the array are nodes on the array level network and are configurable to implement the processing graphs and various processing nodes of various sections discussed herein.
In this example, the array of configurable units 1400 includes a plurality of types of configurable units, which are to execute the various processing nodes of various sections of processing graphs discussed herein. The types of configurable units in this example, include Pattern Compute Units (PCUs), Pattern Memory Units (PMUs), Switch units (S), and Address Generation and Coalescing Units ( each includeing two address generators AG and a shared CU [by routing the input tensor data to the PE data memories that store the plurality of tiles]).
Yu and Name are considered to be analogous to the claimed invention because they are in the same field of machine learning. In view of the teachings of Yu, it would have been obvious for a person of ordinary skill in the art to apply the teachings of Name to Yu before the effective filing date of the claimed invention in order to store the tiles only once in memory (Name col 20 line 27-32, For example, unlike the tiles 834 of the tensor 820 of FIG. 8A, in FIG. 8B individual tiles 830 of the tensor 810 are materialized and stored in the memory 140 in an overlapping manner. Thus, for example, the overlap region 831 between the tiles 830a and 830b is stored merely once in the memory 140.).
Regarding claim 2, Yu, as modified by Nama, teaches The computing system of claim 1.
Yu and Nama are combinable for the same rationale as set forth above with respect to claim 1.
Yu teaches further comprising: an input handler configured to provide an indication of the determined split to the data router (Yu Col. 3, Lines 23-28, The [an input handler configured to provide an indication of the determined split to the data router] memory allocation, tensorization (converting a large tensor operation into multiple smaller tensor operations), and data transfer may be determined by a compiler that generates and schedules instructions to be executed by the neural network processor and other execution engines of a computing system to implement a neural network model.).
Regarding claim 3, Yu, as modified by Nama, teaches The computing system of claim 1.
Yu and Nama are combinable for the same rationale as set forth above with respect to claim 1.
Yu teaches wherein each PE is associated with a PE convolution engine configured to perform a convolution on a respective portion of a tile stored in the associated PE data memory (Yu Col. 10, Line 64-Col. 11, Line 8, Processing element array 310 is the computation matrix of accelerator 302. each PE is associated with a PE convolution engine configured to perform a convolution on a respective portion of a tile stored in the associated PE data memory Processing element array 310 can, for example, execute parallel integration, convolution, correlation, and/or matrix multiplication, among other things. Processing element array 310 may include multiple processing elements 311, arranged in rows and columns, such that results output by one processing element 311 can be input directly into another processing element 311. Processing elements 311 that are not on the outside edges of processing element array 310 thus can receive data to operate on from other processing elements 311, rather than from memory subsystem 304.).
Regarding claim 4, Yu, as modified by Nama, teaches The computing system of claim 3.
Yu and Nama are combinable for the same rationale as set forth above with respect to claim 1.
Yu teaches further comprising a systolic controller configured to control each of the PE convolution engines to perform the convolution on the respective portion of the tile stored in the associated PE data memory based on the split (Yu Col. 15, Lines 1-20, a systolic controller configured to control each of the PE convolution engines to perform the convolution on the respective portion of the tile stored in the associated PE data memory based on the split The compiler may traverse the data flow graph and perform shape inference on the neural network model, for example, to determine the sizes of the data used for each operation. The compiler may add, to the neural network model, operations for padding the input feature map for each input channel, based on parameters of a convolution operation, such as the size of an original input feature map, the size of a filter (e.g., kernel), the stride used for the convolution, the memory alignment, and the size of the processing element array. Optionally, the compiler may add to the neural network model operations for dividing input data into multiple partitions and dividing the convolution operation into multiple sub-operations. The compiler may map the operations of the neural network model to the computing system, such as memory subsystem 304 and processing element array 310 in accelerator 302, pooling engines 318, activation engines 316, DMA engines (not shown in FIG. 3), and the like, and generate and schedule instructions to be performed by these different components of the computing system.; [Col. 3, Lines 23-37] The memory allocation, tensorization (transforming a large tensor operation to multiple smaller tensor operations), and data transfer may be determined by a compiler that generates and schedules instructions to be executed by the neural network processor and other execution engines of a computing system to implement a neural network model. The instructions may generally include instructions for memory load operations that read input data (e.g., input feature maps) and static variables (e.g., weights, such as filter tensors for a convolutional neural network), instructions for computation operations that use the input data and the static variables to perform arithmetic operations, and memory save operations that save outputs (e.g., intermediate results) of the computation operations to the system memory to make room for other input tensors for other operations.).
Regarding claim 5 and analogous claims 13, 19, Yu, as modified by Nama, teaches The computing system of claim 3, The method of claim 12, and The NPU of claim 18, respectively.
Yu and Nama are combinable for the same rationale as set forth above with respect to claim 1.
Name teaches wherein the PE convolution engine is configured to perform the convolution on respective portions of multiple tiles stored in the associated PE data memory by reusing data in the associated PE data memory that overlaps the shared edge (Name Col 20 line 27-32, For example, unlike the tiles 834 of the tensor 820 of FIG. 8A, in FIG. 8B individual tiles 830 of the tensor 810 are materialized and stored in the memory 140 in an overlapping manner. Thus, for example, the overlap region 831 between the tiles 830a and 830b is stored merely once in the memory 140.
Col 26 line 29-36, The section 900 of FIG. 9B has been discussed with 190 generates the tile 924a, the data flow logic 286 causes respect to FIG. 9A. Section 930 of FIG. 9B has layers that are at least in part similar to the corresponding layers of section 900. For example, section 930 comprises a plurality of processing nodes 934, 938, 942, which includes an input processing node 934 configured to receive an input tensor 932, and convolve the input tensor 932 to generate an intermediate tensor 936 [memory by reusing data in the associated PE data memory that overlaps the shared edge].).
Regarding claim 6 and analogous claims 14, 20, Yu, as modified by Nama, teaches The computing system of claim 1, The method of claim 11, and The NPU of claim 17, respectively.
Yu and Nama are combinable for the same rationale as set forth above with respect to claim 1.
Nama teaches wherein duplicated row or column of the shared edge includes and the first tile and in a second PE data memory associated with a second PE and the second tile. (Name
Col 20 line 28-43, in FIG. 8B individual tiles 830 of the tensor 810 are materialized and stored in the memory 140 in an overlapping manner. Thus, for example, the overlap region 831 between the tiles 830a and 830b is stored merely once in the memory 140.
Thus, the left section of FIG. 8B and the middle section of FIG. 8B have same dimensions and configuration. For example, because the overlapping 20x20 tiles 830a, ... , 830d of the 36x36 tensor 810 are stored in the overlapping manner in the memory 140, the tensor 810 occupies 36x36 storage space in the memory 140. The reasons for storing the tiles 834 of the tensor 820 of FIG. 8A in a non-overlapping manner in the memory 140, while storing the tiles 830 of the tensor 810 of FIG. 8B in an overlapping manner in the memory 140, will be discussed herein in further detail in turn.
Col 20 line 44-55 The right section of FIG. 8B illustrates a notation which describes the manner in which the tensor 810 is materialized. For example, the notation corresponding to the tensor 810 includes a size 36x36(F), where "(F)" indicates that the tensor 810 has an actual or full size of 36x36 (as discussed) with respect to the left section of FIG. 8B). The notation corresponding to the tensor 810 further includes a size 20x20(T), which indicates that the tensor 810 has tiles of size 20x20. The notation corresponding to the tensor 810 further includes a size 36x36(M), which indicates that the tensor 810 has a size of 36x36 when stored in the memory 55 140, as discussed with respect to the middle section of FIG. 8B [wherein duplicated row or column of the shared edge includes].
Col 20 line 59-62, The notation corresponding to the tensor 810 further includes a size 4x4(MO), where "MO" indicates a size of overlap among the tiles, when the tiles are stored in the memory 140 [along the shared edge duplicated in a first PE data memory associated with a first PE and the first tile and in a second PE data memory associated with a second PE and the second tile]).
Regarding claim 7 and analogous claim 15, Yu, as modified by Nama, teaches The computing system of claim 1, The method of claim 11, respectively.
Yu and Nama are combinable for the same rationale as set forth above with respect to claim 1.
Name teaches wherein the data router is further configured to transpose the plurality of tiles in the PE data memories by storing tile rows as columns in the PE data memories (Nama Col 7 line 41-51, At operation 241, the compiler 106 compiles the application 204 to generate one or more configuration files 216. The configuration files 216 include a plurality of functions. Examples of functions in the plurality of functions include, but are not limited to, non-linearities like Rectified Linear 45 Unit (ReLU) and its variants (e.g., leaky ReLU), convolution, transpose convolution, hyperbolic tangent, sigmoid, and softmax, element-wise addition, matrix multiplication (e.g., General Matrix Multiply (GeMM)), layer normalization (e.g., batch normalization), loss functions like crossentropy, and tensor shape modifiers like transpose.
Col 20 line 28-43, in FIG. 8B individual tiles 830 of the tensor 810 are materialized and stored in the memory 140 in an overlapping manner. Thus, for example, the overlap region 831 between the tiles 830a and 830b is stored merely once in the memory 140.
Thus, the left section of FIG. 8B and the middle section of FIG. 8B have same dimensions and configuration. For example, because the overlapping 20x20 tiles 830a, ... , 830d of the 36x36 tensor 810 are stored in the overlapping manner in the memory 140, the tensor 810 occupies 36x36 storage space in the memory 140. The reasons for storing the tiles 834 of the tensor 820 of FIG. SA in a non-overlapping manner in the memory 140, while storing the tiles 830 of the tensor 810 of FIG. 8B in an overlapping manner in the memory 140, will be discussed herein in further detail in [by storing tile rows as columns in the PE data memories].
Col 30 line 6-13, For example, the layers in the sequence of layers of each of the sections 900 and 1000 can include one or more of convolution layers, max pooling layers, min pooling layers, average pooling layers, non-linearity layers, normalization layers, dropout layers, concatenation layers, transpose convolution layers, fully connected layers, softmax layers, and/or loss layers, although not all such operations are illustrated in FIG. 10A [transpose the plurality of tiles in the PE data memories].).
Regarding claim 8 and analogous claim 16, Yu, as modified by Nama, teaches The computing system of claim 1, The method of claim 11, respectively.
Yu and Nama are combinable for the same rationale as set forth above with respect to claim 1.
Yu teaches wherein each PE is further associated with a PE weight memory and wherein the data router is further configured to route weights to the PE weight memories based on the routing of the input tensor to store the plurality of tiles in the PE data memories (Yu Col. 11, Lines 35-50, An example of a each PE processing element 311 is illustrated in an inset diagram in FIG. 3. As illustrated by this example, processing element 311 can include a multiplier-accumulator circuit. Inputs from the left can include, for example, is further associated with a PE weight memory input data i and a weight value w, where the wherein the data router is further configured to route weights to the PE weight memories based on the routing of the input tensor to store the plurality of tiles in the PE data memories input data is a value taken from either a set of input data or a set of intermediate results, and the weight value is from a set of weight values that connect one layer of the neural network to the next. A set of input data can be, for example, an image being submitted for identification or object recognition, an audio clip being provided for speech recognition, a string of text for natural language processing or machine translation, or the current state of a game requiring analysis to determine a next move, among other things. In some examples, the input data and the weight value are output to the right, for input to the next processing element 311.).
Regarding claim 9, Yu, as modified by Nama, teaches The computing system of claim 1.
Yu and Nama are combinable for the same rationale as set forth above with respect to claim 1.
Yu teaches wherein the data router comprises a hardware-implemented algorithm (Yu Col. 3, Line 63-Col. 4, Line 24, According to certain embodiments, to reduce the overhead of performing a large number of DMA transfers, a compiler may, after the resource allocation and instruction generation of a typical compiling process, data router comprises a hardware-implemented algorithm analyze the local memory allocation and usage by the instructions and modify the instructions by, for example, replacing some DMA operations with tensor copy instructions that copy tensors within the local memory without involving the DMA controller. For example, at least some DMA operations for state buffer spilling and reloading may be replaced by tensor copy instructions that can reshape the dimensions (e.g., changing the number of partitions and the number of elements per partition but not the total number of elements) of the tensors to be spilled, and save the reshaped tensors in unused regions of the state buffer that may not have dimensions suitable for storing the original tensors. When the original tensor of a reshaped tensor needs to be used in subsequent instructions, a second tensor copy instruction may be used to change the reshaped tensor in the state buffer back to its original shape and save the tensor to the state buffer for use by the subsequent instructions. In some embodiments, the subsequent instructions may use the tensor with dimensions different from the dimensions of the original tensor and the reshaped tensor, and thus the second tensor copy instruction may read the reshaped tensor, reshape it into the desired dimensions, and save the tensor with the desired dimensions to the state buffer for use by the subsequent instruction. In this way, some DMA save instructions and DMA load instructions may be replaced by tensor copy instructions.).
Regarding claim 10, Nama, as modified by Nama, teaches The computing system of claim 1.
Yu and Nama are combinable for the same rationale as set forth above with respect to claim 1.
Yu teaches wherein the systolic array comprises a scalable array of interconnected Pes ([Col. 18, Lines 33-45] FIG. 5 illustrates an example of a computing system 500 according to certain embodiments. In the illustrated example, computing system 500 includes a DMA engine 550, system memory 520, and one or more accelerators 502-1 to 502-m. Computing system 500 may include other components not specifically shown, such as a host processor. Accelerators 502-1 may be a neural network accelerator (e.g., a neural network processor or tensor processing unit), and may include a systolic array processing element array 510-1 (e.g., a systolic array), a state buffer 504-1, and a result buffer 512-1 as described above with respect to FIG. 3. Processing element array 510-1 may comprises a scalable array of interconnected Pes include an array of processing elements arranged in rows and columns.).
Regarding claim 12, Yu, as modified by Nama, teaches The method of claim 11.
Yu and Nama are combinable for the same rationale as set forth above with respect to claim 1.
Yu teaches further comprising: performing a convolution on the input tensor by performing, by a PE convolution engine associated with each PE, a convolution on respective portions of the input tiles stored in the associated PE data memory (Yu Col. 10, Line 64-Col.11, Line 8, Processing element array 310 is the computation matrix of accelerator 302. performing a convolution on the input tensor by performing, by a PE convolution engine associated with each PE, a convolution on respective portions of the input tiles stored in the associated PE data memory Processing element array 310 can, for example, execute parallel integration, convolution, correlation, and/or matrix multiplication, among other things. Processing element array 310 may include multiple processing elements 311, arranged in rows and columns, such that results output by one processing element 311 can be input directly into another processing element 311. Processing elements 311 that are not on the outside edges of processing element array 310 thus can receive data to operate on from other processing elements 311, rather than from memory subsystem 304.).
Regarding claim 18, Yu, as modified by Nama, teaches The NPU of claim 17.
Yu and Nama are combinable for the same rationale as set forth above with respect to claim 1.
Yu teaches wherein the data router is further configured to route weights to (Yu Col. 11, Lines 35-50, An example of a processing element 311 is illustrated in an inset diagram in FIG. 3. As illustrated by this example, processing element 311 can include a multiplier-accumulator circuit. Inputs from the left can include, for example, input data i and a weight value w, where the wherein the data router is further configured to route weights to the PE weight memories based on the routing of the input tensor to store the plurality of tiles in the PE data memories input data is a value taken from either a set of input data or a set of intermediate results, and the weight value is from a set of weight values that connect one layer of the neural network to the next. A set of input data can be, for example, an image being submitted for identification or object recognition, an audio clip being provided for speech recognition, a string of text for natural language processing or machine translation, or the current state of a game requiring analysis to determine a next move, among other things. In some examples, the input data and the weight value are output to the right, for input to the next processing element 311.); and
wherein each PE is associated with a PE convolution engine configured to perform a convolution on the input tensor by performing a convolution on respective portions of the (Yu Col. 10, Line 64-Col.11, Line 8, Processing element array 310 is the computation matrix of accelerator 302. Processing element array 310 can, for example, execute parallel integration, convolution, correlation, and/or matrix multiplication, among other things. Processing element array 310 may include multiple processing elements 311, arranged in rows and columns, such that results output by one processing element 311 can be input directly into another processing element 311.).
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALFREDO CAMPOS whose telephone number is (571)272-4504. The examiner can normally be reached 7:00 - 4:00 pm M - F.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michael J. Huntley can be reached at (303) 297-4307. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ALFREDO CAMPOS/Examiner, Art Unit 2129
/MICHAEL J HUNTLEY/Supervisory Patent Examiner, Art Unit 2129