Notice of Pre-AIA or AIA Status
The present application 89/565,996, filed on 11/30/2023 (or after March 16, 2013), is being examined under the first inventor to file provisions of the AIA (First Inventor to File).
In the event the determination of the status of the application as subject to AIA 35
U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
This application is a 371 of PCT/CN2021/133099 11/25/2021
DETAILED ACTION
Response to Amendment
Claims 1-3,5-11,13,15,17-19,21,23,26-28 are pending, canceled claim 4,12,14,16,20,22,24-25 in this application.
Examiner acknowledges applicant’s amendment filed on 8/26/2026
Drawings
The Drawings filed on 11/30/2023 are acceptable for examination purpose.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 3/4/2025, 11/30/2023 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner, PTO-1446 mailed on 5/26/2026
Response to Arguments
Applicant's arguments filed 8/26/2026 with respect to claims 1-13,15,17-21,23 have been fully considered but they are not persuasive, for examiner’s response, see discussion below:
a)At page 10-11, claim 1,9,17, applicant argues:
Applicant submits that the claims integrate any such exception into a practical application at Step 2A Prong Two. The Office Action has alleged that the additional elements "fail to apply the exception with a particular machine. Neural network acceleration hardware that computes at least one of an inner product or convolution over a fixed number of operands per computation cycle is specifically identifiable and is the very machine fed by the claimed slices; it is not a general-purpose computer applying conventional computer functions………….
Second, the claims reflect an improvement to the functioning of a computer. The
specification identifies a technical problem: in the traditional approach, padded zeroes
consume computation resources without contributing to the output………….The components producing this improvement, the per-cycle process
capacity, the remainder-based padding count, slicing sized equal to the process capacity, and sequential feeding of the acceleration hardware……………
Examiner’s response:
Examiner submits that the pending claims (as amended 8/26/2026) should pass the test set forth in the 2019 Revised Patent Subject Matter Eligibility Guidance published on January 7, 2019 (84 Fed. Reg. 50), as updated October 2019, referred to herein as the PEG 2019. Applicant will focus on Prong Two of Step 2A, in evaluating the pending claims using this section of the test set forth in the PEG 2019
As explained in the 2019 PEG, the evaluation of Prong Two of Step 2A requires the use of the considerations (e.g. improving technology, effecting a particular treatment or prophylaxis, implementing with a particular machine, etc.) identified by the Supreme Court and the Federal Circuit, to ensure that the claim as a whole “integrates [the] judicial exception into a practical application [that] will apply, rely on, or use the judicial exception in a manner that imposes a meaningful limit on the judicial exception, such that the claim is more than a drafting effort designed to monopolize the judicial exception”. These considerations are set forth in the 2019 PEG, MPEP 2106.05(a) through (c), and MPEP 2106.05(e) through (h). Note, a specific way of achieving a result is not a stand-alone consideration in Step 2A Prong Two. However, the specificity of the claim limitations is relevant to the evaluation of several considerations including the use of a particular machine, particular transformation and whether the limitations are mere instructions to apply an exception. If the claim integrates the judicial exception into a practical application based upon evaluation of these considerations, the additional limitations impose a meaningful limit on the judicial exception, and the claim is eligible at Step 2A.
For example, if the additional limitations as amended 8/26/2026, (process capacity indicative of a number of data elements the process engine processes in one computation cycle; wherein the process engine is neural network acceleration hardware to calculate at least one of an inner product or convolution of data elements sized equal to the process capacity) do not provide “improvement to another technology or technical field”, for example “calculate……….step itself appears “mental” process but using generic computers and covers mental process of abstract idea(s) because they cover concepts performed in the human mind, including observation, evaluation, judgement and opinion. See MPEP 2106.04(a)(2).
Taking the claim 1,9,17 (8/26/2026) elements separately, the functions performed add nothing that is not already present when the limitations are considered separately. For example, claim 1 does not purport to improve the functioning of the data structure of memory layout, nor does claim 1,9,17 effect an improvement in any other technology or technical field., but mere manipulation and/or data calculation. Instead, claim 1,9,17 (as amended 8/26/2026) amounts to nothing significantly more than an instruction to apply the abstract idea using generic computer components performing routine computer functions That is not enough to transform an abstract idea into a patent-eligible invention. See Alice, 573 US at 225-26; see also Inventor Holdings, LLC v. Bed Bath & Beyond, Inc., 876 F.3d 1372,1378 Fed.Cir.2017) (sequence of receiving, analyzing, modifying, generating, displaying, and transmitting data recited an abstraction)
examiner applies the arguments above of claim 9,17, and claims 2-3,5-8,10-11,13-,15,18-19,21,23,26-28 depend from claim 1,9,17 as such, the pending claims fail Prong Two of Step 2A-2B of the PEG 2019. Therefore, claims 1-20 are rejection under 35 U.S.C. § 101
b)At page 12-13, claim 1, applicant argues:
as amended claim 1, requires slices "sized equal to the process capacity," which the claim defines as the number of data elements the process engine processes in one computation cycle. Ross's tile size is fixed by kernel geometry and is entirely independent of any per-cycle throughput of a process engine. Ross also does not disclose weight data and activation data stored in an NHWC memory layout as set forth in claim 1.
Accordingly, Ross does not teach or suggest the slicing of "all weight data elements belonging to the filter and zeroes padded after the last element of the weight data into weight data slices sized equal to the process capacity"
Examiner’s response:
As best understood by the examiner, the prior art of Ross is directed to transformation of input values convolution particularly weight values by kernel to generate respective output values (Ross: Abstract), Ross teaches kernel formed from one or more dimensional matrix and the matrix includes kernel values (Ross: fig 1-2, 0053). Ross teaches spatial locality convolution engine (fig 1, element 120) defining the data structure having kernel size and respective values(fig 1-2)
The prior art of Guo is directed to memory layouts and conversion particularly in a neural network interface environment that including a batches, height, width and channel or NHWC layout (Guo: Abstract), and these batches based on the input data equal to or greater than the defined Nt threshold (Guo: 001400015). It is however, noted that Ross does not teach weight data slices sized equal to the process capacity, although Ross teaches convolution operation (fig 1) to generate output activations particularly kernel values (element 114) in neural network environment (fig 1-2, 0053-0054). On the other hand, Guo teaches s convolution of data element with respect to kernel computation to generate feature height, refers to feature width performance using the NCHW format, while determining different memory size using thresholds (fig 1,fig 3 0018, 0021-0023, 0028-0030) for example Ct, Nt channel, batch thresholds respectively in NCHW layout that determines and calculates size, process capacity as detailed in 0029-0030)
PNG
media_image1.png
153
243
media_image1.png
Greyscale
PNG
media_image2.png
149
160
media_image2.png
Greyscale
It would have been obvious to a person of ordinary skill in the art at the time of filing the claimed invention memory layouts and conversion particularly for a neural network (NN) of different memory layouts stored in multi-dimensional NN kernel of Guo into coevolve with input tensor particularly in selecting dimension(s) of the kernel spatial locality transform of matrices of Ross et al., because both Ross, Guo directed to memory layout and convolution by kernel (Ross: fig 2, Abstract, 0086-0087; Guo: Abstract, fig 2, 0024) and they both are from the same field of endeavor. Because both Ross, Guo teaches neural network kernel computations, it would have been obvious to one skilled in the art to substitute and/or modify one method for the other memory layout, simulation defining threshold of respective data size(s) particularly kernel computation using expressions (Guo: para 0025) and defining thresholds of particular values in NCHW format (Guo: 0029) to achieve the predictable optimum performance of computations in memory layouts including selecting memory layout based on thresholds from performance simulations of the neural network environment (Guo: 0013-0014).. The exemplary rationales that may support prima facie conclusion of obviousness includes (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention- -KSR, 550 US at 398.
examiner applies the arguments above of claim 9,17, and claims 2-3,5-8,10-11,13-,15,18-19,21,23,26-28 depend from claim 1,9,17
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-13,15,17-21,23 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The judicial exception is not integrated into a practical application.
Claim 1-20 is/are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The judicial exception is not integrated into a practical application. The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. The eligibility analysis in support of these findings is provided below, in accordance with the 2019 Revised Patent Subject Matter Eligibility Guidance, Federal Register (84 FR 50) on January 7, 2019 hereinafter 2019 PEG
Step 1. In accordance with Step 1 of the eligibility inquiry (as explained in MPEP 2106), it is noted that the method of claim 1,9,15, directed to one of the eligible categories of subject matter and therefore satisfy Step 1.
Step 2A. In accordance with Step 2A prong one of the 2019 PEG, the limitations reciting the abstract idea are highlighted, and the limitations directed to additional elements are highlighted, as set forth in exemplary claim 1
Claim 1,9,17
“interface circuitry to receive weight data and activation data, the weight data and the activation data stored in a batch-height-width-channel (NHWC) memory layout; instructions: and
at least one processor circuitry to execute the instructions to:
determine a process capacity of a process engine process capacity indicative of a number of data elements the process engine processes in one computation cycle;
determine an input channel size;
in response to the input channel size not being an integer multiple of the process capacity,
pad a number of zeroes after a last element of weight data belonging to a filter and a last element of corresponding activation data respectively, wherein the number equals to an absolute difference between the process capacity of process engine and a remainder of a product of the input channel size and a kernel width and a kernel height of the filter divided by the process capacity of process engine,
slice all weight data elements belonging to the filter and zeroes padded after the last element of the weight data into weight data slices in a scale of the process capacity, and corresponding activation data elements and zeroes padded after the last element of the corresponding activation data into corresponding activation data slices sized equal to the process capacity, and
feed the process engine with each weight data slice and a corresponding activation data slice sequentially wherein the process engine is neural network acceleration hardware to calculate at least one of an inner product or convolution of data elements sized equal to the process capacity”, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components. For example “input channel size, pad a number of zeroes, weight data, number equals to an absolute difference, kernel width, kernel height, slice all weight data, process capacity, feed the process engine, activation data slice…..”, appears mere data structure manipulation in the context of this claim encompasses the user thinking mere data gathering of the process capacity.
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas set forth in the 2019 PEG. Accordingly, the claim recites an abstract idea.
With respect to Step 2A prong two of the 2019 PEG, the judicial exception is not integrated into a practical application. The additional elements are directed to method steps, however, these elements fail to integrate the abstract idea into a practical application because they fail to provide an improvement to the functioning of a computer or to any other technology or technical field, fail to apply the exception with a particular machine, fail to apply the judicial exception to effect a particular memory data structure of channel size, kernel width, kernel height, filter divided by the process capacity, convolution of data elements sized to effect a transformation of a particular article to a different state or thing, and fail to apply/use the abstract idea in a meaningful way beyond generally linking the use of the judicial exception to a particular technological environment.
Furthermore, although these elements have been fully considered, they are directed to the use of generic computing elements (fig 7, para 76-84,89-90, of the instant specification make it clear that the disclosed functionality is implemented on well-known computing systems and general purpose computing devices) to perform the abstract idea, which is not sufficient to amount to a practical application (as noted in the 2019 PEG) and is amount to simply saying "apply it" using a general purpose computer, which merely serves to tie the abstract idea to a particular technological environment computer based operating environment) by using the computer as a tool to perform the abstract idea.
Claim 1,9,17 (as amended 8/26/2026), limitation “calculate at least one of an inner product or convolution of data elements………………process capacity”, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer component(s) such as by the processor. That is, other than reciting “by a processor”, nothing in the claim element precludes the step from practically being performed in the mind. For example, but for the “by a processor” calculate at least one of an inner product or convolution of data elements in the context of this claim limitation encompasses the user manually calculating sized equal to the process capacity and like, If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “mental processes” grouping of abstract ideas. Accordingly, the claim 1,9,17, recites an abstract idea, which have been determined to be extra-solution activity that does not impose any meaningful limits on practicing the abstract idea. See MPEP 2106.05(b)(I). Even in combination, the additional details recited in these claims do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea.
Since the analysis of Step 2A prong one and prong two results in the conclusion that the claims are directed to an abstract idea, additional analysis under Step 2B of the eligibility inquiry must be conducted in order to determine whether any claim element or combination of elements amount to significantly more than the judicial exception.
Step 2B. The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. The additional method limitations are directed to a generic computer, at a very high level of generality and without imposing meaningful limitations on the scope of the claim. In addition fig 7, para 76-84,89-90 of the instant specification describe generic off-the-shelf computer-based elements for implementing the claimed invention which does not amount to significantly more than the abstract idea and is not enough to transform an abstract idea into eligible subject matter. Such generic, high-level, and nominal involvement of a computer or computer-based elements for carrying out the invention merely serves to tie the abstract idea to a particular technological environment, which is not enough to render the claims patent-eligible, as noted at pg. 74624 of Federal Register/Vol. 79, No. 241, citing Alice, which in turn cites Mayo. Further, See, e.g., Alice Corp. Pty. Ltd. v. CLS Bank Int'l, 134 S. Ct. 2347, 2359-60, 110 USPQ2d 1976, 1984 (2014). See also OIP Techs. v. Amazon.com, 788 F.3d 1359, 1364, 115 USPQ2d 1090, 1093-94 (Fed. Cir. 2015) ("Just as Diehr could not save the claims in Alice, which were directed to 'implement[ing] the abstract idea of intermediated settlement on a generic computer', it cannot save O/P's claims directed to implementing the abstract idea of price optimization on a generic computer.") (citations omitted). See also, Affinity Labs of Texas LLC v. DirecTV LLC, 838 F.3d 1253, 1257-1258 (Fed. Cir. 2016) (mere recitation of a GUI does not make a claim patent-eligible); Intellectual Ventures I LLC v. Capital One Bank, 792 F.3d 1363, 1370 (Fed. Cir. 2015) ("the interactive interface limitation is a generic computer element". )the additional elements are broadly applied to the abstract idea at a high level of generality ("similar to how the recitation of the computer in the claims in Alice amounted to mere instructions to apply the abstract idea of intermediated settlement on a generic computer,") as explained in MPEP § 2106.05(f)) and they operate in a well-understood, routine, and conventional manner.
MPEP § 2106.05 (d)(II) sets forth the following:
The courts have recognized the following computer functions as well-understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g. at a high level of generality) as insignificant extra-solution activity.
Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec...; TLI Communications LLC v. AV Auto. LLC...; OIP Techs., Inc., v. Amazon.com, Inc... ; buySAFE, Inc. v. Google, Inc...;
Performing repetitive calculations, Flook ... ; Bancorp Services v. Sun Life...;
Electronic recordkeeping, Alice Corp...; Ultramercial... ;
Storing and retrieving information in memory, Versata Dev. Group, Inc. v. SAP Am., Inc...;
Electronically scanning or extracting data from a physical document, Content Extraction and Transmission, LLC v. Wells Fargo Bank...; and
A web browser's back and forward button functionality, Internet Patent Corp. v. Active Network, Inc.
Courts have held computer-implemented processes not to be significantly more than an abstract idea (and thus ineligible) where the claim as a whole amounts to nothing more than generic computer functions merely used to implement an abstract idea, such as an idea that could be done by a human analog (i.e., by hand or by merely thinking).
claim 2,10,18, further elaborates “wherein a weight data slice comprises weight data elements from one or more data groups belonging to the filter, and each data group comprises weight data elements of the input channel size”, which have been determined to be extra-solution activity that does not impose any meaningful limits on practicing the abstract idea. See MPEP 2106.05(b)(I). Even in combination, the additional details recited in these claims do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea.
Claim 3,11,19, further elaborates “wherein one or more of the at least one processor circuit is to perform the padding slicing, and feeding operations on weight data belonging to a next filter and corresponding activation data”, which have been determined to be extra-solution activity that does not impose any meaningful limits on practicing the abstract idea. See MPEP 2106.05(b)(I). Even in combination, the additional details recited in these claims do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea.
Claim 4,12,20 (cancelled)
Claim 5,13,21, further elaborates “wherein the process capacity of the process engine is 16, the input channel size is 8, and both the kernel width and the kernel height of the filter are 3, and one or more of the at least one processor circuit is to
:pad 8 zeroes after a last element of weight data belonging to the filter and a last element of corresponding activation data respectively;
slice all 72 weight data elements belonging to the filter and 8 zeroes padded after the last element of the weight data into 5 weight data slices and corresponding 72 activation data elements and 8 zeroes padded after the last element of the corresponding activation data into corresponding 5 activation data slices, each of the 5 weight data slices and the corresponding 5 activation data slices comprising 16 data elements; and
feed the process engine with each of the 5 weight data slices and a corresponding activation data slice sequentially in 5 computation cycles”, which have been determined to be extra-solution activity that does not impose any meaningful limits on practicing the abstract idea. See MPEP 2106.05(b)(I). Even in combination, the additional details recited in these claims do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea.
Claim 6, further elaborates “wherein a memory utilization ratio is 90%”, which have been determined to be extra-solution activity that does not impose any meaningful limits on practicing the abstract idea. See MPEP 2106.05(b)(I). Even in combination, the additional details recited in these claims do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea.
Claim 7,15,23, further elaborates “wherein the process capacity of the process engine is 16, the input channel size is 24, and both the kernel width and the kernel height of the filter are 3, and one or more of the at least one processor circuit
pad 8 zeroes after a last element of weight data belonging to the filter and a last element of corresponding activation data respectively;
slice all 216 weight data elements belonging to the filter and 8 zeroes padded after the last element of the weight data into 14 weight data slices and corresponding 216 activation data elements and 8 zeroes padded after the last element of the corresponding activation data into corresponding 14 activation data slices, each of the 14 weight data slices and the corresponding 14 activation data slices comprising 16 data elements; and
feed the process engine with each of the 14 weight data slices and a corresponding activation data slice sequentially in 14 computation cycles”, which have been determined to be extra-solution activity that does not impose any meaningful limits on practicing the abstract idea. See MPEP 2106.05(b)(I). Even in combination, the additional details recited in these claims do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea.
claim 8, further elaborates “wherein a memory utilization ratio is 96.42%”, which have been determined to be extra-solution activity that does not impose any meaningful limits on practicing the abstract idea. See MPEP 2106.05(b)(I). Even in combination, the additional details recited in these claims do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea.
The examiner suggests that the applicant review the specification of the instant application to find further teachings that if recited in said claims, may provide significantly more to the judicial exception. Such elements / limitations that can be considered as significantly more recite an improvement to another technology or technical field, an improvement to the functioning of a computer itself, or meaningful limitations beyond generally linking the use of an abstract idea to a particular technological environment.
Claim 14 (canceled)
Claim 16 (canceled)
Claim 22 (canceled)
Claim 24-25 (canceled)
As to claim 26., further elaborates “wherein the weight data is stored in the NHWC memory layout in a first memory and the corresponding activation data is stored in the NHWC memory layout in a second memory, and wherein the processor circuitry is to form the weight data slices and the corresponding activation data slices without relocating the weight data elements or the corresponding activation data elements within the first memory or the second memory”, which have been determined to be extra-solution activity that does not impose any meaningful limits on practicing the abstract idea. See MPEP 2106.05(b)(I). Even in combination, the additional details recited in these claims do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea.
As to claim 27, further elaborates “wherein a weight data slice includes a portion of weight data elements of a first data group belonging to the filter and a portion of weight data elements of a second data group belonging to the filter, each data group including weight data elements of the input channel size, and wherein the process engine is to receive the weight data slice and the corresponding activation data slice in a single computation cycle”, which have been determined to be extra-solution activity that does not impose any meaningful limits on practicing the abstract idea. See MPEP 2106.05(b)(I). Even in combination, the additional details recited in these claims do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea.
As to claim 28, further elaborates “wherein the number of zeroes is padded after the last element of the weight data belonging to the filter without zeroes being padded between data groups belong to the filter, each data group including weight data elements of the input channel size”, which have been determined to be extra-solution activity that does not impose any meaningful limits on practicing the abstract idea. See MPEP 2106.05(b)(I). Even in combination, the additional details recited in these claims do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1-3,5-11,13,15,17-19,21,23,26-28 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ross et al., (hereafter Ross), US Pub. No. 2020/0159814 published May, 2020 in view of Guo, US Pub. No. 2020/0349424 published Nov, 2020
As to Claim 1,9,17, .Ross teaches a system which including an apparatus, comprising (Ross: fig 23)
“interface circuitry (0264 – hardware engine comprise dedicated circuitry or logic) to receive weight data and activation data, the weight data (Ross: 0240 – Ross teaches weight data stored in a particular multiplier) and the activation data stored ( Ross: 0155)in a batch-height-width-channel (NHWC) memory layout; instructions(Ross: fig 1, 0052-0053, ,fig 11, 0079, 0155-0158, 0198 – Ross teaches memory layout data structure with respect to current tile, row or column, width/height and depth of a matrix of the input tensor element 102 used in selecting the data values stored. Ross teaches values in all positions are generated as control pattern square matrix having width, height of the kernel as detailed in 0158, it is noted that data storage in multi-dimensional array identifying height and width information stored in the channels and
PNG
media_image3.png
260
233
media_image3.png
Greyscale
PNG
media_image4.png
179
235
media_image4.png
Greyscale
“processor circuitry to execute the instructions to: (Ross: fig 23, 0256-0257 – Ross teaches computer system defining instructions element 2324)
“determine a process capacity of a process engine” (Ross: fig 1, 0052-0053);
“determine an input channel size” (Ross: fig 1, 0059 – Ross teaches multiple tiles defining each channel size);
PNG
media_image5.png
169
142
media_image5.png
Greyscale
“in response to the input channel size not being an integer multiple of the process capacity” (Ross: fig 2, 0053, 0058-0060 – Ross teaches multiple channel size(s) divides the matrix of each input channel for example as shown in fig 1 where each input channel determines the size, values with respect to column-major order, diagonal-major order of respective layer),
“pad a number of zeroes after a last element of weight data belonging to a filter and a last element of corresponding activation data respectively, (Ross: 0054,0057-0058,0067, 0098, fig 4, 0104-0106 – Ross teaches pad the input tensor ith respect to size and generating the respective value reaches every input value in the input tensor reaches including two dimensional matrix of the input tensor, input with padding, it should be noted that padding values may be zero or null or some other values because input tensor has largest index value);
PNG
media_image6.png
129
269
media_image6.png
Greyscale
wherein the number equals to an absolute difference between the process capacity of process engine and a remainder of a product of the input channel size and a kernel width and a kernel height of the filter divided by the process capacity of process engine (Ross: ,fig 4, 0105-106, fig 5C-6, 0116-0118 – Ross teaches multiple input channels are convolved with the kernel and generate an output features, and the input stream generator pads each input channel with respect to the channel size and kernel width)
PNG
media_image7.png
248
185
media_image7.png
Greyscale
PNG
media_image8.png
96
170
media_image8.png
Greyscale
“slice all weight data elements belonging to the filter and zeroes padded after the last element of the weight data into weight data slices sized equal to the process capacity, and corresponding activation data elements and zeroes padded after the last element of the corresponding activation data into corresponding activation data slices sized equal the process capacity” (Ross: fig 5-6, fig 9, 0054,0059,0085,0088-0089,0116 – Ross teaches slice of respective weights associated with the two dimensional kernels of different sizes and convolution operation is performed on each of channel with sub-filter or filter kennel, applying different kernel weights to different parts of inputs, further position of the input values, padding of the width of the matrix determines the size of the kernel as detailed in 0088-0089)
“feed the process engine with each weight data slice and a corresponding activation data slice sequentially” (Ross: fig 13A-B, 0176-0177,0183-0184 – Ross teaches pattern processing generating engine generates position of the row, column and the respective identified tile size of each dimension in the input tensor, and the weight data indicates the tile size)
PNG
media_image9.png
344
224
media_image9.png
Greyscale
PNG
media_image10.png
331
230
media_image10.png
Greyscale
It is however, noted hat Ross does not teach “process engine processes in one computation cycle”, “wherein the process engine is neural network acceleration hardware to calculate at least one of an inner product or convolution of data elements sized equal to the process capacity”, although Ross teaches convolution operation (fig 1) to generate output activations particularly kernel values (element 114) in neural network environment (fig 1-2, 0053-0054). On the other hand, Guo teaches “process engine processes in one computation cycle” (Guo: fig 2, 0024-0025 – Guo teaches kernel computation using NCHW or height, width and channel particularly single instruction multiple data instruction selection of respective cycle along with floating point in layouts used for kernel computation(s)
PNG
media_image11.png
159
252
media_image11.png
Greyscale
The prior art of Guo teaches “wherein the process engine is neural network acceleration hardware to calculate at least one of an inner product or convolution of data elements sized equal to the process capacity” (Guo: fig 1,fig 3 0018, 0021-0023, 0028-0030 – Guo teaches convolution of data element with respect to kernel computation to generate feature height, refers to feature width performance using the NCHW format, while determining different memory size using thresholds for example Ct, Nt channel, batch thresholds respectively in NCHW layout that determines and calculates size, process capacity as detailed in 0029-0030)
PNG
media_image1.png
153
243
media_image1.png
Greyscale
PNG
media_image2.png
149
160
media_image2.png
Greyscale
It would have been obvious to a person of ordinary skill in the art at the time of filing the claimed invention memory layouts and conversion particularly for a neural network (NN) of different memory layouts stored in multi-dimensional NN kernel of Guo into coevolve with input tensor particularly in selecting dimension(s) of the kernel spatial locality transform of matrices of Ross et al., because both Ross, Guo directed to memory layout and convolution by kernel (Ross: fig 2, Abstract, 0086-0087; Guo: Abstract, fig 2, 0024) and they both are from the same field of endeavor. Because both Ross, Guo teaches neural network kernel computations, it would have been obvious to one skilled in the art to substitute and/or modify one method for the other memory layout, simulation defining threshold of respective data size(s) particularly kernel computation using expressions (Guo: para 0025) and defining thresholds of particular values in NCHW format (Guo: 0029) to achieve the predictable optimum performance of computations in memory layouts including selecting memory layout based on thresholds from performance simulations of the neural network environment (Guo: 0013-0014)
As to claim 2,10,18, the combination of Ross, Guo disclosed “wherein a weight data slice comprises weight data elements from one or more data groups belonging to the filter, and each data group comprises weight data elements of the input channel size” (Ross: fig 13A-13B, 0178-0179,0183-0184).
As to Claim 3,11,19, the combination of Ross, Guo disclosed “wherein one or more the at least one processor circuit is to perform the padding slicing, and feeding operations on weight data belonging to a next filter and corresponding activation data” (Ross: fig 4-5, 0104-0108).
PNG
media_image12.png
307
224
media_image12.png
Greyscale
Claim 4,12,20, (cancelled)
As to Claim 5,13,21,the combination of Ross, Guo disclosed “wherein the process capacity of the process engine is 16, the input channel size is 8, and both the kernel width and the kernel height of the filter are 3, and one or more of the at least one processor circuit is to (Ross: Abstract, fig 1)
:”pad 8 zeroes after a last element of weight data belonging to the filter and a last element of corresponding activation data respectively” (Ross: fig 4,0105-017);
“slice all 72 weight data elements belonging to the filter and 8 zeroes padded after the last element of the weight data into 5 weight data slices and corresponding 72 activation data elements and 8 zeroes padded after the last element of the corresponding activation data into corresponding 5 activation data slices, each of the 5 weight data slices and the corresponding 5 activation data slices comprising 16 data elements” (Ross: fig 13A-13B, 0176-0180); and
“feed the process engine with each of the 5 weight data slices and a corresponding activation data slice sequentially in 5 computation cycles” (Ross: 0198-0200).
As to Claim 6, the combination of Ross, Guo disclosed “wherein a memory utilization ratio is 90% (Ross: 0165,0200-0201).
As to Claim 7,15,23, the combination of Ross, Guo disclosed “wherein the process capacity of the process engine is 16, the input channel size is 24, and both the kernel width and the kernel height of the filter are 3, and one or more of the at least one processor circuit is to” (Ross: Abstract, fig 1, fig 4):
pad 8 zeroes after a last element of weight data belonging to the filter and a last element of corresponding activation data respectively” (Ross: fig 4,5A, 0105-0110);
“slice all 216 weight data elements belonging to the filter and 8 zeroes padded after the last element of the weight data into 14 weight data slices and corresponding 216 activation data elements and 8 zeroes padded after the last element of the corresponding activation data into corresponding 14 activation data slices, each of the 14 weight data slices and the corresponding 14 activation data slices comprising 16 data elements” Ross: fig 13A-13B, 0152-0159,0176-0180; and
feed the process engine with each of the 14 weight data slices and a corresponding activation data slice sequentially in 14 computation cycles” (Ross: 0198-0203).
As to claim 8, the combination of Ross, Guo disclosed “wherein a memory utilization ratio is 96.42%” (Ross: 0165,0186-0187,0200-0201.
Claim 14 (canceled)
Claim 16 (canceled)
Claim 22 (canceled)
Claim 24-25 (canceled)
As to claim 26, the combination of Ross, Guo disclosed “wherein the weight data is stored in the NHWC memory layout in a first memory and the corresponding activation data is stored in the NHWC memory layout in a second memory,(Guo: fig 1, 0018-0019) and wherein the processor circuitry is to form the weight data slices and the corresponding activation data slices without relocating the weight data elements or the corresponding activation data elements within the first memory or the second memory” (Geo: 0025-0027).
As to claim 27, the combination of Ross, Guo disclosed “wherein a weight data slice includes a portion of weight data elements of a first data group belonging to the filter and a portion of weight data elements of a second data group belonging to the filter (Ross: fig 1-2, 0080-0084) each data group including weight data elements of the input channel size, and wherein the process engine is to receive the weight data slice (Ross: 0089-0091). On the other hand, Guo disclosed “activation data slice in a single computation cycle” (Guo: fig 2, 0024-0025)
As to claim 28, the combination of Ross, Guo disclosed “wherein the number of zeroes is padded after the last element of the weight data belonging to the filter without zeroes being padded between data groups belong to the filter, each data group including weight data elements of the input channel size” (Ross: 0098-0099, 0105-0107)
Conclusion
The prior art made of record
a. US Pub. No. 2020/0159814
b. US Pub. No. 2020/0349424
Examiner's Note: Examiner has cited particular columns and line numbers in the references applied to the claims above for the convenience of the applicant. Although the specified citations are representative of the teachings of the art and are applied to specific limitations within the individual claim, other passages and figures may apply as well. It is respectfully requested from the applicant in preparing responses, to fully consider the references in entirety as potentially teaching all or part of the claimed invention, as well as the context of the passage as taught by the prior art or disclosed by the Examiner.
SEE MPEP 2141.02 [R-5] VI. PRIOR ART MUST BE CONSIDERED IN ITS ENTIRETY, INCLUDING DISCLOSURES THAT TEACH AWAY FROM THE CLAIMS: A prior art reference must be considered in its entirety, i.e., as a whole, including portions that would lead away from the claimed invention. W.L. Gore & Associates, Inc. v. Garlock, Inc., 721 F.2d 1540, 220 USPQ 303 (Fed. Cir. 1983), cert. denied, 469 U.S. 851 (1984) In re Fulton, 391 F.3d 1195, 1201,73 USPQ2d 1141, 1146 (Fed. Cir. 2004). >See also MPEP §2123.
In the case of amending the Claimed invention, Applicant is respectfully requested to indicate the portion(s) of the specification which dictate(s) the structure relied on for proper interpretation and also to verify and ascertain the metes and bounds of the claimed invention.
The prior art made of record, listed on form PTO-892, and not relied upon, if any, is considered pertinent to applicant's disclosure
Authorization for Internet Communications
The examiner encourages Applicant to submit an authorization to communicate with the examiner via the Internet by making the following statement (from MPEP 502.03):
“Recognizing that Internet communications are not secure, I hereby authorize the USPTO to communicate with the undersigned and practitioners in accordance with 37 CFR 1.33 and 37 CFR 1.34 concerning any subject matter of this application by video conferencing, instant messaging, or electronic mail. I understand that a copy of these communications will be made of record in the application file.”
Please note that the above statement can only be submitted via Central Fax (not Examiner's Fax), Regular postal mail, or EFS Web using PTO/SB/439.
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Srirama Channavajjala whose telephone number is 571-272-4108. The examiner can normally be reached on Monday-Friday from 8:00 AM to 5:30 PM Eastern Time.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Gorney, Boris, can be reached on (571) 270- 5626. The fax phone numbers for the organization where the application or proceeding is assigned is 571-273-8300 Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free)
/Srirama Channavajjala/Primary Examiner, Art Unit 2154