DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim(s) 1 – 21 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception (i.e. an abstract idea) without significantly more.
In step 1, of the 101-analysis set forth in the MPEP 2106, the examiner has determined that the following limitations recite a process that, under the broadest reasonable interpretation, falls within one or more statutory categories (processes).
In step 2A prong 1, of the 101-analysis set forth in MPEP 2106, the examiner has determined that the following limitations recite a process that, under broadest reasonable interpretation, recites abstract idea but for the recitation of generic computer components:
Regarding claim 1,
defining a plurality of layer groups …
(i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mental process: It involves organizing layers into groups based on selected criteria. See (MPEP 2106.04)).
grouping the layer groups into one or more tile groups, …
(i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mental process: It involves evaluating existing layer groups and organizing them into larger groups. See (MPEP 2106.04)).
selecting a layer group that precedes, in execution order, a first layer group in the tile group and determining an input pre-fetch ratio for the layer group, …
(i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mental process: It involves selecting information based on execution order and determining a value for evaluation. See (MPEP 2106.04)).
comparing the input pre-fetch ratio to an input pre-fetch force factor, …
(i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mental process: It involves evaluating two values to determine their relationship. See (MPEP 2106.04)).
in response to the input pre-fetch ratio exceeding the pre-fetch force factor, assessing one or more criteria relating to output data of the layer group, …
(i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mental process: It involves evaluating output data against one or more criteria based on a comparison result. See (MPEP 2106.04)).
… in response to determining that the one or more criteria relating to output data of the layer group are satisfied, merging the layer group into the tile group.
(i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mental process: It involves deciding whether a layer group satisfies specified criteria and grouping it accordingly. See (MPEP 2106.04)).
If the claim limitations, under their broadest reasonable interpretation, covers performance of the limitations as a mental process, but for the recitation of generic computer components, then it falls within the mental process. Accordingly, the claim recites an abstract idea.
Step 2A Prong 2 of the 101-analysis, set forth in MPEP 2106, the examiner has determined that the following additional elements do not integrate this judicial exception into a practical application:
A method of mapping a neural network to hardware, …
(i.e.: deemed insufficient to transform the judicial exception to a patentable invention because the claim recites limitation which does not amount to more than a recitation of the words "apply it" (or an equivalent), such as mere instructions to implement an abstract idea on a computer. See MPEP 2106.05(f)).
… each layer group comprising one or more layers of the neural network that are processed in a single pass through the hardware; and
(i.e.: deemed insufficient to transform the judicial exception to a patentable invention because the claim recites limitation which does not amount to more than a recitation of the words "apply it" (or an equivalent), such as mere instructions to implement an abstract idea on a computer. See MPEP 2106.05(f)).
each tile group comprising a set of layers groups that are evaluated when executing the neural network, wherein grouping the layer groups into a tile group comprises:
(i.e.: deemed insufficient to transform the judicial exception to a patentable invention because the claim recites limitation which does not amount to more than a recitation of the words "apply it" (or an equivalent), such as mere instructions to implement an abstract idea on a computer. See MPEP 2106.05(f)).
the input pre-fetch ratio corresponding to a number of times that input data to the layer group is read from memory;
(i.e.: deemed insufficient to transform the judicial exception to a patentable invention because the claim recites limitation simply links the judicial exception to a field of use and/or technology environment, see MPEP 2106.05(h)).
wherein the input pre-fetch force factor defines a threshold for pre-fetching input data; and
(i.e.: deemed insufficient to transform the judicial exception to a patentable invention because the claim recites limitation simply links the judicial exception to a field of use and/or technology environment, see MPEP 2106.05(h)).
… the criteria including a size of a pre-fetch buffer configured to store the input data …
(i.e.: deemed insufficient to transform the judicial exception to a patentable invention because the claim recites limitation simply links the judicial exception to a field of use and/or technology environment, see MPEP 2106.05(h)).
In Step 2B of the 101-analysis set forth in the 2019 PEG, the examiner has determined that the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception:
Regarding limitation (I, II and III), recite mere application of the abstract idea or mere instructions to implement an abstract idea on a computer are deemed insufficient to transform the judicial exception to a patentable invention because the limitations generally apply the use of a generic computer and/or process with the judicial exception, see MPEP 2106.05(f).
Regarding limitation (IV, V and VI), additional elements are deemed insufficient to transform the judicial exception to a patentable invention to a patentable invention because they generally link the judicial exception to the technology environment, see MPEP 2106.05(h).
As analyzed above, the additional elements, analyzed above, do not integrate the noted judicial exception into a practical application because they do not impose any meaningful limits on practicing the abstract idea. Therefore, the claim is directed to an abstract idea.
Regarding claim 2, dependent upon claim 1, and fail to resolve the deficiencies identified above by integrating the judicial exception into a practical application, or introducing significantly more than the judicial exception. The claim recites:
wherein in response to determining that the one or more criteria relating to output data of the layer group are not satisfied, the layer group is not merged into the tile group.
(i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mental process: It involves deciding not to group information when specified criteria are unmet. See (MPEP 2106.04)).
Claim 13, recite similar subject matter as claim 2, so is rejected under the same rationale.
Regarding claim 3, dependent upon claim 1, and fail to resolve the deficiencies identified above by integrating the judicial exception into a practical application, or introducing significantly more than the judicial exception. The claim recites:
in response to the input pre-fetch ratio exceeding the pre-fetch force factor …
(i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mental process: It involves determining whether a comparison satisfies a threshold condition. See (MPEP 2106.04)).
pre-allocating space in on-chip memory for the input data.
Deemed insufficient to transform the judicial exception to a patentable invention because the limitation is directed to mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and are considered to adding the words “apply it” (or an equivalent) with the judicial exception, See MPEP 2106.05(f).
Limitations directed to using the computer as a tool for implementing an abstract idea cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B.
Claim 14, recite similar subject matter as claim 3, so is rejected under the same rationale.
Regarding claim 4, dependent upon claim 3, and fail to resolve the deficiencies identified above by integrating the judicial exception into a practical application, or introducing significantly more than the judicial exception. The claim recites:
wherein in response to the input pre-fetch ratio exceeding the pre-fetch force factor and further in response to determining that the one or more criteria relating to output data of the layer group are not satisfied …
(i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mental process: It involves evaluating multiple conditions and making a decision based on those conditions. See (MPEP 2106.04)).
… releasing the pre-allocation.
Deemed insufficient to transform the judicial exception to a patentable invention because the limitation is directed to mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and are considered to adding the words “apply it” (or an equivalent) with the judicial exception, See MPEP 2106.05(f).
Limitations directed to using the computer as a tool for implementing an abstract idea cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B.
Claim 15, recite similar subject matter as claim 4, so is rejected under the same rationale.
Regarding claim 5, dependent upon claim 1, and fail to resolve the deficiencies identified above by integrating the judicial exception into a practical application, or introducing significantly more than the judicial exception. The claim recites:
in response to the input pre-fetch ratio equalling the pre-fetch force factor, assessing one or more criteria relating to output data of the layer group, the criteria including a size of a pre-fetch buffer configured to store the input data and in response to determining that the one or more criteria relating to output data of the layer group are satisfied, merging the layer group into the tile group.
(i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mental process: It involves comparing values, evaluating one or more criteria and deciding whether to merge a layer group. See (MPEP 2106.04)).
Claim 16, recite similar subject matter as claim 5, so is rejected under the same rationale.
Regarding claim 6, dependent upon claim 5, and fail to resolve the deficiencies identified above by integrating the judicial exception into a practical application, or introducing significantly more than the judicial exception. The claim recites:
the method further comprising: in response to the input pre-fetch ratio equalling the pre-fetch force factor, pre-allocating space in on-chip memory for the input data.
Deemed insufficient to transform the judicial exception to a patentable invention because the limitation is directed to mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and are considered to adding the words “apply it” (or an equivalent) with the judicial exception, See MPEP 2106.05(f).
Limitations directed to using the computer as a tool for implementing an abstract idea cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B.
Claim 17, recite similar subject matter as claim 6, so is rejected under the same rationale.
Regarding claim 7, dependent upon claim 6, and fail to resolve the deficiencies identified above by integrating the judicial exception into a practical application, or introducing significantly more than the judicial exception. The claim recites:
wherein in response to the input pre-fetch ratio equalling the pre-fetch force factor and further in response to determining that the one or more criteria relating to output data of the layer group are not satisfied, releasing the pre-allocation.
Deemed insufficient to transform the judicial exception to a patentable invention because the limitation is directed to mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and are considered to adding the words “apply it” (or an equivalent) with the judicial exception, See MPEP 2106.05(f).
Limitations directed to using the computer as a tool for implementing an abstract idea cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B.
Claim 18, recite similar subject matter as claim 7, so is rejected under the same rationale.
Regarding claim 8, dependent upon claim 1, and fail to resolve the deficiencies identified above by integrating the judicial exception into a practical application, or introducing significantly more than the judicial exception. The claim recites:
wherein grouping the layer groups into a tile group further comprises, when assessing one or more criteria relating to output data of the layer group, disregarding any pre-allocation of space in on-chip memory for the input data of the first layer group in the tile group.
The recitation in the additional limitation simply links the judicial exception to a field of use and/or technology environment, see MPEP 2106.05(h).
Limitations directed to field of use cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B.
Claim 19, recite similar subject matter as claim 8, so is rejected under the same rationale.
Regarding claim 9, dependent upon claim 1, and fail to resolve the deficiencies identified above by integrating the judicial exception into a practical application, or introducing significantly more than the judicial exception. The claim recites:
wherein assessing one or more criteria relating to output data of the layer group, the criteria including a size of a pre-fetch buffer configured to store the input data, …
The recitation in the additional limitation simply links the judicial exception to a field of use and/or technology environment, see MPEP 2106.05(h).
Limitations directed to field of use cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B.
comprises: splitting the input data into a plurality of passes; and assessing one or more criteria relating to output data of the layer group, the criteria including a size of a pre-fetch buffer configured to store a pass of the input data.
(i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mental process: It involves organizing data into portions and evaluating those portions against specified criteria. See (MPEP 2106.04)).
Claim 20, recite similar subject matter as claim 9, so is rejected under the same rationale.
Regarding claim 10, dependent upon claim 1, and fail to resolve the deficiencies identified above by integrating the judicial exception into a practical application, or introducing significantly more than the judicial exception. The claim recites:
further comprising: storing the mapped neural network in a memory for execution on the hardware.
Deemed insufficient to transform the judicial exception to a patentable invention because the limitation is directed to mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and are considered to adding the words “apply it” (or an equivalent) with the judicial exception, See MPEP 2106.05(f).
Limitations directed to using the computer as a tool for implementing an abstract idea cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B.
Claim 21, recite similar subject matter as claim 10, so is rejected under the same rationale.
Regarding claim 11, dependent upon claim 1, and fail to resolve the deficiencies identified above by integrating the judicial exception into a practical application, or introducing significantly more than the judicial exception. The claim recites:
further comprising: executing the neural network on the hardware.
Deemed insufficient to transform the judicial exception to a patentable invention because the limitation is directed to mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and are considered to adding the words “apply it” (or an equivalent) with the judicial exception, See MPEP 2106.05(f).
Limitations directed to using the computer as a tool for implementing an abstract idea cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B.
Regarding claim 12,
The rest of the limitations are analogues to claim 1, so are rejected under similar rationale.
A computing device comprising: a processor; and a memory arranged to store computer executable instructions that, when executed by the processor, cause the processor to:
Deemed insufficient to transform the judicial exception to a patentable invention because the limitation is directed to mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and are considered to adding the words “apply it” (or an equivalent) with the judicial exception, See MPEP 2106.05(f).
Limitations directed to using the computer as a tool for implementing an abstract idea cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 3 – 12 and 14 – 21 are rejected under 35 U.S.C. 103 as being unpatentable over Desappan et al., Pub. No.: US11748599B2 in view of Tirunagari, Pub. No.: US20120041914A1, Zhu et al., Pub. No.: US9307045B2 and LUO et al., Pub. No.: US11119915B2.
Regarding claim 1, Desappan teaches: A method of: defining a plurality of layer groups, each layer group comprising one or more layers of the neural network that are processed in a single pass through the hardware; and
(Desappan, (col. 5 line [53 – 62]), “In accordance with aspects of the present disclosure, super tiles may be grouped across a set of layers [defining a plurality of layer groups] and these layers may be used to process super tiles. In certain cases, layer groups may be used to help increase processing and memory bandwidth efficiency. For example, a particular CNN may include ten layers where only the third, fourth, and fifth layers are associated with tensors which may not fit into L3. The third, fourth, and fifth layers may be grouped together into a layer group and processed together using super tiles [each layer group comprising one or more layers of the neural network that are processed in a single pass through the hardware], while other layers may be processed one layer at a time.”)
grouping the layer groups into one or more tile groups,
(Desappan, (col. 5 line [53 – 57]), “In accordance with aspects of the present disclosure, super tiles may be grouped across a set of layers [defining a plurality of layer groups] and these layers may be used to process super tiles. In certain cases, layer groups may be used to help increase processing and memory bandwidth efficiency.
each tile group comprising a set of layers groups that are evaluated when executing the neural network, wherein grouping the layer groups into a tile group comprises:
(Desappan, (col. 3 line [43 – 52]), “In certain cases, a tensor may be split into tiles for processing, as shown in tensor 200 of FIG. 2 , where the tiles may be sized based, for example, on the pipeline design of the processor. For example, a tile may include one or more nodes based on a number of parallel pipelines available on a processor [each tile group comprising a set of layers groups that are evaluated when executing the neural network]. Of note, going forward, tensors are shown as two-dimensional structures for the sake of clarity. In common implementations, all tiles of a given tensor are processed by a particular layer before processing starts on the next tensor and layer.”)
selecting a layer group that precedes, in execution order, a first layer group in the tile group and
(Desappan, (col. 7 line [46 – 54]), “In this example, as each prior layer requires more nodes to be calculated than the next, the size, and hence memory space required to calculate the nodes of the first tensor 502A for the first pass, would be a limiting factor to the size of the overall super tile. That is, the size of the super tile may be selected to allow the calculations needed for the first tensor 502A [selecting a layer group that precedes, in execution order, a first layer group in the tile group] in the first pass to fit into a memory, such as the L3 cache.”)
Desappan does not teach:
mapping a neural network to hardware, the method comprising
determining an input pre-fetch ratio for the layer group, the input pre-fetch ratio corresponding to a number of times that input data to the layer group is read from memory;
comparing the input pre-fetch ratio to an input pre-fetch force factor, wherein the input pre-fetch force factor defines a threshold for pre-fetching input data; and
in response to the input pre-fetch ratio exceeding the pre-fetch force factor, assessing one or more criteria relating to output data of the layer group,
the criteria including a size of a pre-fetch buffer configured to store the input data and in response to determining that the one or more criteria relating to output data of the layer group are satisfied,
merging the layer group into the tile group
Tirunagari teaches:
mapping a neural network to hardware, the method comprising:
(Tirunagari, “ [0007] … In some embodiments, replacing an initial caching algorithm with a caching algorithm selected by the neural network may include changing the value of one or more parameters of the current caching algorithm, the hardware configuration (e.g., changing the size of the cache), or changing the value of a parameter of the operating system [mapping a neural network to hardware] ...”)
determining an input pre-fetch ratio for the layer group, the input pre-fetch ratio corresponding to a number of times that input data to the layer group is read from memory;
(Tirunagari, “[0031] As illustrated at 120 in FIG. 1, the method may include a neural network monitoring one or more performance related parameters during execution of the application. For example, an input layer of the neural network may gather and/or receive inputs indicating data accesses made by the application [determining an input pre-fetch ratio for the layer group, the input pre-fetch ratio corresponding to a number of times that input data to the layer group is read from memory], a rate of accesses made by the application, a cache hit rate, a data throughput rate, an access response time (e.g., an average, current, or maximum memory access time experienced in response to one or more memory access requests), various hardware parameter values, various operation system parameters, and/or other information that may reflect and/or affect the performance of the application in the system...”)
comparing the input pre-fetch ratio to an input pre-fetch force factor, wherein the input pre-fetch force factor defines a threshold for pre-fetching input data; and
(Tirunagari, “[0031] … As illustrated at 130, the method may include the neural network selecting a second caching algorithm for the application, dependent on the results of the monitoring. For example, if any of the monitored parameters indicates an unacceptable level of performance, a deterioration of performance, or a change in the hardware or software parameters of the system that may affect the performance of the application [comparing the input pre-fetch ratio to an input pre-fetch force factor, wherein the input pre-fetch force factor defines a threshold for pre-fetching input data], the hidden layer of the neural network may be configured to determine a more suitable caching algorithm for the application, based on the current state and/or the value of one or more of the monitored parameters.”)
in response to the input pre-fetch ratio exceeding the pre-fetch force factor, assessing one or more criteria relating to output data of the layer group,
(Tirunagari, “[0024] In some embodiments, detection of a particular input value may trigger the neural network to select (or to determine whether to select) a replacement caching algorithm for temporarily storing data accessed by a given executing application. For example, if the value of one of the performance related parameters meets or exceeds a pre-determined threshold value (e.g., a cache hit rate below a given acceptable level) [in response to the input pre-fetch ratio exceeding the pre-fetch force factor, assessing one or more criteria relating to output data of the layer group], the neural network may be configured to analyzed the current set of inputs and to output an indication of an appropriate caching algorithm for the current application and execution contexts.”)
Tirunagari and Desappan are related to the same field of endeavor (i.e.: memory optimization). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to combine the teaching of Tirunagari with teachings of Desappan to add dynamic selection and updating a caching algorithm using a neural network based on runtime performance metric to improve memory access efficiency and cache utilization (Tirunagari, Abstract).
Desappan in view of Tirunagari do not teach:
the criteria including a size of a pre-fetch buffer configured to store the input data and in response to determining that the one or more criteria relating to output data of the layer group are satisfied,
merging the layer group into the tile group
Zhu teaches:
the criteria including a size of a pre-fetch buffer configured to store the input data and in response to determining that the one or more criteria relating to output data of the layer group are satisfied,
(Zhu, (col. 15 line [9 – 19]), “In another embodiment, the module 188 accesses a map data specific memory 1308, which is a portion of the memory 1306 allocated specifically to storing map data, in the illustrated example. In either case, the block 1304 determines the amount of available map data memory at the time of inspection, either the size of the memory 1306 or the allocated or available portion of the memory 1308 [the criteria including a size of a pre-fetch buffer configured to store the input data and in response to determining that the one or more criteria relating to output data of the layer group are satisfied]. At a block 1310, the routine or process 1300 accesses the current tile budget settings for the various tile budget criteria and analyzes these in light of the determined available amount of memory.”)
Zhu, Desappan and Tirunagari are related to the same field of endeavor (i.e.: memory optimization). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to combine the teaching of Zhu with teachings of Desappan and Tirunagari to add identifying and per-fetching a selected subset of data based on predicted future use to improve memory efficiency and reduce data access latency by proactively storing only an optimized amount of relevant data. (Zhu, Abstract).
Desappan in view of Tirunagari and Zhu do not teach:
merging the layer group into the tile group
LUO teaches:
merging the layer group into the tile group
(LUO, (col. 19 [13 – 19]), “At block 1008, the processor feeds the one or more tiles written to on-chip memory in block 1006 into the fused layer and the processor performs a merged convolution/bnorm operation on the one or more tiles as described in the present disclosure [merging the layer group into the tile group]. At block 1010, the processor writes the result of the convolution/bnorm operation to on-chip memory as a temporary feature map.”)
LUO, Desappan, Tirunagari and Zhu are related to the same field of endeavor (i.e.: memory optimization). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to combine the teaching of LUO with teachings of Desappan, Tirunagari and Zhu to add mapping neural network feature maps to different levels of a memory hierarchy based on available memory space and access, speed and removing feature maps from memory when they expire to reclaim storage resources. (LUO, Abstract).
Claim 12, recites limitations analogous to claim 1, so is rejected under the same rationale.
Regarding claim 3, Desappan in view of Tirunagari, Zhu and LUO teach the method of claim 1.
Zhu further teaches: the method further comprising: in response to the input pre-fetch ratio exceeding the pre-fetch force factor, pre-allocating space in on-chip memory for the input data.
((Zhu, col.2 line [50 – 59]), “a tile budget for the client device; determining if the tile budget has been met by the received pre-fetch map data tiles, where, if the tile budget has been met, the client device stops receiving additional pre-fetch map data tiles from the map database, and if the tile budget has not been met [in response to the input pre-fetch ratio exceeding the pre-fetch force factor], the client device, continues receiving additional pre-fetch map data tiles from the map database until the tile budget is met [pre-allocating space in on-chip memory for the input data.] or until all pre-fetch map data tiles corresponding to the one or more map points of interest have been received at the client device; ”)
It would have been obvious to one of ordinary skill in the art before the effective filling date of the present application to combine the teachings of Zhu with teachings of Desappan, Tirunagari and LUO for the same reasons disclosed for claim 1.
Claim 14, recites limitations analogous to claim 3, so is rejected under the same rationale.
Regarding claim 4, Desappan in view of Tirunagari, Zhu and LUO teach the method of claim 3.
Zhu further teaches: wherein in response to the input pre-fetch ratio exceeding the pre-fetch force factor and further in response to determining that the one or more criteria relating to output data of the layer group are not satisfied,
((Zhu, col.2 line [50 – 59]), “a tile budget for the client device; determining if the tile budget has been met by the received pre-fetch map data tiles, where, if the tile budget has been met [wherein in response to the input pre-fetch ratio exceeding the pre-fetch force factor and further in response to determining that the one or more criteria relating to output data of the layer group are not satisfied], the client device stops receiving additional pre-fetch map data tiles from the map database, and if the tile budget has not been met, the client device, continues receiving additional pre-fetch map data tiles from the map database until the tile budget is met or until all pre-fetch map data tiles corresponding to the one or more map points of interest have been received at the client device; ”)
LUO further teaches: releasing the pre-allocation.
(LUO, (col. 1 [46 – 50]), “The at least one processor is configured to map a first feature map of a plurality of feature maps to a memory in the memory hierarchy having available memory space and providing quickest access to the first feature map. The at least one processor is also configured to, when the first feature map expires, remove the first feature map from the memory, and unallocate the memory used to store the first feature map.”)
It would have been obvious to one of ordinary skill in the art before the effective filling date of the present application to combine the teachings of Zhu with teachings of Desappan, Tirunagari and LUO for the same reasons disclosed for claim 1.
Claim 15, recites limitations analogous to claim 4, so is rejected under the same rationale.
Regarding claim 5, Desappan in view of Tirunagari, Zhu and LUO teach the method of claim 1.
Tirunagari further teaches: wherein grouping the layer groups into a tile group further comprises: in response to the input pre-fetch ratio equalling the pre-fetch force factor,
(Tirunagari, “[0024] In some embodiments, detection of a particular input value may trigger the neural network to select (or to determine whether to select) a replacement caching algorithm for temporarily storing data accessed by a given executing application. For example, if the value of one of the performance related parameters meets [in response to the input pre-fetch ratio equalling the pre-fetch force factor] or exceeds a pre-determined threshold value (e.g., a cache hit rate below a given acceptable level), the neural network may be configured to analyzed the current set of inputs and to output an indication of an appropriate caching algorithm for the current application and execution contexts.”)
Zhu further teaches: assessing one or more criteria relating to output data of the layer group, the criteria including a size of a pre-fetch buffer configured to store the input data and in response to determining that the one or more criteria relating to output data of the layer group are satisfied,
(Zhu, (col. 15 line [9 – 19]), “In another embodiment, the module 188 accesses a map data specific memory 1308, which is a portion of the memory 1306 allocated specifically to storing map data, in the illustrated example. In either case, the block 1304 determines the amount of available map data memory at the time of inspection, either the size of the memory 1306 or the allocated or available portion of the memory 1308 [assessing one or more criteria relating to output data of the layer group, the criteria including a size of a pre-fetch buffer configured to store the input data and in response to determining that the one or more criteria relating to output data of the layer group are satisfied]. At a block 1310, the routine or process 1300 accesses the current tile budget settings for the various tile budget criteria and analyzes these in light of the determined available amount of memory.”)
It would have been obvious to one of ordinary skill in the art before the effective filling date of the present application to combine the teachings of Tirunagari with teachings of Desappan, Zhu and LUO for the same reasons disclosed for claim 1.
LUO further teaches: merging the layer group into the tile group.
(LUO, (col. 19 [13 – 19]), “At block 1008, the processor feeds the one or more tiles written to on-chip memory in block 1006 into the fused layer and the processor performs a merged convolution/bnorm operation on the one or more tiles as described in the present disclosure [merging the layer group into the tile group]. At block 1010, the processor writes the result of the convolution/bnorm operation to on-chip memory as a temporary feature map.”)
It would have been obvious to one of ordinary skill in the art before the effective filling date of the present application to combine the teachings of LUO with teachings of Desappan, Tirunagari and Zhu for the same reasons disclosed for claim 1.
Claim 16, recites limitations analogous to claim 5, so is rejected under the same rationale.
Regarding claim 6, Desappan in view of Tirunagari, Zhu and LUO teach the method of claim 5.
Desappan further teaches: the method further comprising: in response to the input pre-fetch ratio equalling the pre-fetch force factor, pre-allocating space in on-chip memory for the input data.
(Desappan, (col. 5 line [7 – 17]), “In certain cases, output from the first layer 332 may be dynamically written over corresponding parts of the first portion 328 in the on-chip memory 322 [pre-allocating space in on-chip memory for the input data] as the output is generated. Once generated, the second portion 336 is processed in a second layer 340 in conjunction with second ML network information 342 to produce a second layer output 344, which is written back into the on-chip memory 322, overwriting portions of the on-chip memory 322 which were storing the second portion 336 to obtain a third portion 346 of a third tensor.”)
Claim 17, recites limitations analogous to claim 6, so is rejected under the same rationale.
Regarding claim 7, Desappan in view of Tirunagari, Zhu and LUO teach the method of claim 6.
LUO further teaches: wherein in response to the input pre-fetch ratio equalling the pre-fetch force factor and further in response to determining that the one or more criteria relating to output data of the layer group are not satisfied, releasing the pre-allocation.
(LUO, (col. 1 [46 – 50]), “The at least one processor is configured to map a first feature map of plurality of feature maps to a memory in the memory hierarchy having available memory space and providing quickest access to the first feature map. The at least one processor is also configured to, when the first feature map expires [wherein in response to the input pre-fetch ratio equalling the pre-fetch force factor and further in response to determining that the one or more criteria relating to output data of the layer group are not satisfied], remove the first feature map from the memory, and unallocate the memory used to store the first feature map [releasing the pre-allocation].”)
It would have been obvious to one of ordinary skill in the art before the effective filling date of the present application to combine the teachings of LUO with teachings of Desappan, Tirunagari and Zhu for the same reasons disclosed for claim 1.
Claim 18, recites limitations analogous to claim 7, so is rejected under the same rationale.
Regarding claim 8, Desappan in view of Tirunagari, Zhu and LUO teach the method of claim 1.
Desappan further teaches: wherein grouping the layer groups into a tile group further comprises, when assessing one or more criteria relating to output data of the layer group, disregarding any pre-allocation of space in on-chip memory for the input data of the first layer group in the tile group.
(Desappan, (col. 5 line [7 – 17]), “In certain cases, output from the first layer 332 may be dynamically written over corresponding parts of the first portion 328 in the on-chip memory 322 [disregarding any pre-allocation of space in on-chip memory for the input data of the first layer group in the tile group] as the output is generated. Once generated, the second portion 336 is processed in a second layer 340 in conjunction with second ML network information 342 to produce a second layer output 344, which is written back into the on-chip memory 322, overwriting portions of the on-chip memory 322 which were storing the second portion 336 to obtain a third portion 346 of a third tensor.”)
Claim 19, recites limitations analogous to claim 8, so is rejected under the same rationale.
Regarding claim 9, Desappan in view of Tirunagari, Zhu and LUO teach the method of claim 1.
Desappan further teaches: wherein assessing one or more criteria relating to output data of the layer group, the criteria including a size of a pre-fetch buffer configured to store the input data, comprises:
splitting the input data into a plurality of passes; and
(Desappan, (col. 3 line [43 – 52]), “In certain cases, a tensor may be split into tiles for processing [splitting the input data into a plurality of passes], as shown in tensor 200 of FIG. 2 , where the tiles may be sized based, for example, on the pipeline design of the processor. For example, a tile may include one or more nodes based on a number of parallel pipelines available on a processor [each tile group comprising a set of layers groups that are evaluated when executing the neural network]. Of note, going forward, tensors are shown as two-dimensional structures for the sake of clarity. In common implementations, all tiles of a given tensor are processed by a particular layer before processing starts on the next tensor and layer.”)
Zhu further teaches: assessing one or more criteria relating to output data of the layer group, the criteria including a size of a pre-fetch buffer configured to store a pass of the input data.
(Zhu, (col. 15 line [37 – 44]), “The routine or process 1400 differs with respect to block 1406, where a map data tile memory allocation size is accessed, and block 1408, where it is determined the number of map tile memory slots available on the client device. At a block 1406, the routine or process 1400 accesses a current tile memory slot allocation size. In the illustrated embodiment, a memory 1410 on the client device has a dedicated map buffer memory 1412 [assessing one or more criteria relating to output data of the layer group, the criteria including a size of a pre-fetch buffer configured to store a pass of the input data],”)
It would have been obvious to one of ordinary skill in the art before the effective filling date of the present application to combine the teachings of Zhu with teachings of Desappan, Tirunagari and LUO for the same reasons disclosed for claim 1.
Claim 20, recites limitations analogous to claim 9, so is rejected under the same rationale.
Regarding claim 10, Desappan in view of Tirunagari, Zhu and LUO teach the method of claim 1.
Tirunagari further teaches: further comprising: storing the mapped neural network in a memory for execution on the hardware.
(Tirunagari, “[0036] … For example, in some embodiments, the input layer of a neural network may be implemented, at least in part, using circuitry configured to collect performance related data and provide it to the hidden layer. In other embodiments, the input layer may include one or more software modules configured to gather (e.g., read or otherwise determine) values stored or generated by a hardware component [storing the mapped neural network in a memory for execution on the hardware] (e.g., a performance counter or performance related status register, a snoop circuit that captures bus traffic, an operating system parameter register, or a hardware configuration status indicator), and to provide those values to the hidden layer of the neural network…”)
It would have been obvious to one of ordinary skill in the art before the effective filling date of the present application to combine the teachings of Tirunagari with teachings of Desappan, Zhu and LUO for the same reasons disclosed for claim 1.
Claim 21, recites limitations analogous to claim 10, so is rejected under the same rationale.
Regarding claim 11, Desappan in view of Tirunagari, Zhu and LUO teach the method of claim 1.
Tirunagari further teaches: further comprising: executing the neural network on the hardware.
(Tirunagari, “[0019] … For example, in some embodiments, neural networks may be used to model the relationships between performance related inputs and outputs in the system in order to iteratively and dynamically select a suitable caching algorithm for a given application, resource request, and/or execution context (e.g., hardware and/or operating system configuration) during runtime [executing the neural network on the hardware]. In some embodiments, the methods described herein for applying neural networks to the design of caches may also be highly scalable.”)
It would have been obvious to one of ordinary skill in the art before the effective filling date of the present application to combine the teachings of Tirunagari with teachings of Desappan, Zhu and LUO for the same reasons disclosed for claim 1.
Claim(s) 2 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Desappan in view of Tirunagari, Zhu, LUO and in further view of Sharma et al., Pub. No.: US20200226473A1.
Regarding claim 2, Desappan in view of Tirunagari, Zhu and LUO teach the method of claim 1.
Desappan in view of Tirunagari, Zhu and LUO do not teach:
wherein in response to determining that the one or more criteria relating to output data of the layer group are not satisfied, the layer group is not merged into the tile group.
Sharma teaches:
wherein in response to determining that the one or more criteria relating to output data of the layer group are not satisfied, the layer group is not merged into the tile group.
(Sharma, “[0119] FIG. 9 shows how the operations in a fully-connected layer are split into tiles [the layer group is not merged into the tile group] for each level of memory hierarchy including global memory tile (e.g., URAM tile) at operation 910 and cluster memory tile (e.g., BRAM tile) at operation 930. Using a larger tile size for each level of hierarchy increases the data reuse at that level of hierarchy at operation 940. The tile sizes are constrained by the capacity of memory at that level of hierarchy [wherein in response to determining that the one or more criteria relating to output data of the layer group are not satisfied] (i.e.: a tile cannot be enlarged beyond available memory capacity).”)
Sharma, Desappan, Tirunagari, Zhu and LUO are related to the same field of endeavor (i.e.: memory optimization). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to combine the teaching of Sharma with teachings of Desappan, Tirunagari, Zhu and LUO to selectively merging preceding layer groups based on memory access frequency and on chip memory availability for pre-fetched input data. (Sharma, Abstract).
Claim 13, recites limitations analogous to claim 2, so is rejected under the same rationale.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Hu et al., Pub. No.: US20230120631A1.
Hu teaches obtaining a codeword corresponding to a first weight matrix of a neural network model from a memory; determining, based on the codeword, that a weight matrix of the neural network model is the first weight matrix, and training the first weight matrix by using training data; updating the codeword when a preset stop condition is not met, to obtain an updated codeword; storing the updated codeword in the memory;
Shao et al., Pub. No.: US11270197B2.
Shao teaches a distributed, tile-based architecture includes multiple chips, each with a central processing element, a global memory buffer, and a plurality of additional processing elements. Each additional processing element includes a weight buffer, an activation buffer, and vector multiply-accumulate units to combine, in parallel, the weight values and the activation values using stationary data flows.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MATIYAS T MARU whose telephone number is (571)270-0902 or via email: matiyas.maru@uspto.gov. The examiner can normally be reached Monday 8:00am - Friday 4:00pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a
USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor,
Michelle Bechtold can be reached on (571)431-0762. The fax phone number for the organization were this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from
Patent Center. Unpublished application information in Patent Center is available to registered users.
To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit
https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and
https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional
questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like
assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA)
or 571-272-1000.
/M.T.M./ Examiner, Art Unit 2148
/MICHELLE T BECHTOLD/Supervisory Patent Examiner, Art Unit 2148