Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Specification
The disclosure is objected to because of the following informalities: [0045] “defined in commonly-used dictionaries” should read as “defined in commonly used dictionaries”. Appropriate correction is required.
Claim Objections
Claim 4 objected to because of the following informalities: “either one of a compute-bound layer and a memory-bound layer” should read as “either one of a compute-bounded layer or a memory-bound layer”. Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112:
The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention.
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 3 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 3 recites the limitation “for each of the plurality of layers, determining a weight between a computation amount and a memory access amount in a layer”. It is not clear how this “weight” is determined and what form it is, in both claim and specification, so it renders the claim indefinite. For examination purposes examiner has interpreted “a weight between a computation amount and a memory access amount in a layer” in claim 3 as way to represent the relationship between the computational cost and memory access cost, including but not limited to scaling, correlation, adding a coefficient and regression.
Claim 9 recites the limitation "a component configuring the device". There is insufficient antecedent basis for this limitation in the claim. Claim 7 also states “receiving … information on a component configuring the device”. Therefore, it is unclear which component is being referred to and the scope of the claim is unclear. For examination purposes examiner has interpreted “a component configuring the device” in claim 9 as “the ”.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1:
Claim 1-10 are method (process) claims. Claim 11 is a manufacture claim. Claim 12-20 are machine claims. Therefore, claims 1-20 are directed to either a process, machine, manufacture or composition of matter.
Regarding claim 1:
2A Prong 1:
building a prediction model based on the benchmark execution result and the hardware information (Mental process of evaluation - This step of “building a prediction model” is practically performable in the human mind and is understood to be a recitation of a mental process with the aid of pen and paper.)
extracting layer information respectively corresponding to a plurality of layers configuring the neural network model (Mental process of evaluation - This step of “extracting layer information” is practically performable in the human mind and is understood to be a recitation of a mental process.)
and predicting either one or both of operation performance information and energy efficiency information respectively corresponding to the plurality of layers by inputting the analysis requirement information and the layer information to the prediction model (Mental process of evaluation - This step of “predicting either one or both of operation performance information and energy efficiency information” is practically performable in the human mind and is understood to be a recitation of a mental process with the aid of pen and paper.)
2A Prong 2: This judicial exception is not integrated into a practical application.
Additional elements:
A processor-implemented method comprising: (“A processor-implemented method” is understood as mere instructions to apply the exception using generic computer components as discussed in MPEP 2106.05(f));
obtaining a benchmark execution result (“Obtaining a benchmark execution result” is understood as adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g).)
receiving input data comprising a neural network model subject to prediction and analysis requirement information (“Receiving input data” is understood as adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g).)
receiving information on hardware of a device in which the neural network model is run (“Receiving information on hardware” is understood as adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g)).
The additional elements as disclosed above alone or in combination do not integrate the judicial exception into practical application as they are adding insignificant extra-solution activity to the judicial exception that are implemented to perform the disclosed abstract idea above.
2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
A processor-implemented method comprising: (“A processor-implemented method” is understood as mere instructions to apply the exception using generic computer components as discussed in MPEP 2106.05(f));
An apparatus comprising: one or more processors configured to: (“An apparatus comprising: one or more processors” is understood as mere instructions to apply the exception using generic computer components as discussed in MPEP 2106.05(f));
obtaining a benchmark execution result (well-understood, routine and convention activity (According to MPEP 2106.05(d)(I), the claim “obtaining a benchmark execution result” is merely a court identified computer function of “Receiving or transmitting data over a network”. Thereby, a conclusion that the claimed step is well-understood, routine, conventional activity is supported under Berkheimer).
receiving input data comprising a neural network model subject to prediction and analysis requirement information (According to MPEP 2106.05(d)(I), the claim “receiving input data” is merely a court identified computer function of “Receiving or transmitting data over a network”. Thereby, a conclusion that the claimed step is well-understood, routine, conventional activity is supported under Berkheimer).
receiving information on hardware of a device in which the neural network model is run (According to MPEP 2106.05(d)(I), the claim “receiving information on hardware of a device” is merely court identified computer function of “Receiving or transmitting data over a network”. Thereby, a conclusion that the claimed step is well-understood, routine, conventional activity is supported under Berkheimer).
The additional elements as disclosed above in combination of the abstract idea are not sufficient to amount to significantly more than the judicial exception as they are court recognized well-understood, routine, conventional functions claimed in a merely generic manner that are implemented to perform the disclosed abstract idea above.
Regarding claim 2:
2A Prong 1:
wherein the predicting comprises: predicting the operation performance information and the energy efficiency information respectively corresponding to the plurality of layers for each component configuring the device. (Mental process of evaluation - This step of “predicting the operation performance information and energy efficiency information” is practically performable in the human mind and is understood to be a recitation of a mental process with the aid of pen and paper.)
2A Prong 2 & 2B: This claim does not recite any additional elements.
Regarding claim 3:
2A Prong 1:
wherein the predicting comprises:
for each of the plurality of layers, determining a weight between a computation amount and a memory access amount in a layer (Mental process of evaluation - This step of “for each of the plurality of layers, determining a weight” is practically performable in the human mind and is understood to be a recitation of a mental process by using mathematical calculations with the aid of pen and paper.)
and predicting the operation performance information and the energy efficiency information respectively corresponding to the plurality of layers based on the weight (Mental process of evaluation - This step of “predicting the operation performance information and energy efficiency information” is practically performable in the human mind and is understood to be a recitation of a mental process with the aid of pen and paper.)
2A Prong 2 & 2B This claim does not recite any additional elements.
Regarding claim 4:
2A Prong 1:
wherein the predicting comprises:
classifying the plurality of layers into either one of compute-bound layer and a memory-bound layer; (Mental process of observation - This step of “classifying the plurality of layers” is practically performable in the human mind and is understood to be a recitation of a mental process.)
and predicting the operation performance information and the energy efficiency information respectively corresponding to the plurality of layers based on a result of classifying (Mental process of evaluation - This step of “predicting the operation performance information and energy efficiency information” is practically performable in the human mind and is understood to be a recitation of a mental process with the aid of pen and paper.)
2A Prong 2 & 2B: This claim does not recite any additional elements.
Regarding claim 5:
2A Prong 1:
wherein the extracting of the layer information comprises extracting input data information and output data information respectively corresponding to the plurality of layers (Mental process of evaluation - This step of “extracting input data information and output data information” is practically performable in the human mind and is understood to be a recitation of a mental process.)
2A Prong 2 & 2B: This claim does not recite any additional elements.
Regarding claim 6:
2A Prong 1: Incorporates the rejection of claim 1.
2A Prong 2: This judicial exception is not integrated into a practical application.
Additional elements:
wherein the obtaining of the benchmark execution result comprises:
receiving user setting information (“Receiving user setting information” is understood as adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g).)
generating a benchmark based on the user setting information (“Generating a benchmark” is understood as mere instructions to implement an abstract idea (e.g. generating content) on a computer – see MPEP 2106.05(f).)
and generating the benchmark execution result by executing the benchmark. (“Generating a benchmark execution result” is understood as mere instructions to implement an abstract idea (e.g. generating content) on a computer – see MPEP 2106.05(f).)
The additional elements as disclosed above alone or in combination do not integrate the judicial exception into practical application as they are generic computer functions and adding insignificant extra-solution activity to the judicial exception that are implemented to perform the disclosed abstract idea above.
2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
receiving user setting information (According to MPEP 2106.05(d)(I), the claim “receiving user setting information” is merely court identified computer function of “Receiving or transmitting data over a network”. Thereby, a conclusion that the claimed step is well-understood, routine, conventional activity is supported under Berkheimer).
generating a benchmark based on the user setting information (“Generating a benchmark” is understood as mere instructions to implement an abstract idea (e.g. generating content) on a computer – see MPEP 2106.05(f).)
and generating the benchmark execution result by executing the benchmark. (“Generating a benchmark execution result” is understood as mere instructions to implement an abstract idea (e.g. generating content) on a computer – see MPEP 2106.05(f).)
The additional elements as disclosed above in combination of the abstract idea are not sufficient to amount to significantly more than the judicial exception as they are generic computer functions and court recognized well-understood, routine, conventional functions claimed in a merely generic manner that are implemented to perform the disclosed abstract idea above.
Regarding claim 7:
2A Prong 1: Incorporates the rejection of claim 6.
2A Prong 2: This judicial exception is not integrated into a practical application.
Additional elements:
wherein the receiving of the user setting information comprises receiving information on a preset operating frequency, information on a neural network model subject to benchmark, and information on a component configuring the device (“Receiving information” is understood as adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g).)
The additional elements as disclosed above alone or in combination do not integrate the judicial exception into practical application as they are adding insignificant extra-solution activity to the judicial exception that are implemented to perform the disclosed abstract idea above.
2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
comprises receiving information on a preset operating frequency, information on a neural network model subject to benchmark, and information on a component configuring the device (According to MPEP 2106.05(d)(I), the claim “receiving information” is merely court identified computer function of “Receiving or transmitting data over a network”. Thereby, a conclusion that the claimed step is well-understood, routine, conventional activity is supported under Berkheimer).
The additional elements as disclosed above in combination of the abstract idea are not sufficient to amount to significantly more than the judicial exception as they are court recognized well-understood, routine, conventional functions claimed in a merely generic manner that are implemented to perform the disclosed abstract idea above.
Regarding claim 8:
2A Prong 1: Incorporates the rejection of claim 7.
2A Prong 2: This judicial exception is not integrated into a practical application.
Additional elements:
wherein the generating of the benchmark comprises, for each of the plurality of layers configuring the neural network model subject to the benchmark, generating the benchmark based on an input data size and an operating frequency of a layer (“Generating the benchmark” is understood as mere instructions to implement an abstract idea (e.g. generating content) on a computer – see MPEP 2106.05(f).)
The additional elements as disclosed above in combination of the abstract idea are not sufficient to amount to significantly more than the judicial exception as they are generic computer functions that are implemented to perform the disclosed abstract idea above.
2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
wherein the generating of the benchmark comprises, for each of the plurality of layers configuring the neural network model subject to the benchmark, generating the benchmark based on an input data size and an operating frequency of a layer (“Generating the benchmark” is understood as mere instructions to implement an abstract idea (e.g. generating content) on a computer – see MPEP 2106.05(f).)
The additional elements as disclosed above in combination of the abstract idea are not sufficient to amount to significantly more than the judicial exception as they are generic computer functions that are implemented to perform the disclosed abstract idea above.
Regarding claim 9:
2A Prong 1: Incorporates the rejection of claim 7.
2A Prong 2: This judicial exception is not integrated into a practical application.
Additional elements:
wherein the generating of the benchmark comprises generating the benchmark based on a kernel corresponding to a component configuring the device (“Generating the benchmark” is understood as mere instructions to implement an abstract idea (e.g. generating content) on a computer – see MPEP 2106.05(f).)
The additional elements as disclosed above in combination of the abstract idea are not sufficient to amount to significantly more than the judicial exception as they are generic computer functions that are implemented to perform the disclosed abstract idea above.
2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
wherein the generating of the benchmark comprises generating the benchmark based on a kernel corresponding to a component configuring the device (“Generating the benchmark” is understood as mere instructions to implement an abstract idea (e.g. generating content) on a computer – see MPEP 2106.05(f).)
The additional elements as disclosed above in combination of the abstract idea are not sufficient to amount to significantly more than the judicial exception as they are generic computer functions that are implemented to perform the disclosed abstract idea above.
Regarding claim 10:
2A Prong 1:
wherein the generating of the benchmark execution result comprises:
performing consistency determination on the benchmark execution result by comparing a result of the first measurement and a result of the second measurement (Mental process of evaluation - This step of “performing consistency determination” is practically performable in the human mind and is understood to be a recitation of a mental process.)
2A Prong 2: This judicial exception is not integrated into a practical application.
Additional elements:
performing a first measurement on operation performance and energy efficiency corresponding to the benchmark based on an application programming interface (API) (“Performing a first measurement” is understood as mere instructions to implement an abstract idea (e.g. generating content) on a computer – see MPEP 2106.05(f).)
performing a second measurement on operation performance and energy efficiency corresponding to the benchmark using an external device (“Performing a second measurement” is understood as mere instructions to implement an abstract idea (e.g. generating content) on a computer – see MPEP 2106.05(f).)
The additional elements as disclosed above in combination of the abstract idea are not sufficient to amount to significantly more than the judicial exception as they are generic computer functions that are implemented to perform the disclosed abstract idea above.
2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
performing a first measurement on operation performance and energy efficiency corresponding to the benchmark based on an application programming interface (API) (“Performing a first measurement” is understood as mere instructions to implement an abstract idea (e.g. generating content) on a computer – see MPEP 2106.05(f).)
performing a second measurement on operation performance and energy efficiency corresponding to the benchmark using an external device (“Performing a second measurement” is understood as mere instructions to implement an abstract idea (e.g. generating content) on a computer – see MPEP 2106.05(f).)
The additional elements as disclosed above in combination of the abstract idea are not sufficient to amount to significantly more than the judicial exception as they are generic computer functions that are implemented to perform the disclosed abstract idea above.
Regarding claim 11:
2A Prong 1: Incorporates the rejection of claim 1.
2A Prong 2: This judicial exception is not integrated into a practical application.
Additional elements:
A non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors, configure the one or more processors to perform the method of claim 1 (“A non-transitory computer-readable storage medium” is understood as mere instructions to implement an abstract idea on a computer – see MPEP 2106.05(f).)
The additional elements as disclosed above in combination of the abstract idea are not sufficient to amount to significantly more than the judicial exception as they are generic computer functions that are implemented to perform the disclosed abstract idea above.
2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
A non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors, configure the one or more processors to perform the method of claim 1 (“A non-transitory computer-readable storage medium” is understood as mere instructions to implement an abstract idea on a computer – see MPEP 2106.05(f).)
The additional elements as disclosed above in combination of the abstract idea are not sufficient to amount to significantly more than the judicial exception as they are generic computer functions that are implemented to perform the disclosed abstract idea above.
Regarding claim 12:
2A Prong 1: Incorporates the rejection of claim 1.
2A Prong 2: This judicial exception is not integrated into a practical application.
Additional elements:
An apparatus comprising: one or more processors configured to: (“An apparatus comprising: one or more processors” is understood as mere instructions to apply the exception using generic computer components as discussed in MPEP 2106.05(f));
The additional elements as disclosed above in combination of the abstract idea are not sufficient to amount to significantly more than the judicial exception as they are generic computer functions that are implemented to perform the disclosed abstract idea above.
2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
An apparatus comprising: one or more processors configured to: (“An apparatus comprising: one or more processors” is understood as mere instructions to apply the exception using generic computer components as discussed in MPEP 2106.05(f));
The additional elements as disclosed above in combination of the abstract idea are not sufficient to amount to significantly more than the judicial exception as they are generic computer functions that are implemented to perform the disclosed abstract idea above.
Regarding claim 13-20:
2A Prong 1: Incorporates the rejection of claim 2-8, 10 in addition to the rejection of claim 12.
2A Prong 2 & 2B: These claims do not recite any additional elements.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1, 2, 4-8, 10-13, 15-20 are rejected under 35 U.S.C. 103 as being unpatentable over Uppalapati et al (US 20230195590 A1, hereinafter "Uppalapati") in view of Zhu et al. (US 20240168948 A1, hereinafter "Zhu"), based on the English translation (see attachment).
Regarding claim 1: Uppalapati teaches:
A processor-implemented method comprising: obtaining a benchmark execution result; (Paragraph [0065-0067] describes that the system can generate a set of runtime performance metrics during execution of the target neural network by the processing device based on these performance values.)
receiving input data comprising a neural network model subject to prediction and analysis requirement information; (Paragraph [0011] a method S100 for profiling neural network performance includes: … the target neural network including a set of layers defining a set of operations. Paragraph [0044] in one implementation, the system can access a target neural network. (accessing is basically analogous to receiving) Paragraph [0147] the system can: generate an initial schedule for a target neural network based on an initial set of parameters in Block S102 and Paragraph [0018] enabling the user to pinpoint bottlenecks in the neural network and adjust parameters and/or topology based on this feedback together indicates user gives initial parameters and later can still input adjustments. Both aspects are reflected in Figure 6 as it shows the compiler receives initial parameters and the target neural network as inputs.)
receiving information on hardware of a device in which the neural network model is run; (Paragraph [0030-0034] describes performance monitor register (e.g., performance counters) may be configured to monitor various hardware information (e.g., interrupt latency, idle time, task execution time). Paragraph [0066] describes that the system can access performance values captured by a set of performance counters in the processing device during the execution of the target neural network. In addition, paragraph [0024] indicates different device types can be used in the implementation, e.g., different number of resources, different specification of resources, different architecture, different manufacturer, different model.)
extracting layer information respectively corresponding to a plurality of layers configuring the neural network model; (Paragraph [0043-0046] describes that the system can access various layer information from the target neural network, for example a layer type, a set of input tensor dimensions, and a set of weight tensor dimensions etc.)
and predicting either one or both of operation performance information and energy efficiency information respectively corresponding to the plurality of layers by inputting the analysis requirement information and the layer information to the prediction model (Paragraph [0146-0149] describes the whole process as iterative, such that in a given round, the expected results output from the previous execution of the process are used as inputs to the given round and then various adjustments are made, and updated schedule and corresponding updated performance metrics are generated. Specifically, the system (e.g., a machine learning system), as the prediction model, can generate performance metrics based on an initial set of parameters and the layer classification. Paragraph [0050-0054] predict a set of expected performance metrics for the target neural network … based on the static schedule and the cost model (cost model defines a set of costs including computation costs, memory costs and power costs. The system can generate expected performance metrics at various levels of granularity, e.g., layer-level.)
Uppalapati fails to teach in detail, but Zhu teaches: building a prediction model based on the benchmark execution result and the hardware information; (Fig. 3 and paragraph [0042-0047] describes that the workload profile generator 110 receives a plurality of benchmark queries 302 and receives a plurality of benchmark queries 302, uses them to train a prediction model 112 to predict performance.)
Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention having the teachings of Uppalapati and Zhu in front of him/her to have modified the teachings of Uppalapati (directed to System And Method For Profiling On-Chip Performance Of Neural Network Execution), to incorporate the teachings of Zhu (directed to a prediction model to determine a candidate query sequence based on the determined set of performance characteristics) to build a prediction model based on the benchmark execution result and the hardware information. One of ordinary skill would have been motivated to perform such a modification in order to fulfil challenging analytic needs of the neural network workload via building a trained prediction model based on benchmark and hardware information as described in Zhu (Paragraph [0022] describes that evolving workloads increasing the complexity in benchmarking a database, a single standardized benchmark may not represent the variety of workloads and cover the analytic needs. Paragraph [0024] proposes the solution as a trained prediction model that predicts performance characteristics of a workload based on the workload and hardware configurations. The prediction model is trained using workload profiles generated from benchmark data and configuration data).
Regarding claim 2:
Uppalapati teaches: wherein the predicting comprises predicting the operation performance information and the energy efficiency information respectively corresponding to the plurality of layers for each component configuring the device (Paragraph [0053] Generally, the system can generate expected performance metrics representing performance of the target neural network on the processing device at various levels of granularity (e.g., descriptor-level, layer-level, network-level, device-level). Paragraph [0056] the system can generate the first expected performance metric representing an expected computational cost (in cycles of a processing unit) of executing the first operation by the processing device. Additionally, the system can generate expected performance metrics representing an expected bandwidth cost (in cycles of a DMA core) of the first operation, an expected power cost of the first operation, and/or an expected accuracy (or expected accuracy loss) of the first operation, etc. Paragraph [0063] The system can similarly generate an expected performance metric (or subset of expected performance metrics) for each resource (e.g., DMA core, transfer bus, communication bus, PCIe bus, AXI bus) in the set of resources of the processing device.)
Regarding claim 4:
Uppalapati teaches:
wherein the predicting comprises: classifying the plurality of layers into either one of a compute-bound layer and a memory- bound layer (Paragraph [0131] Generally, the system can identify each layer in the target neural network as a compute-bound layer or a bandwidth-bound layer.)
and predicting the operation performance information and the energy efficiency information respectively corresponding to the plurality of layers based on a result of the classifying (Paragraph [0148] the system (e.g., a machine learning system) can analyze these performance metrics—such as specific layers classified as compute-bound or bandwidth-bound, etc.—to generate an updated static schedule, for the target neural network.)
Regarding claim 5:
Uppalapati teaches: wherein the extracting of the layer information comprises extracting input data information and output data information respectively corresponding to the plurality of layers (Paragraph [0043-0044] the system can identify, for each layer in the set of layers: a layout of the layer relative to other layers in the set of layers; a layer type of the layer; a set of input tensor dimensions of the layer; and a set of weight tensor dimensions of the layer.)
Regarding claim 6:
Uppalapati teaches: wherein the obtaining of the benchmark execution result comprises: generating the benchmark execution result by executing the benchmark (Paragraph [0065-0067] describes that the system can generate a set of runtime performance metrics during execution of the target neural network by the processing device based on these performance values.)
While Uppalapati implies receiving user setting information (Paragraph [0018] and [0115] discuss the user receiving feedback, pinpointing bottlenecks in the neural network, and adjusting parameters and topology based on the feedback.), Zhu teaches directly receiving user setting information via user application (Zhu Fig. 1 and Paragraph [0033] illustrate that a user of computing device 104A may enter input via user application 118A or otherwise interact with the application to execute the workload.), such that, when combined with the teachings of Uppalapati, the user-provided settings of Zhu may additionally include user setting information related to obtaining of the benchmark execution result.
Uppalapati fails to teach, but Zhu teaches:
generating a benchmark based on the user setting information (Fig. 1 and paragraph [0033] illustrate that computing device 104N includes a developer application 118N that enables a user to perform developer operations (e.g., database benchmarking or workload replay). In other words, the benchmark is generated per user’s instruction with user setting information.)
Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention having the teachings of Uppalapati and Zhu in front of him/her to have modified the teachings of Uppalapati (directed to System And Method For Profiling On-Chip Performance Of Neural Network Execution), to incorporate the teachings of Zhu (directed to a prediction model to determine a candidate query sequence based on the determined set of performance characteristics) to receive user setting information and using it in generating benchmark. One of ordinary skill would have been motivated to perform such a modification in order to allow users tuning configurations and allow developers determining improvements as described in Zhu (Paragraph [0021] Database benchmarking and workload replay processes provide insight into performance characteristics, which allows developers to determine improvements in database engines (e.g., in query processing, query optimization, improvements in storage layers) and allows end-users, database administrators, and/or cloud service providers to tune configurations of a database.)
Regarding claim 7:
Uppalapati teaches receiving information on the neural network model subject to benchmark and on a component configuring the device. (Paragraph [0074] the processing device can capture a set of runtime performance metrics. In response to detecting a change in layer execution by the processing device (e.g., initiating execution of a current layer), initiate a performance metric extraction process to store performance metrics (e.g., layer-level performance metrics) captured by the performance monitor registers in the main memory of the processing device. Paragraph [0032] the processing device can include a corresponding set of performance monitor registers coupled to each associated component in order to provide component-specific performance metrics during execution of the scheduled parallel process representing the target neural network.)
Even though Uppalapati does not explicitly disclose receiving the setting information from a user, i.e. “user setting information,” Zhu teaches receiving various benchmark-related settings from a user (Paragraph [0033] For example, computing device 104A as shown in FIG. 1 includes a user application 118A that enables a user to execute a workload (e.g., including one or more queries) against a database. A user of computing device 104A may enter input via user application 118A or otherwise interact with the application to execute the workload. As also shown in FIG. 1, computing device 104N includes a developer application 118N that enables a user to perform developer operations (e.g., database benchmarking or workload replay) with respect to the database and/or workload logs 132. Paragraph [0046] Additional training data may be obtained from user input (e.g., via user application 118A or via developer application 118N)) such that, when combined with the teachings of Uppalapati, the user-provided settings of Zhu may additionally include user setting information related to the neural network and component.
Uppalapati fails to teach, but Zhu teaches:
(receiving) information on a preset operating frequency (Paragraph [0048] prediction model 112 is trained to determine performance characteristics based on one or more of a type of a workload, a hardware and/or software configuration, a query frequency of the workload, a query concurrency of the workload, and a product of the query frequency and the query concurrency of the workload.)
Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention having the teachings of Uppalapati and Zhu in front of him/her to have modified the teachings of Uppalapati (directed to System And Method For Profiling On-Chip Performance Of Neural Network Execution), to incorporate the teachings of Zhu (directed to a prediction model to determine a candidate query sequence based on the determined set of performance characteristics) to receive user setting information at a preset operating frequency in generating benchmark. One of ordinary skill would have been motivated to perform such a modification in order to control the variables during the benchmark generation as described in Zhu (Paragraph [0044] workload profile generator 110 generates workload profiles based on benchmark queries of benchmark queries 302 with performance that is impacted based on query frequency and/or query concurrency.)
Regarding claim 8:
Uppalapati teaches:
wherein the generating of the benchmark comprises, for each of the plurality of layers configuring the neural network model subject to the benchmark, generating the benchmark based on an input data size and an operating frequency of a layer (Paragraph [0044] In one implementation, the system can access a target neural network defined via a deep-learning framework (e.g., CAFFE, TENSORFLOW, or TORCH) to identify a set of layers in the neural network. In this implementation, the system can identify, for each layer in the set of layers: … a set of input tensor dimensions of the layer. Paragraph [0031] In one implementation, each performance monitor register can monitor (or capture) a performance value (or a set of performance values) during execution of a target neural network. For example, the processing device can include a set of performance counters configured to: monitor interrupt latency (in number of cycles of a control processor); monitor idle time (in number of cycles of a control processor); monitor task execution time (in number of cycles of a processing unit or DMA core); monitor task wait time (in number of cycles of a processing unit or DMA core); and/or monitor unallocated time (in number of cycles of a processing unit or DMA core). In this example, each performance counter is communicatively coupled (on silicon) to an associated region of the processing unit.” in number of cycles of …” reads as operating frequency. Paragraph [0092] In another implementation, the system can generate the set of runtime performance metrics including a second subset of runtime performance metrics for the set of layers. In this implementation, the system can generate a runtime performance metric (or a subset of runtime performance metrics) for each layer in the set of layers.)
Regarding claim 10:
Uppalapati teaches:
wherein the generating of the benchmark execution result comprises: performing a first measurement on operation performance and energy efficiency corresponding to the benchmark based on an application programming interface (API);
performing a second measurement on operation performance and energy efficiency corresponding to the benchmark using an external device;
and performing consistency determination on the benchmark execution result by comparing a result of the first measurement and a result of the second measurement (Paragraph [0021-0024] describes in one implementation, the system can include a first processing device and a second processing device characterized by a different device type than the first device, profile execution of the target neural network on both of them, then compare the difference to choose a higher-performing device. Paragraph [0022] specifically mentions that the system can include additional processing devices communicatively coupled to the client device via the processing device interface (or additional processing device interfaces), so the first measurement can be based on the processing device interface, i.e. API.)
Regarding claim 11:
Uppalapati teaches:
A non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors, configure the one or more processors to perform the method of claim 1 (Paragraph [0150] The systems and methods described herein can be embodied and/or implemented at least in part as a machine configured to receive a computer-readable medium storing computer-readable instructions.)
Regarding claim 12:
Incorporating the rejection of claim 1, Uppalapati further teaches:
An apparatus comprising: one or more processors configured to… (Paragraph [0027] Generally, a processing device can include: a set of resources (e.g., processing units, queue processor(s)))
Regarding claim 13, 15 - 20:
Incorporates the rejection of claim 2, 4-8, 10 correspondingly in addition to the rejection of claim 12.
Claims 3 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Uppalapati in view of Zhu, and further in view of Yang et al, ("Designing Energy-Efficient Convolutional Neural Networks Using Energy-Aware Pruning," 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 2017, pp. 6071-6079, doi: 10.1109/CVPR.2017.643., hereinafter “Yang”).
Regarding claim 3:
Uppalapati fails to teach, but Yang teaches wherein the predicting comprises:
for each of the plurality of layers, determining a weight between a computation amount and a memory access amount in a layer; and predicting the operation performance information and the energy efficiency information respectively corresponding to the plurality of layers based on the weight (Section 2.2 describes a method for estimating the energy consumption of a CNN. For each CNN layer, the framework calculates the energy consumption by dividing it into two parts: computation energy consumption, and data movement energy consumption. It further accounts for the impact of data sparsity and bitwidth reduction on energy consumption, and quantifies the impact of bitwidth by scaling the energy cost of different hardware components accordingly, for example, energy consumption and memory access. Scaling reads as assigning a weight in the reference.)
Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention having the teachings of Uppalapati and Yang in front of him/her to have modified the teachings of Uppalapati (directed to System And Method For Profiling On-Chip Performance Of Neural Network Execution), to incorporate the teachings of Yang (directed to Designing Energy-Efficient Convolutional Neural Networks Using Energy-Aware Pruning) to assign a weight between computation and memory costs in a layer and predict the operation performance and the energy efficiency based on this weight. One of ordinary skill would have been motivated to perform such a modification in order to better estimate the energy consumption of neural network at design stage as described in Yang (Section 2.3 With this methodology, we can quantify the difference in energy costs between various popular CNN models and methods, such as increasing data sparsity or aggressive bitwidth reduction (discussed in Sec. 5). More importantly, it provides a gateway for researchers to assess the energy consumption of CNNs at design time, which can be used as a feedback that leads to CNN designs with significantly reduced energy consumption.)
Regarding claim 14:
Incorporates the rejection of claim 3 in addition to the rejection of claim 12.
Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Uppalapati in view of Zhu, and further in view of Du et al, (CN 107168859 A, hereinafter “Du”).
Regarding claim 9:
Uppalapati fails to teach, but Du teaches:
wherein the generating of the benchmark comprises generating the benchmark based on a kernel corresponding to a component configuring the device (Page 2 paragraph 4 In the method of the present invention, the energy consumption feature dataset is obtained by collecting kernel-level energy consumption characteristic values when the component induces a drive call event.)
Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention having the teachings of Uppalapati and Du in front of him/her to have modified the teachings of Uppalapati (directed to System And Method For Profiling On-Chip Performance Of Neural Network Execution), to incorporate the teachings of Du (directed to Energy Analysis Method Of The Android Device) to generate benchmark base on a kernel corresponding to a hardware component. One of ordinary skill would have been motivated to perform such a modification in order to extract as described in Du (Page 2 Paragraph 8 Compared with the prior art, the advantages of the present invention are: it enables kernel-based energy feature acquisition, tracks key state transformation events driven by hardware drivers including CPU, GPU, Flash, LTE, Wi-Fi, and display, and extracts energy consumption features for training energy consumption models.)
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to WENNA DUAN whose telephone number is (571)270-5795. The examiner can normally be reached Monday-Friday (9:00 am - 5:00 pm).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, ABDULLAH KAWSAR can be reached at (571)270-3169. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/W.D./Examiner, Art Unit 2127
/JEREMY L STANLEY/Examiner, Art Unit 2127