DETAILED ACTION
Claims 1-20 are pending in the application.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
Examiner’ s Notes
The Examiner cites particular sections in the references as applied to the claims below for the convenience of the applicant(s). Although the specified citations are representative of the teachings in the art and are applied to the specific limitations within the individual claim, other passages and figures may apply as well. It is respectfully requested that, in preparing responses, the applicant(s) fully consider the references in their entirety as potentially teaching all or part of the claimed invention, as well as the context of the passage as taught by the prior art or disclosed by the Examiner.
Drawings
Figures 1 and 2 should be designated by a legend such as --Prior Art-- because only that which is old is illustrated. See MPEP § 608.02(g). Corrected drawings in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. The replacement sheet(s) should be labeled “Replacement Sheet” in the page header (as per 37 CFR 1.84(c)) so as not to obstruct any portion of the drawing figures. If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
Claim 20 reads as follows:
“An apparatus for distributing execution of at least one artificial intelligence (AI) workload across a plurality of processing cores, comprising:
means for obtaining first characteristics for the plurality of processing cores;
means for obtaining second characteristics for the at least one AI workload;
means for computing one or more metrics for each of the processing cores, based on the first characteristics and the second characteristics; and
means for scheduling the at least one AI workload on at least one of the processing cores, based on the computed metrics and one or more conditions.”
Claim 20 uses the phrase “means for” without being modified by sufficient structure, material, or acts for performing the claimed function. Therefore, the claim limitation fails part c of the three-pronged test as explained in MPEP § 2181, subsection I. In the specification, paragraph [0118] discloses what the means for obtaining, means for computing, and means for scheduling comprise.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step1:
Claim 1 is directed to an apparatus for distributing execution of artificial intelligence workloads across processing cores, the apparatus comprising: a series of physical hardware configured to execute instructions, causing the apparatus to complete a series of steps, and is therefore directed to a product, which is one of the four statutory categories.
Step 2A, Prong One:
Claim 1 recites the limitation:
Compute one or more metrics for each of the processing cores, based on the first characteristics and the second characteristics;
this step can be performed in the human mind through computation, with the aid of pen and paper, and therefore recites a mental process.
Accordingly, claim 1 recites a judicial exception (i.e., an abstract idea).
Step 2A, Prong Two:
The additional elements recited in claim 1 include:
At least one memory comprising computer-executable instructions; and
One or more processors configured to execute the computer-executable instructions and cause the apparatus to:
Obtain first characteristics for the plurality of processing cores;
Obtain second characteristics for the at least one AI workload;
Schedule the at least one AI workload on at least one of the processing cores, based on the computed metrics and one or more conditions.
Regarding the additional elements (i) and (ii), the limitations recited are generic computer components, which have been identified by the courts as well-understood, routine, and conventional. See MPEP 2106.05(d).
Regarding the additional elements (iii), (iv), and (v), the limitations recited are insignificant extra-solution activity. Additional elements (iii) and (iv) recite mere data gathering and additional element (v) recites mere instructions to apply the mental process, based on the data gathering. See MPEP 2106.05(g)
Step 2B:
Regarding the additional elements (i) and (ii), the limitations are reciting generic computing components to perform the steps of mere data gathering, mental process, and mere instructions to apply. The courts have found adding well-understood, routine, and conventional activity is not enough to amount to significantly more than the recited judicial exception. See MPEP 2106.05(g).
Regarding the additional elements (iii), (iv), and (v), the limitations recited are insignificant extra-solution activity. Additional elements (iii) and (iv) recite mere data gathering and additional element (v) recites mere instructions to apply the mental process, based on the data gathering. The courts have found adding insignificant extra-solution activity is not enough to amount to significantly more than the recited judicial exception. See MPEP 2106.05(g)
The combination of these additional elements amounts to a product comprising well understood, routine, and conventional generic computer components that are utilized to perform the steps of mere data gathering, computation based on the data gathering (mental process), and mere instructions to apply the mental process.
Therefore, the additional elements, when considered individually and in combination, fail to add an inventive concept to the claim.
Consequently, claim 1 as a whole does not amount to significantly more than the recited judicial exceptions and the claim is not eligible.
Claim 2 is dependent on claim 1, and therefore inherits the same judicial exception recited in claim 1. Claim 2 recites the limitation wherein the plurality of processing cores comprise at least one of:
A central processing unit (CPU),
A graphics processing unit (GPU),
A neural processor, or
Another type of processing core.
The limitation recited above is still generic computer components that are well-understood, routine, and conventional in the art. See MPEP 2106.05(g).
Claim 2 does not recite any additional elements beyond those recited in claim 1. Accordingly, for the same reasons presented with respect to claim 1, the additional limitations are not indicative of integration into a practical application, nor do they amount to significantly more than the recited judicial exception. Thus, claim 2 is not eligible.
Claim 3 is dependent on claim 1, and therefore inherits the same judicial exception recited in claim 1. Claim 3 recites the limitations wherein:
the at least one AI workload involves multiple stages; and
metrics are computed for each of the processing cores for each of the multiple stages.
Regarding the additional limitation (i), the limitation recited is insignificant extra-solution activity. Additional limitation (i) recites activity that is merely a tangential addition to the claim. See MPEP 2106.05(g)
Regarding the additional limitation (ii), the limitation recited is a judicial exception (i.e., an abstract idea). This step can be performed in the human mind through computation, with the aid of pen and paper, and therefore recites a mental process. See MPEP 2106.05(g)
The additional limitations recited in claim 3 recite insignificant extra-solution activity, followed by an abstract idea. See MPEP 2106.04(a) and 2106.05(d). The combination of the elements of claim 1 and the limitations of claim 3 amounts to a product comprising well understood, routine, and conventional generic computer components that are utilized to perform the steps of mere data gathering, computation based on the data gathering (mental process), and mere instructions to apply the mental process.
Claim 3 does not recite any additional elements beyond those recited in claim 1. Accordingly, for the same reasons presented with respect to claim 1, the additional elements are not indicative of integration into a practical application, nor do they amount to significantly more than the recited judicial exceptions. Thus, claim 3 is not eligible.
Claim 4 is dependent on claim 3, and therefore inherits the same judicial exception recited in claims 1 and 3. Claim 4 recites the limitations wherein in order to schedule the at least one AI workload on at least one of the processing cores, the one or more processors are further configured to schedule different stages to different processing cores based on different conditions.
The additional limitations recited in claim 4 are insignificant extra-solution activity. Further configuring the processors to schedule different stages to different processing cores based on different conditions is mere instructions to apply the judicial exception. See MPEP 2106.05(g)
Claim 4 does not recite any additional elements beyond those recited in claims 1 and 3. Accordingly, for the same reasons presented with respect to claims 1 and 3, the additional elements are not indicative of integration into a practical application, nor do they amount to significantly more than the recited judicial exception. Thus, claim 4 is not eligible.
Claim 5 is dependent on claim 1, and therefore inherits the same judicial exception recited in claim 1. Claim 5 recites the limitations wherein:
the first characteristics comprise at least one of a dynamic power slope, a leakage power, or dynamic frequency and voltage operating points; and
the second characteristics comprise at least one of a number of stages for the at least one AI workload, instructions per cycle (IPC) for the at least one AI workload, or total instructions for the at least one AI workload.
The additional limitations recited in claim 5 are insignificant extra-solution activity. Further specification of the data gathered prior to the judicial exception is merely a tangential addition to the claim. See MPEP 2106.05(g)
Accordingly, for the same reasons presented with respect to claim 1, the additional limitations are not indicative of integration into a practical application, nor do they amount to significantly more than the recited judicial exception. Thus, claim 5 is not eligible.
Claim 6 is dependent on claim 5, and therefore inherits the same judicial exception recited in claim 5. Claim 6 recites the limitation wherein the dynamic frequency and voltage operating points comprise at least one of dynamic clock and voltage scaling (DCVS) points or dynamic voltage and frequency scaling (DVFS) points.
The additional limitation recited in claim 6 is insignificant extra-solution activity. Further specification of the data gathered prior to the judicial exception is merely a tangential addition to the claim. See MPEP 2106.05(g)
Accordingly, for the same reasons presented with respect to claim 5, the additional limitations are not indicative of integration into a practical application, nor do they amount to significantly more than the recited judicial exception. Thus, claim 6 is not eligible.
Claim 7 is dependent on claim 5, and therefore inherits the same judicial exception recited in claim 1. Claim 7 recites the limitations wherein the metrics comprise at least one of a compute time, a total power consumption, or total energy consumption for the at least one AI workload.
The additional limitation recited in claim 7 further specifies over the judicial exception recited in claim 1. This step can be performed in the human mind through computation, with the aid of pen and paper, and therefore recites a mental process. Further specification of the metrics being computed describe a mathematical calculation. See MPEP 2106.05(g).
Accordingly, for the same reasons presented with respect to claims 1 and 5, the additional limitations are not indicative of integration into a practical application, nor do they amount to significantly more than the recited judicial exception. Thus, claim 7 is not eligible.
Claim 8 is dependent on claim 7, and therefore inherits the same judicial exception recited in claim 1. Claim 8 recites the limitation wherein the one or more conditions relate to a desired behavior.
The additional limitation recited in claim 8 refers to taking a desired behavior into consideration when scheduling the AI workload. This step can be performed in the human mind, and therefore recites a mental process. See MPEP 2106.05(g).
Accordingly, for the same reasons presented with respect to claim 7, the additional limitations are not indicative of integration into a practical application, nor do they amount to significantly more than the recited judicial exception. Thus, claim 8 is not eligible.
Claim 9 is dependent on claim 8, and therefore inherits the same judicial exception recited in claim 1. Claim 9 recites the limitation wherein the desired behavior relates to optimization of compute time, total power consumption, or a balance of compute time and power consumption.
The additional limitation recited in claim 9 further limits the scope of the desired behavior considered when scheduling the at least one AI workload. The element of scheduling the at least one AI workload recited in claim 1 is still insignificant extra-solution activity with the further limitation considered, because it is just mere instructions to apply the judicial exception. See MPEP 2106.05(g).
Accordingly, for the same reasons presented with respect to claim 8, the additional limitations are not indicative of integration into a practical application, nor do they amount to significantly more than the recited judicial exception. Thus, claim 9 is not eligible.
Claim 10 is dependent on claim 1, and therefore inherits the same judicial exception recited in claim 1. Claim 10 recites the limitations wherein:
the first characteristics comprise dynamic frequency and voltage operating points and a total bandwidth for each dynamic frequency and voltage operating point; and
the second characteristics comprise a priority level and a bandwidth requirement for a given AI workload.
The additional limitations recited in claim 10 are insignificant extra-solution activity. Further specification of the data gathered prior to the judicial exception is merely a tangential addition to the claim. See MPEP 2106.05(g)
Accordingly, for the same reasons presented with respect to claim 1, the additional limitations are not indicative of integration into a practical application, nor do they amount to significantly more than the recited judicial exception. Thus, claim 10 is not eligible.
Claim 11 is dependent on claim 10, and therefore inherits the same judicial exception recited in claim 1. Claim 11 recites the limitation wherein the dynamic frequency and voltage operating points comprise at least one of dynamic clock and voltage scaling (DCVS) points or dynamic voltage and frequency scaling (DVFS) points.
The additional limitation recited in claim 11 is insignificant extra-solution activity. Further specification of the data gathered prior to the judicial exception is merely a tangential addition to the claim. See MPEP 2106.05(g)
Accordingly, for the same reasons presented with respect to claims 1 and 10, the additional limitations are not indicative of integration into a practical application, nor do they amount to significantly more than the recited judicial exception. Thus, claim 11 is not eligible.
Claim 12 is dependent on claim 10, and therefore inherits the same judicial exception recited in claim 1. Claim 12 recites the limitation wherein the metrics comprise an available bandwidth for the given AI workload for each processing core at one or more of the dynamic frequency and voltage operating points.
The additional limitation recited in claim 12 further specifies over the judicial exception recited in claim 1. This step can be performed in the human mind through computation, with the aid of pen and paper, and therefore recites a mental process. Further specification of the metrics being computed describe a mathematical calculation. See MPEP 2106.05(g).
Accordingly, for the same reasons presented with respect to claims 1 and 10, the additional limitations are not indicative of integration into a practical application, nor do they amount to significantly more than the recited judicial exception. Thus, claim 12 is not eligible.
Claim 13 is dependent on claim 12, and therefore inherits the same judicial exception recited in claim 1. Claim 13 recites the limitation wherein the one or more conditions depend on the priority level for the given AI workload relative to a priority level for a non-AI workload.
The additional limitation recited in claim 13 refers to taking a priority comparison into consideration when scheduling the AI workload. A decision of evaluating priority can be performed in the human mind, and therefore recites a mental process. See MPEP 2106.05(g).
Accordingly, for the same reasons presented with respect to claims 1, 10, and 12, the additional limitations are not indicative of integration into a practical application, nor do they amount to significantly more than the recited judicial exception. Thus, claim 13 is not eligible.
Claim 14 is dependent on claim 13, and therefore inherits the same judicial exception recited in claim 1. Claim 14 recites the limitation wherein in order to schedule the at least one AI workload on at least one of the processing cores when the priority level for the given AI workload is less than the priority level for the non-AI workload, the one or more processors are further configured to schedule the given AI workload on a selected one of the processing cores at a dynamic frequency and voltage operating point where the required bandwidth of the selected processing core for AI workload is less than the available bandwidth after scheduling the non-AI workload.
The additional limitation recited in claim 14 refers to configuring generic computing components to perform scheduling based on a condition. The scheduling is mere instructions to apply the judicial exception recited in claim 1, and the logic behind the scheduling based on known conditions can be performed in the human mind, therefore reciting a mental process. See MPEP 2106.05(g).
Accordingly, for the same reasons presented with respect to claims 1, 10, 12, and 13, the additional limitations are not indicative of integration into a practical application, nor do they amount to significantly more than the recited judicial exception. Thus, claim 14 is not eligible.
Claim 15 is dependent on claim 13, and therefore inherits the same judicial exception recited in claim 1. Claim 15 recites the limitation wherein in order to schedule the at least one AI workload on at least one of the processing cores when the priority level for the given AI workload is greater than the priority level for the non-AI workload, the one or more processors are further configured to: schedule the given AI workload on a selected one of the processing cores at a dynamic frequency and voltage operating point based on the computed metrics and one or more conditions; and allocate remaining bandwidth of the selected processing core for the non-AI workload.
The additional limitation recited in claim 15 refers to configuring generic computing components to perform scheduling based on a condition. As discussed, the scheduling is mere instructions to apply the judicial exception recited in claim 1, and the logic behind the scheduling based on known conditions can be performed in the human mind, therefore reciting a mental process. See MPEP 2106.05(g).
Accordingly, for the same reasons presented with respect to claims 1, 10, 12, and 13, the additional limitations are not indicative of integration into a practical application, nor do they amount to significantly more than the recited judicial exception. Thus, claim 15 is not eligible.
Claim 16 is dependent on claim 1, and therefore inherits the same judicial exception recited in claim 1. Claim 16 recites the limitation wherein the one or more conditions relate to at least one of a desired behavior, a core junction temperature, and a battery capacity condition.
The additional limitation recited in claim 16 limits the scope of the conditions considered when scheduling the at least one AI workload. The element of scheduling the at least one AI workload recited in claim 1 is still insignificant extra-solution activity with the further limitation considered, because it is just mere instructions to apply the judicial exception. See MPEP 2106.05(g).
Accordingly, for the same reasons presented with respect to claim 1, the additional limitations are not indicative of integration into a practical application, nor do they amount to significantly more than the recited judicial exception. Thus, claim 16 is not eligible.
Claim 17 is dependent on claim 16, and therefore inherits the same judicial exception recited in claim 1. Claim 17 recites the limitation wherein:
the desired behavior relates to optimization of compute time, total power consumption, or a balance of compute time and power consumption; and
the desired behavior changes based on at least one of: the core junction temperature relative to a first threshold, or the battery capacity condition relative to a second threshold.
The additional limitations recited in claim 17 further limit the scope of the desired behavior considered when scheduling the at least one AI workload. The element of scheduling the at least one AI workload recited in claim 1 is still insignificant extra-solution activity with the further limitation considered, because it is just mere instructions to apply the judicial exception. See MPEP 2106.05(g).
Accordingly, for the same reasons presented with respect to claims 1 and 16, the additional limitations are not indicative of integration into a practical application, nor do they amount to significantly more than the recited judicial exception. Thus, claim 17 is not eligible.
Claim 18 is dependent on claim 1, and therefore inherits the same judicial exception recited in claim 1. Claim 18 recites the limitations wherein:
the second characteristics comprise at least one of a token input length or a token input characteristic for the at least one AI workload,
the one or more conditions relate to a desired behavior, and
the desired behavior changes based on the token input length relative to at least one threshold.
The additional limitations recited in claim 18 further limit the scope of the characteristics, conditions, and behavior from claim 1.
Claim limitation (i) limits the following element from claim 1: obtain second characteristics for the at least one AI workload. With the further limitation of scope, this claim element is still mere data gathering. See MPEP 2106.05(g).
Claim limitation (ii) limits the following element from claim 1: schedule the at least one AI workload on at least one of the processing cores, based on the computed metrics and one or more conditions. With the further limitation of scope, this claim element is still insignificant extra-solution activity (mere data gathering and mere instructions to apply it). See MPEP 2106.05(g).
Claim limitation (iii) limits the following element from claim 1: schedule the at least one AI workload on at least one of the processing cores, based on the computed metrics and one or more conditions. With the further limitation of scope, this claim element is still insignificant extra-solution activity (mere data gathering and mere instructions to apply it). See MPEP 2106.05(g).
Accordingly, for the same reasons presented with respect to claim 1, the additional limitations are not indicative of integration into a practical application, nor do they amount to significantly more than the recited judicial exception. Thus, claim 18 is not eligible.
Claim 19 recites A method to distribute execution of at least one artificial intelligence (AI) workload across a plurality of processing cores, the method comprising: the steps provided in the product of claim 1. Thus, for the same reason presented with respect to claim 1, claim 19 is rejected because the claimed invention is directed to an abstract idea without significantly more.
Claim 20 recites substantially the same limitations as those recited in claim 1, under the interpretation of 35 U.S.C 112(f). After careful consideration of the specification, for the same reasons presented with respect to claim 1, claim 20 is directed to an abstract idea without significantly more and is not eligible.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
This application currently names joint inventors. In considering the patentability of the claims, the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1-4, 16, and 19-20 are rejected under 35 U.S.C. 102(a)(1) and 102(a)(2) as being anticipated by Nimmagadda et al. (US 2021/0382754 A1; herein after “Nimmagadda”).
With respect to claim 1, Nimmagadda teaches:
An apparatus for distributing execution of at least one artificial intelligence (AI) workload across a plurality of processing cores, comprising:
at least one memory comprising computer-executable instructions ([0070] – “FIG. 10 also illustrates a memory 270 coupled to the processor core 200. The memory 270 may be any of a wide variety of memories (including various layers of memory hierarchy) as are known or otherwise available to the of skill in the art. The memory 270 may include one or more code 213 instruction(s) to be executed by the processor core 200,”); and
one or more processors configured to execute the computer-executable instructions and cause the apparatus to ([0084] – “…and a memory coupled to the processor, the memory including a set of executable program instructions, which when executed by the processor, cause the processor to…):
obtain first characteristics for the plurality of processing cores ([0055] – “In some examples, method 800 further includes identifying data formats supported by available hardware resources, identifying DL operators that are supported by available hardware resources and selecting the plurality of hardware devices from the available hardware resources based on the data formats and the deep learning operators.”);
obtain second characteristics for the at least one AI workload ([0054] – “Illustrated processing block 802 analyzes an input stream and an AI model graph to generate a workload characterization, where the workload characterization characterizes one or more of compute resources or memory resources and the one or more of compute resources or the memory resources is associated with the execution of the AI model graph based on the input stream.”);
compute one or more metrics for each of the processing cores, based on the first characteristics and the second characteristics ([0098] – “wherein the instructions, when executed, further cause the computing system to identify data formats supported by available hardware resources, identify deep learning operators that are supported by the available hardware resources, and select the plurality of hardware devices from the available hardware resources based on the data formats and the deep learning operators.”); and
schedule the at least one AI workload on at least one of the processing cores, based on the computed metrics and one or more conditions ([0097] – “wherein the one or more of the compute resources or the memory resources is associated with execution of the AI model graph based on the input stream, partition the AI model graph into subgraphs based on the workload characterization, and select a plurality of hardware devices to execute the subgraphs.”).
With respect to claim 2, Nimmagadda teaches:
The apparatus of claim 1 (see claim 1), wherein the plurality of processing cores comprise at least one of:
a central processing unit (CPU),
a graphics processing unit (GPU),
a neural processor,
or another type of processing core. ([0069] – “FIG. 10 illustrates a processor core 200 according to one embodiment. The processor core 200 may be the core for any type of processor, such as a micro-processor, an embedded processor, a digital signal processor (DSP), a network processor, or other device to execute code.”)
With respect to claim 3, Nimmagadda teaches:
The apparatus of claim 1 (see claim 1), wherein:
the at least one AI workload involves multiple stages ([0090] – “wherein the one or more of the compute resources or the memory resources is associated with the execution of the AI model graph based on the input stream, partition the AI model graph into subgraphs based on the workload characterization, and select a plurality of hardware devices to execute the subgraphs.”); and
metrics are computed for each of the processing cores for each of the multiple stages ([0033] – “The compute and the intermediate data output sizes may thus be mapped to the AI model. These metrics may be used to determine the AI workload partitioning.”)
For clarity of the record, the examiner would like to point out that the terminology of AI subgraphs and stages are of significant correlation when referring to an AI workload’s lifecycle. In the prior art, the partitioning of an AI model into subgraphs indicates that there are indeed multiple stages correlated to the subgraphs. Additionally, the metrics mapped to the AI model are partitioned into the subgraphs, thus computing metrics for each stage.
With respect to claim 4, Nimmagadda teaches:
The apparatus of claim 3 (see claim 3), wherein in order to schedule the at least one AI workload on at least one of the processing cores, the one or more processors are further configured to schedule different stages to different processing cores based on different conditions. ([0084] – “wherein the one or more of the compute resources or the memory resources is associated with execution of the AI model graph based on the input stream, partition the AI model graph into subgraphs based on the workload characterization, and select a plurality of the hardware devices to execute the subgraphs.”)
With respect to claim 16, Nimmagadda teaches:
The apparatus of claim 1 (see claim 1), wherein the one or more conditions relate to at least one of a desired behavior, a core junction temperature, and a battery capacity condition. ([0041] – “A user requirements manager 310 provides user requirements and/or constraints. The user requirement manager 310 ingests requirements provided by the user regarding performance, accuracy, device of choice, priority of the models etc. The partitioner 306 may partition the input stream and AI model based on the user constraints.”)
With respect to claim 19: Claim 19 is directed to a method corresponding to the active functions implemented by the apparatus recited in claim 1; please see the rejection directed to claim 1 above which also cover the limitations recited in claim 19.
With respect to claim 20: Claim 20 is directed to an apparatus substantially identical to the apparatus of claim 1; please see the rejection directed to claim 1 above which also cover the limitations recited in claim 20.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 5-9 are rejected under 35 U.S.C. 103 as being unpatentable over Nimmagadda in view of Verrilli et al. (US 2022/0317758 A1); herein after “Verrilli”).
With respect to claim 5, Nimmagadda teaches:
The apparatus of claim 1 (see 35 U.S.C. 102 rejection for claim 1), wherein: […]
the second characteristics comprise at least one of a number of stages for the at least one AI workload, instructions per cycle (IPC) for the at least one AI workload, or total instructions for the at least one AI workload. ([0033] – “For example, the AI model may include a series of layers.”)
Nimmagadda fails to teach but Verrilli teaches the first characteristics comprise at least one of a dynamic power slope, a leakage power, or dynamic frequency and voltage operating points. (Verrilli, [0002] – “Dynamic clock and voltage scaling (“DCVS”) is a technique by which the frequency and voltage at which a processor is operated are adjusted dynamically, i.e., in real time in response to changes in operating conditions, to deliver a desired balance or tradeoff between power consumption and performance level.”)
Nimmagadda and Verrilli are analogous art because they are in the same field of endeavor: inter-process communications. Therefore, it would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention to modify Nimmagadda with the teachings of Verrilli. The motivation is taught by Verrilli in paragraph [0002]: “to deliver a desired balance or tradeoff between power consumption and performance level.”
With respect to claim 6, Nimmagadda and Verrilli teach:
The apparatus of claim 5 (see claim 5),
Nimmagadda fails to teach, but Verrilli teaches wherein the dynamic frequency and voltage operating points comprise at least one of dynamic clock and voltage scaling (DCVS) points or dynamic voltage and frequency scaling (DVFS) points. (Verrilli, [0002] – “Dynamic clock and voltage scaling (“DCVS”) is a technique by which the frequency and voltage at which a processor is operated are adjusted dynamically, i.e., in real time in response to changes in operating conditions, to deliver a desired balance or tradeoff between power consumption and performance level.”)
Nimmagadda and Verrilli are analogous art because they are in the same field of endeavor: inter-process communications. Therefore, it would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention to modify Nimmagadda with the teachings of Verrilli. The motivation is taught by Verrilli in paragraph [0002]: “to deliver a desired balance or tradeoff between power consumption and performance level.”
With respect to claim 7, Nimmagadda and Verrilli teach:
The apparatus of claim 5 (see claim 5),
Nimmagadda also teaches wherein the metrics comprise at least one of a compute time, a total power consumption, or total energy consumption for the at least one AI workload. (Nimmagadda, [0033] – “The graph computer analyzer 304b may determine compute (e.g., FLOPS or TOPS) present in each layer and intermediate data output sizes may thus be mapped to the AI model. These metrics may be used to determine the AI workload partitioning.”)
Nimmagadda and Verrilli are analogous art because they are in the same field of endeavor: inter-process communications. Therefore, it would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention to modify Nimmagadda with the teachings of Verrilli. The motivation would be to leverage well known task metrics to make educated scheduling decisions.
With respect to claim 8, Nimmagadda and Verrilli teach:
The apparatus of claim 7 (see claim 7),
Nimmagadda also teaches wherein the one or more conditions relate to a desired behavior. (Nimmagadda, [0041] – “A user requirements manager 310 provides user requirements and/or constraints. The user requirement manager 310 ingests requirements provided by the user regarding performance, accuracy, device of choice, priority of the models etc. The partitioner 306 may partition the input stream and AI model based on the user constraints.”)
Nimmagadda and Verrilli are analogous art because they are in the same field of endeavor: inter-process communications. Therefore, it would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention to modify Nimmagadda with the teachings of Verrilli. The motivation would be to allow for more flexible personalization based on individual desired behavior at a given time.
With respect to claim 9, Nimmagadda and Verrilli teach:
The apparatus of claim 8 (see claim 8),
Nimmagadda also teaches wherein the desired behavior relates to optimization of compute time, total power consumption, or a balance of compute time and power consumption. (Nimmagadda, [0041] – “A user requirements manager 310 provides user requirements and/or constraints. The user requirement manager 310 ingests requirements provided by the user regarding performance, accuracy, device of choice, priority of the models etc. The partitioner 306 may partition the input stream and AI model based on the user constraints.”)
For clarity of the record, the Examiner would like to point out that the optimization of compute time, total power consumption, or a balance of compute time and power consumption are inherently performance indicators.
Nimmagadda and Verrilli are analogous art because they are in the same field of endeavor: inter-process communications. Therefore, it would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention to modify Nimmagadda with the teachings of Verrilli. The motivation would be to allow for more flexible personalization based on individual desired behavior at a given time.
Claims 10-12 are rejected under 35 U.S.C. 103 as being unpatentable over Nimmagadda in view of Kim et al. (US 2019/0384370 A1); herein after “Kim”).
With respect to claim 10, Nimmagadda teaches:
The apparatus of claim 1 (see 35 U.S.C. 102 rejection for claim 1), wherein:
[…]
the second characteristics comprise a priority level (Nimmagadda, [0041] – “A user requirement manager 310 provides user requirements and/or constraints. The user requirement manager 310 ingests requirements provided by the user regarding performance, accuracy, device of choice, priority of the models etc. The partitioner 306 may partition the input stream and AI model based on the user constraints. For example, the partitioner 306 may treat the user requirements as constraints that are to be met while optimizing the resource allocation.”)
and a bandwidth requirement for a given AI workload. (Nimmagadda, [0031] – “The analyzer 304 characterizes the workload of the edge function 302 based on the input stream and the AI model to create a workload profile. The workload profile characterizes compute resources and/or memory resources that will be required during execution of the edge function 302. In some examples, the workload profile represents types of operations in the model, whether specific hardware units are required for execution, input streams, resolution, bitrates, encode formats, decode formats, compute in the AI model (e.g., measured in floating point operations (FLOPS), tera operations (TOPS), etc.), whether certain data formats (e.g., fp16, bfloat16, int16, etc.) are required for execution and so forth.”)
Nimmagadda fails to teach but Kim teaches the first characteristics comprise dynamic frequency and voltage operating points and a total bandwidth for each dynamic frequency and voltage operating point; (Kim, [0029] – “In one example, a DL/AI application can be compiled with power-performance optimization hints. These software hints can be used by hardware to assist dynamic voltage, frequency, and pipeline scaling, and bandwidth utilization control for the executing the DL/AI application. Dynamic voltage, frequency, and pipeline scaling hardware allows execution pipeline utilization, as well as the supply voltages and operating frequencies, to be changed dynamically. Pipelines can be modulated to control system utilization. Also, periodic sampling of performance, power, and temperature monitors may be performed. Collaboration of hardware, firmware, and software maximizes performance based on pre-knowledge of actual workloads and operating conditions.”)
Nimmagadda and Kim are analogous art because they are in the same field of endeavor: inter-process communications. Therefore, it would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention to modify Nimmagadda with the teachings of Kim. The motivation is taught by Kim in paragraph [0029]: “Collaboration of hardware, firmware, and software maximizes performance based on pre-knowledge of actual workloads and operating conditions.”
With respect to claim 11, Nimmagadda and Kim teach:
The apparatus of claim 10 (see 35 U.S.C. 103 rejection for claim 10),
Nimmagadda fails to teach but Kim teaches wherein the dynamic frequency and voltage operating points comprise at least one of dynamic clock and voltage scaling (DCVS) points or dynamic voltage and frequency scaling (DVFS) points. (Kim, [0027] – “One particular approach involves dynamic voltage frequency scaling (DVFS) technology, which is a technique aimed at reducing dynamic power consumption by dynamically adjusting voltage and frequency.”)
Nimmagadda and Kim are analogous art because they are in the same field of endeavor: inter-process communications. Therefore, it would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention to modify Nimmagadda with the teachings of Kim. The motivation would be to effectively reduce dynamic power consumption.
With respect to claim 12, Nimmagadda and Kim teach:
The apparatus of claim 10 (see 35 U.S.C. 103 rejection for claim 10),
Nimmagadda fails to teach but Kim teaches wherein the metrics comprise an available bandwidth for the given AI workload for each processing core at one or more of the dynamic frequency and voltage operating points. (Kim, [0028] – “Deep learning/artificial intelligence (DL/AI) workloads have predictable behavior across compute kernels, topologies, etc. One or more embodiments of an integrated circuit with software assisted power management described herein exploit this information to implement proactive power management policies to optimize performance (e.g., compute rates vs. memory access bandwidth) of the hardware by balancing power consumption between execution units (e.g., matrix processing units (MPUs)) and memory (e.g., high bandwidth memory (HBM)). This can be achieved by distinguishing different resources (e.g., compute or memory) needed for each processing phase of an instruction stream and by prioritizing their usages.”)
Nimmagadda and Kim are analogous art because they are in the same field of endeavor: inter-process communications. Therefore, it would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention to modify Nimmagadda with the teachings of Kim. The motivation would be to allocate resources according to availability.
Claim 13 is rejected under 35 U.S.C. 103 as being unpatentable over Nimmagadda and Kim, further in view of Musleh et al. (US 2021/0092069 A1); herein after “Musleh”).
With respect to claim 13, Nimmagadda and Kim teach:
The apparatus of claim 12 (see 35 U.S.C. 103 rejection for claim 12), wherein
Nimmagadda and Kim fail to teach but Musleh teaches the one or more conditions depend on the priority level for the given AI workload relative to a priority level for a non-AI workload. (Musleh, [0247] – “At 2552, a priority of the data can be set based on whether the data is AI-related or non-AI related and a type of AI-related data. For example, AI-related data can be prioritized over non-AI data. For example, AI-related data that is used for inference can be prioritized over AI-related data that is used for training or re-training. For example, AI-related data that is used for time sensitive inference can be prioritized over AI-related data that is used for non-time sensitive inference. In some examples, an application can specify whether data is AI-related or non-AI related and a type of AI-related data.”)
Nimmagadda, Kim, and Musleh are analogous art because they are in the same field of endeavor: inter-process communications. Therefore, it would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention to modify Nimmagadda and Kim with the teachings of Musleh. The motivation would be to schedule high priority workloads ahead of low priority workloads.
Claims 14-15 are rejected under 35 U.S.C. 103 as being unpatentable over Nimmagadda, Kim, and Musleh, further in view of Zaykov et al. (EP 4 250 108 A1); herein after “Zaykov”).
With respect to claim 14, Nimmagadda, Kim, and Musleh teach:
The apparatus of claim 13 (see 35 U.S.C. 103 rejection for claim 13), wherein
Nimmagadda, Kim, and Musleh fail to teach but Zaykov teaches in order to schedule the at least one AI workload (Zaykov, [0002] – “New safety-critical computing systems for aerospace are expected to host artificial intelligence (AI) or machine learning (ML) applications.”)
on at least one of the processing cores when the priority level for the given AI workload is less than the priority level for the non-AI workload (Zaykov, [0012] – “A safety-critical computing system has two types of workloads for the coprocessor 112-high-priority workloads and low-priority workloads. For high-priority workloads, the computing system 101 shall deliver a timing guarantee for performance of the workload, whereas low-priority workloads are best-effort and executed whenever the computing resources 102 and memory 104 in the computing system 101 are available.”),
the one or more processors are further configured to schedule the given AI workload on a selected one of the processing cores at a dynamic frequency and voltage operating point (Zaykov, [0060] – “Also, as shown in FIG. 6, the DVFS technique, which includes dynamically varying the operational frequency and voltage of the computing resources, provides better performance compared to the naive approach.”)
where the required bandwidth of the selected processing core for AI workload is less than the available bandwidth after scheduling the non-AI workload. (Zaykov, [0020] – “In some examples, the thermal management instructions 108 implement a thermal-aware scheduling policy, adjust an amount of available memory bandwidth to computing resources 102 and/or adjust an amount of memory utilization by one or more applications 118 executed by the computing resources 102 (memory throttling), and/or allocate cache 114 to workloads based on memory demands (cache allocation).”; Zaykov, [0023] – “In such examples, the available memory bandwidth can be utilized on a prioritized basis (for example, more memory bandwidth allocated to safety-critical applications compared to best-effort applications).”)
Nimmagadda, Kim, Musleh, and Zaykov are analogous art because they are in the same field of endeavor: inter-process communications. Therefore, it would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention to modify Nimmagadda, Kim, and Musleh with the teachings of Zaykov. The motivation is taught by Zaykov in paragraph [0010]: “By using one or more of these techniques, the computing system can operate closer to its maximum performance given the environmental conditions while providing consistent operation sufficient for safety-critical applications.”
With respect to claim 15, Nimmagadda, Kim, and Musleh teach:
The apparatus of claim 13 (see 35 U.S.C. 103 rejection for claim 13), wherein
in order to schedule the at least one AI workload on at least one of the processing cores when the priority level for the given AI workload is greater than the priority level for the non-AI workload (see 35 U.S.C. 103 rejection for claim 13), the one or more processors are further configured to:
Nimmagadda, Kim, and Musleh fail to teach but Zaykov teaches:
schedule the given AI workload on a selected one of the processing cores at a dynamic frequency and voltage operating point (Zaykov, [0060] – “Also, as shown in FIG. 6, the DVFS technique, which includes dynamically varying the operational frequency and voltage of the computing resources, provides better performance compared to the naive approach.”)
based on the computed metrics and one or more conditions (Zaykov, [0060] – “However, the DVFS technique still does not provide optimal performance due to the limitations described above. The hybrid and dynamic methods described herein enable the maximum available cooling capacity to be utilized by the computing system during all stages of flight by using the thermal-aware scheduling policy, memory throttling, and/or cache allocation techniques described above.”); and
allocate remaining bandwidth of the selected processing core for the non-AI workload. (Zaykov, [0054] – “The cache allocation 504, which is static in the example shown in FIG. 5, includes approximately 60% of the cache allocated for application 1, approximately 30% of the cache allocated for application 2, and approximately 10% of the cache allocated for application 3. This cache allocation 504 would be utilized, for example, where the workloads associated with application 1 are the most memory intensive and/or prioritized and the workloads associated with application 3 are the least memory intensive and/or prioritized. In other examples, the cache allocation 504 could be modified based on the operational parameter(s).”)
Nimmagadda, Kim, Musleh, and Zaykov are analogous art because they are in the same field of endeavor: inter-process communications. Therefore, it would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention to modify Nimmagadda, Kim, and Musleh with the teachings of Zaykov. The motivation is taught by Zaykov in paragraph [0010]: “By using one or more of these techniques, the computing system can operate closer to its maximum performance given the environmental conditions while providing consistent operation sufficient for safety-critical applications.”
Claim 17 is rejected under 35 U.S.C. 103 as being unpatentable over Nimmagadda in view of Piednoel (US 12,136,002 B1); herein after “Piednoel”).
With respect to claim 17, Nimmagadda teaches:
The apparatus of claim 16 (See claim 16), wherein:
the desired behavior relates to optimization of compute time, total power consumption, or a balance of compute time and power consumption (Nimmagadda, [0051] – “The compute-based subgraph partitioning technology described partitions the DL model graph herein helps to reduce or eliminate inefficiencies such as waiting, memory accesses and excessive power usage during execution of the deep learning model graph.”) and
Nimmagadda fails to teach but Piednoel teaches the desired behavior changes based on at least one of: the core junction temperature relative to a first threshold, or the battery capacity condition relative to a second threshold. (Piednoel, Column 21, Lines 1-7 – “In particular, these components may operate in a low power state in which the components are ready to take over the set of tasks being performed by the first SoC 410. The state information can include whether the components are operating within nominal temperatures and other nominal ranges (e.g., available bandwidth, power, memory, etc.).”)
Nimmagadda and Piednoel are analogous art because they are in the same field of endeavor: inter-process communications. Therefore, it would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention to modify Nimmagadda with the teachings of Piednoel. The motivation would be to protect the hardware components and performance when a threshold is met. This would be considered common knowledge in the art, for example, with most phones having overheat protection and low battery modes.
Claim 18 is rejected under 35 U.S.C. 103 as being unpatentable over Nimmagadda in view of Ramanujan et al. (US 2024/0419493 A1); herein after “Ramanujan”).
With respect to claim 18, Nimmagadda teaches:
The apparatus of claim 1 (see 35 U.S.C. 102 rejection for claim 1), wherein:
Nimmagadda fails to teach but Ramanujan teaches:
the second characteristics comprise at least one of a token input length or a token input characteristic for the at least one AI workload, (Ramanujan, [0014] – “In another example, each token length is unique and determined using various parameters, rules, and/or thresholds. As such, with distinct token lengths for various AI models, utilization process 10 processes 100 the workload data for a plurality of requests for the AI model to determine the utilization for each workload.”)
the one or more conditions relate to a desired behavior, and (Ramanujan, [0014] – “In another example, each token length is unique and determined using various parameters, rules, and/or thresholds. As such, with distinct token lengths for various AI models, utilization process 10 processes 100 the workload data for a plurality of requests for the AI model to determine the utilization for each workload.”)
the desired behavior changes based on the token input length relative to at least one threshold. (Ramanujan, [0014] – “In another example, each token length is unique and determined using various parameters, rules, and/or thresholds. As such, with distinct token lengths for various AI models, utilization process 10 processes 100 the workload data for a plurality of requests for the AI model to determine the utilization for each workload.”)
Nimmagadda and Ramanujan are analogous art because they are in the same field of endeavor: inter-process communications. Therefore, it would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention to modify Nimmagadda with the teachings of Ramanujan. The motivation is taught by Ramanujan in paragraph [0008]: “As such, implementations of the present disclosure provide more accurate determination of GPU utilization by measuring utilization in terms of tokens processed by the AI model.” By estimating an AI workload by its token length, resource pre-allocation can be far more accurate than estimating an AI workload by other means such as power consumption.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to LUKE WILLIAM CULLIP whose telephone number is (571)-270-5733. The examiner can normally be reached Monday - Friday (8:00am - 5:00pm ET).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kevin Young can be reached at (571)-270-3180. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/LUKE WILLIAM CULLIP/Examiner, Art Unit 2194 /KEVIN L YOUNG/Supervisory Patent Examiner, Art Unit 2194