DETAILED ACTION
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This Office Action is in response to claims filed on 07/18/2024.
Claims 1-20 are pending.
Claim Objections
Claims 14 and 16-20 are objected to because of the following informalities: In claim 14, “receiving the control signal” should read “receiving the first control signal”. In claim 16, “a response queue utilization signal from from the first controller” should read “a response queue utilization signal from . Appropriate correction is required.
Claims 17-20 depend, directly or indirectly, from objected claims and do not resolve the deficiencies thereof and are therefore objected to for at least the same reasons.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 5 and 16-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claims 5 and 20 recite the limitations “each flow control unit (FLIT) communicated from the first accelerator device to the host device” and “each FLIT communicated to the host device”. There is insufficient antecedent basis for these limitations in the claims. There is no prior mention of any flow control unit or FLIT communicated to the host device withing the claims 5 and 20 or within claims 1, 4, 16, and 19, from which the claims depend. From the current claim language, it is unclear whether any flow control unit is actually communicated to the host device because these claims do not include a positive recitation of this action.
Claim 16 recites the limitation "based on the at least one utilization signal" in line 10. There is insufficient antecedent basis for this limitation in the claim. There is no prior mention of at least one utilization signal within the claim. Claim 16 does state “to receive at least one of a request queue utilization signal and a response queue utilization signal”; however, from the current claim language it is not clear that the above limitation is referring to these same signals. For the sake of compact prosecution, Examiner will interpret the limitation to mean “based on the at least one of a request queue utilization signal and a response queue utilization signal”.
Claims 17-20 depend, directly or indirectly, from rejected claims and do not resolve the deficiencies thereof and are therefore rejected for at least the same reasons.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention recites a judicial exception, is directed to that judicial exception, an abstract idea, as it has not been integrated into practical application and the claims further do not recite significantly more than the judicial exception. Examiner has evaluated the claims under the framework provided in the 2019 Patent Eligibility Guidance published in the Federal Register 01/07/2019 and has provided such analysis below.
Step 1: Claims 1-15 are directed to a method and fall within the statutory category of process. Claims 16-20 are directed to a system and fall within the statutory category of machine. Therefore, “Are the claims to a process, machine, manufacture or composition of matter?” Yes.
In order to evaluate the Step 2A inquiry “Is the claim directed to a law of nature, a natural phenomenon or an abstract idea?” we must determine, at Step 2A Prong 1, whether the claim recites a law of nature, a natural phenomenon or an abstract idea and further whether the claim recites additional elements that integrate the judicial exception into a practical application.
Step 2A Prong 1:
Claims 1 and 16: The limitations of “determining a device loading metric for the first accelerator device based on the queue utilization signal;” and “and, based on the at least one utilization signal, provide a device utilization-indicating DevLoad signal”, as drafted, is a process that, but for the recitation of generic computing components, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, a person can observe a queue utilization signal and, based on this observation, can mentally determine a device loading metric for a first accelerator device. This may also be done with pencil and paper.
Therefore, Yes, claims 1 and 16 recites a judicial exception.
Step 2A Prong 2:
Claims 1 and 16: The judicial exception is not integrated into a practical application. In particular, the claim recites additional element recitations of “receiving, at a first accelerator device, commands from a host device;”, “at the first accelerator device: receiving a queue utilization signal indicative of a volume of transaction request messages received by, or response messages sent from, the first accelerator device based on the commands from the host device;”, “and providing a first control signal to the host device, wherein the first control signal includes information about the device loading metric for the first accelerator device”, and “provide a device utilization-indicating DevLoad signal to the host device via the first controller” which are merely recitations of data reception and transmission which is insignificant extra solution activity (see MPEP §2106.05(g)) which does not integrate a judicial exception into practical application. Further, the claims recite additional element recitations of “a host device coupled to multiple accelerator devices using an interconnect;”, “and a first accelerator device of the multiple accelerator devices, the first accelerator device including: a first controller configured to manage transactions with the host device via the interconnect;”, “a memory controller configured to manage transactions with a memory;”, “and a telemetry manager configured to receive at least one of a request queue utilization signal and a response queue utilization signal from from the first controller or from the memory controller”, which are merely recitations of generic computing components and technological environment/field of use (see MPEP § 2106.05(f) and 2106.05(h)) which does not integrate a judicial exception into practical application.
Therefore, “Do the claims recite additional elements that integrate the judicial exception into a practical application? No, these additional elements do not integrate the abstract idea into a practical application and they do not impose any meaningful limits on practicing the abstract idea. The claims are directed to an abstract idea.
After having evaluated the inquires set forth in Steps 2A Prong 1 and 2, it has been concluded that claims 1 and 16 not only recite a judicial exception but that the claims are directed to the judicial exception as the judicial exception has not been integrated into practical application.
Step 2B:
Claims 1 and 16: The claims do not include additional elements, alone or in combination, that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements amount to no more than insignificant extra solution activity, generic computing components, and technological environment/field of use which do not amount to significantly more than the abstract idea. Further, the insignificant extra solution activity is well-understood, routine, and conventional in the art. “The courts have recognized the following computer functions as well‐understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity. i. Receiving or transmitting data over a network…iv. Storing and retrieving information in memory” [MPEP§ 2106.05(d)(II)].
Therefore, “Do the claims recite additional elements that amount to significantly more than the judicial exception? No, these additional elements, alone or in combination, do not amount to significantly more than the judicial exception.
Having concluded analysis within the provided framework, Claim 1 does not recite patent eligible subject matter under 35 U.S.C. § 101.
With regard to claim 2, the claim recites additional abstract idea recitations of “at the host device, apportioning subsequent commands to the first accelerator device and to a second accelerator device based on the first control signal from the first accelerator device”, as drafted, is a process that, but for the recitation of generic computing components, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, person can observe a first control signal and, based on this observation, can mentally apportion subsequent commands to a first accelerator device and to a second accelerator device by mentally dividing and assigning the commands. This may also be done with pencil and paper. Further, claim 2, does not recite any further additional elements and for the same reasons as above with regard to integration into practical application and whether additional elements amount to significantly more, claim 2, also fails both Step 2A prong 2, thus the claims are directed to the judicial exception as it has not been integrated into practical application, and fails Step 2B as not amounting to significantly more. Therefore, Claim 2 does not recite patent eligible subject matter under 35 U.S.C. § 101.
With regard to claims 3 and 13, the claims recite additional abstract idea recitations of “and wherein apportioning the subsequent commands is based on the first and second control signals” and “and based on the control signals, selecting a particular one of the first and second accelerator devices to receive a subsequent command”, as drafted, is a process that, but for the recitation of generic computing components, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, person can observe a first and second control signal and, based on these observations, can mentally apportion subsequent commands to a first accelerator device and to a second accelerator device by mentally dividing and assigning the commands. This may also be done with pencil and paper. Further, the claims recite additional element recitations of “receiving a second control signal with information about a device loading metric for the second accelerator device;” and “at the host device: receiving the control signal from the first accelerator device and at least one other control signal from a second accelerator device;”, which are merely recitations of data reception which is insignificant extra solution activity (see MPEP §2106.05(g)) which does not integrate a judicial exception into practical application. Further, the insignificant extra solution activity is well-understood, routine, and conventional in the art. “The courts have recognized the following computer functions as well‐understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity. i. Receiving or transmitting data over a network…iv. Storing and retrieving information in memory” [MPEP§ 2106.05(d)(II)]. Further, claims 3 and 13, do not recite any further additional elements and for the same reasons as above with regard to integration into practical application and whether additional elements amount to significantly more, claims 3 and 13, also fail both Step 2A prong 2, thus the claims are directed to the judicial exception as it has not been integrated into practical application, and fail Step 2B as not amounting to significantly more. Therefore, Claims 3 and 13 do not recite patent eligible subject matter under 35 U.S.C. § 101.
With regard to claim 4, the claim recites additional element recitations of “wherein receiving the commands from the host device includes receiving the commands using a compute express link (CXL) interconnect,” and “and wherein providing the first control signal to the host device includes using the CXL interconnect”, which are merely recitations of technological environment/field of use (see MPEP § 2106.05(h)) which does not integrate a judicial exception into practical application. Further, claim 4, does not recite any further additional elements and for the same reasons as above with regard to integration into practical application and whether additional elements amount to significantly more, claim 4, also fails both Step 2A prong 2, thus the claims are directed to the judicial exception as it has not been integrated into practical application, and fails Step 2B as not amounting to significantly more. Therefore, Claim 4 does not recite patent eligible subject matter under 35 U.S.C. § 101.
With regard to claim 5, the claim recites additional element recitations of “wherein at least a portion of the first control signal is provided to the host device together with each flow control unit (FLIT) communicated from the first accelerator device to the host device”, which are merely recitations of data transmission which is insignificant extra solution activity (see MPEP §2106.05(g)) which does not integrate a judicial exception into practical application. Further, the insignificant extra solution activity is well-understood, routine, and conventional in the art. “The courts have recognized the following computer functions as well‐understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity. i. Receiving or transmitting data over a network…iv. Storing and retrieving information in memory” [MPEP§ 2106.05(d)(II)]. Further, claim 5, does not recite any further additional elements and for the same reasons as above with regard to integration into practical application and whether additional elements amount to significantly more, claim 5, also fails both Step 2A prong 2, thus the claims are directed to the judicial exception as it has not been integrated into practical application, and fails Step 2B as not amounting to significantly more. Therefore, Claim 5 does not recite patent eligible subject matter under 35 U.S.C. § 101.
With regard to claims 6, 7, 9, and 10, the claims recite additional element recitations of “wherein receiving the queue utilization signal includes receiving a CXL response queue utilization signal that indicates a volume of transactions queued for communication from the first accelerator device to the host using the CXL interconnect”, “wherein receiving the queue utilization signal includes receiving a CXL request queue utilization signal that indicates a quantity of transactions queued for further processing by compute resources of the first accelerator device”, “wherein receiving the queue utilization signal includes receiving a memory controller request queue utilization signal from a memory controller that comprises a portion of the first accelerator device”, and “wherein receiving the queue utilization signal includes receiving a memory controller response queue utilization signal from a memory controller that comprises a portion of the first accelerator device”, which are merely recitations of data reception which is insignificant extra solution activity (see MPEP §2106.05(g)) which does not integrate a judicial exception into practical application. Further, the insignificant extra solution activity is well-understood, routine, and conventional in the art. “The courts have recognized the following computer functions as well‐understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity. i. Receiving or transmitting data over a network…iv. Storing and retrieving information in memory” [MPEP§ 2106.05(d)(II)]. Further, claims 6, 7, 9, and 10, do not recite any further additional elements and for the same reasons as above with regard to integration into practical application and whether additional elements amount to significantly more, claims 6, 7, 9, and 10, also fail both Step 2A prong 2, thus the claims are directed to the judicial exception as it has not been integrated into practical application, and fail Step 2B as not amounting to significantly more. Therefore, Claims 6, 7, 9, and 10 do not recite patent eligible subject matter under 35 U.S.C. § 101.
With regard to claim 8, the claim recites additional element recitations of “wherein the CXL request queue utilization signal includes information about a utilization of a cache controller on the first accelerator device”, which are merely recitations of technological environment/field of use (see MPEP § 2106.05(h)) which does not integrate a judicial exception into practical application. Further, claim 8, does not recite any further additional elements and for the same reasons as above with regard to integration into practical application and whether additional elements amount to significantly more, claim 8, also fails both Step 2A prong 2, thus the claims are directed to the judicial exception as it has not been integrated into practical application, and fails Step 2B as not amounting to significantly more. Therefore, Claim 8 does not recite patent eligible subject matter under 35 U.S.C. § 101.
With regard to claim 11, the claim recites additional abstract idea recitations of “determining a read/write ratio for transactions processed by the first accelerator device, the ratio based on a number of data response (DRS) messages and a number of no data response (NDR) messages queued for communication from the first accelerator device to the host device; and wherein determining the device loading metric includes using the determined read/write ratio”, as drafted, is a process that, but for the recitation of generic computing components, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, person can observe a number of data response (DRS) messages and a number of no data response (NDR) messages queued for communication and, based on these observations, can mentally determine a read/write ratio for transactions processed by the first accelerator device by using mental calculation. Further, a person can mentally determine a device loading metric based on this read/write ratio. This may also be done with pencil and paper. Further, claim 11, does not recite any further additional elements and for the same reasons as above with regard to integration into practical application and whether additional elements amount to significantly more, claim 11, also fails both Step 2A prong 2, thus the claims are directed to the judicial exception as it has not been integrated into practical application, and fails Step 2B as not amounting to significantly more. Therefore, Claim 11 does not recite patent eligible subject matter under 35 U.S.C. § 101.
With regard to claim 12, the claim recites additional abstract idea recitations of “and determining the device loading metric about the first accelerator device based on the thermal status signal and the queue utilization signal”, as drafted, is a process that, but for the recitation of generic computing components, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, person can observe a thermal status signal and a queue utilization signal and, based on these observations, can mentally determine a device loading metric about the first accelerator device. This may also be done with pencil and paper. Further, the claim recites additional element recitations of “at a telemetry manager of the first accelerator device, receiving a thermal status signal indicative of a temperature of a portion of the first accelerator device,”, which are merely recitations of data reception which is insignificant extra solution activity (see MPEP §2106.05(g)) which does not integrate a judicial exception into practical application. Further, the insignificant extra solution activity is well-understood, routine, and conventional in the art. “The courts have recognized the following computer functions as well‐understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity. i. Receiving or transmitting data over a network…iv. Storing and retrieving information in memory” [MPEP§ 2106.05(d)(II)]. Further, claim 12, does not recite any further additional elements and for the same reasons as above with regard to integration into practical application and whether additional elements amount to significantly more, claim 12, also fails both Step 2A prong 2, thus the claims are directed to the judicial exception as it has not been integrated into practical application, and fails Step 2B as not amounting to significantly more. Therefore, Claim 12 does not recite patent eligible subject matter under 35 U.S.C. § 101.
With regard to claim 14, the claim recites additional abstract idea recitations of “based on the control signal, classifying the first accelerator device as underutilized, overutilized, or optimally utilized;”, as drafted, is a process that, but for the recitation of generic computing components, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, person can observe a control signal and, based on these observations, can mentally classify the first accelerator device as underutilized, overutilized, or optimally utilized. This may also be done with pencil and paper. Further, the claim recites additional abstract idea recitations of “and selecting the first accelerator device or a different accelerator device coupled to the host device to perform a subsequent command based on the classification of the first accelerator device”, as drafted, is a process that, but for the recitation of generic computing components, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, person can observe a classification of the first accelerator device and, based on these observations, can mentally select the first accelerator device or a different accelerator device coupled to the host device to perform a subsequent command. Further, the claim recites additional element recitations of “at the host device: receiving the control signal from the first accelerator device;”, which are merely recitations of data reception which is insignificant extra solution activity (see MPEP §2106.05(g)) which does not integrate a judicial exception into practical application. Further, the insignificant extra solution activity is well-understood, routine, and conventional in the art. “The courts have recognized the following computer functions as well‐understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity. i. Receiving or transmitting data over a network…iv. Storing and retrieving information in memory” [MPEP§ 2106.05(d)(II)]. Further, claim 14, does not recite any further additional elements and for the same reasons as above with regard to integration into practical application and whether additional elements amount to significantly more, claim 14, also fails both Step 2A prong 2, thus the claims are directed to the judicial exception as it has not been integrated into practical application, and fails Step 2B as not amounting to significantly more. Therefore, Claim 14 does not recite patent eligible subject matter under 35 U.S.C. § 101.
With regard to claim 15, the claim recites additional abstract idea recitations of “wherein classifying the first accelerator device includes using information from the control signal about one or more of a request path loading status, a response path loading status, a read/write transaction ratio for the first accelerator device, and a thermal status for the first accelerator device”, as drafted, is a process that, but for the recitation of generic computing components, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, person can observe information from the control signal about one or more of a request path loading status, a response path loading status, a read/write transaction ratio for the first accelerator device, and a thermal status for the first accelerator device and, based on these observations, can mentally classify the first accelerator device as underutilized, overutilized, or optimally utilized. This may also be done with pencil and paper. Further, claim 15, does not recite any further additional elements and for the same reasons as above with regard to integration into practical application and whether additional elements amount to significantly more, claim 15, also fails both Step 2A prong 2, thus the claims are directed to the judicial exception as it has not been integrated into practical application, and fails Step 2B as not amounting to significantly more. Therefore, Claim 15 does not recite patent eligible subject matter under 35 U.S.C. § 101.
With regard to claim 17, the claim recites additional element recitations of “wherein the first accelerator device includes a cache controller coupled to a cache memory, and wherein the telemetry manager is configured to provide the device utilization-indicating signal based on information about a utilization of the cache controller”, which are merely recitations of technological environment/field of use (see MPEP § 2106.05(h)) which does not integrate a judicial exception into practical application. Further, claim 17, does not recite any further additional elements and for the same reasons as above with regard to integration into practical application and whether additional elements amount to significantly more, claim 17, also fails both Step 2A prong 2, thus the claims are directed to the judicial exception as it has not been integrated into practical application, and fails Step 2B as not amounting to significantly more. Therefore, Claim 17 does not recite patent eligible subject matter under 35 U.S.C. § 101.
With regard to claim 18, the claim recites additional element recitations of “wherein the first accelerator device includes a thermal manager configured to receive temperature information about at least a portion of the first accelerator device, and wherein the telemetry manager is configured to provide information about a temperature of the first accelerator device in the DevLoad signal”, which are merely recitations of technological environment/field of use (see MPEP § 2106.05(h)) which does not integrate a judicial exception into practical application. Further, claim 18, does not recite any further additional elements and for the same reasons as above with regard to integration into practical application and whether additional elements amount to significantly more, claim 18, also fails both Step 2A prong 2, thus the claims are directed to the judicial exception as it has not been integrated into practical application, and fails Step 2B as not amounting to significantly more. Therefore, Claim 18 does not recite patent eligible subject matter under 35 U.S.C. § 101.
With regard to claim 19, the claim recites additional element recitations of “wherein the host device is coupled to the multiple accelerator devices using a compute express link (CXL) interconnect”, which are merely recitations of technological environment/field of use (see MPEP § 2106.05(h)) which does not integrate a judicial exception into practical application. Further, claim 19, does not recite any further additional elements and for the same reasons as above with regard to integration into practical application and whether additional elements amount to significantly more, claim 19, also fails both Step 2A prong 2, thus the claims are directed to the judicial exception as it has not been integrated into practical application, and fails Step 2B as not amounting to significantly more. Therefore, Claim 19 does not recite patent eligible subject matter under 35 U.S.C. § 101.
With regard to claim 20, the claim recites additional element recitations of “wherein the first controller is configured to include information about the DevLoad signal in each FLIT communicated to the host device using the interconnect”, which are merely recitations of technological environment/field of use (see MPEP § 2106.05(h)) which does not integrate a judicial exception into practical application. Further, claim 20, does not recite any further additional elements and for the same reasons as above with regard to integration into practical application and whether additional elements amount to significantly more, claim 20, also fails both Step 2A prong 2, thus the claims are directed to the judicial exception as it has not been integrated into practical application, and fails Step 2B as not amounting to significantly more. Therefore, Claim 20 does not recite patent eligible subject matter under 35 U.S.C. § 101.
Therefore, Claims 1-20 do not recite patent eligible subject matter under U.S.C. §101.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1 and 2 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Ferreira (US 2022/0229695 A1).
With regard to claim 1, Ferreira teaches:
A method comprising: receiving, at a first accelerator device, commands from a host device; “Many data scientists desire to run their GPU-based inference tasks in containers in such an environment. To increase GPU utilization, there has been a desire to schedule some of these concurrent tasks on the same GPU device, effectively sharing the GPU across different containers or pods” [Ferreira ¶ 5]. “This fine grain allocation may for example comprise creating a set of queues for each resource and the assigning tasks to the queues. The fine grain scheduler may assign tasks to queues based on the requirements of the task (e.g., task benefiting from GPU would be assigned to a queue with one or more GPUs allocated to the core)” [Ferreira ¶ 46]. “Instead, fine grain scheduler 340 communicates with coarse scheduler 370 and performs sub-task scheduling by allocating fine grain scheduled processes 420 to computing system nodes 450 directly. Sub-tasks assigned to each node are then scheduled and executed on the node's processor(s) by the node's scheduler 392” [Ferreira ¶ 34].
at the first accelerator device: receiving a queue utilization signal indicative of a volume of transaction request messages received by, or response messages sent from, the first accelerator device based on the commands from the host device; “Each container may execute multiple fine grain processes 344, e.g., using queues that are allocated resources by the fine grain schedulers 340. For example, one fine grain scheduler 340 may create four queues and schedule multiple fine grain tasks to each queue” [Ferreira ¶ 32]. “In one embodiment, the method may comprise allocating coarse blocks of resources to portions of an application that is to be executed, allocating tasks to queues, assigning nodes to the queues with tasks, and monitoring the queue length (queue utilization signal) and resource utilization of the allocated nodes” [Ferreira ¶ 13].
determining a device loading metric for the first accelerator device based on the queue utilization signal; “In one embodiment, the method may comprise allocating coarse blocks of resources to portions of an application that is to be executed, allocating tasks to queues, assigning nodes to the queues with tasks, and monitoring the queue length and resource utilization of the allocated nodes. If the queue length or allocated resource utilization are above a first predetermined threshold, additional resources from the coarse-grained scheduler may be allocated” [Ferreira ¶ 13].
and providing a first control signal to the host device, wherein the first control signal includes information about the device loading metric for the first accelerator device. “In some embodiments, the coarse scheduler may allocate coarse blocks of resources to portions of an application that is to be executed, and the fine grain scheduler may allocate tasks to queues and assign nodes to the queues with tasks. The monitoring may for example include tracking the queue length and resource utilization of the allocated nodes. If the queue length or allocated resource utilization percentage is outside a desired threshold (e.g., either predetermined or as informed by a prediction engine based on historical training data), the fine grain scheduler may request (first control signal) additional resources from the coarse scheduler or release resources back (so they are available again to the coarse scheduler)” [Ferreira ¶ 38].
With regard to claim 2, Ferreira teaches:
The method of claim 1, as referenced above.
comprising, at the host device, apportioning subsequent commands to the first accelerator device and to a second accelerator device based on the first control signal from the first accelerator device. “Performance of the system may be monitored (step 620). For example, queue depth, wait times, and utilization rates may be monitored. If the performance determined to be outside a desired range or is predicted to be outside the desired range in the near future (step 630), the fine grain scheduler may be configured to determine if additional resources are available (e.g., resources that have been allocated at the coarse level to the pod or container). If additional resources are available (step 660), the fine grain scheduler may allocate those (step 670). For example, if a particular queue has many tasks queued up, additional CPUs/GPUs may be allocated the to queue, or a new queue may be created and allocated additional CPUs/GPUs” [Ferreira ¶ 47].
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 3, 9-10, and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Ferreira (US 2022/0229695 A1) in view of Wong (US 2023/0102063 A1).
With regard to claim 3, Ferreira teaches the method of claim 2, as referenced above. Ferreira further teaches and wherein apportioning the subsequent commands is based on the first and second control signals. “If the performance determined to be outside a desired range or is predicted to be outside the desired range in the near future (step 630), the fine grain scheduler may be configured to determine if additional resources are available (e.g., resources that have been allocated at the coarse level to the pod or container). If additional resources are available (step 660), the fine grain scheduler may allocate those (step 670). For example, if a particular queue has many tasks queued up, additional CPUs/GPUs may be allocated the to queue, or a new queue may be created and allocated additional CPUs/GPUs. If additional allocated resources are not available, the fine grain scheduler may be configured to request additional (coarsely allocated) resources from the coarse scheduler (step 680)” [Ferreira ¶ 47].
Ferreira fails to explicitly teach comprising, at the host device, receiving a second control signal with information about a device loading metric for the second accelerator device; and wherein apportioning the subsequent commands is based on the first and second control signals.
However, Wong teaches:
comprising, at the host device, receiving a second control signal with information about a device loading metric for the second accelerator device; “The method also includes inspecting runtime utilization metrics of a plurality of processing resources based on the workload description, where the plurality of processing resources includes at least a first GPU and a second GPU” [Wong ¶ 13].
and wherein apportioning the subsequent commands is based on the first and second control signals. “The method also includes determining, based on the utilization metrics and one or more policies, a workload allocation recommendation” [Wong ¶ 13].
Wong is considered to be analogous to the claimed invention because it is in the same field of considering the load. Therefore, it would be obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ferreira to incorporate the teachings of Wong and include comprising, at the host device, receiving a second control signal with information about a device loading metric for the second accelerator device; and wherein apportioning the subsequent commands is based on the first and second control signals. Doing so would allow for the assignment of commands to take into account metrics across accelerator devices. “Memory subsystem utilization can be expressed by metrics such as the number of read packets and the number of write packets issued over the interface within a current time period, the current utilization of ingress and egress queues or buffers, data transfer times and latency, and so on” [Wong ¶ 36].
With regard to claim 9, Ferreira teaches the method of claim 1, as referenced above. Ferreira further teaches wherein receiving the queue utilization signal includes receiving a memory controller request queue utilization signal “Each container may execute multiple fine grain processes 344, e.g., using queues that are allocated resources by the fine grain schedulers 340. For example, one fine grain scheduler 340 may create four queues and schedule multiple fine grain tasks to each queue” [Ferreira ¶ 32]. “In one embodiment, the method may comprise allocating coarse blocks of resources to portions of an application that is to be executed, allocating tasks to queues, assigning nodes to the queues with tasks, and monitoring the queue length (queue utilization signal) and resource utilization of the allocated nodes” [Ferreira ¶ 13].
Ferreira fails to teach a memory controller request queue utilization signal from a memory controller that comprises a portion of the first accelerator device.
However, Wong teaches a memory controller request queue utilization signal from a memory controller that comprises a portion of the first accelerator device. “The discrete GPU 134 also includes memory controllers 144 and DMA engines 148 for accessing graphics memory 180. In some examples, the memory controllers 144 and DMA engines 148 are configured to access a shared portion of system memory 160” [Wong ¶ 27]. “In some examples, inspecting 220 runtime utilization metrics of a plurality of processing resources can also include collecting values of runtime utilization metrics from additional processing resources including multimedia accelerators such video codecs and audio codecs, display controllers, security processors, memory subsystems such as DMA engines and memory controllers, and bus interfaces such as a PCIe interface… Memory subsystem utilization can be expressed by metrics such as the number of read packets and the number of write packets issued over the interface within a current time period, the current utilization of ingress (request queue utilization) and egress queues or buffers, data transfer times and latency, and so on” [Wong ¶ 36].
Wong is considered to be analogous to the claimed invention because it is in the same field of considering the load. Therefore, it would be obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ferreira to incorporate the teachings of Wong and include a memory controller request queue utilization signal from a memory controller that comprises a portion of the first accelerator device. Doing so would allow for the assignment of commands to take into account metrics across accelerator devices. “Memory subsystem utilization can be expressed by metrics such as the number of read packets and the number of write packets issued over the interface within a current time period, the current utilization of ingress and egress queues or buffers, data transfer times and latency, and so on” [Wong ¶ 36].
With regard to claim 10, Ferreira teaches the method of claim 1, as referenced above. Ferreira further teaches wherein receiving the queue utilization signal includes receiving a memory controller response queue utilization signal “Each container may execute multiple fine grain processes 344, e.g., using queues that are allocated resources by the fine grain schedulers 340. For example, one fine grain scheduler 340 may create four queues and schedule multiple fine grain tasks to each queue” [Ferreira ¶ 32]. “In one embodiment, the method may comprise allocating coarse blocks of resources to portions of an application that is to be executed, allocating tasks to queues, assigning nodes to the queues with tasks, and monitoring the queue length (queue utilization signal) and resource utilization of the allocated nodes” [Ferreira ¶ 13].
Ferreira fails to teach a memory controller response queue utilization signal from a memory controller that comprises a portion of the first accelerator device.
However, Wong teaches a memory controller response queue utilization signal from a memory controller that comprises a portion of the first accelerator device. “The discrete GPU 134 also includes memory controllers 144 and DMA engines 148 for accessing graphics memory 180. In some examples, the memory controllers 144 and DMA engines 148 are configured to access a shared portion of system memory 160” [Wong ¶ 27]. “In some examples, inspecting 220 runtime utilization metrics of a plurality of processing resources can also include collecting values of runtime utilization metrics from additional processing resources including multimedia accelerators such video codecs and audio codecs, display controllers, security processors, memory subsystems such as DMA engines and memory controllers, and bus interfaces such as a PCIe interface… Memory subsystem utilization can be expressed by metrics such as the number of read packets and the number of write packets issued over the interface within a current time period, the current utilization of ingress and egress (response queue utilization) queues or buffers, data transfer times and latency, and so on” [Wong ¶ 36].
Wong is considered to be analogous to the claimed invention because it is in the same field of considering the load. Therefore, it would be obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ferreira to incorporate the teachings of Wong and include a memory controller response queue utilization signal from a memory controller that comprises a portion of the first accelerator device. Doing so would allow for the assignment of commands to take into account metrics across accelerator devices. “Memory subsystem utilization can be expressed by metrics such as the number of read packets and the number of write packets issued over the interface within a current time period, the current utilization of ingress and egress queues or buffers, data transfer times and latency, and so on” [Wong ¶ 36].
With regard to claim 13, Ferreira teaches the method of claim 1, as referenced above. Ferreira fails to explicitly teach further comprising, at the host device: receiving the control signal from the first accelerator device and at least one other control signal from a second accelerator device; and based on the control signals, selecting a particular one of the first and second accelerator devices to receive a subsequent command.
However, Wong teaches:
further comprising, at the host device: receiving the control signal from the first accelerator device and at least one other control signal from a second accelerator device; “The method also includes inspecting runtime utilization metrics of a plurality of processing resources based on the workload description, where the plurality of processing resources includes at least a first GPU and a second GPU” [Wong ¶ 13].
and based on the control signals, selecting a particular one of the first and second accelerator devices to receive a subsequent command. “The method also includes determining, based on the utilization metrics and one or more policies, a workload allocation recommendation” [Wong ¶ 13]. “Based on the operating system's response, the application can select a GPU and assign the workload to that GPU. For example, the application can assign the workload to the integrated GPU because the integrated GPU typically consumes less power than the discrete GPU” [Wong ¶ 10].
Wong is considered to be analogous to the claimed invention because it is in the same field of considering the load. Therefore, it would be obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ferreira to incorporate the teachings of Wong and include further comprising, at the host device: receiving the control signal from the first accelerator device and at least one other control signal from a second accelerator device; and based on the control signals, selecting a particular one of the first and second accelerator devices to receive a subsequent command. Doing so would allow for the assignment of commands to take into account metrics across accelerator devices. “Memory subsystem utilization can be expressed by metrics such as the number of read packets and the number of write packets issued over the interface within a current time period, the current utilization of ingress and egress queues or buffers, data transfer times and latency, and so on” [Wong ¶ 36].
Claims 4 and 7 are rejected under 35 U.S.C. 103 as being unpatentable over Ferreira (US 2022/0229695 A1) in view of Choe (US 2022/0100669 A1).
With regard to claim 4, Ferreira teaches the method of claim 1, as referenced above. Ferreira fails to teach wherein receiving the commands from the host device includes receiving the commands using a compute express link (CXL) interconnect, and wherein providing the first control signal to the host device includes using the CXL interconnect.
However, Choe teaches:
wherein receiving the commands from the host device includes receiving the commands using a compute express link (CXL) interconnect, and wherein providing the first control signal to the host device includes using the CXL interconnect. “The CXL interface is a computer device interconnector standard, and is an interface that may reduce the overhead and waiting time of the host device and the smart storage device 1000 and may allow the storage space of the host memory and the memory device to be shared in a heterogeneous computing environment in which the host device 10 and the smart storage device 1000 operate together. For example, the host device 10 and the system-on-chip, GPU, which performs complex computations, and an acceleration module, such as a field-programmable gate array (FPGA), directly communicate and share memory. The smart storage device 1000 of the present specification is based on the CXL standard” [Choe ¶ 25].
Choe is considered to be analogous to the claimed invention because it is in the same field of PCI express. Therefore, it would be obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ferreira to incorporate the teachings of Choe and include wherein receiving the commands from the host device includes receiving the commands using a compute express link (CXL) interconnect, and wherein providing the first control signal to the host device includes using the CXL interconnect. Doing so would allow for further efficiency in shared memory. “The CXL interface is a computer device interconnector standard, and is an interface that may reduce the overhead and waiting time of the host device and the smart storage device 1000 and may allow the storage space of the host memory and the memory device to be shared in a heterogeneous computing environment in which the host device 10 and the smart storage device 1000 operate together” [Choe ¶ 25].
With regard to claim 7, Ferreira in view of Choe teaches the method of claim 4, as referenced above. Ferreira further teaches wherein receiving the queue utilization signal includes receiving a CXL request queue utilization signal that indicates a quantity of transactions queued for further processing by compute resources of the first accelerator device. “Each container may execute multiple fine grain processes 344, e.g., using queues that are allocated resources by the fine grain schedulers 340. For example, one fine grain scheduler 340 may create four queues and schedule multiple fine grain tasks to each queue” [Ferreira ¶ 32]. “In one embodiment, the method may comprise allocating coarse blocks of resources to portions of an application that is to be executed, allocating tasks to queues, assigning nodes to the queues with tasks, and monitoring the queue length (queue utilization signal) and resource utilization of the allocated nodes” [Ferreira ¶ 13].
Ferreira fails to teach a CXL request.
However, Choe teaches a CXL request “The cards/blades/systems are reachable to the CPUs/servers that use the memory resources through some kind of network infrastructure such as CXL, CAPI, etc” [Macnamara ¶ 72].
Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Ferreira (US 2022/0229695 A1) in view of Choe (US 2022/0100669 A1) in view of Ibrahim (US 2020/0293445 A1).
With regard to claim 5, Ferreira in view of Choe teaches the method of claim 4, as referenced above. Ferreira in view of Choe fails to teach wherein at least a portion of the first control signal is provided to the host device together with each flow control unit (FLIT) communicated from the first accelerator device to the host device.
However, Ibrahim teaches wherein at least a portion of the first control signal is provided to the host device together with each flow control unit (FLIT) communicated from the first accelerator device to the host device. “The exchange of information can be done in an opportunistic way. In other words, in various embodiments, a CU transmits the collected local information to another CU as a separate one-flit packet or piggybacks the collected information on an outgoing request/reply to another CU” [Ibrahim ¶ 65].
Ibrahim is considered to be analogous to the claimed invention because it is in the same field of information transfer on a bus. Therefore, it would be obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ferreira in view of Choe to incorporate the teachings of Ibrahim and include wherein at least a portion of the first control signal is provided to the host device together with each flow control unit (FLIT) communicated from the first accelerator device to the host device. Doing so would allow for further efficiency in the sharing of data. “The exchange of information can be done in an opportunistic way” [Ibrahim ¶ 65].
Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Ferreira (US 2022/0229695 A1) in view of Choe (US 2022/0100669 A1) in view of Rahman (US 2023/0403233 A1).
With regard to claim 6, Ferreira in view of Choe teaches the method of claim 4, as referenced above. Ferreira in view of Choe fails to teach wherein receiving the queue utilization signal includes receiving a CXL response queue utilization signal that indicates a volume of transactions queued for communication from the first accelerator device to the host using the CXL interconnect.
However, Rahman teaches wherein receiving the queue utilization signal includes receiving a CXL response queue utilization signal that indicates a volume of transactions queued for communication from the first accelerator device to the host using the CXL interconnect. “Packet processing device 400 can be coupled to one or more servers using a bus, PCIe, CXL, or Double Data Rate (DDR)” [Rahman ¶ 52]. “Some examples of packet processing device 400 are part of an Infrastructure Processing Unit (IPU) or data processing unit (DPU) or utilized by an IPU or DPU. An xPU can refer at least to an IPU, DPU, GPU, GPGPU, or other processing units (e.g., accelerator devices)” [Rahman ¶ 53]. “Various examples described herein can provide congestion telemetry data for use in CC based on queue depth information that considers one or more other queues that utilize a same egress or output port and, potentially, based on arbiter configurations that allocate egress bandwidth from an egress port according to even or uneven weighting to different queues and associated traffic classes. According to various examples, a network interface device can calculate a utilization (U) of a traffic class and send congestion telemetry data, including the U for the traffic class, to a traffic sender network interface device as feedback. For example, the switch can determine per-queue utilization based on one or more of: outgoing or egress queue length (qleni) (response queue utilization signal) of class i (where i is an integer and different values of i are assigned to different traffic classes), transmission rate (txRatei) for class i, or the egress bandwidth or line or port egress rate (B)” [Rahman ¶ 13].
Rahman is considered to be analogous to the claimed invention because it is in the same field of information transfer on a bus. Therefore, it would be obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ferreira in view of Choe to incorporate the teachings of Rahman and include wherein receiving the queue utilization signal includes receiving a CXL response queue utilization signal that indicates a volume of transactions queued for communication from the first accelerator device to the host using the CXL interconnect. Doing so would allow for congestion control taking into account the egress rate of packets. “For example, utilization of HPCC involves calculating congestion based on queueing in switch 100 and comparing a queue's draining rate (e.g., rate of packet egress from the queue to an egress port) against a target egress port bandwidth (BW)” [Rahman ¶ 11].
Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Ferreira (US 2022/0229695 A1) in view of Choe (US 2022/0100669 A1) in view of Wong (US 2023/0102063 A1) in view of Puranik (US 2023/0052808 A1).
With regard to claim 8, Ferreira in view of Choe teaches the method of claim 7, as referenced above. Ferreira in view of Choe fails to teach wherein the CXL request queue utilization signal includes information about a utilization of a cache controller on the first accelerator device.
However, Wong teaches wherein the CXL request queue utilization signal includes information about a utilization of a (memory) cache controller on the first accelerator device. “In some examples, inspecting 220 runtime utilization metrics of a plurality of processing resources can also include collecting values of runtime utilization metrics from additional processing resources including multimedia accelerators such video codecs and audio codecs, display controllers, security processors, memory subsystems such as DMA engines and memory controllers, and bus interfaces such as a PCIe interface… Memory subsystem utilization can be expressed by metrics such as the number of read packets and the number of write packets issued over the interface within a current time period, the current utilization of ingress and egress queues or buffers, data transfer times and latency, and so on” [Wong ¶ 36].
Wong is considered to be analogous to the claimed invention because it is in the same field of considering the load. Therefore, it would be obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ferreira in view of Choe in view of Macnamara to incorporate the teachings of Wong and include wherein the CXL request queue utilization signal includes information about a utilization of a (memory) cache controller on the first accelerator device. Doing so would allow for the assignment of commands to take into account metrics across accelerator devices. “Memory subsystem utilization can be expressed by metrics such as the number of read packets and the number of write packets issued over the interface within a current time period, the current utilization of ingress and egress queues or buffers, data transfer times and latency, and so on” [Wong ¶ 36].
Ferreira in view of Choe in view of Wong fails to explicitly teach a cache controller on the first accelerator device.
However, Puranik teaches a cache controller on the first accelerator device. “The accelerator 100 can include an accelerator cache 105 managed by a cache controller 110. The accelerator cache 105 can be any of a variety of different types of memory, e.g., an L1 cache, and be of any of a variety of different sizes, e.g., 8 kilobytes to 64 kilobytes. The accelerator cache 105 is local to the accelerator 100. A cache controller (not shown) can read from and write to contents of the accelerator cache 105, which may be updated many times throughout the execution of operations by the accelerator 100” [Puranik ¶ 47].
Puranik is considered to be analogous to the claimed invention because it is in the same field of allocation of resources. Therefore, it would be obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ferreira in view of Choe in view of Wong to incorporate the teachings of Puranik and include a cache controller on the first accelerator device. Doing so would allow for further communication efficiency. “Devices can communicate less data on input/output channels, and more data on memory and cache channels that are more efficient for data transmission” [Puranik Abstract].
Claim 11 is rejected under 35 U.S.C. 103 as being unpatentable over Ferreira (US 2022/0229695 A1) in view of Wong (US 2023/0102063 A1) in view of HSU (US 2022/0164118 A1).
With regard to claim 11, Ferreira teaches the method of claim 1, as referenced above. Ferreira fails to teach the ratio based on a number of data response (DRS) messages and a number of no data response (NDR) messages queued for communication from the first accelerator device to the host device.
However, Wong teaches the ratio based on a number of data response (DRS) messages and a number of no data response (NDR) messages queued for communication from the first accelerator device to the host device; “The discrete GPU 134 also includes memory controllers 144 and DMA engines 148 for accessing graphics memory 180. In some examples, the memory controllers 144 and DMA engines 148 are configured to access a shared portion of system memory 160” [Wong ¶ 27]. “Memory subsystem utilization can be expressed by metrics such as the number of read packets and the number of write packets issued over the interface within a current time period, the current utilization of ingress and egress queues or buffers, data transfer times and latency, and so on” [Wong ¶ 36].
Wong is considered to be analogous to the claimed invention because it is in the same field of considering the load. Therefore, it would be obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ferreira to incorporate the teachings of Wong and include the ratio based on a number of data response (DRS) messages and a number of no data response (NDR) messages queued for communication from the first accelerator device to the host device. Doing so would allow for the management of workloads to take into account utilization of memory subsystems that are specifically relevant to the workload. “Then, the resource manager queries the respective drivers of the plurality of processing resources to for utilization metrics to construct a utilization state of the computing device as it pertains to the workload that will potentially be allocated on those processing resources” [Wong ¶ 36].
Ferreira in view of Wong fails to teach comprising determining a read/write ratio for transactions processed by the first accelerator device, and wherein determining the device loading metric includes using the determined read/write ratio.
However, HSU teaches:
comprising determining a read/write ratio for transactions processed by the first accelerator device, “In one or more embodiments, the plurality of memory hotness metrics includes a ratio of read instances to write instances” [HSU ¶ 181].
and wherein determining the device loading metric includes using the determined read/write ratio. “In one or more embodiments, the hotness evaluator 514 may consider the types of access instances (e.g., reads, writes) in determining an associated hotness metric. For example, a hotness metric may indicate a ratio of read operations versus write operations to provide a characterization as to how the memory segment is being accessed” [HSU ¶ 123].
HSU is considered to be analogous to the claimed invention because it is in the same field of PCI express. Therefore, it would be obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ferreira in view of Wong to incorporate the teachings of HSU and include comprising determining a read/write ratio for transactions processed by the first accelerator device, and wherein determining the device loading metric includes using the determined read/write ratio. Doing so would allow for the management of workloads to take into account a characterization of memory access. “For example, a hotness metric may indicate a ratio of read operations versus write operations to provide a characterization as to how the memory segment is being accessed” [HSU ¶ 123].
Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Ferreira (US 2022/0229695 A1) in view of Ippatapu (US 2020/0356292 A1).
With regard to claim 12, Ferreira teaches the method of claim 1, as referenced above. Ferreira fails to teach further comprising, at a telemetry manager of the first accelerator device, receiving a thermal status signal indicative of a temperature of a portion of the first accelerator device, and determining the device loading metric about the first accelerator device based on the thermal status signal and the queue utilization signal.
However, Ippatapu teaches:
further comprising, at a telemetry manager of the first accelerator device, receiving a thermal status signal indicative of a temperature of a portion of the first accelerator device, “Method 500 typically begins at block 510 where the data storage system monitors the performance of network, devices, and/or resources and collects performance data. The performance data may be associated with usage of computing resources such as CPU utilization, graphics process unit (GPU) utilization, storage device utilization, a random-access memory (RAM) utilization, and other resource utilization data over a particular period of time” [Ippatapu ¶ 82]. “The GPU utilization may include GPU process unit utilization, graphics memory utilization, GPU temperature, GPU fan speed, and other graphics utilization metrics” [Ippatapu ¶ 82].
and determining the device loading metric about the first accelerator device based on the thermal status signal and the queue utilization signal. “The performance data may be used to determine the performance metrics described and used in connection with the techniques herein. The performance data may be used in determining a workload for one or more physical devices, logical devices or volumes (LVs), thin devices or portions thereof, and the like. The workload may also be a measurement or level of "how busy" a device, or portion thereof is for example, in terms of I/O operations such as I/O throughput such as a number of I/O per second, and the like” [Ippatapu ¶ 83].
Ippatapu is considered to be analogous to the claimed invention because it is in the same field of considering the load. Therefore, it would be obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ferreira to incorporate the teachings of Ippatapu and include further comprising, at a telemetry manager of the first accelerator device, receiving a thermal status signal indicative of a temperature of a portion of the first accelerator device, and determining the device loading metric about the first accelerator device based on the thermal status signal and the queue utilization signal. Doing so would allow for the assignment of commands to take into account the temperature of the accelerator devices. “The performance data may be used to determine the performance metrics described and used in connection with the techniques herein” [Ippatapu ¶ 83].
Claim 14 is rejected under 35 U.S.C. 103 as being unpatentable over Ferreira (US 2022/0229695 A1) in view of Marchetti (US 2018/0241812 A1).
With regard to claim 14, Ferreira teaches the method of claim 1, as referenced above. Ferreira further teaches:
further comprising, at the host device: receiving the control signal from the first accelerator device; “In some embodiments, the coarse scheduler may allocate coarse blocks of resources to portions of an application that is to be executed, and the fine grain scheduler may allocate tasks to queues and assign nodes to the queues with tasks. The monitoring may for example include tracking the queue length and resource utilization of the allocated nodes. If the queue length or allocated resource utilization percentage is outside a desired threshold (e.g., either predetermined or as informed by a prediction engine based on historical training data), the fine grain scheduler may request (control signal) additional resources from the coarse scheduler or release resources back (so they are available again to the coarse scheduler)” [Ferreira ¶ 38].
and selecting the first accelerator device or a different accelerator device coupled to the host device to perform a subsequent command based on the classification of the first accelerator device. “… monitor the queue length and resource utilization of the allocated nodes, wherein if the queue length or allocated resource utilization are above a first predetermined threshold, the fine-grained scheduler is configured to allocate additional resources from the coarse-grained scheduler” [Ferreira ¶ 10].
Ferreira fails to explicitly teach based on the control signal, classifying the first accelerator device as underutilized, overutilized, or optimally utilized.
However, Marchetti teaches based on the control signal, classifying the first accelerator device as underutilized, overutilized, or optimally utilized; “Based on the received consumption data 168, the resource allocator 152 can be configured to provision additional or reduce an existing amount of provisioned computing resources for the applications 147 when the received consumption data exceeds an overutilization or underutilization threshold, respectively. In one implementation, only when neither of the overutilization or underutilization threshold is exceeded, the resource allocator 152 can adjust the amount of provisioned computing resources based on input from the predictive autoscaler 126” [Marchetti ¶ 47]. “In other examples, the autoscaler can also monitor for a number of items in a job queue for a computing resource and adjust provisioned capacities based on the monitored number of items” [Marchetti ¶ 4].
Marchetti is considered to be analogous to the claimed invention because it is in the same field of considering the load. Therefore, it would be obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ferreira to incorporate the teachings of Marchetti and include based on the control signal, classifying the first accelerator device as underutilized, overutilized, or optimally utilized. Doing so would allow for the scheduling of commands to address overutilization and underutilization. “Based on the received consumption data 168, the resource allocator 152 can be configured to provision additional or reduce an existing amount of provisioned computing resources for the applications 147 when the received consumption data exceeds an overutilization or underutilization threshold, respectively” [Marchetti ¶ 47].
Claim 15 is rejected under 35 U.S.C. 103 as being unpatentable over Ferreira (US 2022/0229695 A1) in view of Marchetti (US 2018/0241812 A1) in view of Ippatapu (US 2020/0356292 A1).
With regard to claim 15, Ferreira in view of Marchetti teaches the method of claim 14, as referenced above. Ferreira further teaches wherein classifying the first accelerator device includes using information from the control signal “… monitor the queue length and resource utilization of the allocated nodes, wherein if the queue length or allocated resource utilization are above a first predetermined threshold, the fine-grained scheduler is configured to allocate additional resources from the coarse-grained scheduler” [Ferreira ¶ 10].
Ferreira in view of Marchetti fails to teach information about one or more of a request path loading status, a response path loading status, a read/write transaction ratio for the first accelerator device, and a thermal status for the first accelerator device.
However, Ippatapu teaches information about one or more of a request path loading status, a response path loading status, a read/write transaction ratio for the first accelerator device, and a thermal status for the first accelerator device. “The GPU utilization may include GPU process unit utilization, graphics memory utilization, GPU temperature, GPU fan speed, and other graphics utilization metrics” [Ippatapu ¶ 82].
Ippatapu is considered to be analogous to the claimed invention because it is in the same field of considering the load. Therefore, it would be obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ferreira in view of Marchetti to incorporate the teachings of Ippatapu and include information about one or more of a request path loading status, a response path loading status, a read/write transaction ratio for the first accelerator device, and a thermal status for the first accelerator device. Doing so would allow for the assignment of commands to take into account the temperature of the accelerator devices. “The performance data may be used to determine the performance metrics described and used in connection with the techniques herein” [Ippatapu ¶ 83].
Claims 16 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Balle (US 2023/0325265 A1) in view of Ferreira (US 2022/0229695 A1).
With regard to claim 16, Balle teaches:
A system comprising: a host device coupled to multiple accelerator devices using an interconnect; “In some embodiments, a platform 102 may function as a host platform for one or more guest systems 122 that invoke these applications. The platform may be logically or physically subdivided into clusters and these clusters may be enhanced through specialized networking accelerators and the use of Compute Express Link (CXL) memory semantics to make such clusters more efficient, among other example enhancements” [Balle ¶ 18]. “In an improved system implementation, a data center cluster may be implemented utilizing the CXL-based communication channels. For instance, a CXL-based data center cluster may include a number of host computers coupled to a CXL-based switch. Traffic within the cluster and between clusters may be implemented utilizing a network processor device (e.g., a smart network interface controller (NIC), data processing unit (DPU), infrastructure processing unit (IPU), programmable networking device, etc.), which is connected to the CXL-based switch” [Balle ¶ 43]. “CXL enables communication between host processors (e.g., CPUs) and a set of workload accelerators (e.g., graphics processing units (GPUs), field programmable gate array (FPGA) devices, tensor and vector processor units, machine learning accelerators, networking accelerators, purpose-built accelerator solutions, among other examples)” [Balle ¶ 46].
and a first accelerator device of the multiple accelerator devices, the first accelerator device including: a first controller configured to manage transactions with the host device via the interconnect; “CXL provides a rich set of sub-protocols that include I/O semantics similar to PCIe (CXL.io), caching protocol semantics (CXL.cache), and memory access semantics (CXL.mem) over a discrete or on-package link. Based on the particular accelerator usage model, all of the CXL protocols or only a subset of the protocols may be enabled” [Balle ¶ 47]. “CXL multiplexing logic (e.g., 355a-b) may also be provided to enable multiplexing of CXL protocols (e.g., I/O protocol 335a-b (e.g., CXL.io), caching protocol 340a-b (e.g., CXL.cache), and memory access protocol 345a-b (CXL.mem)), thereby enabling data of any one of the supported protocols ( e.g., 335a-b, 340a-b, 345ab) to be sent, in a multiplexed manner, over the link 350 between host processor 305 and accelerator device 310” [Balle ¶ 48, fig. 3 Examiner notes the inclusion of I/O protocol 335a in accelerator 310]. “The CXL I/O protocol, CXL.io, provides a noncoherent load/store interface for I/O devices. Transaction types, transaction packet formatting, credit-based flow control, virtual channel management, and transaction ordering rules in CXL.io may follow all or a portion of the PCIe definition. CXL cache coherency protocol, CXL.cache, defines the interactions between the device and host as a number of requests that have at least one associated response message and sometimes a data transfer. The interface consists of three channels in each direction: Request, Response, and Data” [Balle ¶ 52]. “In various embodiments, an I/O block 1304 may include an I/O controller that is integrated onto the same package as cores 1302 or may simply include interfacing logic to couple to an I/O controller that is located off-chip” [Balle ¶ 101].
a memory controller configured to manage transactions with a memory; “CXL provides a rich set of sub-protocols that include I/O semantics similar to PCIe (CXL.io), caching protocol semantics (CXL.cache), and memory access semantics (CXL.mem) over a discrete or on-package link. Based on the particular accelerator usage model, all of the CXL protocols or only a subset of the protocols may be enabled” [Balle ¶ 47]. “CXL multiplexing logic (e.g., 355a-b) may also be provided to enable multiplexing of CXL protocols (e.g., I/O protocol 335a-b (e.g., CXL.io), caching protocol 340a-b (e.g., CXL.cache), and memory access protocol 345a-b (CXL.mem)), thereby enabling data of any one of the supported protocols ( e.g., 335a-b, 340a-b, 345ab) to be sent, in a multiplexed manner, over the link 350 between host processor 305 and accelerator device 310” [Balle ¶ 48, fig. 3 Examiner notes the inclusion of memory protocol 345a in accelerator 310]. “The CXL memory protocol, CXL.mem, is a transactional interface between the processor and memory and uses the physical and link layers of CXL when communicating across dies. CXL.mem can be used for multiple different memory attach options including when a memory controller is located in the host CPU, when the memory controller is within an accelerator device, or when the memory controller is moved to a memory buffer chip, among other examples” [Balle ¶ 53].
…at least one of a request queue utilization signal and a response queue utilization signal from from the first controller or from the memory controller “CXL cache coherency protocol, CXL.cache, defines the interactions between the device and host as a number of requests that have at least one associated response message and sometimes a data transfer. The interface consists of three channels in each direction: Request, Response, and Data” [Balle ¶ 52]. “Further, the decision to delegate serialization/deserialization to another CXL device can be a result of congestion condition (e.g., queue occupancy above configurable threshold) on the IPU side, among other example features” [Balle ¶ 89]. “Workloads associated with applications, services, containers, and/or virtual machines 132 can be balanced across cores using network load and traffic patterns rather than just CPU and memory utilization metrics” [Balle ¶ 40].
and, based on the at least one utilization signal, provide (another device) a device utilization-indicating DevLoad signal to the host device via the first controller. “Further, the decision to delegate serialization/deserialization to another CXL device can be a result of congestion condition (e.g., queue occupancy above configurable threshold) on the IPU side, among other example features” [ Balle ¶ 89].
Balle fails to explicitly teach and a telemetry manager configured to receive at least one of a request queue utilization signal and a response queue utilization signal from from the first controller or from the memory controller and, based on the at least one utilization signal, provide a device utilization-indicating DevLoad signal to the host device via the first controller.
However, Ferreira teaches:
and a telemetry manager configured to receive at least one of a request queue utilization signal and a response queue utilization signal from from the first controller or from the memory controller “Each container may execute multiple fine grain processes 344, e.g., using queues that are allocated resources by the fine grain schedulers 340 (telemetry manager). For example, one fine grain scheduler 340 may create four queues and schedule multiple fine grain tasks to each queue” [Ferreira ¶ 32]. “In one embodiment, the method may comprise allocating coarse blocks of resources to portions of an application that is to be executed, allocating tasks to queues, assigning nodes to the queues with tasks, and monitoring the queue length (request queue utilization signal) and resource utilization of the allocated nodes” [Ferreira ¶ 13].
and, based on the at least one utilization signal, provide a device utilization-indicating DevLoad signal to the host device via the first controller. “In some embodiments, the coarse scheduler may allocate coarse blocks of resources to portions of an application that is to be executed, and the fine grain scheduler may allocate tasks to queues and assign nodes to the queues with tasks. The monitoring may for example include tracking the queue length and resource utilization of the allocated nodes. If the queue length or allocated resource utilization percentage is outside a desired threshold (e.g., either predetermined or as informed by a prediction engine based on historical training data), the fine grain scheduler may request (DevLoad signal) additional resources from the coarse scheduler or release resources back (so they are available again to the coarse scheduler)” [Ferreira ¶ 38].
Ferreira is considered to be analogous to the claimed invention because it is in the same field of considering the load. Therefore, it would be obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Balle to incorporate the teachings of Ferreira and include a telemetry manager configured to receive at least one of a request queue utilization signal and a response queue utilization signal from from the first controller or from the memory controller and, based on the at least one utilization signal, provide a device utilization-indicating DevLoad signal to the host device via the first controller. Doing so would allow for improved system performance. “In one embodiment, the system comprises a coarse scheduler for scheduling at the cluster/pod, a set of one or more containers for and a set of fine grain schedulers configured within each container configured to schedule at the process level. This hierarchical o multi-level coarse and fine-grained scheduling system and method may be particularly helpful to improve performance in systems where a very large number of small tasks need to be performed and can be performed in parallel” [Ferreira ¶ 7].
With regard to claim 19, Balle in view of Ferreira teaches the system of claim 16, as referenced above. Balle further teaches wherein the host device is coupled to the multiple accelerator devices using a compute express link (CXL) interconnect. “In an improved system implementation, a data center cluster may be implemented utilizing the CXL-based communication channels. For instance, a CXL-based data center cluster may include a number of host computers coupled to a CXL-based switch. Traffic within the cluster and between clusters may be implemented utilizing a network processor device (e.g., a smart network interface controller (NIC), data processing unit (DPU), infrastructure processing unit (IPU), programmable networking device, etc.), which is connected to the CXL-based switch” [Balle ¶ 43]. “CXL enables communication between host processors (e.g., CPUs) and a set of workload accelerators (e.g., graphics processing units (GPUs), field programmable gate array (FPGA) devices, tensor and vector processor units, machine learning accelerators, networking accelerators, purpose-built accelerator solutions, among other examples)” [Balle ¶ 46].
Claims 17 is rejected under 35 U.S.C. 103 as being unpatentable over Balle (US 2023/0325265 A1) in view of Ferreira (US 2022/0229695 A1) in view of Wong (US 2023/0102063 A1).
With regard to claim 17, Balle in view of Ferreira teaches the system of claim 16, as referenced above. Balle further teaches:
wherein the first accelerator device includes a cache controller coupled to a cache memory, … the cache controller. “CXL provides a rich set of sub-protocols that include I/O semantics similar to PCIe (CXL.io), caching protocol semantics (CXL.cache), and memory access semantics (CXL.mem) over a discrete or on-package link. Based on the particular accelerator usage model, all of the CXL protocols or only a subset of the protocols may be enabled” [Balle ¶ 47]. “CXL multiplexing logic (e.g., 355a-b) may also be provided to enable multiplexing of CXL protocols (e.g., I/O protocol 335a-b (e.g., CXL.io), caching protocol 340a-b (e.g., CXL.cache), and memory access protocol 345a-b (CXL.mem)), thereby enabling data of any one of the supported protocols ( e.g., 335a-b, 340a-b, 345ab) to be sent, in a multiplexed manner, over the link 350 between host processor 305 and accelerator device 310” [Balle ¶ 48, fig. 3 Examiner notes the inclusion of cache protocol 340a in accelerator 310]. “Transaction types, transaction packet formatting, credit-based flow control, virtual channel management, and transaction ordering rules in CXL.io may follow all or a portion of the PCIe definition. CXL cache coherency protocol, CXL.cache, defines the interactions between the device and host as a number of requests that have at least one associated response message and sometimes a data transfer” [Balle ¶ 52]. “The size of cache that can be supported for such devices depends on the host's snoop filtering capacity. CXL supports such devices using its optional CXL.cache link over which an accelerator can use CXL.cache protocol for cache coherency transactions” [Balle ¶ 57].
Balle fails to explicitly teach and wherein the telemetry manager is configured to provide the device utilization-indicating signal based on information about a utilization.
However, Ferreira teaches and wherein the telemetry manager is configured to provide the device utilization-indicating signal based on information about a utilization “Each container may execute multiple fine grain processes 344, e.g., using queues that are allocated resources by the fine grain schedulers 340. For example, one fine grain scheduler 340 may create four queues and schedule multiple fine grain tasks to each queue” [Ferreira ¶ 32]. “In one embodiment, the method may comprise allocating coarse blocks of resources to portions of an application that is to be executed, allocating tasks to queues, assigning nodes to the queues with tasks, and monitoring the queue length and resource utilization of the allocated nodes” [Ferreira ¶ 13].
Balle in view of Ferreira fails to explicitly teach information about a utilization of the cache controller.
However, Wong teaches information about a utilization of the cache (memory) controller. “In some examples, inspecting 220 runtime utilization metrics of a plurality of processing resources can also include collecting values of runtime utilization metrics from additional processing resources including multimedia accelerators such video codecs and audio codecs, display controllers, security processors, memory subsystems such as DMA engines and memory controllers, and bus interfaces such as a PCIe interface… Memory subsystem utilization can be expressed by metrics such as the number of read packets and the number of write packets issued over the interface within a current time period, the current utilization of ingress and egress queues or buffers, data transfer times and latency, and so on” [Wong ¶ 36].
Wong is considered to be analogous to the claimed invention because it is in the same field of considering the load. Therefore, it would be obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Balle in view of Ferreira to incorporate the teachings of Wong and include information about a utilization of the cache controller. Doing so would allow for the assignment of commands to take into account metrics across accelerator devices. “Memory subsystem utilization can be expressed by metrics such as the number of read packets and the number of write packets issued over the interface within a current time period, the current utilization of ingress and egress queues or buffers, data transfer times and latency, and so on” [Wong ¶ 36].
Claim 18 is rejected under 35 U.S.C. 103 as being unpatentable over Balle (US 2023/0325265 A1) in view of Ferreira (US 2022/0229695 A1) in view of Ippatapu (US 2020/0356292 A1).
With regard to claim 18, Balle in view of Ferreira teaches the system of claim 16, as referenced above. Balle further teaches temperature information about at least a portion of the first accelerator device, “In some embodiments, as workloads are distributed among the cores, the hypervisor 120 may steer a greater number of workloads to the higher performing cores than the lower performing cores. In certain instances, cores that are exhibiting problems such as overheating or heavy loads may be given less tasks than other cores or avoided altogether (at least temporarily). Workloads associated with applications, services, containers, and/or virtual machines 132 can be balanced across cores using network load and traffic patterns rather than just CPU and memory utilization metrics” [Balle ¶ 40].
Balle in view of Ferreira fails to explicitly teach wherein the first accelerator device includes a thermal manager configured to receive temperature information about at least a portion of the first accelerator device, and wherein the telemetry manager is configured to provide information about a temperature of the first accelerator device in the DevLoad signal.
However, Ippatapu teaches:
wherein the first accelerator device includes a thermal manager configured to receive temperature information about at least a portion of the first accelerator device, “Method 500 typically begins at block 510 where the data storage system monitors the performance of network, devices, and/or resources and collects performance data. The performance data may be associated with usage of computing resources such as CPU utilization, graphics process unit (GPU) utilization, storage device utilization, a random-access memory (RAM) utilization, and other resource utilization data over a particular period of time” [Ippatapu ¶ 82]. “The GPU utilization may include GPU process unit utilization, graphics memory utilization, GPU temperature, GPU fan speed, and other graphics utilization metrics” [Ippatapu ¶ 82].
and wherein the telemetry manager is configured to provide information about a temperature of the first accelerator device in the DevLoad signal. “The performance data may be used to determine the performance metrics described and used in connection with the techniques herein. The performance data may be used in determining a workload for one or more physical devices, logical devices or volumes (LVs), thin devices or portions thereof, and the like. The workload may also be a measurement or level of "how busy" a device, or portion thereof is for example, in terms of I/O operations such as I/O throughput such as a number of I/O per second, and the like” [Ippatapu ¶ 83].
Ippatapu is considered to be analogous to the claimed invention because it is in the same field of considering the load. Therefore, it would be obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Balle in view of Ferreira to incorporate the teachings of Ippatapu and include wherein the first accelerator device includes a thermal manager configured to receive temperature information about at least a portion of the first accelerator device, and wherein the telemetry manager is configured to provide information about a temperature of the first accelerator device in the DevLoad signal. Doing so would allow for the assignment of commands to take into account the temperature of the accelerator devices. “The performance data may be used to determine the performance metrics described and used in connection with the techniques herein” [Ippatapu ¶ 83].
Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Balle (US 2023/0325265 A1) in view of Ferreira (US 2022/0229695 A1) in view of Ibrahim (US 2020/0293445 A1).
With regard to claim 20, Balle in view of Ferreira teaches the system of claim 19, as referenced above. Balle in view of Ferreira fails to teach wherein the first controller is configured to include information about the DevLoad signal in each FLIT communicated to the host device using the interconnect.
However, Ibrahim teaches wherein the first controller is configured to include information about the DevLoad signal in each FLIT communicated to the host device using the interconnect. “The exchange of information can be done in an opportunistic way. In other words, in various embodiments, a CU transmits the collected local information to another CU as a separate one-flit packet or piggybacks the collected information on an outgoing request/reply to another CU” [Ibrahim ¶ 65].
Ibrahim is considered to be analogous to the claimed invention because it is in the same field of information transfer on a bus. Therefore, it would be obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Balle in view of Ferreira to incorporate the teachings of Ibrahim and include wherein the first controller is configured to include information about the DevLoad signal in each FLIT communicated to the host device using the interconnect. Doing so would allow for further efficiency in the sharing of data. “The exchange of information can be done in an opportunistic way” [Ibrahim ¶ 65].
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ARI F RIGGINS whose telephone number is (571)272-2772. The examiner can normally be reached Monday-Friday 7:00AM-4:30PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bradley Teets can be reached at (571) 272-3338. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/A.F.R./Examiner, Art Unit 2197
/BRADLEY A TEETS/Supervisory Patent Examiner, Art Unit 2197