DETAILED ACTION
This Office Action is in response to claims filed on 02/02/2026.
Claims 1-20 are pending.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant’s arguments, see pg. 8-10 of applicant's remark, filed 02/02/2026, with respect to 35 U.S.C. 112(b) rejection of claim 1 and dependent claims 2-20 rejected under similar rationale have been fully considered and are persuasive. The 112(b) rejection of 11/05/2025 has been withdrawn.
Applicant's arguments filed 02/02/2026 have been fully considered but they are not persuasive. Applicant argues in substance:
The characterization or interpretation of Ranganathan does not support a disclosure of claim 1 for “perform a background operation to select second DPUs of the plurality of DPUs based, at least in part, on capabilities associated with the first DPU of the plurality of DPUs being within a threshold”, and instead, uses metadata from the workload. The secondary reference, Sindhu, is not described as curing such deficiency in Ranganathan.
With respect to point (a), Examiner respectfully disagrees. In summary, Applicant argues that Ranganathan fails to teach the recited limitation presented above because selection of a second set of accelerator devices is based on workload metadata. However, Examiner submits that this argument of distinction is not supported by the claim language.
Examiner notes that the claim does not recite any limitation that defines how the capabilities associated with a first DPU come to be known and what mechanism establishes or communicates such associations between capabilities and selected DPUs. Thus, the Examiner appeals to the broadest reasonable interpretation of the claim, wherein the claim merely requires that capabilities associated with a first DPU serve as the basis, at least in part, for selection of a second set of DPUs.
As a matter of context, Ranganathan discloses a compute sled “receiving the availability data” in which describes a plurality of descriptions for each accelerator device within the compute sled including device architectures, device kernels specifying supported operations, and quality of service parameters associated with particular devices (Ranganathan, [0082]). As such, Ranganathan teaches, further in paragraph [0082], that “the compute 1614 selects, as a function of the availability data (e.g., the availability data obtained in block 1704), one or more target accelerator devices to assist in the execution of the workload (e.g., to perform one or more operations associated with the corresponding application 1682, 1684)” such that reasonably fulfills the limitation of “receive a selection of a first data processing unit (DPU) of a plurality of DPUs to perform a workload” recited in claim 1.
In view of the teaching above, Ranganathan reasonably teaches “selecting one or more target accelerator devices” wherein “the compute sled 1614 may determine compatible types of accelerator devices for the workload” (Ranganathan, [0083]). Herein, Ranganathan further teaches that “the compute sled 1614 may read metadata associated with the workload … indicative of accelerator device types capable of performing operations within the workload”, thus identifying the capabilities of a DPU required for operation (Ranganathan, [0083]). Moreover, Ranganathan teaches “using the metadata, the compute sled 1614 may match the available types of accelerator devices … to the accelerator types and/or kernels indicated in the metadata” such that leads to “select[ing], as a function of a target quality of service … and the quality of service data … one or more of the available accelerator devices”, wherein a quality of service would designate the threshold of capabilities. The workload metadata constitutes a reasonable expression of capabilities associated with the first DPU, such that metadata sufficient to inform the selection of the first DPU is equally sufficient to select second DPUs possessing the same capabilities.
Further, Ranganathan discloses that “the operations described reference to the method … could be performed in a different order and/or concurrently (e.g., the compute sled 1614 may continually obtain availability data while the compute sled 1614 concurrently executes a workload in cooperation with one or more target accelerator devices)”, thereby suggesting that accelerator discovery may occur concurrently (e.g., in the background) in relation to workload execution (Ranganathan, [0084]).
Examiner directs Applicant’s attention to the instant specification in support of this position. The instant specification recites “when a select DPU is to perform RegEx, for instance, the background operation is able to query a registry of enumerated DPUs and is able to determine other DPUs that support RegEx capabilities with similar performance threshold, for instance, as the select DPU” (Specification, [0024]). It is understood that the selection of the second DPUs in view of the capabilities associated with the first DPU is necessarily informed by the nature of the workload to be executed (e.g., the system identifies that a workload requires RegEx operations and thus selects DPUs that support RegEx). As such, the workload metadata of Ranganathan to identify and select devices comprising of common capabilities within a threshold is consistent with the scope of the limitation in view of the instant specification.
Examiner notes that the purported distinction between capabilities expressed as attributes of the selected DPU and capabilities derived from the metadata of the workload that the selected DPU is to perform is not a distinction supported by the claim language. The claim does not explicitly limit how the capabilities are associated with the first selected DPU. Thus, the workload metadata of Ranganathan may reasonably be interpreted as “capabilities associated with the first DPU … within a threshold” and that such metadata is used, at least in part, in “select[ing] second DPUs of the plurality of DPUs”. As illustrated by the instant specification, an association between a DPU and its capabilities is defined by what the workload requires the DPU to perform within a performance threshold. Accordingly, the Examiner maintains that the reference teaches the claimed limitation of “perform a background operation to select second DPUs of the plurality of DPUs based, at least in part, on capabilities associated with the first DPU of the plurality of DPUs being within a threshold” as reasonably interpreted.
Argument has not been found to be persuasive.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 5, 6, 11, 15, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Ranganathan et al. Pub. No. US 2020/0341810 A1 (hereinafter Ranganathan) in view of Sindhu et al. Pub. No. US 2019/0012350 A1 (hereinafter Sindhu).
With regard to claim 1, Ranganathan teaches a system comprising at least one processor and memory comprising instructions that when executed by the at least one processor cause the system to ([0022], The disclosed embodiments may be implemented in some cases, hardware, firmware, software, or any combination thereof. The disclosed embodiments mays also be implemented as instructions carried by or stored on a transitory or non-transitory machine-readable (e.g., computer-readable) storage medium, which may be read and executed by one or more processors; [0074], a system for executing one or more workloads):
receive a selection of a first [accelerator] data processing unit (DPU) of a plurality of [accelerators] DPUs to perform a workload ([0082], the compute sled 1614 selects as a function of the availability data (e.g., the availability data obtained in block 1704), one or more target accelerator devices to assist in the execution of the workload (e.g., to perform one or more operations associated with the corresponding application 1682, 1684);
perform a background operation to select second DPUs of the plurality of [accelerators] DPUs based, at least in part, on capabilities associated with the first DPU of the plurality of DPUs begin within a threshold ([0083], the compute sled 1614 may determine compatible types of accelerators for the workload, as indicated in block 1734. For example, and as indicated in block 1736, the compute sled 1614 may read metadata associated with the workload (e.g., with the application 1682, 1684) indicative of accelerator device types capable of performing operations within the workload … Using the metadata, the compute sled 1614 may match the available types of accelerator devices 162, 1622, 1624, 1626 and kernels (e.g., as indicated in the availability data from block 1718, 1720, 1722) to the accelerator types and/or kernels indicated in the metadata); and
cause the first DPU and the second DPUs of the plurality of [accelerators] DPUs to be in a load balancing arrangement to perform the workload ([0083], As indicated in block 1742, the compute sled 1614 may partition the workload (e.g., the set of operations to be performed in association with the application 1682, 1684) into multiple sections (e.g., subsets of operations) to be performed by multiple sections (e.g., subsets of operations) to be performed by multiple target accelerator devices. In doing so, the compute sled 1614 may partition the workload as a function of the compatibility of the types of accelerator devices 1620, 1622, 1624, 1626 to the operations in the workload (e.g., match operations associated with the workload with accelerator devices that are capable of performing those operations).
However, Ranganathan does not explicitly teach that accelerators are data processing units.
Sindhu teaches plurality of data processing units (DPUs) to perform a workload ([0062], DPU 13 includes a plurality of programmable processing cores 140A-140N (“cores 140”) … DPU 130 also includes networking, one or more host units 146, a memory controller 144, and one or more accelerators 148)
It would have been obvious to one of ordinary skill in the art at the time the invention was filed to apply the teachings of Sindhu with the teachings of Ranganathan in order to provide a system that teaches accelerators and data processing units as entities capable of balancing and performing a workload. The motivation for applying Sindhu teaching with Ranganathan teaching is to provide a system that teaches simple substitution of the accelerator of Ranganathan with the data processing unit of Sindhu as networked offload acceleration devices capable of performing a plurality of workloads and that the substitution would reasonably be expected to obtain predictable results. Ranganathan and Sindhu are analogous art directed towards partitioning of resources in accordance with a workload and data center resource management. Therefore, it would have been obvious for one of ordinary skill in the art to combine Sindhu with Ranganathan to teach the claimed invention in order to provide data processing units capable of handling offloaded acceleration workloads.
With regard to claim 5, Ranganathan teaches wherein the capabilities and the threshold include two or more of hardware revision, binning information, clock frequency, or throughput ([0074], In the illustrative embodiment, the orchestrator server 1520 may selectively allocate and/or deallocate physical resources 620 from the sleds 400 and/or add or remove one or more sleds 400 from the managed node 1570 as a function of quality of service (QoS) targets (e.g., performance targets associated with a throughput, latency, instruction per second (Examiner notes: where IPS is a function of clock frequency), etc.) associated with a service level agreement for the workload).
With regard to claim 6, Ranganathan teaches wherein the memory comprising instructions that when executed by the at least one processor further cause the system to:
determine first hardware capabilities ([0083], The compute sled 1614 may partition the workload as a function of the quality of service data (e.g., the quality of service parameters from 1726 of FIG. 17) associated with the available accelerator devices 1620, 1622, 1624, 1626 and target qualities of service for one or more sets of operations within the workload, as indicated in block 1746) of the first DPU of the plurality of DPUs ([0083], To do so, the compute sled 1614 may select the VPU 1626 to perform object recognition operations, as the quality of service data indicates that the VPU 1626 performs those operations with lower latency (e.g., in a shorter time period) than an FPGA) and second hardware capabilities of the second DPUs of the plurality of DPUs ([0083], the FPGA 1620 should perform the corresponding decision-making operations, as it performs those operations faster (e.g., with lower latency) than the VPU 1626); and
enable the workload to be distributed, in the load balancing arrangement, to the first DPU and the second DPUs of the plurality of DPUs according to the first and the second hardware capabilities ([0083], Subsequently, the method 1700 advances to block 1748 of FIG. 19, in which the compute sled 1614 executes the workload with the one or more target accelerator devices (e.g., the accelerator devices 1620, 1626).
With regard to claim 11, Ranganathan teaches a method for seamless offload of a workload to a plurality of data processing units (DPUs) ([0081], the compute sled 1614, in operation, may perform a method of utilizing the discovery service 1618 to select and communicate with one or more accelerator devices). Claim 11 is a method having similar limitations as claim 1. Thus, claim 11 is rejected for the same rationale as applied to claim 1.
With regard to claim 15, it is a method having similar limitations as claim 6. Thus, claim 15 is rejected for the same rationale as applied to claim 6.
With regard to claim 18, Ranganathan teaches a system for seamless offload of a workload to a plurality of data processing units (DPUs) ([0078], a system 1600 for providing an accelerator device discovery service includes multiple accelerator sleds 1610, 1612, and a compute sled 1614 in communication with each other and with an orchestrator server 1616) Claim 18 is a system having similar limitations as claim 1. Thus, claim 18 is rejected for the same rationale as applied to claim 1.
Claims 2, 3, 7, 12, 13, 16, 19, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Ranganathan in view of Sindhu as applied to claim 1, 11, and 18 respectively above, and further in view of Sukhomlinov et al. Pub. No. US 2019/0004871 A1 (hereinafter Sukhomlinov).
With regard to claim 2, Ranganathan teaches perform the selection using user input, a reference application, or an application programming interface (API) ([0080], The orchestrator server 1616, is the illustrative embodiment, executes a discovery service 1618. The discovery service 1618, in the illustrative embodiment, is a set of operations in which the orchestrator server 1616 receives queries from other devices in the system 1600, such as from an accelerator device selection logic unit … of the compute sled 1614, requesting availability data indicative of accelerator devices in the system 1600 that are available to a particular tenant (e.g., a customer for whom an application 1682, 1684 is being executed by one or more processors 160 of the compute sled 1614),
However, Ranganathan does not explicitly teach a support library associated with a registry to enable discovery of data processing unit accelerators associated with their capabilities.
Sukhomlinov teaches wherein a library definition associated with a support library ([0030], As instances of a microservice or a microservice-capable resource come online (Examiner notes: DPU accelerators), they may register with a service discovery function (SDF). The SDF can then maintain a catalog of available microservices, which may include mappings for translating the standard microservices API calls to an API call usable by a particular instance of the microservice) enables the background operation that comprises a query to the plurality of DPUs based, at least in part, on the capabilities and the threshold ([0031], When the microservices driver (Examiner notes: the discovery service of Ranaganathan) receives a request for a new microservice, it can query the SDF, and identify the availability of microservices instances; [0032], Similar procedures can be used for … allocating a hardware accelerator).
It would have been obvious to one of ordinary skill in the art at the time the invention was filed to apply the teachings of Sukhomlinov with the teachings of Ranganathan and Sindhu in order to provide a system that teaches a support service to enable discovery of accelerator devices based on their capabilities to perform a workload in a distributed environment. The motivation for applying Sukhomlinov teaching with Ranganathan and Sindhu teaching is to provide a system that allows for data centers delivering important service functions related to encryption and compression to incorporate accelerators, thereby increasing flexibility and speed in executing these functions (Sukhomlinov, [0018]-[0019]). Ranganathan and Sindhu and Sukhomlinov are analogous art directed towards allocation of resources in a data center environment. Therefore, it would have been obvious for one of ordinary skill in the art to combine Sukhomlinov with Ranganathan and Sindhu to teach the claimed invention in order to provide support servicing in discovering and coordinating accelerator devices.
With regard to claim 12, it is a method having similar limitations as claim 2. Thus, claim 12 is rejected for the same rationale as applied to claim 2.
With regard to claim 19, it is a system having similar limitations as claim 2. Thus, claim 19 is rejected for the same rationale as applied to claim 2.
With regard to claim 3, Ranganathan teaches wherein the memory comprising instructions that when executed by the at least one processor further causes the system to:
perform the background operation to query a plurality of capabilities of the plurality of DPUs based, at least in part, on the capabilities and the threshold associated with the first DPU of the plurality of DPUs ([0081], In response to a determination to utilize the discovery service 1618, the method 1700 advances to block 1704 in which the compute sled 1614 obtains, from a discovery service (e.g., the discovery service 1618), availability data indicative of a set of the accelerator devices 1620, 1622, 1624, 1626 that are available to assist in the execution of a workload (e.g., a set of operations associated with one of the applications 1682, 1684); and
However, Ranganathan does not explicitly teach registering DPU accelerators associated with their capabilities to a support library registry data structure.
Sukhomlinov teaches register the plurality of capabilities of the plurality of DPUs in a support library ([0090], When each resource comes online, it may register with a service discovery function (SDF); [0092], Similarly, resource pool 424, FPGA 420, and hardware accelerator 416 may also register their capabilities SDF 440; SDF 440 stores all of these in catalog 448)
It would have been obvious to one of ordinary skill in the art at the time the invention was filed to apply the teachings of Sukhomlinov with the teachings of Ranganathan and Sindhu in order to provide a system that teaches an accelerator registry data structure managed by the support service. The motivation for applying Sukhomlinov teaching with Ranganathan and Sindhu teaching is to provide a system that allows for device selection abstraction such that a registry enables a simple and flexible addressing and invocation method for defining and discovering accelerator devices based on desired functionality (Sukhomlinov, [0129]). Ranganathan and Sindhu and Sukhomlinov are analogous art directed towards allocation of resources in a data center environment. Therefore, it would have been obvious for one of ordinary skill in the art to combine Sukhomlinov with Ranganathan and Sindhu to teach the claimed invention in order to provide an accelerator registry used to query device availability and capabilities.
With regard to claim 13, it is a method having similar limitations as claim 3. Thus, claim 13 is rejected for the same rationale as applied to claim 3.
With regard to claim 20, it is a system having similar limitations as claim 3. Thus, claim 20 is rejected for the same rationale as applied to claim 3.
With regard to claim 7, Ranganathan teaches wherein the memory comprising instructions that when executed by the at least one processor further cause the system to:
poll, in periodic intervals, a plurality of capabilities associated with the plurality of DPUs ([0078], Additionally, the accelerator sled 1650 includes a discovery data logic unit 1650 which may be embodied as any device or circuitry… configured to provide data … indicative of quality of service metrics associated with each accelerator device 1620, 1622 (e.g., a present latency a present bandwidth, a present load on the accelerator device which may affect latency, etc.) In the illustrative embodiment, the discovery data logic unit 1640, in operation, continually provides updated data as described above to the orchestrator server 1616)
However, Ranganathan does not explicitly teach DPU accelerators and their capabilities associated with the support library registry data structure.
Sukhomlinov teaches the plurality of capabilities are associated in a support library ([0090], When each resource comes online, it may register with a service discovery function (SDF); [0092], Similarly, resource pool 424, FPGA 420, and hardware accelerator 416 may also register their capabilities SDF 440; SDF 440 stores all of these in catalog 448)
Rationale to claim 3 applied here.
With regard to claim 16, it is a method having similar limitations as claim 7. Thus, claim 16 is rejected for the same rationale as applied to claim 7.
Claims 4 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Ranganathan in view of Sindhu as applied to claims 1 and 11 respectively above, and further in view of Sukhomlinov et al. Pub. No. US 2019/0004871 A1 (hereinafter Sukhomlinov) in view of Burger et al. Pub. No. US 2016/0364271 A1 (hereinafter Burger).
With regard to claim 4, Ranganathan teaches wherein the memory comprising instructions that when executed by the at least one processor further causes the system to:
start a [service] support library in a discovery phase ([0080], The orchestrator server 1616, in the illustrative embodiment, executes a discovery service 1618);
…
enable the [service] support library to be in an operation phase ([0083], Subsequently, the method 1700 advances to block 1748 of FIG. 19, in which the compute sled 1614 executes the workload with the one or more target accelerator devices (e.g., the accelerator devices 1620, 1626);
perform the background operation to query further capabilities of the plurality of DPUs using the [service] support library in the discovery phase or the operation phase ([0081], the compute sled 1615, in the illustrative embodiment, receives, from the discovery service 1618, availability data indicative of a set of accelerator devices 1620, 1622, 1623, 1626 available to a tenant (e.g., available for use by one of the applications 1682, 1684) associated with the request. In doing so, and as indicated in block 1712, the compute sled 1614 may receive identifiers (e.g., a set of numbers and/or letters that uniquely identify the corresponding accelerator device 162, 1622, 1624, 1626 (Examiner notes: performed during the discovery phase); and
build a load balancing [arrangement] table to allocate the workload for the first DPU and second DPUs of the plurality of DPUs ([0083], As indicated in block 1742, the compute sled 1614 may partition the workload (e.g., the set of operations to be performed in association with the application 1682, 1684) into multiple sections (e.g., subsets of operations) to be performed by multiple target accelerator devices. In doing so, the compute sled 1614 may partition the workload as function of the compatibility of the types of the accelerator devices 162, 1622, 1624, 1626 to the operations in the workload (e.g., match operations associated with the workload with accelerator devices that are capable of performing those operations).
However, Ranganathan does not explicitly teach a support library that is associated with a registry data structure capable of registering DPU accelerator devices.
Sukhomlinov teaches enable the support library ([0030], As instances of a microservice or microservice-capable resource come online, they may register with a service discovery function (SDF). The SDF can maintain a catalog of available microservices, which may include mappings for translating the standard microservices API calls to an API call usable by a particular instance of the microservice. This architecture enables the specialization of certain architecture capabilities, such as “bump in the wire” acceleration, FPGA function sets invoked from processing cores, or purpose optimized processors with highly specialized software; [0032], Similar procedure can be used for … allocating a hardware accelerators)
register individual DPUs of the plurality of DPUs in the support library ([0090], When each resource (Examiner notes: DPU accelerator) comes online, it may register with a service discovery function (SDF) 440. SDF 400 may be separate VM running in the data center, may be a module or function of an orchestrator (e.g., orchestrator 260 of FIG. 2);
It would have been obvious to one of ordinary skill in the art at the time the invention was filed to apply the teachings of Sukhomlinov with the teachings of Ranganathan and Sindhu in order to provide a system that teaches an accelerator registry data structure managed by the support service. The motivation for applying Sukhomlinov teaching with Ranganathan and Sindhu teaching is to provide a system that allows for device selection abstraction such that a registry enables a simple and flexible addressing and invocation method for defining and discovering accelerator devices based on desired functionality (Sukhomlinov, [0129]). Ranganathan and Sindhu and Sukhomlinov are analogous art directed towards allocation of resources in a data center environment. Therefore, it would have been obvious for one of ordinary skill in the art to combine Sukhomlinov with Ranganathan and Sindhu to teach the claimed invention in order to provide an accelerator registry used to query device availability and capabilities.
However, the combination does not explicitly teach that a load balancing table is maintained to distribute the workload.
Burger teaches a load balancing table to allocate the workload for the first DPU and second DPUs of the plurality of DPUs (Fig. 3, Hardware Accelerator Table 330 and a distribution of functionality across networked hardware accelerators for a workflow of server blades 322-352; Abstract, Request for hardware acceleration from workflows being executed by general central processing units of server computing devices are directed to hardware accelerators in accordance with a table associating available hardware accelerators with the computing operations they are optimized to perform. Load balancing, as well as dynamic modifications in available hardware accelerators, is accomplished through updates to such a table; [0046], As another example, multiple server blade computing devices can be executing workflows that can seek to have the same fucntiaonlity hardware accelerated … In such an example, dynamic updates to a hardware accelerator table, such as exemplary hardware accelerator table 330, can provide load-balancing and otherwise enable multiple workflows, such as exemplary workflows 271 and 341, to share the hardware acceleration capabilities of the exemplary network hardware accelerators)
It would have been obvious to one of ordinary skill in the art at the time the invention was filed to apply the teachings of Burger with the teachings of Ranganathan and Sindhu in order to provide a system that teaches a load balancing table of available DPU accelerators for allocation to a workload. The motivation for applying Burger teaching with Ranganathan and Sindhu teaching is to provide a system that allows for a particular load balancing algorithm to be selected best suited to the nature of the workload and available accelerators, thereby maximizing the utilization of hardware accelerators and improving overall workflow performance. Ranganathan and Sindhu and Burger are analogous art directed towards allocation of resources considering the load. Therefore, it would have been obvious for one of ordinary skill in the art to combine Burger with Ranganathan and Sindhu to teach the claimed invention in order to provide a load balancing table to manage and allocate data processing accelerators.
With regard to claim 14, it is a method having similar limitations as claim 4. Thus, claim 14 is rejected for the same rationale as applied to claim 4.
Claims 8, 9, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Ranganathan in view of Sindhu as applied to claims 1 and 11 respectively above, and further in view of Burger et al. Pub. No. US 2016/0364271 A1 (hereinafter Burger).
With regard to claim 8, Ranganathan teaches wherein the memory comprising instructions that when executed by the at least one processor further cause the system to:
build a load balancing [arrangement] table of the first DPU and the second DPUs of the plurality of DPUs to allocate the workload ([0083], As indicated in block 1742, the compute sled 1614 may partition the workload (e.g., the set of operations to be performed in association with the application 1682, 1684) into multiple sections (e.g., subsets of operations) to be performed by multiple target accelerator devices. In doing so, the compute sled 1614 may partition the workload as function of the compatibility of the types of the accelerator devices 162, 1622, 1624, 1626 to the operations in the workload (e.g., match operations associated with the workload with accelerator devices that are capable of performing those operations), wherein at least one of the second DPUs of the plurality of DPUs comprises features that are in a first threshold range within the threshold of the capabilities associated with the first DPU of the plurality of DPUs ([0083], the compute sled 1614 may select the VPU 1626 to perform object recognition operations, as the quality of service data indicates that the VPU 1626 performs those operations with lower latency) and wherein another of the second DPUs of the plurality of DPUs comprise features that are in a second threshold range within the threshold of capabilities associated with the first DPU of the plurality of DPUs ([0083], the FPGA 1620 should perform the corresponding decision-making operations, as it performs those operations faster (e.g., with lower latency) … which the compute sled 1614 executes the workload with the one or more target accelerator devices (e.g., the accelerator devices 1620, 1626)
However, the combination does not explicitly teach that a load balancing table is maintained to distribute the workload.
Burger teaches a load balancing table to allocate the workload for the first DPU and second DPUs of the plurality of DPUs (Fig. 3, Hardware Accelerator Table 330 and a distribution of functionality across networked hardware accelerators for a workflow of server blades 322-352; Abstract, Request for hardware acceleration from workflows being executed by general central processing units of server computing devices are directed to hardware accelerators in accordance with a table associating available hardware accelerators with the computing operations they are optimized to perform. Load balancing, as well as dynamic modifications in available hardware accelerators, is accomplished through updates to such a table; [0046], As another example, multiple server blade computing devices can be executing workflows that can seek to have the same functionality hardware accelerated … In such an example, dynamic updates to a hardware accelerator table, such as exemplary hardware accelerator table 330, can provide load-balancing and otherwise enable multiple workflows, such as exemplary workflows 271 and 341, to share the hardware acceleration capabilities of the exemplary network hardware accelerators)
Rationale to claim 4 applied here.
With regard to claim 17, it is a method having similar limitations as claim 8. Thus, claim 17 is rejected for the same rationale as applied to claim 8.
With regard to claim 9, Ranganathan teaches wherein the memory comprising instructions that when executed by the at least one processor further cause the system to:
build a load balancing [arrangement] table of the first DPU and the second DPUs of the plurality of DPUs to allocate the workload ([0083], As indicated in block 1742, the compute sled 1614 may partition the workload (e.g., the set of operations to be performed in association with the application 1682, 1684) into multiple sections (e.g., subsets of operations) to be performed by multiple target accelerator devices. In doing so, the compute sled 1614 may partition the workload as function of the compatibility of the types of the accelerator devices 162, 1622, 1624, 1626 to the operations in the workload (e.g., match operations associated with the workload with accelerator devices that are capable of performing those operations), wherein a first time to perform the workload is allocated to at least one of the second DPUs of the plurality of DPUs ([0075], the orchestrator server 1520 may identify trends in resource utilization of the workload (e.g., the application 1532, such as by identifying phases of execution (e.g., time periods in which different operations, each having different resource utilization characteristics, are performed) of the workload (e.g., the application 1532) and pre-emptively identifying available resources in the data center 100 and allocating them to the managed node 1570 (e.g., within a predefined time period of the associated phase beginning) and a second time to perform the workload is allocated to another of the second DPUs of the plurality of DPUs, and wherein the first time is more than the second time ([0083], the compute sled 1614 may determine that a set of object recognition operations are to be performed within a particular time period in order for a set of corresponding decision-making operations to be performed, based on the identified objects, within a total time period (Examiner notes: such that the particular time to allocate the second accelerator occurs after the corresponding decision-making operation of a first time).
However, the combination does not explicitly teach that a load balancing table is maintained to distribute the workload.
Burger teaches a load balancing table to allocate the workload for the first DPU and second DPUs of the plurality of DPUs (Fig. 3, Hardware Accelerator Table 330 and a distribution of functionality across networked hardware accelerators for a workflow of server blades 322-352; Abstract, Request for hardware acceleration from workflows being executed by general central processing units of server computing devices are directed to hardware accelerators in accordance with a table associating available hardware accelerators with the computing operations they are optimized to perform. Load balancing, as well as dynamic modifications in available hardware accelerators, is accomplished through updates to such a table; [0046], As another example, multiple server blade computing devices can be executing workflows that can seek to have the same functionality hardware accelerated … In such an example, dynamic updates to a hardware accelerator table, such as exemplary hardware accelerator table 330, can provide load-balancing and otherwise enable multiple workflows, such as exemplary workflows 271 and 341, to share the hardware acceleration capabilities of the exemplary network hardware accelerators)
Rationale to claim 4 applied here.
Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Ranganathan in view of Sindhu as applied to claim 1 above, and further in view of Goyal et al. Pub. No. US 2020/0159568 A1 (hereinafter Goyal).
With regard to claim 10, Goyal teaches wherein the capabilities and the thresholds are associated with hardware and software features for one or more of data compression, data encryption, or regular expression operations ([0037], As further described herein, in one example, each access node 17 is a highly programmable I/O processor (referred to as a data processing unit, or DPU) specially designed for offloading certain functions from servers 12. In one example, each access node 17 includes a number of internal processor clusters, each including two or more processing cores and equipped with hardware engines (also referred to herein as accelerators) that offload cryptographic functions, compression and decompression, regular expression (RegEx) processing, data storage functions, and networking operations).
It would have been obvious to one of ordinary skill in the art at the time the invention was filed to apply the teachings of Goyal with the teachings of Ranganathan and Sindhu in order to provide a system that teaches hardware and software features associated with the accelerators include data compression, data encryption and regular expression operations. The motivation for applying Goyal teaching with Ranganathan and Sindhu teaching is to provide a system that allows for offloading data compression, data encryption, and regular expression operations to accelerators, thereby freeing general-purpose processors to dedicate resources to application workloads (Goyal, [0037]). Ranganathan and Sindhu and Goyal are analogous art directed towards partitioning of resources in accordance with a workload and data center resource management. Therefore, it would have been obvious for one of ordinary skill in the art to combine Goyal with Ranganathan and Sindhu to teach the claimed invention in order to provide improved resource management through workload offloading of specific operations.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to IVAN A CASTANEDA whose telephone number is (571)272-0465. The examiner can normally be reached Monday-Friday 9:30AM-5:30PM EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Aimee Li can be reached at (571) 272-4169. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/I.A.C./Examiner, Art Unit 2195
/Aimee Li/Supervisory Patent Examiner, Art Unit 2195