DETAILED ACTION
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 1-20 are pending for examination.
Claim objections
In claims 1, 8 and 15 (line# refers to claim 1), line 13, it recites abbreviations “AI cluster”. Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception (i.e., a law of nature, a natural phenomenon, or an abstract idea) without significantly more.
Claim 1 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1, Statutory Category: Yes, the claim 1 is a computer-implemented method that recites a series of steps and therefore falls in the statutory category of a process.
Step 2A- Prong 1: Judicial Exception Recited: Yes, the claim recites: “determining a resource limit for performing the operation based on the metadata;; analyzing the at least one attribute associated with each GPU resource with respect to the resource limit; identifying a set of GPU resources from the plurality of GPU resources based on the analysis; generating a dedicated AI cluster by patching the set of GPU resources within a single cluster,; and allocating the dedicated AI cluster to the client associated with the client ID.” As drafted, the claim as a whole recites a method including steps that could be performed in the human mind, but for the recitation of generic computing components. The human mind can easily judging/evaluating/determining a resource limit for performing the operation based on the metadata; analyzing/determining/evaluating the at least one attribute associated with each GPU resource with respect to the resource limit, identifying/determining a set of GPU resources from the plurality of GPU resources based on the analysis; generating/creating/establishing a dedicated AI cluster by patching the set of GPU resources within a single cluster, and allocating/assigning the dedicated AI cluster to the client associated with the client ID. Therefore, but for the recitation of generic computing components, these steps may be a Mental Processes that can be performed in the human mind (including an observation, evaluation, judgment, opinion).
Therefore, yes, the claims do recite judicial exceptions.
Step 2A- Prong 2: Integrated into a practical Application: No, this judicial exception is not integrated into a practical application. In particular, the claim recites an additional limitations that “receiving a request for allocating graphical processing unit (GPU) resources for performing an operation, wherein the request includes metadata identifying a client identifier (ID) associated with a client, a target throughput and a target latency of the operation” and “obtaining at least one attribute associated with each GPU resource of a plurality of GPU resources available for assignment in a computing system, wherein the at least one attribute indicates capacity of a corresponding GPU resource” which are insignificant pre-solution data gathering (see MPEP § 2106.05(g)). In addition, the limitation of “wherein the dedicated AI cluster reserves a portion of a computation capacity of the computing system for a period of time” which is directed to Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a generic computer as a tool to perform an abstract idea (see MPEP 2106.05(f)). Accordingly, even in combination, these additional elements do not integrate the abstract idea into a practical application because they not impose any meaningful limits on practicing the abstract idea. Therefore, the claim is directed to the abstract idea.
Step 2B: Claim provides an Inventive Concept: No. The additional element “wherein the dedicated AI cluster reserves a portion of a computation capacity of the computing system for a period of time” which is directed to Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a generic computer as a tool to perform an abstract idea (see MPEP 2106.05(f)). In addition, the limitation of “receiving a request for allocating graphical processing unit (GPU) resources for performing an operation, wherein the request includes metadata identifying a client identifier (ID) associated with a client, a target throughput and a target latency of the operation” and “obtaining at least one attribute associated with each GPU resource of a plurality of GPU resources available for assignment in a computing system, wherein the at least one attribute indicates capacity of a corresponding GPU resource” which are insignificant pre-solution data gathering (see MPEP § 2106.05(g)) which are well understood, routine, conventional activity (see MPEP § 2106.05(d)). Courts have identified “receiving and transmitting data, storing and retrieving information”, et cetera as well understood, routine, conventional and mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (see MPEP 2106.05(f))). These additional elements and combination of the elements does not amount to significant more than the exception itself or provide an inventive concept in Step 2B.
Under the 2019 PEG, a conclusion that an additional element is insignificant extra-solution activity in Step 2A should be re-evaluated in Step 2B. Here, the “receiving” and “obtaining” steps were considered to be extra-solution activity in Step 2A as insignificant data gathering and communication and are well understood, routine, conventional activity in the field. The “receiving” and “obtaining” steps are for the purpose of “communication” and “transmitting the data” and these can be reached on one of court case (Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TLI Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016) see MPEP § 2106.05(d) II). Accordingly, a conclusion that “receiving” and “obtaining” are well understood, routine, conventional activity is supported under Berkheimer options 2.
For these reasons, there is no inventive concept in the claim, and thus the claim is ineligible.
Independent claims 8 and 15 are rejected for the same reason as claim 1 above. Claim 8 further recites “one or more processors; and a memory coupled to the one or more processors, the memory storing a plurality of instructions, executable by the one or more processors, which, when executed by the one or more processors cause the one or more processors to perform a set of operations comprising”. Claim 15 further recites “A non-transitory computer-readable medium storing a plurality of instructions executable by one or more processors to cause the one or more processors to perform a set of operations comprising”. These additional elements are directed to generic computing components/functions merely applying the abstract idea (MPEP § 2106.05(f)).
With respect to the dependent claim 2, the claim elaborates that authenticating, prior to the allocation of the dedicated AI cluster, the request based on the client ID associated with the client, wherein the request is authenticated using a private key extracted from an asymmetric key pair associated with the client ID (“authenticating” are being treated as part of abstract idea and is analogous to Mental processes, such that concept can be performed in the human mind. Further, the claim as a whole is a Mental Processes that can be performed in the human mind (including an observation, evaluation, judgment, opinion)).
With respect to the dependent claim 3, the claim elaborates that comparing a set of performance parameters corresponding to each GPU resource of the set of GPU resources with a pre-defined set of performance parameters; determining an anomaly in a first GPU resource of the set of GPU resources based on the comparison, wherein the anomaly indicates a deviation in the set of performance parameters from the pre-defined set of performance parameters; and replacing the first GPU resource with a second GPU resource within the dedicated AI cluster, wherein a hash value of the second GPU resource is the same as a hash value of the first GPU resource (“comparing”, “determining” and “replacing” are being treated as part of abstract idea and is analogous to Mental processes, such that concept can be performed in the human mind. Further, the claim as a whole is a Mental Processes that can be performed in the human mind (including an observation, evaluation, judgment, opinion)).
With respect to the dependent claim 4, the claim elaborates that determining a pre-approved quota associated with the request; determining whether the pre-approved quota exceeds a pre-defined request limit corresponding to the client ID; and blocking the request based on the determination that the pre-approved quota exceeds the pre-defined request limit. (“determining a pre-approved quota” and “blocking the request” are being treated as part of abstract idea and is analogous to Mental processes, such that concept can be performed in the human mind. Further, the claim as a whole is a Mental Processes that can be performed in the human mind (including an observation, evaluation, judgment, opinion)).
With respect to the dependent claim 5, the claim elaborates that determining a type of the operation based on the request; and selecting, based on the type of the operation, the set of GPU resources from one of a single node or multiple nodes, to generate the dedicated AI cluster (“determining” and “selecting” are being treated as part of abstract idea and is analogous to Mental processes, such that concept can be performed in the human mind. Further, the claim as a whole is a Mental Processes that can be performed in the human mind (including an observation, evaluation, judgment, opinion)).
With respect to the dependent claim 6, the claim elaborates that based on determining that the request indicates a fine-tuning operation: obtaining a data model to be fine-tuned; and executing a fine-tuning logic on the data model using the dedicated AI cluster, wherein the dedicated AI cluster is generated using the set of GPU resources selected from the single node (“obtaining” which are insignificant pre-solution data gathering (see MPEP § 2106.05(g)). In addition, “executing a fine-tuning logic on the data model using the dedicated AI cluster, wherein the dedicated AI cluster is generated using the set of GPU resources selected from the single node” which is directed to Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a generic computer as a tool to perform an abstract idea (see MPEP 2106.05(f)).
With respect to the dependent claim 7, the claim elaborates that identifying at least one GPU resource, from the set of GPU resources of the dedicated AI cluster, that is underutilized; and executing, in response to the identification of the at least one GPU resource, a dummy operation on the at least one GPU resource, wherein the dummy operation is exactly the same as the operation performed on the at least one GPU resource (“identifying” are being treated as part of abstract idea and is analogous to Mental processes, such that concept can be performed in the human mind. In addition, “executing, in response to the identification of the at least one GPU resource, a dummy operation on the at least one GPU resource…” which is directed to Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a generic computer as a tool to perform an abstract idea (see MPEP 2106.05(f)).
Dependent claims 9-14 recite the same features as applied to claims 2-7 respectively above, therefore they are also rejected under the same rationale.
Dependent claims 16-20 recite the same features as applied to claims 2-5 and 7 respectively above, therefore they are also rejected under the same rationale.
Claim Rejections - 35 USC § 103
The following is a quotation of pre-AIA 35 U.S.C. 103(a) which forms the basis for all obviousness rejections set forth in this Office action:
(a) A patent may not be obtained though the invention is not identically disclosed or described as set forth in section 102, if the differences between the subject matter sought to be patented and the prior art are such that the subject matter as a whole would have been obvious at the time the invention was made to a person having ordinary skill in the art to which said subject matter pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 8 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Sun et al. (US Pub. 2019/0197655 A1) in view of Cheng et al. (US Pub. 2021/0224665 A1).
Sun was cited in the IDS filed on 03/04/2025.
As per claim 1, Sun teaches the invention substantially as claimed including A computer-implemented method comprising:
receiving a request for allocating graphical processing unit (GPU) resources for performing an operation (Sun, Abstract, The control server receives a service request from a client system for GPU processing services, allocates multiple GPU servers within the cluster to handle GPU processing tasks specified by the service request) , wherein the request includes metadata identifying a client identifier (ID) associated with a client, (Sun, [0064] the GPU service request comprises a client request for GPU server allocation which is processed by the GPU server allocation and scheduling module 142. The client request for GPU server allocation will include associated information such as, e.g., an identifier of the client system 310 and/or GPU-accelerated application requesting the GPU service, the GPU processing task(s) to be executed, priority level information, quality of service (QoS) information, preferred GPU server capabilities, and a requested number of GPU devices and/or processing resources (e.g., GPUs, virtual central processing units (vCPUs), etc.) for handling server-side execution of GPU-accelerated application program code, and/or other types of relevant information that can utilized by GPU service platform 130 to allocate GPU resources to support GPUaaS);
determining a resource limit for performing the operation based on the metadata (Sun, [0028] the client request may specify a number of GPU devices for handling the GPU processing tasks associated with the GPU service request, wherein the allocation of one or more GPU server nodes within the server cluster 150 is determined so that the allocated GPU server nodes comprise a total number of available GPU devices that meet the specified number of GPU devices as requested in the service request; [0066] the GPU server allocation and scheduling module 142 will determine the amount of resources (e.g., GPU devices) which are needed to handle the GPU processing task(s), and proceed to allocate one or more available GPU server nodes to handle the GPU processing task(s) requested by the client system);
obtaining at least one attribute associated with each GPU resource of a plurality of GPU resources available for assignment in a computing system (Sun, Fig. 3, 146; [0026] Each GPU server node 150-1, 150-2, . . . , 150-s within the GPU server cluster 150 registers with the GPU service controller 140, wherein the GPU server node registration information is maintained in the GPU server registration information database 146. For example, when a given GPU server node is booted, the GPU server node will acquire information regarding all available GPU devices and resources on the given GPU server node. The GPU server node will then register itself and all available GPU devices and resources with the GPU service controller 140, and provide various types of registration information that enables connection to the GPU server node and utilization of the available GPU devices and resources on the given GPU server node; [0066] the GPU server allocation and scheduling module 142 will access the database of GPU server registration information 146 to determine all available GPU resources and GPU sever nodes within the current GPU resource pool of the GPU service platform 130, and determine all pending jobs that are currently scheduled for execution (or which are being executed) by the GPU server nodes. Then, based on the available GPU server nodes and GPU resources, and based on the nature of the GPU processing task(s) requested by the client system, and allocation requests (e.g., number of GPU devices) specified by the client system in the GPU service request, the GPU server allocation and scheduling module 142 will determine the amount of resources (e.g., GPU devices) which are needed to handle the GPU processing task(s), and proceed to allocate one or more available GPU server nodes to handle the GPU processing task(s) requested by the client system),
analyzing the at least one attribute associated with each GPU resource with respect to the resource limit (Sun, [0029] based on information contained within the databases 144 and 146, and information contained in the GPU service request received from the client system 110, the GPU server allocation and scheduling module 142 will have knowledge of all available GPU devices and resources within the GPU server cluster 150, knowledge of all currently mapped client-to-GPU server node connections, as well as knowledge of the required GPU processing resources based on the client service request. Using this knowledge, the GPU server allocation and scheduling module 142 will survey all available (registered) GPU server nodes, resources and currently connected jobs, and then allocate one or more registered GPU server nodes within the server cluster 150 to handle the service request; [0030] he GPU server allocation and scheduling module 142 can allocate a single GPU server node within the server cluster 150 if the single GPU server node has an amount of available GPU devices and resources which is deemed sufficient to handle execution of the GPU processing tasks associated with the service request. When the GPU processing tasks of the service request cannot be handled using the GPU devices and resources of a single GPU server node within the server cluster 150, the GPU server allocation and scheduling module 142 will select two or more GPU server nodes within the server cluster 150 which collectively have an amount of available GPU devices and resources which is sufficient to handle execution of the GPU processing tasks associated with the service request);
identifying a set of GPU resources from the plurality of GPU resources based on the analysis (Sun, [0030] the GPU server allocation and scheduling module 142 can allocate a single GPU server node within the server cluster 150 if the single GPU server node has an amount of available GPU devices and resources which is deemed sufficient to handle execution of the GPU processing tasks associated with the service request. When the GPU processing tasks of the service request cannot be handled using the GPU devices and resources of a single GPU server node within the server cluster 150, the GPU server allocation and scheduling module 142 will select two or more GPU server nodes within the server cluster 150 which collectively have an amount of available GPU devices and resources which is sufficient to handle execution of the GPU processing tasks associated with the service request; also see [0070]);
generating a dedicated AI cluster by patching the set of GPU resources within a single cluster (Sun, [0013] the GPU service platform can dynamically scale the amount of GPU resources that can be allocated to handle the service request of the client system by logically binding two or more GPU server nodes in a peer-to-peer or master/slave configuration. The logical binding of multiple GPU server nodes presents a single logical GPU server node which logically combines the GPU devices and resources of the logically bound GPU server nodes to create a pool of GPU devices and resources that are mapped to the client system for consumption.(as dedicated AI cluster)), wherein the dedicated AI cluster reserves a portion of a computation capacity of the computing system for a period of time (Sun, [0060] enables spatial sharing of the GPU resources by multiple client systems. For example, spatial sharing of a given GPU by two different client systems allows pending tasks of the different client systems to be concurrently executed using the same GPU device, but using different portions (e.g., different sets of cores) of the GPU device. Therefore, when a given GPU device is allocated to a first client system, and the first client system cannot fully utilize the given GPU device, then the same GPU device can be allocated to a second client system to allow the second client system to utilize another portion (e.g., set of cores) of the given GPU device at the same time (the different portions are reserved for processing by different client system for a period of time (i.e., during processing)); and
allocating the dedicated AI cluster to the client associated with the client ID (Sun, [0013] when the GPU processing tasks of a given service request cannot be handled using the GPU devices and resources of a single GPU server node within the cluster, the GPU service platform can dynamically scale the amount of GPU resources that can be allocated to handle the service request of the client system by logically binding two or more GPU server nodes in a peer-to-peer or master/slave configuration. The logical binding of multiple GPU server nodes presents a single logical GPU server node which logically combines the GPU devices and resources of the logically bound GPU server nodes to create a pool of GPU devices and resources that are mapped to the client system for consumption. In this configuration, the single logical GPU server node collectively utilizes the pool of GPU devices and resources across the logically bound; [0014] GPU server nodes to execute the GPU processing tasks associated with the service request of the client system as if the pool of GPU devices and resources resided in a single GPU server node dedicated to the client system for handling the service request; also see [0031], [0072-0073]).
Sun fails to specifically teach wherein the request includes metadata identifying a target throughput and a target latency of the operation; and wherein the at least one attribute indicates capacity of a corresponding GPU resource.
However, Cheng teaches wherein the request includes metadata identifying a target throughput and a target latency of the operation (Cheng, [0100] receiving a request for system resources for an inference job. For example, system resource manager 152 may receive a request for system resources for an inference job. As an example, system resource manager 152 may receive a request for system resources for an inference job associated with the machine learning model. In such an example, the request for system resources for the inference job may include a quality of service requirements associated with the inference job; [0103] a quality of service requirement includes at least one of a latency requirement (e.g., an average latency, a minimum latency, a maximum latency, etc.) for an inference job associated with a machine learning model and a throughput requirement (e.g., an average throughput, a minimum throughput, a maximum throughput, etc.) for the inference job associated with the machine learning model); and
wherein the at least one attribute indicates capacity of a corresponding GPU resource (Cheng, Fig. 1B, 156a system resources, 160a 160b, 160n GPUs; [0097] a performance profile for a system resource includes a latency (e.g., an average latency, a minimum latency, a maximum latency, etc.) associated with a machine learning model for that system resource, a throughput (e.g., an average throughput, a minimum throughput, a maximum throughput, etc.) associated with a machine learning model for that system resource, and an availability of that system resource for processing an inference job associated with the machine learning model (as attribute indicates capacity of a corresponding GPU resource). A performance profile associated with a machine learning model for a system resource may be updated in response to that system resource being used to process an inference job using that machine learning model. An initial performance profile associated with a machine learning model for a system resource may be determined based on benchmarks associated with the system resource).
It would have been obvious to one having ordinary skill in the art before the effective filling date of the claimed invention to have combined the teaching of Sun with Cheng because Cheng’s teaching of determining the performance profile associated with resource for processing the request would have provided Sun’s system with the advantage and capability to allow the system to easily identifying the correct matching resource for processing the request which improving the resource utilization and system performance.
As per claim 8, it is a system claim of claim 1 above. Therefore, it is rejected for the same reason as claim 1 above.
As per claim 15, it is a non-transitory computer-readable medium claim of claim 1 above. Therefore, it is rejected for the same reason as claim 1 above.
Claims 2, 9 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Sun and Cheng, as applied to claims 1, 8 and 15 above, and further in view of Doloff (US Patent. 10,505,925 B1).
As per claim 2, Sun and Cheng teach the invention according to claim 1 above. Sun further teaches authenticating, prior to the allocation of the dedicated AI cluster, the request based on the client ID associated with the client, wherein the request is authenticated (Sun, [0055] the server frontend 322 implements methods to handle requests that are received from the GPU API 314 of the client system 310. For example, the server frontend 322 comprises methods to process client credentials, perform authentication, and perform other types of functions to authenticate a user of the client system 310 and establish a communications session with the client system 310. For GPU service requests that require the processing of GPU code (e.g., compute kernels) passed from the GPU API 314, the server frontend 322 passes the service requests to the task queue service 324).
Sun and Cheng fail to specifically teach wherein the request is authenticated using a private key extracted from an asymmetric key pair associated with the client ID.
However, Doloff teaches wherein the request is authenticated using a private key extracted from an asymmetric key pair associated with the client ID (Doloff, Col 5, lines 23-40, Once on the network, however, the client device may still not be allowed access to most data without appropriate authentication. This can include identification of users as the users access data, for auditing and other purposes, as well as authorization of actions taken by those users on the various systems. In order to provide for such authentication, approaches in accordance with various embodiments can involve at least two separate authentication steps. These steps can include a first step to authenticate with the VPN server, in order to obtain a VPN connection; Col 11, lines 16-29, the request information can be encrypted using a specified key that indicates the information was encrypted by the first authentication provider. In some embodiments a private key from an asymmetric key pair, or a symmetric key generated during authentication and provided to the client, can also be used for verification of authentication by signing the request. The user would not need to provide the same bytes every time, but can encrypt the message or a hash of the message to prove the user has authenticated because a key is being used that was provided by the authentication mechanism. These can be time-based as well, such as by keeping track of when the corresponding public or symmetric key was generated).
It would have been obvious to one having ordinary skill in the art before the effective filling date of the claimed invention to have combined the teaching of Sun and Cheng with Doloff because Doloff’s teaching of a private key from an asymmetric key pair during authentication for verification of authentication by signing the request would have provided Sun and Cheng’s system with the advantage and capability to allow the system to improving the security in order to protected through software-based security to prevent external network traffic from reaching into the network, as well as preventing unauthenticated users from accessing private data stored in the environment.
As per claim 9, it is a system claim of claim 2 above. Therefore, it is rejected for the same reason as claim 2 above.
As per claim 16, it is a non-transitory computer-readable medium claim of claim 2 above. Therefore, it is rejected for the same reason as claim 2 above.
Claims 3 , 10 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Sun and Cheng, as applied to claims 1, 8 and 15 above, and further in view of Ranjan et al. (US Pub. 2019/0087762 A1), Hornok, Jr. et al. (US Patent. 7,093,013 B1) and JOHNSON (US Pub. 2014/0337835 A1).
As per claim 3, Sun and Cheng teach the invention according to claim 1 above. Sun teaches each GPU resource of the set of GPU resources and GPU resource within the dedicated AI cluster (Sun, [0013] the GPU service platform can dynamically scale the amount of GPU resources that can be allocated to handle the service request of the client system by logically binding two or more GPU server nodes in a peer-to-peer or master/slave configuration. The logical binding of multiple GPU server nodes presents a single logical GPU server node which logically combines the GPU devices and resources of the logically bound GPU server nodes to create a pool of GPU devices and resources that are mapped to the client system for consumption.(as dedicated AI cluster)). In addition, Cheng teaches a set of performance parameters corresponding to each GPU resource of the set of GPU resources (Cheng, [0007] a plurality of performance profiles associated with a plurality of system resources, each performance profile being associated with a machine learning model; receiving, with at least one processor, a request for system resources for an inference job associated with the machine learning model; determining, with at least one processor, a system resource of the plurality of system resources for processing the inference job associated with the machine learning model based on the plurality of performance profiles and a quality of service requirement associated with the inference job).
Sun and Cheng fail to specifically teach comparing a set of performance parameters corresponding to each GPU resource of the set of GPU resources with a pre-defined set of performance parameters; determining an anomaly in a first GPU resource of the set of GPU resources based on the comparison, wherein the anomaly indicates a deviation in the set of performance parameters from the pre-defined set of performance parameters; and replacing the first GPU resource with a second GPU resource within the dedicated AI cluster, wherein a hash value of the second GPU resource is the same as a hash value of the first GPU resource.
However, Ranjan teaches comparing a set of performance parameters corresponding to each GPU resource of the set of GPU resources with a pre-defined set of performance parameters; determining an anomaly in a first GPU resource of the set of GPU resources based on the comparison, wherein the anomaly indicates a deviation in the set of performance parameters from the pre-defined set of performance parameters; (Ranjan, [0007] a computer-implemented method of managing resource consumption is provided, comprising: receiving from a facility, by one or more servers, data produced by one or more sensors installed at the facility; detecting, by the one or more servers, an anomaly by comparing the received data with one or more data values describing expected values for the received data and identifying a deviation of the received data from the expected values; and in response to detecting the anomaly, calculating, by the one or more servers, an expected cost savings associated with correction of the anomaly, said calculation comprising: determining a usage price of at least one resource associated with the anomaly; determining, based on the received data, an amount of the at least one resource utilized by the facility as a result of the anomaly; and determining the expected cost savings based on the usage price and the amount of the at least one resource utilized by the facility as a result of the anomaly).
It would have been obvious to one having ordinary skill in the art before the effective filling date of the claimed invention to have combined the teaching of Sun and Cheng with Ranjan because Ranjan’s teaching of detecting an anomaly by comparing the received data with one or more data values describing expected values for the received data and identifying a deviation would have provided Sun and Cheng’s system with the advantage and capability to allow the system to easily identify the anomaly based on the comparing which improving the system performance and efficiency.
Sun, Cheng and Ranjan fail to specifically teach replacing the first GPU resource with a second GPU resource within the dedicated AI cluster.
However, Hornok teaches replacing the first GPU resource with a second GPU resource within the dedicated AI cluster (Hornok, Col 1, lines 49-53, monitors groups of resources controlled by "clusters" of computer systems. In the event of a failure in a resource, CLUSTER SERVER.TM. can deactivate the resource and replace it with another "backup" resource (i.e., it performs a failover of the resource; please note: GPU resources and dedicated AI cluster was taught by Sun).
It would have been obvious to one having ordinary skill in the art before the effective filling date of the claimed invention to have combined the teaching of Sun, Cheng and Ranjan with Hornok because Hornok’s teaching of replace the failed resource with another "backup" resource would have provided Sun, Cheng and Ranjan’s system with the advantage and capability to preventing any potential system failure in order to improving the system reliability and performance.
Sun, Cheng, Ranjan and Hornok fail to specifically teach wherein a hash value of the second GPU resource is the same as a hash value of the first GPU resource.
However, JOHNSON teaches wherein a hash value of the second GPU resource is the same as a hash value of the first GPU resource (JOHNSON, [0020] VGPU resource manager 127 retrieves the contents of the graphics object and sends content (or a pointer to content) for the graphics object received from the VM to resource key generator 134 of HGPU resource manager 130 as shown in FIG. 1 by arrow 122. In operation 158, resource key generator 134 uses content from the graphics resource to generate the resource key. Resource key generator 134 may use a hash algorithm to generate the resource key from all or part of the graphics object contents so that identical resources will have identical keys. For example, the hash algorithm receives or otherwise accesses a series of bytes of content from the graphics resource and generates a unique number corresponding to those bytes. With typical hash algorithms, there is no guarantee that two different graphics resources will not be mapped to the same hash value such that resource key collisions are possible, but extremely unlikely. Thus, the resource key is considered "substantially unique" to the graphics resource).
It would have been obvious to one having ordinary skill in the art before the effective filling date of the claimed invention to have combined the teaching of Sun, Cheng, Ranjan and Hornok with JOHNSON because JOHNSON’s teaching of use a hash algorithm to generate the resource key from all or part of the graphics object contents so that identical resources will have identical keys would have provided Sun, Cheng, Ranjan and Hornok’s system with the advantage and capability to allow the system ensuring the same resources are utilized for processing in order to improving the system performance and efficiency.
As per claim 10, it is a system claim of claim 3 above. Therefore, it is rejected for the same reason as claim 3 above.
As per claim 17, it is a non-transitory computer-readable medium claim of claim 3 above. Therefore, it is rejected for the same reason as claim 3 above.
Claims 4, 11 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Sun and Cheng, as applied to claims 1, 8 and 15 above, and further in view of Chakraborty et al. (US Patent. 10,152,268 B1).
As per claim 4, Sun and Cheng teach the invention according to claim 1 above. Sun and Cheng fail to specifically teach determining a pre-approved quota associated with the request; determining whether the pre-approved quota exceeds a pre-defined request limit corresponding to the client ID; and blocking the request based on the determination that the pre-approved quota exceeds the pre-defined request limit.
However, Chakraborty teaches determining a pre-approved quota associated with the request (Chakraborty, Fig. 7, 710; Col 14, lines 14-22, at block 705, a target storage system receives a replication request to replicate data from a source storage system, where the target storage system stores data replicated from one or more source storage systems. At block 710, in response to the replication request, the target storage system identifies a replication resource limit (e.g., capacity soft quota limit, capacity hard quota limit, stream quota limit, etc.) associated with the data to be replicated from the source storage system);
determining whether the pre-approved quota exceeds a pre-defined request limit corresponding to the client ID (Chakraborty, Fig. 7, 725, Yes, Col 14, lines 31-33, processing logic of the target storage system determines whether a hard quota limit has been exceeded); and
blocking the request based on the determination that the pre-approved quota exceeds the pre-defined request limit (Chakraborty, Fig. 7, 730; Col 14, lines 31-40, processing logic of the target storage system determines whether a hard quota limit has been exceeded. At block 730, if the processing logic of the target storage system determines that the hard quota limit has been exceeded, then the target storage system denies the replication request and raises a hard quota exceeded alert).
It would have been obvious to one having ordinary skill in the art before the effective filling date of the claimed invention to have combined the teaching of Sun and Cheng with Chakraborty because Chakraborty’s teaching of determines that the hard quota limit has been exceeded, then the target storage system denies the replication request would have provided Sun and Cheng’s system with the advantage and capability to allow the system to preventing any pontifical failures due to exceeded quota limit which improving the system performance and efficiency.
As per claim 11, it is a system claim of claim 4 above. Therefore, it is rejected for the same reason as claim 4 above.
As per claim 18, it is a non-transitory computer-readable medium claim of claim 4 above. Therefore, it is rejected for the same reason as claim 4 above.
Claims 5, 12 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Sun and Cheng, as applied to claims 1, 8 and 15 above, and further in view of KIM et al. (US Pub. 2013/0198755 A1).
As per claim 5, Sun and Cheng teach the invention according to claim 1 above. Sun teaches generate the dedicated AI cluster (Sun, [0013] the GPU service platform can dynamically scale the amount of GPU resources that can be allocated to handle the service request of the client system by logically binding two or more GPU server nodes in a peer-to-peer or master/slave configuration. The logical binding of multiple GPU server nodes presents a single logical GPU server node which logically combines the GPU devices and resources of the logically bound GPU server nodes to create a pool of GPU devices and resources that are mapped to the client system for consumption).
Sun and Cheng fail to specifically teach determining a type of the operation based on the request; and selecting, based on the type of the operation, the set of GPU resources from one of a single node or multiple nodes, to generate the dedicated AI cluster.
However, KIM teaches determining a type of the operation based on the request; and selecting, based on the type of the operation, the set of GPU resources from one of a single node or multiple nodes, to generate the dedicated AI cluster (KIM, [0011] improvement of overall operation performance can be achieved only when applications capable of efficiently using resources based on the characteristics of each node are executed. That is, as shown in FIG. 2, the resource agent node 140 may include nodes 141 including performance computation acceleration apparatuses such as a Graphics Processing Unit (GPU); [0039] the resource management unit 211 may analyze received parameters, transfer information necessary to execute a lower block, and control a task so that an operation is appropriately performed; [0041] may manage a resource allocation policy and a quantitative resource allocation policy based on the characteristics of resources, and determine a resource allocation policy based on the characteristic of a task which is performed in a cluster system. That is, the resource policy management unit 212 may determine allocated nodes and resources by applying a policy in which a characteristic has been taken into consideration to the management of resource allocation policies and the selection of allocated resources; Fig. 7, S411, s412, s416 select node; also see [0075] to [0076]).
It would have been obvious to one having ordinary skill in the art before the effective filling date of the claimed invention to have combined the teaching of Sun and Cheng with KIM because KIM’s teaching of determining the resource allocation based on the characteristic of the task/request would have provided Sun and Cheng’s system with the advantage and capability to allow the system to ensuring the resource allocated are best fit for the request in order to improving the resource utilization and system performance.
As per claim 12, it is a system claim of claim 5 above. Therefore, it is rejected for the same reason as claim 5 above.
As per claim 19, it is a non-transitory computer-readable medium claim of claim 5 above. Therefore, it is rejected for the same reason as claim 5 above.
Claims 6 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Sun, Cheng and KIM, as applied to claims 5 and 12 above, and further in view of Ramakrishnan et al. (US Pub. 2024/0177049 A1).
As per claim 6, Sun, Cheng and KIM teach the invention according to claim 5 above. Sun further teaches using the dedicated AI cluster, wherein the dedicated AI cluster is generated using the set of GPU resources selected from the single node (Sun, Fig. 4A, 418, 420; [0068] In this allocation determination process, the GPU server allocation and scheduling module 142 can allocate a single registered GPU server (either local or remote to the GPU service platform 130) to handle the GPU processing task(s) associated with current GPU service request, if a single registered GPU with sufficient GPU processing resources is available to execute the GPU processing task(s)).
Sun, Cheng and KIM fail to specifically teach based on determining that the request indicates a fine-tuning operation: obtaining a data model to be fine-tuned; and executing a fine-tuning logic on the data model using the dedicated AI cluster.
However, Ramakrishnan teaches based on determining that the request indicates a fine-tuning operation (Ramakrishnan, [0019] Machine learning system 110 may receive a fine tuning request, as indicated at 104. The request may specify a data set for tuning (e.g., by providing a location and access credentials to obtain the data set). Like the pre-trained model, the data set may have user data access restrictions, which may only allow access to the data set for tuning purposes by machine learning system 110 (e.g., and not to a provider of the pre-trained model or any other entity). In some embodiments, the fine tuning request 104 may include parameters for controlling or effecting the fine tuning (e.g., number of training epochs, stop criteria, or various other features or hyperparameters). In some embodiments, as discussed in detail below with regard to FIGS. 7 and 9, the fine tuning request 104 may include analyses to perform before, during, and/or after fine tuning the machine learning model; [0020] implement confidential model tuning 130 to perform the fine tuning request 104. Confidential model tuning 130 may implement the various techniques discussed below with regard to FIGS. 3 and 5-9, including providing tuning analyses and reports and, if permitted or enabled, a debugging mode to share some data with the provider only for development assistance and/or model information with a model user);
obtaining a data model to be fine-tuned; and executing a fine-tuning logic on the data model using the dedicated AI cluster (Ramakrishnan, [0035] Confidential model development 216 may perform various ingestion or other processing steps to translate, format, otherwise package the pre-trained model for confidential use, including tuning. For example, if an image is provided, the image may be registered, format, or otherwise prepared for execution on a host for training or deployment. The packaged confidential pre-trained ML model may be stored 342 in data store(s) 310, as model and tuning instructions 345. In some embodiments, access restrictions on data store(s) 310 may limit access to machine learning service 210 components (and not requests associated with a model provider account that submitted upload 340 or a model user account, which may submit a fine-tuning request 350). A catalog, index, or other information describing the availability of pre-trained models may be updated in order to make the model and tuning instructions 345 available for selection in a fine tuning request 340; [0036] A fine tuning request 350 may be received through interface 211. The request may specify a tuning data set 316 (e.g., by specifying a model identifier for a listing of a pre-trained machine learning model) and, in some embodiments, tuning parameters (which may be used together with the tuning instructions provided by the model provider). In some embodiments, interface 211 may provide search criteria, filter criteria, or other user interface elements for identifying a pre-trained machine learning model (e.g., by model type, by model provider, task, or other category (or combination of categories); [0037] Tuning instructions for the pre-trained machine learning model may then be performed and the fine-tuned model generated; also see [0042]).
It would have been obvious to one having ordinary skill in the art before the effective filling date of the claimed invention to have combined the teaching of Sun, Cheng and KIM with Ramakrishnan because Ramakrishnan’s teaching of performing tuning instructions for the pre-trained machine learning model and the fine-tuned model generated based on fin-turning request would have provided Sun, Cheng and KIM’s system with the advantage and capability to allow the system to improve pre-trained machine learning model performance to perform particular tasks for particular model users.
As per claim 13, it is a system claim of claim 6 above. Therefore, it is rejected for the same reason as claim 6 above.
Claims 7, 14 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Sun and Cheng, as applied to claims 1, 8 and 15 above, and further in view of JULIEN et al. (US Pub. 2022/0214912 A1) and Kamaraj et al. (US Pub. 2024/0231900 A1).
As per claim 7, Sun and Cheng teach the invention according to claim 1 above. Sun and Cheng fail to specifically teach identifying at least one GPU resource, from the set of GPU resources of the dedicated AI cluster, that is underutilized; and executing, in response to the identification of the at least one GPU resource, a dummy operation on the at least one GPU resource, wherein the dummy operation is exactly the same as the operation performed on the at least one GPU resource.
However, JULIEN teaches identifying at least one GPU resource, from the set of GPU resources of the dedicated AI cluster, that is underutilized and performing, in response to the identification of the at least one GPU resource, an eviction (JULIEN, [0060] At operation 518, the monitoring agent 126A of the GPGPU agent 124A monitors resources of the GPGPU 102A.sub.1 and/or the workload of the application 104A being processed by the GPGPU 102A.sub.1. In particular, the monitoring agent 126A monitors and profiles the GPGPU 102A.sub.1, including associated resources and workloads/applications 104 being processed by the GPGPU 102A.sub.1. The monitoring information produced by the monitoring agent 126A can be used to form usage/performance profiles for the workload/application 104A that describe the performance/operation of workload of the application 104A on the GPGPU 102A.sub.1, including respective GPGPU memory 112A usage. Although shown as a single operation, the monitoring agent 126A may continually generate monitoring information to update the usage/performance profiles of the workload/application 104; [0061] At operation 520, the proxy agent 122 may determine if an idle period for the workload/application 104A on the GPGPU 102A.sub.1 is below a threshold usage/idle value. In particular, the proxy agent 122 may determine whether the workload/application 104A, which is being processed by the GPGPU 102A.sub.1, is efficiently using resources of the GPGPU 102A.sub.1 or is underutilizing resources of the GPGPU 102A.sub.1. In some embodiments, the proxy agent 122 may use a usage/performance profile of the workload/application 104A, which was generated based on monitoring information from the monitoring agent 126A, to determine whether an idle period for the workload/application 104A on the GPGPU 102A.sub.1 is below the threshold usage/idle value. In response to determining at operation 520 that an idle period for the workload/application 104A on the GPGPU 102A.sub.1 is not below the threshold usage/idle value, the method 500 may return to operation 516 to continue processing the workload/application 104A. Conversely, in response to determining at operation 520 that an idle period for the workload/application 104A on the GPGPU 102A.sub.1 is below the threshold usage/idle value, the method 500 may move to operation 522; [0086] At operation 814, the proxy agent 122 adds an identifier of the first workload to a candidate list of workloads for eviction in response to determining that the performance profile of the first workload indicates that the usage of resources of the first GPGPU 102A.sub.1 by the first workload is below the threshold usage value).
It would have been obvious to one having ordinary skill in the art before the effective filling date of the claimed invention to have combined the teaching of Sun and Cheng with JULIEN because JULIEN’s teaching of adds an identifier of the first workload to a candidate list of workloads for eviction in response to determining that the performance profile of the first workload indicates that the usage of resources of the first GPGPU by the first workload is below the threshold usage value would have provided Sun and Cheng’s system with the advantage and capability to allow the system to efficiently utilizing the resource in order to improving the system performance and efficiency.
Sun, Cheng and JULIEN fail to specifically teach executing, in response to the identification of the at least one GPU resource, a dummy operation on the at least one GPU resource, wherein the dummy operation is exactly the same as the operation performed on the at least one GPU resource.
However, Kamaraj teaches executing, in response to the identification of the at least one GPU resource, a dummy operation on the at least one GPU resource, wherein the dummy operation is exactly the same as the operation performed on the at least one GPU resource (Kamaraj, [0003] ne execution unit can also be used to run a redundant instance of a thread run on another execution unit in order to check that the system is working as expected; [0006] “critical” for the present purposes means critical to the desired application for which the thread in question is being run. Particularly, “critical” may refer herein to any thread for which it is desired to run a duplicate instance—the check thread—and check at least one result of at least one operation performed by the critical thread against a corresponding result of the same operation performed by the check thread; [0012] being a duplicate of the critical thread, to be executed on a second one of the plurality of execution units other than the first execution unit; [0015] if one of the execution units is detected to be idle upon the reading of the request from the request buffering storage to thereupon select an idle one of the execution units as the second execution unit to begin the execution of the check thread; [0070] Based on the idle flag(s) 1030 . . . 103_K−1, the safety thread scheduler (STS) 608 is configured to detect when at least one of the cores 102_0 . . . 102_K−1 is idle; also see [0074]-[0079]).
It would have been obvious to one having ordinary skill in the art before the effective filling date of the claimed invention to have combined the teaching of Sun, Cheng and JULIEN with Kamaraj because Kamaraj’s teaching of performing duplicated check thread on the idle resource, such that the check thread performs the same operation as the original thread would have provided Sun, Cheng and JULIEN’s system with the advantage and capability to allow the system to improving the resource utilization and computational performance (see Kamaraj, [0131] “The performance improvements may include one or more of increased computational performance, reduced latency, increased throughput, and/or reduced power consumption”).
As per claim 14, it is a system claim of claim 7 above. Therefore, it is rejected for the same reason as claim 7 above.
As per claim 20, it is a non-transitory computer-readable medium claim of claim 7 above. Therefore, it is rejected for the same reason as claim 7 above.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ZUJIA XU whose telephone number is (571)272-0954. The examiner can normally be reached M-F 9:30-5:30 EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Aimee J Li can be reached at (571) 272-4169. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ZUJIA XU/Primary Examiner, Art Unit 2195