DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim(s) 1-20 is/are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 1 recite: “receiving orchestration logic corresponding to a machine learning model and computing resource data corresponding to a plurality of computing resources available; based on the orchestration logic and the computing resource data, generating the machine learning model”.
The examiner is unclear how an orchestration logic can be received corresponding to a machine learning model and the machine learning model is subsequently generated. There is no way to determine if a received orchestration logic corresponds to any machine learning model prior to a machine learning model being generated.
Claims 12 and 20 recites similarly recites convention and/or orchestration logic. Therefore, are rejected based on similar rationale.
Claim 5 recite: “local computing resources”. The examiner is unclear how the limitation “local” should be interpreted. For example, local to particular geographic location, local to a particular network, local to users’ physical location, etc.
Claim 8 (similarly claim 19) recite: “increase efficiency of utilization”. The term “increase efficiency” is ambiguous. The examiner is unclear how efficiency is measured and/or determined. For example, maximum use of resource, minimum use of resource, maximum cost, minimum cost, etc.
Claim 13 recite: “computing resource efficiency rules configured to optimize usage”. The examiner is unclear how “optimize” is measured and/or determined. For example, maximum use of resource, minimum use of resource, maximum cost, minimum cost, etc.
Claims 2-11 and 13-19 are rejected based on rejection of its corresponding dependent claim.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claim(s) 1-20 is/are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Ramanujan et al. (Pub 20240143414) (hereafter Ram).
As per claim 1, Ram teaches:
A method comprising:
receiving orchestration logic corresponding to a machine learning model and computing resource data corresponding to a plurality of computing resources available; ([Paragraph 36], The computing resources 124 are any computing hardware that facilitates the execution of the representative workload 114 via a graphics processing units (GPUs), central processing units (CPUs), memory, storage, and so forth. As mentioned above, the model endpoint 120 consists of an artificial intelligence model 126. In various examples, the artificial intelligence model 126 can be a large language transformer-based model that is deployed via the model config 104. This utilizes the computing resources 124 to process the representative workload 114. Stated another way, some, or all of the representative workload 114 is scheduled to the computing resources 124 for execution by the artificial intelligence model 126 available at the model endpoint 120.)
based on the orchestration logic and the computing resource data, generating the machine learning model and an allocation of one or more computing resources of the plurality of computing resources available for the machine learning model, in which the machine learning model conforms to the orchestration logic; and ([Paragraph 30], In various examples, the model configuration 104 defines various parameters of an artificial intelligence model that is to be deployed. For example, the model configuration 104 defines output type, cache size, the amount of available computing resources, and so forth.)
deploying the machine learning model and the allocation of the one or more computing resources of the plurality of computing resources available for the machine learning model. ([Paragraph 21], The techniques described herein enhance the operation of cloud computing platforms for deploying artificial intelligence models… In various examples, the model configuration 104 defines various parameters of an artificial intelligence model that is to be deployed. For example, the model configuration 104 defines output type, cache size, the amount of available computing resources, and so forth. [Paragraph 25], As the workload is executed by the artificial intelligence model, an analytics layer extracts various performance metrics from the model deployment layer. Alternatively, the analytics layer extracts the performance metrics after the artificial intelligence model is finished executing the representative workload (e.g., via a log file). In one example, the analytics layer is configured to track the latency of the artificial intelligence model in response to computing demand imposed by the representative workload. In another example, the analytics layer is configured to monitor data throughput in response to increasing demand. Moreover, the analytics layer can track multiple performance metrics simultaneously (e.g., latency and data throughput).)
As per claim 2, rejection of claim 1 is incorporated:
Ram teaches wherein the orchestration logic includes one or more of computing resource allocation logic, workflow management logic, scaling logic, performance optimization logic, failover and recovery logic, or cost management logic. ([Paragraph 21], The techniques described herein enhance the operation of cloud computing platforms for deploying artificial intelligence models… In various examples, the model configuration 104 defines various parameters of an artificial intelligence model that is to be deployed. For example, the model configuration 104 defines output type, cache size, the amount of available computing resources, and so forth.)
As per claim 3, rejection of claim 1 is incorporated:
Ram teaches further comprising receiving convention logic pertaining to the machine learning model, and wherein the generating of the machine learning model and the allocation of the one or more computing resources of the plurality of computing resources available is based in part on the convention logic pertaining to the machine learning model. ([Paragraph 25], As the workload is executed by the artificial intelligence model, an analytics layer extracts various performance metrics from the model deployment layer. Alternatively, the analytics layer extracts the performance metrics after the artificial intelligence model is finished executing the representative workload (e.g., via a log file). In one example, the analytics layer is configured to track the latency of the artificial intelligence model in response to computing demand imposed by the representative workload. In another example, the analytics layer is configured to monitor data throughput in response to increasing demand. Moreover, the analytics layer can track multiple performance metrics simultaneously (e.g., latency and data throughput). [Paragraph 35], Furthermore, the model endpoint 120 also handles communication between the model deployment layer 118 and the source of the request (e.g., a user). [Paragraph 53], Secondly, the performance measurement 402 identifies an inflection point 410 at which the computing resources 124 are heavily utilized and latency begins to increase significantly. Beyond the inflection point 410, there is an exponential increase in latency 404 for relatively little gain in data throughput 406. In the event the computing resources 124 for an artificial intelligence model 126 reach the inflection point 410, an automatic resource scaling process is triggered.)
As per claim 4, rejection of claim 1 is incorporated:
Ram teaches wherein the plurality of computing resources available includes one or more of Graphics Processing Units (“GPUs”), Central Processing Units (“CPUs”), or Tensor Processing Units (“TPUs”). ([Paragraph 73], Processing unit(s), such as processing unit(s) of processing system 602, can represent, for example, a CPU-type processing unit, a GPU-type processing unit, a field-programmable gate array (FPGA), another class of digital signal processor (DSP), or other hardware logic components that may, in some instances, be driven by a CPU. For example, illustrative types of hardware logic components that can be used include Application-Specific Integrated Circuits (ASICs), Application-Specific Standard Products (ASSPs), System-on-a-Chip Systems (SOCs), Complex Programmable Logic Devices (CPLDs), and the like.)
As per claim 5, rejection of claim 1 is incorporated:
Ram teaches wherein the plurality of computing resources available includes one or more of cloud computing resources and local computing resources. ([Paragraph 4], The techniques described herein provide systems for enhancing cloud computing platforms by introducing load testing and performance benchmarking for artificial intelligence models. [Paragraph 30], FIG. 1 illustrates a system 100 that enables load testing and/or performance benchmarking for artificial intelligence models in cloud computing platforms. As shown in FIG. 1, a setup layer 102 is responsible for initializing a load test environment by selecting a model configuration 104, an input dataset 106, and a load profile 108. In various examples, the model configuration 104 defines various parameters of an artificial intelligence model that is to be deployed. For example, the model configuration 104 defines output type, cache size, the amount of available computing resources, and so forth. In addition, the model configuration 104 can also define a geographic region in which the artificial intelligence model is deployed as the location of a user relative to a cloud computing center can have an impact on latency. [Paragraph 72], Stated another way, one processing unit of the processing system 602 may be located in a first location (e.g., a rack within a datacenter) while another processing unit of the processing system 602 is located in a second location separate from the first location.)
As per claim 6, rejection of claim 1 is incorporated:
Ram teaches receiving updated computing resource data corresponding to the plurality of computing resources available;
based on the updated computing resource data, generating an updated allocation of one or more computing resources of the plurality of computing resources available for the machine learning model; and
deploying the updated allocation of one or more computing resources of the plurality of computing resources available for the machine learning model. ([Paragraph 86-105], A, a method for evaluating a performance of an artificial intelligence model for a cloud computing platform, the method comprising: selecting a load profile defining a plurality of parameters; generating a representative workload configured to evaluate the performance of the artificial intelligence model for the cloud computing platform under a computing demand imposed by a workload context represented by the load profile; executing the representative workload on a computing resource utilizing the artificial intelligence model; extracting a plurality of performance metrics from the computing resource following execution of the representative workload on the computing resource utilizing the artificial intelligence model; synthesizing a performance measurement based on the plurality of performance metrics extracted from the computing resource; updating the load profile using the performance measurement by modifying the plurality of parameters defined by the load profile; and generating, utilizing the modified plurality of parameters in the updated load profile, an updated representative workload configured to further evaluate the performance of the artificial intelligence model of the cloud computing platform under a different computing demand imposed by an updated workload context.)
As per claim 7, rejection of claim 6 is incorporated:
Ram teaches wherein the generating of the updated allocation is based on the updated computing resource data indicating a utilization amount of the one or more computing resources not exceeding a threshold utilization amount. ([Paragraph 44], In still another example, the feedback layer 138 can identify load thresholds at which a given workload must be directed to new and/or additional computing resources 124 for proper processing often known as a failover. In typical scenarios, a failover is a high severity issue requiring significant effort to address. By identifying concrete load thresholds at which a failover is likely to occur, the system 100 can gracefully handle a failover. For instance, by automatically managing the workload and/or by issuing an advance warning to a system administrator or engineer. In still another example, by iteratively modifying and testing input datasets 106 and load profiles 108, the system 100 can identify various tiers of customer experience to offer users. For example, a user paying for a premium subscription to the cloud computing platform may receive lower latencies at higher data throughputs than a free user. [Paragraph 53], Secondly, the performance measurement 402 identifies an inflection point 410 at which the computing resources 124 are heavily utilized and latency begins to increase significantly. Beyond the inflection point 410, there is an exponential increase in latency 404 for relatively little gain in data throughput 406. In the event the computing resources 124 for an artificial intelligence model 126 reach the inflection point 410, an automatic resource scaling process is triggered. Stated another way, the inflection point 410 defines a threshold latency 404 and data throughput 406 at which additional computing resources 124 are assigned to the artificial intelligence model 126 for processing various workloads. Furthermore, the inflection point 410 can also be referred to as an “elbow state”, indicating a zone of data throughput 406 at which the latency 404 increases exponentially.)
As per claim 8, rejection of claim 6 is incorporated:
Ram teaches wherein the updated allocation of one or more computing resources of the plurality of computing resources available for the machine learning model is adapted to the updated computing resource data of the plurality of computing resources to increase efficiency of utilization of the plurality of computing resources available by the machine learning model. ([Paragraph 43], Furthermore, the system 100 can perform custom tuning of the artificial intelligence model 126 to maximize efficiency for available computing capacity. [Paragraph 86-105], A, a method for evaluating a performance of an artificial intelligence model for a cloud computing platform, the method comprising: selecting a load profile defining a plurality of parameters; generating a representative workload configured to evaluate the performance of the artificial intelligence model for the cloud computing platform under a computing demand imposed by a workload context represented by the load profile; executing the representative workload on a computing resource utilizing the artificial intelligence model; extracting a plurality of performance metrics from the computing resource following execution of the representative workload on the computing resource utilizing the artificial intelligence model; synthesizing a performance measurement based on the plurality of performance metrics extracted from the computing resource; updating the load profile using the performance measurement by modifying the plurality of parameters defined by the load profile; and generating, utilizing the modified plurality of parameters in the updated load profile, an updated representative workload configured to further evaluate the performance of the artificial intelligence model of the cloud computing platform under a different computing demand imposed by an updated workload context.)
As per claim 9, rejection of claim 1 is incorporated:
Ram teaches further comprising: receiving one or more performance metrics corresponding to performance of the machine learning model; based on the one or more performance metrics, generating an updated allocation of one or more computing resources of the plurality of computing resources available for the machine learning model; and deploying the updated allocation of one or more computing resources of the plurality of computing resources available for the machine learning model. ([Paragraph 43], Furthermore, the system 100 can perform custom tuning of the artificial intelligence model 126 to maximize efficiency for available computing capacity. [Paragraph 86-105], A, a method for evaluating a performance of an artificial intelligence model for a cloud computing platform, the method comprising: selecting a load profile defining a plurality of parameters; generating a representative workload configured to evaluate the performance of the artificial intelligence model for the cloud computing platform under a computing demand imposed by a workload context represented by the load profile; executing the representative workload on a computing resource utilizing the artificial intelligence model; extracting a plurality of performance metrics from the computing resource following execution of the representative workload on the computing resource utilizing the artificial intelligence model; synthesizing a performance measurement based on the plurality of performance metrics extracted from the computing resource; updating the load profile using the performance measurement by modifying the plurality of parameters defined by the load profile; and generating, utilizing the modified plurality of parameters in the updated load profile, an updated representative workload configured to further evaluate the performance of the artificial intelligence model of the cloud computing platform under a different computing demand imposed by an updated workload context.)
As per claim 10, rejection of claim 9 is incorporated:
Ram teaches wherein the one or more performance metrics include computing resource usage metrics. ([Paragraph 43], Furthermore, the system 100 can perform custom tuning of the artificial intelligence model 126 to maximize efficiency for available computing capacity. [Paragraph 86-105], A, a method for evaluating a performance of an artificial intelligence model for a cloud computing platform, the method comprising: selecting a load profile defining a plurality of parameters; generating a representative workload configured to evaluate the performance of the artificial intelligence model for the cloud computing platform under a computing demand imposed by a workload context represented by the load profile; executing the representative workload on a computing resource utilizing the artificial intelligence model; extracting a plurality of performance metrics from the computing resource following execution of the representative workload on the computing resource utilizing the artificial intelligence model; synthesizing a performance measurement based on the plurality of performance metrics extracted from the computing resource; updating the load profile using the performance measurement by modifying the plurality of parameters defined by the load profile; and generating, utilizing the modified plurality of parameters in the updated load profile, an updated representative workload configured to further evaluate the performance of the artificial intelligence model of the cloud computing platform under a different computing demand imposed by an updated workload context.)
As per claim 11, rejection of claim 9 is incorporated:
Ram teaches wherein the generating of the updated allocation of one or more computing resources of the plurality of computing resources available for the machine learning model is based on at least one performance metric of the one or more performance metrics not exceeding a threshold amount. ([Paragraph 44], In still another example, the feedback layer 138 can identify load thresholds at which a given workload must be directed to new and/or additional computing resources 124 for proper processing often known as a failover. In typical scenarios, a failover is a high severity issue requiring significant effort to address. By identifying concrete load thresholds at which a failover is likely to occur, the system 100 can gracefully handle a failover. For instance, by automatically managing the workload and/or by issuing an advance warning to a system administrator or engineer. In still another example, by iteratively modifying and testing input datasets 106 and load profiles 108, the system 100 can identify various tiers of customer experience to offer users. For example, a user paying for a premium subscription to the cloud computing platform may receive lower latencies at higher data throughputs than a free user. [Paragraph 53], Secondly, the performance measurement 402 identifies an inflection point 410 at which the computing resources 124 are heavily utilized and latency begins to increase significantly. Beyond the inflection point 410, there is an exponential increase in latency 404 for relatively little gain in data throughput 406. In the event the computing resources 124 for an artificial intelligence model 126 reach the inflection point 410, an automatic resource scaling process is triggered. Stated another way, the inflection point 410 defines a threshold latency 404 and data throughput 406 at which additional computing resources 124 are assigned to the artificial intelligence model 126 for processing various workloads. Furthermore, the inflection point 410 can also be referred to as an “elbow state”, indicating a zone of data throughput 406 at which the latency 404 increases exponentially.)
As per claims 12, 14, 15 and 17-19, these are system claims corresponding to the method claims 1, 2, 3, 6, 8 and 9. Therefore, rejected based on similar rationale.
As per claim 13, rejection of claim 12 is incorporated:
Ram teaches wherein the convention logic includes one or more computing resource efficiency rules configured to optimize usage of the plurality of computing resources available for the machine learning model. ([Paragraph 43], Furthermore, the system 100 can perform custom tuning of the artificial intelligence model 126 to maximize efficiency for available computing capacity. [Paragraph 86-105], A, a method for evaluating a performance of an artificial intelligence model for a cloud computing platform, the method comprising: selecting a load profile defining a plurality of parameters; generating a representative workload configured to evaluate the performance of the artificial intelligence model for the cloud computing platform under a computing demand imposed by a workload context represented by the load profile; executing the representative workload on a computing resource utilizing the artificial intelligence model; extracting a plurality of performance metrics from the computing resource following execution of the representative workload on the computing resource utilizing the artificial intelligence model; synthesizing a performance measurement based on the plurality of performance metrics extracted from the computing resource; updating the load profile using the performance measurement by modifying the plurality of parameters defined by the load profile; and generating, utilizing the modified plurality of parameters in the updated load profile, an updated representative workload configured to further evaluate the performance of the artificial intelligence model of the cloud computing platform under a different computing demand imposed by an updated workload context.)
As per claim 16, rejection of claim 12 is incorporated:
Ram teaches wherein the receiving of the convention logic is via user input via a user interface of a client device. ([Paragraph 2], Within an AI platform, a user can dynamically scale computing resources to suit their needs. For example, the user may scale resources vertically by selecting a more or less powerful computing instance (e.g., a virtual machine). Similarly, the user may scale resources horizontally by selecting the number of computing instances to deploy. From the perspective of the user, there are unlimited computing resources at their disposal to meet any demand. In reality, however, the available set of computing resources is finite and shared among multiple users. [Paragraph 35], various examples, the model endpoint 120 is the component which enables a user to interact with the cloud computing platform.)
As per claim 20, this is a non-transitory computer-readable storage medium claim corresponding to the method claim 1. Therefore, rejected based on similar rationale.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DONG U KIM whose telephone number is (571)270-1313. The examiner can normally be reached 9:00am - 5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bradley Teets can be reached at 5712723338. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/DONG U KIM/Primary Examiner, Art Unit 2197