DETAILED ACTION
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This Office Action is in response to claims filed 04/11/2024.
Claims 1-20 are pending.
Claim Objections
Claims 1, 7, 10, and 18 objected to because of the following informalities: “data engineering pipeline” should be “a data engineering pipeline”. Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention recites a judicial exception, is directed to that judicial exception, an abstract idea, as it has not been integrated into practical application and the claims further do not recite significantly more than the judicial exception. Examiner has evaluated the claims under the framework provided in the 2019 Patent Eligibility Guidance published in the Federal Register 01/07/2019 and has provided such analysis below.
Step 1:
Claims 1-9 are directed to a system and falls within the statutory category of machines; Claims 10-17 are directed to a non-transitory machine-readable medium and falls within the statutory category of manufacture; Claims 18-20 are directed to methods and fall within the statutory category of processes. Therefore, “Are the claims to a process, machine, manufacture or composition of matter?” Yes.
In order to evaluate the Step 2A inquiry “Is the claim directed to a law of nature, a natural phenomenon or an abstract idea?” we must determine, at Step 2A Prong 1, whether the claim recites a law of nature, a natural phenomenon or an abstract idea and further whether the claim recites additional elements that integrate the judicial exception into a practical application.
Step 2A Prong 1:
Claims 1, 10 and 18: The limitations “provisioning a head node and a first group of worker nodes based on the processing capacity, the memory capacity, the processing capacity threshold, and the memory capacity threshold; provisioning a first portion of data engineering pipeline on each of the first group of worker nodes; and provisioning a first portion of the machine learning model on each of the first group of worker nodes”, as drafted, is a process that, but for the recitation of generic computing components, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, but for the recitation of generic computing components being used as a tool to perform the functionality, a person may make a mental determination of what resources are necessary for completing a job and which portions of that job require certain resources. The act of provisioning is merely an assignment/plan for the resource’s used which can be performed as a mental determination.
Therefore, yes, Claims 1, 10 and 18 recite judicial exceptions.
The claims have been identified to recite judicial exceptions, Step 2A Prong 2 will evaluate whether the claims are directed to the judicial exception.
Step 2A Prong 2:
Claims 1, 10 and 18: The judicial exceptions are not integrated into practical applications. In particular, the claims recite the following additional elements – “obtaining a processing capacity associated with a machine learning model and obtaining a memory capacity associated with the machine learning model; obtaining a processing capacity threshold and obtaining a memory capacity threshold;”, merely recite insignificant extra-solution data gathering and data storage which do not integrate the judicial exception into a practical application. See MPEP § 2106.05(g). Further, “executable instructions that, when executed by the processing system, facilitate performance of operations”, is a recitation of generic computing components and functions merely being used as a tool to apply the abstract idea (see MPEP § 2106.05(f)). Lastly, Claim 1 recites “a processing system including a processor; and a memory that stores executable instructions”, is a further recitation of generic computing components and functions merely being used as a tool to apply the abstract idea (see MPEP § 2106.05(f)).
Therefore, “Do the claims recite additional elements that integrate the judicial exception into a practical application? No, these additional elements do not integrate the abstract idea into a practical application and they do not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
After having evaluating the inquires set forth in Steps 2A Prong 1 and 2, it has been concluded that the Claims 1, 10 and 18 not only recite a judicial exception but that the claims are directed to a judicial exception as a judicial exception has not been integrated into a practical application.
Step 2B:
Claims 1, 10 and 18: The claims do not include additional elements, alone or in combination, that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements amount to no more than generic computing components, field of use/technological environment, and insignificant extra-solution activity which do not amount to significantly more than the abstract idea. Further, the insignificant extra-solution activity is Well-Understood, Routine, and Conventional. “The courts have recognized the following computer functions as well‐understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity. i. Receiving or transmitting data over a network … iv. Storing and retrieving information in memory”. See MPEP § 2106.05(d)(II).
Therefore, “Do the claims recite additional elements that amount to significantly more than the judicial exception? No, these additional elements, alone or in combination, do not amount to significantly more than the judicial exception.
Having concluded analysis within the provided framework, Claims 1, 10 and 18 do not recite patent eligible subject matter under 35 U.S.C. § 101.
Claims 2 and 14: “the operations comprise managing, by the head node, the first group of worker nodes”, as drafted, is a process that, but for the recitation of generic computing components, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, but for the recitation of generic computing components being used as a tool to perform the functionality, a person may mentally manage a group of resources. With regard to integration into practical application and whether additional elements amount to significantly more, Claims 2 and 14 fail both prongs of Step 2A, thus the claims are directed to the judicial exception as it has not been integrated into practical application, and fails Step 2B as not amounting to significantly more. Therefore, Claims 2 and 14 do not recite patent eligible subject matter under 35 U.S.C. § 101.
Claims 3, 15 and 19: “the managing of the first group of worker nodes comprises loading, by the head node, a portion of a group of trained machine learning model artifacts on each of the first group of worker nodes” reciting insignificant extra-solution data gathering activity, MPEP § 2106.05(g). With regard to integration into practical application and whether additional elements amount to significantly more, Claims 3, 15 and 19 fail both prongs of Step 2A, thus the claims are directed to the judicial exception as it has not been integrated into practical application, and fails Step 2B as not amounting to significantly more, performing a well understood, routine, and conventional task of data gathering. “The courts have recognized the following computer functions as well‐understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity. i. Receiving or transmitting data over a network … iv. Storing and retrieving information in memory”. See MPEP § 2106.05(d)(II). Therefore, Claims 3, 15 and 19 do not recite patent eligible subject matter under 35 U.S.C. § 101.
Claims 4, 16 and 20: “the managing of the first group of worker nodes comprises loading, by the head node, inference code on each of the first group of worker nodes” reciting insignificant extra-solution data gathering activity, MPEP § 2106.05(g). With regard to integration into practical application and whether additional elements amount to significantly more, Claims 4, 16 and 20 fail both prongs of Step 2A, thus the claims are directed to the judicial exception as it has not been integrated into practical application, and fails Step 2B as not amounting to significantly more, performing a well understood, routine, and conventional task of data gathering. “The courts have recognized the following computer functions as well‐understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity. i. Receiving or transmitting data over a network … iv. Storing and retrieving information in memory”. See MPEP § 2106.05(d)(II). Therefore, Claims 4, 16 and 20 do not recite patent eligible subject matter under 35 U.S.C. § 101.
Claim 5: “the operations comprise obtaining an adjusted processing capacity associated with the machine learning model and obtaining an adjusted memory capacity associated with the machine learning model”, reciting insignificant extra-solution data gathering activity, MPEP § 2106.05(g). With regard to integration into practical application and whether additional elements amount to significantly more, Claim 5 fail both prongs of Step 2A, thus the claims are directed to the judicial exception as it has not been integrated into practical application, and fails Step 2B as not amounting to significantly more, performing a well understood, routine, and conventional task of data gathering. “The courts have recognized the following computer functions as well‐understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity. i. Receiving or transmitting data over a network”. See MPEP § 2106.05(d)(II). Therefore, Claim 5 does not recite patent eligible subject matter under 35 U.S.C. § 101.
Claim 6: “the operations comprise provisioning a second group of worker nodes based on the adjusted processing capacity, the adjusted memory capacity, the processing capacity threshold, and the memory capacity threshold” as drafted, is a process that, but for the recitation of generic computing components, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, but for the recitation of generic computing components being used as a tool to perform the functionality, a person can mentally evaluate what resources are necessary for the completion of a job. With regard to integration into practical application and whether additional elements amount to significantly more, Claim 6 fail both prongs of Step 2A, thus the claims are directed to the judicial exception as it has not been integrated into practical application, and fails Step 2B as not amounting to significantly more. Therefore, Claim 6 does not recite patent eligible subject matter under 35 U.S.C. § 101.
Claim 7: “provisioning a second portion of data engineering pipeline on each of the second group of worker nodes; and provisioning a second portion of the machine learning model on each of the second group of worker nodes”, as drafted, is a process that, but for the recitation of generic computing components, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, but for the recitation of generic computing components being used as a tool to perform the functionality, a person can mentally evaluate what portions of resources are necessary for the completion of a portion job. With regard to integration into practical application and whether additional elements amount to significantly more, Claim 7 fail both prongs of Step 2A, thus the claims are directed to the judicial exception as it has not been integrated into practical application, and fails Step 2B as not amounting to significantly more. Therefore, Claim 7 does not recite patent eligible subject matter under 35 U.S.C. § 101.
Claims 8 and 17: “each of the first group of worker nodes is implemented by a virtual machine”, recites a field of use which generally links the use of a judicial exception to a particular technological environment (MPEP § 2106.05(h)). With regard to integration into practical application and whether additional elements amount to significantly more, Claims 8 and 17 fail both prongs of Step 2A, thus the claims are directed to the judicial exception as it has not been integrated into practical application, and fails Step 2B as not amounting to significantly more. Therefore, Claims 8 and 17 do not recite patent eligible subject matter under 35 U.S.C. § 101.
Claims 9 and 12: “the obtaining of the processing capacity threshold and obtaining the memory capacity threshold comprises receiving the processing capacity threshold via a first user-generated input and receiving the memory capacity threshold via a second user-generated input”, reciting insignificant extra-solution data gathering activity, MPEP § 2106.05(g). With regard to integration into practical application and whether additional elements amount to significantly more, Claims 9 and 12 fail both prongs of Step 2A, thus the claims are directed to the judicial exception as it has not been integrated into practical application, and fails Step 2B as not amounting to significantly more, performing a well understood, routine, and conventional task of data gathering. “The courts have recognized the following computer functions as well‐understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity. i. Receiving or transmitting data over a network”. See MPEP § 2106.05(d)(II). Therefore, Claims 9 and 12 do not recite patent eligible subject matter under 35 U.S.C. § 101.
Claims 11 and 13: “obtaining a processing capacity threshold and obtaining a memory capacity threshold”, reciting insignificant extra-solution data gathering activity, MPEP § 2106.05(g). Further, “the provisioning of the head node and the first group of worker nodes comprises provisioning the head node and the first group of worker nodes based on the processing capacity threshold and the memory capacity threshold”, as drafted, is a process that, but for the recitation of generic computing components, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, but for the recitation of generic computing components being used as a tool to perform the functionality, a person can mentally evaluate what resources are necessary for the completion of job based on that resources observed performance. With regard to integration into practical application and whether additional elements amount to significantly more, Claims 11 and 13 fail both prongs of Step 2A, thus the claims are directed to the judicial exception as it has not been integrated into practical application, and fails Step 2B as not amounting to significantly more, performing a well understood, routine, and conventional task of data gathering. “The courts have recognized the following computer functions as well‐understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity. i. Receiving or transmitting data over a network”. See MPEP § 2106.05(d)(II). Therefore, Claims 11 and 13 do not recite patent eligible subject matter under 35 U.S.C. § 101.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-20 are rejected under 35 U.S.C. 103(a) as being unpatentable over Sathe et al. (US 20220114019 A1) (hereinafter Sathe) in further view of Sidhu et al. (US 20200034710 A1) (hereinafter Sidhu).
Regarding Claim 1, Sathe teaches:
A device, comprising: a processing system including a processor; and a memory that stores executable instructions that, when executed by the processing system, facilitate performance of operations,
“Devices used herein may include one or more processors 02, one or more computer-readable RAMs 04, one or more computer-readable ROMs 06, one or more computer readable storage media 08”, (Sathe: ¶45), “Computer readable program instructions for carrying out operations of the present invention may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions” … “The computer readable program instructions may execute entirely on the user's computer” … “execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry”, (Sathe: ¶79).
provisioning a head node and a first group of worker nodes based on the processing capacity, the memory capacity, the processing capacity threshold, and the memory capacity threshold;
“According to the exemplary embodiments, the pipeline training system 100 may include one or more worker nodes 120A-K and a pipeline training server 130, which all may be interconnected via a network 108”, (Sathe: ¶21), “the one or more worker nodes 120A-K may each be an enterprise server, a laptop computer, a notebook, a tablet computer, a netbook computer, a personal computer (PC), a desktop computer, a server, a personal digital assistant (PDA), a rotary phone, a touchtone phone, a smart phone, a mobile phone, a virtual device, a thin client, an IoT device, or any other electronic device or computing system capable of sending and receiving data to and from other computing devices”, (Sathe: ¶23), “the pipeline training server 130 includes a joint optimizer 132, a performance predictor 134, and a load balancer 136” … “the pipeline training server 130 may be comprised of a cluster or plurality of computing devices”, (Sathe: ¶24), “The claimed system may be implemented in, for example, a Kubernetes and Docker platform wherein the one or more worker nodes 120A-K are Docker containers and heartbeat features can be obtained using kubectl” … “Moreover, the system can be scaled using AutoScaler or manually creating pods using the output of the ML/DL mode”, (Sathe: ¶42), “The joint optimizer 132 may select worker nodes (step 210). In embodiments, the joint optimizer 132 may select at least one of the one or more worker nodes 120A-K for executing the pipeline based on the predicted pipeline training resources required”, (Sathe: ¶37), “With reference to the previously introduced example, the joint optimizer 132 selects the worker node 120A to train the first pipeline and worker node 120B to train the second pipeline”, (Sathe: ¶38). Examiner notes: the pipeline training server 130 is interpreted as the head node, one or more worker nodes 120 initially selected are the first group of worker nodes, the pipeline training server 130 distributes work to worker nodes that may be implemented as Docker Containers which is being interpreted as the provisioning.
provisioning a first portion of data engineering pipeline on each of the first group of worker nodes; and provisioning a first portion of the machine learning model on each of the first group of worker nodes.
“A machine learning pipeline is a series of operations (such as data preprocessing, outlier detection, feature engineering, etc.) followed by an estimator”, (Sathe: ¶17), “Each of the one or more worker nodes 120A-K may be configured to train one or more machine learning pipelines. In the example embodiment, it is assumed that each of the one or more worker nodes 120A-K have access to a same dataset and each pipeline can be trained on a single worker node 120 of the one or more worker nodes 120A-K”, (Sathe: ¶23), “the pipeline features may include type of estimator, type of pre-processor, type of feature engineering, and parameter settings thereof…”, (Sathe: ¶31), “The joint optimizer 132 may select worker nodes (step 210). In embodiments, the joint optimizer 132 may select at least one of the one or more worker nodes 120A-K for executing the pipeline based on the predicted pipeline training resources required”, (Sathe: ¶37), “the joint optimizer 132 is configured to train two pipelines: 1) principal component analysis (PCA) to random forest (RF); and 2) outlier detection (OD) to support vector machine (SVM), on any of four worker nodes 120A, 120B, 120C, and 120D”, (Sathe: ¶30), “the joint optimizer 132 may predict required performance measures for each of the one or more worker nodes 120A-K to train a respective pipeline via the performance predictor 134”, (Sathe: ¶35).
Further regarding Claim 1, Sathe fails to teach:
the operations comprising: obtaining a processing capacity associated with a machine learning model and obtaining a memory capacity associated with the machine learning model; obtaining a processing capacity threshold and obtaining a memory capacity threshold; provisioning a head node and a first group of worker nodes based on the processing capacity, the memory capacity, the processing capacity threshold, and the memory capacity threshold;
However, Sidhu teaches: “the performance evaluator estimates the naïve FLOPS, naïve memory allocation, and naïve memory bandwidth”, (Sidhu: ¶93), “The performance evaluator 280 maththematically [sic] determines a number of naïve floating point operations (FLOPS) the total number of FLOPS used by the model when implemented with default kernels. The naïve FLOPS are estimated using the model description and model parameters generated by the model generator”, (Sidhu: ¶63), “the performance evaluator 280 determines a number of operations to be performed and compares the determined number of operations to a maximum number of operations per second the embedded processor 270 is capable of performing” … “a GPU has 1.8 TFLOPS (or 1.8×10.sup.12 floating point operations per second) of computing capability, and the model is performed using 20×10.sup.9 floating point operations”, (Sidhu: ¶61), “The performance evaluator 280 mathematically determines a naïve memory allocation as the total memory used by the model for all the model-parameters and temporary variables”, (Sidhu: ¶65), “The performance evaluator 280 mathematically determines a naïve memory bandwidth as the total memory bandwidth used by the model for all the model-parameters and temporary variables”, (Sidhu: ¶67), “The performance evaluator 280 determines an optimized memory allocation as the amount of memory used by the model after the allocation of the model-parameters and temporary variables have been scheduled…”, (Sidhu: ¶66), “The performance evaluator 280 empirically determines an optimized memory bandwidth as the memory bandwidth used by the model after the allocation of the model-parameters and temporary variables, as well as the operations executed by the model, have been scheduled…”, (Sidhu: ¶68), “the tensor scheduler 340 receives a maximum amount of memory available in the target platform system”, (Sidhu: ¶73), “If the estimated performance on any of these 3 metrics is lower than a specified performance, the process advances to step 860, where a new model is generated by the model generator 210 based on the performance of the previous model”, (Sidhu: ¶93).
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine the operations comprising: obtaining a processing capacity associated with a machine learning model and obtaining a memory capacity associated with the machine learning model; obtaining a processing capacity threshold and obtaining a memory capacity threshold; provisioning a head node and a first group of worker nodes based on the processing capacity, the memory capacity, the processing capacity threshold, and the memory capacity threshold of Sidhu with the methods and systems of Sathe resulting in system that can adjust its evaluation of resource usage limits. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “may determine which operations not to schedule concurrently so as to reduce the amount of memory that is being used concurrently at any given point in time.”, (Sidhu: ¶73), “optimized memory bandwidth as the memory bandwidth used by the model after the allocation of the model-parameters and temporary variables, as well as the operations executed by the model, have been scheduled”, (Sidhu: ¶68), “the system is able to test multiple models without having to wait for the model to be trained, thus reducing the amount of computing power and time used to test and select a model for using in the target system.”, (Sidhu: ¶94).
Regarding Claim 2, Sathe teaches:
managing, by the head node, the first group of worker nodes.
“the pipeline training system 100 may include one or more worker nodes 120A-K and a pipeline training server 130, which all may be interconnected via a network 108”, (Sathe: ¶21), “the pipeline training server 130 includes a joint optimizer 132, a performance predictor 134, and a load balancer 136”, (Sathe: ¶24), “the joint optimizer 132 fails to receive heartbeat features from any of the one or more worker nodes 120A-K, the joint optimizer 132 marks the one or more unresponsive worker nodes 120A-K as unresponsive and omits training predictions therefor until heartbeat features are again received” … “the joint optimizer 132 may train a model for determining which of the worker nodes 120A-K can train the pipeline in a least amount of resources based on the heartbeat features collected herein, along with pipeline features and dataset features described below”, (Sathe: ¶29), “the joint optimizer 132 may select at least one of the one or more worker nodes 120A-K for executing the pipeline based on the predicted pipeline training resources required” … “he joint optimizer 132 may select the one or more worker nodes 120A-K based on a ε-Greedy or Multi Arm Bandit problem approach…”, (Sathe: ¶37).
Regarding Claim 5, Sathe fails to teach:
obtaining an adjusted processing capacity associated with the machine learning model and obtaining an adjusted memory capacity associated with the machine learning model.
However Sidhu teaches: “The performance evaluator 280 maththematically [sic] determines a number of naïve floating point operations (FLOPS)”, (Sidhu: ¶63), “The performance evaluator 280 uses static analysis to determine a number of optimized FLOPS as the number of FLOPS used by the model based on the machine code 245 generated by the virtual machine 240”, (Sidhu: ¶64), “The performance evaluator 280 mathematically determines a naïve memory allocation as the total memory used by the model for all the model-parameters and temporary variables”, (Sidhu: ¶65), “The performance evaluator 280 determines an optimized memory allocation as the amount of memory used by the model after the allocation of the model-parameters and temporary variables have been scheduled” … “The optimized memory allocation is lower than the naïve memory allocation, for example, when the memory to store temporary variables are re-allocated once the temporary variables are no longer needed”, (Sidhu: ¶66), “The performance evaluator 280 empirically determines an optimized memory bandwidth as the memory bandwidth used by the model after the allocation of the model-parameters and temporary variables, as well as the operations executed by the model, have been scheduled”, (Sidhu: ¶68). Examiner notes: The optimized FLOPS and memory allocation is interpreted as the adjusted capacities.
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine obtaining an adjusted processing capacity associated with the machine learning model and obtaining an adjusted memory capacity associated with the machine learning model of Sidhu with the methods and systems of Sathe resulting in system that can adjust its evaluations of resource usage limits. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “may determine which operations not to schedule concurrently so as to reduce the amount of memory that is being used concurrently at any given point in time.”, (Sidhu: ¶73), “optimized memory bandwidth as the memory bandwidth used by the model after the allocation of the model-parameters and temporary variables, as well as the operations executed by the model, have been scheduled”, (Sidhu: ¶68), “the system is able to test multiple models without having to wait for the model to be trained, thus reducing the amount of computing power and time used to test and select a model for using in the target system.”, (Sidhu: ¶94).
Regarding Claim 6, Sathe teaches:
provisioning a second group of worker nodes based on the adjusted processing capacity, the adjusted memory capacity, the processing capacity threshold, and the memory capacity threshold.
“In a Multi-Armed Bandit approach, the joint optimizer 132 may train the model by first selecting three random workers and executing a pipeline for n iterations, e.g., n=1000” … “If the joint optimizer 132 determines that the performance of the best performing working node 120A-K degrades as a result, the joint optimizer 132 may then revert back to randomly identifying a best performing worker node 120A-K and repeating the process”, (Sathe: ¶37). Examiner notes: system picks a group of nodes and assess performance and then picks a different based on that assessment creating distinct provision groups.
Further regarding Claim 6, Sathe fails to teach:
provisioning a second group of worker nodes based on the adjusted processing capacity, the adjusted memory capacity, the processing capacity threshold, and the memory capacity threshold.
However Sidhu teaches: “The performance evaluator 280 maththematically [sic] determines a number of naïve floating point operations (FLOPS)”, (Sidhu: ¶63), “The performance evaluator 280 uses static analysis to determine a number of optimized FLOPS as the number of FLOPS used by the model based on the machine code 245 generated by the virtual machine 240”, (Sidhu: ¶64), “The performance evaluator 280 mathematically determines a naïve memory allocation as the total memory used by the model for all the model-parameters and temporary variables”, (Sidhu: ¶65), “The performance evaluator 280 determines an optimized memory allocation as the amount of memory used by the model after the allocation of the model-parameters and temporary variables have been scheduled” … “The optimized memory allocation is lower than the naïve memory allocation, for example, when the memory to store temporary variables are re-allocated once the temporary variables are no longer needed”, (Sidhu: ¶66), “The performance evaluator 280 empirically determines an optimized memory bandwidth as the memory bandwidth used by the model after the allocation of the model-parameters and temporary variables, as well as the operations executed by the model, have been scheduled”, (Sidhu: ¶68). Examiner notes: The optimized FLOPS and memory allocation is interpreted as the adjusted capacities.
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine obtaining an adjusted processing capacity associated with the machine learning model and obtaining an adjusted memory capacity associated with the machine learning model of Sidhu with the methods and systems of Sathe resulting in system that can adjust its evaluations of resource usage limits. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “may determine which operations not to schedule concurrently so as to reduce the amount of memory that is being used concurrently at any given point in time.”, (Sidhu: ¶73), “optimized memory bandwidth as the memory bandwidth used by the model after the allocation of the model-parameters and temporary variables, as well as the operations executed by the model, have been scheduled”, (Sidhu: ¶68), “the system is able to test multiple models without having to wait for the model to be trained, thus reducing the amount of computing power and time used to test and select a model for using in the target system.”, (Sidhu: ¶94).
Regarding Claim 7, Sathe teaches:
provisioning a second portion of data engineering pipeline on each of the second group of worker nodes; and provisioning a second portion of the machine learning model on each of the second group of worker nodes.
“A machine learning pipeline is a series of operations (such as data preprocessing, outlier detection, feature engineering, etc.) followed by an estimator”, (Sathe: ¶17), “Each of the one or more worker nodes 120A-K may be configured to train one or more machine learning pipelines. In the example embodiment, it is assumed that each of the one or more worker nodes 120A-K have access to a same dataset and each pipeline can be trained on a single worker node 120 of the one or more worker nodes 120A-K”, (Sathe: ¶23), “the pipeline features may include type of estimator, type of pre-processor, type of feature engineering, and parameter settings thereof…”, (Sathe: ¶31), “The joint optimizer 132 may select worker nodes (step 210). In embodiments, the joint optimizer 132 may select at least one of the one or more worker nodes 120A-K for executing the pipeline based on the predicted pipeline training resources required” … “In a Multi-Armed Bandit approach, the joint optimizer 132 may train the model by first selecting three random workers and executing a pipeline for n iterations, e.g., n=1000” … “If the joint optimizer 132 determines that the performance of the best performing working node 120A-K degrades as a result, the joint optimizer 132 may then revert back to randomly identifying a best performing worker node 120A-K and repeating the process”, (Sathe: ¶37), “the joint optimizer 132 is configured to train two pipelines: 1) principal component analysis (PCA) to random forest (RF); and 2) outlier detection (OD) to support vector machine (SVM), on any of four worker nodes 120A, 120B, 120C, and 120D”, (Sathe: ¶30), “the joint optimizer 132 may predict required performance measures for each of the one or more worker nodes 120A-K to train a respective pipeline via the performance predictor 134”, (Sathe: ¶35).
Regarding Claim 8, Sathe teaches:
each of the first group of worker nodes is implemented by a virtual machine.
“the one or more worker nodes 120A-K may each be an” … “a virtual device, a thin client, an IoT device, or any other electronic device or computing system capable of sending and receiving data to and from other computing devices” … “The one or more worker nodes 120A-K are described in greater detail as a hardware implementation with reference to FIG. 4, as part of a cloud implementation”, (Sathe: ¶23, Fig 4), “Cloud computing is a model of service delivery for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services)”, (Sathe: ¶53).
Regarding Claim 10, Sathe teaches:
A non-transitory machine-readable medium, comprising executable instructions that, when executed by a processing system including a processor, facilitate performance of operations,
“Devices used herein may include one or more processors 02, one or more computer-readable RAMs 04, one or more computer-readable ROMs 06, one or more computer readable storage media 08”, (Sathe: ¶45), “Computer readable program instructions for carrying out operations of the present invention may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions” … “The computer readable program instructions may execute entirely on the user's computer” … “execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry”, (Sathe: ¶79).
provisioning a head node and a first group of worker nodes based on the processing capacity, and the memory capacity;
“According to the exemplary embodiments, the pipeline training system 100 may include one or more worker nodes 120A-K and a pipeline training server 130, which all may be interconnected via a network 108”, (Sathe: ¶21), “the one or more worker nodes 120A-K may each be an enterprise server, a laptop computer, a notebook, a tablet computer, a netbook computer, a personal computer (PC), a desktop computer, a server, a personal digital assistant (PDA), a rotary phone, a touchtone phone, a smart phone, a mobile phone, a virtual device, a thin client, an IoT device, or any other electronic device or computing system capable of sending and receiving data to and from other computing devices”, (Sathe: ¶23), “the pipeline training server 130 includes a joint optimizer 132, a performance predictor 134, and a load balancer 136” … “the pipeline training server 130 may be comprised of a cluster or plurality of computing devices”, (Sathe: ¶24), “The claimed system may be implemented in, for example, a Kubernetes and Docker platform wherein the one or more worker nodes 120A-K are Docker containers and heartbeat features can be obtained using kubectl” … “Moreover, the system can be scaled using AutoScaler or manually creating pods using the output of the ML/DL mode”, (Sathe: ¶42), “The joint optimizer 132 may select worker nodes (step 210). In embodiments, the joint optimizer 132 may select at least one of the one or more worker nodes 120A-K for executing the pipeline based on the predicted pipeline training resources required”, (Sathe: ¶37), “With reference to the previously introduced example, the joint optimizer 132 selects the worker node 120A to train the first pipeline and worker node 120B to train the second pipeline”, (Sathe: ¶38). Examiner notes: the pipeline training server 130 is interpreted as the head node, one or more worker nodes 120 initially selected are the first group of worker nodes, the pipeline training server 130 distributes work to worker nodes that may be implemented as Docker Containers which is being interpreted as the provisioning.
provisioning a first portion of data engineering pipeline on each of the first group of worker nodes; and provisioning a first portion of the machine learning model on each of the first group of worker nodes.
“A machine learning pipeline is a series of operations (such as data preprocessing, outlier detection, feature engineering, etc.) followed by an estimator”, (Sathe: ¶17), “Each of the one or more worker nodes 120A-K may be configured to train one or more machine learning pipelines. In the example embodiment, it is assumed that each of the one or more worker nodes 120A-K have access to a same dataset and each pipeline can be trained on a single worker node 120 of the one or more worker nodes 120A-K”, (Sathe: ¶23), “the pipeline features may include type of estimator, type of pre-processor, type of feature engineering, and parameter settings thereof…”, (Sathe: ¶31), “The joint optimizer 132 may select worker nodes (step 210). In embodiments, the joint optimizer 132 may select at least one of the one or more worker nodes 120A-K for executing the pipeline based on the predicted pipeline training resources required”, (Sathe: ¶37), “the joint optimizer 132 is configured to train two pipelines: 1) principal component analysis (PCA) to random forest (RF); and 2) outlier detection (OD) to support vector machine (SVM), on any of four worker nodes 120A, 120B, 120C, and 120D”, (Sathe: ¶30), “the joint optimizer 132 may predict required performance measures for each of the one or more worker nodes 120A-K to train a respective pipeline via the performance predictor 134”, (Sathe: ¶35).
Further regarding Claim 10, Sathe fails to teach:
the operations comprising: obtaining a processing capacity associated with a machine learning model and obtaining a memory capacity associated with the machine learning model; provisioning a head node and a first group of worker nodes based on the processing capacity, and the memory capacity;
However, Sidhu teaches: “the performance evaluator estimates the naïve FLOPS, naïve memory allocation, and naïve memory bandwidth”, (Sidhu: ¶93), “The performance evaluator 280 maththematically [sic] determines a number of naïve floating point operations (FLOPS) the total number of FLOPS used by the model when implemented with default kernels. The naïve FLOPS are estimated using the model description and model parameters generated by the model generator”, (Sidhu: ¶63), “the performance evaluator 280 determines a number of operations to be performed and compares the determined number of operations to a maximum number of operations per second the embedded processor 270 is capable of performing” … “a GPU has 1.8 TFLOPS (or 1.8×10.sup.12 floating point operations per second) of computing capability, and the model is performed using 20×10.sup.9 floating point operations”, (Sidhu: ¶61), “The performance evaluator 280 mathematically determines a naïve memory allocation as the total memory used by the model for all the model-parameters and temporary variables”, (Sidhu: ¶65), “The performance evaluator 280 mathematically determines a naïve memory bandwidth as the total memory bandwidth used by the model for all the model-parameters and temporary variables”, (Sidhu: ¶67), “The performance evaluator 280 determines an optimized memory allocation as the amount of memory used by the model after the allocation of the model-parameters and temporary variables have been scheduled…”, (Sidhu: ¶66), “The performance evaluator 280 empirically determines an optimized memory bandwidth as the memory bandwidth used by the model after the allocation of the model-parameters and temporary variables, as well as the operations executed by the model, have been scheduled…”, (Sidhu: ¶68), “the tensor scheduler 340 receives a maximum amount of memory available in the target platform system”, (Sidhu: ¶73), “If the estimated performance on any of these 3 metrics is lower than a specified performance, the process advances to step 860, where a new model is generated by the model generator 210 based on the performance of the previous model”, (Sidhu: ¶93).
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine the operations comprising: obtaining a processing capacity associated with a machine learning model and obtaining a memory capacity associated with the machine learning model; provisioning a head node and a first group of worker nodes based on the processing capacity, and the memory capacity of Sidhu with the methods and systems of Sathe resulting in system that can adjust its evaluation of resource usage limits. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “may determine which operations not to schedule concurrently so as to reduce the amount of memory that is being used concurrently at any given point in time.”, (Sidhu: ¶73), “optimized memory bandwidth as the memory bandwidth used by the model after the allocation of the model-parameters and temporary variables, as well as the operations executed by the model, have been scheduled”, (Sidhu: ¶68), “the system is able to test multiple models without having to wait for the model to be trained, thus reducing the amount of computing power and time used to test and select a model for using in the target system.”, (Sidhu: ¶94).
Regarding Claim 11, Sathe fails to teach:
obtaining a processing capacity threshold and obtaining a memory capacity threshold.
However, Sidhu teaches: “the performance evaluator estimates the naïve FLOPS, naïve memory allocation, and naïve memory bandwidth”, (Sidhu: ¶93), “The performance evaluator 280 maththematically [sic] determines a number of naïve floating point operations (FLOPS) the total number of FLOPS used by the model when implemented with default kernels. The naïve FLOPS are estimated using the model description and model parameters generated by the model generator”, (Sidhu: ¶63), “the performance evaluator 280 determines a number of operations to be performed and compares the determined number of operations to a maximum number of operations per second the embedded processor 270 is capable of performing” … “a GPU has 1.8 TFLOPS (or 1.8×10.sup.12 floating point operations per second) of computing capability, and the model is performed using 20×10.sup.9 floating point operations”, (Sidhu: ¶61), “The performance evaluator 280 mathematically determines a naïve memory allocation as the total memory used by the model for all the model-parameters and temporary variables”, (Sidhu: ¶65), “The performance evaluator 280 mathematically determines a naïve memory bandwidth as the total memory bandwidth used by the model for all the model-parameters and temporary variables”, (Sidhu: ¶67), “The performance evaluator 280 determines an optimized memory allocation as the amount of memory used by the model after the allocation of the model-parameters and temporary variables have been scheduled…”, (Sidhu: ¶66), “The performance evaluator 280 empirically determines an optimized memory bandwidth as the memory bandwidth used by the model after the allocation of the model-parameters and temporary variables, as well as the operations executed by the model, have been scheduled…”, (Sidhu: ¶68), “the tensor scheduler 340 receives a maximum amount of memory available in the target platform system”, (Sidhu: ¶73), “If the estimated performance on any of these 3 metrics is lower than a specified performance, the process advances to step 860, where a new model is generated by the model generator 210 based on the performance of the previous model”, (Sidhu: ¶93).
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine obtaining a processing capacity threshold and obtaining a memory capacity threshold of Sidhu with the methods and systems of Sathe resulting in system that can adjust its evaluation of resource usage limits. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “may determine which operations not to schedule concurrently so as to reduce the amount of memory that is being used concurrently at any given point in time.”, (Sidhu: ¶73), “optimized memory bandwidth as the memory bandwidth used by the model after the allocation of the model-parameters and temporary variables, as well as the operations executed by the model, have been scheduled”, (Sidhu: ¶68), “the system is able to test multiple models without having to wait for the model to be trained, thus reducing the amount of computing power and time used to test and select a model for using in the target system.”, (Sidhu: ¶94).
Regarding Claim 13, Sathe teaches:
provisioning the head node and the first group of worker nodes based on the processing capacity threshold and the memory capacity threshold.
“According to the exemplary embodiments, the pipeline training system 100 may include one or more worker nodes 120A-K and a pipeline training server 130, which all may be interconnected via a network 108”, (Sathe: ¶21), “the one or more worker nodes 120A-K may each be an enterprise server, a laptop computer, a notebook, a tablet computer, a netbook computer, a personal computer (PC), a desktop computer, a server, a personal digital assistant (PDA), a rotary phone, a touchtone phone, a smart phone, a mobile phone, a virtual device, a thin client, an IoT device, or any other electronic device or computing system capable of sending and receiving data to and from other computing devices”, (Sathe: ¶23), “the pipeline training server 130 includes a joint optimizer 132, a performance predictor 134, and a load balancer 136” … “the pipeline training server 130 may be comprised of a cluster or plurality of computing devices”, (Sathe: ¶24), “The claimed system may be implemented in, for example, a Kubernetes and Docker platform wherein the one or more worker nodes 120A-K are Docker containers and heartbeat features can be obtained using kubectl” … “Moreover, the system can be scaled using AutoScaler or manually creating pods using the output of the ML/DL mode”, (Sathe: ¶42), “The joint optimizer 132 may select worker nodes (step 210). In embodiments, the joint optimizer 132 may select at least one of the one or more worker nodes 120A-K for executing the pipeline based on the predicted pipeline training resources required”, (Sathe: ¶37), “With reference to the previously introduced example, the joint optimizer 132 selects the worker node 120A to train the first pipeline and worker node 120B to train the second pipeline”, (Sathe: ¶38). Examiner notes: the pipeline training server 130 is interpreted as the head node, one or more worker nodes 120 initially selected are the first group of worker nodes, the pipeline training server 130 distributes work to worker nodes that may be implemented as Docker Containers which is being interpreted as the provisioning.
Further regarding Claim 13, Sathe fails to teach:
provisioning the head node and the first group of worker nodes based on the processing capacity threshold and the memory capacity threshold.
However, Sidhu teaches: “the performance evaluator estimates the naïve FLOPS, naïve memory allocation, and naïve memory bandwidth”, (Sidhu: ¶93), “The performance evaluator 280 maththematically [sic] determines a number of naïve floating point operations (FLOPS) the total number of FLOPS used by the model when implemented with default kernels. The naïve FLOPS are estimated using the model description and model parameters generated by the model generator”, (Sidhu: ¶63), “the performance evaluator 280 determines a number of operations to be performed and compares the determined number of operations to a maximum number of operations per second the embedded processor 270 is capable of performing” … “a GPU has 1.8 TFLOPS (or 1.8×10.sup.12 floating point operations per second) of computing capability, and the model is performed using 20×10.sup.9 floating point operations”, (Sidhu: ¶61), “The performance evaluator 280 mathematically determines a naïve memory allocation as the total memory used by the model for all the model-parameters and temporary variables”, (Sidhu: ¶65), “The performance evaluator 280 mathematically determines a naïve memory bandwidth as the total memory bandwidth used by the model for all the model-parameters and temporary variables”, (Sidhu: ¶67), “The performance evaluator 280 determines an optimized memory allocation as the amount of memory used by the model after the allocation of the model-parameters and temporary variables have been scheduled…”, (Sidhu: ¶66), “The performance evaluator 280 empirically determines an optimized memory bandwidth as the memory bandwidth used by the model after the allocation of the model-parameters and temporary variables, as well as the operations executed by the model, have been scheduled…”, (Sidhu: ¶68), “the tensor scheduler 340 receives a maximum amount of memory available in the target platform system”, (Sidhu: ¶73), “If the estimated performance on any of these 3 metrics is lower than a specified performance, the process advances to step 860, where a new model is generated by the model generator 210 based on the performance of the previous model”, (Sidhu: ¶93).
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine provisioning the head node and the first group of worker nodes based on the processing capacity threshold and the memory capacity threshold of Sidhu with the methods and systems of Sathe resulting in system that can adjust its evaluation of resource usage limits. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “may determine which operations not to schedule concurrently so as to reduce the amount of memory that is being used concurrently at any given point in time.”, (Sidhu: ¶73), “optimized memory bandwidth as the memory bandwidth used by the model after the allocation of the model-parameters and temporary variables, as well as the operations executed by the model, have been scheduled”, (Sidhu: ¶68), “the system is able to test multiple models without having to wait for the model to be trained, thus reducing the amount of computing power and time used to test and select a model for using in the target system.”, (Sidhu: ¶94).
Regarding Claim 14, Sathe teaches:
managing, by the head node, the first group of worker nodes.
“the pipeline training system 100 may include one or more worker nodes 120A-K and a pipeline training server 130, which all may be interconnected via a network 108”, (Sathe: ¶21), “the pipeline training server 130 includes a joint optimizer 132, a performance predictor 134, and a load balancer 136”, (Sathe: ¶24), “the joint optimizer 132 fails to receive heartbeat features from any of the one or more worker nodes 120A-K, the joint optimizer 132 marks the one or more unresponsive worker nodes 120A-K as unresponsive and omits training predictions therefor until heartbeat features are again received” … “the joint optimizer 132 may train a model for determining which of the worker nodes 120A-K can train the pipeline in a least amount of resources based on the heartbeat features collected herein, along with pipeline features and dataset features described below”, (Sathe: ¶29), “the joint optimizer 132 may select at least one of the one or more worker nodes 120A-K for executing the pipeline based on the predicted pipeline training resources required” … “he joint optimizer 132 may select the one or more worker nodes 120A-K based on a ε-Greedy or Multi Arm Bandit problem approach…”, (Sathe: ¶37).
Regarding Claim 17, Sathe teaches:
each of the first group of worker nodes is implemented by a virtual machine.
“the one or more worker nodes 120A-K may each be an” … “a virtual device, a thin client, an IoT device, or any other electronic device or computing system capable of sending and receiving data to and from other computing devices” … “The one or more worker nodes 120A-K are described in greater detail as a hardware implementation with reference to FIG. 4, as part of a cloud implementation”, (Sathe: ¶23, Fig 4), “Cloud computing is a model of service delivery for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services)”, (Sathe: ¶53).
Claims 3-4, 9, 12, 15-16 and 18-20 are rejected under 35 U.S.C. 103(a) as being unpatentable over Sathe in view of Sidhu, in further view of Mueller et al. (US 20210326717 A1) (hereinafter Mueller)
Regarding Claim 3, Sathe in view of Sidhu fails to teach:
loading, by the head node, a portion of a group of trained machine learning model artifacts on each of the first group of worker nodes.
However, Meuller teaches: “the CIVIL service 103 (e.g., ML orchestrator 115) or AMPGS 102 may send one or more commands (e.g., API calls) to a model hosting system 140 described further herein to “host” the pipeline—e.g., launch or reserve one or more compute instances, run pipeline code 126, configure endpoints associated with the pipeline” … “the commands may include a “create model” API call that combines code for the model (e.g., inference code implemented within a container) along with model artifacts (e.g., data describing weights associated with various aspects of the model)”, (Mueller: ¶73), “the model hosting system 140 initializes ones or more ML scoring containers 1050 in one or more hosted virtual machine instance 1042” … “the model hosting system 140 forms the ML scoring container(s) 1050 from the identified container image(s)”, (Mueller: ¶123), “The model hosting system 140 further forms the ML scoring container(s) 1050 by retrieving model data corresponding to the identified trained machine learning model(s)”, (Mueller: ¶124), “The model hosting system 140 can insert the model data files into the same ML scoring container 1050, into different ML scoring containers 1050 initialized in the same virtual machine instance 1042, or into different ML scoring containers 1050 initialized in different virtual machine instances 1042”, (Mueller: ¶125), “The ML scoring containers 1050 each include a runtime 1054, code 1056, and dependencies 1052” … “The code 1056 can also include model data that represent characteristics of the defined machine learning model” … “the code 1056 results in the generation of outputs (e.g., predicted or “inferred” results”, (Mueller: ¶119).
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine loading, by the head node, a portion of a group of trained machine learning model artifacts on each of the first group of worker nodes of Mueller with the methods and systems of Sathe in view of Sidhu resulting in a system that can load ML artifacts on to nodes. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “By parallelizing the training process, the model training system 120 can significantly reduce the training time”, (Mueller: ¶109).
Regarding Claim 4, Sathe in view of Sidhu fails to teach:
loading, by the head node, inference code on each of the first group of worker nodes.
However, Meuller teaches: “the CIVIL service 103 (e.g., ML orchestrator 115) or AMPGS 102 may send one or more commands (e.g., API calls) to a model hosting system 140 described further herein to “host” the pipeline—e.g., launch or reserve one or more compute instances, run pipeline code 126, configure endpoints associated with the pipeline” … “the commands may include a “create model” API call that combines code for the model (e.g., inference code implemented within a container) along with model artifacts (e.g., data describing weights associated with various aspects of the model)”, (Mueller: ¶73), “the model hosting system 140 initializes ones or more ML scoring containers 1050 in one or more hosted virtual machine instance 1042” … “the model hosting system 140 forms the ML scoring container(s) 1050 from the identified container image(s)”, (Mueller: ¶123), “The model hosting system 140 further forms the ML scoring container(s) 1050 by retrieving model data corresponding to the identified trained machine learning model(s)”, (Mueller: ¶124), “The model hosting system 140 can insert the model data files into the same ML scoring container 1050, into different ML scoring containers 1050 initialized in the same virtual machine instance 1042, or into different ML scoring containers 1050 initialized in different virtual machine instances 1042”, (Mueller: ¶125), “The ML scoring containers 1050 each include a runtime 1054, code 1056, and dependencies 1052” … “The code 1056 can also include model data that represent characteristics of the defined machine learning model” … “the code 1056 results in the generation of outputs (e.g., predicted or “inferred” results”, (Mueller: ¶119).
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine comprises loading, by the head node, a portion of a group of trained machine learning model artifacts on each of the first group of worker nodes of Mueller with the methods and systems of Sathe in view of Sidhu resulting in a system that can load inference code on to nodes. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “By parallelizing the training process, the model training system 120 can significantly reduce the training time”, (Mueller: ¶109).
Regarding Claim 9, Sathe in view of Sidhu fails to teach:
receiving the processing capacity threshold via a first user-generated input and receiving the memory capacity threshold via a second user-generated input.
However, Mueller teaches: “a user device 1002 can provide a training request to the frontend 1029” … “information describing the computing machine on which to train a machine learning model (e.g., a graphical processing unit (GPU) instance type, a central processing unit (CPU) instance type, an amount of memory to allocate, a type of virtual machine instance to use for training, etc.)”, (Mueller: ¶94), “settings can be exposed in this manner to users, including but not limited to resource utilization limits”, (Mueller: ¶45).
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine receiving the processing capacity threshold via a first user-generated input and receiving the memory capacity threshold via a second user-generated input of Mueller with the methods and systems of Sathe in view of Sidhu resulting in a system that allows users to define thresholds. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “thereby relieving the user from the burden of having to worry about over-utilization (e.g., acquiring too little computing resources and suffering performance issues) or under-utilization (e.g., acquiring more computing resources than necessary to train the machine learning models, and thus overpaying)”, (Mueller: ¶98).
Regarding Claim 12, Sathe in view of Sidhu fails to teach:
receiving the processing capacity threshold via a first user-generated input and receiving the memory capacity threshold via a second user-generated input.
However, Mueller teaches: “a user device 1002 can provide a training request to the frontend 1029” … “information describing the computing machine on which to train a machine learning model (e.g., a graphical processing unit (GPU) instance type, a central processing unit (CPU) instance type, an amount of memory to allocate, a type of virtual machine instance to use for training, etc.)”, (Mueller: ¶94), “settings can be exposed in this manner to users, including but not limited to resource utilization limits”, (Mueller: ¶45).
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine receiving the processing capacity threshold via a first user-generated input and receiving the memory capacity threshold via a second user-generated input of Mueller with the methods and systems of Sathe in view of Sidhu resulting in a system that allows users to define thresholds. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “thereby relieving the user from the burden of having to worry about over-utilization (e.g., acquiring too little computing resources and suffering performance issues) or under-utilization (e.g., acquiring more computing resources than necessary to train the machine learning models, and thus overpaying)”, (Mueller: ¶98).
Regarding Claim 15, Sathe in view of Sidhu fails to teach:
loading, by the head node, a portion of a group of trained machine learning model artifacts on each of the first group of worker nodes.
However, Meuller teaches: “the CIVIL service 103 (e.g., ML orchestrator 115) or AMPGS 102 may send one or more commands (e.g., API calls) to a model hosting system 140 described further herein to “host” the pipeline—e.g., launch or reserve one or more compute instances, run pipeline code 126, configure endpoints associated with the pipeline” … “the commands may include a “create model” API call that combines code for the model (e.g., inference code implemented within a container) along with model artifacts (e.g., data describing weights associated with various aspects of the model)”, (Mueller: ¶73), “the model hosting system 140 initializes ones or more ML scoring containers 1050 in one or more hosted virtual machine instance 1042” … “the model hosting system 140 forms the ML scoring container(s) 1050 from the identified container image(s)”, (Mueller: ¶123), “The model hosting system 140 further forms the ML scoring container(s) 1050 by retrieving model data corresponding to the identified trained machine learning model(s)”, (Mueller: ¶124), “The model hosting system 140 can insert the model data files into the same ML scoring container 1050, into different ML scoring containers 1050 initialized in the same virtual machine instance 1042, or into different ML scoring containers 1050 initialized in different virtual machine instances 1042”, (Mueller: ¶125), “The ML scoring containers 1050 each include a runtime 1054, code 1056, and dependencies 1052” … “The code 1056 can also include model data that represent characteristics of the defined machine learning model” … “the code 1056 results in the generation of outputs (e.g., predicted or “inferred” results”, (Mueller: ¶119).
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine loading, by the head node, a portion of a group of trained machine learning model artifacts on each of the first group of worker nodes of Mueller with the methods and systems of Sathe in view of Sidhu resulting in a system that can load ML artifacts on to nodes. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “By parallelizing the training process, the model training system 120 can significantly reduce the training time”, (Mueller: ¶109).
Regarding Claim 16, Sathe in view of Sidhu fails to teach:
loading, by the head node, inference code on each of the first group of worker nodes.
However, Meuller teaches: “the CIVIL service 103 (e.g., ML orchestrator 115) or AMPGS 102 may send one or more commands (e.g., API calls) to a model hosting system 140 described further herein to “host” the pipeline—e.g., launch or reserve one or more compute instances, run pipeline code 126, configure endpoints associated with the pipeline” … “the commands may include a “create model” API call that combines code for the model (e.g., inference code implemented within a container) along with model artifacts (e.g., data describing weights associated with various aspects of the model)”, (Mueller: ¶73), “the model hosting system 140 initializes ones or more ML scoring containers 1050 in one or more hosted virtual machine instance 1042” … “the model hosting system 140 forms the ML scoring container(s) 1050 from the identified container image(s)”, (Mueller: ¶123), “The model hosting system 140 further forms the ML scoring container(s) 1050 by retrieving model data corresponding to the identified trained machine learning model(s)”, (Mueller: ¶124), “The model hosting system 140 can insert the model data files into the same ML scoring container 1050, into different ML scoring containers 1050 initialized in the same virtual machine instance 1042, or into different ML scoring containers 1050 initialized in different virtual machine instances 1042”, (Mueller: ¶125), “The ML scoring containers 1050 each include a runtime 1054, code 1056, and dependencies 1052” … “The code 1056 can also include model data that represent characteristics of the defined machine learning model” … “the code 1056 results in the generation of outputs (e.g., predicted or “inferred” results”, (Mueller: ¶119).
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine loading, by the head node, a portion of a group of trained machine learning model artifacts on each of the first group of worker nodes of Mueller with the methods and systems of Sathe in view of Sidhu resulting in a system that can load inference code on to nodes. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “By parallelizing the training process, the model training system 120 can significantly reduce the training time”, (Mueller: ¶109).
Claims 18-20 are rejected under 35 U.S.C. 103(a) as being unpatentable over Sidhu in view of Sathe, in further view of Mueller.
Regarding Claim 18, Sidhu teaches:
A method, comprising: obtaining, by a processing system including a processor, a processing capacity associated with a machine learning model and obtaining, by the processing system, a memory capacity associated with the machine learning model; provisioning, by the processing system, a head node and a first group of worker nodes based on the processing capacity, the memory capacity, the processing capacity threshold, and the memory capacity threshold;
“2. The method of claim 1 wherein determining an amount of resources used by the model comprises: determining a number of floating point operations used by the untrained model when implemented with default kernels;”, (Sidhu: Claim 2), “4. The method of claim 1 wherein determining an amount of resources used by the model comprises: determining a total amount of memory used by the untrained model; and determining a total memory bandwidth used by the model”, (Sidhu: Claim 4), “the performance evaluator estimates the naïve FLOPS, naïve memory allocation, and naïve memory bandwidth”, (Sidhu: ¶93), “The performance evaluator 280 maththematically [sic] determines a number of naïve floating point operations (FLOPS) the total number of FLOPS used by the model when implemented with default kernels. The naïve FLOPS are estimated using the model description and model parameters generated by the model generator”, (Sidhu: ¶63), “the performance evaluator 280 determines a number of operations to be performed and compares the determined number of operations to a maximum number of operations per second the embedded processor 270 is capable of performing” … “a GPU has 1.8 TFLOPS (or 1.8×10.sup.12 floating point operations per second) of computing capability, and the model is performed using 20×10.sup.9 floating point operations”, (Sidhu: ¶61), “The performance evaluator 280 mathematically determines a naïve memory allocation as the total memory used by the model for all the model-parameters and temporary variables”, (Sidhu: ¶65), “The performance evaluator 280 mathematically determines a naïve memory bandwidth as the total memory bandwidth used by the model for all the model-parameters and temporary variables”, (Sidhu: ¶67), “The performance evaluator 280 determines an optimized memory allocation as the amount of memory used by the model after the allocation of the model-parameters and temporary variables have been scheduled…”, (Sidhu: ¶66), “The performance evaluator 280 empirically determines an optimized memory bandwidth as the memory bandwidth used by the model after the allocation of the model-parameters and temporary variables, as well as the operations executed by the model, have been scheduled…”, (Sidhu: ¶68), “the tensor scheduler 340 receives a maximum amount of memory available in the target platform system”, (Sidhu: ¶73), “If the estimated performance on any of these 3 metrics is lower than a specified performance, the process advances to step 860, where a new model is generated by the model generator 210 based on the performance of the previous model”, (Sidhu: ¶93).
Further regarding Claim 18, Sidhu fails to teach:
provisioning, by the processing system, a head node and a first group of worker nodes based on the processing capacity, the memory capacity, the processing capacity threshold, and the memory capacity threshold;
“According to the exemplary embodiments, the pipeline training system 100 may include one or more worker nodes 120A-K and a pipeline training server 130, which all may be interconnected via a network 108”, (Sathe: ¶21), “the one or more worker nodes 120A-K may each be an enterprise server, a laptop computer, a notebook, a tablet computer, a netbook computer, a personal computer (PC), a desktop computer, a server, a personal digital assistant (PDA), a rotary phone, a touchtone phone, a smart phone, a mobile phone, a virtual device, a thin client, an IoT device, or any other electronic device or computing system capable of sending and receiving data to and from other computing devices”, (Sathe: ¶23), “the pipeline training server 130 includes a joint optimizer 132, a performance predictor 134, and a load balancer 136” … “the pipeline training server 130 may be comprised of a cluster or plurality of computing devices”, (Sathe: ¶24), “The claimed system may be implemented in, for example, a Kubernetes and Docker platform wherein the one or more worker nodes 120A-K are Docker containers and heartbeat features can be obtained using kubectl” … “Moreover, the system can be scaled using AutoScaler or manually creating pods using the output of the ML/DL mode”, (Sathe: ¶42), “The joint optimizer 132 may select worker nodes (step 210). In embodiments, the joint optimizer 132 may select at least one of the one or more worker nodes 120A-K for executing the pipeline based on the predicted pipeline training resources required”, (Sathe: ¶37), “With reference to the previously introduced example, the joint optimizer 132 selects the worker node 120A to train the first pipeline and worker node 120B to train the second pipeline”, (Sathe: ¶38). Examiner notes: the pipeline training server 130 is interpreted as the head node, one or more worker nodes 120 initially selected are the first group of worker nodes, the pipeline training server 130 distributes work to worker nodes that may be implemented as Docker Containers which is being interpreted as the provisioning.
provisioning, by the processing system, a first portion of data engineering pipeline on each of the first group of worker nodes; and provisioning, by the processing system, a first portion of the machine learning model on each of the first group of worker nodes.
“A machine learning pipeline is a series of operations (such as data preprocessing, outlier detection, feature engineering, etc.) followed by an estimator”, (Sathe: ¶17), “Each of the one or more worker nodes 120A-K may be configured to train one or more machine learning pipelines. In the example embodiment, it is assumed that each of the one or more worker nodes 120A-K have access to a same dataset and each pipeline can be trained on a single worker node 120 of the one or more worker nodes 120A-K”, (Sathe: ¶23), “the pipeline features may include type of estimator, type of pre-processor, type of feature engineering, and parameter settings thereof…”, (Sathe: ¶31), “The joint optimizer 132 may select worker nodes (step 210). In embodiments, the joint optimizer 132 may select at least one of the one or more worker nodes 120A-K for executing the pipeline based on the predicted pipeline training resources required”, (Sathe: ¶37), “the joint optimizer 132 is configured to train two pipelines: 1) principal component analysis (PCA) to random forest (RF); and 2) outlier detection (OD) to support vector machine (SVM), on any of four worker nodes 120A, 120B, 120C, and 120D”, (Sathe: ¶30), “the joint optimizer 132 may predict required performance measures for each of the one or more worker nodes 120A-K to train a respective pipeline via the performance predictor 134”, (Sathe: ¶35).
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine provisioning, by the processing system, a head node and a first group of worker nodes based on the processing capacity, the memory capacity, the processing capacity threshold, and the memory capacity threshold; provisioning, by the processing system, a first portion of data engineering pipeline on each of the first group of worker nodes; and provisioning, by the processing system, a first portion of the machine learning model on each of the first group of worker nodes of Sathe with the methods and systems of Sidhu resulting in a system that can allocate certain parts of a recourse to certain portions of a job. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “capable of distributing a set of tasks over a set of resources with the aim of making their overall processing more efficient”, (Sathe: ¶27), “the claimed invention can predict the resource requirements of training a pipeline and continuously learns to improve the predictions using data of previous pipeline executions”, (Sathe: ¶20), “improved performance over time through backpropagation of loss, generation of a variety of training data using a Multi Arm Bandit approach, and the use of an Random Forest system that continuously predicts, gathers training data, learns, and predicts better”, (Sathe: ¶41).
Lastly regarding Claim 18, Sidhu in view of Sathe fails to teach:
receiving, by the processing system, a processing capacity threshold via first user-generated input and receiving, by the processing system, a memory capacity threshold via second generated input;
However, Mueller teaches: “a user device 1002 can provide a training request to the frontend 1029” … “information describing the computing machine on which to train a machine learning model (e.g., a graphical processing unit (GPU) instance type, a central processing unit (CPU) instance type, an amount of memory to allocate, a type of virtual machine instance to use for training, etc.)”, (Mueller: ¶94), “settings can be exposed in this manner to users, including but not limited to resource utilization limits”, (Mueller: ¶45).
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine receiving the processing capacity threshold via a first user-generated input and receiving the memory capacity threshold via a second user-generated input of Mueller with the methods and systems of Sidhu in view of Sathe resulting in a system that allows users to define thresholds. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “thereby relieving the user from the burden of having to worry about over-utilization (e.g., acquiring too little computing resources and suffering performance issues) or under-utilization (e.g., acquiring more computing resources than necessary to train the machine learning models, and thus overpaying)”, (Mueller: ¶98).
Regarding Claim 19, Sidhu in view of Sathe fails to teach:
loading, by the processing system including the head node, a portion of a group of trained machine learning model artifacts on each of the first group of worker nodes.
However, Meuller teaches: “the CIVIL service 103 (e.g., ML orchestrator 115) or AMPGS 102 may send one or more commands (e.g., API calls) to a model hosting system 140 described further herein to “host” the pipeline—e.g., launch or reserve one or more compute instances, run pipeline code 126, configure endpoints associated with the pipeline” … “the commands may include a “create model” API call that combines code for the model (e.g., inference code implemented within a container) along with model artifacts (e.g., data describing weights associated with various aspects of the model)”, (Mueller: ¶73), “the model hosting system 140 initializes ones or more ML scoring containers 1050 in one or more hosted virtual machine instance 1042” … “the model hosting system 140 forms the ML scoring container(s) 1050 from the identified container image(s)”, (Mueller: ¶123), “The model hosting system 140 further forms the ML scoring container(s) 1050 by retrieving model data corresponding to the identified trained machine learning model(s)”, (Mueller: ¶124), “The model hosting system 140 can insert the model data files into the same ML scoring container 1050, into different ML scoring containers 1050 initialized in the same virtual machine instance 1042, or into different ML scoring containers 1050 initialized in different virtual machine instances 1042”, (Mueller: ¶125), “The ML scoring containers 1050 each include a runtime 1054, code 1056, and dependencies 1052” … “The code 1056 can also include model data that represent characteristics of the defined machine learning model” … “the code 1056 results in the generation of outputs (e.g., predicted or “inferred” results”, (Mueller: ¶119).
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine comprises loading, by the head node, a portion of a group of trained machine learning model artifacts on each of the first group of worker nodes of Mueller with the methods and systems of Sidhu in view of Sathe resulting in a system that can load ML artifacts on to nodes. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “By parallelizing the training process, the model training system 120 can significantly reduce the training time”, (Mueller: ¶109).
Regarding Claim 20, Sidhu in view of Sathe fails to teach:
loading, by the processing system including the head node, inference code on each of the first group of worker nodes.
However, Meuller teaches: “the CIVIL service 103 (e.g., ML orchestrator 115) or AMPGS 102 may send one or more commands (e.g., API calls) to a model hosting system 140 described further herein to “host” the pipeline—e.g., launch or reserve one or more compute instances, run pipeline code 126, configure endpoints associated with the pipeline” … “the commands may include a “create model” API call that combines code for the model (e.g., inference code implemented within a container) along with model artifacts (e.g., data describing weights associated with various aspects of the model)”, (Mueller: ¶73), “the model hosting system 140 initializes ones or more ML scoring containers 1050 in one or more hosted virtual machine instance 1042” … “the model hosting system 140 forms the ML scoring container(s) 1050 from the identified container image(s)”, (Mueller: ¶123), “The model hosting system 140 further forms the ML scoring container(s) 1050 by retrieving model data corresponding to the identified trained machine learning model(s)”, (Mueller: ¶124), “The model hosting system 140 can insert the model data files into the same ML scoring container 1050, into different ML scoring containers 1050 initialized in the same virtual machine instance 1042, or into different ML scoring containers 1050 initialized in different virtual machine instances 1042”, (Mueller: ¶125), “The ML scoring containers 1050 each include a runtime 1054, code 1056, and dependencies 1052” … “The code 1056 can also include model data that represent characteristics of the defined machine learning model” … “the code 1056 results in the generation of outputs (e.g., predicted or “inferred” results”, (Mueller: ¶119).
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine comprises loading, by the head node, a portion of a group of trained machine learning model artifacts on each of the first group of worker nodes of Mueller with the methods and systems of Sidhu in view of Sathe resulting in a system that can load inference code on to nodes. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “By parallelizing the training process, the model training system 120 can significantly reduce the training time”, (Mueller: ¶109).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHIHAB ALAM whose telephone number is (571)272-8705. The examiner can normally be reached Mon - Fri 7:30am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bradley Teets can be reached at (571) 272-3338. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/S.A./Examiner, Art Unit 2197
/BRADLEY A TEETS/Supervisory Patent Examiner, Art Unit 2197