DETAILED ACTION
This non-final office action is responsive to application 18/438,584 as submitted on February 12th 2024.
Claim status is currently pending and under examination for claims 1-20 of which independent claims are 1, 8 and 15.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
The following are the references relied upon in the rejections below:
Mohanty (US 20250190212 A1)
Claims 1-2, 8-9 and 15-16 are rejected under 35 U.S.C. 102a(2) as being anticipated by Mohanty.
Regarding Claim 1, Mohanty teaches:
A system for dynamic allocation of container session network resources via a machine learning model ([0032] “Workspaces may be deployed in a container-based environment for fault tolerance and resiliency and may be managed by container orchestration tools … These orchestration tools may support dynamic automatic scaling (also referred to herein as “elastic auto-scaling”) of a container cluster to automatically scale a workspace's infrastructure in an effort to meet demands associated with configuring a workspace based on predicted resource sizes in connection with the execution of a recommended machine learning algorithm for a given task”
[0052] “Using one or more machine learning models, in response to the request, the ML algorithm type and workspace configuration prediction engine 230 predicts a machine learning algorithm and values for one or more workspaces including, for example, the number of containers, an amount of compute resources and an amount of memory resources needed in each container to host the predicted algorithm”
Machine learning models are used to predict an appropriate machine learning algorithm and amounts of compute and memory resources needed for a workspace in a container. A workspace running in a container is a container session and a container can be scaled based on predicted resource sizes (predicted compute, memory resources), therefore, compute and memory resources (‘container session network resources’) for a container are dynamically allocated via machine learning models.),
the system comprising: a processing device ([0006] processor);
a non-transitory storage device containing instructions when executed by the processing device, causes the processing device to perform the steps of ([0006] “a non-transitory computer-readable storage medium having embodied therein executable program code that when executed by a processor causes the processor to perform the above steps”):
train a machine learning model to form a trained machine learning model ([0035] “training a multi-target classification and regression machine learning model with historical machine learning workspace metrics”);
receive a request to initiate a container session comprising a container having a containerized version of an application model ([0041] “in response to a request to predict a machine learning algorithm and a corresponding workspace configuration received from a user via the workspace provisioning engine 140, the algorithm size and type prediction layer 132 predicts a best fit machine learning algorithm and an optimal configuration of a machine learning workspace based on a variety of features specified in the request.”
[0052] “in response to the request, the ML algorithm type and workspace configuration prediction engine 230 predicts a machine learning algorithm and values for one or more workspaces including, for example, the number of containers, an amount of compute resources and an amount of memory resources needed in each container to host the predicted algorithm”
A request from a user is received to predict a machine learning algorithm and an optimal workspace configuration. The optimal workspace configuration is used to determine an appropriate amount of compute and memory resources to host a predicted algorithm (‘application model’) in a container. Therefore, the container hosting a predicted algorithm is a container having a containerized version of an application model, and a request from a user is a request to initiate a container session (since the predicted algorithm runs in the container).);
input, to the trained machine learning model, input data of at least one selected from the group consisting of: a volume of data to be used by the application model, complexity of the application model, an identifier of a user of the application model, expected runtime of an application model, and performance of the application model ([0041] “the algorithm size and type prediction layer 132 predicts a best fit machine learning algorithm and an optimal configuration of a machine learning workspace based on a variety of features specified in the request. The features specified in the request include, for example, information regarding the type of machine learning needed (e.g., regression, classification, NLP, image classification, recommendation, etc.), required or desired size of a training dataset, required or desired feature dimension size, domain type, a number of users working with or accessing the workspace, and a type of usage (e.g., production/non-production).”
[0040] “the historical workspace metrics data is used to train the machine learning models used by the algorithm size and type prediction layer 132 to learn different combinations of metrics that correspond to particular machine learning algorithms and resource configurations.”
Trained machine learning models predict an optimal configuration of a machine learning workspace based on a required or desired size of a training dataset (therefore inputting input data to a trained machine learning model). A user sends a request to predict a corresponding workspace configuration for a predicted algorithm (see [0041]), therefore the required or desired size of a training dataset is a ‘volume of data to be used by an application model’ (the application model being the predicted algorithm to be hosted in a container).);
determine, from an output of the machine learning model based on the input data, network resource requirements of the container session for running the container (An optimal configuration of a machine learning workspace (‘output’) is predicted by trained machine learning models based on required or desired size of a training dataset (‘input data’), see [0041]. An optimal configuration describes an amount of compute resources and an amount of memory resources needed in a container to host a predicted algorithm (see [0052]), therefore the amount of compute and memory resources are network resource requirements of a container session (container hosting the predicted algorithm) for running the container.);
generate the container session comprising the container ([0052] “in a scenario where containers for workspaces A and B are needed, the workspace provisioning engine 240 predicts resource sizes for containers hosting workspaces A and B and, generates containers 1 to 6 255-1 to 255-6 with predicted resource sizes for instances of workspaces A and B 256-1 to 256-6.”
After resource sizes are predicted, containers are generated with predicted resource sizes to host instances of workspaces, therefore a container hosting an instance of a workspace is a container session.);
allocate network resources to the container session based on the output of the machine learning model ([0051] “the workspace provisioning engine 240 forwards the request to the ML algorithm type and workspace configuration prediction engine 230”
[0052] “Using one or more machine learning models, in response to the request, the ML algorithm type and workspace configuration prediction engine 230 predicts a machine learning algorithm and values for one or more workspaces including, for example, the number of containers, an amount of compute resources and an amount of memory resources needed in each container to host the predicted algorithm”
A workspace provisioning engine forwards a user request to a workspace configuration prediction engine to obtain predicted (‘output’) amount of compute and memory resources needed for a container by using machine learning models. The predicted amounts are used by the workspace provisioning engine to generate a container with the predicted amounts of resources (see [0052]), thereby allocating network resources to a container hosting a predicted algorithm (‘container session’).);
and initiate the application model in the container session ([0030] “A host device 103 may comprise one or more workspace instances 106 configured to execute machine learning algorithms and corresponding tasks. For example, a plurality of workspace instances 106 respectively corresponding to different workspaces may collectively correspond to services associated with execution of a machine learning algorithm, with each workspace instance corresponding to an independently deployable service associated with the execution of the machine learning algorithm. In illustrative embodiments, each function or a plurality of functions of a machine learning application are executed by an autonomous, independently-running workspace. As explained in more detail herein, a workspace may run in a container”
A workspace executing a machine learning algorithm (‘application model’) runs in a container, thereby initiating an application model in a container session.).
Regarding Claims 2, 9 and 16, Mohanty teaches:
The system of claim 1, wherein training the machine learning model comprises: tagging known network resource requirements for the input data to form a dataset ([0039] “The workspace metrics data may be collected from the host devices 103 and/or from applications used for monitoring workspace and host component metrics, … The workspace metrics data comprises, for example, for respective ones of a plurality of workspaces, workspace identifiers (e.g., workspace names), the type of machine learning operations executed in a given workspace … The workspace metrics data further comprises, … an amount of central processing unit (CPU) utilization (e.g., number of CPU cores (millicores)), an amount of memory utilization and an amount of input/output (IO) utilization for respective ones of a plurality of workspaces … the workspace metrics data comprises average CPU, memory and IO utilization values of a workspace.”
[0048] “the training data identifies the machine learning workspace name, the workspace domain (e.g., support, sales, marketing, supply chain), machine learning type (e.g., image classification, regression, recommendation, NLP, classification), size of feature dimensions (e.g., low, medium, high), usage (e.g., production, non-production) and size of training datasets (e.g., MiB). The training data further includes four possible ones of the multiple outputs including, but not necessarily limited to, machine learning algorithm, number of containers, compute size (e.g., number of CPU cores (millicores)) and ephemeral storage size of host components (e.g., host devices, containers, pods, VMs, etc.) (e.g., MiB).”
Workspace metrics data is collected from host devices. The collected workspace metrics data comprises inputs (workspace domain, size of training sets) and outputs (compute size, amount of memory and CPU utilization). Therefore, when collecting metrics data, the outputs are ‘tags’ describing network resource requirements (amount of memory/CPU utilization) for input data (workspace domain, size of training sets) to form a dataset (workspace metrics data).);
transforming the dataset by preprocessing the dataset ([0061] “the historical machine learning workspace metrics data 336 is read and a Pandas data frame is generated, which contains all the columns including independent variables and the dependent/target variable columns (e.g., four columns representing ML algorithm, number of containers, compute size and memory size). The pre-processing component 335 performs pre-processing of data to handle any null or missing values in the columns. For example, null/missing values in numerical columns can be replaced by the median value of that column.”);
creating a first training set comprising the dataset ([0060] “FIG. 7 depicts example pseudocode 700 for loading historical machine learning workspace metrics data into a Pandas data frame for building training data.”);
and training the machine learning model using the first training set ([0048] “ historical machine learning workspace metrics are used for training the multi-target classification and regression models.”).
Regarding Claim 8, the rejection of claim 1 is incorporated. The difference in scope being:
A computer program product … ([0109] “computer program products comprising processor-readable storage media can be used.”),
the computer program product comprising a non-transitory computer-readable medium comprising code causing an apparatus to ([0006] “non-transitory computer-readable storage medium having embodied therein executable program code that when executed by a processor causes the processor to perform the above steps.).
Regarding Claim 15, the rejection of claim 1 is incorporated. The difference in scope being:
A method … the method comprising ([Abstract] “A method comprises receiving a request to predict at least one machine learning algorithm to perform one or more tasks and to predict a configuration of one or more workspaces in which the at least one machine learning algorithm is to be executed”).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The following are the references relied upon in the rejections below:
Ghosh (US 20190066016 A1)
Claims 3-6, 10-13 and 17-19 are rejected under 35 U.S.C. 103 as being unpatentable over Mohanty / Ghosh.
Regarding Claims 3, 10 and 17, Mohanty teaches the system of claim 1, however, Mohanty does not teach monitoring a volume of data to be used an application model, which is taught by Ghosh:
wherein the instructions further cause the processing device to perform the steps of: monitor the volume of data to be used by the application model ([0077] “benchmarking analytics platform 220 may provide data (e.g., ticket assignment application usage statistics, manual ticket assignment measurement data, ticket complexity data, ticket priority data, etc.) to a computing resource, may allocate computing resources to implement the computing resource to perform the benchmarking. For example, based on providing a threshold size data set of project data (e.g., millions of data points, billions of data points, etc.) to a module for processing, benchmarking analytics platform 220 may dynamically reallocate computing resources … benchmarking analytics platform 220 may process the data using machine learning, artificial intelligence, heuristics, natural language processing, or another big data technique.”
[0083] “benchmarking analytics platform 220 may determine the recommendation based on using a machine learning technique”
[0076] “benchmarking analytics platform 220 may process the project data using an algorithm”
A benchmarking analytics platform compares the size of project data to a threshold to dynamically reallocate computing resources. The benchmarking analytics platform also performs benchmarking by using a machine learning algorithm and the project data, therefore, the benchmarking analytics platform monitors the size of project data (‘volume of data’) to reallocate resources used by the machine learning algorithm (‘application model’).).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the resource allocation method of Mohanty with the technique disclosed by Ghosh to dynamically reallocate computing resources based on the size of data to be processed. By dynamically reallocating computing resources based on the size of data to be processed, data size can be used to determine how much resources to allocate for a machine learning model, thereby ensuring optimal computing resource utilization.
Regarding Claims 4, 11 and 18, the combined resource allocation method of Mohanty / Ghosh teaches:
The system of claim 3, wherein the instructions further cause the processing device to perform the steps of: compare the volume of data to be used by the application model to a predetermined threshold ([0077] “benchmarking analytics platform 220 may provide data (e.g., ticket assignment application usage statistics, manual ticket assignment measurement data, ticket complexity data, ticket priority data, etc.) to a computing resource, may allocate computing resources to implement the computing resource to perform the benchmarking. For example, based on providing a threshold size data set of project data (e.g., millions of data points, billions of data points, etc.) to a module for processing, benchmarking analytics platform 220 may dynamically reallocate computing resources (e.g., processing resources or memory resources) to increase/decrease a resource allocation … benchmarking analytics platform 220 may process the data using machine learning, artificial intelligence, heuristics, natural language processing, or another big data technique.”
[0076] “benchmarking analytics platform 220 may process the project data using an algorithm”
A benchmarking analytics platform compares the size of project data to a threshold to dynamically reallocate computing resources. The benchmarking analytics platform also performs benchmarking by using a machine learning algorithm (‘application model’) and project data, therefore comparing project data size (‘volume of data to be used by the application model’) to a threshold. The threshold is predetermined since the threshold size is based on either the size of project data being either millions of data points or billions of data points.).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined resource allocation method of Mohanty / Ghosh with the technique disclosed by Ghosh to compare the size of data to be processed by a machine learning model to a threshold. By comparing the size of data to be processed by a machine learning model to a threshold, it can be determined if computing resources need to be increased or decreased, thereby ensuring optimal computing resource utilization.
Regarding Claims 5, 12 and 19, the combined resource allocation method of Mohanty / Ghosh teaches:
The system of claim 4, wherein upon a condition where the volume of data to be used by the application model is above the predetermined threshold, the allocated network resources increase ([0077] “For example, based on providing a threshold size data set of project data (e.g., millions of data points, billions of data points, etc.) to a module for processing, benchmarking analytics platform 220 may dynamically reallocate computing resources (e.g., processing resources or memory resources) to increase/decrease a resource allocation to the module … resource allocation for benchmarking analytics platform 220 better matches a task that is being performed, enabling more efficient resource allocation, while freeing up resources for other processing tasks when not needed. In this way, processing tasks are completed faster, and less memory is needed based on dynamically reallocating memory relative to a static allocation of resources for project management”
Resource allocation is increased or decreased based on a threshold. Resources are freed (decreased) when they are not needed by a processing task, therefore the size of project data is below the threshold (there is less data to process). When resource allocation is increased, there is more data to process and therefore the size of project data (‘volume of data’) exceeds the threshold.).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined resource allocation method of Mohanty / Ghosh with the technique disclosed by Ghosh to increase the allocation of computing resources when a machine learning model processes large datasets. By increasing the allocation of computing resources when a machine learning model processes large datasets, more processing and memory resources can be allocated to perform the large data processing task, thereby allowing the model to process the large datasets faster.
Regarding Claims 6 and 13, the combined resource allocation method of Mohanty / Ghosh teaches:
The system of claim 4, wherein upon a condition where the volume of data to be used by the application model is below the predetermined threshold, the allocated network resources decrease ([0077] “For example, based on providing a threshold size data set of project data (e.g., millions of data points, billions of data points, etc.) to a module for processing, benchmarking analytics platform 220 may dynamically reallocate computing resources (e.g., processing resources or memory resources) to increase/decrease a resource allocation to the module … resource allocation for benchmarking analytics platform 220 better matches a task that is being performed, enabling more efficient resource allocation, while freeing up resources for other processing tasks when not needed. In this way, processing tasks are completed faster, and less memory is needed based on dynamically reallocating memory relative to a static allocation of resources for project management”
Resource allocation is increased or decreased based on a threshold. Resources are freed (decreased) when they are not needed by a processing task, therefore the size of project data (‘volume of data’) is below the threshold (there is less data to process).).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined resource allocation method of Mohanty / Ghosh with the technique disclosed by Ghosh to decrease the allocation of computing resources when a machine learning model processes small datasets. By decreasing the allocation of computing resources when a machine learning model processes small datasets, excess processing and memory resources can be freed since they are not in use by the machine learning model, thereby freeing computing resources to be used for other data processing tasks.
The following are the references relied upon in the rejections below:
Noguchi (US 20250348345 A1)
Claims 7, 14 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Mohanty / Noguchi.
Regarding Claims 7, 14 and 20, Mohanty teaches: the system of claim 1, however Mohanty does not teach receiving a signal to decommission a container session, which is taught by Noguchi:
wherein the instructions further cause the processing device to perform the steps of: receive a signal to decommission the container session ([0052] “the container operation reception unit 111 receives a request for deleting the container 252 from the user's terminal 870”
[0070] “Upon receiving a request for deleting the container 252, the virtual computing resource deployment device 100 deletes the container 252 and reduces the amount of the resources used by the container 252 from the resources of the virtual machine 250 in which the container 252 has been running.”
A request (‘signal’) from a user is received to delete a container that has been running, therefore decommissioning (deleting) a container session (a running container).);
deallocate the network resources from the container session ([0038] “the container allocation unit 112 determines the number of CPU cores and the memory size used by the container 252 and the application 251, and instructs the virtual machine management unit 113 to delete the resources from the virtual machine 250.”);
and terminate the container session (When a container that has been running is deleted and the container’s resources are also deleted, the container is ‘terminated’ since it is not running anymore and its resources are no longer in use (therefore terminating a container session).).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the resource allocation method of Mohanty with the container deletion technique disclosed by Noguchi to delete containers. By deleting containers, computing resources used by a container can be freed, thereby allowing the freed resources to be used by other containers or other processing tasks.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
You et al. (US 20240370307 A1) teaches using a machine learning model to predict a resource limit for a container (pod), and monitoring resource usage to continuously adjust a container’s resource limit.
Ma et al. (US 20190286486 A1) teaches dynamically allocating resources for a container at a future time by using a machine learning model trained on runtime data.
Srinivasan et al. (US 20180300653 A1) teaches using a distributed machine learning system to determine the resources required for a container based on the size of a volume of training data.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PEDRO J MORALES whose telephone number is (571)272-6106. The examiner can normally be reached 8:30 AM - 6:00 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, MIRANDA M HUANG can be reached at (571)270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PEDRO J MORALES/Examiner, Art Unit 2124
/MIRANDA M HUANG/Supervisory Patent Examiner, Art Unit 2124