DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 1-20 are pending.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention recites a judicial exception and is directed to that judicial exception, an abstract idea, as it has not been integrated into practical application. The claims further do not recite significantly more than the judicial exception. Examiner has evaluated the claims under the framework provided in the 2019 Patent Eligibility Guidance published in the Federal Register 01/07/2019 and has provided such analysis below.
Step 1:
Claims 1-13 are directed towards a machine (i.e., apparatus). Claims 14-20 are directed towards a machine (i.e., apparatus).
Step 2A Prong 1:
In order to evaluate the Step 2A inquiry “Is the claim directed to a law of nature, a natural
phenomenon or an abstract idea?” we must determine, at Step 2A Prong 1, whether the claim recites a law of nature, a natural phenomenon or an abstract idea.
Claims 1 and 14 recite judicial exceptions in the form of abstract ideas. Claims 1 and 14 recite the mental process:
“determining, based on the first information, the second information and the third information, at least one deployment option of at least one inference step to offload on at least one host node of the plurality of host nodes”.
Step 2A Prong 2:
Do the claims recite additional elements that integrate the judicial exception into a practical application?
Claim 1 recites the additional elements:
“one processor, and at least one memory including computer program code, wherein the at least one memory and the computer program code are configured, with the at least one processor, to cause the apparatus at least to perform”
Such additional element(s) represent adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea MPEP 2106.05(f) which does not integrate the judicial exception into a practical application or amount to significantly more.
“receiving, at a first entity from a second entity, a request to offload at least one inference step of a machine learning model for an application to a host node of a network, the network comprising a plurality of host nodes; acquiring first information related to the application; acquiring second information related to the machine learning model; acquiring third information related to at least one host node of the plurality of host nodes, wherein the third information comprises at least an indication of a carbon emission footprint associated with at least one host node of the plurality of host nodes;”.
Such additional element(s) represent insignificant extra-solution data gathering activity MPEP 2106.05(g) and is well-understood, routine, and conventional in the art ((buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network)) or (Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015) (storing and retrieving information in memory))) which does not integrate the judicial exception into a practical application or amount to significantly more than the abstract idea.
“providing an indication of the determined at least one deployment option to the second entity from the first entity.”
Such additional element(s) represent insignificant extra-solution data transmitting activity MPEP 2106.05(g) and is well-understood, routine, and conventional in the art ((buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network)) or (Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015) (storing and retrieving information in memory))) which does not integrate the judicial exception into a practical application or amount to significantly more than the abstract idea.
Claim 14 recites similar additional elements as claim 1.
Step 2B:
Do the claims recite additional elements that amount to significantly more than the judicial exception?
As per claims 1 and 14, the claims do not include additional elements, alone or in combination, that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements amount to no more than generic computing components as tools to apply the abstract idea MPEP 2106.05(f), merely represent insignificant extra-solution data gathering/transmitting activity MPEP 2106.05(g), or generally link the use of the judicial exception to a particular technological environment or field of use MPEP 2106.05(h) which do not amount to significantly more than the abstract idea. Furthermore, the additional element: “2106.05(g)” is well-understood, routine, and conventional in the field because it is merely [data storing, transmitting, etc] (buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network)).
Having concluded analysis within the provided framework, independent claims 1 and 14 do not recite patent eligible subject matter under 35 U.S.C. § 101.
Dependent claims 2-13 and 15-20 do not recite additional elements that integrate the judicial exception into a practical application or amount to significantly more than the abstract idea. Therefore, claims 2-13 and 15-20 are not eligible subject matter under 35 U.S.C § 101.
Claim 2 recites the additional element: “wherein the acquiring the first information comprises receiving the first information at the first entity from the second entity”. Such additional element(s) represent insignificant extra-solution data gathering activity MPEP 2106.05(g) and is well-understood, routine, and conventional in the art ((buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network)) or (Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015) (storing and retrieving information in memory))) which does not integrate the judicial exception into a practical application or amount to significantly more than the abstract idea.
Claim 3 recites the additional element: “wherein the acquiring the second information comprises requesting the second information from a network function and receiving the second information from the network function”. Such additional element(s) represent insignificant extra-solution data gathering activity MPEP 2106.05(g) and is well-understood, routine, and conventional in the art ((buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network)) or (Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015) (storing and retrieving information in memory))) which does not integrate the judicial exception into a practical application or amount to significantly more than the abstract idea.
Claim 4 recites the additional element: “wherein the acquiring the third information comprises requesting the third information from a further network function, and receiving the third information from the further network function”. Such additional element(s) represent insignificant extra-solution data gathering activity MPEP 2106.05(g) and is well-understood, routine, and conventional in the art ((buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network)) or (Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015) (storing and retrieving information in memory))) which does not integrate the judicial exception into a practical application or amount to significantly more than the abstract idea.
Claims 5 and 16 recite the additional element: “wherein the first information comprises an indication of at least one inference step of the machine learning model to be offloaded to at least one of the plurality of host nodes”. Such additional element(s) represent linking the judicial exception to the technological environment or field of use MPEP 2106.05(h) which does not integrate the judicial exception into a practical application or amount to significantly more than the abstract idea.
Claims 6 and 17 recite the additional element: “wherein the first information comprises a quality of service indicator”. Such additional element(s) represent linking the judicial exception to the technological environment or field of use MPEP 2106.05(h) which does not integrate the judicial exception into a practical application or amount to significantly more than the abstract idea.
Claims 7 and 18 recite the additional element: “wherein the first information comprises at least one of time sensitivity information, hardware requirements of at least one inference step and data privacy requirements of at least one inference step”. Such additional element(s) represent linking the judicial exception to the technological environment or field of use MPEP 2106.05(h) which does not integrate the judicial exception into a practical application or amount to significantly more than the abstract idea.
Claims 8 and 19 recite the additional element: “wherein the request comprises the first information”. Such additional element(s) represent linking the judicial exception to the technological environment or field of use MPEP 2106.05(h) which does not integrate the judicial exception into a practical application or amount to significantly more than the abstract idea.
Claims 9 and 20 recite the additional element: “wherein the third information further comprises an indication of computing capacity of at least one host node of the plurality of host nodes”. Such additional element(s) represent linking the judicial exception to the technological environment or field of use MPEP 2106.05(h) which does not integrate the judicial exception into a practical application or amount to significantly more than the abstract idea.
Claim 10 recites the additional element: “wherein the second information comprises at least one of input data size, output data size, computing complexity and parameter size of at least one inference step of the machine learning model”. Such additional element(s) represent linking the judicial exception to the technological environment or field of use MPEP 2106.05(h) which does not integrate the judicial exception into a practical application or amount to significantly more than the abstract idea.
Claim 11 recites the additional element: “wherein the first entity comprises a management service producer hosted on a network function”. Such additional element(s) represent linking the judicial exception to the technological environment or field of use MPEP 2106.05(h) which does not integrate the judicial exception into a practical application or amount to significantly more than the abstract idea.
Claim 12 recites the additional element: “wherein the second entity comprises at least one of a user equipment, a management services consumer hosted on a network function or a machine learning entity”. Such additional element(s) represent linking the judicial exception to the technological environment or field of use MPEP 2106.05(h) which does not integrate the judicial exception into a practical application or amount to significantly more than the abstract idea.
Claim 13 recites the additional element: “wherein the plurality of host nodes comprise at least one of edge cloud servers, centre cloud servers and network function servers”. Such additional element(s) represent linking the judicial exception to the technological environment or field of use MPEP 2106.05(h) which does not integrate the judicial exception into a practical application or amount to significantly more than the abstract idea.
Claim 15 recites the additional element: “wherein the at least one memory and the computer program code are configured, with the at least one processor, to cause the apparatus to perform: providing the first information to the first entity from the second entity.” Such additional element(s) represent insignificant extra-solution data transmitting activity MPEP 2106.05(g) and is well-understood, routine, and conventional in the art ((buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network)) or (Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015) (storing and retrieving information in memory))) which does not integrate the judicial exception into a practical application or amount to significantly more than the abstract idea.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-6, 8-17, and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Haykal et al. US 20230409889 A1 in view of Wong et al. US 20230138727 A1.
Regarding claim 1, Haykal teaches the invention substantially as claimed including:
An apparatus comprising:
at least one processor ([0006] one or more processors), and at least one memory including computer program code, wherein the at least one memory and the computer program code are configured, with the at least one processor, to cause the apparatus at least to ([0006] one or more storage devices coupled to the one or more processors and storing instructions that, when executed by the one or more processors, causes the one or more processors to perform operations for machine learning model disaggregation) perform:
receiving, at a first entity from a second entity, a request to offload at least one inference step of a machine learning model for an application to a host node of a network, the network comprising a plurality of host nodes (Fig 3 architecture 300; [0055] The coordinator 316 can confirm the partition determined by the model optimizer 316 is possible based on the resource threshold from the resource optimizer 314. Once confirmed, the coordinator 316 can perform the partition of the model graph into a plurality of host nodes 326 and one or more accelerator nodes 328 of a tenant project 330; Examiner notes: user requires partitioning of their models on host nodes via service control plane 310);
acquiring first information related to the application ([0053] The resource optimizer 314 can determine a resource threshold for performing a machine learning application, such as image classification, object detection, speech recognition, natural language processing. The resource threshold can be selected based on the particular application or can be determined from results of the model profiler 312);
acquiring second information related to the machine learning model ([0054] The model optimizer 316 can determine how to partition the machine learning model graph based on factors such as data transfer threshold, connectivity topology, and/or distributional statistics over a model state, each factor to be described further below. The factors can be predetermined based on the particular machine learning application or can be determined from results of the model profiler 312);
acquiring third information related to at least one host node of the plurality of host nodes ([0028] The machine learning model graph can further be partitioned based on a hardware resource threshold, e.g., a cost of the machine learning hardware itself. The machine learning model partition can consider different generations and/or types of accelerators and general purpose processors to reach a target throughput subject to latency constraints; [0062] The model optimizer can also partition the machine learning model graph based on its statistical distribution. For example, top power law embedding table rows, e.g., rows that make up a top 10-20% of a power law distribution that are referenced by at least 80% of inference traffic, can be allocated to HBM while other embedding table rows, e.g., the torso and/or tail of the power law distribution, can be allocated to general purpose memory. The machine learning model graph can also be partitioned based on various batch-sizes and load demands of a machine learning application. For example, an analysis can be performed for various batch sizes to determine the best configuration based on the load demand for a machine learning application;),
determining, based on the first information, the second information and the third information, at least one deployment option of at least one inference step to offload on at least one host node of the plurality of host nodes ([0063] the coordinator can confirm the partition determined by the model optimizer is possible based on the resource threshold from the resource optimizer. The coordinator can confirm that the determined partition includes sufficient accelerators, high bandwidth memory, general purpose processors, and general purpose memory for performing a machine learning application); and
providing an indication of the determined at least one deployment option to the second entity from the first entity ([0063] As shown in block 440, the coordinator can confirm the partition determined by the model optimizer is possible based on the resource threshold from the resource optimizer. The coordinator can confirm that the determined partition includes sufficient accelerators, high bandwidth memory, general purpose processors, and general purpose memory for performing a machine learning application; [0064] As shown in block 450, the coordinator can perform the partition of the machine learning model graph into a plurality of host nodes and accelerator nodes when confirmed).
Haykal does not explicitly teach wherein the third information comprises at least an indication of a carbon emission footprint associated with at least one host node of the plurality of host nodes;
However, Wong teaches wherein the third information comprises at least an indication of a carbon emission footprint associated with at least one host node of the plurality of host nodes ([0003] The method further includes based on the SLA requirement, the criticality level, the peak load duration, and the previous success rates, the computer system selecting an optimized configuration of one or more cloud resources and one or more cloud service providers providing the one or more cloud resources for the workload. The one or more cloud resources are selected from the list of cloud resources and have a carbon footprint that does not exceed the carbon footprint cap at a given load level).
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to have combined Wong’s factoring of carbon footprint in resource allocation with the existing system. A person of ordinary skill in the art would have been motivated to make this combination to provide the resulting system with the advantage of satisfying consumer carbon footprint goals by reducing carbon footprints (see Wong [0002] Industries are advocating for greener data centers and are targeting a reduced carbon footprint as part of their corporate social responsibility. Companies are working towards green computer environments and a reduced carbon footprint while executing their business applications. Since cloud companies do not provide carbon footprint details of cloud resources, companies using the cloud resources do not know the carbon footprint and the impact on sustainable environment caused by processing their business logic, applications, and workload).
Regarding claim 2, Haykal and Wong teach the apparatus according to claim 1.
Haykal further teaches wherein the acquiring the first information comprises receiving the first information at the first entity from the second entity ([0032] a coordinator to confirm the partition decided by the model optimizer is possible based on the resource threshold from the resource optimizer).
Regarding claim 3, Haykal and Wong teach the apparatus according to claim 1.
Haykal further teaches wherein the acquiring the second information comprises requesting the second information from a network function and receiving the second information from the network function ([0061] The model optimizer can further partition the machine learning model graph based on connectivity topologies, such as an arrangement of host nodes, accelerator nodes, and interconnects. The connectivity topology can include slicing of accelerators, hierarchies of network topologies, and placement of nodes; [0034] The network 120 can facilitate interactions between participant devices. Example networks include the Internet, a local network, a network fabric, or any other local area or wide area network. The network 120 can be composed of multiple connected sub-networks or autonomous networks).
Regarding claim 4, Haykal and Wong teach the apparatus according to claim 1.
Haykal further teaches wherein the acquiring the third information comprises requesting the third information from a further network function, and receiving the third information from the further network function ([0030] When a model is deployed for inference, a user can pass a sample dataset from and the model profiler will run the model graph on that set and collect information regarding various resources used, e.g., CPU, RAM, TPU, HBM, network bandwidth. This information can be used to make better informed graph partitioning decisions; [0034] The network 120 can facilitate interactions between participant devices. Example networks include the Internet, a local network, a network fabric, or any other local area or wide area network. The network 120 can be composed of multiple connected sub-networks or autonomous networks).
Regarding claim 5, Haykal and Wong teach the apparatus according to claim 1.
Haykal further teaches wherein the first information comprises an indication of at least one inference step of the machine learning model to be offloaded to at least one of the plurality of host nodes ([0053] the resource optimizer 314 can determine a resource threshold for performing a machine learning application, such as image classification, object detection, speech recognition, natural language processing. The resource threshold can be selected based on the particular application or can be determined from results of the model profiler 312; [0063] As shown in block 440, the coordinator can confirm the partition determined by the model optimizer is possible based on the resource threshold from the resource optimizer. The coordinator can confirm that the determined partition includes sufficient accelerators, high bandwidth memory, general purpose processors, and general purpose memory for performing a machine learning application).
Regarding claim 6, Haykal and Wong teach the apparatus according to claim 1.
Haykal further teaches wherein the first information comprises a quality of service indicator ([0060] The model optimizer can partition the machine learning model graph based on a data transfer threshold, such as a cost of exchanging data between nodes over a network. The data transfer threshold can include thresholds for network bandwidth, latency, and/or throughput to reduce hops between host nodes and accelerator nodes. The amount of bytes being transferred should minimize disaggregation overhead. For example, a flow of bytes between two resulting sub-graphs of a partition should not cause a bottleneck due to network bandwidth).
Regarding claim 8, Haykal and Wong teach the apparatus according to claim 1.
Haykal further teaches wherein the request comprises the first information (Fig 3 Coordinator 318; [0055] The coordinator 316 can confirm the partition determined by the model optimizer 316 is possible based on the resource threshold from the resource optimizer 314).
Regarding claim 9, Haykal and Wong teach the apparatus according to claim 1.
Haykal further teaches wherein the third information further comprises an indication of computing capacity of at least one host node of the plurality of host nodes ([0032] Horizontal auto-scaling can be performed based on real-time resource usage of the disaggregated resource pools. For example, if TPU compute is close to maximum which results in capping throughput of the model, the coordinator can spin up more accelerator nodes to better balance the load; [0055] The coordinator 316 can also dynamically auto-scale the disaggregated resource pools of host nodes 326 and accelerator nodes 328 based on real-time resource usage from user requests 332. For example, if accelerator compute is close to maximum, resulting in capping throughput of the partitioned model, the coordinator 316 can spin up more accelerator nodes 328 to better balance the load).
Regarding claim 10, Haykal and Wong teach the apparatus according to claim 1.
Haykal further teaches wherein the second information comprises at least one of input data size, output data size, computing complexity and parameter size of at least one inference step of the machine learning model ([0029] The machine learning model graph can also be partitioned based on various batch-sizes when considering a load demand. For example, running a matrix multiplication function with a low batch-size might not be as efficient, so an analysis can be performed for various batch sizes to determine the best configuration based on the demand of a received load).
Regarding claim 11, Haykal and Wong teach the apparatus according to claim 1.
Haykal further teaches wherein the first entity comprises a management service producer hosted on a network function (Fig 3 Service control plane 310; [0051] FIG. 3 depicts a block diagram of an example architecture 300 for partitioning a machine learning model graph. The architecture can include a service control plane 310 to allow for degrees of freedom to compile and/or rewrite model graphs as well as automatically tune runtime parameters, such as operation placement, threading, and accelerator-specific knobs. The service control plane 310 can include a model profiler 312, resource optimizer 314, model optimizer 316, and coordinator 318).
Regarding claim 12, Haykal and Wong teach the apparatus according to claim 1.
Haykal further teaches wherein the second entity comprises at least one of a user equipment, a management services consumer hosted on a network function or a machine learning entity (Fig 3 User Project, Tenant Project; [0052] A user project 320 can be a project in a user space that is visible to end users of the elements of the service control plane 310; [0055] A tenant project 330 can be a project parallel to the user project 320 where most logic of the elements of the service control plane 310 run.).
Regarding claim 13, Haykal and Wong teach the apparatus according to claim 1.
Haykal further teaches wherein the plurality of host nodes comprise at least one of edge cloud servers, centre cloud servers and network function servers ([0023] Each host node can include a general purpose processor, e.g., CPU, and a general purpose memory, e.g., RAM. The processor can include parsing and/or lookup operations and the memory can include embedding tables. Each accelerator node can include a machine learning accelerator, e.g., tensor processing unit (TPU). The machine learning accelerator can include deep neural network (DNN) operations and a high bandwidth memory (HBM). The HBM can include additional embedding tables and/or model parameters. Each accelerator node can also include a general purpose processor and general purpose memory. The host nodes and accelerator nodes can include a network interface card that connects to a network so that the nodes can interact with each other, e.g., transferring data).
Regarding claim 14, Haykal teaches the invention substantially as claimed including:
An apparatus comprising:
at least one processor ([0006] one or more processors), and at least one memory including computer program code, wherein the at least one memory and the computer program code are configured, with the at least one processor, to cause the apparatus at least to ([0006] one or more storage devices coupled to the one or more processors and storing instructions that, when executed by the one or more processors, causes the one or more processors to perform operations for machine learning model disaggregation) perform:
providing, to a first entity from a second entity, a request to offload at least one inference step of a machine learning model for an application to a network, the network comprising a plurality of host nodes (Fig 3 architecture 300; [0055] The coordinator 316 can confirm the partition determined by the model optimizer 316 is possible based on the resource threshold from the resource optimizer 314. Once confirmed, the coordinator 316 can perform the partition of the model graph into a plurality of host nodes 326 and one or more accelerator nodes 328 of a tenant project 330; Examiner notes: user requires partitioning of their models on host nodes via service control plane 310); and
receiving, from the first entity at the second entity an indication of a determined at least one deployment option of at least one inference step to offload on at least one host node of the plurality of host nodes ([0063] As shown in block 440, the coordinator can confirm the partition determined by the model optimizer is possible based on the resource threshold from the resource optimizer. The coordinator can confirm that the determined partition includes sufficient accelerators, high bandwidth memory, general purpose processors, and general purpose memory for performing a machine learning application; [0064] As shown in block 450, the coordinator can perform the partition of the machine learning model graph into a plurality of host nodes and accelerator nodes when confirmed), the at least one deployment option determined by the first entity based on first information related to the application ([0053] The resource optimizer 314 can determine a resource threshold for performing a machine learning application, such as image classification, object detection, speech recognition, natural language processing. The resource threshold can be selected based on the particular application or can be determined from results of the model profiler 312), second information related to the machine learning model ([0054] The model optimizer 316 can determine how to partition the machine learning model graph based on factors such as data transfer threshold, connectivity topology, and/or distributional statistics over a model state, each factor to be described further below. The factors can be predetermined based on the particular machine learning application or can be determined from results of the model profiler 312) and third information related to at least one host node of the plurality of host nodes ([0028] The machine learning model graph can further be partitioned based on a hardware resource threshold, e.g., a cost of the machine learning hardware itself. The machine learning model partition can consider different generations and/or types of accelerators and general purpose processors to reach a target throughput subject to latency constraints; [0062] The model optimizer can also partition the machine learning model graph based on its statistical distribution. For example, top power law embedding table rows, e.g., rows that make up a top 10-20% of a power law distribution that are referenced by at least 80% of inference traffic, can be allocated to HBM while other embedding table rows, e.g., the torso and/or tail of the power law distribution, can be allocated to general purpose memory. The machine learning model graph can also be partitioned based on various batch-sizes and load demands of a machine learning application. For example, an analysis can be performed for various batch sizes to determine the best configuration based on the load demand for a machine learning application;).
Haykal does not explicitly teach wherein the third information comprises at least an indication of a carbon emission footprint associated with at least one host node of the plurality of host nodes;
However, Wong teaches wherein the third information comprises at least an indication of a carbon emission footprint associated with at least one host node of the plurality of host nodes ([0003] The method further includes based on the SLA requirement, the criticality level, the peak load duration, and the previous success rates, the computer system selecting an optimized configuration of one or more cloud resources and one or more cloud service providers providing the one or more cloud resources for the workload. The one or more cloud resources are selected from the list of cloud resources and have a carbon footprint that does not exceed the carbon footprint cap at a given load level).
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to have combined Wong’s factoring of carbon footprint in resource allocation with the existing system. A person of ordinary skill in the art would have been motivated to make this combination to provide the resulting system with the advantage of satisfying consumer carbon footprint goals by reducing carbon footprints (see Wong [0002] Industries are advocating for greener data centers and are targeting a reduced carbon footprint as part of their corporate social responsibility. Companies are working towards green computer environments and a reduced carbon footprint while executing their business applications. Since cloud companies do not provide carbon footprint details of cloud resources, companies using the cloud resources do not know the carbon footprint and the impact on sustainable environment caused by processing their business logic, applications, and workload).
Regarding claims 15-17 and 19-20, they are the apparatus with elements according to claims 2, 5-6, and 8-9. Therefore, they are rejected for the same reasons as claims 2, 5-6, and 8-9 respectively.
Claims 7 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Haykal et al. US 20230409889 A1 in view of Wong et al. US 20230138727 A1 in view of Beisiegel et al. US 20090313319 A1.
Regarding claim 7, Haykal and Wong teach the apparatus according to claim 6.
Haykal further teaches wherein the first information comprises at least one of time sensitivity information, hardware requirements of at least one inference step ([0028] The machine learning model graph can further be partitioned based on a hardware resource threshold, e.g., a cost of the machine learning hardware itself. The machine learning model partition can consider different generations and/or types of accelerators and general purpose processors to reach a target throughput subject to latency constraints).
Haykal and Wong do not explicitly teach the first information comprising data privacy requirements of at least one inference step.
However, Beisiegel teaches the first information comprising data privacy requirements of at least one inference step ([0029] a constraint may be an abstract constraint, in which the system automatically decides at runtime where to run a scope or resource based on a developer's specified non-functional and/or QoS requirements for a scope or a resource. The non-functional and/or QoS requirements may comprise at least one requirement including, but not limited to, data sharing, security, reliability, privacy, trusted code, confidentiality, or user identity).
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to have combined Beisiegel’s data privacy QoS requirement with the existing system. A person of ordinary skill in the art would have been motivated to make this combination to provide the resulting system with the advantage of specifying requirements of data privacy and confidentiality in partitioned applications (see Beisiegel [0048] According to the present invention, a web application may comprise a constraint comprising at least one privacy requirement. The at least one privacy requirement may comprise sensitive information including, but are not limited to, passwords, user names, financial information, or proprietary information; [0049] At runtime, the application is preferably partitioned so that any sensitive application logic/data stays on a server. Thus, the present invention allows a web application to run in a secure environment. In embodiments, an application may provide a confidentiality guarantee as a result of the partitioning.).
Regarding claim 18, it is the apparatus with elements according to claim 7. Therefore, it is rejected for the same reasons as claim 7 respectively.
Conclusion
Any inquiry concerning this communication or earlier communications from the
examiner should be directed to HARRISON LI whose telephone number is (703) 756-1469. The
examiner can normally be reached Monday-Friday 9:00am-5:30pm ET.
Examiner interviews are available via telephone, in-person, and video conferencing
using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is
encouraged to use the USPTO Automated Interview Request (AIR) at
http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s
supervisor, Aimee Li can be reached on (571) 272-4169. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/H.L./
Examiner, Art Unit 2195
/Aimee Li/Supervisory Patent Examiner, Art Unit 2195