Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Examiner Notes
Examiner cites particular columns and line numbers in the references as applied to the claims below for convenience of the applicant. Although the specified citations are representative of the teachings in the art and are applied to the specific limitations within the individual claim, other passages and figures may apply as well. It is respectfully requested that, in preparing responses, the applicant fully consider the references cited in their entirety as potentially teaching all or part of the claimed invention, as well as the context of the passage as taught by the prior art or disclosed by the examiner.
Response to Amendment
The amendment filed 8/20/2025 has been entered. The Amendments to the Specification and Claims overcome all Objections set forth in the previous Office Action. Claims 1-20 remain pending in the application.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 11-19 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 11 recites “a processor configured to:” in line 2. This limitation identifies an intended use for the processor. Accordingly, it is unclear if the processor actually performs the functions following the “to:” (e.g. “establish…”, “allow…”) or if it is only intended to perform such functions (e.g. the processor has the hardware configuration capable of performing the functions, but does not necessarily perform them).
For the following analysis, the Examiner will consider the processor actually performing the functions recited in the claim.
Claims 12-19 inherit the deficiencies of claim 11 and are rejected accordingly. Claims 12-13 and 18-19 recite additional limitations containing the phrase “the processor is configured to” or “the processor is further configured to” and thus recite intended use for the same reasons presented with respect to claim 11.
To overcome this rejection, the Examiner recommends amending the claims to specify the apparatus comprises a memory containing instructions, and that the processor executes the instructions to cause the processor to perform the recited functions.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1, 6-7, 9-11, 15-16 and 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over SENGUPTA et al. (U.S. Pub. No. 2020/0004596), hereinafter SENGUPTA, in view of Fong (U.S. Pub. No. 2023/0273837).
Regarding claim 1, SENGUPTA teaches a method, comprising:
establishing, utilizing a processor of a computing device (FIG. 2, CPU 212), a connection ([0044] – “Communication between an application instance 211 and an accelerator 221 happens on multiple network connections with each connection being initiated by an application instance 211.”) between a first application ([0037]-[0038] – “An application instance 211 is a virtual computing environment that uses a particular configuration of CPU 212, memory, storage, and networking capacity that is to execute an application 213. In some embodiments, the configuration is called an instance type. A template for an application instance (including an operating system) is called a machine image. The application 213 may be called "inference application" below to highlight that a part of the application 213 makes inference calls to at least one accelerator slot. However, the application 213 typically includes other code and the inference call is usually one aspect of an application pipeline. The instance type that is specified by a user determines the hardware of the host computer to be used for the instance within the web services provider 101.”; [0028] – “The elastic inference 105 is utilized to execute an application that includes at least some portion of code (such as a model) to be executed on an accelerator.”; [0030] – “one or more applications including one or more models to be executed on the elastic inference service 105. Applications may be hosted by the elastic inference service 105 as containers or a run as a part of a virtual machine.”; Client component (application instance 211) on a host computer with CPU 212 (“processor of a computing device”) is analogous to the “first application pod running at the computing device”.) and a machine learning inference ([0039] – “Similarly, an accelerator appliance 221 (another compute instance) uses a particular configuration of CPU 224, memory, storage, and networking capacity that is to execute a machine learning model of application 213. The accelerator appliance 221 additionally has access to one or more accelerators 222. Each accelerator 222 is comprised of one or more accelerator slots 223. […] Code executing on the CPU 224 orchestrates the on-board accelerators, runs an inference engine runtime, and offloads computation to the accelerators 222. […] An accelerator slot 223 handle the application instance's 211 calls.”; [0085] – “inference engine 325 running on accelerator slot 223”; [0091] – “In some embodiments, a ML engine on the accelerator appliance 221 is utilized. The ML engine takes in the model as input and executes it during inferencing.”; [0097] – “The accelerator appliance manager 343 performs a launch/configure of each accelerator slot 223 to use at circle 2. In some embodiments, this includes launching of a container, or virtual machine for the machine learning model.” Accelerator appliance 221 compute instance on which a container is launched for the machine learning model and an inference engine runtime is run is analogous to the “machine learning inference pod”), wherein the first application ([0028] – “The elastic inference 105 is utilized to execute an application that includes at least some portion of code (such as a model) to be executed on an accelerator.”; [0030] – “one or more applications including one or more models”; [0037] – “application 213 typically includes other code and the inference call is usually one aspect of an application pipeline.” Application 213 (and implicitly the container in which application 213 is executed) includes the code (“states representative of”) for the machine learning model called by the application in addition to other code for the application pipeline.) and
utilizing the processor of the computing device, allowing the machine learning inference ([0055] – “The application 213 itself uses a client library 315 to make calls to the inference engine 325 of the accelerator 223.”; [0058] – “eia.loadModel(name, model-config.xml, runtimeParameters) - load a model with the configuration given in "model-config.xml". The framework, version, location and other details related to the model could be passed using "model-config.xml"”; [0077] – “An inference engine (IE) 325 handles model loading and inference execution. As shown, this engine 325 is a separate process per accelerator slot 223. The IE 325 receives requests from the client library 315 on customer instance via a front-end receiver library.”; [0086] – “The EI interface (the interface to the EIA) can be attached to an application instance during instance launch or dynamically attached/detached (to/from) a live instance. The EI interface can be accessed directly using the client library 315”; [0088] – “Typically, a model is loaded into an application instance and accelerator via one or more files and/or objects.”; [0089] – “During EI interface initialization, the trained model (computation graph) is provided as input”. The model, e.g., the file “model-config.xml”, is provided as input in an API call sent from the application instance 211 (analogous to the “application pod” and comprising the client library 315 – see FIG. 3 and [0139]) to the inference engine of the attached accelerator slot of the accelerator appliance 221 (analogous to the “machine learning inference pod”) in order to load the model. Thus, this API call “allows” the inference engine to load the machine learning model “directly” from the application instance 211 via the EI interface connecting the accelerator slot and the application.) to thereby enable the machine learning inference ([0059] – exemplary API call of client library 315 sent to the inference engine to perform inference, see [0086] and [0089) based at least in part on the machine learning model ([0077] – “An inference engine (IE) 325 handles model loading and inference execution. As shown, this engine 325 is a separate process per accelerator slot 223. The IE 325 receives requests from the client library 315 on customer instance via a front-end receiver library. The IE 325 encompasses the run-times needed for the inference to work.”; [0049] – remotely attached accelerator slot 223 is also referred to as an elastic inference accelerator (EIA); [0086] – “The EI interface (the interface to the EIA) can be attached to an application instance during instance launch or dynamically attached/detached (to/from) a live instance. The EI interface can be accessed directly using the client library 315 or via model frameworks (such TensorFlow and MXNet frameworks). In a typical use case, an application instance 211 runs a larger machine learning pipeline, out of which only the accelerator appliance 221 bound calls will be sent using the EI interface API and the rest executed locally.”; [0089] – “During EI interface initialization, the trained model (computation graph) is provided as input and profiled, and, subsequently, the application 213 makes inferencing calls via the client library 315 on the application instance 211.”).
SENGUPTA fails to expressly teach using pods, nor the machine learning inference pod executing at an edge node.
However, Fong teaches using pods ([0045] – “In general, for a Kubernetes environment, one or more containers are part of a pod. […] Furthermore, a pod is typically considered the smallest execution unit in the Kubernetes container orchestration environment. A pod encapsulates one or more containers. […] Furthermore, pods typically represent the respective processes running on a cluster. A pod may be configured as a single process wherein one or more containers execute one or more functions that operate together to implement the process. […] Configuration information (configuration objects) indicating how a container executes can be specified for each pod.”; [0031] – “virtualized computing resources, such as virtual machines (VMs), containers, etc.”) and a machine learning inference pod running at an edge node ([0045] – “One or more pods are executed on a worker node.”; [0030] – “an information processing center may comprise an edge-based system that includes one or more edge computing platforms with edge devices and/or edge gateways that operate in accordance with an edge computing paradigm. […] Nodes 102 in computing environment 100 are intended to represent any one or more of the types of devices illustratively mentioned above, as well as other types of devices.”; [0060] – “Each inference executable 213 is implemented differently in terms of its corresponding ML framework, resource (CPU, accelerator) dependency, ML model format, etc. However, regardless of the particular ML framework, when an inference executable container is warming up, it is assumed that all necessary libraries required by the corresponding ML framework are loaded.”).
SENGUPTA and FONG are considered to be analogous art to the claimed invention because they are in the same field of computing environments for providing machine learning inference as a service. As discussed above, SENGUPTA teaches the application which has an inference model may be run as containers or a virtual machine (see [0030]), and a container or virtual machine may be launched for the machine learning model at the accelerator slot (see [0097]). Thus, it would have been obvious to one of ordinary skill in the art that these containers for the application containing an inference model and container for the machine learning model taught by SENGUPTA could be implemented in Kubernetes pods as disclosed by Fong. Fong teaches Kubernetes is a prevalent and well-known container orchestration system for managing containerized workloads, and that pods are the smallest execution unit for Kubernetes such that one or more containers are part of a pod (Fong: [0044]-[0045]). Further, it would have been obvious to one of ordinary skill in the art to have implemented the machine learning inference model taught by SENGUPTA as a pod running on an edge node as taught by FONG, in order to move data computation closer to the source of the data and reduce the time needed to handle inference requests (Fong: [0017], [0021], and [0027]).
Regarding claim 6, the combination of SENGUPTA in view of Fong teaches the method of claim 1, wherein the machine learning inference pod comprises one or more application programming interfaces and an inference engine (SENGUPTA: [0086] – “The EI interface (the interface to the EIA) can be attached to an application instance during instance launch or dynamically attached/detached (to/from) a live instance. The EI interface can be accessed directly using the client library 315 or via model frameworks (such TensorFlow and MXNet frameworks). In a typical use case, an application instance 211 runs a larger machine learning pipeline, out of which only the accelerator appliance 221 bound calls will be sent using the EI interface API and the rest executed locally.” The EIA (“elastic inference accelerator”/accelerator slot 223) hosted by the accelerator appliance 221 (see [0049]) has an interface API which allows the application instance 211 to send calls to the accelerator appliance 221. [0077] – “An inference engine (IE) 325 handles model loading and inference execution. As shown, this engine 325 is a separate process per accelerator slot 223. The IE 325 receives requests from the client library 315 on customer instance via a front-end receiver library. The IE 325 encompasses the run-times needed for the inference to work.” The calls made to inference engine 325 from client library 315 to the receiver library may be API calls (see [0055]-[0056]).).
Regarding claim 7, the combination of SENGUPTA in view of Fong teaches the method of claim 6, wherein the edge node performs the inference operation at least in part by executing the inference engine (SENGUPTA: [0077] – “An inference engine (IE) 325 handles model loading and inference execution. As shown, this engine 325 is a separate process per accelerator slot 223. The IE 325 receives requests from the client library 315 on customer instance via a front-end receiver library. The IE 325 encompasses the run-times needed for the inference to work.”; [0125] – “As scoring data is received by the inference application 213 is directed to the inference engine 325 and the result(s) are passed back.”).
Regarding claim 9, the combination of SENGUPTA in view of Fong teaches the method of claim 1, wherein the establishing the connection between the first application pod and the machine learning inference pod comprises establishing the connection at least in part in accordance with a hypertext transport protocol (HTTP) and/or at least in part in accordance with a remote procedure call framework (SENGUPTA: [0053] – “When the control plane 351 attaches an accelerator 223, it injects into the instance metadata service (IMDS) 371 of the application instance 211 information on how to contact the accelerator 223 […] the AIM 317 uses the IMDS 371 to check if an accelerator is attached. […] If an accelerator is attached, the AIM 317 tries to connect to an accelerator slot manager (ASM) 329 in some embodiments. The communication happens through an endpoint served by the ASM 329 and initiated by AIM 317 in some embodiments. […] the IMDS 371 is a http server”; [0055] – “The application 213 itself uses a client library 315 to make calls to the inference engine 325 of the accelerator 223. In some embodiments, the client library 315 uses gRPC for the remote procedure calls to the inference engine 325.”).
Regarding claim 10, the combination of SENGUPTA in view of Fong teaches the method of claim 1, further comprising:
establishing one or more additional connections between one or more additional application pods and the machine learning inference pod (SENGUPTA: [0049] – “a remotely attached accelerator 223 is presented as an elastic inference accelerator (EIA) (or simply accelerator) attached to the user's compute instance. The user' s compute instance is labeled as an application instance (AI) 211 and the compute instance which hosts the EIA on the service side is labeled accelerator appliance (AA) 221. Each EIA is mapped to an accelerator slot (AS), which is a fraction of an accelerator is and managed by the AA 221.”; [0050] – “An AA 221 may host multiple EIAs 223 and supports multi-tenancy (such that it allows attachments from different application instances 211 belonging to different users). Each accelerator slot can only be attached to one application instance at any given time.”; [0051] – “The data plane 301 enables users to run Deep Learning inference applications using one or more remotely attached EIAs”); and
allowing the machine learning inference pod to directly access one or more additional machine learning models pertaining respectively to the one or more additional application pods (SENGUPTA: [0039] – “Similarly, an accelerator appliance 221 (another compute instance) uses a particular configuration of CPU 224, memory, storage, and networking capacity that is to execute a machine learning model of application 213. The accelerator appliance 221 additionally has access to one or more accelerators 222. Each accelerator 222 is comprised of one or more accelerator slots 223”; [0049]-[0051] – multiple application instances 211 (with respective inference applications 213 and associated inference models) may be attached to respective accelerator slots 223 of the accelerator appliance 221; [0077] – there is a separate inference engine 325 process to handle model loading and execution for each accelerator slot 223; [0087] – “An EI interface API allows for loading models, making inference calls against them (such as tensor in/tensor out), and unloading models. Multiple models can be loaded via an EI interface at any given time.”), wherein the machine learning inference pod prevents the first software application from accessing the one or more additional machine learning models and also prevents one or more additional software applications pertaining respectively to the one or more additional application pods from accessing the machine learning model of the first application pod ([0039] – “Each accelerator 222 is comprised of one or more accelerator slots 223 . In particular, the compute resources of the accelerator appliance 221 are partitioned, allocated, resource governed and isolated across accelerator slots 223 for consistent, sustained performance in a multitenant environment.”; [0040] – “Resource governance and isolation may reduce/mitigate interference of requests across accelerator slots 223 through static partitioning of resources across accelerator slots 223 using appliance management components 241.”; [0050] – “Each accelerator slot can only be attached to one application instance at any given time.” Only the attached application instance 211 from which the model is loaded can be attached to a particular accelerator slot 223 on which the model is loaded and executed.; [0071] – “The AA 221 comprises several components including accelerator slot(s) 223, disk 333, and appliance management components 241. AA 221 components are responsible for bootstrapping, provisioning of isolated accelerator slots, monitoring events from the control plane 351 (such as attachment/recycling of accelerator slots), updating slots and appliance status (such as health and network connectivity) to the control plane 351, uploading metrics/logs.”; [0074] – “As noted above, accelerator slot 223 components run on every accelerator slot. All the components in an accelerator slot are isolated in terms of resources such as CPU, RAM, GPU compute, GPU memory, disk, and network. Accelerator slot components serve the attached customer instance for that accelerator slot 223.”; [0163] – “In some embodiments, a network namespace is utilized to isolate network interfaces among accelerator slots on the same physical accelerator appliance such that each accelerator slot's network interface resides in its own namespace. Moving an accelerator slot's network interface to its own namespace allows for different virtual networks to have overlapping IP addresses.”).
Claim 11 is directed to an apparatus, comprising: a processor (SENGUPTA: [0194] – “FIG. 19 illustrates a logical arrangement of a set of general components of an example computing device 1900 such as a web services provider, etc. Generally, a computing device 1900 can also be referred to as an electronic device. The techniques shown in the figures and described herein can be implemented using code and data stored and executed on one or more electronic devices (e.g., a client end station and/or server end station). […] such electronic devices include hardware, such as a set of one or more processors 1902”) configured to: perform the method of claim 1. Accordingly, claim 11 is rejected as being unpatentable over SENGUPTA in view of Fong for the same reasons presented with respect to claim 1.
Claims 15-16 and 18-19 recite the same limitations as claims 6-7 and 9-10, respectively; accordingly, claims 15-16 and 18-19 are rejected as being unpatentable over SENGUPTA in view of Fong for the same reasons presented with respect to claims 6-7 and 9-10.
Claim 20 is directed to an article, comprising: a non-transitory computer-readable medium having
stored thereon one or more instructions executable by a computing device (SENGUPTA: [0194] – “The techniques shown in the figures and described herein can be implemented using code and data stored and executed on one or more electronic devices […] the non-transitory machine-readable storage media (e.g., memory 1904) of a given electronic device typically stores code (e.g., instructions 1914) for execution on the set of one or more processors 1902 of that electronic device. One or more parts of various embodiments may be implemented using different combinations of software, firmware, and/or hardware.”) to: perform the method of claim 1. Accordingly, claim 20 is rejected as being unpatentable over SENGUPTA in view of Fong for the same reasons presented with respect to claim 1.
Claims 2-4 and 12-14 are rejected under 35 U.S.C. 103 as being unpatentable over SENGUPTA in view of Fong as applied to claims 1 and 11 above, and further in view of Wells (U.S. Pub. No. 2021/0218750).
Regarding claim 2, the combination of SENGUPTA in view of Fong teaches the method of claim 1, but fails to teach wherein the allowing the machine learning inference pod to load the machine learning model directly from the first application pod includes allowing the machine learning inference pod to load the machine learning model over a reverse mount of a local file namespace for the first application pod.
However, Wells teaches allowing the machine learning inference pod to load the machine learning model over a reverse mount of a local file namespace for the first application pod ([0012] – “the method includes executing, on a computing device, a process running in a first container assigned to a first namespace, the first container being assigned a first privilege that restricts access by the first container to the first namespace. […] The method further includes sending, to the namespace service and from the process, a request to create a socket in the second namespace to allow the process to communicate with a second container assigned to the second namespace. The method further includes receiving, at the process and from the namespace service, a file descriptor associated with the socket and a third privilege that allows the process to utilize the socket.”; [0045] – “At 202, a process 110(1) may execute, on a computing device 106 in the cloud computing network 102. In some examples, the process 110(1) may be running in a container 108(1) that is assigned to a first namespace (namespace A 114).”; [0046] – “the namespace service 116 may execute on each of the physical server(s) 106 in the cloud computing environment 102.”; [0047] – “At 206, the namespace service 116 may receive a request from the process 110(1) executing in the container 108(1) in the first namespace. In some examples, the request may include a request to create a socket in the second namespace, such that the process 110(1) in the first namespace may communicate with a container 108(2) in the second namespace without escalating the privileges of the process 110(1) and/or the container 108(1) with respect to the system.”; [0049] – “At 210, the namespace service 116 may provide the process 110(1) with information required to utilize the socket. In some examples, the namespace service 116 may send the file descriptor associated with the socket to the process 110(1), via the UDS 118(1). In some examples, the namespace service 116 may provide the process 110(1) with a rights object associated with the socket, such that the rights object allows the process 110(1) to utilize the socket in the second namespace.”; [0075] – “A process contained by a container 108 may utilize rights and file descriptors associated with a socket in a separate namespace received from the namespace service 116, allowing the process 108 to send and receive communications to multiple namespaces.” In FIG. 1, the first process 110(1) in a first container 108(1) (analogous to the “machine learning inference pod”) may be allowed to access the namespace 114 of second container 108(2) in which a second process 110(2) is executing (“first application pod”) by a namespace service 116 which executes on each physical device.
For clarity of the record, the Examiner would like to point to paragraph [0052] of the specification of the instant application which recites “’Reverse mount’ and/or the like refers to a sending component, such as an application pod, providing content (e.g., signals and/or signal packets) to a receiving component, such as a ML inference engine running at an edge node, to enable the receiving component to load particular content, such as a serialized ML model, directly from the sending component via access to the sending component's local namespace.” Thus, a container accessing the namespace of another container using a socket to send and receive communications may be analogous to a “reverse mount” as claimed.).
Wells is considered to be analogous art to the claimed invention because it is reasonably pertinent to the problem faced by the inventor of loading a machine learning model from the application pod to the machine learning inference pod. As presented in the rejection of claim 1, SENGUPTA teaches loading a model from a container in which an application is executed to another container where the model is executed, and Fong teaches containers in the Kubernetes environment are executed within pods. Wells teaches the Kubernetes architecture uses namespaces for containers to prevent containerized applications from interfering with each other and to deny traffic from different namespaces, which ensures security of the applications (Wells: [0014]-[0015]). Therefore, it would have been obvious to one of ordinary skill in the art to modify the method taught by SENGUPTA in view of Fong to incorporate the teachings of Wells such that the machine learning pod loads the model from the first application pod over a reverse mount of a local file namespace for the first application pod, i.e., by accessing the local namespace for the first application pod. The methods of Wells provides the advantages of enabling unprivileged containerized applications to access multiple namespaces without escalating the privileges of the application to an unsafe level, while also maintaining security between workloads in different containers and avoiding network interference (Wells: [0023]).
Regarding claim 3, the combination of SENGUPTA in view of Fong teaches the method of claim 1, but fails to teach wherein the allowing the machine learning inference pod to load the machine learning model directly from the first application pod includes allowing the machine learning inference pod executing at the edge node to access a local namespace from the first application pod.
However, Wells teaches allowing the machine learning inference pod executing at the edge node to access a local namespace from the first application pod ([0012] – “the method includes executing, on a computing device, a process running in a first container assigned to a first namespace, the first container being assigned a first privilege that restricts access by the first container to the first namespace. […] The method further includes sending, to the namespace service and from the process, a request to create a socket in the second namespace to allow the process to communicate with a second container assigned to the second namespace. The method further includes receiving, at the process and from the namespace service, a file descriptor associated with the socket and a third privilege that allows the process to utilize the socket.”; [0045] – “At 202, a process 110(1) may execute, on a computing device 106 in the cloud computing network 102. In some examples, the process 110(1) may be running in a container 108(1) that is assigned to a first namespace (namespace A 114).”; [0046] – “the namespace service 116 may execute on each of the physical server(s) 106 in the cloud computing environment 102.”; [0047] – “At 206, the namespace service 116 may receive a request from the process 110(1) executing in the container 108(1) in the first namespace. In some examples, the request may include a request to create a socket in the second namespace, such that the process 110(1) in the first namespace may communicate with a container 108(2) in the second namespace without escalating the privileges of the process 110(1) and/or the container 108(1) with respect to the system.”; [0049] – “At 210, the namespace service 116 may provide the process 110(1) with information required to utilize the socket. In some examples, the namespace service 116 may send the file descriptor associated with the socket to the process 110(1), via the UDS 118(1). In some examples, the namespace service 116 may provide the process 110(1) with a rights object associated with the socket, such that the rights object allows the process 110(1) to utilize the socket in the second namespace.”; [0075] – “A process contained by a container 108 may utilize rights and file descriptors associated with a socket in a separate namespace received from the namespace service 116, allowing the process 108 to send and receive communications to multiple namespaces.” In FIG. 1, the first process 110(1) in a first container 108(1) (analogous to the “machine learning inference pod”) may be allowed to access the namespace 114 of second container 108(2) in which a second process 110(2) is executing (“first application pod”) by a namespace service 116 which executes on each physical device.).
Wells is considered to be analogous art to the claimed invention because it is reasonably pertinent to the problem faced by the inventor of loading a machine learning model from the application pod to the machine learning inference pod. As presented in the rejection of claim 1, SENGUPTA teaches loading a model from a container in which an application is executed to another container where the model is executed, and Fong teaches containers in the Kubernetes environment are executed within pods. Wells teaches the Kubernetes architecture uses namespaces for containers to prevent containerized applications from interfering with each other and to deny traffic from different namespaces, which ensures security of the applications (Wells: [0014]-[0015]). Therefore, it would have been obvious to one of ordinary skill in the art to modify the method taught by SENGUPTA in view of Fong to incorporate the teachings of Wells such that the machine learning pod loads the model from the first application pod by accessing the local namespace for the first application pod. The methods of Wells provides the advantages of enabling unprivileged containerized applications to access multiple namespaces without escalating the privileges of the application to an unsafe level, while also maintaining security between workloads in different containers and avoiding network interference (Wells: [0023]).
Regarding claim 4, the combination of SENGUPTA in view of Fong and Wells teaches the method of claim 3, wherein the machine learning model comprises a serialized machine learning model (SENGUPTA: [0088] – “Typically, a model is loaded into an application instance and accelerator via one or more files and/or objects. For example, a model may be loaded into storage (such as storage 361) and then made available to the application instance”.
For clarity of the record, the Examiner would like to point to paragraph [0056] of the Specification of the instant application which recites “’Serialized,’ ‘serialization’ and/or the like refers to a process of translating a data structure and/or object state, for example, into a format that can be stored and/or transmitted and reconstructed later.” Thus, in light of the specification, the model first being loaded into storage 361 and then being made available to the application instance 211 (comprising the “first application pod” is analogous to the model being “serialized”.) and wherein, to load the machine learning model directly from the first application pod, the machine learning inference pod deserializes the machine learning model directly from the first application pod and stores the deserialized machine learning model in an inference engine memory allocated to the machine learning inference pod (SENGUPTA: [0058] – the API call to load the model comprises the model file, e.g. input parameter “model-config.xml”; [0089] – “During EI interface initialization, the trained model (computation graph) is provided as input and profiled, and, subsequently, the application 213 makes inferencing calls via the client library 315 on the application instance 211.”; [0090] – “during the model loading on the EI interface, the trained model is compiled into target (hardware) code. This involves two sub-steps. First, a frontend compiler converts from the trained model file format to an intermediate representation while incorporating target-independent optimizations and analyses. Second, a backend compiler converts from the intermediate representation to machine code with target-dependent optimizations and analyses. This ‘AOT' compilation allows for whole program analysis. The target hardware is a combination of a CPU and accelerator device on the accelerator appliance 221 with the compilation is done on the accelerator appliance 221. […] The runtime on the accelerator appliance 221 instantiates this as an ‘inference runtime object’ in memory and uses it to execute future inferencing calls on the EI interface.”; [0039]-[0040] – “an accelerator appliance 221 (another compute instance) uses a particular configuration of CPU 224, memory, storage, and networking capacity that is to execute a machine learning model of application 213. […] the compute resources of the accelerator appliance 221 are partitioned, allocated, resource governed and isolated across accelerator slots 223 for consistent, sustained performance in a multitenant environment. […] runs an inference engine runtime”. The resources, including memory (and separately recited storage), are partitioned, allocated to, and isolated across accelerator slots each comprising an inference engine (see FIG. 3).
For clarity of the record, the Examiner would like to point to paragraph [0056] of the Specification of the instant application which recites “’Deserialized,’ ‘deserialization’ and/or the like refers to a process of reconstructing a previously serialized data structure or object state, for example, to instantiate the object for consumption, execution, etc.”. Thus, the runtime on the accelerator appliance 221 instantiating the model as an “inference runtime object” in memory is analogous to deserializing the model and storing it in memory allocated to the machine learning inference pod.) wherein the inference engine memory […] and is separate from a local model storage ([0039] – resources allocated include both memory and storage (i.e., memory is separate from storage); [0074] – “memory” components of the accelerator slot include RAM and GPU memory; [0090] – “Additionally, the representation [of the model] can also be serialized to storage if needed, so that the compile phase can be avoided for future instantiations”. This serializing the compiled model to storage, in addition to instantiating the compiled model in memory, is disclosed as an optional additional step to avoid having to compile the model again in the future. Thus, the model may be instantiated in memory separate from a model storage, e.g. the storage of accelerator appliance 221 partitioned and allocated in [0039] for use by the accelerator slots.).
Fong further teaches inference engine memory is local to the edge node ([0049] – “computational resource layer 203 comprises physical resources deployed on and/or otherwise available to worker node 202 such as, but not limited to, CPU, random-access memory, accelerators, etc.”; [0045] – “One or more pods are executed on a worker node.”; [0030] – “an information processing center may comprise an edge-based system that includes one or more edge computing platforms with edge devices and/or edge gateways that operate in accordance with an edge computing paradigm. […] Nodes 102 in computing environment 100 are intended to represent any one or more of the types of devices illustratively mentioned above, as well as other types of devices.”; [0034] – “nodes 102 may comprise one or more servers, gateways, or other types of devices forming systems including, but not limited to, edge computing platforms”).
SENGUPTA teaches the application which has an inference model may be run as containers or a virtual machine (see [0030]), and a container or virtual machine may be launched for the machine learning model at the accelerator slot (see [0097]). Thus, it would have been obvious to one of ordinary skill in the art that these containers for the application containing an inference model and container for the machine learning model taught by SENGUPTA could be implemented in Kubernetes pods as disclosed by Fong. Fong teaches Kubernetes is a prevalent and well-known container orchestration system for managing containerized workloads, and that pods are the smallest execution unit for Kubernetes such that one or more containers are part of a pod (Fong: [0044]-[0045]). Further, it would have been obvious to one of ordinary skill in the art to have implemented the accelerator slot with containerized machine learning inference model taught by SENGUPTA as a pod running on an edge node as taught by FONG, where the resources used such as the memory, e.g., RAM, in which the model is stored is local to the edge node, in order to move data computation closer to the source of the data and reduce the time needed to handle inference requests (Fong: [0017], [0021], and [0027]).
Claims 12-14 recite the same limitations as claims 2-4, respectively; accordingly, claims 12-14 are rejected as being unpatentable over SENGUPTA in view of Fong as applied to claim 11, and further in view of Wells for the same reasons presented with respect to claims 2-4.
Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over SENGUPTA in view of Fong and Wells as applied to claim 4 above, and further in view of Adibowo (U.S. Pub. No. 2022/0215008).
Regarding claim 5, the combination of SENGUPTA in view of Fong and Wells teaches the method of claim 4, but fails to teach wherein the allowing the machine learning inference pod to load the machine learning model directly from the first application pod includes performing a hashing algorithm on the machine learning model prior to deserialization at least in part to prevent duplicate machine learning models from being stored in the inference engine memory.
However, Adibowo teaches performing a hashing algorithm on the machine learning model prior to deserialization at least in part to prevent duplicate machine learning models from being stored in the inference engine memory ([0046] – “A model server to be used for inference is determined (308). […] The ML model identifier can be used to determine whether the ML model is already loaded in memory”; [0049] – “the API server 232 can execute the process 400 to determine the model server 234 that is to be used to execute inference (e.g., 308 of FIG. 3).”; [0050] – “A list of all in-use model servers is provided (402). […] A hash value for each model server in the list is calculated (404). The hash values are calculated using any appropriate hash function (e.g., secure hash algorithm (SHA) 1 (SHA-1 )). In some examples, for each model server 324, the hash value is determined by a concatenating the respective node identifier and the model identifier, and processing the concatenation through the hash function. In this manner, a list of hash values is provided, each hash value corresponding to an in-use model server.” Step 308 involves a hashing algorithm performed on an identifier of the model to select the appropriate model server to send the inference request to. The selected model server then determines if the model is already loaded in memory of the model server at step 314. [0046]-[0047] – “A remote call is executed to the model server (310). For example, the API server 232 makes a network call to the model server 234 to perform the inference. In some examples, the network call to the model server 234 can include a ML model identifier of the identified ML model along with the inference data. […] The network call is received (312). For example, the model server 234 receives the network call and the information regarding the inference request and selected ML model. It is determined whether the ML model is loaded (314). For example, the model server 234 determines whether the identified ML model is already loaded in memory. If the ML model is not loaded in memory, a cache miss is recorded (316), the ML model is retrieved (e.g., from the model storage 240) and is loaded into memory (318), and inference is executed (322). […] If the ML model is loaded in memory, a cache hit is recorded and inference is executed (322).” If the model is not loaded at the selected model server, then the model is retrieved and loaded from a storage at steps 316 and 318 (which involves deserialization – see [0002]). However, if the model is already loaded at the selected model server, the inference is executed with the already loaded model.).
Adibowo is considered to be analogous art to the claimed invention because it is in the same field of computing environments for providing machine learning inference as a service. Adibowo teaches each model server in the machine learning inference platform may be provided as a Kubernetes container (Adibowo: [0036]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of SENGUPTA in view of Fong and Wells to perform a hashing algorithm prior to model loading (and deserialization) as taught by Adibowo. The methods of Adibowo, including performing a hashing algorithm on a machine learning model of an inference request, enables the dynamic adjustment of the number of model servers (i.e., containers) in use for ML model inference to balance user experience and service cost (Adibowo: [0030] and [0036]).
Claims 8 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over SENGUPTA in view of Fong as applied to claims 7 and 16 above, and further in view of GAZZETTI et al. (U.S. Pub. No. 2021/0065063), hereinafter GAZZETTI.
Regarding claim 8, the combination of SENGUPTA in view of Fong teaches the method of claim 7, but fails to teach wherein the edge node executes the inference engine at least in part by establishing a connection with a cloud-based inference service to have the inference operation performed, at least in part, by the cloud-based inference service.
However, GAZZETTI teaches wherein the edge node executes the inference engine at least in part by establishing a connection with a cloud-based inference service to have the inference operation performed, at least in part, by the cloud-based inference service ([0102] – “Each of the cameras may be connected to or integrated within an edge computing device that is used to (locally) process received/detected images. In particular, each edge device may be configured with an edge inference service that is equipped to utilize an object detection model (or other machine learning model) and suitably detect objects captured in the images (or video frames). The edge inference service may also include a request handler (or advisor, such as that described above) that may evaluate the results returned from the object detection model.”; [0103] – “If the confidence results do not meet the threshold (or defined QoS parameters), the request handler may forward requests to a cloud connector that causes a cloud inference service (and/or second object detection model) to be utilized to evaluate or process the images.”; [0106] – “The edge inference model 1614 may include a machine learning model that is capable of, for example, inferring or detecting objects appearing in the input data (e.g., video frames) and providing insight about the position of the objects. The cloud connector 1616 may connect the local service to the cloud platform 1606, forward the input data to the cloud platform 1606 (e.g., in case of unsatisfying result from the edge inference model 1614), and receive inference results from the cloud platform (or a machine learning model deployed thereon).”).
GAZZETTI is considered to be analogous art to the claimed invention because it is in the same field of computing environments for providing machine learning inference as a service. Therefore, it would have been obvious to one of ordinary skill in the art to have modified the teachings of SENGUPTA in view of Fong such that the edge node is to execute the inference engine at least in part by establishing a connection with a cloud-based inference service to have the inference operation performed, at least in part, by the cloud-based inference service. Cloud-based services have nearly unlimited resources, and are therefore able to generate more accurate inferences than what may be generated at the edge node (GAZZETTI: [0105]).
Claim 17 recites the same limitations as claim 8; accordingly, claim 17 is rejected as being unpatentable over SENGUPTA in view of Fong as applied to claim 11, and further in view of GAZZETTI for the same reasons presented with respect to claim 8.
Response to Arguments
Applicant's arguments filed 8/20/2025 have been fully considered but they are not persuasive.
Regarding the rejections under 35 U.S.C. 112(b):
Applicant has amended claims 11-19, but the amendments do not overcome the rejection.
Specifically, “An apparatus comprising: a processor configured to: establish […]; and allow […]” is still reciting intended use for the processor. Thus, it is unclear if the processor actually performs the functions (e.g. “establish…”, “allow…”) or if it is only intended to perform such functions (e.g. the processor has a hardware configuration capable of performing the functions, but does not necessarily perform them).
As such the rejection of claims 11-19 under 35 U.S.C. 112(b) is maintained.
Regarding the rejections under 35 U.S.C. 103:
On page 19, Applicant alleges “Sengupta does not state that it shows allowing a machine learning inference pod, executing at an edge node, to load a machine learning model directly from a first application pod executing at a computing device to enable the machine learning inference pod to perform an inference operation base at least in part on the machine learning model” as recited in claim 1.
Examiner disagrees. Applicant alleges that application 213 of application instance 211 making API calls to the inference engine 325, of accelerator appliance 221, by way of the client library 315 is not analogous to the claim limitations “utilizing the processor of the computing device, allowing the machine learning inference pod, executing at the edge node, to load the machine learning model directly from the first application pod to thereby enable the machine learning inference pod to perform an inference operation based at least in part on the machine learning model.”
Examiner disagrees. The claims do not recite how the processor of the computing device, on which the first application pod (analogous to the application instance 211) is executing, allows the machine learning pod (analogous to the accelerator appliance 221) to load the model. Thus, an API call from the application 213 via the client library 315 (both of which are contained in the application instance 211 executing on the computing device – see FIG. 3, and [0139]) which causes the inference engine 325 of the accelerator appliance 221 to handle loading the model ([0077] – “inference engine (IE) 325 handles model loading and inference execution […] IE 325 receives requests from the client library 315 on customer instance”) is considered to be the application instance 211 “allowing” the inference engine 325 to load the model.
Regarding the arguments that Sengupta does not teach the accelerator appliance 221 loading the model directly from application instance 211, the Examiner points to paragraphs [0058], [0078], [0086] and [0088]-[0089] of Sengupta.
Specifically, [0088] teaches the models may be loaded as files or objects. Regarding Applicant’s arguments that the models are loaded in storage 361 and then made available as recited in [0088], the Examiner would like to point out the language recited is exemplary (“For example, a model may be loaded into storage (such as storage 361)”). Paragraph [0089] further supports this by reciting “the trained model (computation graph) is provided as input and profiled, and, subsequently, the application 213 makes inferencing calls via the client library 315 on the application instance 211. There are many ways to implement this approach.”
As seen in the exemplary API call of paragraph [0058], the model to be loaded is provided as input, see input parameter “model-config.xml”. Further evidence that the model itself may be provided as input to the inference engine can be found in paragraphs [0078] and [0124], which teaches the inference application 213 (part of application instance 211) loads a model to an inference engine 325, and the inference engine 325 provides the uploaded model to a model validator 327 (see [0124]), wherein the model validator 327 “checks user provided model file syntax” and may “convert the provided model to a different format (such as serializing MXNET to JSON).”
Lastly, paragraph [0086] teaches “The EI interface (the interface to the EIA) can be attached to an application instance during instance launch or dynamically attached/detached (to/from) a live instance. The EI interface can be accessed directly using the client library 315”. Thus, the client library 315 in application instance 211 enables direct communication between the application instance 211 and the interface of the EIA (the “elastic inference accelerator”/accelerator slot 223, see [0049], with inference engine 325) hosted by the accelerator appliance 221, as evidenced by the arrows of FIG. 3 connecting the client library 315 directly to inference engine 325. Since the model may be provided as input in an API call sent from the application instance 211 using the client library 315 via the EI interface which allows direct access to the accelerator slot 223 comprising the inference engine 325 of the accelerator appliance 221, the model loaded by the inference engine 325 is loaded “directly” from the application instance 211.
As shown in the rejection of claim 1, Fong is relied upon to teach pods and a machine learning inference pod executing on an edge node. Specifically, Fong discusses the use of pods which comprise one or more containers and are the smallest unit of execution in the “prevalent”, i.e., widespread and commonly known, Kubernetes environment, and discusses implementing pods comprising containers to which inference models are loaded and perform inference operations on nodes in an edge-based system (see [0030]-[0031], [0045], and [0060] as cited in the rejection of claim 1, and [0005], [0044], and FIG. 5, steps 508-510 for further evidence). Even further, Fong teaches pods are able to communicate with one another (see [0045]).
SENGUPTA teaches the application which has an inference model may be run as containers or a virtual machine (see [0030]), and a container or virtual machine may be launched for the machine learning model at the accelerator slot (see [0097]). Thus, it would have been obvious to one of ordinary skill in the art that these containers for the application containing an inference model and container for the machine learning model taught by SENGUPTA could be implemented in Kubernetes pods as disclosed by Fong. Fong teaches Kubernetes is a prevalent and well-known container orchestration system for managing containerized workloads, and that pods are the smallest execution unit for Kubernetes such that one or more containers are part of a pod (Fong: [0044]-[0045]). Further, it would have been obvious to one of ordinary skill in the art to have implemented the machine learning inference model taught by SENGUPTA as a pod running on an edge node as taught by FONG, in order to move data computation closer to the source of the data and reduce the time needed to handle inference requests (Fong: [0017], [0021], and [0027]).
Thus, in combination, SENGUPTA in view of Fong teaches “utilizing the processor of the computing device, allowing the machine learning inference pod, executing at the edge node, to load the machine learning model directly from the first application pod to thereby enable the machine learning inference pod to perform an inference operation based at least in part on the machine learning model” as recited in claim 1, wherein using Kubernetes pods to implement the containerized application instance 211 and the containerized machine learning model used at each accelerator slot of the accelerator appliance 221 taught by SENGUPTA, provides the benefit of enabling management of the containerized workloads in a well-known container orchestration environment, and implementing the machine learning inference pod at an edge node provides the benefit of reducing the time needed to handle inference requests as taught by Fong.
No additional arguments were presented with respect to claims 2-20.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Poothia et al. (U.S. Pub. No. 2022/0083389) teaches executing a machine learning model in a clustered edge system, where GPUs and other hardware accelerators on an edge node may be assigned when a customer is creating a Kubernetes container, and machine learning models are uploaded by an end user (see Abstract, [0052], [0055]-[0056]).
Vittal et al. (U.S. Pub. No. 2023/0108778) teaches Kubernetes pods provide advantages such as easy deployment of applications across a wide variety of computing resources and ensuring service consistency by allowing rapid and efficient re-deployment when a cluster executing a pod crashes (see [0002]). Further, a service in Kubernetes may be an abstraction of multiple pods, allowing organizations to divide up aspects of the same service into discrete containerized applications (see [0065]).
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JENNIFER MARIE GUTMAN whose telephone number is (703)756-1572. The examiner can normally be reached M-F: 9:00 am - 5:00 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kevin Young can be reached at 571-270-3180. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JENNIFER MARIE GUTMAN/Examiner, Art Unit 2194 /KEVIN L YOUNG/Supervisory Patent Examiner, Art Unit 2194