Prosecution Insights
Last updated: October 02, 2026
Application No. 18/448,192

MANAGING INSTANCES OF SERVERLESS FUNCTIONS IN A CLOUD COMPUTING SYSTEM

Final Rejection §101§103
Filed
Aug 11, 2023
Examiner
BLACKBURN, CONNOR IMIOLA
Art Unit
2194
Tech Center
2100 — Computer Architecture & Software
Assignee
International Business Machines Corporation
OA Round
2 (Final)
100%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 100% — above average
100%
Career Allowance Rate
1 granted / 1 resolved
+45.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
10 currently pending
Career history
9
Total Applications
across all art units

Statute-Specific Performance

§101
17.5%
-22.5% vs TC avg
§103
50.9%
+10.9% vs TC avg
§102
24.6%
-15.4% vs TC avg
§112
7.0%
-33.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1 resolved cases

Office Action

§101 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment The Amendment filed 3/19/2026 has been entered. Claims 1-20 are pending in the present Office Action. Claim Rejections - 35 USC § 103 The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action. Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Gunasekaran et al. (Fifer: Tackling Resource Underutilization in the Serverless Era), hereinafter Gunasekaran in view of Ibryam et al. (US 11018965 B1), herinafter Ibryam, and further in view of Deval et al. (US 2019/0107965 A1), hereinafter Deval. Regarding Claim 1, Gunasekaran teaches: A method for managing instances of serverless functions in a cloud computing system, comprising: (see e.g., page [003], lines [034]-[038], left column, "In this paper, we present, Fifer, which to the best of our knowledge, is the first work that employs stage-aware container provisioning and management of function chains for serverless platforms.") obtaining, [by a cloud computing system], a service level objective for a serverless function, wherein the service level objective specifies a maximum number of concurrent requests a compute node of the cloud computing system can process for the serverless function; (see e.g., page [003], lines [033-035], left column, "Leveraging slack allows individual functions to be queued in batches at existing containers without violating the application-level SL Os") (see e.g., page [005], lines [032-038], right column, "On top of these RM frameworks, one can additionally batch the requests by queuing them at every stage of an application, which we name as Request Batching RM (RBRM}. The number of requests which can be queued in a container is defined as the batch size (B_size) of the container. Essentially, B_Size is the length of the processing queue each container.") obtaining, [by a cloud computing system], a command queue length for a graphical processing unit (GPU) disposed on each of a plurality of compute nodes in the cloud computing system; (see e.g., page [009], lines [009-011], left column, "Each container has a local queue of length equal to the number of free-slots in the container.") obtaining, [by a cloud computing system], a request queue length of the serverless function; (see e.g., page [006], lines [039-044], right column, "Fifer utilizes a request queue, which holds all the incoming tasks for each stage. We design a load balancer along with a load monitor for efficiently scaling containers for the application. Since we know the execution time and available slack, the LB can calculate the batch size (B_size) for each stage.") calculating, [by a cloud computing system], a number of instances of the serverless function to deploy in the cloud computing system, wherein the number of instances is determined based on the service level objective and the request queue length of the serverless function; (see e.g., page [006], lines [045-052], right column, "To accurately determine the number of containers needed at every stage which is a function of B_size and queue length, we need to periodically measure the queuing delay due to batching of requests. As shown in Algorithm la, for a given monitoring interval at every stage, the LM monitors the scheduled requests in the last 10s to determine if there are any SLO violations due to queuing delays." )(see e.g., page [006], lines [010-021], right column, "by knowing the available slack and execution time at each stage, we can accurately determine the number of requests that can be executed in a batch in one container.") (see e.g., page [006], lines [037-042], right column, "Fifer utilizes a request queue, which holds all the incoming tasks for each stage 1. We design a load balancer (LB) 2 along with a load monitor that are integrated to each stage {LM} 3 for efficiently scaling containers for the application. Since we know the execution time and available slack, the LB can calculate the batch size (B_size) for each stage.") identifying, [based on the queue for each computer resource], compute nodes from the plurality of compute nodes to deploy each of the number of instances of the serverless function, (see e.g., page [007], lines [010-013], right column, "In Fifer, we design a scheduling policy such that, each stage will submit the request to the container with the least remaining free-slots where the number of free-slots is calculated using the container's batch-size.") The examiner would like to note that the container can be considered a type of computing node. and creating, [by a cloud computing system], an instance of the serverless function on each of the identified compute nodes. (see e.g., page [009], lines [039-048], left column, "In order to efficiently bin-pack containers into fewer nodes, we make modifications to[. .. }") (see e.g., pages [008-009], lines [038-002], "by default, [the framework] creates a worker pod for each job, which in turn handles container creation and scheduling of tasks within the job and destroys the containers after job completion.") Gunasekaran fails to explicitly teach: Use of items including a metrics collector and a scaler collecting and aggregating, by the metrics collector, metrics from each of the plurality of compute nodes in the cloud computing system; maintaining, by the metrics collector, a request queue for each of the serverless functions, wherein the request queue includes the pending requests for the serverless function, across all of the instances of the serverless function; maintaining, by the metrics collector, a command queue for each of the GPUs in the cloud computing system, wherein the command queue tracks the number of pending requests for each GPU from each of the instances of the serverless functions that the GPU services; based at least in part on the command queue length for the GPU of each of the plurality of compute nodes; wherein the compute nodes are identified as the compute nodes that have GPUs with the shortest command queues; However, lbryam teaches, in the context of the scaling of serverless functions, that the various processor types are functional if interchanged between one another. (see e.g., paragraph [043], "All the disclosed methods and procedures described herein can be implemented using one or more computer programs or components. [ ... ] The instructions may be provided as software or firmware, and/or may be implemented in whole or in part in hardware components such as GPUs, ASICs, or any other similar devices.") Use of items including a metrics collector and a scaler (see e.g., column [01], rows [20-36], “The processor is also configured to measure one or more performance metrics of the first serverless function while implementing one or more instances of the first serverless function with one or more services in accordance with the first concurrency limit. Additionally, the processor is configured to determine whether each of the one or more performance metrics meets a respective predetermined threshold, and increase the first concurrency limit to a second concurrency limit in response to determining that each of the one or more performance metrics meets the respective predetermined threshold.”) maintaining, by the metrics collector, a request queue for each of the serverless functions, wherein the request queue includes the pending requests for the serverless function, across all of the instances of the serverless function; (see e.g., columns [07-08], rows [65-10], "For example, once a respective instance of the serverless function 150 completes a request, the respective instance may be executed again to complete another request from the queue 116 of the serverless platform 110.") Gunasekaran and lbryam are considered to be analogous art to the claimed invention as they are reasonably pertinent to the problem faced by the inventor of optimizing instances of serverless functions. Therefore, it would have been prima facie obvious to one of ordinary skill in the art, before the effective filing date, to attempt to combine the systems for managing serverless functions taught by Gunasekaran with the computer elements to run such systems for running serverless functions as taught by lbryam in order to utilize the strengths of the GPU (being able to run the same instruction on many pieces of different data extremely efficiently) in a setting that might call for it (running a group of commands with various input values repeatedly and rapidly) and to increase the potential batch size or otherwise increase throughput. One aspect that could be easily substituted is the container with a GPU based environment. GPUs have an analogous request queue to a container in this situation. Gunasekaran in view of Ibryam fails to explicitly teach: maintaining, by the metrics collector, a command queue for each of the GPUs in the cloud computing system, wherein the command queue tracks the number of pending requests for each GPU from each of the instances of the serverless functions that the GPU services; wherein the compute nodes are identified as the compute nodes that have GPUs with the shortest command queues; However, Deval teaches: maintaining, by the metrics collector, a command queue for each of the GPUs in the cloud computing system, wherein the command queue tracks the number of pending requests for each GPU from each of the instances of the serverless functions that the GPU services; (see e.g., page [], paragraph [0073], Assignable Device Interfaces (ADIs) 428, 430, 432, 434 refer to a set of device backend resources 330 that are allocated, configured and organized as an isolated unit, forming the unit of device sharing. The type and number of backend resources grouped to compose an ADI is device specific. For example, for a network controller device (such as an Ethernet NIC), an ADI may be composed of a set of TX/RX queues and resources associated with a Virtual Switch Interface (VSI). An ADI on a storage controller may be the set of command queues and completion queues associated with a storage namespace. Similarly, an ADI on a GPU may be organized as a set of graphics or compute contexts created on behalf of a virtual-GPU device instance. Depending on the design, ADI on an FPGA device may be an entire Accelerator Function Unit (AFU) or a context on a multi-context capable AFU.”) A compute context of a GPU is very similar to a command, but with more information. It describes all the state information that is required to perform a specific task and information about the task to be performed and is eventually performed based off those details. Maintaining a queue of contexts, therefore is something that is being used to anticipate maintaining a queue of commands. Gunasekaran and lbryam and Deval are considered to be analogous art to the claimed invention as they are reasonably pertinent to the problem faced by the inventor of organizing commands to be completed by computer systems. Therefore, it would have been prima facie obvious to one of ordinary skill in the art, before the effective filing date, to attempt to combine the systems for managing serverless functions taught by Gunasekaran with the method to organize commands as taught by Deval in order to ensure that the various active nodes have their commands properly managed, so that they can be organized to fulfill various service level objectives. Gunasekaran in view of Ibryam and further in view of Deval fails to teach: wherein the compute nodes are identified as the compute nodes that have GPUs with the shortest command queues; However, Cardei teaches: wherein the compute nodes are identified as the compute nodes that have GPUs with the shortest command queues; (see e.g., column [05], rows [09-21], “The particular image processing resource may be selected based on the availability of this and other image processing resources. For example, the system may include multiple instances of a particular type of image processing resource. The scheduler may select an instance that has a shortest processing queue or a processing queue containing therein only lower priority tasks, among other considerations.”) Gunasekaran and lbryam and Deval and Cardei are considered to be analogous art to the claimed invention as they are reasonably pertinent to the problem faced by the inventor of organizing commands to be completed by computer systems. Therefore, it would have been prima facie obvious to one of ordinary skill in the art, before the effective filing date, to attempt to combine the command queue as taught by Gunasekaran and Ibryam and Deval with the organization parameter as taught by Cardei in order to ensure that the various active nodes are all somewhat equally burdened, as that would help to raise efficiency. Regarding claim 2, Gunasekaran teaches: The method of claim 1, further comprising monitoring, by the scaler, the request queue length of the serverless function. (see e.g., page [006], figure 5, an image showing a load monitor for the request queue, as seen below) PNG media_image1.png 190 312 media_image1.png Greyscale Regarding claim 3, Gunasekaran teaches: The method of claim 2, further comprising: determining, by the scaler and based on the request queue length of the serverless function, that the number of compute nodes is insufficient to meet the service level objective; (see e.g., pages [006-007], lines [050-006], “This is because there are not enough containers to handle all the queued requests. In that case, we estimate the additional containers needed using the Estimate_Containers function. By knowing the B_size and number of pending requests in the Queue (PQten), the function can estimate the number of containers”) identifying, by the scaler and based at least in part on the command queue length for the GPU of each of the plurality of compute nodes, an additional compute node from the plurality of compute nodes to deploy an additional instance of the serverless function; (see e.g., page [009], lines [040 – 048], left column, “we make modifications to the MostRequestedPriority scheduling policy in Kubernetes such that it always chooses the node with the least-available-resources to satisfy the Pod requirements. […] We determine idle cores in a node by calculating the difference between number of cores in a node and the sum of cpu-shares for all allocated pods in that node.”) and creating, by the scaler, the additional instance of the serverless function on each of the additional compute node. (see e.g., page [006], lines [003-004], left column, “Additional containers would be spawned if the arrival rate [of requests] increases”) Regarding claim 4, Gunasekaran teaches: The method of claim 3, wherein the additional compute node is further identified based on an available memory capacity of the GPU of each of the plurality of compute nodes. (see e.g., page [009], lines [039-048], left column, “In order to efficiently bin-pack containers into fewer nodes, we make modifications to the Most Requested Priority scheduling policy in Kubernetes such that it always chooses the node with the least-available-resources to satisfy the Pod requirements. For our experiments, each container requires 0.5 CPU-core and memory within 1GB. Hence, we set the CPU limit for all containers to be 0.5. We determine idle cores in a node by calculating the difference between number of cores in a node and the sum of cpu-shares for all allocated pods in that node.”) Regarding claim 5, Gunasekaran teaches: The method of claim 1, wherein the compute nodes are identified from the plurality of compute nodes by ranking the command queue lengths for the GPU of each of the plurality of compute nodes and selecting the number of lowest ranked compute nodes. (see e.g., page [009], lines [039-048], left column, “In order to efficiently bin-pack containers into fewer nodes, we make modifications to the Most Requested Priority scheduling policy in Kubernetes such that it always chooses the node with the least-available-resources to satisfy the Pod requirements. For our experiments, each container requires 0.5 CPU-core and memory within 1GB. Hence, we set the CPU limit for all containers to be 0.5. We determine idle cores in a node by calculating the difference between number of cores in a node and the sum of cpu-shares for all allocated pods in that node.”) Regarding claim 6, Gunasekaran teaches: The method of claim 1, wherein the compute nodes are identified from the plurality of compute nodes based at least in part on one or more of a GPU utilization of each of the plurality of compute nodes and an available memory capacity of the GPU of each of the plurality of compute nodes. (see e.g., page [007], lines [010-013], right column, “In Fifer, we design a scheduling policy such that, each stage will submit the request to the container with the least remaining free-slots where the number of free-slots is calculated using the container’s batch-size.”) Regarding claim 7, Gunasekaran teaches: The method of claim 1, wherein the number of instances is calculated by rounding up a result of dividing the request queue length of the serverless function by the service level objective of the serverless function to a next whole integer. (see e.g., page 8, Algorithm 1, line 12, “current_req <- len(stage.containers) * batchSize”) Current_req is equivalent to the request queue length, while stage.containers.length is equivalent to the number of instances. The service level objective of a serverless function (the number of concurrent functions runnable) is equivalent to a batch size. Regarding claim 8, Gunasekaran teaches: The method of claim 1, wherein the number of instances is calculated by rounding up a result of dividing the request queue length of the serverless function by the service level objective of the serverless function to a next whole integer and adding a margin constant that is a positive integer. (see e.g., page 8, Algorithm 1, line 12, “current_req <- len(stage.containers) * batchSize”) Current_req is equivalent to the request queue length, while stage.containers.length is somewhat equivalent to the number of instances. The service level objective of a serverless function (the number of concurrent functions runnable) is equivalent to a batch size. (see e.g., page 8, Algorithm 1, line 15, “est_containers <- (PQ_LEN – current_req)” This takes the previous output and uses an integer (presumably positive, as it is a length value) as further processing. Regarding claim 9, Gunasekaran teaches: A computing system having a memory having computer readable instructions and one or more processors for executing the computer readable instructions, the computing system including a metrics collector and a scaler, the computer readable instructions controlling the one or more processors to perform operations comprising: the method of claim 1. Accordingly, claim 9 is rejected as being unpatentable over Gunasekaran in view of lbryam and further in view of Deval and further in view of Cardei for the same reasons presented with respect to claim 1. Claims 10-16 recite substantially the same limitations as those recited in claims 2-8, applied to the system of claim 9. Accordingly, claims 10-16 are rejected as being unpatentable over Gunasekaran in view of lbryam and further in view of Deval and further in view of Cardei for the same reasons presented with respect to claims 2-8. Regarding claim 17, Gunasekaran teaches: A computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a processor to cause the processor to perform operations in a cloud computing system including a metrics collector and a scaler, the operations comprising: the method of claim 1. Accordingly, claim 17 is rejected as being unpatentable over Gunasekaran and lbryam and Deval and Cardei for the same reasons presented with respect to claim 1. Claims 18-20 recite substantially the same limitations as those recited in claims 2-4, applied to the system of claim 17. Accordingly, claims 18-20 are rejected as being unpatentable over Gunasekaran in view of lbryam and further in view of Deval and further in view of Cardei for the same reasons presented with respect to claims 2-4. Response to Arguments Applicant’s arguments with respect to the rejections under 35 U.S.C. 101 have been fully considered and are persuasive. The rejections under 35 U.S.C. 101 have been withdrawn. Applicant’s arguments with respect to the rejections under 35 U.S.C 103, claims 1-20 have been considered but are not persuasive, either due to them being moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument, or because the argument itself was not persuasive as follows: Applicant argues the references cited in the previous Office Action fail to teach “maintaining, by the metrics collector, a request queue for each of the serverless functions, wherein the request queue includes the pending requests for the serverless function, across all of the instances of the serverless function” as recited in amended claim 1. Examiner disagrees. Ibryam (US 11018965 B1) teaches maintaining, by the metrics collector, a request queue for each of the serverless functions, wherein the request queue includes the pending requests for the serverless function, across all of the instances of the serverless function (see e.g., columns [07-08], rows [65-10], "For example, once a respective instance of the serverless function 150 completes a request, the respective instance may be executed again to complete another request from the queue 116 of the serverless platform 110.") Although a direct reference to a “metrics collector” is not recited as being the source of the maintenance, Ibryam teaches the request queue, including the pending requests, managed across the serverless platform by the same system that collects metrics, and can be considered a “metrics collector” (see e.g., column [01], rows [20-36], “The processor is also configured to measure one or more performance metrics of the first serverless function while implementing one or more instances of the first serverless function with one or more services in accordance with the first concurrency limit. Additionally, the processor is configured to determine whether each of the one or more performance metrics meets a respective predetermined threshold, and increase the first concurrency limit to a second concurrency limit in response to determining that each of the one or more performance metrics meets the respective predetermined threshold.”) (for clarity’s sake, something that increases the amount (i.e., threshold) of resources available is what the Examiner is considering a “scaler”. The rejection in full is above for convenience, modified to reject the amended claims. Applicant claims that CPU/container/pod resource allocation that uses queues does not read on GPU resource allocation that uses queues. A GPU is a specialized processor, and therefore because Ibryam teaches using both, one of ordinary skill in the art would know that it would be advantageous to apply the method to a GPU as well. However, the full rejection is described under the 103 heading. Applicant claims Examiner has not provided a reason why a person of ordinary skill would have modified previous references to arrive at the amended claims. This is provided in the 103 section. The other arguments pointed towards amended claims are considered persuasive, however the newly provided references in the 103 section do teach the missing teachings. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Connor Imiola Blackburn whose telephone number is (571)272-6547. The examiner can normally be reached M-Th 7-5. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kevin Young can be reached at (571) 270 - 3180. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /C.I.B./Examiner, Art Unit 2194 /KEVIN L YOUNG/Supervisory Patent Examiner, Art Unit 2194
Read full office action

Prosecution Timeline

Aug 11, 2023
Application Filed
Mar 16, 2026
Non-Final Rejection mailed — §101, §103
May 28, 2026
Interview Requested
Jun 11, 2026
Applicant Interview (Telephonic)
Jun 11, 2026
Examiner Interview Summary
Jun 12, 2026
Response Filed
Sep 04, 2026
Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748644
EFFICIENT GENERATION OF APPLICATION PROGRAMMING INTERFACE CALLS USING LANGUAGE MODELS, DATA TYPES, AND ENRICHED SCHEMA
2y 8m to grant Granted Sep 29, 2026
Study what changed to get past this examiner. Based on 1 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
100%
Grant Probability
99%
With Interview (+0.0%)
2y 10m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 1 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month