Prosecution Insights
Last updated: October 02, 2026
Application No. 17/933,343

CLOUD COMPUTING QOS METRIC ESTIMATION USING MODELS

Final Rejection §103§112
Filed
Sep 19, 2022
Examiner
TRAN, KENNETH PHUOC
Art Unit
2197
Tech Center
2100 — Computer Architecture & Software
Assignee
Dell Products L.P.
OA Round
2 (Final)
31%
Grant Probability
At Risk
3-4
OA Rounds
0m
Est. Remaining
98%
With Interview

Examiner Intelligence

Grants only 31% of cases
31%
Career Allowance Rate
4 granted / 13 resolved
-24.2% vs TC avg
Strong +67% interview lift
Without
With
+66.7%
Interview Lift
resolved cases with interview
Typical timeline
3y 8m
Avg Prosecution
27 currently pending
Career history
51
Total Applications
across all art units

Statute-Specific Performance

§101
15.5%
-24.5% vs TC avg
§103
68.5%
+28.5% vs TC avg
§102
4.1%
-35.9% vs TC avg
§112
11.4%
-28.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 13 resolved cases

Office Action

§103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This action is responsive to the Applicant’s amendments filed on 09/26/2025. Claims 1-20 remain pending in the application. Claims 1, 5, 7-8, 11, 15, and 17-18 have been amended. Any examiner’s note, objection, and rejection not repeated is withdrawn due to Applicant’s amendment. Information Disclosure Statement The information disclosure statement (IDS) submitted on 09/19/2022 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Examiner’s Note The Examiner cites particular columns, paragraphs, figures, and line numbers in the references as applied to the claims below for the convenience of the applicant. Although the specified citations are representative of the teachings in the art and are applied to the specific limitations within the individual claim, other passages and figures may also apply. It is respectfully requested that, in preparing responses, the Applicant fully consider the references in its entirety as potentially teaching all or part of the claimed invention, as well as the context of the passage as taught by the prior art or disclosed by the Examiner. Claim Objections Claim 8-9 and 18-19 are objected to because of the following informalities: Claims 8 and 18 recite “... wherein an input to the physical machine vector comprises...”. The use of “physical machine vector” is inconsistent with the claim set language reciting a “physical machine generator” and a “physical machine queue”. Any claim not explicitly mentioned is objected to due to dependency on an objected claim. Appropriate correction is required. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-2 and 11-12 are rejected under 35 U.S.C. 103 as being unpatentable over Colmenares Dias et al. (US 20230141570 A1), hereafter Colmenares, in view of Liu et al. (US 20220237838 A1) hereafter Liu, further in view of Liu et al. (US 20250335257 A1) hereafter Liu2, further in view of Tateno et al. (US 20230421653 A1) hereafter Tateno, further in view of Young et al. (US 10409649 B1) hereafter Young. Regarding claim 1, Colmenares teaches: A method comprising: receiving a request at an infrastructure (Paragraph 1, Abstract; “The present disclosure generally relates to admission control for online data systems”; “Upon receiving a current server query from a client”.); determining an occupancy status of a load balancing engine in the infrastructure (Abstract; A current queue wait time is estimated based on a number of queries currently in the queue and the estimated processing times of query types for each of the queries currently in the queue; and Paragraph 22; “The broker host server(s) 145 receive query requests and broadcast one or more sub-queries to the shard host servers”, where the broker host server corresponds to the load balancer and the number of queries currently queued is considered as occupancy status of the load balancer because per Figure 2 where Query Queue 220 is within the broker host server.); estimating a first metric based on the occupancy status of the load balancing engine with a first estimating engine (Paragraph 11; The admission controller estimates a current queue wait time based on a number of queries currently in the queue, where the admission controller corresponds to a first estimating engine, and the queue wait time is considered as the first metric.); estimating a second metric with a second estimating engine (Paragraph 11; The admission controller estimates the estimated processing times, where the admission controller is also corresponds to the second estimating engine, and the processing time is considered as the second metric.); determining an estimated total metric from the first metric and the second metric (Paragraph 30; the response time of a query, RT(Q), is the sum of the processing time of the query, PT(Q), the wait time between enqueuing and dequeuing the query, WT(Q), and any additional time the server host takes to handle the query, which will be treated as negligible/zero for the sake of this illustration.); and performing an action when the estimated total metric is below a quality of service value (Abstract; “The server query is rejected from being added to the queue in response to determining the estimated response time does not satisfy a service level objective”, where the service level objective is considered as the quality of service value, and the rejection of adding the query to queue is the action performed). Colmenares does not teach determining an occupancy status of physical machines; estimating a first metric based on a first noise vector with a first estimating engine; estimating a second metric based on the occupancy status of the physical machines and a second noise vector with a second estimating engine. However, Liu teaches: determining an occupancy status of physical machines (Paragraph 365; “In at least one embodiment, a scheduler may be used to track resource requirements of applications or containers, current usage or planned usage of these resources, and resource availability”, where the resource requirements of applications or containers, current usage or planned usage of these resources, and resource availability are considered as occupancy status of physical machines.); estimating a first metric based on a first noise vector with a first estimating engine (Paragraph 56; “In at least one embodiment, a set of structural feature vectors are obtained 402... these structural and appearance feature vectors are provided 406 as input to a slot attention transformer that is able to generate a set of transformed feature vectors from these input feature vectors. these transformed feature vectors are provided 408 as input to a generator, such as a generative adversarial network (GAN)”, where the transformed structural feature vectors are considered as a first noise vector, and generator is considered as the estimating engine.); estimating a second metric based on the occupancy status of the physical machines and a second noise vector with a second estimating engine (Paragraph 365; “In at least one embodiment, a scheduler may be used to track resource requirements of applications or containers, current usage or planned usage of these resources, and resource availability”; and Paragraph 56, “a set of appearance feature vectors can also be obtained 404. these structural and appearance feature vectors are provided 406 as input to a slot attention transformer that is able to generate a set of transformed feature vectors from these input feature vectors. these transformed feature vectors are provided 408 as input to a generator, such as a generative adversarial network (GAN)”, where the resource requirements of applications or containers, current usage or planned usage of these resources, and resource availability are considered as occupancy status of physical machines, same as above; the transformed appearance feature vectors are considered as a second noise vector; and generator is considered as the estimating engine.). Colmenares and Liu are considered to be analogous to the claimed invention because they are in the same field of load/resource availability and QOS metric evaluation. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Colmenares to incorporate the teachings of Liu and combine the introduction of occupancy status of the physical machines and the noise vectors as inputs to determine the metrics presented in Colmenares. Specifically, supplying a noise vector as an input to estimate the first metric, and a second noise vector and occupancy status as inputs to estimate the second metric. This combination results in utilizing more information and introducing Generative Adversarial Network (GAN) models to help determine the metrics, given noise vectors are inputs that are particular to Machine Learning models like GAN. A person having ordinary skill in the art before the effective filing date of the claimed invention would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of a more accurate estimation of the metrics by leveraging GAN through introducing inputs such as noise vectors. Specifically, improving the estimation that it is highly realistic and couldn't be discriminated between real-world data through the feedback of the discriminator of GAN (Liu: this generator can be a generative adversarial network (GAN) trained to generate images that a discriminator of that GAN cannot determine to be synthetic and not a "real" image or representation. [0056]). Colmenares in view of Liu does not teach that the first metric is further based on the occupancy status of the physical machines; the first estimating engine comprising a first conditional generative adversarial network; the second metric is further based on the occupancy status of the load balancing engine; and the second metric is further based on a size of the physical machines; the second estimating engine comprising a second conditional generative adversarial network. However, Liu2 teaches: a first and second neural network (Paragraph 76; “obtaining an interference model between operators in each of the second data processing models. Referring to FIG. 6B, the interference model is constructed by quantifying shared resource requirements, performance testing, and building an analytical model... In the building of the analytical model, linear regression models and neural network models can be employed to model the interference status between the operators, obtaining the interference model between the operators in each of the second data processing models.”, which explicitly teaches construction of an interference model that may be implemented using neural network “models”, explicitly disclosed in plural form, which is used to learn patterns from performance testing and resource sharing data to predict resource behavior.). Colmenares, Liu, and Liu2 are considered to be analogous to the claimed invention because they are in the same field of load balancing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Colmenares in view of Liu to incorporate the teachings of Liu2 and have the first and second estimating engines of Colmenares in view of Liu be neural networks as taught by Liu2. A person of ordinary skill in the art before the effective filing date of the claimed invention would have recognized the use of neural network models to perform estimations to be a known machine learning approach for modeling relationships in resource behavior, yielding the predictable result of more accurate and adaptive estimation of resources for QoS scheduling decisions. Colmenares in view of Liu, further in view of Liu2 does not teach that the neural network comprises a conditional generative adversarial network; the first metric is further based on the occupancy status of the physical machines; the second metric is further based on the occupancy status of the load balancing engine; and the second metric is further based on a size of the physical machines. However, Tateno teaches: a conditional generative adversarial network (Paragraph 164; “In a case where the above two sections are reconfigured as one processing section, the intervention analysis section 43 and the intervention material generation section 45 is configured, for example, by Conditional GAN (Generative Adversarial Nets)”). Colmenares, Liu, Liu2, and Tateno are considered to be analogous to the claimed invention because they are in the same field of load balancing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Colmenares in view of Liu, further in view of Liu2 to incorporate the teachings of Tateno and have the neural networks of the first and second estimating engines be conditional GANs. A person of ordinary skill in the art before the effective filing date of the claimed invention would have recognized that implementing the neural network-based estimating engines using the conditional GAN architecture of Tateno would be a known neural network architecture for implementing predictive processing, yielding the predictable result of performing the same predictive estimation of a neural network model using a known alternative neural network implementation capable of generating predictions conditioned on input information. Colmenares in view of Liu, further in view of Liu2, further in view of Tateno does not teach the first metric is further based on the occupancy status of the physical machines; the second metric is further based on the occupancy status of the load balancing engine; and the second metric is further based on a size of the physical machines. However, Young teaches: occupancy status of the physical machines (Col. 4, lines 11-20; “The requests 118 may be received by the load balancer 120 or by one or more other systems of the computing resource service provider 104, such as a request listener not illustrated in FIG. 1 for simplicity, and directed to the load balancer 120. The load balancer 120 may be a computer system or virtual computer system configured to distribute the request 118 to one or more computer systems, supported by physical hosts 142, in order to optimize resource utilization and/or avoid overloading a particular computer system.”, in which Col. 8, lines 25-33 further discloses “a load balancing event may include an increase in request traffic, a decrease in request traffic, the addition or removal of capacity to one or more of the computer systems assigned to the load balancer”, which explicitly discloses monitoring the utilization state of the load balancer and capacity of the computer systems assigned to the load balancer, while Col. 4, lines 11-20 discloses distributing requests based on resource utilization and avoiding overload of computer systems supported by physical hosts. A person of ordinary skill in the art would have recognized that making load balancing decisions would require information indicative of the utilization state of the physical machines, corresponding to the claimed occupancy status of the physical machines.); occupancy status of the load balancing engine (Col. 8, lines 25-33; “For example, a load balancing event may include an increase in request traffic, a decrease in request traffic, the addition or removal of capacity to one or more of the computer systems assigned to the load balancer, a customer request to increase or decrease the load balancer capacity, an over-utilization or under-utilization of load balancer capacity, or other event that may necessitate allocation or deallocation of computing resources to a particular load balancer.”, where utilization of the load balancer capacity reflects the extent to which the available resources of the load balancer are occupied by current demands, corresponding to the claimed occupancy status of the load balancing engine. It further discloses that utilization information is used to determine when computing resources are to be allocated or deallocated, demonstrating that the occupancy status represents an operational state of the load balancing engine.); size of the physical machines (Col. 13, lines 31-35; “Information corresponding to the backend computer systems may include the utilization and maximum capacity of the backend computer systems, such as the processing power and network bandwidth of the backend computing systems.”, where the backend computing systems are described in Col. 11, lines 39-45 as “The physical hosts 542 may include any computer system or virtual computer system described above. A virtualization layer 544 operated by the computing resources service provider 502 enables the physical hosts 542 to be used to provide computational resources upon which one or more load balancers 520 may operate.”.). Colmenares, Liu, Liu2, Tateno, and Young are considered to be analogous to the claimed invention because they are in the same field of load balancing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Colmenares in view of Liu, further in view of Liu2, further in view of Tateno to incorporate the teachings of Young and have the first metric additionally comprise the occupancy status of the physical machines, and have the second metric additionally comprise the occupancy status of the load balancing engine, and a size of the physical machines. A person of ordinary skill in the art before the effective filing date of the claimed invention would have recognized that incorporating the occupancy status of the physical machines into the first metric and the occupancy status of the load balancing engine together with the capacity of the physical machines into the second metric as known infrastructure characteristics that affect system workload, resource availability, and request distribution, whose implementation in their respective estimating engines would yield the predictable result of generating complementary metrics that more accurately characterize the operating state of the system, enabling a more informed combined metric for QoS resource allocation decisions. Claim 11 recites similar limitations as those of claim 1, additionally reciting a non-transitory storage medium having stored therein instructions that are executable by one or more hardware processors to perform operations. Colmenares teaches: A non-transitory storage medium having stored therein instructions that are executable by one or more hardware processors to perform operations (Paragraph 86; “carries out the computer-implemented methods 300 and 400 in response to its processor executing a computer program (e.g., a sequence of instructions) contained in a memory or other non-transitory machine-readable storage medium.”). Claim 11 is rejected for similar reasons as those of claim 1. Regarding claim 2, Colmenares in view of Liu, further in view of Liu2, further in view of Tateno, further in view of Young teach the method of claim 1. Colmenares teaches: wherein the first metric relates to a response time measured from receiving the request to assigning the request to a physical machine, wherein the request is moved from a queue of the load balancing engine to a queue of the physical machine (Paragraphs 30-32; “the wait time between enqueuing and dequeuing the query, WT(Q)”; “dequeued from the query queue 220 for processing by one or more shard hosts 160”, where the time between enqueuing and dequeuing the query is considered as the time the request is moved from a queue of the load balancing engine to a queue of the physical machine because dequeuing means the request is moved to the queue of the physical machine per reference). Claim 12 recites similar limitations as those of claim 2. Claim 12 is rejected for similar reasons as those of claim 2. Claims 3-7 and 13-17 are rejected under 35 U.S.C. 103 as being unpatentable over Colmenares in view of Liu, further in view of Liu2, further in view of Tateno, further in view of Young, further in view of Ferguson et al. (US 5504894 A), hereafter Ferguson. Regarding claim 3, Colmenares in view of Liu, further in view of Liu2, further in view of Tateno, further in view of Young teach the method of claim 1. Colmenares in view of Liu, further in view of Liu2, further in view of Tateno, further in view of Young does not teach wherein the second metric relates to a response time measured from receiving the request at the physical machine to assigning the request to a virtual machine operating on the physical machine or to sending a response to the request. However, Ferguson teaches: wherein the second metric relates to a response time measured from receiving the request at the physical machine to assigning the request to a virtual machine operating on the physical machine or to sending a response to the request (Col. 9, lines 32-35; “This component predicts the effects of a proposed routing decision on the response times of all transactions currently in the system.”, where the response time for any particular completed transaction is a Liu2th of time which has elapsed from an arrival time when said any particular completed transaction arrived at said computer system for handling and a completion time when said any particular completed transaction was completed by the computer system.). Colmenares, Liu, Liu2, Tateno, Young, and Ferguson are considered to be analogous to the claimed invention because they are in the same field of load balancing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Colmenares in view of Liu, further in view of Liu2, further in view of Tateno, further in view of Young to incorporate the teachings of Ferguson and combine the response time of the physical machine with the system and methods of Colmenares in view of Liu, further in view of Liu2, further in view of Tateno, further in view of Young. This combination results in the total metric of Claim 1 being a metric of response time, which part of it is constituted by the time the physical machine takes to respond/process when the request arrives the physical machine. A person having ordinary skill in the art before the effective filing date of the claimed invention would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of assuring the total estimated metric is not below the quality of service value through the estimated response time of each of the physical machines and evaluating each routing choices to the different physical machines (Ferguson: Whenever a transaction arrives, the workload manager considers a number of different possible transaction servers to which that arriving transaction could be routed and predicts estimated new values for the class performance indices for each of the considered routing choices. An overall goal satisfaction index is determined for each one and the routing choice corresponding to the best overall goal satisfaction index is selected as the routing choice. [Abstract]). Claim 13 recites similar limitations as those of claim 3. Claim 13 is rejected for similar reasons as those of claim 3. Regarding claim 4, Colmenares in view of Liu, further in view of Liu2, further in view of Tateno, further in view of Young teach the method of claim 1. Colmenares teaches: the first metric (Paragraph 11; The admission controller estimates a current queue wait time based on a number of queries currently in the queue, where the admission controller corresponds to a first estimating engine, and the queue wait time is considered as the first metric). the second metric (Paragraph 11; The admission controller estimates the estimated processing times, where the admission controller is also corresponds to the second estimating engine, and the processing time is considered as the second metric). Young teaches: the load balancing engine (Col. 8, lines 25-33; “For example, a load balancing event may include an increase in request traffic, a decrease in request traffic, the addition or removal of capacity to one or more of the computer systems assigned to the load balancer, a customer request to increase or decrease the load balancer capacity, an over-utilization or under-utilization of load balancer capacity, or other event that may necessitate allocation or deallocation of computing resources to a particular load balancer.”). Colmenares in view of Liu, further in view of Liu2, further in view of Tateno, further in view of Young does not teach relating to a response time of the load balancing engine, and relating to a response time of the physical machines. However, Ferguson teaches: relating to a response time of the load balancing engine, and relating to a response time of the physical machines (Col. 9, lines 32-35; “This component predicts the effects of a proposed routing decision on the response times of all transactions currently in the system”, where response time for any particular completed transaction being a Liu2th of time which has elapsed from an arrival time when said any particular completed transaction arrived at said computer system for handling and a completion time when said any particular completed transaction was completed by the computer system.). Colmenares, Liu, Liu2, Tateno, Young, and Ferguson are considered to be analogous to the claimed invention because they are in the same field of load balancing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Colmenares in view of Liu, further in view of Liu2, further in view of Tateno, further in view of Young to incorporate the teachings of Ferguson and have the first metric relate to a response time of the load balancing engine, and the second metric relate to a response time of the physical machines. A person of ordinary skill in the art before the effective filing date of the claimed invention would have recognized that defining the first metric as relating to response time of the load balancing engine and the second metric as relating to the response time of physical machines are a known approach in distributed systems for decomposing latency into performance contributions, whose implementation would yield the predictable result of enabling finer identification of performance bottlenecks and more effective QoS management across the load balancing layer and compute infrastructure. Claim 14 recites similar limitations as those of claim 4. Claim 14 is rejected for similar reasons as those of claim 4. Regarding claim 5, Colmenares in view of Liu, further in view of Liu2, further in view of Tateno, further in view of Young, further in view of Ferguson teach the method of claim 4. Liu2 teaches: the estimating engine (Paragraph 76; “obtaining an interference model between operators in each of the second data processing models. Referring to FIG. 6B, the interference model is constructed by quantifying shared resource requirements, performance testing, and building an analytical model... In the building of the analytical model, linear regression models and neural network models can be employed to model the interference status between the operators, obtaining the interference model between the operators in each of the second data processing models.”, which explicitly teaches construction of an interference model that may be implemented using neural network “models”, explicitly disclosed in plural form, which is used to learn patterns from performance testing and resource sharing data to predict resource behavior); a physical machine generator (Paragraph 76; “neural network models can be employed to model the interference status between the operators... based on performance testing and shared resource requirements”, which teaches a neural network based model that predicts system performance using physical resource inputs and makes decisions accordingly, corresponding to a physical machine generator that makes estimates of physical system behavior based on resource conditions.). Young teaches: a load balancing generator (Col. 4, lines 11-20; “The requests 118 may be received by the load balancer 120 or by one or more other systems of the computing resource service provider 104, such as a request listener not illustrated in FIG. 1 for simplicity, and directed to the load balancer 120. The load balancer 120 may be a computer system or virtual computer system configured to distribute the request 118 to one or more computer systems, supported by physical hosts 142, in order to optimize resource utilization and/or avoid overloading a particular computer system.”, the load balancer that evaluates system condition and makes request distribution decisions corresponding to a load balancing generator.). Ferguson teaches: estimating the response time (Col. 9, lines 32-35; “This component predicts the effects of a proposed routing decision on the response times of all transactions currently in the system.”). A person of ordinary skill in the art before the effective filing date of the claimed invention would have recognized that implementing the estimating engines of Colmenares in modular form incorporating both a physical machine generator and load balancing generator as respectively suggested by Liu2 and Young as a known method in the art because distributed computing systems commonly decompose performance modeling and control functions into multiple modular components that contribute to system performance estimation. Based on Young teaching load balancing decision making based on resource utilization to manage request distribution across compute resources and Liu2 teaching neural network-based modeling of physical machine behavior for estimating system performance, in view of Ferguson, which teaches estimating response time based on combined contributions from multiple components, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to include both the load balancing generator and physical machine generator within each estimating engine such that each engine generates complementary estimates reflecting both request distribution behavior and physical machine performance, yielding the predictable result of a unified response time estimation across the distributed system. Further, a person of ordinary skill in the art would have recognized that the load balancing generator and physical machine generator of the estimating engines would be trained using respective datasets corresponding to their respective prediction tasks. Because Young teaches making load balancing decisions based on utilization conditions and Liu2 teaches using machine learning models trained on performance and resource-related data to estimate physical system behavior, it would have been obvious that each predictive component requires training on data relevant to its respective function in order to generate accurate estimates, yielding the predictable result of enabling each generator to produce more accurate outputs for response time estimation and resource management. Claim 15 recites similar limitations as those of claim 5. Claim 15 is rejected for similar reasons as those of claim 5. Regarding claim 6, Colmenares in view of Liu, further in view of Liu2, further in view of Tateno, further in view of Young, further in view of Ferguson teach the method of claim 5. Colmenares teaches: real response times associated with a number of requests in a load balancing queue (Paragraph 11; “The admission controller estimates a current queue wait time based on a number of queries currently in the queue”, in which the number of queries currently in the queue is considered as a number of requests in a load balancing queue which the occupancy values include. Paragraph 30 further discloses “Separate query types often have different processing time distributions that vary over time. As such, admission controller 150 maintains approximations for these distributions in histograms which admission controller 150 periodically updates at run time by tracking timing metrics during the processing of queries or similar measurements of processing times”, and Paragraph 65 discloses “As a result, the admission controller 150 can track/measure and update actual processing times for query types in corresponding histograms”.). Young teaches: wherein the load balancing model implicitly learns a distribution associated with occupancy values associated with the load balancing engine, the occupancy values including a number of active virtual machines in the infrastructure (Col. 4, lines 11-20; “The requests 118 may be received by the load balancer 120 or by one or more other systems of the computing resource service provider 104, such as a request listener not illustrated in FIG. 1 for simplicity, and directed to the load balancer 120. The load balancer 120 may be a computer system or virtual computer system configured to distribute the request 118 to one or more computer systems, supported by physical hosts 142, in order to optimize resource utilization and/or avoid overloading a particular computer system.”, in which Col. 8, lines 25-33 further discloses “a load balancing event may include an increase in request traffic, a decrease in request traffic, the addition or removal of capacity to one or more of the computer systems assigned to the load balancer”, which explicitly discloses monitoring the utilization state of the load balancer and capacity of the computer systems assigned to the load balancer, while Col. 4, lines 11-20 discloses distributing requests based on resource utilization and avoiding overload of computer systems supported by physical hosts.). A person of ordinary skill in the art before the effective filing date of the claimed invention would have recognized that the load balancing behavior of Young, which distributes requests based on optimizing resource utilization and avoiding overload of computer systems, involves processing varying utilization conditions of the load balancing environment. As a result, the neural network-based implementation of the load balancing model must learn a distribution over occupancy states of the load balancing engine, yielding the predictable benefit of enabling the load balancing model to more accurately route requests under dynamic system load conditions by capturing variations in system utilization over time. Claim 16 recites similar limitations as those of claim 6. Claim 16 is rejected for similar reasons as those of claim 6. Regarding claim 7, Colmenares in view of Liu, further in view of Liu2, further in view of Tateno, further in view of Young, further in view of Ferguson teach the method of claim 5. Colmenares teaches: real response times associated with a number of requests in a physical machine queue (Paragraph 11; “The admission controller estimates a current queue wait time based on a number of queries currently in the queue”, in which the number of queries currently in the queue is considered as a number of requests in a load balancing queue which the occupancy values include. Paragraph 30 further discloses “Separate query types often have different processing time distributions that vary over time. As such, admission controller 150 maintains approximations for these distributions in histograms which admission controller 150 periodically updates at run time by tracking timing metrics during the processing of queries or similar measurements of processing times”, and Paragraph 65 discloses “As a result, the admission controller 150 can track/measure and update actual processing times for query types in corresponding histograms”.). Liu2 teaches: a number of requests in a data structure (Paragraph 53; “when the number of the second tasks is greater than one, a coordinated resource allocation sub-method is executed to obtain a resource allocation scheme of the server. In addition, when there is only one second task, mutual interactions between different tasks will not be considered during the resource allocation", explicitly disclosing an evaluation of the number of tasks currently pending in the data structure.). Young teaches: physical machines (Col. 4, lines 11-20; “The requests 118 may be received by the load balancer 120 or by one or more other systems of the computing resource service provider 104, such as a request listener not illustrated in FIG. 1 for simplicity, and directed to the load balancer 120. The load balancer 120 may be a computer system or virtual computer system configured to distribute the request 118 to one or more computer systems, supported by physical hosts 142, in order to optimize resource utilization and/or avoid overloading a particular computer system.”, in which Col. 8, lines 25-33 further discloses “a load balancing event may include an increase in request traffic, a decrease in request traffic, the addition or removal of capacity to one or more of the computer systems assigned to the load balancer”, which explicitly discloses monitoring the utilization state of the load balancer and capacity of the computer systems, corresponding to the physical machines, assigned to the load balancer); a size of the physical machines (Col. 13, lines 31-35; “Information corresponding to the backend computer systems may include the utilization and maximum capacity of the backend computer systems, such as the processing power and network bandwidth of the backend computing systems.”, where the backend computing systems are described in Col. 11, lines 39-45 as “The physical hosts 542 may include any computer system or virtual computer system described above. A virtualization layer 544 operated by the computing resources service provider 502 enables the physical hosts 542 to be used to provide computational resources upon which one or more load balancers 520 may operate.”.). number of active virtual machines in a physical machine (Col. 10, lines 24-35; “The load balancer data may include various metrics data such as the amount of traffic, number of load balancers, the utilization of one or more load balancers, the number of requests processed by the load balancers, virtual machine memory, virtual disk memory, virtual CPU state, virtual CPU memory, virtual GPU memory, operating system cache, operating system page files, virtual machine ephemeral storage, virtual interfaces, virtual devices, number of processors, number of network interfaces, storage space, number of distributed systems, or any other information corresponding to the operation of the load balancers and the computer systems assigned to the load balancers.”. A person of ordinary skill in the art before the effective filing date of the claimed invention would have recognized that the load balancer metrics disclosed in Young correspond to operational status information of VMs assigned to the particular machine hosting the load balancer, and that determining system resource availability for load distribution purposes would have required maintaining metrics indicative of the number of active virtual machines in the infrastructure, thereby making it obvious to derive or maintain such a count based on those VM-level operational metrics to enable more accurate resource allocation and load balancing decisions in varying system conditions.) Ferguson teaches: a queue (Col. 8, lines 12-15; “The BEP scheduling algorithm simply dispatches a transaction from the class C.sub.i that has at least one queued transaction, and which has higher priority than all other classes with queued transactions”). A person of ordinary skill in the art before the effective filing date of the claimed invention would have recognized that the load balancing behavior of Young, which distributes requests based on optimizing resource utilization and avoiding overload of computer systems, involves processing varying utilization conditions of the load balancing environment. As a result, the neural network-based implementation of the load balancing model must learn a distribution over occupancy states of the load balancing engine, applied on the physical machines associated with the load balancer of Young, yielding the predictable benefit of enabling the load balancing model to more accurately route requests under dynamic system load conditions by capturing variations in system utilization over time. Claim 17 recites similar limitations as those of claim 7. Claim 17 is rejected for similar reasons as those of claim 7. Claims 8-9 and 18-19 are rejected under 35 U.S.C. 103 as being unpatentable over Colmenares in view of Liu, further in view of Liu2, further in view of Tateno, further in view of Young, further in view of Ferguson, further in view of Lu et al. (CN 110310344 A), hereafter Lu. Regarding claim 8, Colmenares in view of Liu, further in view of Liu2, further in view of Tateno, further in view of Young, further in view of Ferguson teach the method of claim 5. Liu teaches: wherein an input to the load balancing generator comprises a tensor including a noise vector, (Paragraph 56; “In at least one embodiment, a set of structural feature vectors are obtained 402... these structural and appearance feature vectors are provided 406 as input to a slot attention transformer that is able to generate a set of transformed feature vectors from these input feature vectors these transformed feature vectors are provided 408 as input to a generator, such as a generative adversarial network (GAN)”, and Paragraph 365; “In at least one embodiment, a scheduler may be used to track resource requirements of applications or containers, current usage or planned usage of these resources, and resource availability”, where the transformed feature vectors corresponds to a tensor, the transformed structural vector corresponds to a noise vector, the resource requirements of applications or containers, current usage or planned usage of these resources, and resource availability is considered as number of active virtual machines, and the generator corresponds to the load balancing generator); Liu2 teaches: the number of requests in a load balancing data structure (Paragraph 53; “when the number of the second tasks is greater than one, a coordinated resource allocation sub-method is executed to obtain a resource allocation scheme of the server. In addition, when there is only one second task, mutual interactions between different tasks will not be considered during the resource allocation", explicitly disclosing an evaluation of the number of tasks currently pending in the data structure); input to the physical machine vector (Paragraph 76; “obtaining an interference model between operators in each of the second data processing models. Referring to FIG. 6B, the interference model is constructed by quantifying shared resource requirements, performance testing, and building an analytical model... In the building of the analytical model, linear regression models and neural network models can be employed to model the interference status between the operators, obtaining the interference model between the operators in each of the second data processing models.”, which explicitly teaches construction of an interference model that may be implemented using neural network “models”, explicitly disclosed in plural form, which is used to learn patterns from performance testing and resource sharing data to predict resource behavior. Further, “neural network models can be employed to model the interference status between the operators... based on performance testing and shared resource requirements”, which teaches a neural network based model that predicts system performance using physical resource inputs and makes decisions accordingly, corresponding to a physical machine generator that makes estimates of physical system behavior based on resource conditions.). A person of ordinary skill in the art before the effective filing date of the claimed invention would have recognized that the neural network model of Liu2 receives input data representative of system and resource conditions, corresponding to operational characteristics of underlying resources and corresponds to inputs to inputs to a physical machine generator used to estimate system behavior, yielding the predictable benefit of enabling the model to generate more accurate estimations of system resource behavior. The use of multi-dimensional feature inputs, such as operational parameters and resource metrics within the field of neural networks would have been understood by a person of ordinary skill in the art before the effective filing date of the claimed invention to be capable of being represented as a vector. Young teaches: and the number of active virtual machines (Col. 10, lines 24-35; “The load balancer data may include various metrics data such as the amount of traffic, number of load balancers, the utilization of one or more load balancers, the number of requests processed by the load balancers, virtual machine memory, virtual disk memory, virtual CPU state, virtual CPU memory, virtual GPU memory, operating system cache, operating system page files, virtual machine ephemeral storage, virtual interfaces, virtual devices, number of processors, number of network interfaces, storage space, number of distributed systems, or any other information corresponding to the operation of the load balancers and the computer systems assigned to the load balancers.”. A person of ordinary skill in the art before the effective filing date of the claimed invention would have recognized that the load balancer metrics disclosed in Young correspond to operational status information of VMs assigned to the particular machine hosting the load balancer, and that determining system resource availability for load distribution purposes would have required maintaining metrics indicative of the number of active virtual machines in the infrastructure, thereby making it obvious to derive or maintain such a count based on those VM-level operational metrics to enable more accurate resource allocation and load balancing decisions in varying system conditions). Ferguson teaches: a queue (Col. 8, lines 12-15; “The BEP scheduling algorithm simply dispatches a transaction from the class C.sub.i that has at least one queued transaction, and which has higher priority than all other classes with queued transactions”). Colmenares in view of Liu, further in view of Liu2, further in view of Tateno, further in view of Young, further in view of Ferguson does not teach a one hot encoding. However, Lu teaches: A one hot encoding (Paragraph 68; “being inputted in the generator after the first noise vector Z and only hot vector C are spliced Decoder so that the generator export dummy copy collection”, where vector C corresponds to the one-hot encoding vector.). Colmenares, Liu, Liu2, Tateno, Young, Ferguson, and Lu are considered to be analogous to the claimed invention because they are in the same field of load balancing using models. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Colmenares in view of Liu, further in view of Liu2, further in view of Tateno, further in view of Young, further in view of Ferguson to incorporate the teachings of Lu and combine a one-hot encoding format with the input fields of machine learning models. This combination results in the model able to receive its input fields, namely the number of requests in the load balancing and physical machine queue, the number of active virtual machines (on a physical machine), and the size of the physical machine as a one-hot encoding vector. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of classifying unclassified input data that traditional GAN models are not able to, and hence generate more controlled and accurate outputs of estimated response time (Lu: Specifically, traditional conditional GAN needs data with conditional tag (classification) to train the network, but the disclosed method automatically cluster the data with no label using Noise Characteristic jump and the amplification of one-hot encoding offset, and generate data with conditional tags [0073]; The original generative adversarial Nets are unable to control the content of generated sample when sampling, generated samples are completely are randomly generated. [0003]). Claim 18 recites similar limitations as those of claim 8. Claim 18 is rejected for similar reasons as those of claim 8. Regarding claim 9, Colmenares in view of Liu, further in view of Liu2, further in view of Tateno, further in view of Young, further in view of Ferguson, further in view of Lu teach the method of claim 8. Liu2 teaches: the load balancing model and the physical machine model (Paragraph 76; “obtaining an interference model between operators in each of the second data processing models. Referring to FIG. 6B, the interference model is constructed by quantifying shared resource requirements, performance testing, and building an analytical model... In the building of the analytical model, linear regression models and neural network models can be employed to model the interference status between the operators, obtaining the interference model between the operators in each of the second data processing models.”, which explicitly teaches construction of an interference model that may be implemented using neural network “models”, explicitly disclosed in plural form, which is used to learn patterns from performance testing and resource sharing data to predict resource behavior for both load balancing and for physical machine resource allocation.). a physical machine model (Paragraph 76; “neural network models can be employed to model the interference status between the operators... based on performance testing and shared resource requirements”, which teaches a neural network based model that predicts system performance using physical resource inputs and makes decisions accordingly.). Young teaches: a load balancing model (Col. 4, lines 11-20; “The requests 118 may be received by the load balancer 120 or by one or more other systems of the computing resource service provider 104, such as a request listener not illustrated in FIG. 1 for simplicity, and directed to the load balancer 120. The load balancer 120 may be a computer system or virtual computer system configured to distribute the request 118 to one or more computer systems, supported by physical hosts 142, in order to optimize resource utilization and/or avoid overloading a particular computer system.”.). Lu teaches: a discriminator configured to determine whether an input to the discriminator is real or fake (Paragraph 73; “arbiter D is to differentiate true sample set Xreal With dummy copy collection Xfake, to obtain the differentiated result. Wherein, arbiter can exactly determine an image input one as true sample set or dummy copy collection”, where D corresponds to the discriminator.). It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have the discriminator configured to determine whether an input to the discriminator is real or fake be applied to both of the neural network models. A person of ordinary skill in the art before the effective filing date of the claimed invention would have recognized that applying such a discriminator is a common application to evaluate the validity of generated data against real data samples, corresponding to the implementation of the known GAN-based architecture and yielding the predictable result of improving accuracy of generated estimates by enabling the models to distinguish between real system states and artificially created system states. Claim 19 recites similar limitations as those of claim 9. Claim 19 is rejected for similar reasons as those of claim 9. Claims 10 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Colmenares in view of Liu, further in view of Liu2, further in view of Tateno, further in view of Young, further in view of Ferguson, further in view of Lu, further in view of Redford et al. (US 20230222336 A1), hereafter Redford. Regarding claim 10, Colmenares in view of Liu, further in view of Liu2, further in view of Tateno, further in view of Young, further in view of Ferguson, further in view of Lu teach the method of claim 9. Liu2 teaches: a physical machine model (Paragraph 76; “neural network models can be employed to model the interference status between the operators... based on performance testing and shared resource requirements”, which teaches a neural network based model that predicts system performance using physical resource inputs and makes decisions accordingly.). Tateno teaches: training using ground truth data (Paragraph 169; “Using “real (true)” or “fake (false)” as training data”, explicitly discloses the option of using true data as training data, corresponding to ground truth data.). Young teaches: a load balancing model (Col. 4, lines 11-20; “The requests 118 may be received by the load balancer 120 or by one or more other systems of the computing resource service provider 104, such as a request listener not illustrated in FIG. 1 for simplicity, and directed to the load balancer 120. The load balancer 120 may be a computer system or virtual computer system configured to distribute the request 118 to one or more computer systems, supported by physical hosts 142, in order to optimize resource utilization and/or avoid overloading a particular computer system.”.). Colmenares in view of Liu, further in view of Liu2, further in view of Tateno, further in view of Young, further in view of Ferguson, further in view of Lu does not teach ground truth data that is discretized and binned. However, Redford teaches: ground truth data that is discretized and binned (Paragraph 365; “This functional form can be approximated by discretising each confounder. In this representation, categorical confounders (such as vehicle type) are mapped to bins. Continuous confounders (such as distance from detector) are sliced into ranges and each range mapped to a bin.”, and Paragraph 368; “Training a model requires ground truth and stack predictions (actual perception outputs), collected as described in Section 3.1.1. The mean and covariance of the normal distribution are fitted (e.g. using a maximum a posteriori method to incorporate a prior) to the observations in that bin.”, where the continuous cofounders discretized and binned corresponds to ground truth data). Colmenares, Liu, Liu2, Tateno, Young, Ferguson, Lu, and Redford are considered to be analogous to the claimed invention because they are in the same field of model evaluation. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Colmenares in view of Liu, further in view of Liu2, further in view of Tateno, further in view of Young, further in view of Ferguson, further in view of Lu to incorporate the teachings of Redford and have ground truth data that is discretized and binned. This combination results in continuous input fields being classified and categorized and enables the input to be represented as one-hot encoding vectors. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of making the model interpretable and be analyzed and tuned to optimize the accuracy of the GAN model, further improving the response time prediction (Redford: The advantages of the PCM approach are that it accounts for global heteroskedasticity, gives a unified framework to capture confounders of different types, and it utilizes simple probability distributions. In addition, the model is interpretable: the distribution in a bin can be examined, the training data can be directly inspected and there are no hidden transforms. Moreover, the parameters can be fitted analytically, meaning uncertainty from lack of convergence in optimisation routines can be avoided. [0372]). Claim 20 recites similar limitations as those of claim 10. Claim 20 is rejected for similar reasons as those of claim 10. Response to Arguments Applicant's arguments filed 09/26/2025 have been fully considered. Applicant’s arguments are summarized below: Amendments to the specification overcome the objections to the specification. Amendments made to claims 7 and 17 as suggested by the Examiner overcome the objections made for minor informalities. Amendments to claims 5 and 15 now particularly point out and distinctly claim the subject matter regarded as the invention and overcome the rejections under 35 U.S.C. 112(b). The prior art of record fails to teach amended limitations of the occupancy status of the physical machines, size of the physical machines, or conditional GANs. Dependent claims are submitted as allowable for at least the above reasons. Examiner’s response: The Examiner agrees that the amendments to the specification remedying the element identifier in [0049] overcome the objections to the specification. Accordingly, the objection to the specification is withdrawn. The Examiner agrees that the amendments to claims 7 and 17 remedy the minor informalities previously objected to. Accordingly, the objections to claims 7 and 17 are withdrawn. The Examiner agrees that the amendments to claims 5 and 15 provides antecedent basis for the limitation in the claim by tying the step to both the first and second estimating engines. Accordingly, the rejections of claims 5 and 15 under 35 U.S.C. 112(b) are withdrawn. The Examiner agrees that the prior art of record fails to teach amended limitations of the occupancy status of the physical machines, size of the physical machines, or conditional GANs. Accordingly, the previous rejections of claims 1 and 11 under 35 U.S.C. 103 are withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of Colmenares, Liu, Liu2, Tateno, and Young, under 35 U.S.C. 103. Independent claims 1 and 11 remain rejected for the reasons stated above. Therefore, contrary to Applicant's arguments, because the dependent claims depend from an unpatentable claim and does not add limitations that overcome the rejection, it likewise remains rejected. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Basu et al. (US 20230066080 A1) discusses determining, by a memory system, a plurality of resource parameters associated with operation of the memory system and determining respective time intervals associated with usage patterns corresponding to the memory system, the respective time intervals being associated with one or more sets of the plurality of resource parameters. The method further includes determining, using the plurality of resource parameters, one or more weights for hidden layers of a neural network for the respective time intervals associated with the usage patterns and allocating computing resources within the memory system for use in execution of workloads based on the determined one or more weights for hidden layers of the neural network. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to KENNETH P TRAN whose telephone number is (571)272-6926. The examiner can normally be reached M-TH 4:30 a.m. - 12:30 p.m. PT, F 4:30 a.m. - 8:30 a.m. PT, or at Kenneth.Tran@uspto.gov. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, April Blair can be reached at (571) 270-1014. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /KENNETH P TRAN/Examiner, Art Unit 2196 /APRIL Y BLAIR/Supervisory Patent Examiner, Art Unit 2196
Read full office action

Prosecution Timeline

Sep 19, 2022
Application Filed
Jun 30, 2025
Non-Final Rejection mailed — §103, §112
Sep 26, 2025
Response Filed
Jul 16, 2026
Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743309
LCS RESOURCE DEVICE PRESENTATION SYSTEM
4y 2m to grant Granted Sep 22, 2026
Patent 12602250
LCS RESOURCE DEVICE UTILIZATION SYSTEM
3y 9m to grant Granted Apr 14, 2026
Study what changed to get past this examiner. Based on 2 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
31%
Grant Probability
98%
With Interview (+66.7%)
3y 8m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 13 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month