Prosecution Insights
Last updated: October 02, 2026
Application No. 18/171,689

SYSTEM AND METHOD FOR API RESOURCE PREDICTION

Final Rejection §103
Filed
Feb 21, 2023
Priority
Feb 22, 2022 — IN 202221009466
Examiner
MILLS, FRANK D
Art Unit
2194
Tech Center
2100 — Computer Architecture & Software
Assignee
Jio Platforms Limited
OA Round
2 (Final)
70%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
92%
With Interview

Examiner Intelligence

Grants 70% — above average
70%
Career Allowance Rate
424 granted / 610 resolved
+14.5% vs TC avg
Strong +23% interview lift
Without
With
+22.7%
Interview Lift
resolved cases with interview
Typical timeline
3y 4m
Avg Prosecution
23 currently pending
Career history
631
Total Applications
across all art units

Statute-Specific Performance

§101
16.5%
-23.5% vs TC avg
§103
52.4%
+12.4% vs TC avg
§102
12.0%
-28.0% vs TC avg
§112
12.8%
-27.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 610 resolved cases

Office Action

§103
DETAILED ACTION This action is in response to the reply received 04/02/2026. After consideration of applicant's amendments and/or remarks: Examiner withdraws rejections under 35 USC § 112. Claims 1-18 rejected under 35 USC § 103. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 3-4, 9-10, 12, and 17-18 are rejected under 35 U.S.C. 103 as being unpatentable over Adibowo, U.S. PG-Publication No. 2022/0215008 A1 (hereinafter ADIBOWO), in view of Lin et al., U.S. Patent No. 8,626,791 B1 (hereinafter LIN), further in view of Gao et al., U.S. PG-Publciation No. 2023/0035451 A1 (hereinafter GAO). Claim 1 ADIBOWO discloses a system … comprising: a processor operatively coupled with a memory, wherein said memory stores instructions which when executed by the processor causes the processor to. ¶ 0009: The system includes one or more processors and a coupled computer-readable storage medium storing instructions that cause the processors to perform the disclosed operations. ADIBOWO discloses receive one or more requests from one or more computing device. ¶ 0020: An API server receives “a prediction request from a client system,” selects a model server, calls the model server to execute inference using an ML model loaded in memory, receives the inference result, and returns the result to the client. ¶ 0031: Customer system 210 transmits inference requests to server system 220 and receives corresponding responses. ADIBOWO discloses wherein one or more users operate the one or more computing devices. ¶ 0021: Architecture 100 includes client devices 102 and server system 104; respective “users 110 interact with the client devices 102” to obtain inference using one or more ML models. ¶ 0022: Client devices include desktop computer, tablets, cellular telephones, smartphones, and other data processing devices. ADIBOWO discloses wherein the received one or more requests are based on a training of one or more model via a machine learning (ML) engine operatively coupled to the processor. ¶ 0028: The server hosts multiple disparate ML models “generated and/or trained with different data or trained for different purposes or clients.” The models are loaded into memory upon receiving corresponding prediction requests. ¶ 0039: Each inference request includes input data and identifies the trained ML model to be used. A customer may “train one or more ML models for the request based on historical data.” ADIBOWO discloses unload a plurality of least recently used models from the one or more trained models based on the received one or more requests. ¶ 0034: Model servers load ML models required by received requests and apply an LRU replacement algorithm that replaces an in-use model that “has not been used for the longest period of time relative to other in-use ML models.” ¶ 0047: When memory is full, the model server “can unload one or more in-use ML models that are not recently used based on the LRU replacement algorithm.” ADIBOWO discloses utilize a caching mechanism to optimize a memory space associated with the memory. ¶ 0034: Custom node selection maximized locality and minimizes the overhead required to load or unload ML models from memory. In-use models remain in memory and are managed using LRU replacement. ¶ 0035: The combination of LRU replacement and server selection “optimizes the server system” when providing ML inference services. ¶ 0047: A cache miss causes the requested model to be retrieved and loaded into memory; a full memory causes LRU unloading; and a cache hit permits immediate inference. ADIBOWO does not expressly disclose determine a most recently used model from the one or more trained models; utilize a caching mechanism to optimize a memory space associated with the memory by selectively loading the most recently used trained model; predict, via the most recently used trained model from the one or more trained models. LIN discloses determine a most recently used model from the one or more trained models. Col. 2, ll. 8-18: Determining the set of model identifiers may include evaluating model size against a threshold and “identifying a most recently used predictive model.” LIN discloses utilize a caching mechanism to optimize a memory space associated with the memory by selectively loading the most recently used trained model. Col. 12, ll. 51-65: A trained scheduling model predicts which trained predictive models are likely to receive requests, and the scheduler loads the identified models from secondary memory into the primary memory model cache. Col. 13, ll.1-16: When the cache lacks sufficient resources, less frequently used models are removed. A model that recently received a request remains cached regardless of other factors. Col. 13, ll. 17-25: The scheduler balances model size against available primary memory, stores as many predictive models in the cache as possible,” and compares model size with likelihood of access to decide whether a model should be cached. Col. 14, ll. 3-18: Selected models, are stored in a RAM or DRAM cache; a requested model absent from the cache is selectively loaded from secondary memory. LIN discloses predict, via the most recently used trained model from the one or more trained models. Col. 11, ll. 30-39: A predictive request is forwarded to a trained predictive model, and that model produces predictive output. Col. 14, ll.19-30. When a predictive request is received, the cache determines whether a model capable of satisfying the request is present. The model is accessed from primary memory if present or loaded from secondary memory if absent. It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to modify the request responsive LRU model serving system of ADIBOWO to incorporate the MRU identification, predictive model preloading, and memory aware model retention techniques taught by LIN. One of ordinary skill in the art would be motivated to integrate LIN’s MRU identification, predictive preloading, and memory-aware retention techniques into ADIBOWO, with a reasonable expectation of success, in order to reduce prediction latency, efficiency utilize memory resources, and improve computer-resource utilization, as taught by LIN at col. 2 ll. 30-34, col. 12 ll.51-65, and col. 14, ll.19-30. ADIBOWO -LIN does not expressly disclose a system for resource prediction; predict, via the most recently used trained model from the one or more trained models, resource data based on the optimized memory space; and enable the resource prediction based on the predicted resource data. GAO discloses a system for resource prediction. ¶ 0016: The disclosed system predicts resource usage of a deep-learning model based on model information, the operating environment, static resource usage, and runtime strategy. GAO discloses predict, via the most recently used trained model from the one or more trained models, resource data based on the optimized memory space. ¶ 0029: Resource prediction result 190 may identify computation power consumption, main memory consumption, GPU memory consumption, I/O resource consumption, execution time, power consumption, and budget. ¶ 0036: The prediction sues the model’s computation graph, resource prediction models, and exaction information. The execution information includes computing device specifications, numbers of computing devices, and execution strategy. ¶ 0052: Runtime memory management and optimization affect estimated memory consumption. Memory-allocation strategy, hardware specifications, the number of computing devices, and execution strategy influence the memory consumption estimate. ¶ 0053: The prediction unit determines the runtime memory allocation and execution strategies and adjusts static resource usage according to those strategies to predict runtime resource usage. ¶ 0054: Potential resource consumption during runtime “may be predicted using a trained machine learning model.” ¶ 0055: The ML estimate model is trained using features concerning resource consumption, the computation graph, and the execution environment. The trained model determines potential resource consumption, including predicted memory usage. GAO discloses enable the resource prediction based on the predicted resource data. ¶ 0017: The prediction result identifies model bottlenecks, supports model parameter adjustment, tailors the AutoML search space, and facilitates optimization of job execution strategy and resource utilization. ¶ 0056: The resource-prediction output can improve model performance, tailor the model parameter search space, and optimize AI-platform job scheduling. It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to modify the MRU managed predictive model cache of ADIBOWO-LIN to incorporate the trained resource usage prediction functionality taught by GAO. One of ordinary skill in the art would be motivated to integrate GAO’s trained resource usage prediction functionality into ADIBOWO-LIN, with a reasonable expectation of success, in order to avoid failure caused by insufficient resources, identify performance bottlenecks, optimize model parameters and job scheduling, and enhance resource utilization, as taught by GAO ¶¶ 0014, 0017, 0054, and 0056. Claim 3 ADIBOWO discloses wherein the processor (202) is configured to utilize a least recently used (LRU) technique as the caching mechanism to optimize the memory space. API servers 232 “can each execute a custom node selection to determine which … model server 234 is chosen to perform inference service in response to an inference request.” The model servers 234 “execute a least-recently-used (LRU) replacement algorithm to manage replacement of ML models,” wherein “the LRU replacement algorithm is executed to replace an in-use ML model that has not been used for the longest period of time relative to other in-use ML models.” Adibowo, ¶ 34. The LRU replacement algorithm “optimizes the server system 230 in providing ML inference services using heterogeneous ML models to heterogenous clients.” Id. at ¶ 35 Claim 4 ADIBOWO discloses wherein the memory (204) is a random access memory (RAM) for storing the one or more trained models. Instructions and data are “received from … a random access memory.” Adibowo, ¶ 64; See Also ¶ 25 (“main memory includes random access memory”). Claim 9 ADIBOWO discloses wherein the processor (202) is configured with a conditional lock functionality to process, in a successive order, at least a model from the one or more trained models based on the received one or more requests. The method “provides for selection of the same model node for inference that calls for a particular ML model in combination with LRU caching within each node,” wherein the routing “is achieved by hashing a combination of the node’s identity … with the ML model’s identity and then ordering the hashed combination” (ordering the hashed node/model combination → process in a successive order).” The “ordering is fixed as long as the pool of nodes is made constant,” and “this fixed set of ordering makes effective use of the LRU cache within a node.” In embodiments, “the selection can be based on hashing a combination of model server identifier and the ML model identifier and ordering the model servers based on respective hash values,” such that “the inference request for a particular ML model would more frequently be sent to the same model server 234.” Adibowo, ¶¶ 48-49. Claims 10, 12, and 17 Claims 10, 12, and 17 are rejected utilizing the aforementioned rationale for Claims 1, 3, and 9; the claims are directed to a method performed by the system. Claim 18 Claim 18 is rejected utilizing the aforementioned rationale for Claim 1; the claim is directed to a “user equipment” comprising the same elements of the system. Claims 2 and 11 are rejected under 35 U.S.C. 103 as being unpatentable over ADIBOWO, in view of LIN, further in view of GAO, further in view of Braz et al., U.S. PG-Publication No. 2020/0175387 A1 (hereinafter BRAZ). Claim 2 BRAZ discloses wherein the processor (202) is configured to utilize a versioning logic mechanism to determine the at least recently used trained model from the one or more trained models and process the received one or more requests. Braz discloses a method “for managing and deploying AI models, including storing at least one artificial intelligence (AI) model in a model store memory in a plurality of different versions, each different version having a different level of fidelity; receiving a prediction request to process the AI model; determining … which version of the AI model to use for processing the received prediction request … using the determined version of the AI model; and responding to the received prediction request with a result of the processing … using the determined AI model version.” Braz, ¶ 5. Further, this versioning method uses “eviction/loading policies for which AI model to evict that is currently residing in the memory used for storing models available for immediate execution, when a determination is made to move another AI model into that memory for execution,” including a “least-recently-used (LRU) policy” or “a policy that considers a potential gain in confidence level or fidelity level between different versions of AI models.” Id. at ¶ 39. It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to modify the multi-model inference services of ADIBOWO-LIN-GAO to incorporate multi-model versioning services as taught by BRAZ. One of ordinary skill in the art would be motivated to integrate multi-model versioning services into ADIBOWO-LIN-GAO, with a reasonable expectation of success, in order to reduce latency in responding to inference requests, by allowing “for serving a large number of models with low delay at a temporary cost of performance, as well as a mechanism that permits accuracy to improve with subsequent requests from [a] client.” See Braz, ¶¶ 24-25. Claim 11 Claim 11 is rejected utilizing the aforementioned rationale for Claim 2; the claim is directed to a method performed by the system. Claims 5-8 and 13-16 are rejected under 35 U.S.C. 103 as being unpatentable over ADIBOWO, in view of LIN, further in view of GAO, in view of Brand et al., U.S. PG-Publication No. 2016/0071027 A1 (hereinafter BRAND). Claim 5 BRAND discloses wherein the processor (202) comprises a base manager to store one or more trained models, and enable one or more parallel processes to utilize the one or more trained models. Brand discloses “methods … used to perform real-time analysis and modeling of large events streams of events.” The event stream is partitioned “among multiple local modelers that perform the same set of operations,” so that “processing of the event stream can be performed in parallel, increasing throughput.” Brand, ¶¶ 35-36. The method “initializes each central modeler, e.g., registers each local modeler in communication with the central modeler, obtains machine learning models identified in the configuration file, and so on,” wherein “a central modeler can provide the machine learning model to local modelers in communication with the central modeler” (central modeler → base manager). The method provides “a machine learning model to local modelers by executing a central modeler function to provide the machine learning model, and executing a local modeler function to receive and store the machine learning model” (i.e., enable one or more parallel processes to utilize the trained models). Id. at ¶ 94; See Also ¶ 92 (“multiple local modelers can execute in parallel to increase throughput”), ¶ 99 (“each local modeler performing the same operations in parallel on received events”). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to modify multi-model inference services of ADIBOWO-LIN-GAO to incorporate parallel processing using trained models as taught by BRAND. One of ordinary skill in the art would be motivated to integrate parallel processing using trained models into ADIBOWO-LIN-GAO, with a reasonable expectation of success, in order to improve performance by using “multiple local modelers that perform the same set of operations,” enabling “processing of the event stream … in parallel, increasing throughput.” See Brand, ¶ 36; See Also ¶ 91 (“multiple local modelers can execute in parallel to increase throughput”). Claim 6 BRAND discloses wherein the processor (202) is configured with a common loading functionality for loading the one or more trained models. Figure 6 illustrates system 600 “for processing an event stream by an example routing strategy using context data,” wherein system 600 “includes multiple local modelers 610a-n- of a stream processing system in communication with a central modeler 620.” System 600 further comprises a “routing node 604” that “receives an event stream and routes each event in the event stream according to the routing strategy, e.g., to a particular local modeler that stores context data related to the processing of the event” (routing strategy using context data → common loading functionality). Brand, ¶ 111. System 600 “can partition context data so that particular context data related to a particular event is likely to be located on a same local modeler as other context data related to the particular event, e.g., context data needed for processing the event.” Id. at ¶ 113. It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to modify multi-model inference services of ADIBOWO-LIN-GAO to incorporate parallel processing using trained models as taught by BRAND. One of ordinary skill in the art would be motivated to integrate parallel processing using trained models into ADIBOWO-LIN-GAO, with a reasonable expectation of success, in order to improve performance by using “multiple local modelers that perform the same set of operations,” enabling “processing of the event stream … in parallel, increasing throughput.” See Brand, ¶ 36; See Also ¶ 91 (“multiple local modelers can execute in parallel to increase throughput”). Claim 7 BRAND discloses wherein the base manager is configured to use a race condition avoidance solution to prevent a read or write operation of the one or more parallel processes in a directory. Brand discloses that since “the context data is maintained in operational memory, the system can quickly obtain the requested context data, avoid data locking issues and race conditions, and avoid having to call and obtain context data from an outside database.” The system performs an operation using the context data.” Brand, ¶ 129; See Also ¶ 23 (system partitions context data into local memories of local modelers of the stream processing system to “reduce latency due to data locking issues, and race conditions”). Claim 8 BRAND discloses wherein the common loading functionality is configured with a multi-processing lock functionality to prevent the one or more parallel processes from utilizing the one or more trained models simultaneously. Brand discloses that since “the context data is maintained in operational memory, the system can quickly obtain the requested context data, avoid data locking issues and race conditions, and avoid having to call and obtain context data from an outside database.” The system performs an operation using the context data.” Brand, ¶ 129; See Also ¶ 23 (system partitions context data into local memories of local modelers of the stream processing system to “reduce latency due to data locking issues, and race conditions”). Claims 13-16 Claims 13-16 are rejected utilizing the aforementioned rationale for Claims 5-8; the claims are directed to a method performed by the system. Response to Arguments Applicant’s arguments with respect to claim(s) 1 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. See Lin et al., U.S. Patent No. 8,626,791 B1; Gao et al., U.S. PG-Publication No. 2023/0035451 A1. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to FRANK D MILLS whose telephone number is (571)270-3194. The examiner can normally be reached M-F 9-5:30 CT. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, KEVIN YOUNG can be reached at (571)270-3180. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /FRANK D MILLS/Primary Examiner, Art Unit 2194 September 4, 2026
Read full office action

Prosecution Timeline

Feb 21, 2023
Application Filed
Jan 02, 2026
Non-Final Rejection mailed — §103
Apr 02, 2026
Response Filed
Sep 09, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743305
METHOD AND APPARATUS FOR COLLABORATIVE TASK PLANNING FOR ARTIFICIAL INTELLIGENCE AGENTS
3y 4m to grant Granted Sep 22, 2026
Patent 12737205
COMPUTER DEVICE INCLUDING PROCESS ISOLATED CONTAINERS WITH ASSIGNED VIRTUAL FUNCTIONS
4y 8m to grant Granted Sep 15, 2026
Patent 12710985
COUPLED COMPUTE AND STORAGE RESOURCE AUTOSCALING
3y 10m to grant Granted Aug 18, 2026
Patent 12705099
PROVIDING AI-GENERATED CONTENT
3y 6m to grant Granted Aug 11, 2026
Patent 12699606
ELECTRONIC DEVICE, CONTROL METHOD, AND STORAGE MEDIUM
3y 7m to grant Granted Aug 04, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
70%
Grant Probability
92%
With Interview (+22.7%)
3y 4m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 610 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month