Prosecution Insights
Last updated: August 17, 2026
Application No. 18/518,925

OPTIMIZING PLACEMENT OF FINE-TUNED MACHINE LEARNING MODELS AT HOST SYSTEMS

Non-Final OA §102§103§112
Filed
Nov 24, 2023
Examiner
CHUANG, SU-TING
Art Unit
Tech Center
Assignee
Amazon Technologies Inc.
OA Round
1 (Non-Final)
51%
Grant Probability
Moderate
1-2
OA Rounds
1y 9m
Est. Remaining
90%
With Interview

Examiner Intelligence

Grants 51% of resolved cases
51%
Career Allowance Rate
55 granted / 108 resolved
-9.1% vs TC avg
Strong +40% interview lift
Without
With
+39.5%
Interview Lift
resolved cases with interview
Typical timeline
4y 6m
Avg Prosecution
21 currently pending
Career history
135
Total Applications
across all art units

Statute-Specific Performance

§101
26.3%
-13.7% vs TC avg
§103
47.6%
+7.6% vs TC avg
§102
11.4%
-28.6% vs TC avg
§112
12.3%
-27.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 108 resolved cases

Office Action

§102 §103 §112
DETAILED ACTION Claims 1-20 are pending and have been examined. -- Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statements (IDS) submitted on 01/28/2025, 02/25/2025, 04/09/2025 and 02/24/2026 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statements are being considered by the examiner. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 1-4, 10, 13 and 19-20 are rejected under 35 U.S.C. 112(b) or pre-AIA 35 U.S.C. 112, second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor, or for pre-AIA the applicant regards as the invention. Claim 1 recites the limitation “… including the machine learning model… determine that the at least one machine learning model….” There is insufficient antecedent basis for the limitation “the machine learning model” in the claim, and inconsistent for the limitation “the at least one machine learning model.” For examination purposes examiner has interpreted “the machine learning model” to be “a machine learning model,” and “the at least one machine learning model” to be “the machine learning model.” Claim 1 recites the limitation “for at least the one of the plurality of different machine learning models.” There is insufficient antecedent basis for the limitation “the one of the plurality of different machine learning models” in the claim. For examination purposes examiner has interpreted “the one of the plurality of different machine learning models” to be “one of the plurality of different machine learning models.” Claim 1 recites the limitation “on the host system.” There is insufficient antecedent basis for the limitation “the host system” in the claim. For examination purposes examiner has interpreted “the host system” to be “a host system.” Claims 3, 10 and 19 recite the limitation “the respective delta model is loaded… for the respective delta model….” There is insufficient antecedent basis for the limitation “the respective delta model” in the claim. For examination purposes examiner has interpreted “the respective delta model is loaded… for the respective delta model…” to be “the respective delta models are loaded… for the respective delta models....” Claims 13 and 20 recites the limitation “the plurality of delta models.” There is insufficient antecedent basis for the limitation “the plurality of delta models” in the claim. For examination purposes examiner has interpreted “the plurality of delta models” to be “a plurality of delta models.” Claims 2 and 4 are also rejected due to their dependency on a rejected claim. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention. Claims 5, 10, 12, 14, 19 rejected under 35 U.S.C. 102 (a)(1) as being anticipated by Torres (CA3129720C) In regard to claims 5 and 14, Torres teaches: A method, comprising: receiving, at a machine learning service, a request to place a machine learning model on a host system of the machine learning service, wherein the machine learning model is a base model for a fine-tuned machine learning model; (Torres, [14] "These elements may work together to automatically route prediction requests, [receiving a request] configure ML systems, serve NLP models, and/or generate NLP predictions."; [20] "As shown in FIG. 2, the ML system 130 may serve a variety of NLP applications by integrating one or more light-weight model artifacts 242 into a base model 220 to dynamically adapt the base model 220 [the machine learning model (the base model)] to a particular NLP task."; [21] "The Adapter Service architecture includes a model deployment service 240 and other components that dynamically augment a base model 220 [the machine learning model (the base model)] with one or more model artifacts 242. By integrating the one or more model artifacts 242 into a base model 220, the model deployment service 240 [at a machine learning service] generates a plurality of models for performing one or more NLP tasks.") identifying, by the machine learning service, one or more different machine learning models that are respective delta models with respect to the base model, wherein respective combinations of the delta models with the base model produce respective versions of the fine-tuned machine learning model; and (Torres, [20] "As shown in FIG. 2, the ML system 130 may serve a variety of NLP applications by integrating one or more light-weight model artifacts 242 into a base model 220 [respective delta models (artifacts) with respect to the base model, respective combinations of the delta models with the base model] to dynamically adapt the base model 220 to a particular NLP task."; [21] "The Adapter Service architecture may be a machine learning system architecture that facilitates generating a plurality of specific models [one or more different machine learning models] from a single base model instance. The Adapter Service architecture includes a model deployment service 240 and other components that dynamically augment a base model 220 with one or more model artifacts 242. [respective delta models (artifacts) with respect to the base model, respective combinations of the delta models with the base model] By integrating the one or more model artifacts 242 into a base model 220, the model deployment service 240 generates a plurality of models [respective versions of the fine-tuned machine learning model] for performing one or more NLP tasks."; [23] "A state of the art transfer learning NLP model can be used as a base model 220 that may be augmented with one or more lightweight adapter PNG media_image1.png 692 634 media_image1.png Greyscale PNG media_image2.png 888 634 media_image2.png Greyscale layers 226a-n. [respective delta models]") placing, by the machine learning service, both the base model and the respective delta models on the host system, (Torres, [56] "At 604, the model deployment service loads the model artifact into its proper location within the base model [placing both the base model and the respective delta models] architecture to generate a fine-tuned model."; [63] "After inserting the adapter layers and/or output layer included in the new model artifact within the original model artifact, the new model artifact is loaded into memory at 604 [placing both the base model and the respective delta models on the system] to generate the fine-tuned model.") wherein the host system generates respective inferences for requests that invoke one of the respective versions of the fine-tuned machine learning model. (Torres, [21] "By integrating the one or more model artifacts 242 into a base model 220, the model deployment service 240 generates a plurality of models for performing one or more NLP tasks. Using the Adapter Service architecture, a single deployment of the ML system 130 may dynamically serve a variety of NLP models generating predictions [generates respective inferences] for a wide variety of NLP tasks."; [57] "input text is processed using a combination of the base model parameters and the cached parameters of the model artifact specified by the fine-tuned model. NLP predictions [respective inferences] generated by the fine-tuned model are then distributed to an NLP application at 608. ") Claim 14 recites substantially the same limitation as claim 5, therefore the rejection applied to claim 5 also apply to claim 14. In addition, Torres teaches: One or more non-transitory, computer-readable storage media, storing program instructions that when executed on or across one or more computing devices cause the one or more computing devices to implement a machine learning service that implements: (Torres, [64] "the computing device 700 may include one or more processors 702, one or more input devices 704, one or more display devices 706, one or more network interfaces 708, and one or more computer-readable mediums 712") In regard to claims 10 and 19, Torres teaches: further comprising: receiving a request to generate an inference using one of the respective versions of the fine-tuned machine learning model at the host system; (Torres, [14] "These elements may work together to automatically route prediction requests, [receiving a request] configure ML systems, serve NLP models, and/or generate NLP predictions. [generate an inference]"; [55] "FIG. 6 illustrates an example of generating NLP predictions at runtime 600. At 602, the ML system receives input data. Input data may include a payload comprising input text, an NLP task identifier, and/or the model artifact location corresponding to the model artifact for the specified NLP task identifier. [using one of the respective versions ] The NLP task identifier may define an NLP task type that specifies one or more target NLP task types to be performed on the input text... NLP predictions [generate an inference] generated by the fine-tuned model are then distributed to an NLP application at 608. ") determining that the respective delta model is loaded into a memory of the host system; (Torres, [56] "At 604, the model deployment service loads the model artifact into its proper location within the base model architecture to generate a fine-tuned model."; [63] "After inserting the adapter layers and/or output layer included in the new model artifact within the original model artifact, the new model artifact is loaded into memory at 604 to generate the fine-tuned model.") generating delta values for the respective delta model for given input for the request; and (Torres, [40] "The training method 500 may be used to update adapter parameter sets [generating delta values] to generate model artifacts [for the respective delta model] optimized for a particular NLP task. In various embodiments, adapter parameters are trained [generating delta values] by inserting the untrained adapter layers and output layer into the base model, freezing the base model weights, and training only the adapter parameters."; [1a] "the generating comprising dynamically exchanging the first model artifact [generating delta values] with a second model artifact comprising one or more adapter layers specific to the second task type;") using the delta values to complete generation of the inference using base values generated by the base model in combination with the delta values. (Torres, [57] "At 606... input text is processed using a combination of the base model parameters and the cached parameters of the model artifact specified by the fine-tuned model. NLP predictions generated by the fine-tuned model are then distributed to an NLP application at 608."; [34] "Integrating the adapter parameters [using the delta values] into the base model 220 transfers knowledge captured in the base model parameters [using base values generated by the base model] to the fine-tuned models. The adapter parameters may then augment the knowledge provided by base parameters to optimize the fine-tuned models for a particular NLP task."; [32] "Adapter parameters included in the one or more adapter layers 226 may introduce, eliminate, and/or re-arrange data points included in the learned manifold by adding additional weights and/or modifying weights provided by the base parameters. By adapting the learned manifold to fit a dataset specific to one or more specific NLP tasks, the one or more adapter layers 226 tune the base model 220 to generate a fine-tuned model that performs one or more particular NLP tasks."; using adapter parameters and base parameters in combination for NLP prediction (inference)) In regard to claim 12, Torres teaches: wherein the one or more different machine learning models are trained using different respective tuning data sets. (Torres, [32] "By adapting the learned manifold to fit a dataset [using different respective tuning data sets] specific to one or more specific NLP tasks, the one or more adapter layers 226 tune the base model 220 to generate a fine-tuned model that performs one or more particular NLP tasks."; [34] "The adapter parameters may then augment the knowledge provided by base parameters to optimize the fine-tuned models for a particular NLP task. For example, the adapter parameters may change the learned manifold of the base model 220 to fit a dataset for a particular NLP task.") Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1-4, 6-9, 11, 13, 15-18 and 20 rejected under 35 U.S.C. 103 as being unpatentable over Torres in view of Clement (US 20220398462 A1) In regard to claim 1, Torres teach: A system, comprising: a plurality of computing devices, respectively comprising at least one processor and a memory, that implement a machine learning service, (Torres, [64] "The computing device 700 may be implemented on any electronic device that runs software applications derived from compiled instructions, including without limitation personal computers, servers, [computing devices] smart phones, media players, electronic tablets, game consoles, email devices, etc. In some implementations, the computing device 700 may include one or more processors 702, one or more input devices 704, one or more display devices 706, one or more network interfaces 708, and one or more computer-readable mediums 712"; [21] "By integrating the one or more model artifacts 242 into a base model 220, the model deployment service 240 [a machine learning service] generates a plurality of models for performing one or more NLP tasks.") wherein the machine learning service is configured to:… wherein the managed network endpoint provides access to a plurality of different machine learning models hosted at one or more of a plurality of computing resources associated with the managed network endpoint, including the machine learning model, (Torres, [16] "NLP application 150 may be a hardware and/or software component accessible to client 160 through network 140 (e.g., an application hosted by an application server [a network endpoint] or other server computer)."; [18] "one or more of client 160, NLP application instance 150, and/or ML system 130 may communicate with one another through network 140. For example, communication between the elements may be facilitated by one or more application programming interfaces (APIs)."; a server, listens for requests, and lets a remote client reach it across a network connection via API; [21] "The Adapter Service architecture may be a machine learning system architecture that facilitates generating a plurality of specific models [a plurality of different machine learning models] from a single base model instance... By integrating the one or more model artifacts 242 into a base model 220, [the machine learning model] the model deployment service 240 generates a plurality of models [a plurality of different machine learning models] for performing one or more NLP tasks."; [26] "The Adapter Service architecture may also more efficiently modify processing, memory, and network resources [a plurality of computing resources] in response to changes in demand for predictions generated by different NLP models.") via requests to invoke specified ones of the plurality of different machine learning models received from one or more clients of the machine learning service; (Torres, [21] "To process a subsequent request for an alternative NLP task, [requests to invoke specified ones of the plurality of different machine learning models] for example, document classification, the ML system 130 may exchange the sentiment analysis adapter layers and output layer for the document classification specific adapter layers and output layers to generate a fine-tuned model for document classification. The ML system 130 may then process the document classification request and output the document classification prediction."; [16] "NLP application 150 may be a hardware and/or software component accessible to client 160 through network 140") for at least the one of the plurality of different machine learning models, the machine learning service is configured to: (Torres, [21] "The Adapter Service architecture may be a machine learning system architecture that facilitates generating a plurality of specific models [the plurality of different machine learning models] from a single base model instance...") determine that the at least one machine learning model is a base model for a fine-tuned machine learning model; (Torres, [20] "As shown in FIG. 2, the ML system 130 may serve a variety of NLP applications by integrating one or more light-weight model artifacts 242 into a base model 220 to dynamically adapt the base model 220 [the machine learning model (the base model)] to a particular NLP task."; [21] "The Adapter Service architecture includes a model deployment service 240 and other components that dynamically augment a base model 220 [the machine learning model (the base model)] with one or more model artifacts 242. By integrating the one or more model artifacts 242 into a base model 220, the model deployment service 240 generates a plurality of models for performing one or more NLP tasks.") identify one or more of the plurality of different machine learning models that are respective delta models with respect to the base model, wherein respective combinations of the delta models with the base model produce respective versions of the fine-tuned machine learning model; and (Torres, [20] "As shown in FIG. 2, the ML system 130 may serve a variety of NLP applications by integrating one or more light-weight model artifacts 242 into a base model 220 [respective delta models (artifacts) with respect to the base model, respective combinations of the delta models with the base model] to dynamically adapt the base model 220 to a particular NLP task."; [21] "The Adapter Service architecture may be a machine learning system architecture that facilitates generating a plurality of specific models [one or more different machine learning models] from a single base model instance. The Adapter Service architecture includes a model deployment service 240 and other components that dynamically augment a base model 220 with one or more model artifacts 242. [respective delta models (artifacts) with respect to the base model, respective combinations of the delta models with the base model] By integrating the one or more model artifacts 242 into a base model 220, the model deployment service 240 generates a plurality of models [respective versions of the fine-tuned machine learning model] for performing one or more NLP tasks."; [23] "A state of the art transfer learning NLP model can be used as a base model 220 that may be augmented with one or more lightweight adapter layers 226a-n. [respective delta models]") cause placement of both the base model and the respective delta models on the host system, (Torres, [56] "At 604, the model deployment service loads the model artifact into its proper location within the base model [placing both the base model and the respective delta models] architecture to generate a fine-tuned model."; [63] "After inserting the adapter layers and/or output layer included in the new model artifact within the original model artifact, the new model artifact is loaded into memory at 604 [placing both the base model and the respective delta models on the system] to generate the fine-tuned model.") wherein the host system generates respective inferences for requests that invoke one of the respective versions of the fine-tuned machine learning model. (Torres, [21] "By integrating the one or more model artifacts 242 into a base model 220, the model deployment service 240 generates a plurality of models for performing one or more NLP tasks. Using the Adapter Service architecture, a single deployment of the ML system 130 may dynamically serve a variety of NLP models generating predictions [generates respective inferences] for a wide variety of NLP tasks."; [57] "input text is processed using a combination of the base model parameters and the cached parameters of the model artifact specified by the fine-tuned model. NLP predictions [respective inferences] generated by the fine-tuned model are then distributed to an NLP application at 608. ") Torres does not teach, but Clement teaches: host a managed network endpoint, (Clement, [0028] "The model deployment service 110 generates the infrastructure needed to deploy the model and provides an endpoint for the user to run the fine-tuned model."; [0090] "the service provides the user with an endpoint of the deployed model for the user to use to run the model with its inference dataset"; [0031] "The REST APIs are service endpoints [a managed network endpoint] that support a set of HyperText Transfer Protocol (HTTP) operations or methods to create, retrieve, update, delete or access resources of a web service.") It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified Torres to incorporate the teachings of Clement by including APIs as service endpoints. Doing so would provide an endpoint for the user to run a model and to create, retrieve, update, delete or access resources of a web service. (Clement, [0028] "The model deployment service 110 generates the infrastructure needed to deploy the model and provides an endpoint for the user to run the fine-tuned model.";[0031] "The REST APIs are service endpoints that support a set of HyperText Transfer Protocol (HTTP) operations or methods to create, retrieve, update, delete or access resources of a web service.") In regard to claim 2, Torres teach: wherein the placement is caused in response to a scaling event or a rebalancing event detected for the managed network endpoint. (Torres, [26] "The Adapter Service architecture may also more efficiently modify processing, memory, and network resources in response to changes in demand [model placement (in memory) dynamically changes in response to events] for predictions generated by different NLP models."; [21] "By integrating the one or more model artifacts 242 into a base model 220, the model deployment service 240 generates a plurality of models for performing one or more NLP tasks. Using the Adapter Service architecture, a single deployment of the ML system 130 may dynamically serve a variety of NLP models [dynamically serving multiple models, a scaling event] generating predictions for a wide variety of NLP tasks."; [26] "By allowing the ML system 130 to serve multiple models at scale [a scaling event, increases a number of replicas of the machine learning model] within a single deployment, while processing a high volume of requests for each different NLP task... "; [55] "At runtime, the ML system may dynamically adapt the base model [a rebalance event] based on input data to generate one or more fine-tuned models."; the system allow serving multiple models at scale, which include multiple base models with their adapters) In regard to claim 3, Torres teach: wherein the host system is configured to: receive a request to generate an inference using one of the respective versions of the fine-tuned machine learning model at the host system; (Torres, [14] "These elements may work together to automatically route prediction requests, [receiving a request] configure ML systems, serve NLP models, and/or generate NLP predictions. [generate an inference]"; [55] "FIG. 6 illustrates an example of generating NLP predictions at runtime 600. At 602, the ML system receives input data. Input data may include a payload comprising input text, an NLP task identifier, and/or the model artifact location corresponding to the model artifact for the specified NLP task identifier. [using one of the respective versions ] The NLP task identifier may define an NLP task type that specifies one or more target NLP task types to be performed on the input text... NLP predictions [generate an inference] generated by the fine-tuned model are then distributed to an NLP application at 608. ") determine that the respective delta model is loaded into a memory of the host system; (Torres, [56] "At 604, the model deployment service loads the model artifact into its proper location within the base model architecture to generate a fine-tuned model."; [63] "After inserting the adapter layers and/or output layer included in the new model artifact within the original model artifact, the new model artifact is loaded into memory at 604 to generate the fine-tuned model.") generate delta values for the respective delta model for given input for the request; and (Torres, [40] "The training method 500 may be used to update adapter parameter sets [generating delta values] to generate model artifacts [for the respective delta model] optimized for a particular NLP task. In various embodiments, adapter parameters are trained [generating delta values] by inserting the untrained adapter layers and output layer into the base model, freezing the base model weights, and training only the adapter parameters."; [1a] "the generating comprising dynamically exchanging the first model artifact [generating delta values] with a second model artifact comprising one or more adapter layers specific to the second task type;") use the delta values to complete generation of the inference using base values generated by the base model in combination with the delta values. (Torres, [57] "At 606... input text is processed using a combination of the base model parameters and the cached parameters of the model artifact specified by the fine-tuned model. NLP predictions generated by the fine-tuned model are then distributed to an NLP application at 608."; [34] "Integrating the adapter parameters [using the delta values] into the base model 220 transfers knowledge captured in the base model parameters [using base values generated by the base model] to the fine-tuned models. The adapter parameters may then augment the knowledge provided by base parameters to optimize the fine-tuned models for a particular NLP task."; [32] "Adapter parameters included in the one or more adapter layers 226 may introduce, eliminate, and/or re-arrange data points included in the learned manifold by adding additional weights and/or modifying weights provided by the base parameters. By adapting the learned manifold to fit a dataset specific to one or more specific NLP tasks, the one or more adapter layers 226 tune the base model 220 to generate a fine-tuned model that performs one or more particular NLP tasks."; using adapter parameters and base parameters in combination for NLP prediction (inference)) In regard to claim 4, Torres does not teach, but Clement teaches: wherein the managed network endpoint is created in response to one or more requests to create the managed network endpoint and add the plurality of different machine learning models to the managed network endpoint, received via an interface of the machine learning service. (Clement, [0028] "The data management service 106 receives registration requests [in response to requests] for usage of a pre-trained model. The model fine-tuning service 108 generates the infrastructure needed to perform the fine-tuning task and executes the requested fine-tuning on a selected pre-trained model. The model deployment service 110 generates the infrastructure needed to deploy the model and provides an endpoint for the user to run the fine-tuned model. [endpoint is created, and add the fine-tuned/deployed model]"; [0090] "The model deployment service receives a request [in response to requests] to deploy the selected model. In response the model deployment service constructs the infrastructure to run the fine-tuned model and executes the deployment script. Upon successful completion of the deployment, the service provides the user with an endpoint of the deployed model [endpoint is created, and add the deployed model] for the user to use to run the model with its inference dataset"; [0031] "The user interacts with the services of the cloud platform through Representational State Transfer (REST) Application Programming Interfaces (APIs). [via an interface of the service] The REST APIs are service endpoints that support a set of HyperText Transfer Protocol (HTTP) operations or methods to create, retrieve, update, delete or access resources of a web service.") The rationale for combining the teachings of Torres and Clement is the same as set forth in the rejection of claim 1. In regard to claims 6 and 15, Torres teach: wherein the placement is requested in response to a scaling event for the managed network endpoint that increases a number of replicas of the machine learning model at the managed network endpoint. (Torres, [26] "The Adapter Service architecture may also more efficiently modify processing, memory, and network resources in response to changes in demand [model placement (in memory) dynamically changes in response to events] for predictions generated by different NLP models."; [21] "By integrating the one or more model artifacts 242 into a base model 220, the model deployment service 240 generates a plurality of models for performing one or more NLP tasks. Using the Adapter Service architecture, a single deployment of the ML system 130 may dynamically serve a variety of NLP models [dynamically serving multiple models, a scaling event] generating predictions for a wide variety of NLP tasks."; [26] "By allowing the ML system 130 to serve multiple models at scale [a scaling event, increases a number of replicas of the machine learning model] within a single deployment, while processing a high volume of requests for each different NLP task... "; the system allow serving multiple models at scale, which include multiple base models with their adapters) Torres does not teach, but Clement teaches: wherein the host system is associated with a managed network endpoint, and (Clement, [0028] "The model deployment service 110 generates the infrastructure needed to deploy the model and provides an endpoint for the user to run the fine-tuned model."; [0090] "the service provides the user with an endpoint of the deployed model for the user to use to run the model with its inference dataset"; [0031] "The REST APIs are service endpoints [a managed network endpoint] that support a set of HyperText Transfer Protocol (HTTP) operations or methods to create, retrieve, update, delete or access resources of a web service.") The rationale for combining the teachings of Torres and Clement is the same as set forth in the rejection of claim 1. In regard to claims 7 and 16, Torres teach: wherein the placement is requested in response to a rebalance event for the managed network endpoint that moves the machine learning model from a current host system. (Torres, [26] "The Adapter Service architecture may also more efficiently modify processing, memory, and network resources in response to changes in demand [model placement (in memory) dynamically changes in response to events] for predictions generated by different NLP models."; [55] "At runtime, the ML system may dynamically adapt the base model [moves the machine learning model from a current host system, a rebalance event] based on input data to generate one or more fine-tuned models."; [32] "Adapter parameters included in the one or more adapter layers 226 may introduce, eliminate, and/or re-arrange data points included in the learned manifold by adding additional weights and/or modifying weights provided by the base parameters. [moves the machine learning model from a current host system]"; the original base parameters in the base model were changed or replaced, i.e. moving the base model from the system) Torres does not teach, but Clement teaches: wherein the host system is associated with a managed network endpoint, and (Clement, [0028] "The model deployment service 110 generates the infrastructure needed to deploy the model and provides an endpoint for the user to run the fine-tuned model."; [0090] "the service provides the user with an endpoint of the deployed model for the user to use to run the model with its inference dataset"; [0031] "The REST APIs are service endpoints [a managed network endpoint] that support a set of HyperText Transfer Protocol (HTTP) operations or methods to create, retrieve, update, delete or access resources of a web service.") The rationale for combining the teachings of Torres and Clement is the same as set forth in the rejection of claim 1. In regard to claims 8 and 17, Torres teach: wherein a model registry is updated to include the placement of the base model and the respective delta models on the host system to route subsequent requests to requests that invoke one of the respective versions of the fine- tuned machine learning model to the host system. (Torres, [21] "To process a subsequent request for an alternative NLP task, for example, document classification, the ML system 130 may exchange the sentiment analysis adapter layers and output layer for the document classification specific adapter layers and output layers [a model registry is updated] to generate a fine-tuned model for document classification. The ML system 130 may then process the document classification request and output the document classification prediction."; a single model registry is updated with new adapter/output layers for routing subsequent (alternative/different) requests) Torres does not teach, but Clement teaches: wherein the host system is associated with a managed network endpoint, and (Clement, [0028] "The model deployment service 110 generates the infrastructure needed to deploy the model and provides an endpoint for the user to run the fine-tuned model."; [0090] "the service provides the user with an endpoint of the deployed model for the user to use to run the model with its inference dataset"; [0031] "The REST APIs are service endpoints [a managed network endpoint] that support a set of HyperText Transfer Protocol (HTTP) operations or methods to create, retrieve, update, delete or access resources of a web service.") The rationale for combining the teachings of Torres and Clement is the same as set forth in the rejection of claim 1. In regard to claims 9 and 18, Torres teach: wherein the placement is requested in response to add the machine learning model to the managed network endpoint. (Torres, [26] "The Adapter Service architecture may also more efficiently modify processing, memory, and network resources in response to changes in demand [model placement (in memory) dynamically changes in response to events] for predictions generated by different NLP models."; [20] "As shown in FIG. 2, the ML system 130 may serve a variety of NLP applications by integrating one or more light-weight model artifacts 242 into a base model 220 to dynamically adapt the base model 220 [add the machine learning model (the base model)] to a particular NLP task."; [21] "The Adapter Service architecture includes a model deployment service 240 and other components that dynamically augment a base model 220 with one or more model artifacts 242. By integrating the one or more model artifacts 242 into a base model 220, the model deployment service 240 generates a plurality of models for performing one or more NLP tasks.") Torres does not teach, but Clement teaches: wherein the host system is associated with a managed network endpoint, and (Clement, [0028] "The model deployment service 110 generates the infrastructure needed to deploy the model and provides an endpoint for the user to run the fine-tuned model."; [0090] "the service provides the user with an endpoint of the deployed model for the user to use to run the model with its inference dataset"; [0031] "The REST APIs are service endpoints [a managed network endpoint] that support a set of HyperText Transfer Protocol (HTTP) operations or methods to create, retrieve, update, delete or access resources of a web service.") The rationale for combining the teachings of Torres and Clement is the same as set forth in the rejection of claim 1. In regard to claim 11, Torres teach: wherein the placement is requested in response to a scaling event to scale up from no replicas of the machine learning model to at least one replica of the machine learning model. (Torres, [26] "The Adapter Service architecture may also more efficiently modify processing, memory, and network resources in response to changes in demand [model placement (in memory) dynamically changes in response to events] for predictions generated by different NLP models."; [40] "the base model can be trained separately if no appropriate pre-trained base model exists. [scale up from zero (no replicas of the machine learning model or the base model)] As shown in FIG. 5, 502-508 describe training the base parameters of the base model and 510-516 describe training the adapter parameters of the adapter layers and/or output layer.") Torres does not teach, but Clement teaches: wherein the host system is associated with a managed network endpoint, and (Clement, [0028] "The model deployment service 110 generates the infrastructure needed to deploy the model and provides an endpoint for the user to run the fine-tuned model."; [0090] "the service provides the user with an endpoint of the deployed model for the user to use to run the model with its inference dataset"; [0031] "The REST APIs are service endpoints [a managed network endpoint] that support a set of HyperText Transfer Protocol (HTTP) operations or methods to create, retrieve, update, delete or access resources of a web service.") The rationale for combining the teachings of Torres and Clement is the same as set forth in the rejection of claim 1. In regard to claims 13 and 20, Torres teach: wherein the host system is one of a plurality of different host systems (Torres, [21] "The Adapter Service architecture may be a machine learning system architecture that facilitates generating a plurality of specific models [a plurality of different host systems] from a single base model instance... the model deployment service 240 generates a plurality of models [a plurality of different host systems] for performing one or more NLP tasks."; in light of specification [0028] 'placed across different hosts, such as hosts 132a, 132b, and 132c' where different hosts are host instances including models. Torres teaches that the service can generate multiple models (host instances), where each of the models includes a base model and a lightweight model (artifact/adapter layers)) associated with a managed network endpoint of the machine learning service, (Torres, [16] "NLP application 150 may be a hardware and/or software component accessible to client 160 through network 140 (e.g., an application hosted by an application server [a network endpoint] or other server computer)."; [18] "one or more of client 160, NLP application instance 150, and/or ML system 130 may communicate with one another through network 140. For example, communication between the elements may be facilitated by one or more application programming interfaces (APIs)."; a server, listens for requests, and lets a remote client reach it across a network connection via API) wherein the base model and the plurality of delta models are included in a plurality of different machine learning models (Torres, [20] "As shown in FIG. 2, the ML system 130 may serve a variety of NLP applications by integrating one or more light-weight model artifacts 242 into a base model 220 [the base model and the plurality of delta models] to dynamically adapt the base model 220 to a particular NLP task."; [21] "The Adapter Service architecture may be a machine learning system architecture that facilitates generating a plurality of specific models [a plurality of different machine learning models] from a single base model instance. The Adapter Service architecture includes a model deployment service 240 and other components that dynamically augment a base model 220 with one or more model artifacts 242. [the base model and the plurality of delta models] By integrating the one or more model artifacts 242 into a base model 220, the model deployment service 240 generates a plurality of models [a plurality of different machine learning models] for performing one or more NLP tasks.") associated with the managed network endpoint, and (Torres, [16] "NLP application 150 may be a hardware and/or software component accessible to client 160 through network 140 (e.g., an application hosted by an application server [a network endpoint] or other server computer)."; [18] "one or more of client 160, NLP application instance 150, and/or ML system 130 may communicate with one another through network 140. For example, communication between the elements may be facilitated by one or more application programming interfaces (APIs)."; a server, listens for requests, and lets a remote client reach it across a network connection via API) Torres does not teach, but Clement teaches: wherein the managed network endpoint is created in response to one or more requests to create the managed network endpoint and add the plurality of different machine learning models to the managed network endpoint, received via an interface of the machine learning service. (Clement, [0028] "The data management service 106 receives registration requests [in response to requests] for usage of a pre-trained model. The model fine-tuning service 108 generates the infrastructure needed to perform the fine-tuning task and executes the requested fine-tuning on a selected pre-trained model. The model deployment service 110 generates the infrastructure needed to deploy the model and provides an endpoint for the user to run the fine-tuned model. [endpoint is created, and add the fine-tuned/deployed model]"; [0090] "The model deployment service receives a request [in response to requests] to deploy the selected model. In response the model deployment service constructs the infrastructure to run the fine-tuned model and executes the deployment script. Upon successful completion of the deployment, the service provides the user with an endpoint of the deployed model [endpoint is created, and add the deployed model] for the user to use to run the model with its inference dataset"; [0031] "The user interacts with the services of the cloud platform through Representational State Transfer (REST) Application Programming Interfaces (APIs). [via an interface of the service] The REST APIs are service endpoints that support a set of HyperText Transfer Protocol (HTTP) operations or methods to create, retrieve, update, delete or access resources of a web service.") The rationale for combining the teachings of Torres and Clement is the same as set forth in the rejection of claim 1. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Sheng ("S-LoRA: Serving Thousands of Concurrent LoRA Adapters" 20231107) teaches (Sheng, p. 1, Abstract "we present S-LoRA, a system designed for the scalable serving of many LoRA adapters.") Any inquiry concerning this communication or earlier communications from the examiner should be directed to SU-TING CHUANG whose telephone number is (408)918-7519. The examiner can normally be reached Monday - Thursday 8-5 PT. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Usmaan Saeed can be reached at (571) 272-4046. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /S.C./Examiner, Art Unit 2146 /USMAAN SAEED/Supervisory Patent Examiner, Art Unit 2146
Read full office action

Prosecution Timeline

Nov 24, 2023
Application Filed
Aug 03, 2026
Non-Final Rejection mailed — §102, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12645997
INDIVIDUALIZED CLASSIFICATION THRESHOLDS FOR MACHINE LEARNING MODELS
3y 3m to grant Granted Jun 02, 2026
Patent 12626164
SYSTEM AND METHOD FOR REDUCTION OF DATA TRANSMISSION BY DATA RECONSTRUCTION
4y 0m to grant Granted May 12, 2026
Patent 12626106
MACHINE LEARNING MODELS FOR BEHAVIOR UNDERSTANDING
3y 11m to grant Granted May 12, 2026
Patent 12626140
SYSTEMS AND METHODS FOR ONLINE TIME SERIES FORCASTING
3y 9m to grant Granted May 12, 2026
Patent 12619890
LEARNING PATTERN DICTIONARY FROM NOISY NUMERICAL DATA IN DISTRIBUTED NETWORKS
6y 6m to grant Granted May 05, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
51%
Grant Probability
90%
With Interview (+39.5%)
4y 6m (~1y 9m remaining)
Median Time to Grant
Low
PTA Risk
Based on 108 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month