Prosecution Insights
Last updated: August 16, 2026
Application No. 18/542,676

MACHINE LEARNING MODEL ADMINISTRATION AND OPTIMIZATION

Non-Final OA §101§103§112
Filed
Dec 16, 2023
Priority
Dec 16, 2022 — provisional 63/433,124 +2 more
Examiner
COLEMAN, PAUL
Art Unit
Tech Center
Assignee
C3.ai Inc.
OA Round
1 (Non-Final)
68%
Grant Probability
Favorable
1-2
OA Rounds
11m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 68% — above average
68%
Career Allowance Rate
13 granted / 19 resolved
+8.4% vs TC avg
Strong +43% interview lift
Without
With
+42.9%
Interview Lift
resolved cases with interview
Typical timeline
3y 7m
Avg Prosecution
17 currently pending
Career history
39
Total Applications
across all art units

Statute-Specific Performance

§101
34.5%
-5.5% vs TC avg
§103
43.9%
+3.9% vs TC avg
§102
4.7%
-35.3% vs TC avg
§112
16.9%
-23.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 19 resolved cases

Office Action

§101 §103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statements (IDS) submitted on 01/11/2025, 05/23/2025, 09/11/2025, 09/24/2025, 11/07/2025, 01/15/2026, and 02/27/2026 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statements are being considered by the examiner. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 9, 11-15, and 17-20 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 9 recites the limitation "one or more additional model processing units" in “monitoring utilization of one or more additional model processing units for the multiple instances of the versioned model”. There is insufficient antecedent basis for this limitation in the claim. In particular, neither claim 9 nor the claims from which claim 9 depends previously recite a model processing unit relative to which the recited model processing units are “additional”. It is therefore unclear whether the “additional model processing units” are additional to a processing unit used to deploy the versioned model, additional to processing units used for the multiple instances, or additional to some other unrecited processing resource. Accordingly, the scope of the monitoring and load-balancing operations is indeterminate. Claim 9 further recites “a threshold condition of the computing environment”. There is insufficient antecedent basis for “the computing environment” because neither claim 9 nor the claims from which claim 9 depends previously identify a particular computing environment. It is unclear whether the recited computing environment refers to the configured run-time instance, the computing resources on which the model instances execute, a distributed computing environment, or another environment. Consequently, the boundary of the claimed threshold condition cannot be determined. Claim 11 recites the limitation "memory storing instructions that" in “A system comprising: memory storing instructions that, when executed by the one or more processors, cause the system to perform:”. There is insufficient antecedent basis for this limitation in the claim. The claim does not previously recite one or more processors as an element of the claimed system. The omission renders the scope of claim 11 indeterminate because it is unclear whether the one or more processors are components of the claimed system, external processors that execute instructions stored in the claimed memory, processors associated with the model inference service, or processors associated with the run-time instances. Claim 11 further recites that the instructions cause the system “to perform: a model inference service for instantiating different versioned model”. The phrase “perform: a model inference service” is grammatically and structurally unclear. A “model inference service” appears to identify a software or system component rather than an operation that the system “performs”. It is unclear whether the claim requires the instructions to: implement a model inference service; cause a model inference service to perform the subsequently recited operations; perform a method identified as a model inference service; or merely instantiate different versioned models. These interpretations impose different limitations on the claimed system. Further, the singular phrase “different versioned model” does not clearly identify whether one versioned model or multiple different versioned models are instantiated. Accordingly, the metes and bounds of claim 11 cannot be determined. Claims 12-15 depend from claim 11 and incorporate all the limitations of claim 11, including the indefinite limitations identified above. Accordingly, claims 12-15 are likewise indefinite under 35 U.S.C. § 112(b). Claim 13 recites the limitation "the hierarchical repository" in “The system of claim 11, wherein the hierarchical repository comprises a catalogue of additional baseline models pretrained on datasets from different domains”. There is insufficient antecedent basis for this limitation in the claim. Claim 11 recites “a model registry comprising a hierarchical structure”, but does not recite a “hierarchical repository”. It is unclear whether “the hierarchical repository” refers to the model registry, the hierarchical structure within the model registry, a separate repository containing the hierarchical structure, or another repository. Although the terms may be related, the claim does not establish that they identify the same claimed element. Claim 13 further recites “the additional model records associated with each additional baseline model”. There is insufficient antecedent basis for “the additional model records”. Claim 11 recites “one or more child model records” and “new model records”, but does not recite “additional model records”. It is therefore unclear whether the “additional model records” are: additional child model records; the new model records generated by capturing changes; records associated with the additional baseline models; or a separate category of model record. Because the alternative interpretations impose different relationships between the baseline models and model records, the scope of claim 13 is indeterminate. Claim 14 recites the limitation “the instantiating different versioned are capable of multiple generative tasks including conversational, summarizing, computational, predictive, visualization”. The phrase “the instantiating different versioned” is grammatically incomplete and does not identify a claimed subject that is capable of performing the recited generative tasks. For example, it is unclear whether Applicant intends to recite: that the different versioned models are capable of the generative tasks; that the model inference service is capable of the generative tasks; that the act of instantiating different versioned models enables the generative tasks; or that the run-time instances are capable of the generative tasks. The Office cannot select among these materially different interpretations without improperly rewriting the claim. Additionally, the listed terms “conversational, summarizing, computational, predictive, visualization” do not have parallel grammatical form and do not clearly identify the claimed tasks. For example, “conversational”, “computational”, and “predictive” are adjectives, “summarizing” may identify an act, and “visualization” is a noun. It is therefore unclear what particular functionality is required by each listed term. Accordingly, claim 14 does not clearly identify the claimed element performing the generative tasks or the boundaries of the required generative functionality. Claim 17 reciters “wherein one or more versioned models are selected and replaced at run-time”. There is insufficient antecedent basis for “one or more versioned models” because claim 16 recites model configuration records, but does not recite a versioned model or establish that a versioned model is generated from the retrieved model configuration records. It is therefore unclear whether the “versioned models” are: represented by the retrieved model configuration records; assembled using the retrieved model configuration records; previously deployed models; models selected independently of the retrieved records; or models associated with records other than those retrieved in claim 16. The term “replaced” also fails to identify what is replaced or what performs the replacement. The claim may mean that a selected versioned model replaces a previously executing model, that one of the selected versioned models is replaced by another selected versioned model, or that model configuration records are replaced at run-time. These alternatives define materially different methods. Accordingly, the relationships between the retrieved model configuration records, the selected versioned models, and the replacement operation is unclear. Claim 18 reciters “the domain-specific dataset”. There is insufficient antecedent basis for this limitation because neither claim 18 nor the claims form which claim 18 depends previously recite a domain-specific dataset. It is therefore unclear whether “the domain-specific dataset” is associated with a baseline model, associated with a particular customer, identified by one of the model configuration records, or obtained from another source. Claim 18 further recites that “each of the selected one or more models are pre-trained on customer-specific data subsequent to being trained on the domain-specific dataset”. The recitation is unclear because “pre-trained” ordinarily identifies training that precedes later training, whereas the claim expressly requires the customer-specific pre-training to occur “subsequent to” training on the domain-specific dataset. It is therefore unclear whether the customer-specific operation is intended to be: pre-training; additional training; fine-tuning; adapter training; or another modification performed after domain-specific training. The different interpretations may result in materially different model-training sequences. Accordingly, the scope of claim 18 is indeterminate. Claim 19 recites “compressing at least a portion of the plurality of model parameters of the model”. There is insufficient antecedent basis for both “the plurality of model parameters” and “the model”. Claim 16 recites a plurality of model configuration records and retrieval of one or more of those records, but does not recite a model, a plurality of model parameters, or a relationship in which the retrieved records constitute or identify the parameters of a particular model. It is therefore unclear whether “the model” is: a model represented by one retrieved model configuration record; a model assembled form multiple retrieved model configuration records; a versioned model as recited in claim 17; a baseline model; a model selected independently of the retrieved records; or another unrecited model. It is also unclear whether “the plurality of model parameters” refers to parameters contained in one model configuration record, parameters distributed among multiple records, all parameters of the model, or a separately obtained parameter set. Because the claim does not establish which model or parameters are compressed, deployed, and decompressed, the scope of claim 19 is indeterminate. Claim 20 recites “the compressing”, “the decompressing”, “the plurality of model parameters”, and “the plurality of quantized model parameters”. There is insufficient antecedent basis for these terms in claim 16, from which claim 20 directly depends. Claim 16 does not recite a compressing operation, a decompressing operation, a plurality of model parameters, or a plurality of quantized model parameters. These limitations are instead introduced in claim 19. Consequently, when claim 20 is read according to its stated dependency from claim 16, it is unclear: what subject matter undergoes “the compressing”; what subject matter undergoes “the decompressing”; which model contains “the plurality of model parameters”; and how “the plurality of quantized model parameters” relates to any subject matter recited in claim 16. The Office notes that claim 20 appears to have been intended to depend from claim 19. However, the Office cannot correct the dependency by interpretation because doing so would alter the express dependency and substantive limitations of the claim. Accordingly, claim 20 is indefinite under 35 U.S.C. § 112(b). Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-18 are rejected under 35 U.S.C. 101 for reciting an abstract idea without significantly more. Regarding claim 1 Claim 1 – Step 1 – Is the claim to a process, machine, manufacture or composition of matter? Yes, the claim is to a process. Claim 1 – Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon? Yes, the claim recites an abstract idea. “selecting, by one or more processing devices, a baseline model and one or more child model records from a hierarchical structure based on the request, wherein the baselines model and the one more child model records include model metadata with parameters describing dependencies, access control, and deployment configurations;” – this limitation recites evaluating information contained in a request and stored model metadata and, based on that evaluation selecting corresponding model records. Comparing received information to stored information and selecting records based on dependencies, access-control criteria, and deployment criteria constitutes an evaluation or judgment that may be practically performed in the human mind, or with pen and paper, by reviewing a list or hierarchy of model records and identifying the records satisfying the request criteria. See MPEP § 2106.04(a)(2)(III). Claim 1 – Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application? No. There are no additional elements that integrate the judicial exception into a practical application. The additional elements: “receiving a request associated with a machine learning application, wherein the request includes application information, user information, and execution information;” – this limitation recites obtaining information. Merely receiving or gathering information does not impose a meaningful technological limitation. Data gathering that supplies information used in the subsequently recited selection process, therefore, constitutes insignificant extra-solution activity under MPEP § 2106.05(g). “assembling a versioned model of the baseline model using the one more child model records and associated dependencies;” – this limitation applies the result of the abstract selection by generically assembling a versioned model using the selected records and dependencies. The claim does not recite how the child model records alter the baseline model, how the associated dependencies are resolved, how conflicting model parameters are reconciled, how the model records are combined, or any particular model-assembly algorithm. The limitation therefore does not reflect a specific improvement to computer functionality or machine-learning technology and instead constitutes generic computer implementation of the selected information. See MPEP § 2106.05(f). “and deploying the versioned model in a configured run-time instantiation for use by the application based on the associated metadata.” – this limitation generically apples the result of the preceding selection and assembly by deploying the model in a runtime environment. The claim does not recite a particular runtime architecture, model-loading procedure, hardware configuration, resource-allocation technique, container configuration, virtual-machine configuration, memory-management technique, or deployment protocol. Nor does the claim recite a particular improvement in model inference, latency, memory use, processor utilization, or deployment reliability. The terms “configured run-time instantiation” and “machine learning application” therefore merely identify the technological environment or field in which the selected and assembled model is used. Generally linking an abstract idea to a particular technological environment does not integrate the abstract idea into a practical application. See MPEP § 2106.05(h). Claim 1 – Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception? No. There are no additional elements that amount to significantly more than the judicial exception. The additional elements are: “receiving a request associated with a machine learning application, wherein the request includes application information, user information, and execution information;” – this limitation recites the generic computer function of receiving electronic information for subsequent processing. Receiving input data is well-understood, routine, and conventional (WURC) computer activity and is also insignificant data-gathering activity. See MPEP §§ 2106.05(d) and 2106.05(g). “assembling a versioned model of the baseline model using the one more child model records and associated dependencies;” – this limitation recites model assembly only at a result-oriented level. It does not recite an unconventional model-assembly algorithm, particular parameter-delta operation, improved memory arrangement, specialized hardware, or non-generic interaction among computer components. The specification similarly states that a model inference system or model deployment module assembles the versioned model, but does not identify in connection with claim 1 a particular algorithm required to perform the broadly recited assembly. See spec. ¶[0112]. The broad function of assembling software or model components according to stored records and dependency information does not, as recited, add an inventive concept beyond using generic computer components to apply the selected information. “and deploying the versioned model in a configured run-time instantiation for use by the application based on the associated metadata.” – this limitation recites generic software or model deployment in a runtime environment at a high level of generality. The claim does not recite unconventional deployment hardware, an improved runtime architecture, a particular memory-management mechanism, or any non-generic arrangement of computing components. The versioned model is merely deployed using the associated metadata for its ordinary intended use. See MPEP §§ 2106.05(f), and 2106.05(h). Regarding claim 2 Claim 2 – Step 1 – Is the claim to a process, machine, manufacture or composition of matter? Yes, the claim is to a process. Claim 2 – Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon? Yes, the claim recites an abstract idea. “determining compatibility between the application information and execution information of the request with dependencies and deployment configurations from model metadata,” – this limitation recites reviewing the application information and execution information contained in the request, comparing that information to dependencies and deployment configuration described by the model metadata, and determining whether the information is compatible. Comparing information to stored requirements or criteria and determining whether the information satisfies those requirements constitutes an evaluation or judgment that may be practically performed in the human mind, or with pen and paper, by reviewing the request information and the listed dependencies and deployment configurations and deciding whether they are compatible. See MPEP § 2106.04(a)(2)(III). “and further determining access control of the model metadata and the user information of the request.” – this limitation recites reviewing the user information contained in the request and access-control information associated with the model metadata, and determining whether access should be permitted. Determining whether a user satisfies access-control requirements is an evaluation or judgement that may be practically performed in the human mind, or with pen and paper, by comparing the user’s identity, role, authorization, or privileges to stored access-control criteria. See MPEP § 2106.04(a)(2)(III). Claim 2 – Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application? No. There are no additional elements that integrate the judicial exception into a practical application. Claim 2 – Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception? No. There are no additional elements that amount to significantly more than the judicial exception. Regarding claim 3 Claim 3 – Step 1 – Is the claim to a process, machine, manufacture or composition of matter? Yes, the claim is to a process. Claim 3 – Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon? Yes, the claim recites an abstract idea. Claim 3 – Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application? No. There are no additional elements that integrate the judicial exception into a practical application. The additional elements: “wherein the child model records comprise intermediate representations of the baseline model with changed parameters from a previous instantiation of the baseline model.” – this limitation broadly requires storing or representing changed model parameters relative to a previous model instantiation. The limitation describes the content of the model records and the desired result of representing parameter changes, without reciting a specific technical mechanism that improves computer or machine-learning functionality. See MPEP § 2106.05(a). Claim 3 – Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception? No. There are no additional elements that amount to significantly more than the judicial exception. The additional elements are: “wherein the child model records comprise intermediate representations of the baseline model with changed parameters from a previous instantiation of the baseline model.” – this limitation recites storing or representing differences between versions of a model at a high level of generality. The claim does not recite an unconventional representation format, parameter-difference algorithm, storage architecture, or interaction among non-generic computer components. Instead, the limitation uses model records for their ordinary informational function of identifying parameter changes associated with a prior model version. See MPEP §§ 2106.05(d) and 2106.05(f). Regarding claim 4 Claim 4 – Step 1 – Is the claim to a process, machine, manufacture or composition of matter? Yes, the claim is to a process. Claim 4 – Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon? Yes, the claim recites an abstract idea. Claim 4 – Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application? No. There are no additional elements that integrate the judicial exception into a practical application. The additional elements: “wherein the baseline model is pre-trained on a general domain dataset, and wherein the one or more child model records comprise intermediate representations with changed parameters of the baseline model trained on an enterprise specific dataset.” – this limitation broadly recites training a baseline model on one dataset and representing changed parameters associated with training on another dataset. The claim does not recite a particular model architecture, training algorithm, parameter-update technique, intermediate-representation format, or technical mechanism that improves model accuracy, memory use, training efficiency, or computer operation. Instead, the limitation merely specifies the source and type of data used to train the models and the resulting parameter information. See MPEP §§ 2106.04(d)(1) and 2106.05(a). Claim 4 – Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception? No. There are no additional elements that amount to significantly more than the judicial exception. The additional elements are: “wherein the baseline model is pre-trained on a general domain dataset, and wherein the one or more child model records comprise intermediate representations with changed parameters of the baseline model trained on an enterprise specific dataset.” – this limitation recites model training and storage of resulting parameter changes only at a high level of generality. The claim does not require unconventional training hardware, a particular non-generic training technique, an improved parameter-storage architecture, or a specific technical implementation. The models and records are used according to their ordinary functions of learning from data and representing trained parameter values. See MPEP §§ 2106.05(d) and 2106.05(f). Regarding claim 5 Claim 5 – Step 1 – Is the claim to a process, machine, manufacture or composition of matter? Yes, the claim is to a process. Claim 5 – Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon? Yes, the claim recites an abstract idea. Claim 5 – Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application? No. There are no additional elements that integrate the judicial exception into a practical application. The additional elements: “wherein the deployment configurations determine a set of computing requirements for the run-time instance of the versioned model.” – this limitation merely uses stored deployment information to identify resource requirements. It does not recite a particular resource-allocation algorithm, improved runtime architecture, dynamic provisioning technique, or specific manner of configured computing resources. The limitation states the desired result of determining computing requirements without reciting a specific technological improvement to computer operation or runtime deployment. See MPEP §§ 2106.04(d)(1) and 2106.05(a). It also generally limits the abstract determination to the technological environment of a runtime instance. See MPEP § 2106.05(h). Claim 5 – Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception? No. There are no additional elements that amount to significantly more than the judicial exception. The additional elements are: “wherein the deployment configurations determine a set of computing requirements for the run-time instance of the versioned model.” – this limitation recites generic evaluation of configuration information to identify computing-resource requirements at a high level of generality. It does not require unconventional hardware, a non-generic resource-allocation mechanism, or an improved configuration technique. The added limitation does not provide an inventive concept beyond the abstract model-selection and deployment process. See MPEP §§ 2106.05(d) and 2106.05(f). Regarding claim 6 Claim 6 – Step 1 – Is the claim to a process, machine, manufacture or composition of matter? Yes, the claim is to a process. Claim 6 – Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon? Yes, the claim recites an abstract idea. Claim 6 – Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application? No. There are no additional elements that integrate the judicial exception into a practical application. The additional elements: “pre-loading a set of model configurations comprising at least one or more of: model weights, adapter instructions.” – this limitation broadly recites loading model weights or adapter instructions before they are used. The claim does not recite a particular pre-loading technique, memory arrangement, caching mechanism, loading sequence, adapter architecture, or specific improvement in latency, memory use, or modal execution. The limitation merely uses generic computer functionality to prepare model information for later assembly and generally applies the abstract model-selection process in the machine-learning environment. See MPEP §§ 2106.05(f) and 2106.05(h). Claim 6 – Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception? No. There are no additional elements that amount to significantly more than the judicial exception. The additional elements are: “pre-loading a set of model configurations comprising at least one or more of: model weights, adapter instructions.” – this limitation recites generic loading of stored model data or instructions at a high level of generality. It does not require unconventional hardware, a non-generic memory architecture, or a particular improved model-loading technique. The limitation does not provide an inventive concept beyond using generic computing components to implement the abstract model-selection and deployment process. See MPEP §§ 2106.05(d) and 2106.05(f). Regarding claim 7 Claim 7 – Step 1 – Is the claim to a process, machine, manufacture or composition of matter? Yes, the claim is to a process. Claim 7 – Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon? Yes, the claim recites an abstract idea. “wherein the hierarchical structure comprises a catalogue of different baseline models that are pre-trained with different domain specific datasets, and child model records associated with each different baseline model are generated based on an intermediate record.” – this limitation characterizes how model information is organized and associated within the hierarchical structure. Organizing baseline models in a catalogue and associating child records with respective baseline models constitutes organizing and managing information. To the extent the limitation requires determining which child records corresponds to which baseline model or intermediate record, it involves evaluation and categorization that may be practically performed in the human mind or with pen and paper. See MPEP § 2106.04(a)(2)(III). Claim 7 – Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application? No. There are no additional elements that integrate the judicial exception into a practical application. Claim 7 – Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception? No. There are no additional elements that amount to significantly more than the judicial exception. Regarding claim 8 Claim 8 – Step 1 – Is the claim to a process, machine, manufacture or composition of matter? Yes, the claim is to a process. Claim 8 – Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon? Yes, the claim recites an abstract idea. “capturing changes to the versioned model as new model records with new model metadata in the hierarchical repository.” – this limitation recites identifying changes to model information and recording those changes as new records and metadata. Identifying, organizing, and recording changed information are evaluations and information-management activities that may be practically performed in the human mind or with pen and paper. See MPEP § 2106.04(a)(2)(III). Claim 8 – Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application? No. There are no additional elements that integrate the judicial exception into a practical application. The additional elements: “receiving multiple requests received for one or more additional instances of the versioned model;” – this limitation merely gathers requests for later processing and constitutes insignificant data-gathering activity. See MPEP § 2106.05(g). “deploying multiple instances of the versioned model;” – this limitation generically deploys additional model instances without reciting a particular deployment architecture, resource-allocation technique, or improvement to model execution. It merely applies the abstract model-management process in a machine-learning runtime environment. See MPEP §§ 2106.05(f) and 2106.05(h). Claim 8 – Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception? No. There are no additional elements that amount to significantly more than the judicial exception. The additional elements are: “receiving multiple requests received for one or more additional instances of the versioned model;” – this limitation recites generic receipt of electronic requests and insignificant data gathering. See MPEP §§ 2106.05(d) and 2106.05(g). “deploying multiple instances of the versioned model;” – this limitation recites generic deployment of software or model instances at a high level of generality. It does not require unconventional hardware, an improved runtime architecture, or a non-generic deployment technique. See MPEP §§ 2106.05(d) and 2106.05(f). Regarding claim 9 Claim 9 – Step 1 – Is the claim to a process, machine, manufacture or composition of matter? Yes, the claim is to a process. Claim 9 – Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon? Yes, the claim recites an abstract idea. “monitoring utilization of one or more additional model processing units for the multiple instances of the versioned model;” – this limitation recites observing resource-utilization information. Monitoring or observing information, without more, constitutes a mental process that may be practically performed by reviewing utilization values or records. See MPEP § 2106.04(a)(2)(III). Claim 9 – Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application? No. There are no additional elements that integrate the judicial exception into a practical application. The additional elements: “and executing one or more load-balancing operations to terminate execution of the one or more additional instances of the versioned model based on a threshold condition of the computing environment.” – this limitation applies the utilization determination by generically terminating model instances when a threshold condition is met. The claim does not recite a particular load-balancing algorithm, threshold-calculation technique, processor-allocation method, or improved computing architecture. It therefore states the desired result of load balancing without reciting a specific technological improvement to resource management or computer operation. See MPEP §§ 2106.04(d)(1), 2106.05(a), and 2106.05(f). Claim 9 – Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception? No. There are no additional elements that amount to significantly more than the judicial exception. The additional elements are: “and executing one or more load-balancing operations to terminate execution of the one or more additional instances of the versioned model based on a threshold condition of the computing environment.” – this limitation recites generic load balancing and termination of computing instances at a high level of generality. It does not require unconventional hardware, a non-generic threshold mechanism, or a particular improved load-balancing technique. The added limitation does not provide an inventive concept beyond implementing the abstract monitoring and model-management process using generic computing functionality. See MPEP §§ 2106.05(d) and 2106.05(f). Regarding claim 10 Claim 10 – Step 1 – Is the claim to a process, machine, manufacture or composition of matter? Yes, the claim is to a process. Claim 10 – Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon? Yes, the claim recites an abstract idea. Claim 10 – Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application? No. There are no additional elements that integrate the judicial exception into a practical application. The additional elements: “wherein deploying the versioned model further comprises the machine learning application executing instructions to transmit control system commands for one or more industrial devices.” – this limitation merely applies the abstract model-selection and deployment process in the particular technological environment of industrial-device control. The claim does not recite any particular industrial device, control algorithm, control protocol, actuator operation, or improvement to computer or industrial-device functionality. Instead, it merely limits use of the judicial exception to the field of industrial control. Such a limitation merely links the judicial exception to a particular technological environment or field of use and therefore does not integrate the exception into a particular application. See MPEP § 2106.05(h). Claim 10 – Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception? No. There are no additional elements that amount to significantly more than the judicial exception. The additional elements are: “wherein deploying the versioned model further comprises the machine learning application executing instructions to transmit control system commands for one or more industrial devices.” – this limitation recites the well-understood, routine, and conventional (WURC) transmission of control-system commands following the abstract model-selection and deployment process. The claim does not require a particular control architecture, unconventional command-generation technique, or improvement to industrial-device operation. Rather, the limitation merely uses the result of the abstract idea in the conventional field of industrial-device control and therefore does not provide an inventive concept beyond implementing the abstract idea in a particular field of use. See MPEP §§ 2106.05(d), 2106.05(f), and 2106.05(h). Regarding claim 11 Claim 11 – Step 1 – Is the claim to a process, machine, manufacture or composition of matter? Yes, the claim is to a machine. Claim 11 – Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon? Yes, the claim recites an abstract idea. “wherein a model registry comprises a hierarchical structure with a baselines model and child model records that include model metadata with parameters describing dependencies and deployment configurations” – this limitation recites organizing model information in a hierarchy and associating that information with metadata describing dependencies and deployment configurations. Organizing, classifying, and associating information are mental processes because they may be practically performed in the human mind, or with pen and paper, by reviewing model records and organizing them according to dependency and deployment information. See MPEP § 2106.04(a)(2)(III). “and wherein the model registry is updated with new model records based on the changes to the baseline model from multiple run-time instances.” – this limitation recites identifying changes to model information and recording those changes as new model records. Identifying changed information and updating a record based on those changes constitutes information management and may be practically performed mentally or with pen and paper by reviewing changes and entering corresponding new records. See MPEP § 2106.04(a)(2)(III). Claim 11 – Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application? No. There are no additional elements that integrate the judicial exception into a practical application. The additional elements: “memory storing instructions that, when executed by the one or more processors, cause the system to perform:” – the recited memory and processors merely implement the claimed functions using generic computer components. The claim does not recite a particular processor, particular memory architecture, or non-generic arrangement of computer components. See MPEP § 2106.05(f). “a model inference service for instantiating different versioned model to service a machine-learning application” – this limitation generally links the abstract model-record management process to the machine-learning environment. The claim does not recite a particular inference-service architecture, model-instantiation technique, or technical improvement in inference operation. See MPEP §§ 2106.05(f) and 2106.05(h). “to assemble the different versioned model” and “wherein each versioned model is assembled with the baseline model using the one more child model records and associated dependencies” – these limitations generically apply the organized model-record information to assemble a versioned model. The claim does not recite how the child model records modify the baseline model, how dependencies are resolved, how conflicting parameters are handled, or any particular model-assembly algorithm. Thus, the limitations state the desired result of assembling a model version without reciting a specific improvement to computer functionality or machine-learning technology. See MPEP §§ 2106.04(d)(1), 2106.05(a), and 2106.05(f). “wherein the model inference service concurrently deploys multiple run-time instances with different versions of the model for different user sessions” – this limitation recites concurrent deployment of model instances at a high level of generality. The claim does not recite a particular runtime architecture, resource-allocation mechanism, load-balancing technique, memory-management technique, or deployment protocol. It therefore does not impose a meaningful technological limitation beyond generally applying the abstract model-record management process in a runtime machine-learning environment. See MPEP §§ 2106.05(f) and 2106.05(h). Claim 11 – Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception? No. There are no additional elements that amount to significantly more than the judicial exception. The additional elements are: “memory storing instructions that, when executed by the one or more processors, cause the system to perform:” – this limitation recites generic memory and processors performing ordinary computer functions. Generic processors and memory executing instructions are well-understood, routine, and conventional (WURC) computer components. See MPEP §§ 2106.05(d) and 2106.05(f). “a model inference service for instantiating different versioned model to service a machine-learning application” – this limitation recites an inference service at a high level of generality. The claim does not require unconventional inference-service hardware, a particular service architecture, or a non-generic implementation. See MPEP §§ 2106.05(d) and 2106.05(f). “to assemble the different versioned model” and “wherein each versioned model is assembled with the baseline model using the one more child model records and associated dependencies” – these limitations recite model assembly using stored model records for dependency information at a functional level. The claim does not require an unconventional assembly algorithm, particular dependency-resolution technique, parameter-delta implementation, conflict-resolution process, improved data structure, or non-generic interaction among computer components. These limitations therefore amount to generic computer implementation of the abstract model-record management process and do not supply an inventive concept. See MPEP §§ 2106.05(d) and 2106.05(f). “wherein the model inference service concurrently deploys multiple run-time instances with different versions of the model for different user sessions” – recites deploying multiple runtime instances at a high level of generality. The claim does not require unconventional deployment hardware, a particular runtime architecture, a specific resource-allocation technique, load balancing, memory management, or other improved deployment mechanism. The limitation merely applies the abstract model-management process in the technological environment of machine-learning runtime deployment. See MPEP §§ 2106.05(d), 2106.05(f), and 2106.05(h). Regarding claim 15 Claim 15 – Step 1 – Is the claim to a process, machine, manufacture or composition of matter? Yes, the claim is to a process. Claim 15 – Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon? Yes, the claim recites an abstract idea. Claim 15 – Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application? No. There are no additional elements that integrate the judicial exception into a practical application. The additional elements: “wherein the machine-learning application utilizes the versioned model, and wherein deploying the versioned model further comprises the machine learning application executing instructions to transmit control system commands for one or more industrial devices.” – this limitation merely applies the abstract model-selection and deployment process in the technological environment of industrial-device control. The claim does not recite a particular industrial-control technique or improvement to computer functionality, but instead merely limits the use of the judicial exception to a particular field of use. See MPEP § 2106.05(h). Claim 15 – Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception? No. There are no additional elements that amount to significantly more than the judicial exception. The additional elements are: “wherein the machine-learning application utilizes the versioned model, and wherein deploying the versioned model further comprises the machine learning application executing instructions to transmit control system commands for one or more industrial devices.” – this limitation merely recites well-understood, routine, and conventional (WURC) transmission of control-system commands following the abstract model-selection and deployment process. The claim does not recite an unconventional control mechanism or improvement to industrial-device operation, and therefore does not provide an inventive concept beyond the implementation of the abstract idea in a particular technological environment. See MPEP §§ 2106.05(d), 2106.05(f), and 2106.05(h). Regarding claim 16 Claim 16 – Step 1 – Is the claim to a process, machine, manufacture or composition of matter? Yes, the claim is to a process. Claim 16 – Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon? Yes, the claim recites an abstract idea. “storing a plurality of model configuration records in a hierarchical structure of a model registry;” – this limitation recites storing and organizing model-configuration information in a hierarchy. Organizing information into records and arranging the records within a hierarchy is an information-management activity that may be practically performed in the human mind, or with pen and paper, by listing model configuration records and arranging them according to hierarchical relationships. See MPEP § 2106.04(a)(2)(III). “and retrieving, based on the model request, one or more model configuration records from the hierarchical structure of the model registry.” – this limitation recites reviewing a request and selecting or retrieving corresponding model configuration records form stored information. Comparing request information to stored records and identifying responsive records is an evaluation or judgment that may be practically performed in the human mind, or with pen and paper. See MPEP § 2106.04(a)(2)(III). Claim 16 – Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application? No. There are no additional elements that integrate the judicial exception into a practical application. The additional elements: “receiving a model request;” – this limitation merely gathers information used in the abstract retrieval of model configuration records. The claim does not recite an improved communication technique, an improved input mechanism, or a particular technical manner of receiving the model request. The limitation is insignificant data gathering. See MPEP § 2106.05(g) Claim 16 – Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception? No. There are no additional elements that amount to significantly more than the judicial exception. The additional elements are: “receiving a model request;” – this limitation recites generic receipt of information for later use in retrieving records. Receiving electronic information is well-understood, routine, and conventional (WURC) computer activity and constitutes insignificant data gathering. See MPEP §§ 2106.05(d) and 2106.05(g). Regarding claims 12-14 Claims 12-14 depend from claim 11 and add access-control criteria, model-training and catalogue information, and generative-task capabilities. These limitations do not materially alter the eligibility analysis of claim 11. Under Step 2A, Prong One, claims 12-14 recite the same abstract idea identified for claim 11: Claim 12 recites: “wherein the versioned model for each user session of the different users is based at least on the users access control privileges of each user session.” – this limitation recites evaluating a user’s access-control privileges and selecting or configuring a corresponding versioned model based on that evaluation. Comparing user information to access-control criteria and determining the appropriate model is an evaluation or judgment that may be practically performed in the human mind or with pen and paper. See MPEP § 2106.04(a)(2)(III). Claim 13 recites: “wherein the hierarchical repository comprises a catalogue of additional baseline models pretrained on datasets from different domains.” – this limitation recites organizing baseline-model information according to the domains of the datasets on which the models were trained. Organizing and categorizing information constitutes a mental process under MPEP § 2106.04(a)(2)(III). Under Step 2A, Prong Two, the additional elements do not integrate the abstract idea into a practical application: Claim 13 further recites: “wherein the additional model records associated with each additional baseline model is fine-tuned using local enterprise datasets.” – this limitation recites model fine-tuning only at a high level of generality. It does not recite a particular training algorithm, model architecture, parameter-update technique, or other mechanism that improves model operation or computer functionality. The limitation does not reflect a specific technological improvement under MPEP §§ 2106.04(d)(1) and 2106.05(a). Claim 14 recites: “the instantiating different versioned are capable of multiple generative tasks including conversational, summarizing, computational, predictive, visualization.” – under the reasonable interpretation that the limitation refers to the different versioned models, it merely identifies the intended functions or uses of those models. The claim does not recite a particular model architecture, algorithm, or technical mechanism for performing the listed tasks. The limitation therefore generally links the claimed system to generative-model applications without imposing a meaningful technological limit. See MPEP § 2106.05(h). Under Step 2B, the additional elements, individually and as an ordered combination, do not amount to significantly more than the judicial exception. Claim 12 adds only further abstract evaluation criteria and therefore does not add an element capable of supplying an inventive concept. Claims 13 and 14 recite generic model fine-tuning and model capabilities at a functional, result-oriented level. They do not require unconventional hardware, a non-generic model architecture, a particular training technique, or a non-conventional arrangement of components. The added limitations do not alter the conclusion for claim 11 that the claimed system merely implements the abstract model-selection and administration process using generic computer components. See MPEP §§ 2106.05(d), 2106.05(f), and 2106.05(h). Regarding claims 17-18 Claims 17-18 depend from claim 16 and add runtime model replacement and model-training limitations. These limitations do not materially alter the eligibility analysis of claim 16. Under Step 2A, Prong One, claims 17-18 recite the same abstract idea identified for claims 16: Claim 17 recites: “wherein one or more versioned models are selected” – this limitation recites selecting model information based on the model request and retrieved configuration records. Reviewing available models and selecting one or more models constitutes an evaluation or judgment that may be practically performed in the human mind or with pen and paper. See MPEP § 2106.04(a)(2)(III). Under Step 2A, Prong Two, the additional elements do not integrate the abstract idea into a practical application: Claim 17 further recites: “and replaced at run-time.” – this limitation generically applies the abstract model selection by replacing a model during runtime. The claim does not recite a particular model-swapping mechanism, memory-management technique, resource-allocation process, or runtime architecture. It therefore states the desired result of runtime replacement without reciting a specific improvement to computer or model operation. See MPEP §§ 2106.04(d)(1), 2106.05(a), and 2106.05(f). Claim 18 recites: “wherein each of the selected one or more models are pre-trained on customer-specific data subsequent to being trained on the domain-specific dataset.” – this limitation merely specifies the datasets and sequence used to train the selected models. It does not recite a particular training algorithm, model architecture, parameter-update technique, or technical improvement resulting from the training. The limitation therefore does not integrate the exception into a practical application. Considered as an ordered combination with claim 16, the added limitations merely select a model, generically replace the model at runtime, and identify the data used to train the model. They do not impose a meaningful technological limitation on the abstract storage, retrieval, and selection of model-configuration information. Under Step 2B, the additional elements, individually and as an ordered combination, do not amount to significantly more than the judicial exception. The runtime replacement and model-training limitations are recited at a high level of generality and do not require unconventional hardware, a particular runtime implementation, a non-generic model-training technique, or a non-conventional arrangement of computer components. Accordingly, the added limitations merely use generic computer functionality to implement the abstract model-selection and registry-management process and do not integrate an inventive concept. See MPEP §§ 2106.05(d) and 2106.05(f). Regarding claims 19, and 20 Claims 19, and 20 are not rejected under 35 U.S.C. § 101 because the additional limitations integrate the judicial exception into a practical application. Claims 19 and 20 recite compressing model parameters, deploying the compressed model to an edge device, and decompressing the model at runtime, with claim 20 further reciting quantization. These limitations apply the abstract model-management concepts in a specific technological context and impose meaningful technical limits on the claims. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-8, 11-14, and 16-18 are rejected under 35 U.S.C. 103 as being unpatentable over Chongyuan Xiang (US20240161018A1) in view of Weizhu Chen (US20220383126A1). Regarding claim 1, Xiang in view of Chen, teach a method comprising: “receiving a request associated with a machine learning application, wherein the request includes application information, user information, and execution information;” – Xiang teaches this limitation. Xiang teaches receiving a user request associated with a machine-learning prediction application, where the request includes model/application information and execution/promotion information. Xiang further teaches that the client applications may authenticate a user, thereby teaching or suggesting user information associated with the request: “An application processing interface service provides access to the version of the access module 121 that, in turn, provides access to the version of the machine learning model 123 that, in tum, generates one or more predictions responsive to a request.” (Xiang, pg. 2, ¶[0019]) “If the user decides to deploy the second version of the machine learning model, then the user may utilize the user interface to promote the second version of the machine learning model. For example, the user may request the machine learning model registry system to promote the second version of the machine learning model.” (Xiang, pg. 1, ¶[0017]) “the user may execute a ‘promote’ command. The promote command receives a machine learning model identifier that identifies the machine learning model 123 and a version of the machine learning model 123 that, in tum, is communicated to the machine learning model registry system 118” (Xiang, pg. 4, ¶[0035]) “At operation 352, the registry module 140, at the registry server machine 132, receives input information including the model name information 202 (e.g., ‘Home Valuation’ and the model version identifier 220 (e.g., ‘Version 2’) via a user interface, from the user. For example, the user may have executed the ‘promote’ command.” (Xiang, pg. 6, ¶[0047]) “one or more client applications 126 may be included in a given one of the client device 122, and configured to locally provide the user interface and at least some of the functionalities, with the client applications 126 configured to communicate with other entities in the networked system 120, on an as needed basis, for data and/or processing capabilities not locally available (e.g., access location information, access market information related to homes, to authenticate a user, to verify a method of payment, etc.).” (Xiang, pg. 2, ¶[0021]) “selecting, by one or more processing devices, ” – Xiang teaches this limitation in part. Xiang teaches selecting a machine-learning model version and an associated access-module version from a machine-learning model registry based on receiving request information: “Responsive to receiving the input information, the registry module 140 utilizes the input information to search the deployment information 144 in the machine learning model registry 125 to identify the access module 121.” (Xiang, pg. 6, ¶[0047]) Xiang further teaches: “the registry module 140 searches the deployment information 144 in the machine learning model registry 125 to match model version identifier 220 received from the user with the model version identifier 220 in an element of the deployment information 144. Responsive to identifying matching model version identifiers 220, the registry module 140 utilizes the associated access module version identifier 252 in the element of the deployment information 144 to identify the appropriate access module 121.” (Xiang, pg. 6, ¶[0047]) Xiang teaches that the model registry stores model-version information: “The model version information 206 may include one or more elements of the machine learning model 123. Each element of the machine learning model 123 corresponds to a version of the machine learning model 123 identified by the model identifier 200 and trained with a specific set of training data (e.g., model artifact).” (Xiang, pg. 4-5, ¶[0039]) Thus, Xiang teaches selecting, by one or more processing devices, model-version information from a model registry based on the request. “wherein the ” – Xiang teaches this limitation in part. Xiang teaches model metadata including model-version information, path information, hyperparameter information, deployability information, access-module information, and deployment information associating model versions with access-module versions for deployment: “At operation ‘6,’ registry module 140, at the registry server machine 132, registers (e.g., stores) the training data information and other meta data (e.g., model identifier, model name information, and the like) as model training information 138 in the machine learning model registry 125.” (Xiang, pg. 3, ¶[0028]) Xiang further teaches: “The machine learning model 123 may include the model version identifier 220, date information 222, feature information 224, training data information 226, path information 228, hyperparameter information 230, performance metrics information 232, and a deployable indicator 234.” (Xiang, pg. 5, ¶[0040]) Xiang further teaches: “The path information 228 stores a network path ( e.g., universal resource locater) to the machine learning model 123 stored in the distributed server machine 130. The hyperparameter information 230 stores the parameters utilized for configuring the version of the machine learning model 123.” (Xiang, pg. 5, ¶[0040]) Xiang further teaches access—module information: “The access module information 142 may include an access module version identifier 252 (e.g., Git sha), an access module name 254, date information 222, feature information 224, path information 228, and the deployable indicator.” (Xiang, pg. 5, ¶[0041]) Xiang further teaches deployment information associating access-module versions with model versions: “The deployment information 144 includes an access module version identifier 252 and one or more model version identifiers 220, as previously described. Accordingly, each element of deployment information 144 associates the identified version of the access module 121 with one or more identified versions of the machine learning models 123.” (Xiang, pg. 5, ¶[0042]) Xiang further teaches: “The deployment information 144 may be utilized to identify parings of a version of an access module 121 to one or more versions of machine learning models 123 that may be deployed to provide the predicting service on the API serving server machines 134.” (Xiang, pg. 5, ¶[0042]) Thus, Xiang teaches model metadata with parameters describing dependencies, access control, and deployment configuration. “assembling a versioned model of the baseline model using the one more child model records and associated dependencies;” – Xiang teaches this limitation in part. Xiang teaches identifying interoperable model/access-model versions using registry deployment information and deploying the identified components together: “The machine learning model registry system 1) responds to the request by: automatically identifying the first version of the access module interoperating with the second version of the machine learning model, and 2) automatically deploying the first version of the access module and the second version of the machine learning model to the server machines…” (Xiang, pg. 1, ¶[0017]) Xiang further teaches: “At operation ‘X,’ the registry module 140, at the registry server machine 132, responds to an execution of the ‘promote’ command by automatically identifying a version of the access module 121 that interoperates with the version of the machine learning model 123 that is being promoted.” (Xiang, pg. 4, ¶[0036]) Xiang further teaches: “At operation 354, the registry module 140, at the registry server machine 132, responds to an execution of the ‘promote’ command by automatically deploying the identified version of the access module 121 and the identified version of the machine learning model 123.” (Xiang, pg. 6, ¶[0048]) Thus, Xiang teaches assembling or pairing deployable model/service components based on associated dependencies, but does not expressly teach assembling a versioned model of a baseline model using child model records. “and deploying the versioned model in a configured run-time instantiation for use by the application based on the associated metadata.” – Xiang teaches this limitation. Xiang teaches deploying the identified version of the machine-learning model and associated access module to API service server machines to provide a prediction service based on deployment information: “At operation ‘Y,’ the registry module 140, at the registry server machine 132, responds to an execution of the ‘promote’ command by automatically deploying the machine learning model 123 identified with the promote command and the access module 121 identified via the deployment information 144 in the machine learning model registry 125 to the API server machines 134.” (Xiang, pg. 4, ¶[0037]) Xiang further teaches: “the registry module 140, at the registry server machine 132, deploys the appropriate machine learning model 123 and the appropriate access module 121 to the one or more API service server machines 134 to provide the prediction service.” (Xiang, pg. 4, ¶[0037]) Xiang further teaches: “Responsive to receiving the predictive service software, each of the API service server machines 134 installs the predictive service software and provides the predicting service based on the predictive service software including the ‘Version 01’ of the machine learning model 123 and the ‘Version 01’ of the access module 121.” (Xiang, pg. 5, ¶[0043]) Xiang further teaches: “At operation "Z," the one or more API service server machines 134 provide the prediction service based on the promoted version of the machine learning model 123 and the identified version of the access module 121 that were deployed together.” (Xiang, pg. 4, ¶[0038]) Thus, Xiang teaches deploying the versioned model in a configured runtime instantiation for use by the application based on the associated metadata. Xiang does not teach these limitations and/or portions of: “baseline model and one or more child model records” “wherein the baselines model and the one more child model records include model metadata” “assembling a versioned model of the baseline model using the one more child model records” Chen, however, teaches these remaining limitations and/or portions of: “baseline model and one or more child model records” – Chen teaches a general pretrained model corresponding to the claimed baseline model and LoRA modules or low-rank factorization matrices corresponding to the claimed child model records: “An improved system utilizes low rank adaptation (LoRA) for neural network-based models to adapt a general model for a specific task or domain. The weights in the general model are frozen, and small low-rank factorization matrices are injected into all or some weight matrices of the layers of the general model to form a specific model adapted to the specific task or domain.” (Chen, pg. 2, ¶[0016]) Chen further teaches: “A single pretrained model can be shared and used to build many small adaptations for different tasks.” (Chen, pg. 2, ¶[0019]) Chen further teaches: “Each task produces a single LoRA module, which usually occupies much less space than the pre-trained model.” (Chen, pg. 4, ¶[0038]) Chen further teaches: “During deployment, the service loads the pre-trained model into memory and store (potentially hundreds of) LoRA modules, each corresponding to a particular task, on stand by. A task can also be specialized to different customers and stores in different LoRA modules.” (Chen, pg. 4, ¶[0039]) Thus, Chen teaches a baseline model and one or more child model records. “wherein the baselines model and the one more child model records include model metadata” – Chen teaches that the child model records include parameter information in the form of adaption matrices treated as trainable parameters for adapting the baseline model: “Matrices A and B may be referred to as adaptation matrices, as they adapt the general model to the specific task or domain.” (Chen, pg. 2, ¶[0017]) Chen further teaches: “LoRA allows the training of each of multiple dense layers in the neural network indirectly by injecting and optimizing their rank decomposition matrices A and B, while keeping the original matrices of pretrained weights 110, unchanged.” (Chen, pg. 2, ¶[0018]) Chen further teaches: “During training, W is fixed and does not receive gradient updates, while A and B are treated as trainable parameters.” (Chen, pg. 3, ¶[0026]) Thus, Chen teaches child model records including parameter information corresponding to model metadata, and Xiang teaches that the model metadata describes dependencies, access control, and deployment configurations. “assembling a versioned model of the baseline model using the one more child model records” – Chen teaches assembling an adapted versioned model by injecting, adding, or combining low-rank factorization matrices with base-model weight matrices: “The weights in the general model are frozen, and small low-rank factorization matrices are injected into all or some weight matrices of the layers of the general model to form a specific model adapted to the specific task or domain.” (Chen, pg. 2, ¶[0016]) Chen further teaches: “For a pre-trained weight matrix W … the rank-deficiency constraint is achieved by representing the update matrices with their rank decomposition Δ W=AB” (Chen, pg. 3, ¶[0026]) Chen further teaches: “Both W and Δ W are multiplied to the same input, and their respective output vectors are summed coordinate-wise. For f(x)=Wx, our modified forward pass yields: f x = W x + Δ W x = W x + A B x ” (Chen, pg. 3, ¶[0026]) Chen further teaches: “First low-rank factorization matrices treated as trainable parameters are added to the base model weight matrices at operation 220 to form a first domain language model.” (Chen, pg. 3, ¶[0030]) Thus, Chen teaches assembling a versioned model of the baseline model using one or more child model records. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Xiang’s machine-learning model registry and deployment system to use Chen’s baseline pretrained model and LoRA/adaptation-module approach for representing, assembling, and deploying model versions. Xiang teaches a machine-learning model registry for storing model metadata, model versions, deployment information, and interoperable components for deployment to prediction-service servers. Chen teaches that a single pretrained model can be shared and used to build many small adaptations for different tasks, thereby making training more efficient, reducing serving costs, improving processor utilization, and enabling efficient switching among task- or customer-specific modules. The modification would have predictably improved Xiang’s model registry and deployment system by allowing Xiang’s registry to manage a shared baseline model and smaller task- or customer-specific child/adaptation records, thereby reducing storage and deployment overhead while preserving Xiang’s registry-based model-version selection and deployment workflow. Regarding claim 2, Xiang in view of Chen, teach the method of claim 1, wherein selecting comprises: “determining compatibility between the application information and execution information of the request with dependencies and deployment configurations from model metadata, and further determining access control of the model metadata and the user information of the request.” – Xiang teaches this limitation. Xiang teaches receiving model/application information and execution information in a promote command, and using the received information to search deployment information in the model registry to identify an access module that interoperates with the selected model version: “The promote command receives a machine learning model identifier that identifies the machine learning model 123 and a version of the machine learning model 123 that, in tum, is communicated to the machine learning model registry system 118” (Xiang, pg. 4, ¶[0035]) Xiang further teaches: “At operation 352, the registry module 140, at the registry server machine 132, receives input information including the model name information 202 (e.g., ‘Home Valuation’ and the model version identifier 220 (e.g., ‘Version 2’) via a user interface, from the user. For example, the user may have executed the ‘promote’ command.” (Xiang, pg. 6, ¶[0047]) Xiang teaches determining compatibility with dependencies and deployment configurations form model metadata by searching deployment information and identifying an access module associated with, and interoperable with, the requested model version: “Responsive to receiving the input information, the registry module 140 utilizes the input information to search the deployment information 144 in the machine learning model registry 125 to identify the access module 121.” (Xiang, pg. 6, ¶[0047]) Xiang further teaches: “the registry module 140 searches the deployment information 144 in the machine learning model registry 125 to match model version identifier 220 received from the user with the model version identifier 220 in an element of the deployment information 144. Responsive to identifying matching model version identifiers 220, the registry module 140 utilizes the associated access module version identifier 252 in the element of the deployment information 144 to identify the appropriate access module 121.” (Xiang, pg. 6, ¶[0047]) Xiang further teaches that deployment information associates an access-module version with one or more model versions that may be deployed together: “The deployment information 144 includes an access module version identifier 252 and one or more model version identifiers 220, as previously described. Accordingly, each element of deployment information 144 associates the identified version of the access module 121 with one or more identified versions of the machine learning models 123.” (Xiang, pg. 5, ¶[0042]) Xiang further teaches: “The deployment information 144 may be utilized to identify parings of a version of an access module 121 to one or more versions of machine learning models 123 that may be deployed to provide the predicting service on the API serving server machines 134.” (Xiang, pg. 5, ¶[0042]) Xiang further teaches determining compatibility by performing an interoperability test between the access module and model versions: “At operation ‘C,’ the deploy module 146 performs an interoperability test and registers the results in the corresponding element of the deployment information 144 in the machine learning model registry 125.” (Xiang, pg. 4, ¶[0033]) Xiang further teaches: “At operation ‘X,’ the registry module 140, at the registry server machine 132, responds to an execution of the ‘promote’ command by automatically identifying a version of the access module 121 that interoperates with the version of the machine learning model 123 that is being promoted.” (Xiang, pg. 4, ¶[0036]) Xiang further teaches determining access control using user information because Xiang teaches that client applications may authenticate a user and provide access to the model through an access module: “one or more client applications 126 may be included in a given one of the client device 122, and configured to locally provide the user interface and at least some of the functionalities … (e.g., access location information, access market information related to homes, to authenticate a user, to verify a method of payment, etc.).” (Xiang, pg. 2, ¶[0023]) Xiang further teaches: “An application processing interface service provides access to the version of the access module 121 that, in turn, provides access to the version of the machine learning model 123 that, in tum, generates one or more predictions responsive to a request.” Regarding claim 3, Xiang in view of Chen, teach the method of claim 1, wherein “the child model records comprise intermediate representations of the baseline model with changed parameters from a previous instantiation of the baseline model.” – Xiang does not teach this limitation. Chen, however, teaches this limitation. Chen teaches low-rank adaptation matrices, corresponding to the claimed child model records, that represent changed parameters for adapting a baseline pretrained model while keeping the baseline model weights unchanged: “An improved system utilizes low rank adaptation (LoRA) for neural network-based models to adapt a general model for a specific task or domain. The weights in the general model are frozen, and small low-rank factorization matrices are injected into all or some weight matrices of the layers of the general model to form a specific model adapted to the specific task or domain.” (Chen, pg. 2, ¶[0016]) Chen further teaches that the adaptation matrices are trained parameter changes relative to the pretrained baseline model: “LoRA allows the training of each of multiple dense layers in the neural network indirectly by injecting and optimizing their rank decomposition matrices A and B, while keeping the original matrices of pretrained weights 110, unchanged.” (Chen, pg. 2, ¶[0018]) Chen further teaches: “For a pre-trained weight matrix W … the rank-deficiency constraint is achieved by representing the update matrices with their rank decomposition Δ W=AB” (Chen, pg. 3, ¶[0026]) “During training, W is fixed and does not receive gradient updates, while A and B are treated as trainable parameters.” (Chen, pg. 3, ¶[0026]) Chen further teaches that the changed parameters may be combined with the baseline model during deployment: “During deployment, the original weight matrix W can be replaced with W+AB and used to perform inference as usual.” (Chen, pg. 3, ¶[0029]) Thus, Chen teaches child model records comprising intermediate representations of the baseline model with changed parameters from a previous instantiation of the baseline model. The motivation to combine Xiang and Chen remains the same as set forth with respect to claim 1. In particular, it would have been obvious to modify Xiang’s model-registry and deployment system to use Chen’s shared baseline pretrained model and smaller LoRA/adaptation records, thereby reducing storage and deployment overhead while preserving Xiang’s registry-based model-version selection and deployment workflow. Regarding claim 4, Xiang in view of Chen, teach the method of claim 1, wherein “the baseline model is pre-trained on a general domain dataset,” – Xiang does not teach this limitation. Chen, however, teaches this limitation. Chen teaches a pre-trained general model, corresponding to the claimed baseline model, that is adapted for specific tasks or domains: “An improved system utilizes low rank adaptation (LoRA) for neural network-based models to adapt a general model for a specific task or domain. The weights in the general model are frozen, and small low-rank factorization matrices are injected into all or some weight matrices of the layers of the general model to form a specific model adapted to the specific task or domain.” (Chen, pg. 2, ¶[0016]) Chen further teaches: “A single pretrained model can be shared and used to build many small adaptations for different tasks.” (Chen, pg. 2, ¶[0019]) Thus, Chen teaches a baseline model pre-trained on a general domain dataset. “and wherein the one or more child model records comprise intermediate representations with changed parameters of the baseline model trained on an enterprise specific dataset.” – Xiang does not teach this limitation. Chen, however, teaches this limitation. Chen teaches LoRA modules or low-rank factorization matrices, corresponding to the claimed child model records, that are trained parameters adapting the baseline pretrained model to a specific task, domain, or customer: “LoRA allows the training of each of multiple dense layers in the neural network indirectly by injecting and optimizing their rank decomposition matrices A and B, while keeping the original matrices of pretrained weights 110, unchanged.” (Chen, pg. 2, ¶[0018]) Chen further teaches: “For a pre-trained weight matrix W … the rank-deficiency constraint is achieved by representing the update matrices with their rank decomposition Δ W=AB” (Chen, pg. 3, ¶[0026]) Chen further teaches: “During training, W is fixed and does not receive gradient updates, while A and B are treated as trainable parameters.” (Chen, pg. 3, ¶[0026]) Chen further teaches that each task produces a LoRA module: “The service asks the user to define a task by providing a number of examples, which may be used directly or after data augmentation for training a LoRA module.” (Chen, pg. 4, ¶[0038]) Chen further teaches customer-specific specialization: “During deployment, the service loads the pre-trained model into memory and store (potentially hundreds of) LoRA modules, each corresponding to a particular task, on stand by. A task can also be specialized to different customers and stores in different LoRA modules.” (Chen, pg. 4, ¶[0039]) Thus, Chen teaches one or more model records comprising intermediate representations with changed parameters of the baseline model trained on task-, domain-, or customer-specific data, which teaches or suggests the claimed enterprise-specific dataset. It would have been obvious to modify Xiang’s model-registry and deployment system to use Chen’s shared baseline pretrained model and smaller task-, domain-, or customer-specific LoRA/adaptation records, thereby reducing storage and deployment overhead while preserving Xiang’s registry-based model-version selection deployment workflow. Regarding claim 5, Xiang in view of Chen, teach the method of claim 1, wherein “the deployment configurations determine a set of computing requirements for the run-time instance of the versioned model.” – Xiang does not teach this limitation. Chen, however, teaches this limitation. Chen teaches deployment of a pretrained model into memory with LoRA modules stored on standby, and further teaches that the deployment arrangement reduces servicing cost, improves processor utilization, and lowers hardware requirements: “During deployment, the service loads the pre-trained model into memory and store (potentially hundreds of) LoRA modules, each corresponding to a particular task, on stand by. A task can also be specialized to different customers and stores in different LoRA modules.” (Chen, pg. 4, ¶[0039]) Chen further teaches: “The shared original model may be kept in VRAM (volatile random access memory) or other selected memory while efficiently switching the significantly smaller LoRA model comprising stacked matrices A and B, greatly improving processor utilization.” (Chen, pg. 2, ¶[0019]) Chen further teaches: “Compared to conventional fine-tuning, this lowers the hardware barrier for training and significantly reduces the serving cost, without adding inference latency.” (Chen, pg. 2, ¶[0020]) Thus, Chen teaches that the deployed configuration of the baseline model and LoRA modules determines runtime computing requirements, including memory use, processor utilization, and serving-resource requirements. It would have been obvious to modify Xiang’s model-registry and deployment system to use Chen’s deployment arrangement for a shared baseline model and smaller LoRA/adaptation records, thereby reducing storage and deployment overhead while preserving Xiang’s registry-based model-version selection and deployment workflow. Regarding claim 6, Xiang in view of Chen, teach the method of claim 1, wherein “assembling the versioned model further comprises: pre-loading a set of model configurations comprising at least one or more of: model weights, adapter instructions.” – Xiang does not teach this limitation. Chen, however teaches this limitation. Chen teaches loading a pretrained model into memory during deployment and storing LoRA modules on standby, wherein the pretrained model includes model weights and the LoRA modules corresponding to adapted instructions or adapter configurations used to adapt the pretrained model: “During deployment, the service loads the pre-trained model into memory and store (potentially hundreds of) LoRA modules, each corresponding to a particular task, on stand by. A task can also be specialized to different customers and stores in different LoRA modules.” (Chen, pg. 4, ¶[0039]) Chen further teaches that the pretrained model includes weight matrices: “For a pre-trained weight matrix W … the rank-deficiency constraint is achieved by representing the update matrices with their rank decomposition Δ W=AB” (Chen, pg. 3, ¶[0026]) Chen further teaches that the LoRA adaptation matrices are injected into the general model: “The weights in the general model are frozen, and small low-rank factorization matrices are injected into all or some weight matrices of the layers of the general model to form a specific model adapted to the specific task or domain.” (Chen, pg. 2, ¶[0016]) Chen further teaches: “During deployment, the original weight matrix W can be replaced with W+AB and used to perform inference as usual.” (Chen, pg. 3, ¶[0029]) It would have been obvious to modify Xiang’s model-registry and deployment system to use Chen’s preloaded shared baseline model and smaller LoRA/adaptation modules, thereby reducing storage and deployment overhead while preserving Xiang’s registry-based model-version selection and deployment workflow. Regarding claim 7, Xiang in view of Chen, teach the method of claim 1, wherein “the hierarchical structure comprises a catalogue of different baseline models that are pre-trained with different domain specific datasets,” – Xiang does not teach this limitation. Chen, however, teaches this limitation. Chen teaches pre-trained general/domain models and teaches adapting such models to specific domains: “The dominant paradigm of deep learning consists of large-scale pre-training on general domain data and adaptation to particular tasks or domains.” (Chen, pg. 2, ¶[0015]) Chen further teaches: “An improved system utilizes low rank adaptation (LoRA) for neural network-based models to adapt a general model for a specific task or domain.“ (Chen, pg. 2, ¶[0016]) Chen further teaches: “In one example, the language model may be a transformer-based deep learning language model. The pre-trained weights 110 are in the form of a matrix having dimension of d x d resulting from the overall network being trained on general domain data.” Chen further teaches: “FIG. 2 is a flowchart illustrating computer implemented method 200 of adapting a base model to a domain specific task according to an example embodiment.” (Chen, pg. 3, ¶[0030]) Thus, Chen teaches baseline models pre-trained on general or domain-specific data and adapted for different domain-specific tasks. “and child model records associated with each different baseline model are generated based on an intermediate record.” – Xiang does not teach this limitation. Chen, however, teaches this limitation. Chen teaches generating low-rank factorization matrices associated with a base model, where the low-rank matrices represent intermediate/update matrices added to the base-model weight matrices to form a domain model. The low-rank matrices correspond to child model records generated from an intermediate update representation: “The weights in the general model are frozen, and small low-rank factorization matrices are injected into all or some weight matrices of the layers of the general model to form a specific model adapted to the specific task or domain.” (Chen, pg. 2, ¶[0016]) Chen further teaches: “Matrices A and B may be referred to as adaptation matrices, as they adapt the general model to the specific task or domain.” (Chen, pg. 2, ¶[0017]) Chen further teaches: “LoRA allows the training of each of multiple dense layers in the neural network indirectly by injecting and optimizing their rank decomposition matrices A and B, while keeping the original matrices of pretrained weights 110, unchanged.” (Chen, pg. 2, ¶[0018]) Chen further teaches: “For a pre-trained weight matrix W … the rank-deficiency constraint is achieved by representing the update matrices with their rank decomposition Δ W=AB” (Chen, pg. 3, ¶[0026]) Chen further teaches: “First low-rank factorization matrices treated as trainable parameters are added to the base model weight matrices at operation 220 to form a first domain language model.” (Chen, pg. 3, ¶[0030]) Chen further teaches: “The first domain language model is trained at operation 230 with first domain specific training data without modifying base model weight matrices.” (Chen, pg. 3, ¶[0031]) Chen further teaches a second domain model generated in a similar manner: “Second low-rank factorization matrices are added to the base model weight matrices at operation 320. The second low-rank factorization matrices were obtained in a manner similar first low-rank factorization matrices by training with second domain specific training data without modifying base model weight matrices.” (Chen, pg. 3, ¶[0033]) Thus, Chen teaches child model records associated with a baseline model generated from intermediate/update representations, namely low-rank factorization matrices representing update matrices for adapting the baseline model to different domain-specific models. It would have been obvious to modify Xiang’s machine-learning model registry to catalogue model versions using Chen’s shared baseline model and smaller domain-specific LoRA/adaptation records, thereby reducing storage and deployment overhead while preserving Xiang’s registry-based model-version selection and deployment workflow. Regarding claim 8, Xiang in view of Chen, teach the method of claim 1, further comprising: “receiving multiple requests received for one or more additional instances of the versioned model;” – Xiang teaches this limitation. Xiang teaches replaying prediction requests to a machine-learning model being validated and further teaches receiving a user command to promote a version of the machine-learning model: “In addition, the validation may include replaying past prediction requests and monitoring predictions of the machine learning model 123 being validated.” (Xiang, pg. 3, ¶[0027]) Xiang further teaches: “If the access module 121 is found, the cron job replays past prediction requests to the identified access module 121 that in turn, communicates the requests to the machine learning model 123 being validated that, in tum, generates the predictions.” (Xiang, pg. 3, ¶[0027]) Xiang further teaches: “the cronjob may submit the same requests to the respective machine learning models 123 and compare the distributions.” (Xiang, pg. 3, ¶[0027]) Xiang further teaches receiving a user request to promote a model version: “the registry server machine 128 promotes the second version of the machine learning model 123 responsive to receiving a ‘promote’ command from a user via the user interface to promote the second version of the machine learning model 123.” (Xiang, pg. 5, ¶[0046]) Thus, Xiang teaches receiving multiple requests for one or more additional model instances. “deploying multiple instances of the versioned model;” – Xiang teaches this limitation. Xiang teaches deploying an identified machine-learning model version and access-module version to multiple API service server machines: “Responsive to receiving the "promote" command, the registry module 140, at the registry server machine 132, automatically identifies the first version of the access module 121 as being interoperable with the second version of the machine learning model 123 and automatically deploys the first version of the access module and the second version of the machine learning model to multiple API service server machines 134.” (Xiang, pg. 5-6, ¶[0046]) Xiang further teaches: “At operation ‘Y,’ the registry module 140, at the registry server machine 132, responds to an execution of the ‘promote’ command by automatically deploying the machine learning model 123 identified with the promote command and the access module 121 identified via the deployment information 144 in the machine learning model registry 125 to the API server machines 134.” (Xiang, pg. 4, ¶[0037]) Xiang further teaches: “the registry module 140, at the registry server machine 132, deploys the appropriate machine learning model 123 and the appropriate access module 121 to the one or more API service server machines 134 to provide the prediction service.” (Xiang, pg. 4, ¶[0037]) Xiang further teaches deploying to three API service server machines: “the registry server machine 132 may identify three API service server machines 134 in the machine learning model registry 125 as currently providing a ‘Home Valuation’ (e.g., model name information 202) prediction service and communicate the ‘Version 01’ of the appropriate machine learning model 123 and ‘Version 01’ of the access module 121 (e.g., predictive service software) to each of the API service server machines 134” (Xiang, pg. 5, ¶[0043]) Thus, Xiang teaches deploying instances of the versioned model. “capturing changes to the versioned model as new model records with new model metadata in the hierarchical repository.” – Xiang teaches this limitation. Xiang teaches retraining or updating a machine-learning model to generate a second version, registering metadata and training-data information in the machine-learning model registry, and chronicling different versions or artifacts of the machine-learning model: “At operation ‘6,’ registry module 140, at the registry server machine 132, registers (e.g., stores) the training data information and other meta data (e.g., model identifier, model name information, and the like) as model training information 138 in the machine learning model registry 125.” (Xiang, pg. 3, ¶[0028]) Xiang further teaches: “Accordingly, the registry module 140 may be utilized to chronicle different versions (e.g., artifacts) of the machine learning model 123 in the distributed server machine 130 and the machine learning model registry 125.” (Xiang, pg. 3, ¶[0028]) Xiang further teaches generating a new version based on updated model/training/hyperparameter information: “At operation 304, the training server machine 128 trains (e.g., retrains) the machine learning model 123 to generate a second version of the machine learning model 123. For example, the training server machine 128 may receive an updated version of the machine learning model 123, updated training data information 226, or updated hyperparameter information 230 and so forth.” (Xiang, pg. 5, ¶[0044]) Xiang further teaches: “a cron job, executing at the training server machine 128, retrains the machine learning model 123 by utilizing the updated version of the machine learning model 123 and/or by utilizing the updated training data information 226 and/or by utilizing the updated hyperparameter information 230 to generate the second version (e.g., ‘Version 02’) of the machine learning model 123.” (Xiang, pg. 5, ¶[0044]) Xiang further teaches storing validation and interoperability results as metadata in the registry: “The scheduler module 135 ( e.g., cron job) communicates the results of the validation including the interoperability test and the prediction test to the registry server machine 132 that, in tum, stores the results (e.g., metadata) in the appropriate elements of machine training information 138 and deployment information 144 in the machine learning model registry 125.” (Xiang, pg. 3, ¶[0027]) Thus, Xiang teaches capturing changes to a versioned model as new model records with new model metadata in the machine-learning model registry. The motivation to combine Xiang and Chen remains the same as set forth with respect to claim 1. Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Xiang in view of Chen and further in view of Stefano Stefani (US11126927B2). Regarding claim 9, Xiang in view of Chen and further in view of Stefani, teach the method of claim 8, further comprising: “monitoring utilization of one or more additional model processing units for the multiple instances of the versioned model; and” – Xiang does not teach this limitation. Stefani, however, teaches this limitation. Stefani teaches monitoring utilization of model-processing resources, including CPU, GPU, and memory resources, for a fleet of model instances: “the auto-scaling system 106 includes an auto-scaling monitor 108 that can trigger an auto scaling---e.g., an addition and/or removal of model instances from a fleet 116 by an auto scaling engine 114-based on monitoring ( or obtaining) operational metric values 110 associated with operating conditions of the fleet.” (Stefani, col. 4, lines 50-55) Stefani further teaches: “a variety of operational metric values 110 are shown that can be monitored and potentially be used to determine whether the current fleet 116 of model instances 118A-118N serving a model 120 is over- or under-provisioned and thus, whether to add or remove capacity from the fleet.” (Stefani, col. 4, lines 61-66) Stefani further teaches utilization metrics including CPU, GPU, and memory utilization: “utilization metrics 320 (e.g., how busy is the central processing unit (CPU) or ‘CPU utilization’ 322, how busy is the graphics processing unit (GPU) or ‘GPU utilization’ 324, what is the current or recent memory usage 326, etc.) of the virtual machine and/or underlying host device.” (Stefani, col. 5, lines 9-13) Thus, Stefani teaches monitoring utilization of model-processing units, including CPU, GPU, and memory resources, for multiple machine-learning model instances. “executing one or more load-balancing operations to terminate execution of the one or more additional instances of the versioned model based on a threshold condition of the computing environment.” – Xiang does not teach this limitation. Stefani, however, teaches this limitation. Stefani teaches auto-scaling hosted machine-learning model instances by adding or removing model instances from a fleet based on monitored operational metrics. Stefani further teaches scale-down conditions based on CPU utilization thresholds and removing capacity by shutting down or terminating model instances: “the auto-scaling monitor 108 can determine whether to add or remove machines from the fleet, and send requests (e.g., API requests, function calls, etc.) to an auto-scaling engine 114 to perform scaling.” (Stefani, col. 5, lines 32-35) Stefani further teaches scale-down metric conditions based on utilization thresholds: “In this illustrated example, the user interface 402B also shows a ‘scale down’ condition where, if CPU utilization is less than three percent (3%) for 2 consecutive periods of time, one or more model instances 118A-118N are to be removed from the fleet” (Stefani, col. 6, lines 23-27) Stefani further teaches: “Removing capacity, in some embodiments, includes shutting down ( or otherwise terminating) one or more model instances.” (Stefani, col. 9, lines 28-30) Stefani further teaches that the amount of capacity removed may be based on a metric condition: “The amount of capacity to be added or removed may be determined based on an indicator of the customer specified metric condition, or could be based on a statically configured increment amount, etc.” (Stefani, col. 9, col. 30-33) Thus, Stefani teaches executing load-balancing or auto-scaling operations to terminate execution of one or more additional model instances based on a threshold condition of the computing environment. It would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention to modify the Xiang-Chen model-registry and deployment system to include Stefani’s auto-scaling operations for hosted machine-learning model instances. Xiang teaches deploying machine-learning model versions to multiple API service server machines to provide a prediction service, Chen teaches an efficient baseline-model and adaptation-module architecture for representing and deploying model versions, and Stefani teaches monitoring utilization and other operational metrics for a fleet of hosted machine-learning model instances and adding or removing, including terminating, model instances based on metric threshold conditions. The modification would have predictably improved the Xiang-Chen deployment system by allocating the multiple deployed runtime model instances to scale according to demand and resource utilization, thereby reducing wasted computing resources during low demand while maintaining prediction-service performance during workload changes. Claims 10 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Xiang in view of Chen and further in view of Henry P. Sepulveda (US20230304664). Regarding claim 10, Xiang in view of Chen and further in view of Sepulveda, teach the method of claim 1, wherein “deploying the versioned model further comprises the machine learning application executing instructions to ” – Xiang teaches this limitation in part. Xiang teaches deploying a machine-learning model and access module to API service server machines so that the deployed predictive-service software provides a machine-learning prediction service: “At operation ‘Y,’ the registry module 140, at the registry server machine 132, responds to an execution of the ‘promote’ command by automatically deploying the machine learning model 123 identified with the promote command and the access module 121 identified via the deployment information 144 in the machine learning model registry 125 to the API server machines 134.” (Xiang, pg. 4, ¶[0037]) Xiang further teaches: “the registry module 140, at the registry server machine 132, deploys the appropriate machine learning model 123 and the appropriate access module 121 to the one or more API service server machines 134 to provide the prediction service.” (Xiang, pg. 4, ¶[0037]) Xiang further teaches: “Responsive to receiving the predictive service software, each of the API service server machines 134 installs the predictive service software and provides the predicting service based on the predictive service software including the ‘Version 01’ of the machine learning model 123 and the ‘Version 01’ of the access module 121.” (Xiang, pg. 5, ¶[0043]) Thus, Xiang teaches deploying the versioned model for use by a machine-learning application. Xiang does not teach these limitations and/or portions of: “transmit control system commands for one or more industrial devices.” Sepulveda, however, teaches these remaining limitations and/or portions of: “transmit control system commands for one or more industrial devices.” – Sepulveda teaches machines including turbomachines, sensors, and an electronic control unit that collects operating parameters and controls the machines: “Each machine 120 may be a turbomachine, such as a gas turbine engine or gas compressor (e.g., in a pipeline). Alternatively, machine 120 may be another type of machine.” (Sepulveda, pg. 2, ¶[0015]) Sepulveda further teaches: “It is generally contemplated that machine 120 comprises an engine 122 that outputs emissions. Various operating parameters of engine 122 may be sensed by one or more sensors 124 within machine 120 and collected by an electronic control unit (ECU) 126. In addition, control parameters for controlling one or more subsystems of machine 120 may be supplied to or derived by (e.g., based on the collected operating parameters) ECU 126.” (Sepulveda, pg. 2, ¶[0015]) Sepulveda further teaches deploying the model as a microservice that may be called over a network: “Each statistical model may be a machine-learning model that accepts a feature vector of one or a plurality of time-correlated parameter values for a machine 120 as input and outputs the amount of predicted emissions, output by machine 120, given the parameter values in the feature vector.” (Sepulveda, pg. 6, ¶[0048]) Sepulveda further teaches deploying the model as a microservice that may be called over a network: “In an embodiment, interface 322 may enable deployment of a model as a microservice. In this case, subprocess 336 could call the deployed model directly, for example, over a network of platform 110 and/or from an external system via a microservice APL It should be understood that the call to the deployed model may include the dataset as input parameter(s).” (Sepulveda, pg. 7, ¶[0059]) Sepulveda further teaches using predicted emissions for downstream control of the machine: “Predicted emissions 350, which may comprise the predicted amount of emissions, classification of emissions, emissions score, or the like for a machine 120 and a confidence value for that prediction, may be used for one or more downstream functions 360. Downstream function(s) 360 may include, without limitation, compliance tracking, health monitoring, alerts, notifications, visualization of and/or interaction with predicted emissions 350 via a graphical user interface …, reports, control, and/or the like.” (Sepulveda, pg. 7, ¶[0061]) Sepulveda further teaches generating and transmitting control commands to an ECU of the machine: “In an embodiment, predicted emissions 350 may trigger control of machine 120. For example, if the logic of downstream function(s) 360 determines that predicted emissions 350 is indicative of non-compliant and/or abnormal operation (e.g., satisfies one or more criteria, such as exceeding a threshold), the logic may generate and transmit a control command to ECU 126 of machine 120 (e.g., over network( s) 140) to initiate a transition of machine 120 to a different engine state (e.g., a low-emissions mode, a shutdown state, an idle state, etc.).” (Sepulveda, pg. 8, ¶[0064]) Sepulveda further teaches: “This logic may continually and automatically generate and transmit control commands to ECU 126 of machine 120 ( e.g., over network(s) 140) to transition machine 120 between various engine states in accordance with a determined optimal operation.” (Sepulveda, pg. 8, ¶[0064]) Thus, Sepulveda teaches a machine-learning application executing instructions to transmit control system commands for one or more industrial devices, namely generating and transmitting control commands to an ECU of a turbomachine to transition the machine between engine states. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the Xiang-Chen model-registry and deployment system to use Sepulveda’s machine-learning -based industrial control functionality. Xiang teaches deploying machine-learning model versions to prediction-service servers for use by an application, Chen teaches an efficient baseline-model and adaptation-module architecture for representing and deploying model versions, and Sepulveda teaches using output from a deployed machine-learning model to generate and transmit control commands to an industrial machine. The modification would have predictably extended the Xiang-Chen deployed model service from providing prediction outputs to using those outputs for known downstream industrial control operations, as taught by Sepulveda. The modification would have enabled the deployed machine-learning application to automatically control industrial equipment based on model output while preserving Xiang’s registry-based model selection and deployment workflow. Regarding claims 11-15 Claims 11-15 recite system/apparatus limitations corresponding to the method limitations of claims 1, 2, 4, 7, 8, and 10. Claim 11 recites a system a system including memory and processors configured to perform substantially the same model-registry, versioned-model assembly, runtime deployment, and registry-update operations discussed above with respect to claims 1 and 8. Claims 12-15 further configure the system to perform substantially the same additional operations recited in claims 2, 4/7, and 10, respectively. Xiang and Chen teach the underlying machine-learning model registry, model metadata, deployment information, versioned-model selection, baseline-model/child-record assembly, and runtime deployment operations for the reasons discussed above with respect to claim 1. Xiang implements those operations using a registry server machine, API service server machines, registry modules, deployment modules, and predictive-service software. It would have been obvious to store instructions for performing the disclosed model-registry and deployment operations in memory and execute the instructions using one or more processors because Xiang’s disclosed registry/deployment operations are computer-implemented. Expressing the previously established method operations as processor-configured system functions does not, without more, impart a patentable distinction over the prior-art combination. Accordingly, claim 11 is obvious over Xiang in view of Chen for the same reasons discussed above with respect to claims 1 and 8. Claim 12 recites system limitations corresponding to the access-control/user-session determination operations of claim 2. Xiang teaches authenticating a user and providing access to a machine-learning model through an access module for the reasons discussed above with respect to claim 2. Accordingly, claim 12 is obvious over Xiang in view of Chen. Claim 13 recites system limitations corresponding to the baseline-model catalogue, domain-specific pretraining, and enterprise/customer-specific model-record limitations of claims 4 and 7. Chen teaches a pretrained general model, domain-specific adaptation, and task/customer-specific LoRA modules stored as separate adaption modules for the reasons discussed above with respect to claims 4 and 7. Accordingly, claim 13 is obvious over Xiang in view of Chen. Claim 14 recites system limitations corresponding to instantiating different versioned models capable of different task-specific operations. Chen teaches that a single pretrained model may be shared and used to build many small adaptations for different tasks, and Xiang teaches a predictive machine-learning service, for the reasons discussed above with respect to claims 1 and 7. Accordingly, claim 14 is obvious over Xiang in view of Chen. Claim 15 recites system limitations corresponding to the industrial-device control-command operations of claim 10. Xiang teaches deploying a versioned machine-learning model for use by a machine-learning application, and Sepulveda teaches generating and transmitting control commands to an ECU of an industrial machine based on machine-learning model output, for the reasons discussed above with respect to claim 10. Accordingly, claim 15 is obvious over Xiang in view of Chen and further in view of Sepulveda. Therefore: Claims 11-14 are rejected under 35 U.S.C. § 103 as being unpatentable over Xiang in view of Chen. Claim 15 is rejected under 35 U.S.C. § 103 as being unpatentable over Xiang in view of Chen, and further in view of Sepulveda. Regarding claims 16-18 Claims 16-18 recite method limitations corresponding to the model-registry storage, request-based retrieval, runtime replacement, and customer/domain-specific model-training operations discussed above with respect to claims 1, 4, 6, 7, and 13. Claim 16 recites storing model configuration records in a model registry, receiving a model request, and retrieving model configuration records form the model registry based on the request. Xiang teaches a machine-learning model registry storing model training information, model-version information, access-module information, and deployment information, and further teaches receiving model/version information and searching the deployment information in the registry to identify the appropriate access module and model version, for the reasons discussed above with respect to claim 1. Accordingly, claim 16 is obvious over Xiang. Claim 17 recites selecting and replacing one or more versioned models at runtime. Xiang teaches selecting and deploying a promoted machine-learning model version using registry deployment information, and Chen teaches runtime switching between task-specific models by swapping the LoRA module in use, for the reasons discussed above with respect to claims 1 and 6. Accordingly, claim 17 is obvious over Xiang in view of Chen. Claim 18 recites selected models pre-trained on customer-specific data subsequent to being trained on a domain-specific dataset. Xiang teaches model versions trained with specific sets of training data, and Chen teaches adapting a general model to a domain-specific task and further specializing tasks for different customers using different LoRA modules, for the reasons discussed above with respect to claims 4, 7, and 13. Accordingly, claim 18 is obvious over Xiang in view of Chen. Therefore: Claim 16 is rejected under 35 U.S.C. § 103 as being unpatentable over Xiang. Claims 17-18 are rejected under 35 U.S.C. § 103 as being unpatentable over Xiang in view of Chen. Claim 19 is rejected under 35 U.S.C. 103 as being unpatentable over Xiang in view of Joydeep Ray (US10546393B2) and further in view of Wei Wang (US20210125070A1). Regarding claim 19, Xiang in view of Ray and further in view of Wang, teach the method of claim 16, further comprising: “compressing at least a portion of the plurality of model parameters of the model, thereby generating a compressed model;” – Xiang does not teach this limitation. Ray, however, teaches this limitation. Ray teaches that trained machine-learning/deep-learning models may be compressed so that the models take up fewer resources, including memory and compute resources, and further teaches using lower-precision data types for model parameters: “When trained machine learning or deep learning models are deployed on, for example, edge devices (an edge device being a device providing an entry point into a system), it sometimes becomes necessary to compress (shrink) the machine learning or deep learning models such that the models take up less resources (such as memory and compute resources).” (Ray, col. 40, lines 9-15) Ray further teaches compressing model parameters using lower-precision data types: “One possible way to achieve compression is to use a lower precision data type for the model parameters.” (Ray, col. 40, lines 21-22) Ray further teaches: “Quantize the parameters to lower precision and retrain the model at lower precision using the pseudo-labeled dataset constructed in process (1 ).” (Ray, col. 40, lines 51-53) Ray additionally teaches generating a compressed model: “Compress the original model as required to generate a compressed model, such as by reducing the number of layers or making layers less wide in the original model.” (Ray, col. 41, lines 62-64) Thus, Ray teaches compressing at least a portion of model parameters to generate a compressed model. “deploying the compressed model to an edge device of an enterprise network;” – Xiang does not teach this limitation. Ray, however, teaches this limitation. Ray teaches deploying trained deep-learning models on edge devices: “In order to deploy trained deep learning models on edge devices such as smartphones, cameras and drones, it is sometimes necessary to compress the models to a more compact representation.” (Ray, col. 40, lines 16-19) Ray explains why such compression is used for edge deployment: “This is due to the lower amount of memory and compute capabilities available on such devices.” (Ray, col. 40, lines 19-20) Thus, Ray teaches deploying the compressed model to an edge device of a networked model-deployment system. Xiang and Ray do not teach these limitations and/or portions of: “decompressing the compressed model at run-time.” Wang, however, teaches these remaining limitations and/or portions: “decompressing the compressed model at run-time.” – Wang teaches compressing neural networks for storage/transmission and decompressing them by the computing device using the neural network: “Often times, in order to limit the size of the neural network for storage or transmission, the neural network may be compressed for storage and transmission, and decompressed by the computing device using the neural network.” (Wang, pg. 1, ¶[0004]) Wang further teaches forming a compressed neural network and transmitting it to a target system for decompression and use: “A compressed neural network is formed by encoding the data and the compressed representation of the neural network is transmitted to a target system for decompression and use.” (Wang, pg. 2, ¶[0036]) Wang additionally teaches decoding and dequantizing a compressed neural-network weight tensor: “generating a dequantized column swapped weight tensor by dequantizing the column swapped quantized weight tensor, by using the decoded codebook if the encoded codebook is received, or by using direct dequantization otherwise;” (Wang, pg. 16, ¶[0190]) Thus, Wang teaches decompressing the compressed model at run-time by decoding/dequantizing the compressed neural-network representation for use by the target computing device. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Xiang’s machine-learning model registry and deployment system to use Ray’s model-compression technique when deploying model versions to edge devices. Xiang teaches registry-based model-version storage, retrieval, and deployment, while Ray teaches that trained deep-learning models deployed on edge devices may be compressed because such devices have lower memory and compute capabilities. It further would have been obvious to use Wang’s decompression technique in the Xiang-Ray system so that the compressed neural-network representation could be decompressed and used by the target device at runtime. Wang teaches transmitting a compressed neural-network representation to a target device at runtime. Wang teaches transmitting a compressed neural-network representation to a target system for decompression and use. The modification would have predictably reduced memory, storage, and compute requirements for deploying Xiang’s model versions to resource-constrained edge devices while still allowing the deployed model to be used for inference after decompression. Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Xiang in view of Ray further in view of Wang and further in view of Hyeongseok Yu (US11531893B2). Regarding claim 20, Xian in view of Ray further in view of Wang and further in view of Yu, teach the method of claim 16, “wherein the compressing comprises a quantization of at least a portion of the plurality of model parameters,” – Xiang does not teach this limitation. Yu, however, teaches this limitation. Yu teaches determining quantization values by performing log quantization on neural-network parameters, including weight values: “A processor-implemented method includes determining a first quantization value by performing log quantization on a parameter from one of input activation values and weight values in a layer of a neural network,” (Yu, § Abstract) Yu further teaches: “quantizing the parameter to a value in which the first quantization value and the second quantization value are grouped.” (Yu, § Abstract) Yu’s Figure 5 further states: “QUANTIZE PARAMETER INTO VALUE IN WHICH FIRST QUANTIZATION VALUE AND SECOND QUANTIZATION VALUE ARE GROUPED” (Yu, FIG. 5) Yu also teaches that the parameter may be a weight value processed in a neural-network layer: “The parameter may include, but is not limited to, at least one of activation values processed in a layer of a neural network and weight values processed in the layer,” (Yu, col. 10, lines 30-32) Thus, Yu teaches quantization of neural-network model parameters, including weight parameters, into quantization model parameters. “and the decompressing comprises a dequantization of the plurality of quantized model parameters.” – Xiang does not teach this limitation. Yu, however, teaches this limitation. Yu teaches dequantizing a quantized neural-network parameter value and using the resulting dequantized value in a neural-network operation: “dequantizing the value in which the first quantization value and the second quantization value are grouped, and performing a convolution operation between a dequantization value obtained by dequantizing the value and the input activation values.” (Yu, col. 2, lines 30-34) Yu further teaches that the dequantizing includes calculating dequantization values for the first and second quantization values: “calculating each of a first dequantization value, which is a value obtained by dequantization of the first quantization value, and a second dequantization value, which is a value obtained by dequantization of the second quantization value, and obtaining the dequantization value by adding the first dequantization value and the second dequantization value.” (Yu, col. 2, lines 35-41) Yu also explains in the specification that when neural-network weight values are quantized, the processor may dequantize the grouped quantized value and perform a convolution operation using the dequantized value: “When the weight values processed in the layer of the neural network are quantized, the processor 1710 may dequantize a value, also referred to as a grouped value, in which a first quantization value and a second quantization value are grouped.” (Yu, col. 20, lines 40-44) Thus, Yu teaches that the decompression/reconstruction operation comprises dequantization of quantized neural-network model parameters. It would have been obvious to one of ordinary skill in the art before the effective filing data of the claimed invention to implement the Xiang-Ray-Wang compressed-model deployment system using Yu’s neural-network parameter quantization and dequantization techniques. Ray and Wang teach compressing neural-network models using quantized model weights/parameters, and Yu teaches a known technique for quantizing neural-network parameters and dequantizing the quantized values for use in neural-network operations. The modification would have predictably provided a concrete quantization/dequantization implementation for the compressed model of the Xiang-Ray-Wang combination. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to Paul Coleman whose telephone number is (571)272-4687. The examiner can normally be reached Mon-Fri. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, David Yi can be reached at (571) 270-7519. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /PAUL COLEMAN/ Examiner, Art Unit 2126 /DAVID YI/ Supervisory Patent Examiner, Art Unit 2126
Read full office action

Prosecution Timeline

Dec 16, 2023
Application Filed
Jul 15, 2026
Non-Final Rejection mailed — §101, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12688400
COMPUTATIONAL NEURAL NETWORK APPARATUS, CARD, METHOD, AND READABLE STORAGE MEDIUM
3y 6m to grant Granted Jul 21, 2026
Patent 12665745
MACHINE LEARNING/ARTIFICIAL INTELLIGENCE (ML/AI) SYSTEM WITH PROTECTED NEURAL NETWORKS
3y 5m to grant Granted Jun 23, 2026
Patent 12620453
METHOD, APPARATUS, AND COMPUTER PROGRAM FOR PREDICTING INTERACTION OF COMPOUND AND PROTEIN
3y 10m to grant Granted May 05, 2026
Patent 12614105
METHOD AND DEVICE FOR USE IN DATA PROCESSING, AND MEDIUM
4y 6m to grant Granted Apr 28, 2026
Patent 12597489
METHOD, DEVICE, AND COMPUTER PROGRAM FOR PREDICTING INTERACTION BETWEEN COMPOUND AND PROTEIN
2y 11m to grant Granted Apr 07, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
68%
Grant Probability
99%
With Interview (+42.9%)
3y 7m (~11m remaining)
Median Time to Grant
Low
PTA Risk
Based on 19 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month