Prosecution Insights
Last updated: October 02, 2026
Application No. 17/643,016

TRAINING ALGORITHMS FOR ONLINE MACHINE LEARNING

Final Rejection §101§103
Filed
Dec 07, 2021
Examiner
PAULA, CESAR B
Art Unit
2145
Tech Center
2100 — Computer Architecture & Software
Assignee
International Business Machines Corporation
OA Round
4 (Final)
34%
Grant Probability
At Risk
5-6
OA Rounds
0m
Est. Remaining
42%
With Interview

Examiner Intelligence

Grants only 34% of cases
34%
Career Allowance Rate
59 granted / 174 resolved
-21.1% vs TC avg
Moderate +8% lift
Without
With
+8.2%
Interview Lift
resolved cases with interview
Typical timeline
4y 6m
Avg Prosecution
10 currently pending
Career history
195
Total Applications
across all art units

Statute-Specific Performance

§101
12.7%
-27.3% vs TC avg
§103
51.3%
+11.3% vs TC avg
§102
18.2%
-21.8% vs TC avg
§112
14.0%
-26.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 174 resolved cases

Office Action

§101 §103
DETAILED ACTION This Final action is in response to the amendment filed on 12/17/2025. Claims 1-23 are pending and have been considered below. Claims 1, 7, 13 and 19 are independent claims. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 101 The rejections of claims 1-23 rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more, have been withdrawn as necessitated by the amendment. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-4, 7-10 and 13-16 are rejected under 35 U.S.C. 103 as being unpatentable over Han et al. (WO 2022/060777 A1) and Tao et al. (US 11,620,576 B1, hereinafter Tao). Regarding claim 1, Han teaches a computer-implemented method for training updated algorithms in an online machine learning environment, the method comprising: (Han, Fig. 5). identifying an updated algorithm related to a current algorithm, (Han, Abstract; “The training host of the non-RT RIC is configured to train initial models and update models based on ML offline learning data and other data,” wherein to “update models based on ML offline learning data” necessarily requires identifying an updated algorithm related to a current algorithm. Note that to update a model is equivalent to identifying an updated related to a current algorithm. This is reinforced by Applicant’s definition of the term model provided at paragraph [0002] of the specification of the claimed invention, “[a] machine learning model is the output generated by a machine learning algorithm trained with data.” As such, an update to a model is necessarily defined by the update to its algorithm.). (Han, Abstract; “The training host of the near-RT RIC is configured to send and receive AI/ML models to the model repository of the non-RT RIC. The training host of the near-RT RIC is configured to replace an AI/ML model being used by the inference host if the performance is below a threshold performance,” wherein to “replace an AI/ML model” when “performance is below a threshold performance” is equivalent to identifying an updated algorithm… the updated algorithm having revised terms with respect to the current algorithm. In other words, an entirely new “AI/ML model” necessarily comprises the updated algorithm having revised terms. Han, [0032]; “ln these implementations, the non-RT RIC 212 may provide discovery mechanism if a particular ML model can be executed in a target ML inference host (MF), and what number and type of ML models can be executed in the target ML inference host… The non-RT RIC 212 may also include and/or operate one or more ML engines, which are packaged software executable libraries that provide methods, routines, data types, etc., used to nm ML models. The non-RT RIC 212 may also implement policies to switch and activate ML model instances under different operating conditions,” deploying the pre-trained updated model concurrently with the current model in the artificial intelligence system (Han, Fig. 5; The “Trained Model 536” corresponding to the pre-trained updated model is deploy[ed] in the “Near-RT RIC (Near Real-Time RAN Intelligent Controller)” as the “Trained Running Model 542.” Han, [0048]; “The training host near-RT 514 associated with or in the near-RT RIC 510 receives a model download notification from the model repository 508, and it receives the model, e.g., train model 536 or updated model 540, from the model repository 508.” As such, the pre-trained updated model concurrently with the current model in the artificial intelligence system, or “Running Model 538” also deployed in the “Near-RT RIC.” That theses running models are built with machine learning algorithms indicates that the “Near-RT RIC” is an artificial intelligence system.) transferring reinforcement learning learned in the online machine learning environment by the current model to the updated model (Han, [0040]; “The performance feedback 530 is training data for online training, e.g., rewards, environment states, performed actions, and so forth, and data for performance monitoring. The training host (online learning) ("training host near-RT") 514 is configured for online learning.” Han, [0050]; “The training host near-RT 514 in the near-RT RIC 510 receives AI/ML performance feedback 530 from the inference host 512,” wherein to “receive AI/ML performance feedback 530” produced by reinforcement learning (“e.g., rewards, environment states, performed actions”) “from the inference host 512” housing the “Running Model 538,” or current model, at “[t]he training host near-RT 514” housing the “Trained Running Model 542,” or updated model, is equivalent to training the updated model by transferring reinforcement learning learned in the online machine learning environment corresponding to “Inference Host 512” by the current model to the updated model.). Han does not explicitly teach pre-training an updated model including the updated algorithm, the pre-training performed offline with training data generated by a current model based on the current algorithm, the training data being in a form of input/output pairs of the current model for training new models, the current model being deployed in an online machine learning environment of an artificial intelligence system in which the current model learns while deployed via reinforcement learning. However, Tao discloses teacher models 115 selected to transfer knowledge to student models. The student models download the teacher model data to train the student model locally (col.4, lines 51-63)-- pre-training an updated model including the updated algorithm, the pre-training performed offline with training data generated by a current model based on the current algorithm. Tao, in the area of knowledge transfer via teacher-student reinforcement learning, teaches (Tao, Fig. 4; Tao, Col. 5, lines 24-27; “In some embodiments, instance transfer may transfer instances-samplings of the inputs and/or outputs of a teacher model-and then re-use the instances (or samples) to train a student model,” wherein “inputs and/or outputs…[used] to train a student model” encompasses the training data being in a form of input/output pairs of the current model for training new models. Tao, Col. 13, lines 27-34; “FIG. 4 shows another example process to train a student model with knowledge transfer, according to some embodiments. In this example, process 400 may commence with determining a loss representing a difference between a policy of a student model and a policy of a teacher model based at least in part on a first set of one or more trajectories obtained with the policy of the student model, according to some embodiments (block 405).”. “the difference between a policy” corresponds to the training data being in a form of input/output pairs of the current model for training new models. Tao teaches using reinforcement learning the network accessible teacher models to perform a number of tasks(fig.1, col.4, lines 37-60)-- the current model being deployed in an online machine learning environment of an artificial intelligence system in which the current model learns while deployed via reinforcement learning. Tao is analogous to the claimed invention as both are from the same field of endeavor, that is, knowledge transfer using a teacher-student paradigm. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the offline to online training and deployment scheme of Han with the method of teacher-student knowledge transfer of Tao. The motivation to do so is to allow the updated model to learn from the experiences of the current model in performing the same or similar functions (Tao, Cols 2-3, lines 61-4; “When the decision-making task(s) of the teacher models share common feature(s) with the decision-making task(s) that the student model is going to solve, it may be possible to improve the learning of the student model by leveraging knowledge acquired by those trained teacher models.”). Regarding claim 2, the combination of Han and Tao teaches the method of claim 1, wherein (and thus the rejection of claim 1 is incorporated). Han does not explicitly teach transferring reinforcement learning of the current model includes: storing a set of input/output pairs for the current model; associating rewards with each of the input/output pairs; selecting a sub-set of input/output pairs according to significance of the associated rewards; and performing supervised learning of the updated model using the selected sub-set of input/output pairs. However Tao, in the area of knowledge transfer via reinforcement learning, teaches these limitations. transferring learning of the current model includes: storing a set of input/output pairs for the current model; (Tao, Col. 5, lines 24-27; “In some embodiments, instance transfer may transfer instances-samplings of the inputs and/or outputs of a teacher model and then re-use the instances (or samples) to train a student model.” Here, “the inputs and/or outputs” are taken from “teacher model,” or current model, “to train “a student model,” which denotes the updated algorithm of the claimed invention.) associating rewards with each of the input/output pairs; (Tao, Col. 5, lines 27-32 “Again, in the exemplary context of reinforcement learning, training system 110 may obtain one or more trajectories (e.g., sequences of states, actions and rewards) of teacher models 115, and use the trajectory samples as instances to facilitate the training of student model 120,” wherein “states” and “actions” are input/output pairs with associating rewards ) selecting a sub-set of input/output pairs according to significance of the associated rewards (Tao, Col. 9, lines 50-55; “The instance transfer may involve training the student model directly with instances, e.g., samplings of the inputs and/or outputs of a teacher model. For instance, the instances may include sampled trajectories (e.g., sequences of states, actions and rewards) following the policy of the teacher model.” Here, “sampled trajectories” encompasses a sub-set of states and actions, or input/output pairs. Tao, Col. 10, lines 9-18; “reinforcement learning model 200 may maintain a buffer of policy parameters and/or corresponding trajectory output (“experience”), with which reinforcement learning model 200 has previously been trained. In some embodiments, the experience may be prioritized. For instance, only experience with a policy loss (e.g., LRL or Lelip) beyond a certain level may be stored in the buffer.” This process corresponds to selecting a sub-set of input/output pairs according to significance of associated rewards. Tao, Col. 9, lines 21-25; “reducing LRL, may cause the policy of the student model to mimic the policy of the teacher model as well as increase the total expected rewards following the updated student model,” thereby indicating that “policy loss” is intrinsically linked with reward. Therefore, the “priority,” or significance is in reference to the associated rewards for each sub-set of “experiences,” or input/output pairs.) and performing supervised learning of the updated model using the selected sub-set of input/output pairs (Tao, Col. 9, lines 50-55; “The instance transfer may involve training the student model directly with instances, e.g., samplings of the inputs and/or outputs of a teacher model. For instance, the instances may include sampled trajectories (e.g., sequences of states, actions and rewards) following the policy of the teacher model,” wherein the state and action corresponding to each input/output pair constitute labels thus indicating that the samples, or sub-set[s], are used for “training the student model” or updated model. Because the training data is labeled, the updated model was trained using supervised learning. Tao, Col 10, lines 17-19; “The prioritized experience in the buffer may be re-used to train reinforcement learning model 200,” wherein the “reinforcement learning model 200,” corresponds to the updated model.). Tao is analogous to the claimed invention as both are from the same field of endeavor, that is, model updates using teacher-student paradigms. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the offline to online training and deployment scheme of Han with the prioritized experience replay and supervised learning of Tao. The motivation to do so is to allow the updated model to learn from the successful actions of the current model (Tao, Col. 10, lines 20-23; “The repeated training with the prioritized experience may strengthen the memory of reinforcement learning model 200 as to what policy shall be avoided or taken.”). Regarding claim 3, the combination of Han and Tao teaches the method of claim 2, wherein (and thus the rejection of claim 2 is incorporated). Han does not explicitly teach the input/output pairs are state and action pairs, the state used as input and the action used as output and label. However, Tao, in the area of knowledge transfer via reinforcement learning, teaches this limitation (Tao, Col. 9, lines 50-55; “The instance transfer may involve training the student model directly with instances, e.g., samplings of the inputs and/or outputs of a teacher model. For instance, the instances may include sampled trajectories (e.g., sequences of states, actions and rewards) following the policy of the teacher model,” wherein the state is the input, the action is the output and the “sampled trajectories” correspond to the sub-set of state and action, or input/output pairs the “teacher” has accumulated. Here, the state and action corresponding to each input/output pair constitute labels). Tao is analogous to the claimed invention as both are from the same field of endeavor, that is, model updates using teacher-student paradigms. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the offline to online training and deployment scheme of Han with the prioritized experience replay and supervised learning using labeled input/output pairs of Tao. The motivation to do so is inherent as the supervised learning process inherited from claim 2, by definition, requires labeled training data. Regarding claim 4, the combination of Han and Tao teaches the method of claim 1, further comprising (and thus the rejection of claim 1 is incorporated). Han further teaches determining an update to the current algorithm is available as the updated algorithm; and (Han, [0067]; “For example, the training host 514 and/or inference host 512 may detect performance degradation that is above a threshold value and send a request for a different model. The training host near-RT 514 may use the trained running model 542,” wherein “detect[ing] a performance degradation” that can be ameliorated by “the trained running model 542” is equivalent to determining an update to the current algorithm is available as the updated algorithm.) retrieving the training data used to train the current model (Han, [0042]; “The training host non-RT 506 collects ML learning data 516 from the E2 nodes O-CU/O-DU 502 ("E2 nodes") over the O1 interface for offline reinforcement learning.” As outlined in the rejection of claim 1, this “learning data 516” is used to train each model in the repository including the model acting as the current model.). Regarding claim 7, Han teaches a computer program product comprising a computer-readable storage medium having a set of instructions stored therein which, when executed by a processor, causes the processor to train updated algorithms in an online machine learning environment by: (Han, [0017]; “The O-Cloud 206 is a cloud computing platform including a collection of physical infrastructure nodes to host the relevant O-RAN functions (e.g., the nearRT RIC 214, O-RAN Central Unit-Control Plane (O-CU-CP) 221, O-RAN 1 Central Unit-User Plane O-CU-UP 222, and the O-RAN Distributed Unit (O-DU) 215, supporting software components (e.g., OSs, VMMs, container runtime engines, ML engines, etc.), and appropriate management and orchestration functions,” wherein the “nearRT RIC 214” train[s] updated algorithms in an online machine learning environment.). The following limitations of claim 7 correspond to the steps of claim 1, and are rejected for the same reason as claim 1. Claim 8 is a computer program product claim corresponding to the steps of claim 2, and is rejected for the same reason as claim 2. Claim 9 is a computer program product claim corresponding to the steps of claim 3, and is rejected for the same reason as claim 3. Claim 10 is a computer program product claim corresponding to the steps of claim 4, and is rejected for the same reason as claim 4. Regarding claim 13, Han teaches a computer system for training updated algorithms in an online machine learning environment, the computer system comprising: a processor set; and a computer readable storage medium; wherein: the processor set is structured, located, connected, and/or programmed to run program instructions stored on the computer readable storage medium; and the program instructions which, when executed by the processor set, cause the processor set to train updated algorithms in an online machine learning environment by: (Han, [0017]; “The O-Cloud 206 is a cloud computing platform including a collection of physical infrastructure nodes to host the relevant O-RAN functions (e.g., the nearRT RIC 214, O-RAN Central Unit-Control Plane (O-CU-CP) 221, O-RAN 1 Central Unit-User Plane O-CU-UP 222, and the O-RAN Distributed Unit (O-DU) 215, supporting software components (e.g., OSs, VMMs, container runtime engines, ML engines, etc.), and appropriate management and orchestration functions,” wherein the “nearRT RIC 214” train[s] updated algorithms in an online machine learning environment.) The following limitations of claim 7 correspond to the steps of claim 1, and are rejected for the same reason as claim 1. The following limitations of claim 13 correspond to the steps of claim 1, and are rejected for the same reason as claim 1. Claim 14 is a system claim corresponding to the steps of claim 2, and is rejected for the same reason as claim 2. Claim 15 is a system claim corresponding to the steps of claim 3, and is rejected for the same reason as claim 3. Claim 16 is a system claim corresponding to the steps of claim 4, and is rejected for the same reason as claim 4. Claims 5, 11 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Han, Tao and Zhu et al. (“Learning by Reusing Previous Advice in Teacher-Student Paradigm,” hereinafter Zhu). Regarding claim 5, the combination of Han and Tao teach the method of claim 1, wherein (and thus the rejection of claim 1 is incorporated). Han does not explicitly teach choosing between a first output of the updated model and a second output of the current model according to a variable probability scheme where the current model output is initially favored over the updated model output with decreasing favor over time. However, Zhu, in the area of teacher-student reinforcement learning, teaches this limitation (Zhu, Algorithm 3; 4.3 Decay Reusing Probability; “Every time agent 𝑖 visits advised states, a more flexible method is to give it the opportunity to choose between reusing previous advice, learning by itself and asking for advice. We name this method as Decay Reusing Probability (Decay).” Here, “agent i” denotes the updated model that is choosing between, “reusing previous advice” (output of the current model) and its own output (“learning by itself,” or output of the updated model). “In our work, 𝑃𝑟𝑒𝑢𝑠𝑒 determines whether learning agent 𝑖 should reuse a teacher’s advice to guide its action selection,” i.e. choose output of the current model, “agent 𝑖 will follow previous advice with probability 𝑃𝑟𝑒𝑢𝑠𝑒…As agent 𝑖 repeatedly performs the latest advice in state 𝑠, 𝑃𝑟𝑒𝑢𝑠𝑒(𝑠) decays exponentially.” Here, "𝑃𝑟𝑒𝑢𝑠𝑒"denotes a variable probability scheme that “decays,” thereby indicating that the current model is initially favored but with decreasing favor over time. Thus, the “agent i,” or updated model, will choose “learning by itself,” or its own output, more often over time.). Zhu is analogous to the claimed invention as both are from the same field of endeavor, that is, methods of knowledge transfer using a teacher-student paradigm. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine offline to online training and deployment scheme of Han with the Decay Reusing Probability method of Zhu. The motivation to do so is to make the updated model robust to sub-optimal teachers while allowing it to learn from past experiences (Zhu, Introduction; “As the advice from teacher may be outdated during learning, we also propose Decay Reusing Probability (Decay) to allow agents learning in usual advising framework while reusing previous advices with decaying probabilities.”). Claim 11 is a computer program product claim corresponding to the steps of claim 5, and is rejected for the same reason as claim 5. Claim 17 is a system claim corresponding to the steps of claim 5, and is rejected for the same reason as claim 5. Claims 6, 12 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Han, Tao, Zhu and Zhan et al. (“Theoretically-grounded policy advice from multiple teachers in reinforcement learning settings with applications to negative transfer,” hereinafter Zhan) Regarding claim 6, the combination of Han, Tao and Zhu teach the method of claim 5, wherein (and thus the rejection of claim 5 is incorporated). Han does not explicitly teach choosing between a first output of the updated model and a second output of the current model is further according to a variable bonus reward scheme where the updated model output matching the current model output is rewarded with decreasing reward over time. However, Zhan, in the area of teacher-student reinforcement learning, teaches this limitation (Zhan, Algorithm 3; 4.2 Multi-Teacher Advice Algorithm, “the MDP can be effectively exploited for successful learning. Unfortunately, such a process is not well modeled using current methods. Here, we remedy this problem by introducing an algorithm which follows the teacher’s advice at the very beginning and then switches to a policy computed by an algorithm operating within the MDP. That is, the teacher guides the student at the beginning of the learning process and as the student gathers more experience, the teacher’s influence diminishes over time,” wherein “the teacher” denotes the current model and “the student” denotes the updated model. To “guide the student” in this context corresponds to matching the output of the student and the teacher. “To leverage both the teacher’s and learned policies, we set a mixed policy of the form 𝜋𝑖+1=𝛽𝑖𝜋𝜏+(1−𝛽𝑖)𝜋̂𝑖, for 0≤𝛽𝑖≤1 to guide the student’s dataset collection while allowing the teacher to fractionally control exploration needed to collect data at the next iteration. β should typically be set so as to decay exponentially over time.” Here, “β” denotes a variable bonus reward that at high values favors “the teacher,” or current model output, initially, but “decreases exponentially” over time to favor “the student,” or updated model output.). Zhan is analogous to the claimed invention as both are from the same field of endeavor, that is, methods of knowledge transfer using a teacher-student paradigm. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the offline to online training and deployment scheme using a decaying experience recall, as taught by the combination of Han, Tao and Zhu, with the mixed student-teacher guiding policy of Zhan. The motivation to do so is to allow the updated model to eventually outperform its sub-optimal teacher (Zhan, 4.2 Multi-Teacher Advice Algorithm; “This decreases the student’s reliance on the teacher and allows it to exploit the knowledge gathered from the MDP to learn better behaving policies than that of the teacher.”). Claim 12 is a computer program product claim corresponding to the steps of claim 6, and is rejected for the same reason as claim 6. Claim 18 is a system claim corresponding to the steps of claim 6, and is rejected for the same reason as claim 6. Claims 19-21 and 23 are rejected under 35 U.S.C. 103 as being unpatentable over Han, in view of Tao, and further in view of Li et al. (US 2019/0287515 A1, hereinafter Li). Regarding claim 19, Han teaches a computer-implemented method for facilitating deploying a new algorithm while concurrently executing a deployed algorithm utilizing a reinforcement learning algorithm comprising: (Han, Title; “Online Reinforcement Learning.” Han, Abstract; “The training host of the near-RT RIC is configured to send and receive AI/ML models to the model repository of the non-RT RIC. The training host of the near-RT RIC is configured to replace an AI/ML model being used by the inference host if the performance is below a threshold performance,” wherein “replac[ing] an AI/ML model being used by the inference host” with a different “AI/ML model” is equivalent to facilitating deploying a new algorithm while concurrently executing a deployed algorithm.). receiving training data used to train a current model comprising the deployed algorithm; performing offline supervised pre-training of a new model comprising the new algorithm utilizing the received training data (Han, “The training host non-RT 506 collects ML learning data 516 from the E2 nodes O-CU/O-DU 502 ("E2 nodes") over the O1 interface for offline reinforcement learning. The training host non-RT 506 trains the initial model based on the offline training. The training host non-RT 506 transfers via move model 518 the initial model 534, which is an offline trained model, to the model 20 repository as a trained model 536,” wherein “learning data 516” corresponds to the received training data and “the initial model 534” corresponds to a new model comprising the new algorithm, which will later become a current model comprising the deployed algorithm. Han, [0032]; “For supervised learning, and the ML training host and/or ML inference host/actor can be part of the non-RT RlC 212 and/or the near-RT RIC 214,” thereby specifying that the offline pre-training is supervised.). the new algorithm including an update to the deployed algorithm, the update including revised terms with respect to the deployed algorithm, (Han, Abstract; “The training host of the near-RT RIC is configured to send and receive AI/ML models to the model repository of the non-RT RIC. The training host of the near-RT RIC is configured to replace an AI/ML model being used by the inference host if the performance is below a threshold performance,” wherein to “replace an AI/ML model” indicates that the new algorithm include[es] an update to the deployed algorithm, the update including revised terms with reference respect to the deployed algorithm. In other words, Han teaches [0032]; “ln these implementations, the non-RT RIC 212 may provide discovery mechanism if a particular ML model can be executed in a target ML inference host (MF), and what number and type of ML models can be executed in the target ML inference host… The non-RT RIC 212 may also include and/or operate one or more ML engines, which are packaged software executable libraries that provide methods, routines, data types, etc., used to nm ML models. The non-RT RIC 212 may also implement policies to switch and activate ML model instances under different operating conditions,” wherein different “type[s] of ML models” or “ML model instances” built from a variety of “libraries” necessarily differ in architecture.) the current model having learned via reinforcement learning while deployed; (Han, Title; “Online Reinforcement Learning.” Han, Fig. 5; Han, [0037]; “Examples disclose a deployment scenario for online reinforcement learning in the Near-RT RIC, which incorporates an online training host and inference host in the Near-RT RIC,” wherein “online reinforcement learning” is equivalent to reinforcement learning while deployed.) deploying the pre-trained new model concurrently with the deployed model in the artificial intelligence system (Han, Fig. 5; The “Trained Model 536” corresponding to the pre-trained new algorithm is deployed in the “Near-RT RIC (Near Real-Time RAN Intelligent Controller)” as the “Trained Running Model 542.” Han, [0048]; “The training host near-RT 514 associated with or in the near-RT RIC 510 receives a model download notification from the model repository 508, and it receives the model, e.g., train model 536 or updated model 540, from the model repository 508.” As such, the pre-trained new model and the deployed model are deployed concurrently, in the “Near-RT RIC.”) Han does not explicitly teach receiving training data generated by a current model comprising the deployed algorithm, the training data being in a form of input/output pairs of the current model for training new models, the current model being deployed in an online machine learning environment of an artificial intelligence system. However, Tao discloses teacher models 115 selected to transfer knowledge to student models. The student models download the teacher model data to train the student model locally (col.4, lines 51-63)-- pre-training an updated model including the updated algorithm, the pre-training performed offline with training data generated by a current model based on the current algorithm. Tao is analogous to the claimed invention as both are from the same field of endeavor, that is, knowledge transfer using a teacher-student paradigm. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the offline to online training and deployment scheme of Han with the method of teacher-student knowledge transfer of Tao. The motivation to do so is to allow the updated model to learn from the experiences of the current model in performing the same or similar functions (Tao, Cols 2-3, lines 61-4; “When the decision-making task(s) of the teacher models share common feature(s) with the decision-making task(s) that the student model is going to solve, it may be possible to improve the learning of the student model by leveraging knowledge acquired by those trained teacher models.”). Han does not explicitly teach transferring reinforcement learning iteratively from the deployed model to the pre-trained new model until at least on acceptance criterion is met to form an accepted new model. However, Li, in the area of adversarial teacher-student learning, teaches this limitation (Li, Fig. 4; [0068]; “FIG. 4 is flowchart of a method 400 for student teacher training,” a step-by-step process that encompasses transferring learning iteratively [0072]; “At operation 410, a check is made to determine if the behavior of the student model 208 converges with the behavior of the teacher model 204…A divergence score converging below a convergence threshold indicates that the student model 208 is able to recognize speech in its given domain almost as well as the teacher model 204 is able to recognize speech in its domain,” wherein “a convergence threshold” denotes an acceptance criteria, “the student model” corresponds to the pre-trained new model, and “the teacher model” corresponds to the deployed model. Li, [0075]; “In response to determining that the student model 208 has converged relative to the teacher model 204, method 400 proceeds to operation 412, where the student model 208 is finalized,” indicating that at least one acceptance criteria [has been] met to form an accepted new model.) Li is analogous to the claimed invention as both are from the same field of endeavor, that is, model updates using teacher-student paradigms. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the offline to online training and deployment scheme of Han with the convergence threshold of Li. The motivation to do so is to have a metric for comparing the accuracy of the new algorithm and the deployed algorithm that can determine requisite optimality (Li, [0073]; “the student model 208 may be more or less accurate than the teacher model 204 in some cases for accurately recognizing speech, but the student model 208 is judged based on the similarity of its results to the results of the teacher model 204.”). deploying the accepted new model (Han, [0057]; “For example, the AI/ML model, e.g., running model 538, trained running model 542, trained model 536, or updated model 540, is deployed to the inference host 512, e.g., xApp, in the near-RT rue 510.”) Regarding claim 20, the combination of Han and Li teaches the computer-implemented method of claim 19, wherein (and thus the rejection of claim 19 is incorporated). Han does not explicitly teach the transferring learning further comprises: storing input and output pairs P (P1, P2, ..., Pn) for the deployed model; associating rewards R (R1, R2, ..., Rn) with the input and output pairs P (Pl, P2, ..., Pn; selecting a subset of the input output pairs P' (P'1, P'2, ..., P'n) from the input and output pairs P (P1, P2, ..., Pn) with rewards R' (R'1, R'2, ..., R'k) deemed significant from the rewards R (R1, R2, ..., Rn); performing supervised training of the new model using the selected subset of input output pairs P’ (P’1, P’2,… Pn). However Tao, in the area of knowledge transfer via reinforcement learning, teaches these limitations. the transferring learning further comprises: storing input and output pairs P (P1, P2, ..., Pn) for the deployed model; (Tao, Col. 5, lines 24-27; “In some embodiments, instance transfer may transfer instances-samplings of the inputs and/or outputs of a teacher model and then re-use the instances (or samples) to train a student model.” Here, “the inputs and/or outputs” are taken from “teacher model,” or deployed model, “to train “a student model,” which denotes the updated algorithm of the claimed invention.) associating rewards R (R1, R2, ..., Rn) with the input and output pairs P (Pl, P2, ..., Pn; (Tao, Col. 5, lines 27-32 “Again, in the exemplary context of reinforcement learning, training system 110 may obtain one or more trajectories (e.g., sequences of states, actions and rewards) of teacher models 115, and use the trajectory samples as instances to facilitate the training of student model 120,” wherein “states” and “actions” are input and output pairs with associating rewards) selecting a subset of the input output pairs P' (P'1, P'2, ..., P'n) from the input and output pairs P (P1, P2, ..., Pn) with rewards R' (R'1, R'2, ..., R'k) deemed significant from the rewards R (R1, R2, ..., Rn); (Tao, Col. 8, lines 4-9; “For instance, when actor 205 navigates different paths, the resultant trajectories - the sequences of states, actions and rewards - may be stored in respective lookup tables,” wherein “the sequences of states, actions and rewards” correspond to input and output pairs P (P1, P2, ..., Pn) and rewards R (R1, R2, ..., Rn). Tao, Col. 9, lines 50-55; “The instance transfer may involve training the student model directly with instances, e.g., samplings of the inputs and/or outputs of a teacher model. For instance, the instances may include sampled trajectories (e.g., sequences of states, actions and rewards) following the policy of the teacher model.” Here, “sampled trajectories” encompasses a subset of states and actions, or input and output pairs P' (P'1, P'2, ..., P'n)…with rewards R' (R'1, R'2, ..., R'k). Tao, Col. 10, lines 9-18; “reinforcement learning model 200 may maintain a buffer of policy parameters and/or corresponding trajectory output (“experience”), with which reinforcement learning model 200 has previously been trained. In some embodiments, the experience may be prioritized. For instance, only experience with a policy loss (e.g., LRL or Lelip) beyond a certain level may be stored in the buffer.” This process corresponds to selecting a subset of input and output pairs… deemed significant from the rewards. Tao, Col. 9, lines 21-25; “reducing LRL, may cause the policy of the student model to mimic the policy of the teacher model as well as increase the total expected rewards following the updated student model,” thereby indicating that “policy loss” is intrinsically linked with reward. Therefore, the “priority,” or the deemed [significance] is in reference to the associated rewards for each sub-set of “experiences,” or input and output pairs.) performing supervised training of the new model using the selected subset of input output pairs P’ (P’1, P’2,… Pn) (Tao, Col. 9, lines 50-55; “The instance transfer may involve training the student model directly with instances, e.g., samplings of the inputs and/or outputs of a teacher model. For instance, the instances may include sampled trajectories (e.g., sequences of states, actions and rewards) following the policy of the teacher model,” wherein the state and action corresponding to each input/output pair constitute labels thus indicating that the samples, or subset[s], are used for “training the student model” or new model. Because the training data is labeled, the new algorithm was trained using supervised learning. Tao, Col 10, lines 17-19; “The prioritized experience in the buffer may be re-used to train reinforcement learning model 200,” wherein the “reinforcement learning model 200,” corresponds to the new algorithm.). Tao is analogous to the claimed invention as both are from the same field of endeavor, that is, model updates using teacher-student paradigms. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the offline to online training and deployment scheme and convergence threshold, as taught by the combination of Han and Li, with the prioritized experience replay and supervised learning of Tao. The motivation to do so is to allow the updated model to learn from the successful actions of the current model (Tao, Col. 10, lines 20-23; “The repeated training with the prioritized experience may strengthen the memory of reinforcement learning model 200 as to what policy shall be avoided or taken.”). Regarding claim 21, the combination of Han, Li and Tao teaches the computer-implemented method of claim 20, wherein (and thus the rejection of claim 20 is incorporated). Han does not explicitly teach the input and output pairs P (Pl, P2, ..., Pn) are state and action pairs wherein the state used as input and the action is used as output and label. However Tao, in the area of knowledge transfer via reinforcement learning, teaches this limitation (Tao, Col. 9, lines 50-55; “The instance transfer may involve training the student model directly with instances, e.g., samplings of the inputs and/or outputs of a teacher model. For instance, the instances may include sampled trajectories (e.g., sequences of states, actions and rewards) following the policy of the teacher model,” wherein the state is the input, the action is the output and the “sampled trajectories” correspond to the sub-set of state and action, or input and output pairs the “teacher” has accumulated. Here, the state and action corresponding to each input/output pair constitute labels). Tao is analogous to the claimed invention as both are from the same field of endeavor, that is, model updates using teacher-student paradigms. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the offline to online training and deployment scheme and convergence threshold, as taught by the combination of Han, Li and Tao, with the prioritized experience replay and supervised learning using labeled input/output pairs, also taught by Tao. The motivation to do so is inherent as the supervised learning process inherited from claim 2, by definition, requires labeled training data. Regarding claim 23, the combination of Han and Li teaches the computer-implemented method of claim 19, wherein (and thus the rejection of claim 19 is incorporated). Han further teaches the deployed model is an artificial intelligence model (Han, Abstract; “The training host of the near-RT RIC is configured to train artificial intelligence (AI)/machine learning (ML) models based on performance and feedback data.”). Claim 22 is rejected under 35 U.S.C. 103 as being unpatentable over Han, in view of Tao, and further in view of Li, Zhu and Zhan. Regarding claim 22, the combination of Han and Li teaches the computer-implemented method of claim 19, wherein (and thus the rejection of claim 19 is incorporated). Han does not explicitly teach choosing between new model outputs and deployed model outputs iteratively with a variable probability of choosing outputs…wherein the variable probability of choosing outputs decreases from initially favoring the deployed model. However, Zhu, in the area of teacher-student reinforcement learning, teaches this limitation (Zhu, Algorithm 3; 4.3 Decay Reusing Probability; “Every time agent 𝑖 visits advised states, a more flexible method is to give it the opportunity to choose between reusing previous advice, learning by itself and asking for advice. We name this method as Decay Reusing Probability (Decay).” Here, “agent i” denotes the new model that is choosing between, “reusing previous advice” (deployed model output), and its own output (“learning by itself” or new model outputs). “In our work, 𝑃𝑟𝑒𝑢𝑠𝑒 determines whether learning agent 𝑖 should reuse a teacher’s advice to guide its action selection,” i.e. choose deployed model outputs, “agent 𝑖 will follow previous advice with probability 𝑃𝑟𝑒𝑢𝑠𝑒…As agent 𝑖 repeatedly performs the latest advice in state 𝑠, 𝑃𝑟𝑒𝑢𝑠𝑒(𝑠) decays exponentially.” Here, "𝑃𝑟𝑒𝑢𝑠𝑒"denotes a variable probability scheme that “decays,” thereby indicating that the variable probability decreases from initially favoring the deployed model. Thus, the “agent i,” or new model, will choose “learning by itself,” or its own output, more often over time.). Zhu is analogous to the claimed invention as both are from the same field of endeavor, that is, model updates using teacher-student paradigms. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the offline to online training and deployment scheme with a convergence threshold, as taught by Han and Li, with the Decay Reusing Probability method of Zhu. The motivation to do so is to make the updated model robust to sub-optimal teachers while allowing it to learn from past experiences (Zhu, Introduction; “As the advice from teacher may be outdated during learning, we also propose Decay Reusing Probability (Decay) to allow agents learning in usual advising framework while reusing previous advices with decaying probabilities.”). Han does not explicitly teach and a variable bonus reward wherein…the variable bonus reward decreases from initially favoring a new model output matching the deployed model output until the variable probability is reduced to zero. However, Zhan, in the area of teacher-student reinforcement learning, teaches this limitation (Zhan, Algorithm 3; 4.2 Multi-Teacher Advice Algorithm, “the MDP can be effectively exploited for successful learning. Unfortunately, such a process is not well modeled using current methods. Here, we remedy this problem by introducing an algorithm which follows the teacher’s advice at the very beginning and then switches to a policy computed by an algorithm operating within the MDP. That is, the teacher guides the student at the beginning of the learning process and as the student gathers more experience, the teacher’s influence diminishes over time,” wherein “the teacher” denotes the deployed model and “the MDP” denotes the new model. To “guide the student” in this context corresponds to matching the output of the student and the teacher. “To leverage both the teacher’s and learned policies, we set a mixed policy of the form 𝜋𝑖+1=𝛽𝑖𝜋𝜏+(1−𝛽𝑖)𝜋̂𝑖, for 0≤𝛽𝑖≤1 to guide the student’s dataset collection while allowing the teacher to fractionally control exploration needed to collect data at the next iteration. β should typically be set so as to decay exponentially over time.” Here, “β” denotes a variable bonus reward that at high values favors “the teacher,” or deployed model output, initially, but “decreases exponentially” over time to favor “the student,” or new model output.). Zhan is analogous to the claimed invention as both are from the same field of endeavor, that is, model updates using teacher-student paradigms. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the offline to online training and deployment scheme with a convergence threshold and decaying experience recall, as taught by the combination of Han, Li and Zhu, with the mixed student-teacher guiding policy of Zhan. The motivation to do so is to allow the updated model to eventually outperform its sub-optimal teacher (Zhan, 4.2 Multi-Teacher Advice Algorithm; “This decreases the student’s reliance on the teacher and allows it to exploit the knowledge gathered from the MDP to learn better behaving policies than that of the teacher.”). Response to Arguments Applicant’s arguments and amendments, filed 12/17/2025, regarding the rejections from the previous office action made under 35 U.S.C. 101 have been fully considered but are not persuasive. Applicant argues that the claimed invention is not directed towards an abstract idea because “…the claims are directed to an improvement to computing technology. Specifically, the claimed invention improves on updating machine learning algorithms by ensuring a deployed new model does not lose accumulated online learning progress of the model being replaced. (See paragraph 35 of the Specification.)…” (page 10). The Examiner agrees with Applicant’s argument and, the rejections under 35 U.S.C. 101 are hereby withdrawn. Additionally, the Applicant argues that “…There is no indication that the initial model 534 is an initial version of running model 538... This is a further reason that claims 1-18, as amended, are patentable over the Applied Art.” (pages 11-12). This limitation has been newly rejected by Tao as shown above. Further, the Applicant states that “Han discloses obtaining a trained or updated model upon detecting performance degradation. For at least the reason that Han triggers use of a trained model upon detecting performance degradation, while the claimed invention performs pre-training when an updated algorithm is identified, Han fails to disclose the pre-training limitation, as claimed. Hans appears to perform offline training on an initial model for storing a trained model in the model repository. This is a further reason that claims 1-18, as amended, are patentable over the Applied Art.” (page 12). The pre-training limitation has been newly rejected at least in light of Tao above. Regarding claims 19-23, the Applicant argues that “Han discloses collecting "ML learning data 516... for offline reinforcement learning" and training "the initial model based on the offline training," but does not disclose "receiving training data generated by a current model comprising the deployed algorithm, the training data being in the form of input/output pairs of the current model for training new models."…” (page 13). Please refer to the new rejection of this limitation at least in light of Tao as shown above. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to CESAR PAULA whose telephone number is (571)272-4128. The examiner can normally be reached Monday - Friday, 6.30am- 4:30 pm ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, David Wiley can be reached at (571)272-3923. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /CESAR B PAULA/Supervisory Patent Examiner, Art Unit 2145
Read full office action

Prosecution Timeline

Show 9 earlier events
Nov 18, 2025
Interview Requested
Dec 15, 2025
Applicant Interview (Telephonic)
Dec 15, 2025
Examiner Interview Summary
Dec 17, 2025
Response Filed
Aug 12, 2026
Final Rejection mailed — §101, §103
Sep 03, 2026
Interview Requested
Sep 23, 2026
Examiner Interview Summary
Sep 23, 2026
Applicant Interview (Telephonic)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737620
Regularised Training of Neural Networks
3y 9m to grant Granted Sep 15, 2026
Patent 12731039
DECISION TREE-ORIENTED VERTICAL FEDERATED LEARNING METHOD
4y 6m to grant Granted Sep 08, 2026
Patent 12699894
THREE-DIMENSIONAL OBJECT DETECTION USING PSEUDO-LABELS
4y 7m to grant Granted Aug 04, 2026
Patent 12670367
APPARATUS AND METHOD WITH NEURAL NETWORK OPERATION
3y 4m to grant Granted Jun 30, 2026
Patent 12596934
PREDICTION-MODEL-BUILDING METHOD, STATE PREDICTION METHOD AND DEVICES THEREOF
4y 0m to grant Granted Apr 07, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
34%
Grant Probability
42%
With Interview (+8.2%)
4y 6m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 174 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month