Prosecution Insights
Last updated: October 02, 2026
Application No. 18/640,007

COMPUTER-READABLE RECORDING MEDIUM STORING REINFORCEMENT LEARNING PROGRAM, REINFORCEMENT LEARNING METHOD, AND INFORMATION PROCESSING APPARATUS

Non-Final OA §101§102§112
Filed
Apr 19, 2024
Priority
May 17, 2023 — JP 2023-081533
Examiner
HOANG, AMY P
Art Unit
Tech Center
Assignee
Fujitsu Limited
OA Round
1 (Non-Final)
72%
Grant Probability
Favorable
1-2
OA Rounds
8m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 72% — above average
72%
Career Allowance Rate
177 granted / 246 resolved
+12.0% vs TC avg
Strong +65% interview lift
Without
With
+65.3%
Interview Lift
resolved cases with interview
Typical timeline
3y 1m
Avg Prosecution
17 currently pending
Career history
270
Total Applications
across all art units

Statute-Specific Performance

§101
17.1%
-22.9% vs TC avg
§103
47.5%
+7.5% vs TC avg
§102
16.9%
-23.1% vs TC avg
§112
13.2%
-26.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 246 resolved cases

Office Action

§101 §102 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This action is responsive to the application filed on 04/19/2024. Claims 1-10 are presented in the case. Claims 1, 4, 5 and 8 are independent claims. Priority Applicant's claim for the benefit of a prior-filed Japanese Patent Application No. 2023-081533, filed on May 17, 2023 is acknowledged. Information Disclosure Statement The information disclosure statement submitted on 04/19/2024 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Specification The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed. The following title is suggested: Computer-readable recording medium storing reinforcement learning program, reinforcement learning method, and information processing apparatus for determining an action for an environment. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. Claims 1, 2, 4, 5, 6, 8 and 9 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor, or for pre-AIA the applicant regards as the invention. Claims 1, 2, 4, 5, 6, 8 and 9 recite the limitation "the environment". There is insufficient antecedent basis for this limitation in the claims. Claims 1, 5 and 8 recite the limitation "the state of the environment". There is insufficient antecedent basis for this limitation in the claims. Claims 2, 6 and 9 recite the limitation "the base station". There is insufficient antecedent basis for this limitation in the claims. Claims 2, 6 and 9 recite the limitation "the first state". There is insufficient antecedent basis for this limitation in the claims. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-10 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1: Claims 1-4 are directed to a medium, claims 5-7 are directed to a method and claims 8-10 are directed to an apparatus. Therefore, the claims are eligible under Step 1 for being directed to a manufacture, a process and a machine respectively. Independent claims 1, 5 and 8: Step 2A Prong 1: Claims recite: calculating a second demand amount after a certain period of time and a reliability of the second demand amount based on a current first demand amount for a service provided in a predetermined environment - Under its broadest reasonable interpretation in light of the specification, this limitation encompasses a mathematical concept of a mathematical calculation of calculating using mathematical methods to calculate a second demand amount and a reliability; determining an action to be performed for the environment in accordance with a machine learning model based on input data that includes the second demand amount, the reliability, and a current first state of the environment - Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mental process of evaluating data and selecting data based on judgement, which is observing, evaluating and judging that is practically capable of being performed in the human mind with the assistance of pen and paper, or is a mathematical concept that is achievable through mathematical computation; updating, based on a second state of the environment after the action is performed and a reward, a parameter of the model by constrained reinforcement learning in which the reward is increased in a range that satisfies a constraint on the state of the environment - Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mental process of evaluating data and selecting data based on judgement, which is observing, evaluating and judging that is practically capable of being performed in the human mind with the assistance of pen and paper, or is a mathematical concept that is achievable through mathematical computation. Step 2A Prong 2: This judicial exception is not integrated into a practical application because they recite the additional elements: A non-transitory computer-readable recording medium storing a reinforcement learning program for causing a computer to execute a process; A reinforcement learning apparatus comprising: a memory, and a processor, coupled to the memory and configured to - These limitations amount to components of a general purpose computer that applies a judicial exception, by use of conventional computer functions (see MPEP § 2106.05(b)). executing the determined action for the environment - the step recited at a high level of generality, and amounts to more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)). Accordingly, these additional elements do not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claims are thus directed to the abstract idea. Step 2B: The claims do not include additional elements that amount to significantly more than the judicial exception. The additional elements: A non-transitory computer-readable recording medium storing a reinforcement learning program for causing a computer to execute a process; A reinforcement learning apparatus comprising: a memory, and a processor, coupled to the memory and configured to - These limitations amount to components of a general purpose computer that applies a judicial exception, by use of conventional computer functions (see MPEP § 2106.05(b)). executing the determined action for the environment - the step recited at a high level of generality, and amounts to more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)). Accordingly, these additional elements do not amount to significantly more than the judicial exception. As such, the claims are ineligible. Dependent claims 2, 6 and 9: Step 2A Prong 1: Claims recite: in the calculating the second demand amount and the reliability, a communication environment of a wireless access network is set as the environment, a current first communication traffic amount of the wireless access network is set as the first demand amount, and based on the first communication traffic amount, a second communication traffic amount after a certain period of time in the wireless access network is calculated as the second demand amount - Under its broadest reasonable interpretation in light of the specification, this limitation encompasses a mathematical concept of a mathematical calculation of calculating using mathematical methods to calculate a second communication traffic amount. in the determining the action, whether to cause the base station to be active or sleep is determined as the action by using a load of the base station in the wireless access network as the first state - Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mental process of evaluating data and selecting data based on judgement, which is observing, evaluating and judging that is practically capable of being performed in the human mind with the assistance of pen and paper, or is a mathematical concept that is achievable through mathematical computation, and in the updating the parameter of the model, a penalty is generated when a second load of the base station after controlling the base station in accordance with the determined action exceeds a threshold related to the load of the base station, a larger value is set as the reward as power consumption of the base station after controlling the base station is smaller, and the parameter of the model is updated so as to increase the reward without generating the penalty - Under its broadest reasonable interpretation in light of the specification, this limitation encompasses a mathematical concept of a mathematical calculation of calculating using mathematical methods to calculate a penalty, a reward and update a parameter. Step 2A Prong 2 & Step 2B: There are no additional elements recited so the claims do not provide a practical application and is not considered to be significantly more. As such, the claims are ineligible. Dependent claims 3, 7 and 10: Step 2A Prong 1: Claims recite: in the calculating the second demand amount and the reliability, a variance of the second demand amount is calculated as the reliability - Under its broadest reasonable interpretation in light of the specification, this limitation encompasses a mathematical concept of a mathematical calculation of calculating using mathematical methods to calculate a variance. Step 2A Prong 2 & Step 2B: There are no additional elements recited so the claims do not provide a practical application and is not considered to be significantly more. As such, the claims are ineligible. Independent claim 4: Step 2A Prong 1: Claim recites: calculating a second demand amount after a certain period of time and a reliability of the second demand amount based on a current first demand amount for a service provided in a predetermined environment - Under its broadest reasonable interpretation in light of the specification, this limitation encompasses a mathematical concept of a mathematical calculation of calculating using mathematical methods to calculate a second demand amount and a reliability; determining, based on input data that includes a third demand amount obtained by adding a value according to the reliability to the second demand amount and a current first state of the environment, an action to be performed for the environment in accordance with a model generated by constrained reinforcement learning that increases a reward in a range that satisfies a constraint on a state of the environment - Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mental process of evaluating data and selecting data based on judgement, which is observing, evaluating and judging that is practically capable of being performed in the human mind with the assistance of pen and paper, or is a mathematical concept that is achievable through mathematical computation. Step 2A Prong 2: This judicial exception is not integrated into a practical application because they recite the additional elements: A non-transitory computer-readable recording medium storing a reinforcement learning program for causing a computer to execute a process - These limitations amount to components of a general purpose computer that applies a judicial exception, by use of conventional computer functions (see MPEP § 2106.05(b)). executing the determined action for the environment - the step recited at a high level of generality, and amounts to more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)). Accordingly, these additional elements do not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claims are thus directed to the abstract idea. Step 2B: The claims do not include additional elements that amount to significantly more than the judicial exception. The additional elements: A non-transitory computer-readable recording medium storing a reinforcement learning program for causing a computer to execute a process - These limitations amount to components of a general purpose computer that applies a judicial exception, by use of conventional computer functions (see MPEP § 2106.05(b)). executing the determined action for the environment - the step recited at a high level of generality, and amounts to more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)). Accordingly, these additional elements do not amount to significantly more than the judicial exception. As such, the claims are ineligible. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claims 1-10 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Singh et al. (hereinafter Singh), US 20230319707 A1. Regarding independent claim 1, Singh teaches a non-transitory computer-readable recording medium storing a reinforcement learning program for causing a computer to execute a process ([0020]-[0021]; [0172]) comprising: calculating a second demand amount after a certain period of time and a reliability of the second demand amount based on a current first demand amount for a service provided in a predetermined environment ([0015] In embodiments, the method further comprises performing the determining in response to receiving, as a new load estimate, a new load prediction from a second trained model comprised in the apparatus or in another apparatus, the second trained model outputting periodically, using at least measured load data from the radio access network as input, load predictions; [0054] In the example illustrated in FIG. 2, there are two different models, named in the example a load prediction model and Q learning, without limiting the models to the specific examples. In the example illustrated in FIG. 2, training the load prediction model (block 201) is performed in the non real time part 210 of the radio intelligent controller, and the Q learning (block 202), predicting load periodically (block 203) using the trained load prediction model; [0057] The load prediction model may be a machine learning based model, and it may be called also a load estimation model. Further, it should be appreciated that in some implementations, no machine learning based load prediction model is used to have a predicted load but a load is estimated based on measured load; [0066] In the context of an optimization algorithm, the function used to evaluate a candidate solution (i.e. a set of weights) is referred to as the objective function. Typically, with neural networks, where the target is to minimize the error, the objective function is often referred to as a cost function or a loss function. In adjusting weights 400, any suitable method may be used as a loss function, some examples are mean squared error (MSE), maximum likelihood (MLE), and cross entropy; [0069] The load data may also include a measure of the cell throughput, and/or a measure of device throughput such as geometric mean of devices throughputs. The load may comprise a vector or a tuple comprising one or more of the various metrics of load. It should be appreciated that any one of the load metrics may be measured over a certain time interval; [0120] In both examples it is assumed that the historical data is offline data that comprises a plurality of time series providing evolution of load data (for example by means of number of active served devices, and/or physical resource blocks used), power consumption data and cell throughput data. The load data is used to identify the state, and the throughput/power related metrics to identify the reward attainable. In some implementations, the historical data may comprise a plurality of time series providing evolution of served device(s) throughput data. The historical data may be, for example, in time series of one hour duration, one time series comprising a plurality of time steps, for example a plurality of one minute granularity load samples. For example, having historical data collected during a week will result with one hour time series to 168 time series. Further assumption made is that per a time step throughput is also known, or determinable based on the historical data. For example, per a sample in the load time series may comprise, per a cell, may comprise a tuple representing number of active served devices and physical resource block (PRB) utilization at a given time period, which may be a time step within the time interval, the time interval, plurality of time intervals. The tuple may comprise also cell throughput at the given time interval. The tuple for load may be expanded to for example a mean and a variance of the load (tuple) components, or a mean and Xth percentiles of the load components); determining an action to be performed for the environment in accordance with a machine learning model based on input data that includes the second demand amount, the reliability, and a current first state of the environment ([0052] FIG. 2 illustrates a neural network based solution to decide whether to change status of one or more cells, i.e. switch on or off one or more cells, or to retain them in their current status; [0054] In the example illustrated in FIG. 2, there are two different models, named in the example a load prediction model and Q learning, without limiting the models to the specific examples. In the example illustrated in FIG. 2, training the load prediction model (block 201) is performed in the non real time part 210 of the radio intelligent controller, and the Q learning (block 202), predicting load periodically (block 203) using the trained load prediction model, and determining optimal action (block 204) are performed in the near real time part 220 of the radio intelligent controller. A radio access network node 230 performs determined action (block 205), measures load and throughput metrics (block 206) and measures power consumed (block 207). More precisely, the radio access network node 230 performs the determined action to its transceiver(s), or transmitter(s), or receiver(s), or other radio part(s), or radio head(s) that provide a cell/cells and whose status is to be changed or power settings modified. However, herein term “cell” is used for the sake of clarity to cover the different electronic devices for transmitting and/or receiving data in radio waves, and thereby providing served devices access to communications network. Depending on an implementation, the determined action performed may be a status change (switching off cell(s) or switching on cell(s)), or one of the status change and modifying power settings of cell(s). Further, the radio access network node 230 reports measured load and throughput metrics to blocks 201, 202 and 203. It should be appreciated that training the load prediction model (block 201) may be performed in the near real time part 220 of the radio intelligent controller and/or in the radio access network node, and/or the determining optimal action (block 204) may be performed in the radio access network node, for example); executing the determined action for the environment ([0054] A radio access network node 230 performs determined action (block 205), measures load and throughput metrics (block 206) and measures power consumed (block 207)); and updating, based on a second state of the environment after the action is performed and a reward, a parameter of the model by constrained reinforcement learning in which the reward is increased in a range that satisfies a constraint on the state of the environment ([0064] Initial weights 400 of the model can be set in various alternative ways. During the training phase they are adapted to improve the accuracy of the process based on analyzing errors in decision making. Training a model is basically a trial and error activity. In principle, each node 304, 306, 308, 310, 312 of the neural network 330 makes a decision (input*weight) and then compares this decision to collected data to find out the difference to the collected data. In other words, it determines the error, based on which the weights 400 are adjusted. Thus, the training of the model may be considered a corrective feedback loop; [0065] Typically, a neural network model is trained using a stochastic gradient descent optimization algorithm for which the gradients are calculated using the backpropagation algorithm. The gradient descent algorithm seeks to change the weights 400 so that the next evaluation reduces the error, meaning the optimization algorithm is navigating down the gradient (or slope) of error. It is also possible to use any other suitable optimization algorithm if it provides sufficiently accurate weights 400. Consequently, the trained parameters 332 of the neural network 330 may comprise the weights 400; [0066] In the context of an optimization algorithm, the function used to evaluate a candidate solution (i.e. a set of weights) is referred to as the objective function. Typically, with neural networks, where the target is to minimize the error, the objective function is often referred to as a cost function or a loss function. In adjusting weights 400, any suitable method may be used as a loss function, some examples are mean squared error (MSE), maximum likelihood (MLE), and cross entropy; [0071] FIG. 5 illustrates basic functionality of the trained power saving model or an apparatus comprising the trained power saving model, for example, to decide when to switch on or off one or more cells so that long term reward is maximized, i.e. a long term objective of power savings versus ensuring throughput (capacity) is balanced; [0075] The tradeoff function may define for each possible action a long term reward and the optimal action is the action providing the biggest reward). Regarding dependent claim 2, Singh teaches all the limitations as set forth in the rejection of claim 1 that is incorporated. Singh further teaches wherein in the calculating the second demand amount and the reliability, a communication environment of a wireless access network is set as the environment ([0038]-[0039] FIG. 1 shows devices 100 and 102. The devices 100 and 102 may, for example, be user devices. The devices 100 and 102 are configured to be in a wireless connection on one or more communication channels with a node 104. The node 104 is further connected to a core network 110), a current first communication traffic amount of the wireless access network is set as the first demand amount ([0069] The historical load data may be real-time messaging protocol (RTMP) data collected on streaming audio, video and/or data, for example. Load data, including the historical load data, may comprise various metrics of load. A non-limiting list of examples of various metrics includes a volume of a traffic arriving at or delivered by various cells in downlink and/or uplink (measured in bytes or megabytes, for example), air interface resources, for example physical resource blocks (PRBs) or data channel resources, or control channel resources, required to deliver the traffic, fraction of time/frequency resources consumed by uplink or downlink transmissions, number of devices connected to various cells, number of active devices, an active device being a device that have data ready to deliver, ratio of active devices to system bandwidth, expressed in Megahertz or in PRBs, an effective number of devices that may take into account the distribution or load-balancing of devices across multiple cells. The load data may also include a measure of the cell throughput, and/or a measure of device throughput such as geometric mean of devices throughputs. The load may comprise a vector or a tuple comprising one or more of the various metrics of load), and based on the first communication traffic amount, a second communication traffic amount after a certain period of time in the wireless access network is calculated as the second demand amount ([0015] In embodiments, the method further comprises performing the determining in response to receiving, as a new load estimate, a new load prediction from a second trained model comprised in the apparatus or in another apparatus, the second trained model outputting periodically, using at least measured load data from the radio access network as input, load predictions), in the determining the action, whether to cause the base station to be active or sleep is determined as the action by using a load of the base station in the wireless access network as the first state ([0052] FIG. 2 illustrates a neural network based solution to decide whether to change status of one or more cells, i.e. switch on or off one or more cells, or to retain them in their current status; [0054] In the example illustrated in FIG. 2, there are two different models, named in the example a load prediction model and Q learning, without limiting the models to the specific examples. In the example illustrated in FIG. 2, training the load prediction model (block 201) is performed in the non real time part 210 of the radio intelligent controller, and the Q learning (block 202), predicting load periodically (block 203) using the trained load prediction model, and determining optimal action (block 204) are performed in the near real time part 220 of the radio intelligent controller. A radio access network node 230 performs determined action (block 205), measures load and throughput metrics (block 206) and measures power consumed (block 207). More precisely, the radio access network node 230 performs the determined action to its transceiver(s), or transmitter(s), or receiver(s), or other radio part(s), or radio head(s) that provide a cell/cells and whose status is to be changed or power settings modified. However, herein term “cell” is used for the sake of clarity to cover the different electronic devices for transmitting and/or receiving data in radio waves, and thereby providing served devices access to communications network. Depending on an implementation, the determined action performed may be a status change (switching off cell(s) or switching on cell(s)), or one of the status change and modifying power settings of cell(s). Further, the radio access network node 230 reports measured load and throughput metrics to blocks 201, 202 and 203. It should be appreciated that training the load prediction model (block 201) may be performed in the near real time part 220 of the radio intelligent controller and/or in the radio access network node, and/or the determining optimal action (block 204) may be performed in the radio access network node, for example), and in the updating the parameter of the model, a penalty is generated when a second load of the base station after controlling the base station in accordance with the determined action exceeds a threshold related to the load of the base station, a larger value is set as the reward as power consumption of the base station after controlling the base station is smaller, and the parameter of the model is updated so as to increase the reward without generating the penalty ([0064] Initial weights 400 of the model can be set in various alternative ways. During the training phase they are adapted to improve the accuracy of the process based on analyzing errors in decision making. Training a model is basically a trial and error activity. In principle, each node 304, 306, 308, 310, 312 of the neural network 330 makes a decision (input*weight) and then compares this decision to collected data to find out the difference to the collected data. In other words, it determines the error, based on which the weights 400 are adjusted; [0075]-[0085] The tradeoff function may define for each possible action a long term reward and the optimal action is the action providing the biggest reward. The tradeoff function takes into account conflicting objectives relating to switching on or off one or more cells. On the one hand, switching off one or more cells may reduce power consumption. On the other hand, switching off one or more cells may reduce air interface resources available for transmissions to/from served devices, and thereby reduce the throughput experienced by users of the served devices … In an implementation, the tradeoff function may be calculated as a function of the throughput achieved, the power consumed, and a relative weight representing the relative importance of the throughput function and the power consumption function. In an implementation, the tradeoff function may be calculated as a function of the throughput achieved, the power consumed, and a relative weight representing the relative importance of the throughput function and the power consumption function. In another implementation, the tradeoff function may be provided by the network operator as a policy input, by specifying the function of the throughput to use in calculating the tradeoff, the function of the power consumption, and the relative weight. The function of the throughput may be considered as a benefit function, and the function of the power consumption may be considered as a penalty function … The reward may calculated, for example, using a simple reward function (equation 1)). Regarding dependent claim 3, Singh teaches all the limitations as set forth in the rejection of claim 1 that is incorporated. Singh further teaches wherein in the calculating the second demand amount and the reliability, a variance of the second demand amount is calculated as the reliability ([0066] In the context of an optimization algorithm, the function used to evaluate a candidate solution (i.e. a set of weights) is referred to as the objective function. Typically, with neural networks, where the target is to minimize the error, the objective function is often referred to as a cost function or a loss function. In adjusting weights 400, any suitable method may be used as a loss function, some examples are mean squared error (MSE), maximum likelihood (MLE), and cross entropy; [0069] The load data may also include a measure of the cell throughput, and/or a measure of device throughput such as geometric mean of devices throughputs. The load may comprise a vector or a tuple comprising one or more of the various metrics of load. It should be appreciated that any one of the load metrics may be measured over a certain time interval; [0120] In both examples it is assumed that the historical data is offline data that comprises a plurality of time series providing evolution of load data (for example by means of number of active served devices, and/or physical resource blocks used), power consumption data and cell throughput data. The load data is used to identify the state, and the throughput/power related metrics to identify the reward attainable. In some implementations, the historical data may comprise a plurality of time series providing evolution of served device(s) throughput data. The historical data may be, for example, in time series of one hour duration, one time series comprising a plurality of time steps, for example a plurality of one minute granularity load samples. For example, having historical data collected during a week will result with one hour time series to 168 time series. Further assumption made is that per a time step throughput is also known, or determinable based on the historical data. For example, per a sample in the load time series may comprise, per a cell, may comprise a tuple representing number of active served devices and physical resource block (PRB) utilization at a given time period, which may be a time step within the time interval, the time interval, plurality of time intervals. The tuple may comprise also cell throughput at the given time interval. The tuple for load may be expanded to for example a mean and a variance of the load (tuple) components, or a mean and Xth percentiles of the load components). Regarding independent claim 4, Singh teaches a non-transitory computer-readable recording medium storing a reinforcement learning program for causing a computer to execute a process ([0020]-[0021]; [0172]) comprising: calculating a second demand amount after a certain period of time and a reliability of the second demand amount based on a current first demand amount for a service provided in a predetermined environment ([0015] In embodiments, the method further comprises performing the determining in response to receiving, as a new load estimate, a new load prediction from a second trained model comprised in the apparatus or in another apparatus, the second trained model outputting periodically, using at least measured load data from the radio access network as input, load predictions; [0054] In the example illustrated in FIG. 2, there are two different models, named in the example a load prediction model and Q learning, without limiting the models to the specific examples. In the example illustrated in FIG. 2, training the load prediction model (block 201) is performed in the non real time part 210 of the radio intelligent controller, and the Q learning (block 202), predicting load periodically (block 203) using the trained load prediction model; [0057] The load prediction model may be a machine learning based model, and it may be called also a load estimation model. Further, it should be appreciated that in some implementations, no machine learning based load prediction model is used to have a predicted load but a load is estimated based on measured load; [0066] In the context of an optimization algorithm, the function used to evaluate a candidate solution (i.e. a set of weights) is referred to as the objective function. Typically, with neural networks, where the target is to minimize the error, the objective function is often referred to as a cost function or a loss function. In adjusting weights 400, any suitable method may be used as a loss function, some examples are mean squared error (MSE), maximum likelihood (MLE), and cross entropy; [0069] The load data may also include a measure of the cell throughput, and/or a measure of device throughput such as geometric mean of devices throughputs. The load may comprise a vector or a tuple comprising one or more of the various metrics of load. It should be appreciated that any one of the load metrics may be measured over a certain time interval; [0120] In both examples it is assumed that the historical data is offline data that comprises a plurality of time series providing evolution of load data (for example by means of number of active served devices, and/or physical resource blocks used), power consumption data and cell throughput data. The load data is used to identify the state, and the throughput/power related metrics to identify the reward attainable. In some implementations, the historical data may comprise a plurality of time series providing evolution of served device(s) throughput data. The historical data may be, for example, in time series of one hour duration, one time series comprising a plurality of time steps, for example a plurality of one minute granularity load samples. For example, having historical data collected during a week will result with one hour time series to 168 time series. Further assumption made is that per a time step throughput is also known, or determinable based on the historical data. For example, per a sample in the load time series may comprise, per a cell, may comprise a tuple representing number of active served devices and physical resource block (PRB) utilization at a given time period, which may be a time step within the time interval, the time interval, plurality of time intervals. The tuple may comprise also cell throughput at the given time interval. The tuple for load may be expanded to for example a mean and a variance of the load (tuple) components, or a mean and Xth percentiles of the load components); determining, based on input data that includes a third demand amount obtained by adding a value according to the reliability to the second demand amount and a current first state of the environment, an action to be performed for the environment in accordance with a model generated by constrained reinforcement learning that increases a reward in a range that satisfies a constraint on a state of the environment ([0052] FIG. 2 illustrates a neural network based solution to decide whether to change status of one or more cells, i.e. switch on or off one or more cells, or to retain them in their current status; [0054] In the example illustrated in FIG. 2, there are two different models, named in the example a load prediction model and Q learning, without limiting the models to the specific examples. In the example illustrated in FIG. 2, training the load prediction model (block 201) is performed in the non real time part 210 of the radio intelligent controller, and the Q learning (block 202), predicting load periodically (block 203) using the trained load prediction model, and determining optimal action (block 204) are performed in the near real time part 220 of the radio intelligent controller. A radio access network node 230 performs determined action (block 205), measures load and throughput metrics (block 206) and measures power consumed (block 207). More precisely, the radio access network node 230 performs the determined action to its transceiver(s), or transmitter(s), or receiver(s), or other radio part(s), or radio head(s) that provide a cell/cells and whose status is to be changed or power settings modified. However, herein term “cell” is used for the sake of clarity to cover the different electronic devices for transmitting and/or receiving data in radio waves, and thereby providing served devices access to communications network. Depending on an implementation, the determined action performed may be a status change (switching off cell(s) or switching on cell(s)), or one of the status change and modifying power settings of cell(s). Further, the radio access network node 230 reports measured load and throughput metrics to blocks 201, 202 and 203. It should be appreciated that training the load prediction model (block 201) may be performed in the near real time part 220 of the radio intelligent controller and/or in the radio access network node, and/or the determining optimal action (block 204) may be performed in the radio access network node, for example; [0064] Initial weights 400 of the model can be set in various alternative ways. During the training phase they are adapted to improve the accuracy of the process based on analyzing errors in decision making. Training a model is basically a trial and error activity. In principle, each node 304, 306, 308, 310, 312 of the neural network 330 makes a decision (input*weight) and then compares this decision to collected data to find out the difference to the collected data. In other words, it determines the error, based on which the weights 400 are adjusted. Thus, the training of the model may be considered a corrective feedback loop; [0065] Typically, a neural network model is trained using a stochastic gradient descent optimization algorithm for which the gradients are calculated using the backpropagation algorithm. The gradient descent algorithm seeks to change the weights 400 so that the next evaluation reduces the error, meaning the optimization algorithm is navigating down the gradient (or slope) of error. It is also possible to use any other suitable optimization algorithm if it provides sufficiently accurate weights 400. Consequently, the trained parameters 332 of the neural network 330 may comprise the weights 400; [0066] In the context of an optimization algorithm, the function used to evaluate a candidate solution (i.e. a set of weights) is referred to as the objective function. Typically, with neural networks, where the target is to minimize the error, the objective function is often referred to as a cost function or a loss function. In adjusting weights 400, any suitable method may be used as a loss function, some examples are mean squared error (MSE), maximum likelihood (MLE), and cross entropy; [0071] FIG. 5 illustrates basic functionality of the trained power saving model or an apparatus comprising the trained power saving model, for example, to decide when to switch on or off one or more cells so that long term reward is maximized, i.e. a long term objective of power savings versus ensuring throughput (capacity) is balanced; [0075] The tradeoff function may define for each possible action a long term reward and the optimal action is the action providing the biggest reward); and executing the determined action for the environment ([0054] A radio access network node 230 performs determined action (block 205), measures load and throughput metrics (block 206) and measures power consumed (block 207)). Regarding independent claim 5, it is a method claim that corresponding to the medium of claim 1. Therefore, it is rejected for the same reason as claim 1 above. Regarding dependent claim 6, it is a method claim that corresponding to the medium of claim 2. Therefore, it is rejected for the same reason as claim 2 above. Regarding dependent claim 7, it is a method claim that corresponding to the medium of claim 3. Therefore, it is rejected for the same reason as claim 3 above. Regarding independent claim 8, it is an apparatus claim that corresponding to the medium of claim 1. Therefore, it is rejected for the same reason as claim 1 above. Singh further teaches a reinforcement learning apparatus comprising: a memory, and a processor, coupled to the memory ([0004]; Figs. 11-12; [0160]). Regarding dependent claim 9, it is an apparatus claim that corresponding to the medium of claim 2. Therefore, it is rejected for the same reason as claim 2 above. Regarding dependent claim 10, it is an apparatus claim that corresponding to the medium of claim 3. Therefore, it is rejected for the same reason as claim 3 above. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Applicant is required under 37 C.F.R. § 1.111(c) to consider these references fully when responding to this action. KARAPANTELAKIS et al. (US 20240406835 A1) discloses a method for training a reinforcement learning system for optimising routing for a network including a plurality of Integrated Access and Backhaul (IAB) nodes connected to an IAB donor. It is noted that any citation to specific pages, columns, lines, or figures in the prior art references and any interpretation of the references should not be considered to be limiting in any way. A reference is relevant for all it contains and may be relied upon for all that it would have reasonably suggested to one having ordinary skill in the art. In re Heck, 699 F.2d 1331, 1332-33, 216 U.S.P.Q. 1038, 1039 (Fed. Cir. 1983) (quoting In re Lemelson, 397 F.2d 1006, 1009, 158 U.S.P.Q. 275, 277 (C.C.P.A. 1968)). Any inquiry concerning this communication or earlier communications from the examiner should be directed to AMY P HOANG whose telephone number is (469)295-9134. The examiner can normally be reached M-TH 8:30-5:00PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, JENNIFER WELCH can be reached at 571-272-7212. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /AMY P HOANG/Examiner, Art Unit 2143 /JENNIFER N WELCH/Supervisory Patent Examiner, Art Unit 2143
Read full office action

Prosecution Timeline

Apr 19, 2024
Application Filed
Aug 13, 2026
Non-Final Rejection mailed — §101, §102, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743654
SYSTEM AND METHOD FOR GENERATING REPHRASED ACTIONABLE DATA OF TEXTUAL DATA
4y 3m to grant Granted Sep 22, 2026
Patent 12717243
METHODS OF DETERMINING PROCESS MODELS BY MACHINE LEARNING
5y 6m to grant Granted Aug 25, 2026
Patent 12699917
Capturing Ordinal Historical Dependence in Graphical Event Models with Tree Representations
4y 9m to grant Granted Aug 04, 2026
Patent 12632792
STABLE LOCAL INTERPRETABLE MODEL FOR PREDICTION
4y 2m to grant Granted May 19, 2026
Patent 12619452
INTELLIGENT AUTOMATED ASSISTANT IN A MESSAGING ENVIRONMENT
2y 9m to grant Granted May 05, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
72%
Grant Probability
99%
With Interview (+65.3%)
3y 1m (~8m remaining)
Median Time to Grant
Low
PTA Risk
Based on 246 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month