CTNF 18/563,046 CTNF 82491 Notice of Pre-AIA or AIA Status 07-03-aia AIA 15-10-aia 1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA. 2. Claims 1-15 are pending in this office action. This action is responsive to Applicant’s application filed 11/21/2023. Information Disclosure Statement 3. The references listed in the IDS filed 11/21/2023 has been considered. A copy of the signed or initialed IDS is hereby attached. Claim Rejections - 35 USC § 102 07-07-aia AIA 07-07 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – 07-08-aia AIA (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention. 07-15 AIA 4. Claim s 1-15 are rejected under 35 U.S.C. 102( a1) (a2 ) as being anticipated by Mujumdar et al. (US Patent Publication No. 2024/0195689 A1, hereinafter “Mujumdar”. As to Claim 1, Mujumdar teaches the claimed limitations: “ A learning device comprising:” as a computer program product comprising a computer readable medium, the computer readable medium having computer readable code embodied therein, the computer readable code being configured such that, on execution by a suitable computer or processor, the computer or processor is caused to perform a method according to any one or more of aspects. The processing circuitry is further configured to cause the distributed node to, for each of a plurality of possible actions that may be executed on the environment in its current state, use a Reinforcement Learning, RL, process to obtain predicted values of a reward function representing possible impacts of execution of the possible action on performance of the task by the environment (paragraphs 0017-0018). “A memory storing instructions; and one or more processors configured to execute the instructions to:” as the distributed node may for example receive a plurality of risk contours from a server node, store the received risk contours in a memory and, in later iterations of the method, retrieve previously received risk contours for a possible action from the memory (paragraph 0057). “Accept input of a reward function that defines cumulative reward by a reward term based on a high-level indicator representing a production indicator” as the agent's policy defines the control strategy implemented by the agent, and is a mapping from states to a policy distribution over possible actions, the distribution indicating the probability that each possible action is the most favorable given the current state. An RL interaction proceeds at each time instant the agent finds the environment in a state. The agent selects an action receives a stochastic reward and the environment transitions to a new state. The agent's goal is to find the optimal policy, i.e. a policy that maximizes the expected cumulative reward over a predefined period of time, also known as the policy value function (paragraph 0007). The objective is to maximize some trade-off between capacity and coverage. The environment state comprises representative Key Performance Indicators (KPIs) of each cell. The reward is given by the environment, and is modelled as a weighted sum of capacity, coverage and quality KPIs (paragraphs 0047, 0092, 0119). Using an RL process to obtain predicted values of a reward function representing possible impacts of execution of a possible action on performance of the task by the environment may comprise using the RL process to predict values of the reward function on the basis of the obtained representation of the state of the environment and the possible action. Using a trained Machine Learning (ML) model, which model may be specific to the environment that is managed by the distributed node, one or more environments within a domain may be associated with different reward functions, and with different models for predicting values of reward functions, such that different distributed nodes may obtain different predicted rewards based on the same possible actions and current state representations. Using the trained ML model to predict values of the reward function may comprise inputting the obtained representation of a current state of the environment to the trained ML model, wherein the trained ML model processes the representation in accordance with parameters of the ML model that have been set during training, and outputs a vector of predicted reward values corresponding to the possible future states that may be entered by the environment on execution of the possible action (paragraph 0055). In 5G network slicing, a safety specification may be defined in terms of a high-level intent. The hazard score for individual states can for example be computed on the basis of latencies in individual sub-domains that are part of the slice (paragraph 0130). “Learn a value function for deriving optimal policy for an agent using training data and the reward function” as the agent's goal is to find the optimal policy, i.e. a policy that maximizes the expected cumulative reward over a predefined period of time, also known as the policy value function (paragraph0007). In the context use a Reinforcement Learning (RL), an optimal policy is usually derived in a trial-and-error fashion by direct interaction with the environment (paragraph 0009). Consequently, the standard approach for RL solutions is to employ a simulator as a proxy for the real environment during the training phase (paragraphs 0010-0012). “Output the learned value function” as using the trained Learning algorithms (ML) model to predict values of the reward function may comprise inputting the obtained representation of a current state of the environment to the trained ML model, wherein the trained ML model processes the representation in accordance with parameters of the ML model that have been set during training, and outputs a vector of predicted reward values corresponding to the possible future states that may be entered by the environment on execution of the possible action (paragraph 0055). As to Claim 2, Mujumdar teaches the claimed limitations: “Wherein the processor is configured to execute the instructions to: accept input of a reward function that defines the cumulative reward by multiple reward terms, terms; and learn a value function using the reward function” as (paragraphs 0007-0008, 0015, 0018, 0048, 0050-0051, 0055, 0061, 0066). As to Claim 3, Mujumdar teaches the claimed limitations: “Wherein the processor is configured to execute the instructions to accept input of the reward function with weight set for each reward term” as (paragraphs 0021, 0048, 0055, 0066, 0070, 0136, 0142, 0148-0149). As to Claim 4, Mujumdar teaches the claimed limitations: “Wherein the processor is configured to execute the instructions to: accept input of a reward function that defines cumulative reward by multiple reward terms each having a causal relationship; and learn a value function using the reward function” as (paragraphs 0007-0008, 0015, 0018, 0048, 0050, 0055, 0061, 0066, 0148-0149). As to Claim 5, Mujumdar teaches the claimed limitations: “ Wherein the processor is configured to execute the instructions to: accept input of a reward function that defines cumulative reward by multiple reward terms each having a trade-off relationship, relationship; and learn a value function using the reward function” as (paragraphs 0007-0008, 0015, 0018, 0048, 0050, 0055, 0061, 0066, 0148-0149). As to Claim 6, Mujumdar teaches the claimed limitations: “Wherein the processor is configured to execute the instructions to: accept input of a reward function that defines cumulative reward by a reward term representing stock quantity and a reward term representing production quantity; and learn a value function using the reward function” as (paragraphs 0007-0008, 0015, 0018, 0048, 0050, 0055, 0061, 0066, 0148-0149). As to Claim 7, Mujumdar teaches the claimed limitations: “Wherein the processor is configured to execute the instructions to: accept input of a reward function that defines cumulative reward by a reward term representing a lead time and a reward term representing a throughput; and learn a value function using the reward function” as (paragraphs 0007-0008, 0015, 0018, 0048, 00500055, 0061, 0066, 0148-0149). As to Claim 8, Mujumdar teaches the claimed limitations: “ Wherein the processor is configured to execute the instructions to learn a value function that indicates a policy of an agent using training data that includes high-level indicator, location information of the agent, an action of the agent, and reward information according to the action” as (paragraphs 0007-0008, 0015-0018, 0048, 0050, 0055, 0061, 0066, 0148-0149). As to Claim 9, Mujumdar teaches the claimed limitations: “ Wherein the processor is configured to execute the instructions to: accept input of a reward function that includes a reward term depending on success or failure of delivery of goods during transportation; and learn a value function using the reward function” as (paragraphs 0007-0008, 0015, 0018, 0048, 0050, 0055, 0061, 0066, 0148-0149). As to Claim 10, Mujumdar teaches the claimed limitations: “ A learning system comprising:” as a computer program product comprising a computer readable medium, the computer readable medium having computer readable code embodied therein, the computer readable code being configured such that, on execution by a suitable computer or processor, the computer or processor is caused to perform a method according to any one or more of aspects. The processing circuitry is further configured to cause the distributed node to, for each of a plurality of possible actions that may be executed on the environment in its current state, use a Reinforcement Learning, RL, process to obtain predicted values of a reward function representing possible impacts of execution of the possible action on performance of the task by the environment (paragraphs 0017-0018). “a simulator which outputs data including a high-level indicator, location information of an agent, an action of the agent, and reward information according to the action, from map information which is information indicating operating area of the agent, related agent information which is information of other related agents, the high-level indicator which represents a production indicator, and a route plan of the agent” as a simulator as a proxy for the real environment during the training phase, the simulators are often subject to modelling errors related to inherent environment stochasticity (paragraphs 0010-0011). An RL interaction proceeds at each time instant the agent finds the environment in a state. The agent selects an action receives a stochastic reward and the environment transitions to a new state. The agent's goal is to find the optimal policy, i.e. a policy that maximizes the expected cumulative reward over a predefined period of time, also known as the policy value function (paragraph 0007). The Reinforcement Learning (RL) agent needs to try out different state-action combinations (e.g., cumulative reward) with sufficient frequency to be able to make accurate predictions about the rewards and the transition probabilities of each state-action pair. It is therefore necessary for the agent to repeatedly choose suboptimal actions, which conflict with its goal of maximizing the accumulated reward, in order to sufficiently explore the state-action space. At each time step, the agent must decide whether to prioritize further gathering (e.g., accumulated) of information or to make the best move given current knowledge. Exploration may create opportunities by discovering higher rewards (e.g., high-level indicator) on the basis of previously untried actions. However, exploration also carries the risk that previously unexplored decisions will not provide increased reward, and may instead have a negative impact on the environment (paragraph 0008). The objective is to maximize some trade-off between capacity and coverage (for example 5.sup.th percentile user throughput). The environment state comprises representative Key Performance Indicators (KPIs) of each cell. The reward is given by the environment, and is modelled as a weighted sum of capacity, coverage and quality KPIs (paragraphs 0047, 0092, 0119). “A learning device that uses data output from the simulator as training data for learning” as the agent's goal is to find the optimal policy, i.e. a policy that maximizes the expected cumulative reward over a predefined period of time, also known as the policy value function (paragraph0007). In the context use a Reinforcement Learning (RL), an optimal policy is usually derived in a trial-and-error fashion by direct interaction with the environment (paragraph 0009). Consequently, the standard approach for RL solutions is to employ a simulator as a proxy for the real environment during the training phase (paragraphs 0010-0012). “Wherein the learning system includes: an input unit which accepts input of a reward function that defines cumulative reward by a reward term based on the high-level indicator” as using the trained ML model to predict values of the reward function may comprise inputting the obtained representation of a current state of the environment to the trained ML model, wherein the trained ML model processes the representation in accordance with parameters of the ML model that have been set during training, and outputs a vector of predicted reward values corresponding to the possible future states that may be entered by the environment on execution of the possible action (paragraph 0055). In 5G network slicing, a safety specification may be defined in terms of a high-level intent. The hazard score for individual states can for example be computed on the basis of latencies in individual sub-domains that are part of the slice (paragraph 0130). “A learning unit which learns a value function for deriving optimal policy for the agent using the training data and the reward function” as the agent's goal is to find the optimal policy, i.e. a policy that maximizes the expected cumulative reward over a predefined period of time, also known as the policy value function (paragraph0007). In the context use a Reinforcement Learning (RL), an optimal policy is usually derived in a trial-and-error fashion by direct interaction with the environment (paragraph 0009). Consequently, the standard approach for RL solutions is to employ a simulator as a proxy for the real environment during the training phase (paragraphs 0010-0012). “An output unit which outputs the learned value function” as using the trained Learning algorithms (ML) model to predict values of the reward function may comprise inputting the obtained representation of a current state of the environment to the trained ML model, wherein the trained ML model processes the representation in accordance with parameters of the ML model that have been set during training, and outputs a vector of predicted reward values corresponding to the possible future states that may be entered by the environment on execution of the possible action (paragraph 0055). As to Claim 11, Mujumdar teaches the claimed limitations: “ Wherein the simulator outputs data including the high-level indicator, location information of agent transporting goods, an action of the agent, and reward information according to the action, from route information within a facility showing map information, location and performance of the agent to which the goods are to be transported showing related agent information, the high-level indicator, and route plan of the agent” as (paragraphs 0007, 0017, 0041, 0047, 0055, 0092, 0119). As to Claim 12, Mujumdar teaches the claimed limitations: “ Accepting input of a reward function that defines cumulative reward by a reward term based on a high-level indicator representing a production indicator by a computer; learning a value function for deriving optimal policy for an agent using training data and the reward function by the computer; and outputting the learned value function by the computer” as computer implemented method for managing an environment within a domain, the environment being operable to perform a task. The method, performed by a distributed node, comprises obtaining a representation of a current state of the environment (paragraph 0007,0009-0012, 0015, 0047, 0055, 0092, 0119, 0130). As to Claim 13, Mujumdar teaches the claimed limitations: “ Wherein the computer accepts input of a reward function that defines the cumulative reward by multiple reward terms, and the computer learns a value function using the reward function” as (paragraphs 0002, 0007-0008, 0015, 0018, 0048,0050, 0054-0056, 0059-0061, 0066, 0068-0070, 0093, 0101, 0104, 0113-0014, 0127, 0150). As to Claim 14, Mujumdar teaches the claimed limitations: “ A non-transitory computer readable information recording medium storing a learning program, when executed by a processor, that performs a method for: accepting input of a reward function that defines cumulative reward by a reward term based on a high-level indicator representing a production indicator; learning a value function for deriving optimal policy for an agent using training data and the reward function; and outputting the learned value function” as a method, management node, and computer readable medium which at least partially address one or more of the challenges. It is providing a method, management node and computer readable medium which cooperate to implement a safe Reinforcement Learning process in a distributed environment (paragraph 0007,0009-0012, 0014, 0047, 0055, 0092, 0119, 0130). As to Claim 15, Mujumdar teaches the claimed limitations: “ For causing the computer to further execute: input of a reward function that defines the cumulative reward by multiple reward terms is accepted; and a value function is learned using the reward function” as (paragraphs 0007, 0009-0012, 0047, 0055, 0092, 0119, 0130). Examiner’s Note Examiner has cited particular columns/paragraph and line numbers in the references applied to the claims above for the convenience of the applicant. Although the specified citations are representative of the teachings of the art and are applied to specific limitations within the individual claim, other passages and figures may apply as well. It is respectfully requested from the applicant in preparing responses, to fully consider the references in entirety as potentially teaching all or part of the claimed invention, as well as the context of the passage as taught by the prior art or disclosed by the Examiner. In the case of amending the Claimed invention, Applicant is respectfully requested to indicate the portion(s) of the specification which dictate(s) the structure relied on for proper interpretation and also to verify and ascertain the metes and bounds of the claimed invention. This will assist in expediting compact prosecution. MPEP 714.02 recites: “Applicant should also specifically point out the support for any amendments made to the disclosure. See MPEP § 2163.06. An amendment which does not comply with the provisions of 37 CFR 1.121(b), (c), (d), and (h) may be held not fully responsive. See MPEP § 714.” Amendments not pointing to specific support in the disclosure may be deemed as not complying with provisions of 37 C.F.R. 1.131(b), (c), (d), and (h) and therefore held not fully responsive. Generic statements such as “Applicants believe no new matter has been introduced” may be deemed insufficient. Contact Information Any inquiry concerning this communication or earlier communications from the examiner should be directed to James Hwa whose telephone number is 571-270-1285, email address is james.hwa@uspto.gov . The examiner can normally be reached on 9:00 am – 5:30 pm EST. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Ajay Bhatia can be reached on 571-272-3906. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only, for more information about the PAIR system, see http://pair-direct.uspto.gov . Should you have questions on access to the PAIR system contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SHYUE JIUNN HWA/ Primary Examiner, Art Unit 2156 Application/Control Number: 18/563,046 Page 2 Art Unit: 2156 Application/Control Number: 18/563,046 Page 3 Art Unit: 2156 Application/Control Number: 18/563,046 Page 4 Art Unit: 2156 Application/Control Number: 18/563,046 Page 5 Art Unit: 2156 Application/Control Number: 18/563,046 Page 6 Art Unit: 2156 Application/Control Number: 18/563,046 Page 7 Art Unit: 2156 Application/Control Number: 18/563,046 Page 8 Art Unit: 2156 Application/Control Number: 18/563,046 Page 9 Art Unit: 2156 Application/Control Number: 18/563,046 Page 10 Art Unit: 2156 Application/Control Number: 18/563,046 Page 11 Art Unit: 2156 Application/Control Number: 18/563,046 Page 12 Art Unit: 2156