Prosecution Insights
Last updated: October 02, 2026
Application No. 18/319,472

DYNAMIC INTENT-BASED NETWORK COMPUTING JOB ASSIGNMENT USING REINFORCEMENT LEARNING

Final Rejection §101§103
Filed
May 17, 2023
Priority
May 19, 2022 — provisional 63/343,664
Examiner
LAI, DYLAN HONG
Art Unit
2144
Tech Center
2100 — Computer Architecture & Software
Assignee
NEC Laboratories America Inc.
OA Round
2 (Final)
Grant Probability
Favorable
3-4
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-55.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
13 currently pending
Career history
13
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§101 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Objections Claim 1 is objected to because of the following informalities: The text of any deleted matter must be shown by strike-through except that double brackets placed before and after the deleted characters may be used to show deletion of five or fewer consecutive characters, but not both. See 37 C.F.R. 1.121 (c)(2). The preamble of claim 1 is objected to because of the following informalities: “…for a mobile edge computing infrastructure …” is newly added material to the claim, but is not underlined. See 37 C.F.R. 1.121 (c)(2). Claim 1f) is objected to because of the following informalities: "wherein the current network states comprises remaining computing resources and remaining bandwidth resources of the mobile edge computing infrastructure;" is newly added material to the claim, but is not underlined. See 37 C.F.R. 1.121 (c)(2). Appropriate correction is required. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1- 4 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The analysis of the claims will follow the 2019 Revised Patent Subject Matter Eligibility Guidelines (“2019 PEG”). Step 1: Independent claim 1 (A dynamic, intent-based network computing job assignment method…) is directed towards a method. Therefore, this claim, as well as its dependent claims, are directed towards one of the four statutory categories (process, machine, manufacture, or composition of matter). Claim 1 Step 2A, Prong 1: The claim states, inter alia: a) defining a discounted cumulative reward function; This limitation recites a mental process using evaluation, judgment and opinion, with aid of pen and paper to think of an algorithm to reward successful actions. See MPEP 2106.04(a)(2)(III); b) defining an action space; This limitation recites a mental process using evaluation, judgment and opinion, with aid of pen and paper to think of possible actions. See MPEP 2106.04(a)(2)(III); g) … and predicting via the policy NN, a reward distribution over the action space; This limitation recites a mental process using evaluation, judgment and opinion, with aid of pen and paper to guess the rewards for a set of possible actions. See MPEP 2106.04(a)(2)(III); h) selecting an action that has a maximum predicted reward relative to other actions; This limitation recites a mental process using evaluation, judgment and opinion, with aid of pen and paper to judge which action has the greatest reward. See MPEP 2106.04(a)(2)(III); i) using the selected action, determining whether to accept or reject the request r based on availability of the remaining computing resources and the remaining bandwidth resources; This limitation recites a mental process using evaluation, judgment and opinion, with aid of pen and paper to evaluate if request r has the required available remaining resources and judge whether to accept or reject r based on that evaluation. Step 2A, Prong 2: The additional elements recited in the claim do not integrate the judicial exception into a practical application. Additional elements: c) creating a policy neural network (policy NN); This limitation is recited at a high level of generality and recites use of a generic computer algorithm as mere instructions to apply an abstract idea. Mere instructions that a judicial exception is to be applied using a generic computer algorithm in its ordinary capacity cannot integrate the judicial exception into a practical application. See MPEP 2106.05(f); d) creating a value neural network (value NN); This limitation is recited at a high level of generality and recites use of a generic computer algorithm as mere instructions to apply an abstract idea. Mere instructions that a judicial exception is to be applied using a generic computer algorithm in its ordinary capacity cannot integrate the judicial exception into a practical application. See MPEP 2106.05(f); e) receiving, via a computing network manager, a new intent-based computing job request r specifying a data size and a computing model; This limitation is an insignificant extra-solution activity of merely receiving data/data gathering. See MPEP 2106.05(g); f) adding the request r and a current network state s to a batch, wherein the current network state s comprises remaining computing resources and remaining bandwidth resources of the mobile edge computing infrastructure; This limitation is an insignificant extra-solution activity of selecting a particular data source or type of data to be manipulated. See MPEP 2106.05(g); g) using the request r and the current network state s as input to the policy NN created in step (c) … This limitation is recited at a high level of generality and recites use of generically recited data as input to a generic policy NN as mere instructions to apply an abstract idea. Mere instructions that a judicial exception is to be applied using generically recited data on a generic policy NN in its ordinary capacity cannot integrate the judicial exception into a practical application. See MPEP 2106.05(f); j) if the request r is accepted, deploying the selected action to the mobile edge computing infrastructure and updating the remaining computing resources and the remaining bandwidth resources; This limitation is recited at a high level of generality and recites deploying or applying an abstract idea to generic computer equipment. Mere instructions to apply an exception using generic computer equipment does not integrate the judicial exception into a practical application. See MPEP 2106.05(f); k) if a size of the batch is equal to a threshold, then: i) training the value NN; and ii) training the policy NN; These limitations are recited at a high level of generality and recites use of generic computer algorithms as mere instructions for applying an abstract idea. Mere instruction that a judicial exception is to be applied using generic computer algorithms in its ordinary capacity cannot integrate the judicial exception into a practical application. See MPEP 2106.05(f); l) if the size of the batch is less than the threshold, repeating steps (e) through (j); and This limitation is recited at a high level of generality and recites mere instructions to repeat an abstract idea in response to a generic condition. Mere instruction that a judicial exception is to be applied repetitively cannot integrate the judicial exception into a practical application. See MPEP 2106.05(f); m) outputting a resource allocation decision to at least one edge node. This limitation is an insignificant extra-solution activity of insignificant computer implementation. See MPEP 2106.05(g); Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional elements: c) creating a policy neural network (policy NN); This limitation is recited at a high level of generality and recites use of a generic computer algorithm as mere instructions to apply an abstract idea. Mere instructions that a judicial exception is to be applied using a generic computer algorithm in its ordinary capacity cannot integrate the judicial exception into a practical application. See MPEP 2106.05(f); d) creating a value neural network (value NN); This limitation is recited at a high level of generality and recites use of a generic computer algorithm as mere instructions to apply an abstract idea. Mere instructions that a judicial exception is to be applied using a generic computer algorithm in its ordinary capacity cannot integrate the judicial exception into a practical application. See MPEP 2106.05(f); e) receiving, via a computing network manager, a new intent-based computing job request r specifying a data size and a computing model; MPEP 2106.05(d)(II) indicates that merely receiving or transmitting data over a network is a well-understood, routine, and conventional function when it is claimed in a merely generic manner (as it is in the present claim); f) adding the request r and a current network state s to a batch, wherein the current network state s comprises remaining computing resources and remaining bandwidth resources of the mobile edge computing infrastructure; MPEP 2106.05(d)(II) indicates that merely gathering statistics is a well-understood, routine, and conventional function when it is claimed in a merely generic manner (as it is in the present claim); g) using the request r and the current network state s as input to the policy NN created in step (c) … This limitation is recited at a high level of generality and recites use of generically recited data as input to a generic policy NN as mere instructions to apply an abstract idea. Mere instructions that a judicial exception is to be applied using generically recited data on a generic policy NN in its ordinary capacity cannot integrate the judicial exception into a practical application. See MPEP 2106.05(f); j) if the request r is accepted, deploying the selected action to the mobile edge computing infrastructure and updating the remaining computing resources and the remaining bandwidth resources; This limitation is recited at a high level of generality and recites deploying or applying an abstract idea to generic computer equipment. Mere instructions to apply an exception using generic computer equipment does not integrate the judicial exception into a practical application. See MPEP 2106.05(f); k) if a size of the batch is equal to a threshold, then: i) training the value NN; and ii) training the policy NN; These limitations are recited at a high level of generality and recites use of generic computer algorithms as mere instructions for applying an abstract idea. Mere instruction that a judicial exception is to be applied using generic computer algorithms in its ordinary capacity cannot integrate the judicial exception into a practical application. See MPEP 2106.05(f); l) if the size of the batch is less than the threshold, repeating steps (e) through (j); and This limitation is recited at a high level of generality and recites mere instructions to repeat an abstract idea in response to a generic condition. Mere instruction that a judicial exception is to be applied repetitively cannot integrate the judicial exception into a practical application. See MPEP 2106.05(f); m) outputting a resource allocation decision to at least one edge node. MPEP 2106.06(d)(II) indicates that presenting offers (outputting a resource allocation decision) is well-understood, routine, and conventional when it is claimed in a merely generic manner (as it is in this limitation); Claim 2 Step 2A, Prong 1: This claim recites, inter alia: wherein the current network state s is represented as a one-dimensional array of size 2+V+E, where V is a number of fog nodes and E is a number of physical links in the mobile edge computing infrastructure. This limitation recites a mathematical concept of a mathematical relationship of organizing information about the current network state s through mathematical correlations. See MPEP 2106.04(a)(2)(I)(A); Step 2A, Prong 2: There are no further additional elements recited in this claim. Step 2B: There are no further additional elements recited in this claim. Claim 3 Step 2A, Prong 1: This claim recites, inter alia: wherein the action space comprises a plurality of actions defining a quantity of virtual nodes to use, mapping the virtual nodes to physical fog nodes, and establishing routing paths between the virtual nodes. This limitation recites a mental process using evaluation, judgment and opinion, with aid of pen and paper to decide a number of virtual nodes to use, which physical fog nodes the virtual nodes map to, and which other virtual nodes each of the virtual nodes connect to. See MPEP 2106.04(a)(2)(III); Step 2A, Prong 2: There are no further additional elements recited in this claim. Step 2B: There are no further additional elements recited in this claim. Claim 4 Step 2A, Prong 1: This claim recites, inter alia: wherein defining the discounted cumulative reward function comprises defining a reward of 1 if the request r is accepted and a reward of -1 if the request r is rejected. This limitation recites a mental process using evaluation, judgment and opinion, with aid of pen and paper to judge if a request is accepted or rejected and decide a reward for it based on that judgement. See MPEP 2106.04(a)(2)(III); Step 2A, Prong 2: There are no further additional elements recited in this claim. Step 2B: There are no further additional elements recited in this claim. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claim 1 is rejected under 35 U.S.C. 103 as being unpatentable over US 20210158162 A1 by Hafner et al., hereafter Hafner, in view of US 20240275691 A1 by Kattepur et al., hereafter Kattepur. Regarding independent claim 1, Hafner teaches: A dynamic, intent-based network computing job assignment method using reinforcement learning, the method comprising the steps of: ((Hafner) Paragraph [0021], “This specification describes a training system implemented as computer programs on one or more computers in one or more locations for training a policy neural network that can be used to control a reinforcement learning agent interacting with an environment by, at each of multiple time steps, processing a policy network input derived from data characterizing the current state of the environment at the time step (i.e., an “observation”) to generate an action selection output specifying an action to be performed by the agent.”) a) defining a discounted cumulative reward function; ((Hafner) Paragraph [0027], “To generate each trajectory of latent representations, the system 100 makes use of a reward neural network 130, a value neural network 140, and a transition neural network 150.” A reward neural network is a discounted cumulative reward function.) b) defining an action space; ((Hafner) Paragraph [0021], “…to generate an action selection output specifying an action to be performed by the agent.” An action space is an action selection output.) c) creating a policy neural network (policy NN); ((Hafner) Paragraph [0021], “…training a policy neural network…”) d) creating a value neural network (value NN); ((Hafner) Paragraph [0027], “…a value neural network 140…”) e) receiving, via a computing network manager, a new intent-based computing job request r ((Hafner) Paragraph [0106], “For example, at certain points during the training process, e.g., upon determining that a predetermined number of iterations of processes 400 and 500 have been performed, the system can use the policy neural network and the representation neural network and in accordance with current values of their parameters (as of the iteration) to control the agent to perform a sequence of actions when interacting with the environment, i.e., by generating a latent representation from each new observation and then selecting an action to be performed by the agent in response to the new observation using the action selection output generated from the latent representation.” The system is a computing network manager. Controlling an agent to perform a sequence of actions is an intent-based computing job request.) specifying a data size and a computing model; ((Hafner) Algorithm 1: Dreamer; The algorithm contains batch size, which is specifying a data size, and neural network parameters and model components, which are together specifying a computing model.) f) adding the request r and a current network state s to a batch, wherein the current network state s comprises remaining computing resources and remaining bandwidth resources of the mobile edge computing infrastructure; ((Hafner) Paragraph [0025], "Once trained, the representation neural network 110 and the policy neural network 120 can be deployed and used to control the agent interacting with the environment, i.e., by generating a latent representation from each new observation and then selecting an action to be performed by the agent in response to the new observation using the action selection output generated from the latent representation." A representation neural network is a current network state and controlling an agent is a new request r.) g) using the request r and the current network state s as input to the policy NN created in step (c) and predicting, via the policy NN, a reward distribution over the action space; h) selecting an action that has a maximum predicted reward relative to other actions; ((Hafner) Paragraph [0039], “Either during or after the dynamics learning of the system, the training engine 160 trains, by using reinforcement learning techniques, the policy neural network 120 and the value neural network 130 on the “imagined” trajectory data generated using the representation, reward and transition neural networks and based on processing information contained in the training tuple set. In particular, the training engine 160 trains the policy neural network 120 to generate action selection outputs that can be used to select actions that maximize a cumulative measure of rewards received by the agent and that cause the agent to accomplish an assigned task.” Using a representation and processing information contained in a training tuple set is using a current network state and a request. These are used as input to a policy neural network to generate action selection outputs (an action space) that select actions that maximize cumulative reward which means it must also predict the reward distribution. In addition, the agent accomplishing an assigned task is deploying an action.) k) if a size of the batch is equal to a threshold, then: i) training the value NN; and ii) training the policy NN; ((Hafner) Paragraph [0033], “A training engine 160 can train the neural networks to determine, e.g., from initial values, trained values of the network parameters 158, including respective network parameters of the representation neural network 110, policy neural network 120, value neural network 130, reward neural network 140, and transition neural network 150.”) l) if the size of the batch is less than the threshold, repeating steps (e) through (j); ((Hafner) Paragraph [0088], “In general, the system can repeatedly perform the process 400 until a termination criterion is reached, e.g., after the process 400 have been performed a predetermined number of times” The process being performed a predetermined number of times is the size of the batch being greater than a predetermined threshold.) and m) outputting a resource allocation decision to at least one edge node. ((Hafner) Paragraph [0021], “…to generate an action selection output specifying an action to be performed by the agent.” An action selection is a resource allocation decision. An agent is an edge node. Hafner also teaches: j) […] deploying the selected action to the mobile edge computing infrastructure and updating the remaining computing resources […] ((Hafner) Paragraph [0039], “…that cause the agent to accomplish an assigned task”; Paragraph [0066], “In some further applications, the environment is a real-world environment and the agent manages distribution of tasks across computing resources, e.g., on a mobile device” Causing the agent to accomplish an assigned task is deploying the selected action to the agent. The agent managing distribution of tasks across computing resources on a mobile device is updating computing resources and is on a mobile edge computing infrastructure.) However, Hafner does not explicitly disclose: i) using the selected action, determining whether to accept or reject the request r based on availability of the remaining computing resources and the remaining bandwidth resources; j) if the request r is accepted, […] updating […] the remaining bandwidth resources; Kattepur teaches: Using the selected action, determining to accept or reject the request, ((Kattepur) Paragraph [0116] “…after generating an explanation for the queried action at step 356, the training node sends the explanation to the entity in step 358. In step 360, the training node receives from the entity feedback on the explanation for the queried action, the feedback comprising at least one of acceptance or rejection of the explanation for the queried action.” A request is being interpreted as including an explanation for a queried action from Kattepur.), based on availability of the remaining computing resources and the remaining bandwidth resources; ((Kattepur) Fig. 3b.; Fig. 3c.; Paragraph [0050], [0054], [0069] “Example observations obtained at step 340 may include: …a current network resource allocation … Value of throughput observed on a network link…” An observation of a current network resource allocation used to generate an updated belief state of the environment is remaining computing resources. An observation of a value of throughput observed on a network link is remaining bandwidth resources.) and updating remaining computing resources and remaining bandwidth resources. ((Kattepur) Paragraph [0147], “Referring to FIG. 6, the POMDP agent interacts with the environment, and receives observations and rewards at each time step. The POMDP agent uses this information to update its internal beliefs.” Receiving observations at each time steps is updating remaining computing resources and remaining bandwidth resources.) Kattepur also teaches explanation feedback (accepting and rejecting requests) improves the accuracy of the belief states (action space) and assists with adapting a policy to changes in an environment. Kattepur is in the analogous art of using reinforcement learning with a policy to predict and select actions. Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Kattepur with the teachings of Hafner. One having ordinary skill in the art would have been motivated to combine the feedback mechanism (using a selected action to accept or reject a request) using observations of current network resource allocation and throughput then reobtaining those observations as in Kattepur with the method of defining a reward function, and an action space, creating a policy neural network and value neural network, adding a request and current network state to a batch, using the request and network sate as input to the policy neural network, predicting a reward distribution over the action space, selecting an action that has a maximum predicted reward relative to other actions, and deploying the action as in Hafner in order to improve the accuracy of the action space and assist with adapting the policy to changes in an environment. This combination would cause the predictable result that is the invention specified in claim 1 of the instant application. The Examiner notes that this motivation applies to all dependent and/or otherwise subsequently addressed claims. Claim 2 is rejected under 35 U.S.C. 103 as being unpatentable over Hafner, in view of Kattepur, in further view of US 20220020026 A1 by Wadhwa et al., hereafter Wadhwa. Regarding Claim 2, Hafner teaches the material disclosed in claim 1, and additionally teaches: The current network state being represented as a one dimensional-array. ((Hafner) Paragraph [0021], “In particular, the training system trains the policy network on latent representations generated by a representation neural network from observations. As used throughout this document, a “latent representation” of an observation refers to a representation of a state of the environment as an ordered collection of numerical values, e.g., a vector or matrix of numerical values, and generally has a lower dimensionality than the observation itself.” A vector is a one-dimensional array.) Hafner does not explicitly disclose: wherein the current network state s is represented as a one- dimensional array of size 2+V+E, where V is a number of fog nodes and E is a number of physical links in the mobile edge computing infrastructure. Wadhwa does teach: wherein the current network state s is represented as a one- dimensional array of size 2+V+E, where V is a number of fog nodes and E is a number of physical links in the mobile edge computing infrastructure. ((Wadhwa) Paragraph [0029], “The server system is configured to compute a first vector representation associated with each node of temporal knowledge graph using the node embedding algorithm. The server is also configured to compute second and third vector representations associated with each edge and sub-graph of the temporal knowledge graph using the edge embedding and the subtree graph embedding algorithms, respectively. Additionally, the server system is configured to aggregate the first, second and third vector representations for generating the graph embedding vector.” The graph embedding vector is a one-dimensional array. The first vector has size V as it uses node embedding. The second vector has size E as physical links can be represented by edges and the second vector uses edge embedding. The third vector could have size 2. An aggregate sums the size of the three vectors into the graph embedding vector with size V+E+2.) Wadhwa and Hafner are analogous art because they are in the same area of invention, that being neural networks containing a presentation of the environment. Thus, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention.to have substituted the latent representation vector as taught by Hafner with the graph embedding vector, as taught by Wadhwa. The motivation would have been to convert graph data into a low dimensional space so that graph structural information and graph properties are preserved. This simple substitution of a specific one-dimensional vector taught by Wadhwa into the vector taught by Hafner would obtain the predictable result of the invention specified in claim 2. Claim 3 is rejected under Hafner, in view of Kattepur, in further view of US 20190215381 A1 by Mukund et al., hereafter Mukund. Regarding claim 3, Hafner, in view of Kattepur, teaches the material disclosed in claim 1. Hafner, in view of Kattepur, does not explicitly disclose, but together with Mukund does teach: wherein the action space comprises a plurality of actions defining a quantity of virtual nodes to use ((Mukund) Paragraph [0020], “For example, the application on the cloud may send compute statistics to the orchestrator, which in turn determines an amount of compute resources necessary to handle compute requests. In such a case, the orchestrator application may trigger an instance of the application to be installed on the fog server. The mapping server then assigns the application instance with the EID (i.e., the IP address) of the underlying application on the cloud.” EID are virtual nodes and a quantity are assigned based on compute resources), mapping the virtual nodes to physical fog nodes ((Mukund) Paragraph [0030], “The map server application 134 maintains EID-to-RLOC mappings in a mapping database.” EID are virtual nodes and RLOC are physical fog nodes), and establishing routing paths between the virtual nodes ((Mukund) Paragraph [0041], “Packet traffic for the application 107 that routes through a fog network other than fog network 110 may still be processed at the central cloud network 105.” Packet traffic routing through a central cloud network are routing paths between virtual nodes.). Mukund and Hafner are analogous art because they are in the same area of invention, that being observing a computing environment around mobile edge devices to efficiently allocate computing resources. Thus, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the teachings of Mukund and Hafner before him or her, to have modified the action space as taught by Hafner to include defining a number of virtual node instances to use based on amount of compute resources available, mapping EID virtual node instances to RLOC physical fog nodes, and processing packet traffic routing paths between virtual nodes. The motivation for this would be to efficiently determine when services should be accessed by fog networks and cloud networks, as well as to efficiently determine the appropriate destination for routing traffic. This simple substitution of the definition of the action space in Hafner’s system with the action space defined above from Mukund would obtain the predictable result of the invention specified in claim 3 of the instant application. Claim 4 is rejected under Hafner, in view of Kattepur, in further view of Parallel reward and punishment control in humans and robots: Safe reinforcement learning using the MaxPain algorithm by Elfwing and Seymour, hereafter Elfwing. Regarding claim 4, Hafner, in view of Kattepur, teaches the material disclosed in claim 1. Hafner, in view of Kattepur, does not explicitly disclose: wherein defining the discounted cumulative reward function comprises defining a reward of 1 if the request r is accepted and a reward of -1 if the request r is rejected. Elfwing does teach: wherein defining the discounted cumulative reward function comprises defining a reward of 1 if the request r is accepted and a reward of -1 if the request r is rejected. ((Elfwing) Section II. Equations 2 and 3, “In the MaxPain algorithm, the standard reward R is separated into two parts, the positive reward r≥0: r = max(R,0) and the pain (or punishment) p≥0: p = -min(R,)” The positive reward r is the reward for a positive action such as the request being accepted and r can be 1, making the reward 1. The punishment p is a negative reward for a negative action such as when the request is rejected and p could be 1, making the reward -1.) Elfwing and Hafner are analogous art because they are from the same area of invention, that being reinforcement learning. Thus, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have modified the reward function as taught by Hafner to include the positive reward of 1 if a request is accepted and the punishment negative reward of -1 if a request is rejected as taught by Elfwing. The motivation for this would be to balance the punishment and reward systems and make them separable which allows safer exploration, as well as effective learning and near-optimal long-term performance. This simple substitution of separable positive and negative reward systems taught by Elfwing into the reward system as taught by Hafner would obtain the predictable result of the invention specified in claim 4 of the instant application. Response to Arguments Applicant’s arguments, see page 6 lines 2-17, filed 05/11/2026, with respect to the rejection(s) of claim(s) 1 under 35 U.S.C. 10 after amendment have been fully considered and are not persuasive. The amended limitation of requiring a mobile edge computing infrastructure is mere instructions to apply an abstract idea using a generic mobile edge computing infrastructure and the definition of the network state as comprising remaining computing resources. The amendment to the limitation, adding the request r and a current network state s to a batch, adds defining the current network state s as comprising remaining computing resources and remaining bandwidth resources. This limitation is still mere extra-solution activity, specifically pre-solution activity of selecting the particular data source or type to be manipulated (selecting information, based on remaining/available computing and bandwidth resources, to be analyzed; see MPEP 2106(g)) and is gathering information/statistics on available resources which is understood to be well-understood, routine, and conventional when claimed in a generic manner as it is in this limitation (see MPEP 2106.05(d)(II)). Mere instructions to apply an abstract idea using generic mobile edge computing infrastructure does not integrate the abstract idea into a practical application and is not significantly more than the judicial exception when they are claimed in a generic manner, as they are in this limitation. Selecting a particular data source or type to be manipulated does not integrate the abstract idea into a practical application when it is claimed in a generic manner, as it is in this limitations, while gathering information on available resources, as a well-understood, routine, and conventional function when claimed in a generic manner as it is in this limitation, does not amount to significantly more than the judicial exception. In addition, the limitations of claim 1 make no mention of automated decision making. Applicant’s arguments, see page 7 lines 3-8, filed 05/11/2026, with respect to the rejection(s) of claim(s) 1 under 35 U.S.C. 103 after amendment have been fully considered and are not persuasive. Kattepur teaches, in Paragraph [0147](“Referring to FIG. 6, the POMDP agent interacts with the environment, and receives observations and rewards at each time step. The POMDP agent uses this information to update its internal beliefs”), updating computing and bandwidth resources. Determining to accept or reject a request is shown through Kattepur paragraph [0116] and the determination being based on remaining computing and bandwidth resources is shown in Kattepur through Fig. 3b., Fig. 3c., Paragraph [0050], [0054], and [0069] “Example observations obtained at step 340 may include: …a current network resource allocation … Value of throughput observed on a network link…”. Claim 1 makes no mention of automated determination, so a determination to accept or reject a computing job request by human intent is appropriate. In response to applicant’s argument that there is no teaching, suggestion, or motivation to combine the references, the examiner recognizes that obviousness may be established by combining or modifying the teachings of the prior art to produce the claimed invention where there is some teaching, suggestion, or motivation to do so found either in the references themselves or in the knowledge generally available to one of ordinary skill in the art. See In re Fine, 837 F.2d 1071, 5 USPQ2d 1596 (Fed. Cir. 1988), In re Jones, 958 F.2d 347, 21 USPQ2d 1941 (Fed. Cir. 1992), and KSR International Co. v. Teleflex, Inc., 550 U.S. 398, 82 USPQ2d 1385 (2007). In this case, Kattepur discloses in paragraph [0012], "The explanation feedback also assists with adapting a policy to changes in an environment." This teaching gives the motivation to add the explanation feedback method taught by Kattepur to the hardware resource allocation method in order to adapt to changes in the fog computing environment when solving hardware-bound resource allocation problems. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Patents and/or related publications are cited in the Notice of References Cited (Form PTO-892) attached to this action to further show the state of the art with respect to reinforcement learning, policy networks, value networks, and fog node communication systems. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to DYLAN H LAI whose telephone number is (571)272-8628. The examiner can normally be reached Monday - Friday 7:30am-5:00pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Tamara Kyle can be reached at 5712524241. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. D.H.L. Examiner Art Unit 2144 /TAMARA T KYLE/ Supervisory Patent Examiner, Art Unit 2144
Read full office action

Prosecution Timeline

May 17, 2023
Application Filed
Mar 27, 2026
Non-Final Rejection mailed — §101, §103
May 11, 2026
Response Filed
Jun 03, 2026
Final Rejection (signed) — §101, §103
Aug 05, 2026
Final Rejection mailed — §101, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
Grant Probability
Moderate
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month