Prosecution Insights
Last updated: October 02, 2026
Application No. 18/562,537

INFORMATION PROCESSING APPARATUS, INFORMATION PROCESSING METHOD, INFORMATION PROCESSING SYSTEM, AND STORAGE MEDIUM

Non-Final OA §101§103§112
Filed
Nov 20, 2023
Priority
May 26, 2021 — nonprovisional of PCTJP2021020000
Examiner
HOOVER, BRENT JOHNSTON
Art Unit
Tech Center
Assignee
NEC Corporation
OA Round
1 (Non-Final)
82%
Grant Probability
Favorable
1-2
OA Rounds
7m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 82% — above average
82%
Career Allowance Rate
309 granted / 376 resolved
+22.2% vs TC avg
Strong +22% interview lift
Without
With
+22.0%
Interview Lift
resolved cases with interview
Typical timeline
3y 5m
Avg Prosecution
29 currently pending
Career history
399
Total Applications
across all art units

Statute-Specific Performance

§101
30.7%
-9.3% vs TC avg
§103
37.9%
-2.1% vs TC avg
§102
11.1%
-28.9% vs TC avg
§112
16.7%
-23.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 376 resolved cases

Office Action

§101 §103 §112
CTNF 18/562,537 CTNF 93954 DETAILED ACTION 07-03-aia AIA 15-10-aia The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA. This action is responsive to the original application filed on 11/20/2023. Acknowledgment is made with respect to a claim of priority to PCT Application PCT/JP2021/020000 filed on 5/26/2021. Specification 06-11 AIA The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed. 06-11-01 AIA The following title is suggested: INFORMATION PROCESSING APPARATUS, METHOD, AND SYSTEM FOR PREDICTING A REWARD SUM USING REINFORCEMENT LEARNING Claim Rejections - 35 USC § 112 07-30-02 AIA The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. 07-34-01 Claims 1-11 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 1 recites the limitation “in the determination process, the at least one processor calculating a first function, which predicts a reward sum from a state and an action ” (emphasis added). It is not clear if the claimed “a state” and “an action” in this limitation are the same state and action that were introduced in the previous limitation. Please explain. For examination purposes, the limitation will be interpreted to mean “in the determination process, the at least one processor calculating a first function, which predicts a reward sum from [[a]] the state and [[an]] the action ” (emphasis added). Dependent claims 2-8 depend on indefinite claim 1 , and are also rejected under 35 USC § 112(b) by virtue of this dependency. Independent claims 9 and 11 contain the same indefiniteness issue and are also rejected under 35 USC § 112(b) for the same reasons as claim 1 . Dependent claim 10 depends on indefinite claim 9 , and is also rejected under 35 USC § 112(b) by virtue of this dependency Appropriate correction is required. Claim Rejections - 35 USC § 101 07-04-01 AIA 07-04 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-11 are rejected under 35 U.S.C 101 because the claimed invention is directed to an abstract idea without significantly more. The analysis of the claims will follow the 2019 Revised Patent Subject Matter Eligibility Guidance, 84 Fed. Reg. 50 (“2019 PEG”). Claim 1 Step 1 : The claim recites an apparatus; therefore, it is directed to the statutory category of a machine. Step 2A Prong 1 : The claim recites, inter alia: a determination process of determining an action with reference to the state: Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mental process of determining an action with reference to a state, which is an evaluation or observation that is practically capable of being performed in the human mind with the assistance of pen and paper. For example, one can practically and mentally determine that a ball is being throw by observing the state or position of an arm. in the determination process, the at least one processor calculating a first function, which predicts a reward sum from a state and an action, by carrying out weighting for the learning data: Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mathematical concept of calculating a function by weighting learning data, which is performed through mathematical computation as evidenced by paragraph [0107] pf the originally filed specification. determining the action with use of the first function: Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mental process of determining an action with reference to a function, which is an evaluation or observation that is practically capable of being performed in the human mind with the assistance of pen and paper. For example, one can practically and mentally determine that a ball is being throw looking at a function or some sort of relationship in data. Step 2A Prong 2 : The claim does not recite any additional limitations which integrate the abstract idea into a practical application. Specifically, the additional elements consist of “ at least one processor, the at least one processor carrying out ”, “ an acquisition process of acquiring a state ”, “ an accumulation process of accumulating learning data including (i) the state and (ii) a reward obtained by the action which has been determined in the determination process ”, and “ the at least one processor ”. The additional elements of “ at least one processor, the at least one processor carrying out ” and “ the at least one processor ” amount to generic computer components used as a tool to perform an existing process. Thus, the additional elements amount to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer ( see MPEP § 2106.05(f)). The additional elements “ an acquisition process of acquiring a state ” and “ an accumulation process of accumulating learning data including (i) the state and (ii) a reward obtained by the action which has been determined in the determination proces s” are insignificant extra-solution activities required for any uses of the abstract ideas ( see MPEP § 2106.05(g)). Thus, even when viewed individually and as an ordered combination, these additional elements do not integrate the abstract idea into a practical application and the claim is thus directed to the abstract idea. Step 2B : Finally, the claim taken as a whole does not contain an inventive concept which provides significantly more than the abstract idea. The additional elements of “ at least one processor, the at least one processor carrying out ” and “ the at least one processor ” amount to generic computer components used as a tool to perform an existing process. Thus, the additional elements amount to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer ( see MPEP § 2106.05(f)). The additional elements “ an acquisition process of acquiring a state ” and “ an accumulation process of accumulating learning data including (i) the state and (ii) a reward obtained by the action which has been determined in the determination proces s” are insignificant extra-solution activities required for any uses of the abstract ideas ( see MPEP § 2106.05(g)), and are well-understood, routine, conventional activities (see MPEP § 2106.05(d)(II)(i); “Receiving or transmitting data over a network”). Taken alone or in combination, the additional elements of the claim do not provide an inventive concept and thus the claim is subject-matter ineligible. Claim 2 Step 1 : A machine, as above. Step 2A Prong 1 : The claim recites, inter alia: calculates an index related to variability from one or more values included in the learning data: Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mathematical concept of calculating an index related to variability, which is performed through mathematical computation. calculates the first function by applying a smaller weighting factor to the one or more values as the calculated index related to the variability is higher: Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mathematical concept of calculating a function by applying weighting factors, which is performed through mathematical computation. Step 2A Prong 2, Step 2B : The claim does not recite any additional elements that are sufficient to integrate the judicial exceptions into a practical application or amount to significantly more than the judicial exception. As such, the claim is ineligible. Claim 3 Step 1 : A machine, as above. Step 2A Prong 1 : The claim recites, inter alia: calculates variance of a second function as the index related to the variability, the variance of the second function being obtained with reference to the state and the action: Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mathematical concept of calculating a variance of a function, which is performed through mathematical computation. Step 2A Prong 2, Step 2B : The claim does not recite any additional elements that are sufficient to integrate the judicial exceptions into a practical application or amount to significantly more than the judicial exception. As such, the claim is ineligible. Claim 4 Step 1 : A machine, as above. Step 2A Prong 1 : The claim recites the abstract ideas of the preceding claims from which it depends. Step 2A Prong 2, Step 2B : The additional element “ a display process of displaying (i) at least one selected from the group consisting of the state, the action, the reward, and a value of the first function and (ii) the index related to the variability ” is an insignificant extra-solution activity required for any uses of the abstract ideas ( see MPEP § 2106.05(g)), and is a well-understood, routine, conventional activity (see MPEP § 2106.05(d)(II)(i); “Presenting offers and gathering statistics”; and see Electric Power Group, LLC v. Alstom, S.A., 830 F.3d 1350, 119 USPQ2d 1739 (Fed. Cir. 2016) at pages 9-10: “The claims at issue do not require any nonconventional computer, network, or display components, or even a “non-conventional and non-generic arrangement of known, conventional pieces,” but merely call for performance of the claimed information collection, analysis, and display functions “on a set of generic computer components” and display devices ). Taken alone or in combination, the additional elements of the claim do not provide an inventive concept, integrate the abstract ideas into a practical application, or provide significantly more than the abstract ideas of the claim and thus the claim is subject-matter ineligible. Claim 5 Step 1 : A machine, as above. Step 2A Prong 1 : The claim recites the abstract ideas of the preceding claims from which it depends. Step 2A Prong 2, Step 2B : The additional element “ displays, in an emphasized manner, a value which is of the at least one selected from the state, the action, the reward, and the value of the first function and which satisfies that the index related to the variability is equal to or less than a threshold value ” is an insignificant extra-solution activity required for any uses of the abstract ideas ( see MPEP § 2106.05(g)), and is a well-understood, routine, conventional activity (see MPEP § 2106.05(d)(II)(i); “Presenting offers and gathering statistics”; and see Electric Power Group, LLC v. Alstom, S.A., 830 F.3d 1350, 119 USPQ2d 1739 (Fed. Cir. 2016) at pages 9-10: “The claims at issue do not require any nonconventional computer, network, or display components, or even a “non-conventional and non-generic arrangement of known, conventional pieces,” but merely call for performance of the claimed information collection, analysis, and display functions “on a set of generic computer components” and display devices ). Taken alone or in combination, the additional elements of the claim do not provide an inventive concept, integrate the abstract ideas into a practical application, or provide significantly more than the abstract ideas of the claim and thus the claim is subject-matter ineligible. Claim 6 Step 1 : A machine, as above. Step 2A Prong 1 : The claim recites, inter alia: calculates the first function with use of a feature map that maps the state and the action to a vector: Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mathematical concept of calculating a function using a feature map, which is performed through mathematical computation. Step 2A Prong 2, Step 2B : The claim does not recite any additional elements that are sufficient to integrate the judicial exceptions into a practical application or amount to significantly more than the judicial exception. As such, the claim is ineligible. Claim 7 Step 1 : A machine, as above. Step 2A Prong 1 : The claim recites, inter alia: selects an action that maximizes the first function which includes, as an argument, the state acquired in the acquisition process: Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mental process of selecting an action that maximizes a function, which is an evaluation or observation that is practically capable of being performed in the human mind with the assistance of pen and paper. Step 2A Prong 2, Step 2B : The claim does not recite any additional elements that are sufficient to integrate the judicial exceptions into a practical application or amount to significantly more than the judicial exception. As such, the claim is ineligible. Claim 8 Step 1 : A machine, as above. Step 2A Prong 1 : The claim recites the abstract ideas of the preceding claims from which it depends. Step 2A Prong 2, Step 2B : The additional element of “ an input device ” amounts to a generic computer component used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer ( see MPEP § 2106.05(f)). The additional element “ accept the state and the reward ” is an insignificant extra-solution activity required for any uses of the abstract ideas ( see MPEP § 2106.05(g)), and is a well-understood, routine, conventional activity (see MPEP § 2106.05(d)(II)(i); “Receiving or transmitting data over a network”). Taken alone or in combination, the additional elements of the claim do not provide an inventive concept, integrate the abstract ideas into a practical application, or provide significantly more than the abstract ideas of the claim and thus the claim is subject-matter ineligible. Claim 9 Step 1 : The claim recites a method; therefore, it is directed to the statutory category of a process. Step 2A Prong 1 : The claim recites, inter alia: determining an action with reference to the state: Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mental process of determining an action with reference to a state, which is an evaluation or observation that is practically capable of being performed in the human mind with the assistance of pen and paper. For example, one can practically and mentally determine that a ball is being throw by observing the state or position of an arm. wherein, in the step of determining the action, the action is determined with use of a first function: Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mental process of determining an action with reference to a function, which is an evaluation or observation that is practically capable of being performed in the human mind with the assistance of pen and paper. For example, one can practically and mentally determine that a ball is being throw looking at a function or some sort of relationship in data. the first function predicting a reward sum from a state and an action and being calculated by carrying out weighting for the learning data: Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mathematical concept of calculating a function by weighting learning data, which is performed through mathematical computation as evidenced by paragraph [0107] pf the originally filed specification. Step 2A Prong 2 : The claim does not recite any additional limitations which integrate the abstract idea into a practical application. Specifically, the additional elements consist of “ an information processing apparatus ”, “ acquiring a state ”, “ accumulating learning data including (i) the state and (ii) a reward obtained by the determined action ”, and “ the information processing apparatus ”. The additional elements of “ an information processing apparatus ” and “ the information processing apparatus ” amount to generic computer components used as a tool to perform an existing process. Thus, the additional elements amount to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer ( see MPEP § 2106.05(f)). The additional elements “ acquiring a state ” and “ accumulating learning data including (i) the state and (ii) a reward obtained by the determined action ” are insignificant extra-solution activities required for any uses of the abstract ideas ( see MPEP § 2106.05(g)). Thus, even when viewed individually and as an ordered combination, these additional elements do not integrate the abstract idea into a practical application and the claim is thus directed to the abstract idea. Step 2B : Finally, the claim taken as a whole does not contain an inventive concept which provides significantly more than the abstract idea. The additional elements of “ an information processing apparatus ” and “ the information processing apparatus ” amount to generic computer components used as a tool to perform an existing process. Thus, the additional elements amount to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer ( see MPEP § 2106.05(f)). The additional elements “ acquiring a state ” and “ accumulating learning data including (i) the state and (ii) a reward obtained by the determined action ” are insignificant extra-solution activities required for any uses of the abstract ideas ( see MPEP § 2106.05(g)), and are well-understood, routine, conventional activities (see MPEP § 2106.05(d)(II)(i); “Receiving or transmitting data over a network”). Taken alone or in combination, the additional elements of the claim do not provide an inventive concept and thus the claim is subject-matter ineligible. Claim 10 Step 1 : The claim recites a computer-readable non-transitory storage medium; therefore, it is directed to the statutory category of a manufacture. Step 2A Prong 1 : The claim recites the abstract ideas of claim 1. Step 2A Prong 2, Step 2B : The additional element of “ a program causing a computer to function as the information processing, apparatus according to claim 1, the program causing the computer to carry out the acquisition process, the determination process, and the accumulation process ” amounts to generic computer components used as tools to perform an existing process. Thus, the additional elements amount to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer ( see MPEP § 2106.05(f)). Taken alone or in combination, the additional elements of the claim do not provide an inventive concept, integrate the abstract ideas into a practical application, or provide significantly more than the abstract ideas of the claim and thus the claim is subject-matter ineligible. Claim 11 Step 1 : The claim recites an information processing system; therefore, it is directed to the statutory category of a machine. Step 2A Prong 1 : The claim recites, inter alia: a determination process of determining an action with reference to the state: Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mental process of determining an action with reference to a state, which is an evaluation or observation that is practically capable of being performed in the human mind with the assistance of pen and paper. For example, one can practically and mentally determine that a ball is being throw by observing the state or position of an arm. in the determination process, the at least one processor calculating a first function, which predicts a reward sum from a state and an action, by carrying out weighting for the learning data: Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mathematical concept of calculating a function by weighting learning data, which is performed through mathematical computation as evidenced by paragraph [0107] pf the originally filed specification. determining the action with use of the first function: Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mental process of determining an action with reference to a function, which is an evaluation or observation that is practically capable of being performed in the human mind with the assistance of pen and paper. For example, one can practically and mentally determine that a ball is being throw looking at a function or some sort of relationship in data. Step 2A Prong 2 : The claim does not recite any additional limitations which integrate the abstract idea into a practical application. Specifically, the additional elements consist of “ an information processing apparatus and a terminal apparatus, wherein the information processing apparatus comprises at least one processor, the at least one processor carrying out ”, “ an acquisition process of acquiring a state ”, “ an accumulation process of accumulating learning data including (i) the state and (ii) a reward obtained by the action which has been determined in the determination process ”, “ the at least one processor ”, “ the terminal apparatus comprises at least one processor, the at least one processor of the terminal apparatus carrying out ”, “ a state information provision process of acquiring a state and providing the state to the information processing apparatus ”, and “ a reward information provision process of providing, to the information processing apparatus, reward information indicative of a reward obtained b y executing an action which has been determined by the information processing apparatus ”. The additional elements of “ an information processing apparatus and a terminal apparatus, wherein the information processing apparatus comprises at least one processor, the at least one processor carrying out ”, “ the at least one processor ”, and “ the terminal apparatus comprises at least one processor, the at least one processor of the terminal apparatus carrying out ” amount to generic computer components used as a tool to perform an existing process. The additional element “ executing an action which has been determined by the information processing apparatus ” amount to reciting only the idea of a solution or outcome i.e., the claim fails to recite details of how a solution to a problem is accomplished because it is not clear how the unspecific action is broadly executed. Thus, the additional elements amount to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer ( see MPEP § 2106.05(f)). The additional elements “ an acquisition process of acquiring a state ”, “ an accumulation process of accumulating learning data including (i) the state and (ii) a reward obtained by the action which has been determined in the determination proces s”, “ a state information provision process of acquiring a state and providing the state to the information processing apparatus ”, and “ a reward information provision process of providing, to the information processing apparatus, reward information indicative of a reward ” are insignificant extra-solution activities required for any uses of the abstract ideas ( see MPEP § 2106.05(g)). Thus, even when viewed individually and as an ordered combination, these additional elements do not integrate the abstract idea into a practical application and the claim is thus directed to the abstract idea. Step 2B : Finally, the claim taken as a whole does not contain an inventive concept which provides significantly more than the abstract idea. ng apparatus and a terminal apparatus, wherein the information processing apparatus comprises at least one processor, the at least one processor carrying out ”, “ the at least one processor ”, and “ the terminal apparatus comprises at least one processor, the at least one processor of the terminal apparatus carrying out ” amount to generic computer components used as a tool to perform an existing process. The additional element “ executing an action which has been determined by the information processing apparatus ” amount to reciting only the idea of a solution or outcome i.e., the claim fails to recite details of how a solution to a problem is accomplished because it is not clear how the unspecific action is broadly executed. Thus, the additional elements amount to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer ( see MPEP § 2106.05(f)). The additional elements “ an acquisition process of acquiring a state ”, “ an accumulation process of accumulating learning data including (i) the state and (ii) a reward obtained by the action which has been determined in the determination proces s”, “ a state information provision process of acquiring a state and providing the state to the information processing apparatus ”, and “ a reward information provision process of providing, to the information processing apparatus, reward information indicative of a reward ” are insignificant extra-solution activities required for any uses of the abstract ideas ( see MPEP § 2106.05(g)), and are well-understood, routine, conventional activities (see MPEP § 2106.05(d)(II)(i); “Receiving or transmitting data over a network”). Taken alone or in combination, the additional elements of the claim do not provide an inventive concept and thus the claim is subject-matter ineligible. Claim Rejections - 35 USC § 103 07-06 AIA 15-10-15 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. 07-20-aia AIA The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-3 and 6-11 are rejected under 35 U.S.C. § 103 as being obvious over van Hasselt et al. (US 20170076201 A1, hereinafter “VH”) in view of Cini et al. (Cini et al., “Deep Reinforcement Learning with Weighted Q-Learning”, Mar. 30, 2020, arXiv:2003.09280v2, pp. 1-14, hereinafter “Cini”). Regarding claim 1 , VH discloses [a]n information processing apparatus: comprising at least one processor, the at least one processor carrying out: ([0012-0013]; “This specification generally describes a reinforcement learning system that selects actions to be performed by a reinforcement learning agent interacting with an environment. …In some implementations, the environment is a simulated environment and the agent is implemented as one or more computer programs interacting with the simulated environment” ) an acquisition process of acquiring a state; ([0035]; “The system receives a current observation characterizing the current state of the environment (step 202 ).” ; and Figure 2, Element 202) a determination process of determining an action with reference to the state; and ([0037]; “The system selects an action to be performed by the agent in response to the current observation using the estimated future cumulative rewards (step 206 ).” ; and Figure 2, Element 206) an accumulation process of accumulating learning data including (i) the state and (ii) a reward obtained by the action which has been determined in the determination process, ([0041]; “The system generates an experience tuple that includes the current observation, the selected action, the reward, and the next observation and stores the generated experience tuple in a replay memory for use in training the Q network (step 208 )”, wherein the experience tuple comprises the learning or training data that includes state and reward information ; and Figure 2, Element 208) in the determination process, the at least one processor calculating a first function, which predicts a reward sum from a state and an action, ([0019]; “he Q network 110 is a deep neural network that is configured to receive as input an input observation and an input action and to generate an estimated future cumulative reward from the input in accordance with a set of parameters”, which discloses the first function or NN that predicts a reward sum or cumulative reward from a state and action ) … and determining the action with use of the first function ([0022]; “The reinforcement learning system 100 can then select the action having the highest estimated future cumulative reward as the action to be performed by the agent 102 in response to the observation” ). VH fails to explicitly disclose but Cini discloses by carrying out weighting for the learning data (Page 4, Algorithm 1; “Execute at and add st,at,rt,st+i to D … Use the samples to compute the WDQN weights (Eq.14) and targets (Eq.15)”, which discloses weighting the learning or training data ; and §4.2). VH and Cini are analogous art because both are concerned with training a Q network and reinforcement learning. Before the effective filing date of the claimed invention, it would have been obvious to one skilled in reinforcement and Q-learning to combine the learning data weighting of Cini and the apparatus of VH to yield to the predictable result of in the determination process, the at least one processor calculating a first function, which predicts a reward sum from a state and an action, by carrying out weighting for the learning data, and determining the action with use of the first function . The motivation for doing so would be to obtain calibrated estimates of epistemic uncertainty in DRL (Cini; Abstract). Regarding claim 2 , the rejection of claim 1 is incorporated and VH fails to explicitly disclose but Cini discloses calculates an index related to variability from one or more values included in the learning data, and (§4.2 and Equations 12 and 13; “We can approximate the predictive mean of the process, and the expectation over the posterior distribution of the Q-value estimates, as the average of T stochastic forward passes through the network”, wherein the spread or variance of the T samples are the variability index ) calculates the first function by applying a smaller weighting factor to the one or more values as the calculated index related to the variability is higher (Equation 14; the equation discloses when the Q-value variance is high, this probability is spread evenly such that there is a lower weight per action. When the variance is low, the weight concentrates on the dominant action which results in a higher effective weight ). The motivation to combine VH and Cini is the same as discussed above with respect to claim 1. Regarding claim 3 , the rejection of claims 1 and 2 are incorporated and VH fails to explicitly disclose but Cini discloses in the determination process, the at least one processor calculates variance of a second function as the index related to the variability, the variance of the second function being obtained with reference to the state and the action (§4.2 and Equations 12 and 13; “We can approximate the predictive mean of the process, and the expectation over the posterior distribution of the Q-value estimates, as the average of T stochastic forward passes through the network”, wherein the spread or variance of the T samples are the variability index ; and §4.1; “In fact, a single stochastic forward pass through the BNN can be interpreted as taking a sample from the model’s predictive distribution, while the predictive mean can be computed as the average of multiple samples. This inference technique is known as Monte Carlo (MC) dropout and can be efficiently parallelized in modern GPUs.” ). The motivation to combine VH and Cini is the same as discussed above with respect to claim 1. Regarding claim 6 , the rejection of claim 1 is incorporated and VH discloses calculates the first function with use of a feature map that maps the state and the action to a vector ([0019]; “The Q network 110 is a deep neural network that is configured to receive as input an input observation and an input action and to generate an estimated future cumulative reward from the input in accordance with a set of parameters.” ; and [0016]; “ the observations characterize states of the environment using high-dimensional pixel inputs from one or more images that characterize the state of the environment” ). Regarding claim 7 , the rejection of claim 1 is incorporated and VH discloses selects an action that maximizes the first function which includes, as an argument, the state acquired in the acquisition process ([0036-0037]; “For each action in the set of actions, the system processes the current observation and the action using the Q network in accordance with current values of the parameters of the Q network (step 204 ). … The system selects an action to be performed by the agent in response to the current observation using the estimated future cumulative rewards (step 206 )” ). Regarding claim 8 , the rejection of claim 1 is incorporated and VH discloses an input device configured to accept the state and the reward ([0040]; “The system receives a reward and a next observation (step 206 ). The next observation characterizes the next state of the environment, i.e., the state that the environment transitioned into as a result of the agent performing the selected action, and the reward is a numeric value that is received by the system, e.g., from the environment, as a consequence of the agent performing the selected action” ). Regarding claim 9 , VH discloses [a]n information processing method comprising, in a repeated manner, the steps of: (([0012-0013]; “This specification generally describes a reinforcement learning system that selects actions to be performed by a reinforcement learning agent interacting with an environment. …In some implementations, the environment is a simulated environment and the agent is implemented as one or more computer programs interacting with the simulated environment”; and Abstract) an information processing apparatus acquiring a state; ([0035]; “The system receives a current observation characterizing the current state of the environment (step 202 ).” ; and Figure 2, Element 202) the information processing apparatus determining an action with reference to the state; and ([0037]; “The system selects an action to be performed by the agent in response to the current observation using the estimated future cumulative rewards (step 206 ).” ; and Figure 2, Element 206) the information processing apparatus accumulating learning data including (i) the state and (ii) a reward obtained by the determined action, ([0041]; “The system generates an experience tuple that includes the current observation, the selected action, the reward, and the next observation and stores the generated experience tuple in a replay memory for use in training the Q network (step 208 )”, wherein the experience tuple comprises the learning or training data that includes state and reward information ; and Figure 2, Element 208) wherein, in the step of determining the action, the action is determined with use of a first function, the first function predicting a reward sum from a state and an action ([0019]; “he Q network 110 is a deep neural network that is configured to receive as input an input observation and an input action and to generate an estimated future cumulative reward from the input in accordance with a set of parameters”, which discloses the first function or NN that predicts a reward sum or cumulative reward from a state and action ; and [0022]; “The reinforcement learning system 100 can then select the action having the highest estimated future cumulative reward as the action to be performed by the agent 102 in response to the observation” ). VH fails to explicitly disclose but Cini discloses and being calculated by carrying out weighting for the learning data (Page 4, Algorithm 1; “Execute at and add st,at,rt,st+i to D … Use the samples to compute the WDQN weights (Eq.14) and targets (Eq.15)”, which discloses weighting the learning or training data ; and §4.2). VH and Cini are analogous art because both are concerned with training a Q network and reinforcement learning. Before the effective filing date of the claimed invention, it would have been obvious to one skilled in reinforcement and Q-learning to combine the learning data weighting of Cini and the apparatus of VH to yield to the predictable result of wherein, in the step of determining the action, the action is determined with use of a first function, the first function predicting a reward sum from a state and an action and being calculated by carrying out weighting for the learning data . The motivation for doing so would be to obtain calibrated estimates of epistemic uncertainty in DRL (Cini; Abstract). Regarding claim 10 , VH discloses [a] computer-readable non-transitory storage medium storing a program causing a computer to function as the information processing, apparatus according to claim 1, the program causing the computer to carry out the acquisition process, the determination process, and the accumulation process (Abstract; “Methods, systems, and apparatus, including computer programs encoded on computer storage media, for training a Q network used to select actions to be performed by an agent interacting with an environment” ). Regarding claim 11 , VH discloses [a]n information processing system comprising an information processing apparatus and a terminal apparatus, whereinthe information processing apparatus comprises at least one processor, the at least one processor carrying out: ([0012-0013]; “This specification generally describes a reinforcement learning system that selects actions to be performed by a reinforcement learning agent interacting with an environment. …In some implementations, the environment is a simulated environment and the agent is implemented as one or more computer programs interacting with the simulated environment” ) an acquisition process of acquiring a state; ([0035]; “The system receives a current observation characterizing the current state of the environment (step 202 ).” ; and Figure 2, Element 202) a determination process of determining an action with reference to the state; and ([0037]; “The system selects an action to be performed by the agent in response to the current observation using the estimated future cumulative rewards (step 206 ).” ; and Figure 2, Element 206) an accumulation process of accumulating learning data including (i) the state and (ii) a reward obtained by the action which has been determined in the determination process, ([0041]; “The system generates an experience tuple that includes the current observation, the selected action, the reward, and the next observation and stores the generated experience tuple in a replay memory for use in training the Q network (step 208 )”, wherein the experience tuple comprises the learning or training data that includes state and reward information ; and Figure 2, Element 208) in the determination process, the at least one processor calculating a first function, which predicts a reward sum from a state and an action, ([0019]; “he Q network 110 is a deep neural network that is configured to receive as input an input observation and an input action and to generate an estimated future cumulative reward from the input in accordance with a set of parameters”, which discloses the first function or NN that predicts a reward sum or cumulative reward from a state and action ) … and determining the action with use of the first function ([0022]; “The reinforcement learning system 100 can then select the action having the highest estimated future cumulative reward as the action to be performed by the agent 102 in response to the observation” ) the terminal apparatus comprises at least one processor, the at least one processor of the terminal apparatus carrying out: a state information provision process of acquiring a state and providing the state to the information processing apparatus; and a reward information provision process of providing, to the information processing apparatus, reward information indicative of a reward obtained by executing an action which has been determined by the information processing apparatus (Figure 2, Steps 202 and 208; and [0035-0041]); and Abstract; and Figure 1). VH fails to explicitly disclose but Cini discloses by carrying out weighting for the learning data (Page 4, Algorithm 1; “Execute at and add st,at,rt,st+i to D … Use the samples to compute the WDQN weights (Eq.14) and targets (Eq.15)”, which discloses weighting the learning or training data ; and §4.2). VH and Cini are analogous art because both are concerned with training a Q network and reinforcement learning. Before the effective filing date of the claimed invention, it would have been obvious to one skilled in reinforcement and Q-learning to combine the learning data weighting of Cini and the apparatus of VH to yield to the predictable result of in the determination process, the at least one processor calculating a first function, which predicts a reward sum from a state and an action, by carrying out weighting for the learning data, and determining the action with use of the first function . The motivation for doing so would be to obtain calibrated estimates of epistemic uncertainty in DRL (Cini; Abstract). Claims 4 and 5 are rejected under 35 U.S.C. § 103 as being obvious over VH in view of Cini and further in view of Arel et al. (US 9536191 B1, hereinafter “Arel”). Regarding claim 4 , the rejection of claims 1 and 2 are incorporated and VH fails to explicitly disclose but Arel discloses a display proces s of displaying (i) at least one selected from the group consisting of the state, the action, the reward, and a value of the first function and (ii) the index related to the variability (Column 1, Lines 50-64; “determining, in accordance with a value function representation and from the action and a current state representation for the current state derived from the current observation, a respective value function estimate that is an estimate of a return resulting from the agent performing the action in response to the current observation, wherein the return is a function of future rewards received in response to the agent performing actions to interact with the environment, determining, in accordance with a confidence function representation and from the current state representation and the action, a respective confidence score that is a measure of confidence that the respective value function estimate for the action is an accurate estimate of the return … adjusting the respective value function estimate for the action using the respective confidence score for the action to determine a respective adjusted value function estimate”, which discloses computing and outputting a first function value, state, and action output and a variability index that is the inverse of the confidence score ; and Column 14, Lines 55-61; “To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device” ). VH, Cini, and Arel are analogous art because all are concerned with training a Q network and reinforcement learning. Before the effective filing date of the claimed invention, it would have been obvious to one skilled in reinforcement and Q-learning to combine the display functions and information of Arel and the apparatus of VH and Cini to yield to the predictable result of displaying (i) at least one selected from the group consisting of the state, the action, the reward, and a value of the first function and (ii) the index related to the variability . The motivation for doing so would be to provide for a measure of confidence that the respective value function estimate for the action is an accurate estimate of the return that will result from the agent performing the action (Arel; Abstract). Regarding claim 5 , the rejection of claims 1, 2, and 4 are incorporated and VH fails to explicitly disclose but Arel discloses a display proces s in the display process, the at least one processor displays, in an emphasized manner, a value which is of the at least one selected from the state, the action, the reward, and the value of the first function and which satisfies that the index related to the variability is equal to or less than a threshold value (Claim 7; “selecting an action having a highest adjusted value function estimate as the action to be performed by the agent” ; and Claim 1; “p t (s t ,a t )=(Q(s t ,a t )−Q min )×c(s t ,a t ), “ ; and Claim 8; and Column 14, Lines 55-61; “To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device” ). The motivation to combine VH, Cini, and Arel is the same as discussed above with respect to claim 4. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to Brent Hoover whose telephone number is (303)297-4403. The examiner can normally be reached Monday - Friday 9-5 MST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Abdullah Kawsar can be reached on 571-270-3169. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /BRENT JOHNSTON HOOVER/Primary Examiner, Art Unit 2127 Application/Control Number: 18/562,537 Page 2 Art Unit: 2127 Application/Control Number: 18/562,537 Page 3 Art Unit: 2127 Application/Control Number: 18/562,537 Page 4 Art Unit: 2127 Application/Control Number: 18/562,537 Page 5 Art Unit: 2127 Application/Control Number: 18/562,537 Page 6 Art Unit: 2127 Application/Control Number: 18/562,537 Page 7 Art Unit: 2127 Application/Control Number: 18/562,537 Page 8 Art Unit: 2127 Application/Control Number: 18/562,537 Page 9 Art Unit: 2127 Application/Control Number: 18/562,537 Page 10 Art Unit: 2127 Application/Control Number: 18/562,537 Page 11 Art Unit: 2127 Application/Control Number: 18/562,537 Page 12 Art Unit: 2127 Application/Control Number: 18/562,537 Page 13 Art Unit: 2127 Application/Control Number: 18/562,537 Page 14 Art Unit: 2127 Application/Control Number: 18/562,537 Page 15 Art Unit: 2127 Application/Control Number: 18/562,537 Page 16 Art Unit: 2127 Application/Control Number: 18/562,537 Page 17 Art Unit: 2127 Application/Control Number: 18/562,537 Page 18 Art Unit: 2127 Application/Control Number: 18/562,537 Page 19 Art Unit: 2127 Application/Control Number: 18/562,537 Page 20 Art Unit: 2127 Application/Control Number: 18/562,537 Page 21 Art Unit: 2127 Application/Control Number: 18/562,537 Page 22 Art Unit: 2127 Application/Control Number: 18/562,537 Page 23 Art Unit: 2127 Application/Control Number: 18/562,537 Page 24 Art Unit: 2127 Application/Control Number: 18/562,537 Page 25 Art Unit: 2127 Application/Control Number: 18/562,537 Page 26 Art Unit: 2127 Application/Control Number: 18/562,537 Page 27 Art Unit: 2127 Application/Control Number: 18/562,537 Page 28 Art Unit: 2127 Application/Control Number: 18/562,537 Page 29 Art Unit: 2127 Application/Control Number: 18/562,537 Page 30 Art Unit: 2127
Read full office action

Prosecution Timeline

Nov 20, 2023
Application Filed
May 15, 2026
Non-Final Rejection mailed — §101, §103, §112
Jul 15, 2026
Interview Requested
Aug 10, 2026
Examiner Interview Summary
Aug 10, 2026
Applicant Interview (Telephonic)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12711351
QUANTIZATION METHOD OF NEURAL NETWORK AND APPARATUS FOR PERFORMING THE SAME
4y 0m to grant Granted Aug 18, 2026
Patent 12704655
SEISMIC DATA PROCESSING USING DUnet
5y 2m to grant Granted Aug 11, 2026
Patent 12705458
DATA PROCESSING METHOD FOR RECURRENT NEURAL NETWORK USING NEURAL NETWORK ACCELERATOR BASED ON SYSTOLIC ARRAY AND NEURAL NETWORK ACCELERATOR
3y 8m to grant Granted Aug 11, 2026
Patent 12699907
METHODS AND APPARATUS FOR PROVIDING INFORMATION OF INTEREST TO ONE OR MORE USERS
8y 11m to grant Granted Aug 04, 2026
Patent 12694312
GENERATIVE ADVERSARIAL NETWORKS FOR USE IN REFINING MODELS FOR SYNTHETIC NETWORK TRAFFIC DATA
5y 0m to grant Granted Jul 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
82%
Grant Probability
99%
With Interview (+22.0%)
3y 5m (~7m remaining)
Median Time to Grant
Low
PTA Risk
Based on 376 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month