Prosecution Insights
Last updated: October 02, 2026
Application No. 18/238,337

CONTROL DEVICE, CONTROL SYSTEM, CONTROL METHOD, AND COMPUTER READABLE MEDIUM STORING CONTROL PROGRAM

Final Rejection §103
Filed
Aug 25, 2023
Priority
Mar 11, 2021 — continuation of PCTJP2021009708
Examiner
WENG, PEI YONG
Art Unit
2141
Tech Center
2100 — Computer Architecture & Software
Assignee
Mitsubishi Electric Corporation
OA Round
2 (Final)
79%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 79% — above average
79%
Career Allowance Rate
514 granted / 647 resolved
+24.4% vs TC avg
Strong +23% interview lift
Without
With
+22.8%
Interview Lift
resolved cases with interview
Typical timeline
3y 1m
Avg Prosecution
32 currently pending
Career history
665
Total Applications
across all art units

Statute-Specific Performance

§101
13.0%
-27.0% vs TC avg
§103
55.8%
+15.8% vs TC avg
§102
20.5%
-19.5% vs TC avg
§112
7.0%
-33.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 647 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION This action is responsive to the following communication: Amendment filed Jun. 24, 2026. This Action is made Final. Claims 1-2 and 8-9 are pending in the case. Claims 1 and 7-9 are independent claims. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-2 and 8-9 are rejected under 35 U.S.C. 103 as being unpatentable over “Human-Like Autonomous Vehicle Speed Control by Deep Reinforcement learning with Double Q-Learning” Zhang et al. (hereinafter Zhang) 2018 in view of Vignard et al. (hereinafter Vignard) U.S. Patent Pub No. 2020/0189574. With respect to independent claim 1, Zhang teaches a control device comprising: state data acquisition circuitry to acquire state data indicating a state of a control target (see e.g., Page 2-3); state category identification circuitry to identify a state category to which a state indicated by the state data belongs among a plurality of state categories indicating classifications of states of the control target on the basis of the state data (see e.g., Fig. 1 and Page 2 – “As shown in Fig.1, the general agent-environment interaction modeling (of both traditional RL and the emerging DRL) consists of an agent, an environment, a finite state space S, a set of available actions A, and a reward function: S×A → R. The decision maker is called the agent, and should be trained as the interaction system runs. The agent needs to interact with the outside, which is called the environment. The interaction between the agent and the environment is a continual process. At each decision epoch k, the agent will make” It is implicit that the state categories are different and that they would have to be identified and classified); reward generation circuitry to calculate a reward value of a control detail for the control target on the basis of the state category and the state data (see e.g., Page 2 and 5 – “Reward network: Reward is necessary in almost all reinforcement learning algorithms offering the goal of the reinforcement learning agent. The reward estimates how good the agent performs an action in a given state (or what are the good or bad things for the agent). In this paper, we design a reward network to map each state to a scalar,”); and control learning circuitry to learn the control detail on the basis of the state data and the reward value, wherein the reward generation circuitry includes reward calculation formula selection circuitry to select a reward calculation formula different for each of the plurality of state categories on the basis of the inputted state category (see e.g., Page 5 – the reward function is based on a reward network, where a corresponding reward value is generated in response to a control target the reward network selects a different reward calculation formula reward network e.g. may select +2 for (x, a) E { C-}) based on one of several state categories (e.g. for (x, a) E { C-} )), and reward value calculation circuitry to calculate the reward value using the reward calculation formula selected by the reward calculation formula selection circuitry (see e.g., Page 5). Zhang does not expressly show the feature discussed below. However, Vignard teaches the indicated state category corresponding to a state of the vehicle, and being one of: a state of the vehicle travelling straight, a state of the vehicle turning right, a lane changing state, or a state of the vehicle being parked (see e.g. Para [52][53][68][69][75]-[77]-“A state and maneuver estimation framework may be applied based on a combination of discrete model-based maneuver prediction and discrete-continuous Bayesian filtering. More generally, the probabilistic model proposed can be categorized as a Switching State Space Model (SSSM), in which a high-level layer reasons about the maneuvers being performed by the interacting road users and determines the evolution of the low-level dynamics … A maneuver dynamics database is stored in the data storage device 102 in the present embodiment and comprises a finite plurality of predetermined motion parameter sets, each associated to an alternative maneuver. For example, the database may contain a lane change motion parameter set, corresponding to a road user changing lanes; and a lane keeping motion parameter set, corresponding to a road user staying in the same lane. However, it is envisioned that several other possible maneuvers may be contained in the database.”); the reward generation circuitry includes reward calculation formula selection circuitry having a plurality of reward calculation formulae, each reward calculation formula corresponding to a respective state category, the reward calculation formula selection unit being configured to select, from the plurality of reward calculation formulae, a reward calculation formula corresponding to the inputted state category (see e.g. Para [3][15][16][29][81]-[83]-“ account the at least one prospective state of the at least one road user other than the target road user at the same subsequent time step; aggregating the costs of the prospective states of each alternative subsequent sequence of prospective states of the target road user to obtain an aggregated cost for each alternative subsequent sequence of prospective states of the target road user; averaging the aggregated costs of the alternative sequence of prospective states of the target road user for each maneuver of the finite plurality of alternative maneuvers to obtain an average aggregated cost of each maneuver of the finite plurality of alternative maneuvers … This behavior model balances the (navigational and risk) preferences of road users and enables a planning based prediction of their anticipatory behavior. A target road user will perform a maneuver at the current time step if, given his prediction for the behavior of the surrounding road users, this leads to a sequence of F future states that agree with its own preferences, which are encoded in its behavior model … ”). Both Zhang and Vignard are directed to vehicle control and rewarding system. Accordingly, it would have been obvious to the skilled artisan before the effective filing date of the claimed invention having Zhang and Vignard in front of them to modify the system of Zhang to include the above feature. The motivation to combine Zhang and Vignard comes from Vignard. Vignard discloses the motivation to provide a state and maneuver estimation framework based maneuver prediction so that vehicle control is improved (see e.g. Para [52][53][68][69][75]-[77]). With respect to independent claim 2, the modified Zhang teaches training data generation circuitry to generate training data in which the state data and the control detail are associated with each other (see e.g., Page 3 - ZHANG teaches that the model is trained using correlation (i.e. association) between state and action (i.e. control) data - "DNN to derive the correlation between each state-action pair (s, a) of the system under control"; "DNN ... trained"). Claim 8 is rejected for the similar reasons discussed above with respect to claim 1. Claim 9 is rejected for the similar reasons discussed above with respect to claim 1. Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Zhang in view of Vignard and further in view of “Reinforcement learning is supervised learning on optimized data” Eysenbach et al. (hereinafter Eysenbach) 2020. With respect to independent claim 7, Zhang teaches a control system comprising: state data acquisition circuitry to acquire state data indicating a state of a control target (see e.g., Page 2-3); state category identification circuitry to identify a state category to which a state indicated by the state data belongs among a plurality of state categories indicating classifications of states of the control target on the basis of the state data (see e.g., Fig. 1 and Page 2 – “As shown in Fig.1, the general agent-environment interaction modeling (of both traditional RL and the emerging DRL) consists of an agent, an environment, a finite state space S, a set of available actions A, and a reward function: S×A → R. The decision maker is called the agent, and should be trained as the interaction system runs. The agent needs to interact with the outside, which is called the environment. The interaction between the agent and the environment is a continual process. At each decision epoch k, the agent will make” It is implicit that the state categories are different and that they would have to be identified and classified); reward generation circuitry to calculate a reward value of a control detail for the control target on the basis of the state category and the state data; control learning circuitry to learn the control detail on the basis of the state data and the reward value (see e.g., Page 2 and 5); training data generation circuitry to generate training data in which the state data and the control detail are associated with each other (see e.g., Page 3 - ZHANG teaches that the model is trained using correlation (i.e. association) between state and action (i.e. control) data - "DNN to derive the correlation between each state-action pair (s, a) of the system under control"; "DNN ... trained"); wherein the reward generation circuitry includes reward calculation formula selection circuitry to select a reward calculation formula different for each of the plurality of state categories on the basis of the inputted state category (see e.g., Page 5 – the reward function is based on a reward network, where a corresponding reward value is generated in response to a control target the reward network selects a different reward calculation formula reward network e.g. may select +2 for (x, a) E { C-}) based on one of several state categories (e.g. for (x, a) E { C-} )), and reward value calculation circuitry to calculate the reward value using the reward calculation formula selected by the reward calculation formula selection circuitry (see e.g., Page 5). Zhang does not expressly show the feature discussed below. However, Vignard teaches the indicated state category corresponding to a state of the vehicle, and being one of: a state of the vehicle travelling straight, a state of the vehicle turning right, a lane changing state, or a state of the vehicle being parked (see e.g. Para [52][53][68][69][75]-[77]-“A state and maneuver estimation framework may be applied based on a combination of discrete model-based maneuver prediction and discrete-continuous Bayesian filtering. More generally, the probabilistic model proposed can be categorized as a Switching State Space Model (SSSM), in which a high-level layer reasons about the maneuvers being performed by the interacting road users and determines the evolution of the low-level dynamics … A maneuver dynamics database is stored in the data storage device 102 in the present embodiment and comprises a finite plurality of predetermined motion parameter sets, each associated to an alternative maneuver. For example, the database may contain a lane change motion parameter set, corresponding to a road user changing lanes; and a lane keeping motion parameter set, corresponding to a road user staying in the same lane. However, it is envisioned that several other possible maneuvers may be contained in the database.”); the reward generation circuitry includes reward calculation formula selection circuitry having a plurality of reward calculation formulae, each reward calculation formula corresponding to a respective state category, the reward calculation formula selection unit being configured to select, from the plurality of reward calculation formulae, a reward calculation formula corresponding to the inputted state category (see e.g. Para [3][15][16][29][81]-[83]-“ account the at least one prospective state of the at least one road user other than the target road user at the same subsequent time step; aggregating the costs of the prospective states of each alternative subsequent sequence of prospective states of the target road user to obtain an aggregated cost for each alternative subsequent sequence of prospective states of the target road user; averaging the aggregated costs of the alternative sequence of prospective states of the target road user for each maneuver of the finite plurality of alternative maneuvers to obtain an average aggregated cost of each maneuver of the finite plurality of alternative maneuvers … This behavior model balances the (navigational and risk) preferences of road users and enables a planning based prediction of their anticipatory behavior. A target road user will perform a maneuver at the current time step if, given his prediction for the behavior of the surrounding road users, this leads to a sequence of F future states that agree with its own preferences, which are encoded in its behavior model … ”). Both Zhang and Vignard are directed to vehicle control and rewarding system. Accordingly, it would have been obvious to the skilled artisan before the effective filing date of the claimed invention having Zhang and Vignard in front of them to modify the system of Zhang to include the above feature. The motivation to combine Zhang and Vignard comes from Vignard. Vignard discloses the motivation to provide a state and maneuver estimation framework based maneuver prediction so that vehicle control is improved (see e.g. Para [52][53][68][69][75]-[77]). Zhang does not expressly show supervised learning circuitry to generate a supervised learned model for inferring the control detail from the state data on the basis of the training data generated by the training data generation circuitry; and action inference circuitry to infer the control detail using the supervised learned model. However, Zhang expressly indicates that direct supervised learning of state-action pairs would be obvious to consider (see e.g., Page 2 – “In reference to both Double Q-learning and DQN, we refer to the resulting learning algorithm as Double DQN. In this paper, we use double DQN to build the vehicle speed model. By approximating this function rather than directly learning the state-action pairs in a supervised fashion, one can handle new scenarios better.”) Further, Eysenbach teaches similar feature (see e.g., Page 2-4). Both Zhang and Eysenbach are directed to control and rewarding system. Accordingly, it would have been obvious to the skilled artisan before the effective filing date of the claimed invention having Zhang and Eysenbach in front of them to further modify the modified system of Zhang to include the above feature. The motivation to combine Zhang and Eysenbach comes from Eysenbach. Eysenbach discloses the motivation to implement direct supervised learning to optimize performance of the system (see e.g. Page 2-4). It is noted that any citation to specific pages, columns, lines, or figures in the prior art references and any interpretation of the references should not be considered to be limiting in any way. “The use of patents as references is not limited to what the patentees describe as their own inventions or to the problems with which they are concerned. They are part of the literature of the art, relevant for all they contain.” In re Heck, 699 F.2d 1331, 1332-33, 216 USPQ 1038, 1039 (Fed. Cir. 1983) (quoting In re Lemelson, 397 F.2d 1006, 1009, 158 USPQ 275, 277 (CCPA 1968)). Further, a reference may be relied upon for all that it would have reasonably suggested to one having ordinary skill the art, including nonpreferred embodiments. Merck & Co. v. Biocraft Laboratories, 874 F.2d 804, 10 USPQ2d 1843 (Fed. Cir.), cert. denied, 493 U.S. 975 (1989). See also Upsher-Smith Labs. v. Pamlab, LLC, 412 F.3d 1319, 1323, 75 USPQ2d 1213, 1215 (Fed. Cir. 2005); Celeritas Technologies Ltd. v. Rockwell International Corp., 150 F.3d 1354, 1361, 47 USPQ2d 1516, 1522-23 (Fed. Cir. 1998). Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to PEIYONG WENG whose telephone number is (571)270-1660. The examiner can normally be reached on Mon.-Fri. 8 am to 5 pm. If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Matthew Ell, can be reached on (571) 270-3264. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://portal.uspto.gov/external/portal. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). /PEI YONG WENG/ Primary Examiner, Art Unit 2141
Read full office action

Prosecution Timeline

Aug 25, 2023
Application Filed
Mar 04, 2026
Non-Final Rejection mailed — §103
Jun 11, 2026
Examiner Interview Summary
Jun 11, 2026
Applicant Interview (Telephonic)
Jun 24, 2026
Response Filed
Jul 29, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12722290
CROSS-DOMAIN IMITATION LEARNING USING GOAL CONDITIONED POLICIES
3y 5m to grant Granted Sep 01, 2026
Patent 12718006
METHODS, APPARATUS AND SYSTEMS FOR ANNOTATION OF TEXT DOCUMENTS
2y 3m to grant Granted Aug 25, 2026
Patent 12717464
Systems, methods, and user interfaces for editing digital assets
1y 4m to grant Granted Aug 25, 2026
Patent 12711424
SYSTEMS AND METHODS FOR DETERMINATION, DESCRIPTION, AND USE OF FEATURE SETS FOR MACHINE LEARNING CLASSIFICATION SYSTEMS, INCLUDING ELECTRONIC MESSAGING SYSTEMS EMPLOYING MACHINE LEARNING CLASSIFICATION
3y 6m to grant Granted Aug 18, 2026
Patent 12699878
GENERATING IMPLICIT PLANS FOR ACCOMPLISHING GOALS IN AN ENVIRONMENT USING ATTENTION OPERATIONS OVER PLANNING EMBEDDINGS
4y 0m to grant Granted Aug 04, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
79%
Grant Probability
99%
With Interview (+22.8%)
3y 1m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 647 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month