Prosecution Insights
Last updated: August 16, 2026
Application No. 18/015,222

METHOD AND SYSTEM FOR DEEP REINFORCEMENT LEARNING (DRL) BASED SCHEDULING IN A WIRELESS SYSTEM

Final Rejection §103
Filed
Jan 09, 2023
Priority
Jul 10, 2020 — provisional 63/050,502 +1 more
Examiner
SHAHEED, KHALID W
Art Unit
2643
Tech Center
2600 — Communications
Assignee
Telefonaktiebolaget LM Ericsson
OA Round
5 (Final)
83%
Grant Probability
Favorable
6-7
OA Rounds
0m
Est. Remaining
98%
With Interview

Examiner Intelligence

Grants 83% — above average
83%
Career Allowance Rate
714 granted / 860 resolved
+21.0% vs TC avg
Moderate +15% lift
Without
With
+15.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
32 currently pending
Career history
902
Total Applications
across all art units

Statute-Specific Performance

§101
6.8%
-33.2% vs TC avg
§103
53.2%
+13.2% vs TC avg
§102
28.0%
-12.0% vs TC avg
§112
9.4%
-30.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 860 resolved cases

Office Action

§103
Detailed Action Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Applicant's arguments filed have been fully considered but they are moot in view of new grounds of rejection. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1, 3-5, 8, 10, 14 are rejected under 35 U.S.C. 103 as being unpatentable over Naderializadeh et al. (US 2019/0116560 A1) in view of CN 108470237 A in further view of Youn et al. (US 2010/0309860 A1). Regarding claim 1 & 14, Naderializadeh et al. discloses a method performed by a network node for Deep Reinforcement Learning, DRL, based scheduling, or a network node, the method and network node comprising: performing a DRL-based scheduling procedure (see “see scheduling using …Deep Reinforcement Learning”) using a preference vector (see vector [0149] ..” represents the scheduling decisions”) for a plurality of network performance metrics (see [0056] “relevant network metric or optimization function”) correlated (see [0089], “channel quality indicator (CQI) measured at the UEs, while minimizing the feedback overhead in terms of number of channel metrics ”) to one of a plurality of desired network performance behaviors (see [0151], “state is considered to include an N×N matrix of channel quality indicator (CQI) ”, the performance is behavior is quality), the preference vector defining weights (see “denote the weights” at each scheduling interval, UE weights, [0061]) for the plurality of network performance metrics correlated to the one of the plurality of desired network performance behaviors (see [0143], “defined as a behavior function mapping the state space to the action space. The goal of the agent is to learn a policy which maximizes a value function”, therefore the state or metrics that provide state information are mapped to action or behavior); obtaining a preference vector (see vector [0149] ..” represents the scheduling decisions”) for respective sets of network performance metrics (see “denote the weights” at each scheduling interval, UE weights, [0061]) for the plurality of desired network performance behaviors, respectively (see [0143], “defined as a behavior function mapping the state space to the action space. The goal of the agent is to learn a policy which maximizes a value function”, therefore the state or metrics that provide state information are mapped to action or behavior); Naderializadeh et al does not specifically disclose however CN 108470237 A discloses a plurality of preference vectors (see preference vector set, page 5); It would be obvious to one of ordinary skill in the art at the time of filing to combine to the teachings of Naderializaheh in view of CN 108470237 A. Doing so would conform to well-known conventions in the filed of art. Naderializaheh in view of CN 108470237 A discloses a DRL-Based scheduling procedure specifically however Youn specifically discloses “performs time-domain scheduling of packets for each of a plurality of transmit time intervals (TTIs)” (see “scheduler” and “scheduling time” for “transmit time interval”). Here, it would have been obvious to combine the DRL-Based scheduling procedure of Naderializaheh in view of CN 108470237 A with the TTI time scheduler of Youn. Doing so would conform to well-known standards in the field of invention. Regarding claim 19 and 22, Naderializadeh discloses method performed by a network node for Deep Reinforcement Learning, DRL, based scheduling, and a network node, the method and network node comprising: determining, for each desired network performance behavior of a plurality of desired network performance behaviors, a preference vector (see vector [0149]-[0150]) to apply to a plurality of network performance metrics (see [0056], “relevant network metric or optimization function”) correlated to the desired network performance behavior, during a training phase of a DRL-based scheduling procedure that optimizes a composite reward (see reward [0150]) generated from the plurality of network performance vectors using the preference vector; and during an execution phase of the DRL-based scheduling procedure, performing the DRL-based scheduling procedure (deep reinforcement learning, scheduling [0147]) using the determined preference vector for the plurality of network performance metrics correlated to one of the plurality of desired network performance behaviors ([0143] The RL agent needs to learn a policy π, which is defined as a behavior function mapping the state space to the action space. For example, at each time step t, the agent takes the action). Regarding claim 3, Naderializadeh discloses the method of claim 1 wherein the plurality of network performance metrics comprise: (a) packet size, (b) packet delay (see delay [0098]), (c) Quality of Service, QoS (QoS, [0159]), requirement(s), (d) cell state, or (e) a combination two or more of (a)-(d). Regarding claim 4, Naderializadeh discloses the method of claim 1 , however CN 108470237 A specifically discloses further comprising selecting the preference vector from among a plurality of preference vectors (see preference vector set, page 5) for respective sets of network performance metrics for a plurality of network performance behaviors, respectively (see page 5 “upper and lower bounds, a new preference vector Gc is randomly generated, the external set Q t is updated, and the parent and child populations are mixed with the preference vector set to obtain the mixed population JointP and the mixed preference vector set JointG, and their respective adaptations are calculated. Degree value, truncation selection produces new offspring population and new preference vector set”); It would be obvious to one of ordinary skill in the art at the time of filing to combine to the teachings of Naderializaheh in view of CN 108470237 A. Doing so would conform to well-known conventions in the filed of art. Regarding claim 5, Naderializadeh discloses the method of claim 4 , however CN 108470237 A specifically discloses wherein selecting the preference vector from among the plurality of preference vectors comprises selecting the preference vector from among the plurality of preference vectors based on one or more parameters (see page 5 “upper and lower bounds, a new preference vector Gc is randomly generated, the external set Q t is updated, and the parent and child populations are mixed with the preference vector set to obtain the mixed population JointP and the mixed preference vector set JointG, and their respective adaptations are calculated. Degree value, truncation selection produces new offspring population and new preference vector set”); It would be obvious to one of ordinary skill in the art at the time of filing to combine to the teachings of Naderializaheh in view of CN 108470237 A. Doing so would conform to well-known conventions in the filed of art. Regarding claim 8, Naderializadeh discloses the method of claim 1 wherein the DRL-based scheduling procedure is a Deep Q-Learning Network, DQN, scheduling procedure (see deep Q-network DQN, see [0147]). Regarding claim 10, Naderializadeh discloses the method of claim 1 further comprising, prior to performing the DRL-based scheduling procedure, determining the preference vector for the desired network performance behavior (see joint action vector, representing scheduling decisions) Allowable Subject Matter Claims 19 & 22 are allowed Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to K. WILFORD SHAHEED whose telephone number is (469) 295-9175. The examiner can normally be reached on Monday-Friday 9 am-6pm; CST; ALT Friday. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. The examiner’s Supervisor, Jinsong Hu, can be reached at (571)272-3965, where attempts to reach the examiner are unsuccessful. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /KHALID W SHAHEED/Primary Examiner, Art Unit 2643
Read full office action

Prosecution Timeline

Show 2 earlier events
Jul 10, 2025
Response Filed
Sep 30, 2025
Final Rejection mailed — §103
Nov 26, 2025
Response after Non-Final Action
Dec 08, 2025
Non-Final Rejection mailed — §103
Mar 09, 2026
Response Filed
Apr 16, 2026
Non-Final Rejection mailed — §103
Jul 13, 2026
Response Filed
Aug 11, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12707413
TERMINAL OPERATION METHOD AND DEVICE IN WIRELESS COMMUNICATION SYSTEM
2y 9m to grant Granted Aug 11, 2026
Patent 12707520
MASTER NODE, COMMUNICATION CONTROL METHOD, AND COMMUNICATION APPARATUS
2y 7m to grant Granted Aug 11, 2026
Patent 12707380
SLICE ADMISSION CONTROL METHOD AND COMMUNICATION APPARATUS
2y 7m to grant Granted Aug 11, 2026
Patent 12696161
COMMUNICATION CONTROL METHOD
4y 10m to grant Granted Jul 28, 2026
Patent 12696098
Mining Mobile Communication System and Method
3y 0m to grant Granted Jul 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

6-7
Expected OA Rounds
83%
Grant Probability
98%
With Interview (+15.0%)
2y 7m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 860 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month