Prosecution Insights
Last updated: October 02, 2026
Application No. 18/015,222

METHOD AND SYSTEM FOR DEEP REINFORCEMENT LEARNING (DRL) BASED SCHEDULING IN A WIRELESS SYSTEM

Final Rejection §103
Filed
Jan 09, 2023
Priority
Jul 10, 2020 — provisional 63/050,502 +1 more
Examiner
SHAHEED, KHALID W
Art Unit
2643
Tech Center
2600 — Communications
Assignee
Telefonaktiebolaget LM Ericsson
OA Round
5 (Final)
83%
Grant Probability
Favorable
6-7
OA Rounds
0m
Est. Remaining
98%
With Interview

Examiner Intelligence

Grants 83% — above average
83%
Career Allowance Rate
725 granted / 872 resolved
+21.1% vs TC avg
Moderate +15% lift
Without
With
+14.9%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
26 currently pending
Career history
906
Total Applications
across all art units

Statute-Specific Performance

§101
6.7%
-33.3% vs TC avg
§103
53.6%
+13.6% vs TC avg
§102
27.8%
-12.2% vs TC avg
§112
9.2%
-30.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 872 resolved cases

Office Action

§103
Detailed Action Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Applicant's arguments filed have been fully considered but they are moot in view of new grounds of rejection. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1, 3-5, 8, 10, 14 are rejected under 35 U.S.C. 103 as being unpatentable over Naderializadeh et al. (US 2019/0116560 A1) in view of CN 108470237 A in further view of Katsuki et al. (US 2016/0055416 A1). Regarding claim 1 & 14, Naderializadeh et al. discloses a method performed by a network node for Deep Reinforcement Learning, DRL, based scheduling, or a network node, the method and network node comprising: performing a DRL-based scheduling procedure (see “see scheduling using …Deep Reinforcement Learning”) using a preference vector (see vector [0149] ..” represents the scheduling decisions”) for a plurality of network performance metrics (see [0056] “relevant network metric or optimization function”) correlated (see [0089], “channel quality indicator (CQI) measured at the UEs, while minimizing the feedback overhead in terms of number of channel metrics ”) to one of a plurality of desired network performance behaviors (see [0151], “state is considered to include an N×N matrix of channel quality indicator (CQI) ”, the performance is behavior is quality), the preference vector defining weights (see “denote the weights” at each scheduling interval, UE weights, [0061]) for the plurality of network performance metrics correlated to the one of the plurality of desired network performance behaviors (see [0143], “defined as a behavior function mapping the state space to the action space. The goal of the agent is to learn a policy which maximizes a value function”, therefore the state or metrics that provide state information are mapped to action or behavior); obtaining a preference vector (see vector [0149] ..” represents the scheduling decisions”) for respective sets of network performance metrics (see “denote the weights” at each scheduling interval, UE weights, [0061]) for the plurality of desired network performance behaviors, respectively (see [0143], “defined as a behavior function mapping the state space to the action space. The goal of the agent is to learn a policy which maximizes a value function”, therefore the state or metrics that provide state information are mapped to action or behavior); Naderializadeh et al does not specifically disclose however CN 108470237 A discloses a plurality of preference vectors (see preference vector set, page 5); It would be obvious to one of ordinary skill in the art at the time of filing to combine to the teachings of Naderializaheh in view of CN 108470237 A. Doing so would conform to well-known conventions in the filed of art. Naderializaheh in view of CN 108470237 A does not explicitly disclose however Katsuki discloses preference vector is selected base on results of training (see “[0079] Next, on the 13th line, the learning processing section 150 determines, for each of the sample candidates w•.sup.(m) for the preference vector, whether the sample candidate w•.sup.(m) is selected as the next sample of the preference vector, based on the occurrence probability of the sample candidate w•.sup.(m) in the prior distribution, and the likelihood of the sample candidate w•.sup.(m) for selection in the history data and the preference vector of each selection subject”, therefore based on learning processing – training a from candidates a selected) from among a plurality of candidate preference vectors (see [0062], based on “learning processing section” … “preference vectors”). Here, it would have been obvious to combine the DRL-Based scheduling procedure of Naderializaheh in view of CN 108470237 A with the TTI time scheduler of . Doing so would conform to well-known standards in the field of invention. Regarding claim 19 and 22, Naderializadeh discloses method performed by a network node for Deep Reinforcement Learning, DRL, based scheduling, and a network node, the method and network node comprising: determining, for each desired network performance behavior of a plurality of desired network performance behaviors, a preference vector (see vector [0149]-[0150]) to apply to a plurality of network performance metrics (see [0056], “relevant network metric or optimization function”) correlated to the desired network performance behavior, during a training phase of a DRL-based scheduling procedure that optimizes a composite reward (see reward [0150]) generated from the plurality of network performance vectors using the preference vector; and during an execution phase of the DRL-based scheduling procedure, performing the DRL-based scheduling procedure (deep reinforcement learning, scheduling [0147]) using the determined preference vector for the plurality of network performance metrics correlated to one of the plurality of desired network performance behaviors ([0143] The RL agent needs to learn a policy π, which is defined as a behavior function mapping the state space to the action space. For example, at each time step t, the agent takes the action). Regarding claim 3, Naderializadeh discloses the method of claim 1 wherein the plurality of network performance metrics comprise: (a) packet size, (b) packet delay (see delay [0098]), (c) Quality of Service, QoS (QoS, [0159]), requirement(s), (d) cell state, or (e) a combination two or more of (a)-(d). Regarding claim 4, Naderializadeh discloses the method of claim 1 , however CN 108470237 A specifically discloses further comprising selecting the preference vector from among a plurality of preference vectors (see preference vector set, page 5) for respective sets of network performance metrics for a plurality of network performance behaviors, respectively (see page 5 “upper and lower bounds, a new preference vector Gc is randomly generated, the external set Q t is updated, and the parent and child populations are mixed with the preference vector set to obtain the mixed population JointP and the mixed preference vector set JointG, and their respective adaptations are calculated. Degree value, truncation selection produces new offspring population and new preference vector set”); It would be obvious to one of ordinary skill in the art at the time of filing to combine to the teachings of Naderializaheh in view of CN 108470237 A. Doing so would conform to well-known conventions in the filed of art. Regarding claim 5, Naderializadeh discloses the method of claim 4 , however CN 108470237 A specifically discloses wherein selecting the preference vector from among the plurality of preference vectors comprises selecting the preference vector from among the plurality of preference vectors based on one or more parameters (see page 5 “upper and lower bounds, a new preference vector Gc is randomly generated, the external set Q t is updated, and the parent and child populations are mixed with the preference vector set to obtain the mixed population JointP and the mixed preference vector set JointG, and their respective adaptations are calculated. Degree value, truncation selection produces new offspring population and new preference vector set”); It would be obvious to one of ordinary skill in the art at the time of filing to combine to the teachings of Naderializaheh in view of CN 108470237 A. Doing so would conform to well-known conventions in the filed of art. Regarding claim 8, Naderializadeh discloses the method of claim 1 wherein the DRL-based scheduling procedure is a Deep Q-Learning Network, DQN, scheduling procedure (see deep Q-network DQN, see [0147]). Regarding claim 10, Naderializadeh discloses the method of claim 1 further comprising, prior to performing the DRL-based scheduling procedure, determining the preference vector for the desired network performance behavior (see joint action vector, representing scheduling decisions) Claim(s) 26 is rejected under 35 U.S.C. 103 as being unpatentable over Naderializadeh et al. (US 2019/0116560 A1) in view of CN 108470237 A in further view of Katsuki in further view of Youn et al. (US 2010/0309860 A1). Regarding 26, Naderializaheh in view of CN 108470237 A disclose the method of claim 1; Youn best discloses wherein the DRL-based scheduling procedure performs time-domain scheduling of packets for each of a plurality of transmit time intervals (TTIs) ((see [0006], “scheduler” and “scheduling time” for “transmit time interval”). Here, it would have been obvious to combine the DRL-Based scheduling procedure of Naderializaheh in view of CN 108470237 A with the TTI time scheduler of Youn. Doing so would conform to well-known standards in the field of invention. Allowable Subject Matter Claims 19 & 22 are allowed Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to K. WILFORD SHAHEED whose telephone number is (469) 295-9175. The examiner can normally be reached on Monday-Friday 9 am-6pm; CST; ALT Friday. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. The examiner’s Supervisor, Jinsong Hu, can be reached at (571)272-3965, where attempts to reach the examiner are unsuccessful. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /KHALID W SHAHEED/Primary Examiner, Art Unit 2643
Read full office action

Prosecution Timeline

Show 2 earlier events
Jul 10, 2025
Response Filed
Sep 30, 2025
Final Rejection mailed — §103
Nov 26, 2025
Response after Non-Final Action
Dec 08, 2025
Non-Final Rejection mailed — §103
Mar 09, 2026
Response Filed
Apr 16, 2026
Non-Final Rejection mailed — §103
Jul 13, 2026
Response Filed
Aug 11, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12750167
METHOD FOR SIDELINK MONITORING AND TERMINAL DEVICE
4y 11m to grant Granted Sep 29, 2026
Patent 12748174
CLASSIFYING ACCESS POINT BEACON COMMUNICATIONS USING MACHINE LEARNING
3y 10m to grant Granted Sep 29, 2026
Patent 12745144
METHODS AND NODES FOR MEASURING A SERVING CELL IN L1/L2-CENTRIC MOBILITY
3y 5m to grant Granted Sep 22, 2026
Patent 12739939
METHOD AND APPARATUS FOR TRANSMITTING AND RECEIVING SIGNAL IN WIRELESS COMMUNICATION SYSTEM
3y 0m to grant Granted Sep 15, 2026
Patent 12726260
Dynamic Channel Selection For Unmanned Aerial Vehicles
3y 5m to grant Granted Sep 01, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

6-7
Expected OA Rounds
83%
Grant Probability
98%
With Interview (+14.9%)
2y 7m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 872 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month