Prosecution Insights
Last updated: August 16, 2026
Application No. 18/714,086

DISTRIBUTED REWARD DECOMPOSITION FOR REINFORCEMENT LEARNING

Non-Final OA §103
Filed
May 28, 2024
Priority
Dec 01, 2021 — nonprovisional of PCTIB2021061200
Examiner
ELL, MATTHEW
Art Unit
Tech Center
Assignee
Telefonaktiebolaget LM Ericsson
OA Round
1 (Non-Final)
67%
Grant Probability
Favorable
1-2
OA Rounds
1y 9m
Est. Remaining
89%
With Interview

Examiner Intelligence

Grants 67% — above average
67%
Career Allowance Rate
254 granted / 381 resolved
+6.7% vs TC avg
Strong +22% interview lift
Without
With
+22.2%
Interview Lift
resolved cases with interview
Typical timeline
3y 11m
Avg Prosecution
7 currently pending
Career history
392
Total Applications
across all art units

Statute-Specific Performance

§101
14.0%
-26.0% vs TC avg
§103
49.8%
+9.8% vs TC avg
§102
17.3%
-22.7% vs TC avg
§112
14.6%
-25.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 381 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claims 1-20 are pending, all examined and rejected. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1-10 are rejected under 35 U.S.C. 103 as being unpatentable over Davydov et al., “Q-Mixing Network for Multi-agent Pathfinding in Partially Observable Grid Environments”, October 4, 2021 (hereinafter Davydov) in view of Gemelos, US PGPub #20180376390, published December 27, 2018 With regard to Independent Claim 1, Davydov teaches a method of distributed training of a machine learning model, the method comprising inputting a first set of observations and a first set of actions for a primary agent to generate a first Q function for the primary agent. See e.g., Abstract (multi agent reinforcement learning … many agents working together to accomplish a mission.) See also, Section 2, 2nd paragraph (discussing that each agent generates a Q function from observations (states) and actions.) Davydov further discloses inputting a second set of observations and a second set of actions for a set of secondary agents to generate a set of Q functions for the set of secondary agents. See e.g., Abstract (multi agent reinforcement learning … many agents working together to accomplish a mission.) See also, Section 2, 2nd paragraph (discussing that each agent generates a Q function from observations (states) and actions.) Although not important for this mapping, the examiner notes that the BRI of a second set of agents includes a single second agent. Davydov further discloses generating a Qtot function from the first Q function and the set of Q functions by a mixing network for the primary agent, the Q tot function to [achieve a goal.] See Section 3, 5th paragraph, (Q tot function is generated from all Q-values of the different agents, including primary and secondary agents.) Davydov does not explicitly disclose using this to generate actions or predictions to configure a first node to operate in a telecommunication network. In an analogous art, Gemelos uses reinforcement learning to generate actions or predictions to configure a first node to operate in a telecommunications network. See e.g., Abstract (based on observed states at nodes in device take actions). See Fig. 2, (reinforcement learning). It would have been obvious to a person having ordinary skill in the art having both Davydov and Gemelos before them to modify Davydov to put it to use in the reinforcement system of Gemelos. One would be motivated to do so to use the theories of Davydov in a practical environment to further the reinforcement learning of Nyembe. With regard to Dependent Claim 2, As discussed with Claim 1, Davydov-Gemelos teach all of the limitations. Davydov-Gemelos further teaches deploying the Qtot function to the first node in the telecommunication network to manage actions of the first node. See e.g., Fig. 3, [0038], (discussing reinforcement learning agent which carries out functions can be deployed on the individual devices.) With regard to Dependent Claim 3, As discussed with Claim 2, Davydov-Gemelos teach all of the limitations. Davydov-Gemelos further teaches wherein the primary agent represents the first node in a telecommunications network. See e.g., Fig. 3, [0038], (discussing reinforcement learning agent which carries out functions can be deployed on the individual devices, thus the agents represent the nodes) With regard to Dependent Claim 4, As discussed with regard to Claim 3, Davydov-Gemelos teach all of the limitations. Davydov-Gemelos further teaches wherein the secondary agents represent a set of nodes in the telecommunication network that affect operation of the first node. See e.g., Fig. 3, [0038], (discussing reinforcement learning agent which carries out functions can be deployed on the individual devices, thus the agents represent the nodes.) The examiner notes devices in a telecommunication network naturally “affect operation” of other devices. With regard to Dependent claim 5, As discussed with regard to Claim 1, Davydov-Gemelos teach all of the limitations. Davydov-Gemelos further teaches wherein each of the secondary agents has a respective mixing network to determine a respective Qtot function for each of the secondary agents. The examiner notes at least two ways this claim limitation is met by the art. First, as noted above the BRI of a set of secondary agents includes a single secondary agent. In this situation even if there is only one mixing network then the secondary agent has a “respective mixing network.” Second, even in a situation with a plurality of secondary agents and a plurality of mixing networks this would simply result in more of the same functionality of Davydov-Gemelos and would not result in anything new or unexpected. As such this is an obvious duplication of parts. See MPEP 2144. With regard to Claims 6-10, Claims 6-10 are similar in scope to Claims 1-5 respectively and are rejected under a similar rationale. Claims 11-15 are rejected under 35 U.S.C. 103 as being unpatentable over Davydov in view of Gemelos further in view of Joseph, US PGPub #20210105228, published April 8, 2021 With regard to Claims 11-15, Claims 11-15 are largely similar in scope to Claims 1-5 respectively and are rejected under a similar rationale. Davydov-Gemelos does not explicitly disclose a plurality of virtual machines, the plurality of virtual machines implementing network function virtualization (NFV). In an analogous art, Joseph discloses a telecom network that uses reinforcement learning that comprises a plurality of virtual machines, the plurality of virtual machines implementing network function virtualization (NFV). See e.g., [0003]. It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention having both Davydov-Gemelos and Joseph before them to include the NFV system of Joseph. One would have been motivated to do so as NFV “promises ability to scale and adapt to new technology advancement without significant increase” in costs. See Joseph, [0003]. Claims 16-20 are rejected under 35 U.S.C. 103 as being unpatentable over Davydov in view of Gemelos further in view of Dechene, US PGPub #20220245441, filed January 29, 2021 With regard to Claims 16-20, Claims 16-20 are largely similar in scope to Claims 1-5 respectively and are rejected under a similar rationale. The examiner notes that Gemelos discloses a control plane device (see [0020]). Davydov-Gemelos does not explicitly disclose a software defined networking (SDN) network or that the telecommunication network is managed by the SDN network. In an analogous art, Dechene discloses a telecom network that uses reinforcement learning that is managed by a software defined networking (SDN) network. See e.g., [0034, 0035]. It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention having both Davydov-Gemelos and Dechene before them to include the SDN network of Dechene. One would have been motivated to do so to improve the static architecture of traditional networks. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to MATT ELL whose telephone number is (571)270-3264. The examiner can normally be reached 9-5, M-F. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Christyann Pulliam can be reached at 571-270-1007. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MATTHEW ELL/ Supervisory Patent Examiner, Art Unit 2141
Read full office action

Prosecution Timeline

May 28, 2024
Application Filed
Aug 04, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12699927
Overlapping Gradient Synchronization In Machine Learning
3y 5m to grant Granted Aug 04, 2026
Patent 12670311
AUTOMATIC LAYOUT UPDATES FOR DOCUMENT SPACES IN A GROUP-BASED COMMUNICATION SYSTEM
3y 7m to grant Granted Jun 30, 2026
Patent 12632729
Bilevel Optimization Based Decentralized Framework for Personalized Client Learning
3y 8m to grant Granted May 19, 2026
Patent 12596914
GENERATIVE ADVERSARIAL NEURAL ARCHITECTURE SEARCH
4y 5m to grant Granted Apr 07, 2026
Patent 12596926
SYSTEMS AND METHODS FOR ADJUSTING DATA PROCESSING COMPONENTS FOR NON-OPERATIONAL TARGETS
3y 6m to grant Granted Apr 07, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
67%
Grant Probability
89%
With Interview (+22.2%)
3y 11m (~1y 9m remaining)
Median Time to Grant
Low
PTA Risk
Based on 381 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month