Prosecution Insights
Last updated: October 02, 2026
Application No. 18/388,010

USING A CURRICULUM FOR REINFORCEMENT LEARNING TO TRAIN AN LLM-BASED NETWORK TROUBLESHOOTING AGENT

Final Rejection §103
Filed
Nov 08, 2023
Examiner
PEACH, POLINA G
Art Unit
Tech Center
Assignee
Cisco Technology Inc.
OA Round
2 (Final)
50%
Grant Probability
Moderate
3-4
OA Rounds
10m
Est. Remaining
74%
With Interview

Examiner Intelligence

Grants 50% of resolved cases
50%
Career Allowance Rate
239 granted / 474 resolved
-9.6% vs TC avg
Strong +23% interview lift
Without
With
+23.3%
Interview Lift
resolved cases with interview
Typical timeline
3y 9m
Avg Prosecution
28 currently pending
Career history
510
Total Applications
across all art units

Statute-Specific Performance

§101
19.9%
-20.1% vs TC avg
§103
49.1%
+9.1% vs TC avg
§102
12.7%
-27.3% vs TC avg
§112
13.4%
-26.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 474 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Status of the Claims Claims 1, 7-9, 11, 17-20 have been amended. Claims 2-3, 12-13 have been canceled and claims 21-24 have been added. Claims 1, 4-11 and 14-24 are pending. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1, 4-5, 8-11, 14-15, 18-23 is/are rejected under 35 U.S.C. 103 as being unpatentable over Santhar et al. (US 2023/0214453) in view of Sommers et al. (US 20240378395) and in further view of Wang et al. “VOYAGER: An Open-Ended Embodied Agent with Large Language Models”. Regarding claim 1, Santhar teaches a method comprising: using a Generative Adversarial Network (GAN) an environment scenario regarding an issue ([0034] “allow the agent to solve a problem”) with an environment and an input request regarding the nvironment scenario, the first test having a first difficulty ([0040], [0044] “generated realistic training environments may include a nominal/default predetermined level of difficulty”; [0058] “to simulate additional real time environment data with an appropriate level of difficulty”); injecting the environment scenario into the environment ([0031] “RL is collected via running an agent in the desired environment”, [0032], [0043], [0055] “generate simulated environment data”); causing, as at least part of performing the first test, a environment scenario injected into the environment by providing the input request to the determining, by a device, how well the Generative Adversarial Network (GAN) updating ([0044] “determination that the discriminator is struggling to achieve an accuracy within the predetermined range, a difficulty of the generated environment may be decreased”; “in response to a determination that the first discriminator has correctly determined whether a generated realistic training environments is real or fake for a predetermined number of iterations, the difficulty may be increased”, [0052]), by the device, the GAN GAN selecting, by the device, a second difficulty for a second test based on how well the GAN initiating, by the device, the second test to assess how well the GAN Santhar does not explicitly teach, however Sommers discloses using a large language model ([0018]) to generate a first test that includes a network scenario regarding an issue with a network and an input request regarding the network scenario ([0007]-[0008] “invoking the LLM to produce configuration instructions for the network test”, [0093]) injecting the network scenario into the network ([0008]“using, by the network test system, the configuration instructions to configure a network test system to conduct the network test; and conducting, by the network test system, the network test”, [0028]-[0029], [0049], [0051], [0065]); causing, as at least part of performing the first test, a large language model-based troubleshooting agent to address the network scenario injected into the network by providing the input request to the large language model-based troubleshooting agent ([0020], [0037], [0040]-[0042]) and determine how well a large language model-based troubleshooting agent for a network was able to perform during a first test ([0074] “verifying, correcting, translating, or otherwise augmenting output from LLM system”, [0083], [0088]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Santhar to use a large language model to generate a first test that includes a network scenario regarding an issue with a network and injecting the network scenario into the network as disclosed by Sommers. Doing may help a test operator generate information usable for configuring or programming a network test system (Sommers [0005]). Santhar as modified does not explicitly teach, however Wang discloses using a large language model to generate a first test (p.4 F4, 2.2) … the first test having a first difficulty (p.2 1 “solve progressively harder tasks proposed by the automatic curriculum, which takes into account the exploration progress and the agent’s state. The curriculum is generated by GPT-4”). Wang discloses determine how well a large language model-based troubleshooting agent for a network was able to perform during a first test (p.3 “error trace from the code interpreter (if any); (2) incorporates the feedback into GPT-4’s prompt for another round of code refinement; and (3) repeats the process until a self-verification module confirms the task completion”, p.5 2.3 (3), last par.). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Santhar to using a large language model to generate a first test having a first difficulty as disclosed by Wang. Doing so would provide ability to handle more sophisticated tasks (Wang p.9(2)). Claim 11 recites substantially the same limitations as claim 1, and is rejected for substantially the same reasons. Regarding claims 4, 14 and 21, Santhar as modified teaches the method and the apparatus, wherein the second test has as higher difficulty than that of the first test, based on the large language model-based troubleshooting agent being able to successfully perform the first test (Santhar [0017] “generate a second realistic environment based on the feedback associated with the first realistic environment”, [0044] “determined that the first discriminator has achieved a predetermined accuracy” aka successful, “in response to a determination that the first discriminator has correctly determined … the difficulty may be increased”, [0051], [0054], [0044], [0049], Wang p.3, p.5 2.3 (3)). Regarding claims 5, 15 and 22, Santhar as modified teaches the method and the apparatus, wherein the device selects the second difficulty for the second test further based on a prediction (Santhar [0051]) that performance of the large language model- based troubleshooting agent during the second test will lead to selection of a third difficulty for a third test (Santhar [0051], [0055], [0058], [0048] “a confidence score to be generated that includes a numerical score of difficulty of the first realistic environment and a determination whether a degree of difficulty incorporated into the first realistic environment is correct … confidence score may be backpropagated as the feedback and used to generate a next realistic environment, and therefore the confidence score may serve as a grade of whether a degree of difficulty incorporated into a most recently considered generated realistic environment is correct”, [0049], [0051] “increases a relative probability that the agent of the RL algorithm receives rewards when navigating the second realistic environment … generating another realistic environment, e.g., a second realistic environment, a third realistic environment, a fourth realistic environment, etc.,” [0052]). Regarding claims 8 and 18, Santhar as modified teaches the method and the apparatus, wherein the device generates the second test using a large language model-based generator (Sommers [0028]-[0029], [0049], [0051], [0065], Wang F9). Regarding claims 9 and 19, Santhar as modified teaches the method and the apparatus, wherein selecting the second difficulty for the second test comprises: using a discriminator to compute a predicted reward value for the first test (Santhar [0044], [0046]-[0047]); and comparing the predicted reward value to an actual reward value that is based on how well the large language model-based troubleshooting agent was able to perform during the first test (Santhar [0048]-[0049], [0058], [0061]). Regarding claim 10, Santhar as modified the method as in claim 1, wherein the device selects the second difficulty for the second test based in part on a history of previously performed tests (Santhar [0051]-[0052], [0062]). NOTE a related art XU et al. (US 20210064515) likewise discloses claim 10 and [0025], [0055], and further obviates the teaching of Santhar as modified. Claim 20 recites substantially the same limitations as claim 1, and is rejected for substantially the same reasons. Claim(s) 6, 16 and 23 is/are rejected under 35 U.S.C. 103 as being unpatentable over Santhar as modified and in further view of DENLI et al. (US 20200183047) and Qiu et al. (US 7606165) or Pennarun et al. (US 20190199772). Regarding claims 6, 16 and 23, Santhar as modified teaches the method and the apparatus, wherein the first test and the second test Santhar as modified does not explicitly teach, however DENLI discloses the method and the apparatus, wherein the first test and the second test comprise a same . Claim(s) 7 and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Santhar as modified and in further view of Yelahanka Raghuprasad et al. (US 20240029417), hereafter YR or Schindel et al. (US 12093374) or Pennarun et al. (US 20190199772) or Qiu et al. (US 7606165). Regarding claims 7 and 17, Santhar as modified does not explicitly teach, however YR or Schindel , Pennarun or Qiu discloses the method and the apparatus, wherein the first test evaluates how well the large language model-based troubleshooting agent was able to troubleshoot the network scenario in the network (YR [0076], [0097], Schindel C10L1-6, C16L47-67, Pennarun [0035], [0059], Qiu C4L59-67, C5L22-25, C7L65-67). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Santhar to configure a particular network scenario in the network as disclosed by YR or Schindel or Pennarun or Qiu. Doing so would provide cybersecurity solutions (Schindel C1L53) and provide an objective measures of addressing network impairments caused by congestion, physics, and various physical layer protocol idiosyncrasies (Pennarun [0003]) and allows for modeling the faults in the simulation independent of the other fault and increase performance of a network (Qiu C4L29-51). Claim(s) 24 and alternatively claims 8 and 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Santhar as modified and in further view of Feng et al. “CITING: LARGE LANGUAGE MODELS CREATE CURRICULUM FOR INSTRUCTION TUNING” or DENLI et al. (US 20200183047) Regarding claim 24, Santhar as modified teaches the tangible, non-transitory, computer-readable medium as in claim 20, wherein the process further includes: determining ground truth information based at least in part on instantiating the network scenario in the network (Santhar [0059] “incorporated, e.g., as a Ground Truth, in the process of training the environment generator”, F4:402) Santhar as modified does not explicitly teach, however Feng or DENLI disclose wherein the determining how well the large language model-based troubleshooting agent was able to perform during the first test is based on the ground truth information (Feng p.3, p.5 Algorithm 1, F4, DENLI [0080]. [0083], [0088]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Santhar to include evaluating large language model-based on ground truth as disclosed by Feng. Doing so would provide substantial improvements in terms of articulation, depth, and comprehensiveness for advancing large language models (Feng pp.9-10 ¶5) and may help discern the difference between the generated data and ground truth (DENLI [0083]). Regarding claims 8 and 18, if Santhar as modified teaches claims as disclosed above, Feng further discloses the method and the apparatus, wherein the device generates the second test using a large language model-based generator (Feng Abstract “(1) the teacher LLM crafts the rubrics for evaluating the answers corresponding to various types of questions, and (2) the student LLM learns to follow the rubrics and perform self-correction from the revision made by the teacher”, p.3. F3 “LLM teaches the student LLM to revise its initial response based on the rubric”, p.4 “Curriculum Instruction Tuning for Student LLM. The student LLM learns to revise its initial response from teacher LLM based on the criteria from the first part via instruction turning,” p.4 ¶3.3, p.5 Algorithm 1 CITING). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Santhar to include evaluating large language model-based on ground truth as disclosed by Feng. Doing so would provide substantial improvements in terms of articulation, depth, and comprehensiveness for advancing large language models (Feng pp.9-10 ¶5) Response to Arguments Applicant's arguments, filed 08/07/2026, in regard to the presently amended claims have been fully considered and are addressed in the updated rejections to the claims above. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to POLINA G PEACH whose telephone number is (571)270-7646. The examiner can normally be reached Monday-Friday, 9:30 - 5:30. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Aleksandr Kerzhner can be reached at 571-270-1760. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /POLINA G PEACH/ Primary Examiner, Art Unit 2165 August 27, 2026
Read full office action

Prosecution Timeline

Nov 08, 2023
Application Filed
May 07, 2026
Non-Final Rejection mailed — §103
Jul 13, 2026
Interview Requested
Jul 27, 2026
Applicant Interview (Telephonic)
Jul 27, 2026
Examiner Interview Summary
Aug 07, 2026
Response Filed
Sep 01, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748791
SCALABLE GENERATIVE AI-BASED TOOL INFRASTRUCTURE
2y 4m to grant Granted Sep 29, 2026
Patent 12717811
TRAINING AND UTILIZING LANGUAGE MACHINE LEARNING MODELS TO CREATE STRUCTURED OUTPUTS FOR BUILDING DIGITAL VISUALIZATIONS FROM ANALYTICS DATABASES AND DIGITAL TEXT PROMPTS
2y 7m to grant Granted Aug 25, 2026
Patent 12705277
Method of Modifying Map Data, and Machine-Readable Instruction Code
1y 7m to grant Granted Aug 11, 2026
Patent 12699733
GRAPH-BASED LABELING OF HETEROGENOUS DIGITAL CONTENT ITEMS
4y 10m to grant Granted Aug 04, 2026
Patent 12675458
FILE INDEXING FOR VIRTUAL MACHINE BACKUPS IN A DATA STORAGE MANAGEMENT SYSTEM
1y 7m to grant Granted Jul 07, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
50%
Grant Probability
74%
With Interview (+23.3%)
3y 9m (~10m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 474 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month