Prosecution Insights
Last updated: October 02, 2026
Application No. 17/852,757

EVALUATION AND ADAPTIVE SAMPLING OF AGENT CONFIGURATIONS

Final Rejection §103
Filed
Jun 29, 2022
Examiner
HONORE, EVEL NMN
Art Unit
2142
Tech Center
2100 — Computer Architecture & Software
Assignee
Microsoft Technology Licensing, LLC
OA Round
2 (Final)
52%
Grant Probability
Moderate
3-4
OA Rounds
0m
Est. Remaining
78%
With Interview

Examiner Intelligence

Grants 52% of resolved cases
52%
Career Allowance Rate
14 granted / 27 resolved
-3.1% vs TC avg
Strong +26% interview lift
Without
With
+26.4%
Interview Lift
resolved cases with interview
Typical timeline
4y 2m
Avg Prosecution
27 currently pending
Career history
59
Total Applications
across all art units

Statute-Specific Performance

§101
34.2%
-5.8% vs TC avg
§103
59.0%
+19.0% vs TC avg
§102
5.9%
-34.1% vs TC avg
§112
0.6%
-39.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 27 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION This action is responsive to the Application filed on 05/11/2026 Claims 1-18 and 21-22 are pending in the case. Claims 1, 13 and 22 are independent claims. Claims 19-20 are cancelled . Claims 1-2, 6, 8-9 and 13 are currently amended. Claims 21-22 are newly added. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 2, 7, 10-11, 13-14, 16-17 and 21 are rejected under 35 U.S.C 103 as being unpatentable over Cay et al. (US Patent No. 11,055,639 B2 ), hereinafter referred to as Cay in view of Yu. (US Pub No.: 20190228309 A1), hereinafter referred to as Yu. With respect to claim 1, Cay disclose: A method comprising: performing two or more data gathering iterations comprising (In Fig. 15 and Col. 34, lines 21–27, Cay discloses a system 1500 for executing multiple iterations of the iterative process in parallel. In Cols. 34-35, lines 63-19, Cay discloses multiple instances 1516a-n of a trained machine-learning model running parallel. The optimization manager sends respective sets of candidate values to those ML-model instances; each instance evaluates its respective candidate values, produces output values, and returns those values to the optimization manager. The returned values are then provided back to the optimization model for subsequent use during the parallel iterations.) Distributing experimental units to a plurality of agents having different agent configurations, the experimental units being distributed according to a sampling strategy, wherein the plurality of agents are configured to select actions to take in an environment by executing a machine learning model based at least on the experimental units (In Fig. 15 and Col. 34, lines 21–27, Cay discloses a system 1500 for executing multiple iterations of the iterative process in parallel. In Cols. 34-35, lines 63-19, Cay discloses multiple instances 1516a-n of a trained machine-learning model running parallel. The optimization manager sends respective sets of candidate values to those Ml-model instances; each instance evaluates its respective candidate values, produces output values, and returns those values to the optimization manager. The returned values are then provided back to the optimization model for subsequent use during the parallel iterations.) Based at least on the events in the event log, predicting performance of the plurality of agents with respect to one or more evaluation metrics (based at least on the events in the event log, predicting performance of the plurality of agents with respect to one or more evaluation metrics.) Based at least on predicted performance of the plurality of agents with respect to the one or more evaluation metrics, identifying a selected agent configuration (In Col.33, lines 47–67, Cay discloses that the trained machine learning model can then return the predicted value to the processing device. The processing device determines if the value satisfies a predefined constraint.) With respect to claim 1, Cay do not explicitly disclose: Populating an event log with events representing reactions of the environment to respective actions taken by individual agents in response to individual experimental units Based at least on the events in the event log, adjusting the sampling strategy for use in distributing further experimental units to the plurality of agents for processing with the machine learning model in a subsequent data gathering iteration Deploying a selected agent in the selected agent configuration, the selected agent taking further actions in the environment by executing the machine learning model according to the selected agent configuration However, it is known by Yu disclose: Populating an event log with events representing reactions of the environment to respective actions taken by individual agents in response to individual experimental units (In paragraph [0386], Yu discloses that during each iteration, the system deploys the most recently approved/confirmed set of policies. Those policies interact with the environment and collect trajectories. The trajectories are distributed uniformly across different policies. The trajectories collected from each individual policy are divided and stored in: a training trajectory set, and a testing trajectory set. ) Based at least on the events in the event log, adjusting the sampling strategy for use in distributing further experimental units to the plurality of agents for processing with the machine learning model in a subsequent data gathering iteration (In paragraph [0417], Yu discloses a sampling allocation strategy in which a vector k specifies the respective number of samples selected from each of plurality of proposal distributions, thereby controlling distribution of a total collection of samples among the plurality of proposal distributions.) Deploying a selected agent in the selected agent configuration, the selected agent taking further actions in the environment by executing the machine learning model according to the selected agent configuration (In paragraphs [0374-0380], Yu disclose selecting and deploying a plurality of sage behaviors policies for an autonomous agent based on expected policy performance, and iteratively improving and deploying the policies during successive policy-improvement iterations using diverse exploration strategies. ) Cay in view of Yu are analogous pieces of art because both are directed to machine-learning-based iterative optimization in which candidate behaviors/configuration are evaluated and subsequent is changed based on evaluation. Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Cay, with include selecting a current set of candidate values for the configurable settings from within a current region of a search space defined by the optimization model as taught by Cay, with policies for autonomous agent using reinforcement learning. The motivation for doing so would have been to determine to improve the accuracy of the resultant set of values (See (Col.31, lines 66-67) of Cay.) Regarding claim 2, Cay, in view of Yu disclose the elements of claim 1. In addition, Yu disclose: The method of claim 1, further comprising: training the machine learning model based at least on the reactions of the environment to the respective actions taken by the individual agents in response to the individual experimental units (In paragraphs [0386-0388], Yu disclose collected trajectories from respective deployed behavior policies, wherein the trajectories comprise experience sample including a state, an action, and a reward resulting from an action; portioning the collected trajectories into a training set; and subsequently learning candidate policies from the training trajectories. Thus, the training uses information representing the environment's reaction to action performed according to the respective deployed policies.) Regarding claim 7, Cay, in view of Yu disclose the elements of claim 1. In addition, Cay disclose: The method of claim 1, wherein the sampling strategy is adjusted at each data gathering iteration based at least on the predicted performance of the plurality of agents with respect to the one or more evaluation metrics (In Col. 33, lines 25-43, Cay discloses providing a set of candidate values as input to a trained machine-learning model, which processes the candidate values to predict a value for a target characteristic to be optimized and returns the predicted value to the processing device.) Regarding claim 10, Cay, in view of Yu disclose the elements of claim 7. In addition, Cay disclose: The method of claim 7, further comprising: populating a data structure with predicted aggregate values and corresponding confidence intervals for the one or more evaluation metrics (In Col. 33, lines 25–43, Cay discloses providing a set of candidate values as input to a trained machine-learning model, which processes the candidate values to predict a value for a target characteristic to be optimized and returns the predicted value to the processing device.) Outputting a graphical representation of the data structure (In Col. 32,lines 36–51, Cay discloses that an operator of the manufacturing process can view the graphical user interface on the display device and tune the configurable settings to the recommended set of values, to improve the manufacturing process.) Identifying one or more agent configurations to sample in a subsequent data gathering iteration based at least on user input directed to the graphical representation of the data structure (In Col.33, lines 47–67, Cay discloses that the trained machine learning model can then return the predicted value to the processing device. The processing device determines if the value satisfies a predefined constraint.) Regarding claim 11, Cay, in view of Yu disclose the elements of claim 10. In addition, Cay disclose: The method of claim 10, further comprising: receiving user input specifying two or more evaluation metrics (In Fig. 15 and Col. 34, lines 21–27, Cay discloses a system 1500 for executing multiple iterations of the iterative process in parallel. In Cols. 34-35, lines 63-19, Cay discloses multiple instances 1516a-n of a trained machine-learning model running parallel. The optimization manager sends respective sets of candidate values to those ML-model instances; each instance evaluates its respective candidate values, produces output values, and returns those values to the optimization manager. The returned values are then provided back to the optimization model for subsequent use during the parallel iterations.) Generating the graphical representation based at least on the two or more evaluation metrics specified by the user input (In Col. 32,lines 36–51, Cay discloses that an operator of the manufacturing process can view the graphical user interface on the display device and tune the configurable settings to the recommended set of values, to improve the manufacturing process. In Col. 33, lines 25-43, Cay discloses providing a set of candidate values as input to a trained machine-learning model, which processes the candidate values to predict a value for a target characteristic to be optimized and returns the predicted value to the processing device.) With respect to claim 13, Cay disclose: A system comprising: a processor (In Col. 25, lines 46-49, Cay disclose the computing device includes, but is not limited to, one or more processors and one or more computer-readable mediums operably coupled to the one or more processor) A storage resource storing instructions which, when executed by the processor, cause the system to: perform two or more data gathering iterations comprising (In Fig. 15 and Col. 34, lines 21–27, Cay discloses a system 1500 for executing multiple iterations of the iterative process in parallel. In Cols. 34-35, lines 63-19, Cay discloses multiple instances 1516a-n of a trained machine-learning model running parallel. The optimization manager sends respective sets of candidate values to those ML-model instances; each instance evaluates its respective candidate values, produces output values, and returns those values to the optimization manager. The returned values are then provided back to the optimization model for subsequent use during the parallel iterations.) Distributing experimental units to a plurality of agents having different agent configurations, the experimental units being distributed according to a sampling strategy, wherein the plurality of agents are configured to select actions to take in an environment by executing a machine learning model based at least on the experimental units (In Fig. 15 and Col. 34, lines 21–27, Cay discloses a system 1500 for executing multiple iterations of the iterative process in parallel. In Cols. 34-35, lines 63-19, Cay discloses multiple instances 1516a-n of a trained machine-learning model running parallel. The optimization manager sends respective sets of candidate values to those Ml-model instances; each instance evaluates its respective candidate values, produces output values, and returns those values to the optimization manager. The returned values are then provided back to the optimization model for subsequent use during the parallel iterations.) Based at least on the events in the event log, predicting performance of the plurality of agents with respect to one or more evaluation metrics (based at least on the events in the event log, predicting performance of the plurality of agents with respect to one or more evaluation metrics.) Based at least on predicted performance of the plurality of agents with respect to the one or more evaluation metrics, identifying a selected agent configuration (In Col.33, lines 47–67, Cay discloses that the trained machine learning model can then return the predicted value to the processing device. The processing device determines if the value satisfies a predefined constraint.) With respect to claim 13, Cay do not explicitly disclose: Populating an event log with events representing reactions of [an] the environment to respective actions taken by individual agents in response to individual experimental units Based at least on the events in the event log, adjusting the sampling strategy for use in distributing further experimental units to the plurality of agents for processing with the machine learning model in a subsequent data gathering iteration Deploying a selected agent in the selected agent configuration, the selected agent taking further actions in the environment by executing the machine learning model according to the selected agent configuration However, it is known by Yu disclose: Populating an event log with events representing reactions of the environment to respective actions taken by individual agents in response to individual experimental units (In paragraph [0386], Yu discloses that during each iteration, the system deploys the most recently approved/confirmed set of policies. Those policies interact with the environment and collect trajectories. The trajectories are distributed uniformly across different policies. The trajectories collected from each individual policy are divided and stored in: a training trajectory set, and a testing trajectory set. ) Based at least on the events in the event log, adjusting the sampling strategy for use in distributing further experimental units to the plurality of agents for processing with the machine learning model in a subsequent data gathering iteration (In paragraph [0417], Yu discloses a sampling allocation strategy in which a vector k specifies the respective number of samples selected from each of plurality of proposal distributions, thereby controlling distribution of a total collection of samples among the plurality of proposal distributions.) Deploying a selected agent in the selected agent configuration, the selected agent taking further actions in the environment by executing the machine learning model according to the selected agent configuration (In paragraphs [0374-0380], Yu disclose selecting and deploying a plurality of sage behaviors policies for an autonomous agent based on expected policy performance, and iteratively improving and deploying the policies during successive policy-improvement iterations using diverse exploration strategies. ) Cay in view of Yu are analogous pieces of art because both are directed to machine-learning-based iterative optimization in which candidate behaviors/configuration are evaluated and subsequent is changed based on evaluation. Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Cay, with include selecting a current set of candidate values for the configurable settings from within a current region of a search space defined by the optimization model as taught by Cay, with policies for autonomous agent using reinforcement learning. The motivation for doing so would have been to determine to improve the accuracy of the resultant set of values (See (Col.31, lines 66-67) of Cay.) Regarding claim 14, Cay, in view of Yu disclose the elements of claim 13. In addition, Yu disclose: The system of claim 13, wherein the individual agents include machine learning agents having different hyperparameters or different feature definitions (In paragraph [0523], Yu discloses variance parameters, the empirical return of the trajectory is used as the Q component and V as a baseline. TRPO hyperparameters) Regarding claim 16, Cay, in view of Yu disclose the elements of claim 13. In addition, Yu disclose: The system of claim 13, wherein the sampling strategy is based at least on respective importance weights of the individual agents (In paragraph [0417], Yu discloses a sampling allocation strategy in which a vector k specifies the respective number of samples selected from each of plurality of proposal distributions, thereby controlling distribution of a total collection of samples among the plurality of proposal distributions.) Regarding claim 17, Cay, in view of Yu disclose the elements of claim 13. In addition, Cay disclose: The system of claim 13, wherein the sampling strategy is adjusted based at least on predicted performance of the plurality of agents with respect to the one or more evaluation metrics (In Col. 32, lines 18–22, Cay discloses that the electronic communication can be configured to cause the configurable settings to be adjusted to the recommended set of values.) With respect to claim 21, Cay disclose: A computer-readable storage medium storing instructions which, when executed by a computing device, cause the computing device to perform acts comprising: performing two or more data gathering iterations comprising: (In Col. 25, lines 46-49, Cay disclose the computing device includes, but is not limited to, one or more processors and one or more computer-readable mediums operably coupled to the one or more processor. In Fig. 15 and Col. 34, lines 21–27, Cay discloses a system 1500 for executing multiple iterations of the iterative process in parallel. In Cols. 34-35, lines 63-19, Cay discloses multiple instances 1516a-n of a trained machine-learning model running parallel. The optimization manager sends respective sets of candidate values to those ML-model instances; each instance evaluates its respective candidate values, produces output values, and returns those values to the optimization manager. The returned values are then provided back to the optimization model for subsequent use during the parallel iterations.) Distributing experimental units to a plurality of agents having different agent configurations, the experimental units being distributed according to a sampling strategy, wherein the plurality of agents are configured to select actions to take in an environment by executing a machine learning model based at least on the experimental units (In Fig. 15 and Col. 34, lines 21–27, Cay discloses a system 1500 for executing multiple iterations of the iterative process in parallel. In Cols. 34-35, lines 63-19, Cay discloses multiple instances 1516a-n of a trained machine-learning model running parallel. The optimization manager sends respective sets of candidate values to those Ml-model instances; each instance evaluates its respective candidate values, produces output values, and returns those values to the optimization manager. The returned values are then provided back to the optimization model for subsequent use during the parallel iterations.) Based at least on the events in the event log, predicting performance of the plurality of agents with respect to one or more evaluation metrics (based at least on the events in the event log, predicting performance of the plurality of agents with respect to one or more evaluation metrics.) Based at least on predicted performance of the plurality of agents with respect to the one or more evaluation metrics, identifying a selected agent configuration (In Col.33, lines 47–67, Cay discloses that the trained machine learning model can then return the predicted value to the processing device. The processing device determines if the value satisfies a predefined constraint.) With respect to claim 21, Cay do not explicitly disclose: Populating an event log with events representing reactions of the environment to respective actions taken by individual agents in response to individual experimental units Based at least on the events in the event log, adjusting the sampling strategy for use in distributing further experimental units to the plurality of agents for processing with the machine learning model in a subsequent data gathering iteration Deploying a selected agent in the selected agent configuration, the selected agent taking further actions in the environment by executing the machine learning model according to the selected agent configuration However, it is known by Yu disclose: Populating an event log with events representing reactions of the J[an]l environment to respective actions taken by individual agents in response to individual experimental units (In paragraph [0386], Yu discloses that during each iteration, the system deploys the most recently approved/confirmed set of policies. Those policies interact with the environment and collect trajectories. The trajectories are distributed uniformly across different policies. The trajectories collected from each individual policy are divided and stored in: a training trajectory set, and a testing trajectory set. ) Based at least on the events in the event log, adjusting the sampling strategy for use in distributing further experimental units to the plurality of agents for processing with the machine learning model in a subsequent data gathering iteration (In paragraph [0417], Yu discloses a sampling allocation strategy in which a vector k specifies the respective number of samples selected from each of plurality of proposal distributions, thereby controlling distribution of a total collection of samples among the plurality of proposal distributions.) Deploying a selected agent in the selected agent configuration, the selected agent taking further actions in the environment by executing the machine learning model according to the selected agent configuration (In paragraphs [0374-0380], Yu disclose selecting and deploying a plurality of sage behaviors policies for an autonomous agent based on expected policy performance, and iteratively improving and deploying the policies during successive policy-improvement iterations using diverse exploration strategies. ) Cay in view of Yu are analogous pieces of art because both are directed to machine-learning-based iterative optimization in which candidate behaviors/configuration are evaluated and subsequent is changed based on evaluation. Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Cay, with include selecting a current set of candidate values for the configurable settings from within a current region of a search space defined by the optimization model as taught by Cay, with policies for autonomous agent using reinforcement learning. The motivation for doing so would have been to determine to improve the accuracy of the resultant set of values (See (Col.31, lines 66-67) of Cay.) Claims 3 and 5 are rejected under 35 U.S.C 103 as being unpatentable over Cay in view of Yu and further in view of Ma. (US Pub No.: 20190258904 A1), hereinafter referred to as Ma. Regarding claim 3, Cay in view of Yu disclose element of claim 2. Cay in view of Yu do not explicitly disclose: The method of claim 2, wherein the selected agent configuration is selected automatically or based on user input identifying the selected agent configuration from a graphical representation of the predicted performance of the plurality of agents However, Ma disclose the limitation (In Fig. 3 and paragraph [0155], MA discloses a predictive model selection device that computes performance values for a plurality of prediction models, displays those model performance values in a graph and a table, and uses performance results to decide which model to select.) Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention having the teachings of Cay in view of Yu before them, to include Ma, with a plurality of predictive type models are trained, validated, and scored using the samples to select a best predictive model as taught by Ma. The motivation for doing so would have been to select and process the next predictive type model of the plurality of predictive type models (See [0146] of Ma.) Regarding claim 5, Cay in view of Yu disclose element of claim 4. Cay in view of Yu do not explicitly disclose: The method of claim 4, further comprising: calculating respective sampling probabilities for the individual agents based at least on the importance weights However, Ma disclose the limitation (In Fig. 17 and paragraph [0153], Ma discloses a table of illustrative performance values computed by the predictive model. ) Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention having the teachings of Cay in view of Yu before them, to include Ma, with a plurality of predictive type models are trained, validated, and scored using the samples to select a best predictive model as taught by Ma. The motivation for doing so would have been to select and process the next predictive type model of the plurality of predictive type models (See [0146] of Ma.) Claims 6, 9 and 21 are rejected under 35 U.S.C 103 as being unpatentable over Cay in view of Yu and further in view of Molchanov. (US Pub No.: 20180114114 A1), hereinafter referred to as Molchanov. Regarding claim 6, Cay in view of Yu disclose element of claim 4. Cay in view of Yu do not explicitly disclose: The method of claim 4, wherein adjusting the sampling strategy comprises: removing at least one agent from subsequent data gathering iterations based at least on the importance weights wherein the at least one agent ceases using processing and memory resources to execute the machine learning model after being removed from the subsequent data gathering iterations However, Molchanov disclose the limitation (In paragraph [0042], Molchanov discloses calculating a pruning criterion representing the importance of each neuron, identifies the neuron having the lowest importance, and removes that neuron from the trained neural network.) Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention having the teachings of Cay in view of Yu before them, to include Molchanov, with identifying at least one neuron having a lowest importance and removing the at least one neuron from the trained neural network as taught by Molchanov. The motivation for doing so would have been to fine-tune an existing deep network previously trained on a much larger labeled vision dataset (See [0003] of Molchanov.) Regarding claim 9, Cay in view of Yu disclose element of claim 7. Cay in view of Yu do not explicitly disclose: The method of claim 7, wherein adjusting the sampling strategy comprises: removing at least one agent from subsequent data gathering iterations based at least on the predicted performance wherein the at least one agent ceases using processing and memory resources to execute the machine learning model after being removed from the subsequent data gathering iterations However, Molchanov disclose the limitation (In paragraph [0042], Molchanov discloses calculating a pruning criterion representing the importance of each neuron, identifies the neuron having the lowest importance, and removes that neuron from the trained neural network.) Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention having the teachings of Cay in view of Yu before them, to include Molchanov, with identifying at least one neuron having a lowest importance and removing the at least one neuron from the trained neural network as taught by Molchanov. The motivation for doing so would have been to fine-tune an existing deep network previously trained on a much larger labeled vision dataset (See [0003] of Molchanov.) Regarding claim 21, Cay in view of Yu disclose element of claim 1. Cay in view of Yu do not explicitly disclose: The method of claim 1, wherein adjusting the sampling strategy reduces utilization of processor and memory resources for at least one of the agent configurations in the subsequent data gathering iteration However, Molchanov disclose the limitation (In paragraph [0042], Molchanov discloses calculating a pruning criterion representing the importance of each neuron, identifies the neuron having the lowest importance, and removes that neuron from the trained neural network.) Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention having the teachings of Cay in view of Yu before them, to include Molchanov, with identifying at least one neuron having a lowest importance and removing the at least one neuron from the trained neural network as taught by Molchanov. The motivation for doing so would have been to fine-tune an existing deep network previously trained on a much larger labeled vision dataset (See [0003] of Molchanov.) Claim 8 is rejected under 35 U.S.C 103 as being unpatentable over Cay in view of Yu and further in view of Qiu et al. (US Pub No.: 20190244110 A1), hereinafter referred to as Qiu. Regarding claim 8, Cay in view of Yu disclose element of claim 7. Cay in view of Yu do not explicitly disclose: The method of claim 7, wherein adjusting the sampling strategy comprises: determining respective confidence intervals of the one or more evaluation metrics for the individual agents; and calculating sampling probabilities of the individual agents based at least on upper bounds of the confidence intervals However, Qiu disclose the limitation (In paragraph [0103], Qiu discloses an empirical average reward based on samples taken from the arm. It then derives an upper confidence bound using the Cherboff-Hoeffding inequality. Thus, each candidate/arm has its own performance estimate and corresponding confidence-bound component.) Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention having the teachings of Cay in view of Yu before them, to include Qiu, with quantitative or numerical data type, qualitative data type, discreet data type, continuous data type (with lower and upper bounds) as taught by Qiu. The motivation for doing so would have been to improve the optimization performance in terms of best conversion rate (See [0253] of Qiu) Claim 15 is rejected under 35 U.S.C 103 as being unpatentable over Cay in view of Yu and further in view of Silver et al. (US Patent No. 11,627,165 B2 ), hereinafter referred to as Silver. Regarding claim 15, Cay in view of Yu disclose element of claim 13. Cay in view of Yu do not explicitly disclose: The system of claim 13, wherein the individual agents include at least two different reinforcement learning agents having different reward functions, at least two different supervised learning agents having different loss functions, and at least two different rule-based agents having different rules However, Silver disclose the limitation (In Col. 10, lines 24-36, Silver discloses that each learner policy has a respective matchmaking policy and the matchmaking policies (and corresponding distributions) for two or more of the learner policies are typically different.) Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention having the teachings of Cay in view of Yu before them, to include Silver, with policy parameters define policies to be used in controlling the agent to perform the particular task as taught by Silver. The motivation for doing so would have been to improve overall quality of the training by providing better learning signals, when generating training data for any given one of the learner policies (See (Col. 7, lines 38-40) of Silver.) Claim 18 is rejected under 35 U.S.C 103 as being unpatentable over Cay in view of Yu and further in view of Brooks. (US Pub No.: 20210390401 A1), hereinafter referred to as Brooks. Regarding claim 18, Cay in view of Yu disclose element of claim 13. Cay in view of Yu do not explicitly disclose: The system of claim 13, wherein adjusting the sampling strategy comprises assigning respective probabilities to the individual agents and randomly assigning the experimental units to the individual agents based on the respective probabilities However, Brooks disclose the limitation (In paragraph [0046], Brooks discloses assigning respective allocation probabilities to a plurality of alternatives and performing controlled random assignment of experimental units to the alternatives according to the assigned probability distribution, wherein the relative frequency at which experimental units are assigned to each alternative corresponds to the respective probability specified by an explore/exploit process.) Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention having the teachings of Cay in view of Yu before them, to include Brooks, with repeatedly generate self-organizing experimental units (SOEUs) based on the one or more assumptions as taught by Brooks. The motivation for doing so would have been to improve upon the clustering of experimental units process and explore/exploit management process (See [0062] of Brooks.) Response to Arguments The applicant's arguments filed 05/11/2026 have been fully considered, but in part are not persuasive. Pertaining to Rejection under 101 Rejections for claims 1-18 and 21-22 are withdrawn under 35 USC § 101. Pertaining to Rejection under 103 Applicant’s arguments in regard to the examiner’s rejections under 35 USC 103 are moot in view of the new grounds of rejection. Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to EVEL HONORE whose telephone number is (703)756-1179. The examiner can normally be reached Monday-Friday 8 a.m. -5:30 p.m. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Mariela D Reyes can be reached at (571) 270-1006. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. EVEL HONORE Examiner Art Unit 2142 /Mariela Reyes/Supervisory Patent Examiner, Art Unit 2142
Read full office action

Prosecution Timeline

Jun 29, 2022
Application Filed
Dec 11, 2025
Non-Final Rejection mailed — §103
Mar 04, 2026
Examiner Interview Summary
Mar 04, 2026
Applicant Interview (Telephonic)
May 11, 2026
Response Filed
Sep 14, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737632
METHOD AND DEVICE FOR TRAINING A MACHINE LEARNING SYSTEM
4y 11m to grant Granted Sep 15, 2026
Patent 12725029
INFORMATION PROCESSING APPARATUS, INFORMATION PROCESSING METHOD, AND STORAGE MEDIUM
5y 6m to grant Granted Sep 01, 2026
Patent 12725056
GENERATING MACHINE LEARNING BASED MODELS FOR TIME SERIES FORECASTING
4y 11m to grant Granted Sep 01, 2026
Patent 12694284
NODE, AND METHOD PERFORMED THEREBY, FOR PREDICTING A BEHAVIOR OF USERS OF A COMMUNICATIONS NETWORK
5y 1m to grant Granted Jul 28, 2026
Patent 12657480
SYSTEM AND METHOD FOR REDUCTION OF DATA TRANSMISSION BY INFERENCE OPTIMIZATION AND DATA RECONSTRUCTION
4y 1m to grant Granted Jun 16, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
52%
Grant Probability
78%
With Interview (+26.4%)
4y 2m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 27 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month