Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Detailed Action
The following action is in response to the communication(s) received on 05/16/2024.
As of the claims filed 05/16/2024:
Claims 1-20 are pending.
Claims 1, 19, and 20 are independent claims.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 09/25/2025 was filed in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Objections
Claim 2 is objected to because of the following informalities:
Claim 2 recites “transacting in an item”, which appears to be an incorrect phrase. For purposes of examination, this phrase has been interpreted as “transacting an item”
Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 9 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
The term “accurately training” in claim 9 is a relative term which renders the claim indefinite. The term “accurately training” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-12 and 15-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Claim 1 recites a system, thus a machine, one of the four statutory categories of patentable subject matter (Step 1). However, Claim 1 further recites:
generating a set of reward functions associated with training a reinforcement learning (RL) agent to predict an objective; which is an evaluation or judgement that can be performed in the human mind;
generating a score representing accuracy of the predicted objective for the one or more instances of the RL agent, which is an evaluation or judgement that can be performed in the human mind;
and selecting an individual instance of the one or more instances of the RL agent to predict the objective based on the generated score, which is an evaluation or judgement that can be performed in the human mind.
Thus, the claim recites an abstract idea under Step 2A Prong 1.Under Step 2A Prong 2, the claim recites:
comprising: one or more hardware processors; and at least one machine-storage medium for storing instructions that, when executed by the one or more hardware processors, cause the one or more hardware processors to perform operations comprising, as the performance of an abstract idea on a computer is not more than instructions to 'apply it' on a computer, which by MPEP 2106.05(f) cannot integrate an abstract idea into a practical application;
providing a prompt to a large language model (LLM), the prompt comprising a set of instructions for..., as the performance of an abstract idea on a computer is not more than instructions to 'apply it' on a computer, which by MPEP 2106.05(f) cannot integrate an abstract idea into a practical application;
obtaining the set of reward functions from the LLM, which is merely an insignificant extra-solution activity of data gathering, which by MPEP 2106.05(g) cannot integrate an abstract idea into a practical application;
training one or more instances of the RL agent using the set of reward functions to predict the objective, as the performance of an abstract idea on a computer is not more than instructions to 'apply it' on a computer, which by MPEP 2106.05(f) cannot integrate an abstract idea into a practical application.
Thus, the claim is directed towards an abstract idea.
Further, the additional element(s), alone or in combination, do not provide significantly more than the abstract idea itself, because the activity of data gathering (MPEP 2106.05(g)) cannot provide significantly more, as storing and retrieving information in memory is well understood, routine, and conventional (MPEP 2106.05(d)(II)(iv)); implementation on a computer (MPEP 2106.05(f)) cannot provide significantly more. The combination of these additional elements does not provide an inventive concept; thus, the claim remains ineligible.
Claim 2, dependent on 1, further recites
the objective comprises a risk associated with a user in transacting in an item in an electronic marketplace, which is merely a detail of an abstract idea (predict an objective).
Thus, the claim recites an abstract idea under Step 2A Prong 1.Under Step 2A Prong 2, the claim recites no additional elements which could integrate the abstract idea into a practical application or provide significantly more than the abstract idea itself; thus, the claim remains ineligible.
Claim 3, dependent on 2, further recites
the risk comprises a likelihood of unauthorized chargeback associated with the user, which is merely a detail of an abstract idea (predict an objective).
Thus, the claim recites an abstract idea under Step 2A Prong 1.Under Step 2A Prong 2, the claim recites no additional elements which could integrate the abstract idea into a practical application or provide significantly more than the abstract idea itself; thus, the claim remains ineligible.
Claim 4, dependent on 1, further recites
analyzing a plurality of user features, which is an evaluation or judgement that can be performed in the human mind.
Thus, the claim recites an abstract idea under Step 2A Prong 1.Under Step 2A Prong 2, the claim recites no additional elements which could integrate the abstract idea into a practical application or provide significantly more than the abstract idea itself; thus, the claim remains ineligible.
Claim 5, dependent on 4, further recites
the plurality of user features comprises at least one of velocity of transactions, type of financial instrument being used by a user, type of device being used by the user, a registration date associated with the user, or collusive behavior information between the user and another user, which is merely a detail of an abstract idea (analyzing a plurality of user features).
Thus, the claim recites an abstract idea under Step 2A Prong 1.Under Step 2A Prong 2, the claim recites no additional elements which could integrate the abstract idea into a practical application or provide significantly more than the abstract idea itself; thus, the claim remains ineligible.
Claim 6, dependent on 1, further recite. no additional abstract ideas. However: Under Step 2A Prong 2, the claim recites:
the RL agent comprises a multilayer neural network machine learning (ML) model, as the performance of an abstract idea on a computer is not more than instructions to 'apply it' on a computer, which by MPEP 2106.05(f) cannot integrate an abstract idea into a practical application.
Thus, the claim is directed towards an abstract idea.
Further, the additional element(s), alone or in combination, do not provide significantly more than the abstract idea itself, because implementation on a computer (MPEP 2106.05(f)) cannot provide significantly more. The combination of these additional elements does not provide an inventive concept; thus, the claim remains ineligible.
Claim 7, dependent on 1, further recites
determining that the score representing the accuracy of the predicted objective is greater than a threshold value, which is an evaluation or judgement that can be performed in the human mind.
Thus, the claim recites an abstract idea under Step 2A Prong 1.Under Step 2A Prong 2, the claim recites:
concluding a training process of the RL agent in response, as the performance of an abstract idea on a computer is not more than instructions to 'apply it' on a computer, which by MPEP 2106.05(f) cannot integrate an abstract idea into a practical application.
Thus, the claim is directed towards an abstract idea.
Further, the additional element(s), alone or in combination, do not provide significantly more than the abstract idea itself, because implementation on a computer (MPEP 2106.05(f)) cannot provide significantly more. The combination of these additional elements does not provide an inventive concept; thus, the claim remains ineligible.
Claim 8, dependent on 1, further recites
the objective predicted by the individual instance of the one or more instances of the RL agent comprises a first likelihood of fraudulent activity before authorizing an electronic transaction, a second likelihood of fraudulent activity after authorizing the electronic transaction, and a third likelihood of fraudulent activity associated with delay capture, which is merely a detail of an abstract idea (predict an objective).
Thus, the claim recites an abstract idea under Step 2A Prong 1.Under Step 2A Prong 2, the claim recites no additional elements which could integrate the abstract idea into a practical application or provide significantly more than the abstract idea itself; thus, the claim remains ineligible.
Claim 9, dependent on 1, further recites
removing one or more reward functions from the set of reward functions in response to determining that the one or more reward functions are incapable of accurately training the one or more instances of the RL agent, which is an evaluation or judgement that can be performed in the human mind;
Thus, the claim recites an abstract idea under Step 2A Prong 1.Under Step 2A Prong 2, the claim recites no additional elements which could integrate the abstract idea into a practical application or provide significantly more than the abstract idea itself; thus, the claim remains ineligible.
Claim 10, dependent on 1, further recite. no additional abstract ideas. However: Under Step 2A Prong 2, the claim recites:
training a first instance of the RL agent using a first reward function in the set of reward functions; and training, in parallel with training the first instance, a second instance of the RL agent using a second reward function in the set of reward functions, as the performance of an abstract idea (e.g., generating a set of reward functions associated with training a reinforcement learning (RL) agent to predict an objective) on a computer is not more than instructions to 'apply it' on a computer, which by MPEP 2106.05(f) cannot integrate an abstract idea into a practical application.
Thus, the claim is directed towards an abstract idea.
Further, the additional element(s), alone or in combination, do not provide significantly more than the abstract idea itself, because implementation on a computer (MPEP 2106.05(f)) cannot provide significantly more. The combination of these additional elements does not provide an inventive concept; thus, the claim remains ineligible.
Claim 11, dependent on 10, further recites
predict a first objective associated with the set of training data, which is an evaluation or judgement that can be performed in the human mind;
predict a second objective associated with the set of training data, which is an evaluation or judgement that can be performed in the human mind;
evaluating the first and second objectives based on ground truth information of the set of training data to generate a first score and a second score associated respectively with the first and second instances of the RL agent, which is an evaluation or judgement that can be performed in the human mind.
Thus, the claim recites an abstract idea under Step 2A Prong 1.Under Step 2A Prong 2, the claim recites:
applying the first instance of the RL agent to a set of training data, as the performance of an abstract idea on a computer is not more than instructions to 'apply it' on a computer, which by MPEP 2106.05(f) cannot integrate an abstract idea into a practical application;
applying the second instances of the RL agent to the set of training data, as the performance of an abstract idea on a computer is not more than instructions to 'apply it' on a computer, which by MPEP 2106.05(f) cannot integrate an abstract idea into a practical application.
Thus, the claim is directed towards an abstract idea.
Further, the additional element(s), alone or in combination, do not provide significantly more than the abstract idea itself, because implementation on a computer (MPEP 2106.05(f)) cannot provide significantly more. The combination of these additional elements does not provide an inventive concept; thus, the claim remains ineligible.
Claim 12, dependent on 11, further recites
determining that the second score is greater than the first score; accessing the second reward function used to train the second instance of the RL agent, which is an evaluation or judgement that can be performed in the human mind;
and refining the prompt for the LLM using the second reward function, which is an evaluation or judgement that can be performed in the human mind.
Thus, the claim recites an abstract idea under Step 2A Prong 1.Under Step 2A Prong 2, the claim recites no additional elements which could integrate the abstract idea into a practical application or provide significantly more than the abstract idea itself; thus, the claim remains ineligible.
Claim 15, dependent on 1, further recite. no additional abstract ideas. However: Under Step 2A Prong 2, the claim recites:
the set of instructions comprise code for the RL agent, as the performance of an abstract idea on a computer is not more than instructions to 'apply it' on a computer, which by MPEP 2106.05(f) cannot integrate an abstract idea into a practical application.
Thus, the claim is directed towards an abstract idea.
Further, the additional element(s), alone or in combination, do not provide significantly more than the abstract idea itself, because implementation on a computer (MPEP 2106.05(f)) cannot provide significantly more. The combination of these additional elements does not provide an inventive concept; thus, the claim remains ineligible.
Claim 16, dependent on 15, further recite. no additional abstract ideas. However: Under Step 2A Prong 2, the claim recites:
the set of instructions comprise an initial reward function, as the performance of an abstract idea on a computer is not more than instructions to 'apply it' on a computer, which by MPEP 2106.05(f) cannot integrate an abstract idea into a practical application.
Thus, the claim is directed towards an abstract idea.
Further, the additional element(s), alone or in combination, do not provide significantly more than the abstract idea itself, because implementation on a computer (MPEP 2106.05(f)) cannot provide significantly more. The combination of these additional elements does not provide an inventive concept; thus, the claim remains ineligible.
Claim 17, dependent on 16, further recite. no additional abstract ideas. However: Under Step 2A Prong 2, the claim recites:
a first portion of the set of reward functions comprises a revised version of the initial reward function and a second portion of the set of reward functions comprises a reward function that is entirely different from the initial reward function, as the performance of an abstract idea on a computer is not more than instructions to 'apply it' on a computer, which by MPEP 2106.05(f) cannot integrate an abstract idea into a practical application.
Thus, the claim is directed towards an abstract idea.
Further, the additional element(s), alone or in combination, do not provide significantly more than the abstract idea itself, because implementation on a computer (MPEP 2106.05(f)) cannot provide significantly more. The combination of these additional elements does not provide an inventive concept; thus, the claim remains ineligible.
Claim 18, dependent on 17, further recite. no additional abstract ideas. However: Under Step 2A Prong 2, the claim recites:
the revised version of the initial reward function comprises additional penalty terms that are missing from the initial reward function, as the performance of an abstract idea on a computer is not more than instructions to 'apply it' on a computer, which by MPEP 2106.05(f) cannot integrate an abstract idea into a practical application.
Thus, the claim is directed towards an abstract idea.
Further, the additional element(s), alone or in combination, do not provide significantly more than the abstract idea itself, because implementation on a computer (MPEP 2106.05(f)) cannot provide significantly more. The combination of these additional elements does not provide an inventive concept; thus, the claim remains ineligible.
Claim 19 recites A method, thus a process, one of the four statutory categories of patentable subject matter. However, Claim 19 recites precisely the abstract ideas and additional elements of Claim 1. Therefore, Step 2A Prong 1, Step 2A Prong 2, and Step 2B analyses remain the same. Claim 19 is rejected as subject-matter ineligible for reasons set forth in the rejections of Claim 1.
Claim 20 recites A machine-storage medium, thus an article of manufacture, one of the four statutory categories of patentable subject matter. However, Claim 20 recites for storing instructions that, when executed by one or more hardware processors, cause the one or more hardware processors to perform operations comprising precisely the abstract ideas and additional elements of Claim 1. Therefore, Step 2A Prong 1 analysis remains the same. As for Step 2A Prong 2 and Step 2B: performance on a computer cannot integrate an abstract idea into a practical application (Step 2A Prong 2) nor provide significantly more than the abstract idea itself (Step 2B) (MPEP 2106.05(f)), and thus Claim 20 is rejected as subject-matter ineligible for reasons set forth in the rejections of Claim 1.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1, 6, 15, 16, 19, and 20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Kwon et al., "REWARD DESIGN WITH LANGUAGE MODELS" (hereinafter Kwon)
Regarding Claim 1, Kwon teaches:
A system comprising: one or more hardware processors; and at least one machine-storage medium for storing instructions that, when executed by the one or more hardware processors, cause the one or more hardware processors to perform operations comprising: (Kwon [abstract] In all three tasks, we show that RL agents trained with our framework are well-aligned with the user’s objectives and outperform RL agents trained with reward functions learned via supervised learning. Code and prompts can be found here. https://github.com/minaek/reward_design_with_llms ) (Note: training RL agents with the code requires processors and instructions to perform operations)
providing a prompt to a large language model (LLM), the prompt comprising a set of instructions for generating a set of reward functions associated with training a reinforcement learning (RL) agent to predict an objective; (Kwon [p.4 ¶1] In practice, we do not have access to ground truth user reward functions — this is the function that we are trying to approximate. However, for most of our experiments, we assume access to the true reward by constructing user reward functions that humans have been shown to have inspired by prior work. We use the true rewards only to evaluate our framework’s performance.
[p.2 fig.1]
PNG
media_image1.png
582
1059
media_image1.png
Greyscale
A user provides an example and explanation of desired negotiating behavior (e.g., versatility) before training.)
obtaining the set of reward functions from the LLM; (Kwon [p.2 fig.1] (2-3) We then parse the LLM’s response back into a string and use that as the reward signal for the Alice the RL agent.)
training one or more instances of the RL agent using the set of reward functions to predict the objective; (Kwon [p.2 fig.1]
PNG
media_image2.png
594
1063
media_image2.png
Greyscale
Figure 1: Depiction of our framework on the DEALORNODEAL negotiation task. A user provides an example and explanation of desired negotiating behavior (e.g., versatility) before training.) (Note: the end evaluation of whether trajectory of agent satisfies user objectives corresponds to predicting the objective)
generating a score representing accuracy of the predicted objective for the one or more instances of the RL agent (Kwon [p.4 ¶4] Labeling Accuracy. We construct ground-truth reward functions for each domain. We report the mean accuracy of predictions of the reward value during RL training with respect to the ground-truth reward functions. This assesses how effectively the LLM can produce reward signals that are consistent with the user’s objective.)
and selecting an individual instance of the one or more instances of the RL agent to predict the objective based on the generated score. (Kwon [p.2 fig.1] (4) Alice updates their weights and rolls out a new episode. (5) We parse the episode outcome int a string and continue training. During evaluation, we sample a trajectory from Alice and evaluate whether it is aligned with the user’s objective.)
Regarding Claim 6, Kwon respectively teaches and incorporates the claimed limitations and rejections of Claim 1. Kwon further teaches:
The system of claim 1, wherein the RL agent comprises a multilayer neural network machine learning (ML) model. (Kwon [p.13 last ¶] SL is trained to predict binary labels for a batch of proposed splits. We implemented SL as a multi-layer perceptron (MLP) network that consists of a single hidden layer with depth 32. We also use ReLU activations after our input and hidden layers. We trained the model on the same 10 examples we gave to LLM for 5 epochs with the Adam optimizer.)
Regarding Claim 15, Kwon respectively teaches and incorporates the claimed limitations and rejections of Claim 1. Kwon further teaches:
The system of claim 1, wherein the set of instructions comprise code for the RL agent. (Kwon [abstract] In all three tasks, we show that RL agents trained with our framework are well-aligned with the user’s objectives and outperform RL agents trained with reward functions learned via supervised learning. Code and prompts can be found here. https://github.com/minaek/reward_design_with_llms)
Regarding Claim 16, Kwon respectively teaches and incorporates the claimed limitations and rejections of Claim 15. Kwon further teaches:
The system of claim 15, wherein the set of instructions comprise an initial reward function. (Kwon [p.2 fig.1]
PNG
media_image3.png
657
870
media_image3.png
Greyscale
) (Note: the user’s description of objectives (p2-p3) in the first iteration of RL training corresponds to the initial reward function)
Independent Claim 19 recites A method to perform precisely the limitations of Claim 1. Thus, Claim 19 is rejected for reasons set forth in Claim 1.
Independent Claim 20 recites A machine-storage medium for storing instructions that, when executed by one or more hardware processors, cause the one or more hardware processors to perform operations comprising (Kwon [abstract] In all three tasks, we show that RL agents trained with our framework are well-aligned with the user’s objectives and outperform RL agents trained with reward functions learned via supervised learning. Code and prompts can be found here. https://github.com/minaek/reward_design_with_llms ) (Note: training RL agents with the code requires processors and a machine-storage medium storing instructions to perform operations) to perform precisely the recited methods of Claim 1. Thus, Claim 20 is rejected for reasons set forth in Claim 1.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 2-5 and 8 are rejected under 35 U.S.C. 103 as being unpatentable over Kwon, in view of Li et al., "Predictive Modeling with Delayed Information: a Case Study in E-commerce Transaction Fraud Control" (hereinafter Li).
Regarding Claim 2, Kwon respectively teaches and incorporates the claimed limitations and rejections of Claim 1. Kwon does not teach, but Li further teaches:
The system of claim 1, wherein the objective comprises a risk associated with a user in transacting in an item in an electronic marketplace. (Li [abstract] We studied predictive modeling problems in this research which was motivated by real-world cases that Microsoft data scientists encountered while dealing with e-commerce transaction fraud control decisions using transaction streaming data…These frameworks generated decision environment related features using long-term fully mature and short-term partially mature data, and the values of those features were estimated using varies of learning methods, including… artificial neural network, and recurrent neural network
[p.4 right last¶] A transaction carries user’s account information) (Note: e-commerce fraud corresponds to risk in an electronic marketplace)
Li and Kwon are analogous to the present invention because both are from the same field of endeavor of generating reinforcement learning-based neural networks for classification. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to implement the objective to determining risk of a user transaction from Li into Kwon’s reward designing method for training RL agents. The motivation would be to “motivated by real-world cases that Microsoft data scientists encountered while dealing with e-commerce transaction fraud control decisions using transaction streaming data in an uncertain probabilistic decision environment.” (Li, abstract).
Regarding Claim 3, Kwon/Li respectively teaches and incorporates the claimed limitations and rejections of Claim 2. Li, via Kwon/Li further teaches:
The system of claim 2, wherein the risk comprises a likelihood of unauthorized chargeback associated with the user. (Li [p.4 right last¶] One of the most commonly used transaction fraud labels in e-commerce is “chargeback” which is the return of funds to the credit card holder)
Regarding Claim 4, Kwon respectively teaches and incorporates the claimed limitations and rejections of Claim 1. Kwon does not teach, but Li further teaches:
The system of claim 1, wherein the RL agent comprises a machine learning model that predicts the objective by analyzing a plurality of user features. (Li [abstract] We studied predictive modeling problems in this research which was motivated by real-world cases that Microsoft data scientists encountered while dealing with e-commerce transaction fraud control decisions using transaction streaming data…These frameworks generated decision environment related features using long-term fully mature and short-term partially mature data, and the values of those features were estimated using varies of learning methods, including… artificial neural network, and recurrent neural network
[p.4 right last¶] A transaction carries user’s account information…, product information…, and payment information (e.g. type of payment instrument, location of payment, etc.). The time of occurrence of a purchase transaction is immediately recorded as ”ReceivingTime”. A risk scoring engine then scores this transaction and record it in the risk control database as ”RiskScore”.)
Li and Kwon are analogous to the present invention because both are from the same field of endeavor of generating reinforcement learning-based neural networks for classification. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to implement the objective to determining risk of a user transaction from Li into Kwon’s reward designing method for training RL agents. The motivation would be to “motivated by real-world cases that Microsoft data scientists encountered while dealing with e-commerce transaction fraud control decisions using transaction streaming data in an uncertain probabilistic decision environment.” (Li, abstract).
Regarding Claim 5, Kwon/Li respectively teaches and incorporates the claimed limitations and rejections of Claim 4. Li, via Kwon/Li, further teaches:
The system of claim 4, wherein the plurality of user features comprises at least one of velocity of transactions, type of financial instrument being used by a user, type of device being used by the user, a registration date associated with the user, or collusive behavior information between the user and another user. (Li [p.4 right last¶] A transaction carries user’s account information…, product information…, and payment information (e.g. type of payment instrument, location of payment, etc.). The time of occurrence of a purchase transaction is immediately recorded as ”ReceivingTime”. A risk scoring engine then scores this transaction and record it in the risk control database as ”RiskScore”.)
Regarding Claim 8, Kwon respectively teaches and incorporates the claimed limitations and rejections of Claim 1. Kwon does not teach, but Li further teaches:
The system of claim 1, wherein the objective predicted by the individual instance of the one or more instances of the RL agent comprises a first likelihood of fraudulent activity before authorizing an electronic transaction, (Li [p.4 right fig.2]
PNG
media_image4.png
736
731
media_image4.png
Greyscale
[p.4 right last¶] A transaction carries user’s account information (e.g. Microsoft account information), product information (e.g. a Surface book with a set of specifications, the total price and cost), and payment information (e.g. type of payment instrument, location of payment, etc.). The time of occurrence of a purchase transaction is immediately recorded as ”ReceivingTime”. A risk scoring engine then scores this transaction and record it in the risk control database as ”RiskScore”.) (Note: BankDecision corresponds to authorizing an electronic transaction; RiskScore is generated before BankDecision, thus corresponding to the first likelihood)
a second likelihood of fraudulent activity after authorizing the electronic transaction, (Li [p.4 right last¶] A fraud control engine then makes control decision which is then recorded as “InlineDecision”. The payment issuing bank and the manual review (MR) team make decisions and their decisions are recorded in the database as “BankDecision” and ”MRDecision”.) (Note: MRDecision occurs after BankDecision is generated, thus corresponding to second likelihood)
and a third likelihood of fraudulent activity associated with delay capture. (Li [p.4 right last¶] The fraud label of this transaction is set to “False” by default in “FraudFlag” in the database, and after a stochastic lead time in data maturity, we receive the final fraud label and update transaction’s FraudFlag to “True” if a chargeback returns. The maturity time of a transaction is also recorded simultaneously in the column “MaturityTime”) (Note: the fraud labels depending on the maturity lead time corresponds to the third likelihood)
Li and Kwon are analogous to the present invention because both are from the same field of endeavor of generating reinforcement learning-based neural networks for classification. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to implement the objective to determining risk of a user transaction from Li into Kwon’s reward designing method for training RL agents. The motivation would be to “motivated by real-world cases that Microsoft data scientists encountered while dealing with e-commerce transaction fraud control decisions using transaction streaming data in an uncertain probabilistic decision environment.” (Li, abstract).
Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Kwon, in view of Tam et al., "Managing a PyTorch Training Process with Checkpoints and Early Stopping" (hereinafter Tam).
Regarding Claim 7, Kwon respectively teaches and incorporates the claimed limitations and rejections of Claim 1. Kwon does not teach, but Tam further teaches:
The system of claim 1, wherein the operations comprise: concluding a training process of the RL agent in response to determining that the score representing the accuracy of the predicted objective is greater than a threshold value. (Tam [p.4 last ¶] You can also checkpoint the model per epoch unconditionally together with the best model checkpointing, as you are free to create multiple checkpoint files. Since the code above is the find the best model and make a copy of it, you may usually see a further optimization to the training loop by stopping it early if the hope to see model improvement is slim. This is the early stopping technique that can save time in training.
The code above validates the model with test set at the end of each epoch and keeps the best model found into a checkpoint file. The simplest strategy for early stopping is to set up a threshold of 𝑘 epochs. If you didn’t see the model improved over the last 𝑘 epochs, you terminate the training loop in the middle. This can be implemented as follows:
PNG
media_image5.png
375
454
media_image5.png
Greyscale
)
Li and Tam are analogous to the present invention because both are from the same field of endeavor of building neural network training methods. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to implement the early stopping method from Tam into Lucas’s method of prompting a neural network training pipeline. The motivation would be to “you may usually see a further optimization to the training loop by stopping it early if the hope to see model improvement is slim. This is the early stopping technique that can save time in training” (Tam [p.4 last ¶]).
Claims 9-14 are rejected under 35 U.S.C. 103 as being unpatentable over Kwon, in view of Dubois et al., "AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback" (hereinafter Dubois).
Regarding Claim 9, Kwon respectively teaches and incorporates the claimed limitations and rejections of Claim 1. Kwon does not teach, but Dubois further teaches:
The system of claim 1, wherein the operations comprise removing one or more reward functions from the set of reward functions in response to determining that the one or more reward functions are incapable of accurately training the one or more instances of the RL agent. (Dubois [p.3 ¶2] While there are a rich set of methods that learn from human feedback to directly optimize R (see Section 6), in this work we focus on the setting of learning from pairwise feedback (LPF) due to its central role in recent instruction-following LLMs [46]. The starting point of this process is a model that is fine-tuned on instruction-following demonstrations (x,y), which we denote as pSFT θ (y | x). The LPF process then involves taking pairs of samples y0,y1 from pSFT θ for each instruction x, querying humans for which sample within each pair is better, and learning from this pairwise feedback. Since all methods start from the SFT base, we use pθ for notational simplicity.
[p.6 first ¶] Many LPF methods do not directly operate on pairwise feedback data, but instead first construct a surrogate reward model by fine-tuning a classifier from the SFT base using pairwise feedback. The following LPF methods maximize the continuous-valued reward defined by the logits of this classifier. • Best-of-n sampling. Best-of-n (or re-ranking) [63, 5, 21, 8] is a simple but effective inference-time method that draws n i.i.d. responses from the SFT model and returns the response with the highest surrogate reward. Unless stated otherwise, we use n = 1024 in our experiments. • Expert iteration. Expert iteration [2, 60, 70] is the natural training-time extension of best-of-n: it first generates according to best-of-n on new instructions and then fine-tunes on the best outputs.) (Note: each n sample corresponds to containing the reward function from the n set of reward functions; finding the best output out of n samples corresponds to removing the rest of the n reward functions; the n samples which are not the best output correspond to reward functions that are incapable of accurately training the RL agent)
Li and Dubois are analogous to the present invention because both are from the same field of endeavor of reward shaping using LLM for generating RL agents. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to implement the best-of-n expert iteration feedback mechanism for finding the best-trained RL model from Dubois into Kwon’s method of reward designing for RL agents. The motivation would be to “maximize the continuous-valued reward defined by the logits of this classifier.” (Dubois [p.6 first ¶]).
Regarding Claim 10, Kwon respectively teaches and incorporates the claimed limitations and rejections of Claim 1. Kwon does not teach, but Dubois further teaches:
The system of claim 1, wherein the operations comprise: training a first instance of the RL agent using a first reward function in the set of reward functions; and training, in parallel with training the first instance, a second instance of the RL agent using a second reward function in the set of reward functions. (Dubois [p.3 ¶2] While there are a rich set of methods that learn from human feedback to directly optimize R (see Section 6), in this work we focus on the setting of learning from pairwise feedback (LPF) due to its central role in recent instruction-following LLMs [46]. The starting point of this process is a model that is fine-tuned on instruction-following demonstrations (x,y), which we denote as pSFT θ (y | x). The LPF process then involves taking pairs of samples y0,y1 from pSFT θ for each instruction x, querying humans for which sample within each pair is better, and learning from this pairwise feedback. Since all methods start from the SFT base, we use pθ for notational simplicity.
[p.6 first ¶] Many LPF methods do not directly operate on pairwise feedback data, but instead first construct a surrogate reward model by fine-tuning a classifier from the SFT base using pairwise feedback. The following LPF methods maximize the continuous-valued reward defined by the logits of this classifier. • Best-of-n sampling. Best-of-n (or re-ranking) [63, 5, 21, 8] is a simple but effective inference-time method that draws n i.i.d. responses from the SFT model and returns the response with the highest surrogate reward. Unless stated otherwise, we use n = 1024 in our experiments. • Expert iteration. Expert iteration [2, 60, 70] is the natural training-time extension of best-of-n: it first generates according to best-of-n on new instructions and then fine-tunes on the best outputs.) (Note: each n sample corresponds to each instance of the RL agent; expert iteration involving using new instructions during training time corresponds to applying the first instance to the training data to generate a first objective; evaluation involving selecting the highest surrogate reward corresponds to first generating first and second scores for each of the respective samples (instances))
Li and Dubois are analogous to the present invention because both are from the same field of endeavor of reward shaping using LLM for generating RL agents. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to implement the best-of-n expert iteration feedback mechanism for finding the best-trained RL model from Dubois into Kwon’s method of reward designing for RL agents. The motivation would be to “maximize the continuous-valued reward defined by the logits of this classifier.” (Dubois [p.6 first ¶]).
Regarding Claim 11, Kwon/Dubois respectively teaches and incorporates the claimed limitations and rejections of Claim 10. Dubois, via Kwon/Dubois, further teaches:
The system of claim 10, wherein the operations comprise: applying the first instance of the RL agent to a set of training data to predict a first objective associated with the set of training data; applying the second instances of the RL agent to the set of training data to predict a second objective associated with the set of training data; and evaluating the first and second objectives based on ground truth information of the set of training data to generate a first score and a second score associated respectively with the first and second instances of the RL agent. (Dubois [p.3 ¶2] While there are a rich set of methods that learn from human feedback to directly optimize R (see Section 6), in this work we focus on the setting of learning from pairwise feedback (LPF) due to its central role in recent instruction-following LLMs [46]. The starting point of this process is a model that is fine-tuned on instruction-following demonstrations (x,y), which we denote as pSFT θ (y | x). The LPF process then involves taking pairs of samples y0,y1 from pSFT θ for each instruction x, querying humans for which sample within each pair is better, and learning from this pairwise feedback. Since all methods start from the SFT base, we use pθ for notational simplicity.
[p.6 first ¶] Many LPF methods do not directly operate on pairwise feedback data, but instead first construct a surrogate reward model by fine-tuning a classifier from the SFT base using pairwise feedback. The following LPF methods maximize the continuous-valued reward defined by the logits of this classifier. • Best-of-n sampling. Best-of-n (or re-ranking) [63, 5, 21, 8] is a simple but effective inference-time method that draws n i.i.d. responses from the SFT model and returns the response with the highest surrogate reward. Unless stated otherwise, we use n = 1024 in our experiments. • Expert iteration. Expert iteration [2, 60, 70] is the natural training-time extension of best-of-n: it first generates according to best-of-n on new instructions and then fine-tunes on the best outputs.) (Note: each n sample corresponds to each instance of the RL agent; expert iteration involving using new instructions during training time corresponds to applying the first instance to the training data to generate a first objective; evaluation involving selecting the highest surrogate reward corresponds to first generating first and second scores for each of the respective samples (instances))
Regarding Claim 12, Kwon/Dubois respectively teaches and incorporates the claimed limitations and rejections of Claim 11. Dubois, via Kwon/Dubois further teaches:
The system of claim 11, wherein the operations comprise: determining that the second score is greater than the first score; accessing the second reward function used to train the second instance of the RL agent; and refining the prompt for the LLM using the second reward function. (Dubois [p.6 first ¶] Many LPF methods do not directly operate on pairwise feedback data, but instead first construct a surrogate reward model by fine-tuning a classifier from the SFT base using pairwise feedback. The following LPF methods maximize the continuous-valued reward defined by the logits of this classifier. • Best-of-n sampling. Best-of-n (or re-ranking) [63, 5, 21, 8] is a simple but effective inference-time method that draws n i.i.d. responses from the SFT model and returns the response with the highest surrogate reward. Unless stated otherwise, we use n = 1024 in our experiments. • Expert iteration. Expert iteration [2, 60, 70] is the natural training-time extension of best-of-n: it first generates according to best-of-n on new instructions and then fine-tunes on the best outputs) (Note: the highest surrogate reward corresponds to the second reward function)
Regarding Claim 13, Kwon/Dubois respectively teaches and incorporates the claimed limitations and rejections of Claim 12. Dubois, via Kwon/Dubois, further teaches:
The system of claim 12, wherein the operations comprise: providing the refined prompt to the LLM with an instruction to generate a revised set of reward functions; and training the one or more instances of the RL agent using the revised set of reward functions provided by the LLM to predict the objective. (Dubois [p.6 first ¶] Many LPF methods do not directly operate on pairwise feedback data, but instead first construct a surrogate reward model by fine-tuning a classifier from the SFT base using pairwise feedback. The following LPF methods maximize the continuous-valued reward defined by the logits of this classifier. • Best-of-n sampling. Best-of-n (or re-ranking) [63, 5, 21, 8] is a simple but effective inference-time method that draws n i.i.d. responses from the SFT model and returns the response with the highest surrogate reward. Unless stated otherwise, we use n = 1024 in our experiments. • Expert iteration. Expert iteration [2, 60, 70] is the natural training-time extension of best-of-n: it first generates according to best-of-n on new instructions and then fine-tunes on the best outputs) (Note: the output generated from the expert iteration corresponds to the refined prompt; fin-tuning the agent using the best output corresponds to refining the prompt for the LLM using the second reward function)
Regarding Claim 14, Kwon/Dubois respectively teaches and incorporates the claimed limitations and rejections of Claim 13. Dubois, via Kwon/Dubois, further teaches:
The system of claim 13, wherein the operations comprise: comparing accuracy of predicted objectives generated by the one or more instances of the RL agent using the revised set of reward functions with accuracy of the predicted objectives generated using the second reward function; and selectively updating the prompt in response to comparing the accuracy of predicted objectives generated by the one or more instances of the RL agent using the revised set of reward functions with accuracy of the predicted objectives generated using the second reward function. (Dubois [p.2 fig.1]
PNG
media_image6.png
728
1037
media_image6.png
Greyscale
[p.2 last ¶] We show that the method rankings obtained from developing on AlpacaFarm closely agree with the method rankings obtained from training on actual human data (Spearman correlation of 0.98) and that the best method in AlpacaFarm leads to substantial gains with human feedback.) (Note: training the best method using human feedback corresponds to selectively updating the prompt based on predicted objectives (human feedback))
Claims 17 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Kwon, in view of Adeniji et al., "Language Reward Modulation for Pretraining Reinforcement Learning" (hereinafter Adeniji).
Regarding Claim 17, Kwon respectively teaches and incorporates the claimed limitations and rejections of Claim 16. Kwon does not teach, but Adeniji further teaches:
The system of claim 16, wherein a first portion of the set of reward functions comprises a revised version of the initial reward function and a second portion of the set of reward functions comprises a reward function that is entirely different from the initial reward function. (Adeniji [p.5 last ¶] Objective Because the LAMP reward can be seen as measuring the extent to which the agent is closer to solving the task (see Section 4.1), it can be readily be combined with novelty-seeking unsupervised RL methods that optimize both extrinsic and intrinsic rewards. Therefore, to incentivize exploration, we combine the LAMP reward with the novelty score from a separate exploration technique. Specifically, we consider Plan2Explore [32] that utilizes the disagreement between future latent state predictions as a novelty score. Let this novelty-based score be
PNG
media_image7.png
44
62
media_image7.png
Greyscale
. We then train our pretraining agent to maximize the following weighted sum of rewards:
PNG
media_image8.png
23
433
media_image8.png
Greyscale
where α is a hyperparameter that balances the two rewards. By combining this novelty-based reward with the LAMP reward, we encourage the agent to efficiently explore its environment but with an additional bias towards interacting with the semantically meaningful affordances. We found that an α value of 0.9 works quite well across the tasks evaluated.) (Note: the novelty-based score corresponds to the reward function entirely different from the initial reward function)
Li and Adeniji are analogous to the present invention because both are from the same field of endeavor of reward shaping methods for reinforcement tasks. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to implement the novelty-seeking exploration function from Adeniji into Kwon’s reward shaping method. The motivation would be to “optimizes these rewards in conjunction with standard novelty-seeking exploration rewards with rein forcement learning to acquire a language-conditioned, pretrained policy” (Adeniji, abstract).
Regarding Claim 18, Kwon/Adeniji respectively teaches and incorporates the claimed limitations and rejections of Claim 17. Adeniji, via Kwon/Adeniji, further teaches:
The system of claim 17, wherein the revised version of the initial reward function comprises additional penalty terms that are missing from the initial reward function. (Adeniji [p.5 last ¶] Objective Because the LAMP reward can be seen as measuring the extent to which the agent is closer to solving the task (see Section 4.1), it can be readily be combined with novelty-seeking unsupervised RL methods that optimize both extrinsic and intrinsic rewards. Therefore, to incentivize exploration, we combine the LAMP reward with the novelty score from a separate exploration technique. Specifically, we consider Plan2Explore [32] that utilizes the disagreement between future latent state predictions as a novelty score. Let this novelty-based score be
PNG
media_image7.png
44
62
media_image7.png
Greyscale
. We then train our pretraining agent to maximize the following weighted sum of rewards:
PNG
media_image8.png
23
433
media_image8.png
Greyscale
where α is a hyperparameter that balances the two rewards. By combining this novelty-based reward with the LAMP reward, we encourage the agent to efficiently explore its environment but with an additional bias towards interacting with the semantically meaningful affordances. We found that an α value of 0.9 works quite well across the tasks evaluated.) (Note: the additional bias corresponds to additional penalty terms)
Conclusion
Claim 13-14 are subject-matter eligible.
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Kim et al., US 12608564 B2, which teaches “a model that combines a large language model (LLM) that predicts words and sentences with reinforcement learning from human feedback (RLHF) for accelerating or improving agent learning based on human feedback” (Kim, 2. Description of the Related Art; 3rd ¶); Chandler et al., US 20250315662 A1, which teaches generating optimal prompt combinations based on the corresponding human decisions (Chandler [0008]).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JOSEP HAN whose telephone number is (703)756-1346. The examiner can normally be reached Mon-Fri 9am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kakali Chaki can be reached on (571) 272-3719. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/J.H./Examiner, Art Unit 2122
/MICHAEL H HOANG/PRIMARY EXAMINER, Art Unit 2122