DETAILED ACTION
This non-final office action is responsive to application 18/582,460 as submitted on February 20th 2024.
Claim status is currently pending and under examination for claims 1-25 of which independent claims are 1, 10 and 19.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-25 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Independent Claims 1, 10 and 19
Step 2A Prong One: Does the claim recite an abstract idea, law of nature, or natural phenomenon?
Yes, independent claim 1, under the broadest reasonable interpretation, recites the following limitations that are abstract ideas:
for a first time, predicting a set of propensities for a set of actions for the user by applying a first machine learning model to user features for the user; (mental process)
selecting a first action for the first time for the user based on the set of propensities and treating the user with the selected first action; (mental process)
for the second time, predicting a second set of propensities for the set of actions; (mental process)
selecting a second action for the second time for the user based on the second set of propensities and treating the user with the selected second action; (mental process)
evaluating a second machine learning model by at least: identifying a first importance weight including at least a ratio between a propensity generated by the second machine learning model to the first machine learning model for the first action; (mental process and math)
identifying a second importance weight including at least the first importance weight and a ratio between a propensity generated by the second model to the first model for the second action; (mental process and math)
and combining the first importance weight for the first time with a reward associated with the first action and the second importance weight for the second time with a second reward to evaluate the second machine learning model; (mental process)
and deploying the second machine learning model responsive to identifying that the result of the evaluation meets a criteria. (mental process)
The “for a first time, predicting …” and “for the second time, predicting…” steps involve determining propensities for a set of actions which amounts to no more than observations, evaluations, and judgments that can be performed in the human mind or with the use of a physical aid (e.g., pen and paper). The claim recites the steps of predicting a set of propensities at a high degree of generality, thus the steps are not required to have any specific level of complexity that would preclude the steps from being mental processes. Therefore, the steps are considered to be mental processes, see MPEP § 2106.04(a)(2)(III).
The “selecting” steps involve identifying first and second actions for a user based on propensities which amounts to no more than observations, evaluations, and judgments that can be performed in the human mind or with the use of a physical aid (e.g., pen and paper). The claim recites the steps of selecting actions at a high degree of generality, thus the steps are not required to have any specific level of complexity that would preclude the steps from being mental processes. Therefore, the “selecting” steps are considered to be mental processes, see MPEP § 2106.04(a)(2)(III).
The “evaluating” step involves determining an importance weight that is a ratio between two propensities, which represents a mathematical relationship and amounts to no more than evaluations, observations, and judgments that can be performed in the human mind or with the use of a physical aid (e.g., pen and paper). The claim recites the step of evaluating a second machine learning model at a high degree of generality, thus the step is not required to have any specific level of complexity that would preclude the step from being mental processes. Therefore, the “evaluating” step is considered to be a mathematical concept, see MPEP § 2106.04(a)(2)(I), and mental processes, see MPEP § 2106.04(a)(2)(III).
The “identifying” step involves determining a second importance weight that is a ratio between two propensities, which represents a mathematical relationship and amounts to no more than evaluations, observations, and judgments that can be performed in the human mind or with the use of a physical aid (e.g., pen and paper). The claim recites the step of identifying a second importance weight at a high degree of generality, thus the step is not required to have any specific level of complexity that would preclude the step from being mental processes. Therefore, the “identifying” step is considered to be a mathematical concept, see MPEP § 2106.04(a)(2)(I), and mental processes, see MPEP § 2106.04(a)(2)(III).
The “combining” step involves determining a second model’s performance by aggregating importance weights and rewards for each action which amounts to no more than observations, evaluations, and judgments that can be performed in the human mind or with the use of a physical aid (e.g., pen and paper). The claim recites the step of combining importance weights at a high degree of generality, thus the step is not required to have any specific level of complexity that would preclude the step from being mental processes. Therefore, the “combining” step is considered to be mental processes, see MPEP § 2106.04(a)(2)(III).
The “deploying” step involves identifying if a second model should be deployed based on an evaluation result which amounts to no more than observations, evaluations, and judgments that can be performed in the human mind or with the use of a physical aid (e.g., pen and paper). The claim recites the step of deploying a second machine learning model at a high degree of generality, thus the step is not required to have any specific level of complexity that would preclude the step from being mental processes. Therefore, the “deploying” step is considered to be mental processes, see MPEP § 2106.04(a)(2)(III).
Therefore, the independent claim recites a judicial exception. Independent claims 10 and 19 recite similar limitations corresponding to claim 1, therefore the same subject matter eligibility analysis is applied.
Step 2A Prong Two: Does the claim recite additional elements that integrate the judicial exception into a practical application?
No, the judicial exception recited above is not integrated into a practical application. The claims recite the following additional elements, but these additional elements are not sufficient to integrate the judicial exception into a practical application:
receiving a first indication that a user interacts with an application of an online system at a first time; (MPEP § 2106.05(g) necessary data gathering and insignificant extra-solution activity to the judicial exception)
for a first time, predicting a set of propensities for a set of actions for the user by applying a first machine learning model to user features for the user; (MPEP § 2106.05(f) mere instructions to implement an abstract idea on a computer, or generally links exception to a technological environment)
receiving a second indication the user interacts with the application at a second time; (MPEP § 2106.05(g) necessary data gathering and insignificant extra-solution activity to the judicial exception)
A non-transitory computer-readable storage medium storing instructions that when executed by a computer processor cause the computer processor to perform steps comprising: (claims 10 and 19) (MPEP § 2106.05(f) mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea)
a computer processor; (claim 19) (MPEP § 2106.05(f) mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea)
The “receiving” steps amount to mere data gathering and are recited at a high level of generality, thus adding insignificant extra-solution activity to the judicial exception – see MPEP § 2106.05(g). Under MPEP § 2106.05(d), such additional elements have been found by the courts to not integrate a judicial exception into a practical application.
The “for a first time, predicting …” step requires a “first machine learning model” to predict a set of propensities. The first machine learning model is used to apply the recited judicial exception without placing any limitation on how the first machine learning model operates. The limitation amounts to mere instructions to “apply” the judicial exception on a computer. It can also be viewed as nothing more than an attempt to generally link the use of the judicial exception to the technological environment of computers, see MPEP § 2106.05(f).
The remaining additional elements are recited at a high-level of generality such that they amount to no more than mere instructions to “apply” an exception using a generic component. Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea, see MPEP § 2106.05(f).
Therefore, the above limitations do not integrate the judicial exception into a practical application.
Step 2B: Does the claim recite additional elements that amount to significantly more than the judicial exception?
No. The claims do not include additional elements that are sufficient for the claims to amount to significantly more than the judicial exception.
In regards to the “receiving” steps, these steps add insignificant extra-solution activity. An extra-solution activity is a well-understood, routine and conventional (WURC) activity per MPEP § 2106.05(d)(II), “the courts have recognized the following computer functions as well‐understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity. i. Receiving or transmitting data over a network, e.g., using the Internet to gather data.” The “receiving” steps do not integrate the judicial exception into a practical application and do not amount to significantly more.
In regards to the “first machine learning model” in the “for a first time, predicting …” step, the limitations are recited so generically such that they amount to no more than mere instructions to “apply” the judicial exception on a computer using generic computer components. Mere instructions to apply a judicial exception cannot provide an inventive concept. See MPEP § 2106.05(f).
In regards to the remaining additional elements, the limitations are recited so generically such that they amount to no more than mere instructions to “apply” the judicial exception on a computer using generic computer components. Mere instructions to apply a judicial exception cannot provide an inventive concept. See MPEP § 2106.05(f).
Therefore, independent claims 1, 10 and 19 are not patent eligible.
Dependent Claims 2-9, 11-18 and 20-25
The remaining dependent claims being rejected do not recite additional elements, whether considered individually or in combination, that are sufficient to integrate the judicial exception into a practical application or amount to significantly more than a judicial exception.
Claim limitation
Examiner analysis
2, 11 and 20. The method of claim 1, wherein evaluating the second machine learning model further comprises dividing a time span including the first time and the second time into regular time intervals.
This is a mental process akin to a human evaluation/judgment/observation.
3, 12 and 21. The method of claim 2, wherein evaluating the second machine learning model further comprises for a time where the user did not select an action, assigning a null action to user with a propensity of 1.
This is a mental process akin to a human evaluation/judgment/observation.
4, 13 and 22. The method of claim 1, wherein the second machine learning model has a different set of parameters or a different architecture from the first machine learning model.
The step is recited at a high-level of generality such that the limitations amount to no more than mere instructions to “apply” the judicial exception on a computer. They can also be viewed as nothing more than an attempt to generally link the use of the judicial exception to the technological environment of computers, see MPEP § 2106.05(f).
5, 14 and 23. The method of claim 1, wherein identifying the second importance weight comprises multiplying the first importance weight and a ratio between a propensity generated by the second machine learning model to a propensity generated by the first machine learning model for the second action.
This is a mental process akin to a human evaluation/judgment/observation.
6, 15 and 24. The method of claim 1, wherein receiving the first indication the user interacts with the application of the online system further comprises receiving an indication the user accesses the application on a client device associated with the user by logging into the application.
The limitation represents mere necessary data gathering and is recited at a high level of generality, thus adding insignificant extra-solution activity to the judicial exception - see MPEP § 2106.05(g). The extra-solution activity is a well-understood, routine and conventional (WURC) activity per MPEP § 2106.05(d)(II).
7 and 16. The method of claim 1, wherein training the second machine learning model comprises of training data that may include multiple instances of users, where each instance denotes a user, contextual information obtained for the user, or labels for an action space for the user.
The step is recited at a high-level of generality such that the limitations amount to no more than mere instructions to “apply” the judicial exception on a computer. They can also be viewed as nothing more than an attempt to generally link the use of the judicial exception to the technological environment of computers, see MPEP § 2106.05(f).
8, 17 and 25. The method of claim 1, wherein deploying the second machine learning model comprises deploying the second machine learning model responsive to identifying that the result of the evaluation is above a threshold value.
This is a mental process akin to a human evaluation/judgment/observation.
9 and 18. The method of claim 8, further comprising evaluating a third machine learning model and identifying not to deploy the third machine learning model responsive to identifying the result of the evaluation for the third machine learning model is below the threshold value.
This is a mental process akin to a human evaluation/judgment/observation.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
The following are the references relied upon in the rejections below:
Narita, Yusuke, Shota Yasui, and Kohei Yata. "Debiased off-policy evaluation for recommendation systems." Proceedings of the 15th ACM Conference on Recommender Systems. 2021.
Claims 1-2, 4-5, 7, 10-11, 13-14, 16, 19-20 and 22-23 are rejected under 35 U.S.C. 102a(1) as being anticipated by Narita.
Regarding Claim 1, Narita teaches:
A method comprising ((P. 372, Abstract) “we develop an alternative method, which predicts the performance of algorithms given historical data that may have been generated by a different algorithm.”):
receiving a first indication that a user interacts with an application of an online system at a first time ((P. 377, Sec. 4.2, ¶1-2) “We apply our estimator to empirically evaluate the design of online advertisements. This application uses proprietary data provided by CyberAgent Inc., which we described in the introduction. This company uses bandit algorithms to determine the visual design of advertisements assigned to user impressions in a mobile game. Our data are logged data from a 7-day A/B test on ad “campaigns,” … As the A/B test lasts for 7 days, there are 28 batches in total.”
Online advertisements are assigned to user impressions in a mobile game. When advertisements are assigned to users, a user must be using the mobile game in order to receive the online advertisement, therefore a user using the mobile game is a first indication that a user interacts with an application of an online system at a first time. A mobile game is an ‘application’ and is part of an ‘online system’ since it receives online advertisements.);
for a first time, predicting a set of propensities for a set of actions for the user by applying a first machine learning model to user features for the user (The Examiner interprets “propensities” according to its broadest reasonable interpretation (BRI) in view of the Applicant’s specification as encompassing probabilities. This interpretation is consistent with the illustrative descriptions in the Applicant’s specification at [0025], (see excerpt below).
Applicant’s written description at [0025] “a user may interact with an application of the online concierge system 140, and the ML model outputs a propensity for each action that is a likelihood the action should be assigned to the user to elicit a desired response from the user (e.g., conversion on products, clicking on a content item, etc.).”
(P. 377, Sec. 4.2, ¶3-4) “In the notation of our theoretical framework, T is zero, and a trajectory takes the form of H = (
S
0
,
A
0
,
R
0
). Reward
R
0
is a click while action
A
0
is an advertisement design … State (or context)
S
0
is user and ad characteristics used by the algorithm … we regard the MAB algorithm as the behavior policy
π
b
”
(P. 373, Sec. 2.1, ¶2) “We call a function
π
:
S
→
Δ
(
A
)
a policy, which assigns each state s ∈ S a distribution over actions, where π(a|s) is the probability of taking action a when the state is s. Let H = (
S
0
,
A
0
,
R
0
,
…
,
S
T
,
A
T
,
R
T
) be a trajectory, where
S
t
,
A
t
, and
R
t
are the state, the action, and the reward in step t, respectively, and T denotes the last step and is fixed.”
A behavior policy
π
b
(‘first machine learning model’) is used to assign each state a distribution over actions based on user characteristics (‘user features for the user’). The probability π(a|s) of taking an action a (showing an advertisement design) when the state is s (user characteristics) is a propensity since the probability measures the likelihood of an action a (advertisement design) resulting in a reward (a user clicks the advertisement).
A probability (‘set of propensities’) is calculated (‘predicted’) using the behavior policy for each action in an action space (‘set of actions’) in step t=0 (‘first time’).);
selecting a first action for the first time for the user based on the set of propensities and treating the user with the selected first action ((P. 373, Sec. 2.1, ¶2) “We say that a trajectory H is generated by a policy π, or H ∼ π in short if H is generated by repeating the following process for t = 0, ...,T: … (2) Given
S
t
, the action
A
t
is randomly chosen based on
π
(
⋅
|
S
T
)
. (3) The reward
R
t
is drawn from the conditional reward distribution
P
R
(
⋅
|
S
t
,
A
t
)
.”
A probability
π
(
⋅
|
S
T
)
is the probability of taking an action
"
⋅
” given state
S
t
. Based on the probability (‘set of propensities’) an action
A
t
is chosen at random. For time step t=0 (‘first time’), the randomly chosen action is
A
0
, therefore selecting a first action
A
0
for the first time for the user based on a set of propensities. Based on the randomly chosen action
A
0
, a reward
R
0
is drawn from a conditional reward distribution derived from
A
0
. A reward represents a user click (see P. 377, Sec. 4.2, ¶3), therefore reward
R
0
is treating the user with the selected first action
A
0
since a reward represents a user clicking on the shown advertisement design (action
A
0
).);
receiving a second indication the user interacts with the application at a second time ((P. 377, Sec. 4.2, ¶1-2) “We apply our estimator to empirically evaluate the design of online advertisements. … uses bandit algorithms to determine the visual design of advertisements assigned to user impressions in a mobile game. Our data are logged data from a 7-day A/B test on ad “campaigns,” where each campaign randomly uses either a multi-armed bandit (MAB) algorithm or a contextual bandit (CB) algorithm for each user impression. The parameters of both algorithms are updated every 6 hours. As the A/B test lasts for 7 days, there are 28 batches in total.”
See (P. 377, Sec. 4.2, ¶4) describing MAB algorithm is the behavior policy
π
b
.
Online advertisements are assigned to user impressions in a mobile game (‘application’). When advertisements are assigned to users, a user must be using the mobile game in order to receive the online advertisement, therefore a user using the mobile game is an indication that a user interacts with the application. Since a MAB algorithm (behavior policy) assigns advertisements (actions) at different time steps (see P. 373, Sec. 2.1, ¶2), the algorithm therefore assigns an advertisement at step t=1 (‘second time’) in response to a second indication a user is interacting with the application at a second time.);
for the second time, predicting a second set of propensities for the set of actions ((P. 373, Sec. 2.1, ¶2) “We call a function
π
:
S
→
Δ
(
A
)
a policy, which assigns each state s ∈ S a distribution over actions, where π(a|s) is the probability of taking action a when the state is s. Let H = (
S
0
,
A
0
,
R
0
,
…
,
S
T
,
A
T
,
R
T
) be a trajectory, where
S
t
,
A
t
, and
R
t
are the state, the action, and the reward in step t, respectively, and T denotes the last step and is fixed.”
A probability (‘second set of propensities’) is calculated (‘predicted’) using the behavior policy
π
b
for each action in an action space (‘set of actions’) in step t=1 (‘second time’).);
selecting a second action for the second time for the user based on the second set of propensities and treating the user with the selected second action ((P. 373, Sec. 2.1, ¶2) “We say that a trajectory H is generated by a policy π, or H ∼ π in short if H is generated by repeating the following process for t = 0, ...,T: … (2) Given
S
t
, the action
A
t
is randomly chosen based on
π
(
⋅
|
S
T
)
. (3) The reward
R
t
is drawn from the conditional reward distribution
P
R
(
⋅
|
S
t
,
A
t
)
.”
A probability
π
(
⋅
|
S
T
)
is the probability of taking an action
"
⋅
"
given state
S
t
. Based on the probability (‘second set of propensities’) an action
A
t
is chosen at random. For time step t=1 (‘second time’), the randomly chosen action is
A
1
, therefore selecting a second action
A
1
for the second time for the user based on a second set of propensities. Based on the randomly chosen action
A
1
, a reward
R
1
is drawn from a conditional reward distribution derived from
A
1
. A reward represents a user click, therefore reward
R
1
is treating the user with the selected second action
A
1
since a reward represents a user clicking on the shown advertisement design (action
A
1
).);
evaluating a second machine learning model by at least: identifying a first importance weight including at least a ratio between a propensity generated by the second machine learning model to the first machine learning model for the first action ((P. 374, Sec. 3, ¶2) “Before presenting our estimator, we introduce some notation.
H
t
s
,
a
=
(
S
0
,
A
0
,
…
,
S
t
,
A
t
)
is a trajectory of the state and action up to step t.
ρ
t
π
e
…is the importance weight function: … This equals the probability of H up to step t under the evaluation policy
π
e
divided by its probability under the behavior policy
π
b
”
(P. 372, Abstract) “Efficient methods to evaluate new algorithms are critical for improving interactive bandit and reinforcement learning systems such as recommendation systems … Our estimator has the property that its prediction converges in probability to the true performance of a counterfactual algorithm … We also show a correct way to estimate the variance of our prediction, thus allowing the analyst to quantify the uncertainty in the prediction.”
Narita discloses Equation 1 (reproduced below) on P. 374 describing an importance weight function that calculates a weight for each time step. At time step t’=0 (‘first time’), an importance weight is calculated which is
π
e
(
A
0
|
S
0
)
divided by
π
b
(
A
0
|
S
0
)
, therefore the importance weight for action
A
0
is a ratio for the first action.
π
e
(
A
0
|
S
0
)
is a probability (‘propensity’) generated by an evaluation policy (‘second machine learning model’), and
π
b
A
0
S
0
is a probability generated by the behavior policy (‘first machine learning model’).
PNG
media_image1.png
402
1219
media_image1.png
Greyscale
);
identifying a second importance weight including at least the first importance weight and a ratio between a propensity generated by the second model to the first model for the second action (Narita discloses Equation 1 on P. 374 describing an importance weight function that calculates a weight for each time step. At time step t’=1 (‘second time’), a second importance weight is calculated which is
π
e
(
A
1
|
S
1
)
divided by
π
b
(
A
1
|
S
1
)
, therefore the importance weight for action
A
1
is a ratio for the second action.
π
e
(
A
1
|
S
1
)
is a probability (‘propensity’) generated by an evaluation policy (‘second model’), and
π
b
(
A
1
|
S
1
)
is a probability generated by the behavior policy (‘first model’).
The second importance weight is multiplied with the first importance weight (weight for first action
A
0
at t’=0), therefore, identifying a second importance weight including the first importance weight and a ratio between propensities.);
and combining the first importance weight for the first time with a reward associated with the first action and the second importance weight for the second time with a second reward to evaluate the second machine learning model ((P. 374, Sec. 3, ¶3) “Our estimator is based on the following expression … To give an intuition behind the expression, we arrange the terms as follows:
PNG
media_image2.png
344
767
media_image2.png
Greyscale
The first term is the well-known Inverse Probability Weighting (IPW) estimator.”
The importance weight function
ρ
t
π
e
is multiplied with a reward
R
t
at each time step. At time step t=0 (‘first time’), the importance weight function calculates a first importance weight
ρ
0
π
e
and multiplies it with a reward
R
0
(‘first reward associated with the first action’). At time step t=1 (‘second time’), the importance weight function calculates a second importance weight
ρ
1
π
e
and multiplies it with a reward
R
1
(‘second reward’). An estimator combines the first and second importance weights when they are summed to evaluate a policy (see P. 372, Abstract describing an estimator evaluates a policy).);
and deploying the second machine learning model responsive to identifying that the result of the evaluation meets a criteria ((P. 372, Sec. 1, ¶4) “This leads us to the problem of counterfactual (off-policy, offline) evaluation, where one aims to use batch data collected by a logging policy to estimate the value of a counterfactual policy or algorithm without deploying it. Such evaluation allows us to compare the performance of counterfactual policies to decide which policy should be deployed in the field.”
Counterfactual policies (evaluation policy
π
e
) are evaluated using an estimator. A result of an evaluation performed by the estimator is an estimate value of a counterfactual policy (obtained from summing importance values). The performances of the counterfactual policies (estimate values) are compared to determine which policy performs the best, therefore the best performing policy is selected for deployment based on an evaluation for a policy meeting criteria, the criteria being selecting the best performing policy (the policy with the highest estimate value).).
Regarding Claims 2, 11 and 20, Narita teaches:
The method of claim 1, wherein evaluating the second machine learning model further comprises dividing a time span including the first time and the second time into regular time intervals ((P. 373, Sec. 2.1, ¶2) “We call a function
π
:
S
→
Δ
(
A
)
a policy, which assigns each state s ∈ S a distribution over actions, where π(a|s) is the probability of taking action a when the state is s. Let H = (
S
0
,
A
0
,
R
0
,
…
,
S
T
,
A
T
,
R
T
) be a trajectory, where
S
t
,
A
t
, and
R
t
are the state, the action, and the reward in step t, respectively, and T denotes the last step and is fixed.”
Time (‘a time span’) is divided into T steps going from t=0 to t=T, therefore step t=0 (‘first time’) and step t=1 (‘second time’) are regular time intervals.).
Regarding Claims 4, 13 and 22, Narita teaches:
The method of claim 1, wherein the second machine learning model has a different set of parameters or a different architecture from the first machine learning model ((P. 377, Sec. 4.2, ¶2) “each campaign randomly uses either a multi-armed bandit (MAB) algorithm or a contextual bandit (CB) algorithm for each user impression. The parameters of both algorithms are updated every 6 hours.”
(P. 377, Sec. 4.2, ¶4) “we regard the MAB algorithm as the behavior policy
π
b
and the CB algorithm as the evaluation policy
π
e
”
(P. 376, Sec. 4.1, ¶5) “We use a convolutional neural network to estimate
π
b
… We set the learning rate to 0.0001 for estimating the Q function and to 0.0002 for estimating the behavior policy.”
The parameters of a behavior policy (‘first machine learning model’) and evaluation policy (‘second machine learning model’) are updated every 6 hours and the policies have different learning rates, therefore the policies have a different set of parameters from each other.).
Regarding Claims 5, 14 and 23, Narita teaches:
the method of claim 1, wherein identifying the second importance weight comprises multiplying the first importance weight and a ratio between a propensity generated by the second machine learning model to a propensity generated by the first machine learning model for the second action (Narita discloses Equation 1 on P. 374 describing an importance weight function that calculates a weight for each time step. At time step t’=1 (‘second time’), a second importance weight is calculated which is
π
e
(
A
1
|
S
1
)
divided by
π
b
(
A
1
|
S
1
)
, therefore the second importance weight for action
A
1
is a ratio for the second action.
π
e
(
A
1
|
S
1
)
is a probability (‘propensity’) generated by an evaluation policy (‘second model’), and
π
b
(
A
1
|
S
1
)
is a probability generated by the behavior policy (‘first model’).
The second importance weight is multiplied with the first importance weight (weight for first action
A
0
at t’=0), therefore, multiplying a first importance weight and a ratio between propensities (the second importance weight).).
Regarding Claims 7 and 16, Narita teaches:
The method of claim 1, wherein training the second machine learning model comprises of training data that may include multiple instances of users, where each instance denotes a user, contextual information obtained for the user, or labels for an action space for the user ((P. 377, Sec. 4.2, Last Paragraph) “We obtain the counterfactual policy by training a click prediction model with LightGBM on the data from the first B - 1 batches.”
(P. 377, Sec. 4.2, ¶2-3) “Our data are logged data from a 7-day A/B test on ad “campaigns,” … As the A/B test lasts for 7 days, there are 28 batches in total. … State (or context) S0 is user and ad characteristics used by the algorithm. S0 is high dimensional and has tens of thousands of possible values.”
A counterfactual policy is trained using batches of logged data that include contextual user characteristics (‘contextual information obtained for the user’).).
Regarding Claim 10, the rejection of claim 1 is incorporated. The difference in scope being:
A non-transitory computer-readable storage medium storing instructions that when executed by a computer processor cause the computer processor to perform steps comprising ((P. 377, Sec. 4.2, ¶2) “each campaign randomly uses either a multi-armed bandit (MAB) algorithm or a contextual bandit (CB) algorithm for each user impression. The parameters of both algorithms are updated every 6 hours”
A computer is implied by updating parameters of algorithms, which further implies a non-transitory computer-readable storage medium storing instructions that are executable by a processor.).
Regarding Claim 19, the rejection of claim 1 is incorporated. The difference in scope being:
A computer system, the computer system comprising ((P. 377, Sec. 4.2, ¶2) “The parameters of both algorithms are updated every 6 hours”
Using a computer is implied by updating parameters of algorithms.):
a computer processor (A computer processor is further implied by using a computer.);
and a non-transitory computer-readable storage medium storing instructions that when executed by a computer processor cause the computer processor to perform steps comprising (A computer further implies a non-transitory computer-readable storage medium storing instructions that are executable by a computer processor.).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The following are the references relied upon in the rejections below:
Theocharous, Georgios, et al. "Reinforcement learning for strategic recommendations." arXiv preprint arXiv:2009.07346 (2020).
Claims 3, 12 and 21 are rejected under 35 U.S.C. 103 as being unpatentable over Narita / Theocharous.
Regarding Claims 3, 12 and 21, Narita teaches: The method of claim 2, however Narita does not teach assigning a null action to a user with a propensity of 1, which is taught by Theocharous,
wherein evaluating the second machine learning model further comprises for a time where the user did not select an action, assigning a null action to user with a propensity of 1 ((P. 21, Sec. 8, ¶1) “The ‘recommendation fatigue’ is a problem where people may quickly stop paying attention to recommendations such as ads, if they are presented too often. … then RL would naturally optimize the right sending schedule and thus avoid fatigue. In this section we present experimental results for a Point-of-Interest (POI) recommendation system”
(P. 21, Sec. 8, ¶3) “We trained a PST using the data and performed various experiments to test the ability of our algorithm to quickly optimize the cumulative reward for a given user. … For reward, we used a signal between [0,1] indicating the frequency/desirability of the POIs. The action space was a recommendation for each POI (88 POIs), plus a null action. All actions but the null action had a cost 0.2 of the reward.”
(P. 22-23, Sec. 9.1, Last Paragraph) “The state space of a MOMDP model factors into a fully observable factor x ∈ X and a partially observable factor y ∈ Y, each with their own transition functions, … An observation function
Ω
(
o
|
a
,
y
'
)
exists to inform the decision maker about transitions of the hidden factor. … we derive an equivalent MOMDP … having elements
Ω
(
o
N
U
L
L
|
a
,
θ
'
)
=
1
”
A reinforcement learning algorithm’s action space is comprised of actions (ad recommendations) and a null action. An observation function
Ω
(
o
N
U
L
L
|
a
,
θ
'
)
=
1
is used to inform an RL agent (decision maker) about transitions of a hidden factor, therefore the observation function is the probability of observing a null action given action a and state
θ
'
(observable factor). Since an action is an ad recommendation, a null action represents a user not being shown an ad recommendation (therefore assigning a null action for a time where a user did not select an action (did not interact with an ad)), and therefore an observation function assigns the null action with a probability of 1. The probability is a propensity since it is a likelihood that a null action (not recommending ads) occurs.).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the evaluation method of Narita with the technique disclosed by Theocharous to determine a probability for a null action. By determining a probability for a null action, a null action can be used to account for situations when no action needs to be chosen, thereby preventing sub-optimal actions from being forcefully selected when choosing an action is not necessary for a specific state.
The following are the references relied upon in the rejections below:
Shi, Longxiang, et al. "Session-based interactive recommendation via deep reinforcement learning." 2023 IEEE International Conference on Data Mining (ICDM). IEEE, 2023.
Claims 6, 15 and 24 are rejected under 35 U.S.C. 103 as being unpatentable over Narita / Shi.
Regarding Claims 6, 15 and 24, Narita teaches: The method of claim 1, however Narita does not teach receiving an indication a user accesses an application on a client device associated with the user by logging into the application, which is taught by Shi:
wherein receiving the first indication the user interacts with the application of the online system further comprises receiving an indication the user accesses the application on a client device associated with the user by logging into the application ((P. 1319, Abstract) “A user’s multiple continuous interactions in a given time period (e.g., the time from login to log out) naturally constitute a session. However, existing studies often overlook such valuable session structure and characteristics and instead simply treat them as sequences. As a result, they are not able to capture the complex transitions over users’ interactions within or between sessions, leading to significant information loss. To bridge this significant gap, in this paper, we propose Session-based Interactive Recommendation with Graph Neural Networks (SIR-GNN). SIR-GNN models interaction data as sessions and employs novel graph neural networks to capture rich transition patterns among interactions”
(P. 1319, Sec. 1, ¶1) “Interactive Recommender Systems (IRSs) have become indispensable in providing personalized services that guide users to their desired content, products, and services due to the rapid proliferation of online information and mobile applications. … The state representation module encodes the user’s historical interaction with a sequence of items, and the RL agent selects items for recommendations”).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the evaluation method of Narita with the technique disclosed by Shi to use user login data to identify sessions. By using user login data to identify sessions, recommendation systems can capture user preferences over time from different sessions, thereby helping systems learn how user preferences change over time and improve recommendations.
The following are the references relied upon in the rejections below:
Veneranta, Leevi. "Optimization of web page advertisements using contextual bandits." (2024).
Claims 8-9, 17-18 and 25 are rejected under 35 U.S.C. 103 as being unpatentable over Narita / Veneranta.
Regarding Claims 8, 17 and 25, Narita teaches: The method of claim 1, however Narita does not teach deploying a second machine learning model when a result of an evaluation is above a threshold value, which is taught by Veneranta:
wherein deploying the second machine learning model comprises deploying the second machine learning model responsive to identifying that the result of the evaluation is above a threshold value ((P. 27, Sec. 2.4, ¶3) “Inverse Propensity Scoring (IPS) is a loss type that is used to estimate the performance of a policy based on logged data. IPS can also be used in off-policy learning, that is, when the policy that generated the logged data is different from the policy being evaluated. The basic idea behind IPS is to weight each logged sample by the inverse of the probability of selecting that action under the policy being evaluated. This weighting scheme allows us to estimate the expected reward of the policy being evaluated using the logged data”
(P. 27, Sec. 2.4, Last Paragraph) “The IPS loss can be interpreted as a measure of how well a new policy performs relative to an old policy. A positive IPS loss indicates that the new policy performs better than the old policy, while a negative IPS loss indicates that it performs worse. The magnitude of the IPS loss can be used to compare different policies and select the best one”
An IPS loss is calculated for a new policy (‘second machine learning model’). When the IPS loss (‘result of an evaluation’) is positive (above a threshold value of zero), the new policy is determined to perform better and is selected (‘deployed’).).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the evaluation method of Narita with the technique disclosed by Veneranta to deploy a model when an evaluation result is above a threshold. By deploying a model when an evaluation result is above a threshold, it can be ensured that the model performs at least as well as a desired performance value, thereby guaranteeing that a deployed model meets performance standards.
Regarding Claims 9 and 18, the combined evaluation method of Narita / Veneranta teaches:
The method of claim 8, further comprising evaluating a third machine learning model and identifying not to deploy the third machine learning model responsive to identifying the result of the evaluation for the third machine learning model is below the threshold value ((P. 27, Sec. 2.4, Last Paragraph) “The IPS loss can be interpreted as a measure of how well a new policy performs relative to an old policy. A positive IPS loss indicates that the new policy performs better than the old policy, while a negative IPS loss indicates that it performs worse. The magnitude of the IPS loss can be used to compare different policies and select the best one”
An IPS loss is calculated for a new policy (‘third machine learning model’). When the IPS loss (‘result of an evaluation’) is negative (below a threshold value of zero), the new policy is determined to perform worse than an old policy and is therefore not selected (‘deployed’).).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined evaluation method of Narita / Veneranta with the technique disclosed by Veneranta to not deploy a model when an evaluation result is below a threshold. By not deploying a model when an evaluation result is below a threshold, it can be ensured that a sub-optimal model is not deployed, thereby preventing a model from generating inaccurate predictions.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Ishikawa et al. (US 20240202776 A1) teaches performing off-policy evaluation by using inverse probability weighting to evaluate a policy that recommends advertisements to users each time users request access to a website.
Jeunen et al. (“Joint Policy-Value Learning for Recommendation”) teaches combining inverse propensity score and maximum likelihood estimation to perform off-policy learning for generating optimal recommendations for users.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PEDRO J MORALES whose telephone number is (571)272-6106. The examiner can normally be reached 8:30 AM - 6:00 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, MIRANDA M HUANG can be reached at (571)270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PEDRO J MORALES/Examiner, Art Unit 2124
/MIRANDA M HUANG/Supervisory Patent Examiner, Art Unit 2124