Prosecution Insights
Last updated: October 02, 2026
Application No. 18/423,834

RETRIEVAL STRATEGY SELECTION OPTIMIZATION USING REINFORCEMENT LEARNING

Non-Final OA §101§103
Filed
Jan 26, 2024
Priority
Sep 21, 2023 — provisional 63/584,359
Examiner
MAIDO, MAGGIE T
Art Unit
Tech Center
Assignee
Roku Inc.
OA Round
1 (Non-Final)
67%
Grant Probability
Favorable
1-2
OA Rounds
1y 4m
Est. Remaining
96%
With Interview

Examiner Intelligence

Grants 67% — above average
67%
Career Allowance Rate
37 granted / 55 resolved
+7.3% vs TC avg
Strong +29% interview lift
Without
With
+29.1%
Interview Lift
resolved cases with interview
Typical timeline
4y 0m
Avg Prosecution
22 currently pending
Career history
92
Total Applications
across all art units

Statute-Specific Performance

§101
24.9%
-15.1% vs TC avg
§103
53.5%
+13.5% vs TC avg
§102
3.3%
-36.7% vs TC avg
§112
18.3%
-21.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 55 resolved cases

Office Action

§101 §103
DETAILED ACTION This action is responsive to claims filed on 26 January 2024. Claims 1-20 are pending for examination. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Objections Claim 4 is objected to because of the following informalities: “the top number of the content items” in line 2 should be “the top number of the sampled content items”. Appropriate correction is required. Claim 5 and analogous claim 13 is objected to because of the following informalities: “the top number of the content items” in line 2 should be “the top number of the sampled content items”. Appropriate correction is required. Claim 6 and analogous claims 14, 18 is objected to because of the following informalities: “the top number of content items” in line 2-3 should be “the top number of sampled content items”. Appropriate correction is required. Claim 7 and analogous claims 15, 19 is objected to because of the following informalities: “the top number of content items” in line 3 should be “the top number of sampled content items”. Appropriate correction is required. Claim 8 and analogous claims 16, 20 is objected to because of the following informalities: “the top number of content items” in lines 2, 3 should be “the top number of sampled content items”. Appropriate correction is required. Claim 17 is objected to because of the following informalities: “a top number of content items” in line 15 should be “a top number of sampled content items”. Appropriate correction is required. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception, abstract idea, without significantly more. Step 1: This part of the eligibility analysis evaluates whether the claim(s) falls within any statutory category. MPEP 2106.03: According to the first part of the Alice analysis, in the instant case, the claims were determined to be directed to one of the four statutory categories: an article of manufacture, a method/process (Claims 1-8, 17-20), a machine/system/product (Claims 9-16), and a composition of matter. Based on the claims being determined to be within of the four categories (i.e., process, machine, manufacture, or composition of matter), (Step 1), it must be determined if the claims are directed to a judicial exception (i.e., law of nature, natural phenomenon, and abstract idea). Step 2A Prong One: This part of the eligibility analysis evaluates whether the claim(s) recites a judicial exception. Regarding independent claims 1, 9, 17, the claims recite a judicial exception (i.e., an abstract idea enumerated in the 2019 PEG) without significantly more (Step-2A: Prong One). The applicant's claim limitations under broadest reasonable interpretation covers activities classified under mental processes - concepts performed in the human mind (including an observation, evaluation, judgment, opinion) (see MPEP § 2106.04(a)(2), subsection Ill) and the 2019 PEG. As evaluated below: Claims 1, 9: “determining content item feature vectors corresponding to the sampled content items” (mental process of judgement) “determining, using parameters of an agent model and an embedding of the query, an action vector comprising features corresponding to elements in the content item feature vector” (mental process of judgement) “for each sampled content item, determining a dot product of the action vector and the content item feature vector” (mental process of judgement) “sorting the sampled content items based on dot products” (mental process of evaluation) “computing a reward based on a top number of sampled content items having highest dot products” (mental process of evaluation) If the identified limitation(s) falls within at least one of the groupings of abstract ideas, it is reasonable to conclude that the claim(s) recites an abstract idea in Step 2A Prong One. Step 2A Prong Two: This part of the eligibility analysis evaluates whether the claim(s) as a whole integrates the recited judicial exception into a practical application of the exception. As evaluated below: “obtaining a query” “obtaining, from content items corresponding to the query, a number of sampled content items” “wherein a content item feature vector is generated based on a description of a sampled content item, metadata of the sampled content item, one or more past engagement statistics of the sampled content item, and a retrieval strategy used to retrieve the sampled content item” These recitations are deemed insufficient to transform the judicial exception to a patentable invention because the recitation is directed to instructions for mere data gathering or data output, see MPEP 2106.05(g). “and the action vector is a feature representation of weights given to different retrieval strategies” “updating the parameters of the agent model based on the query, the action vector, and the reward” These recitations are deemed insufficient to transform the judicial exception to a patentable invention because the recitation is directed to instructions merely indicating a field of use or technological environment in which to apply a judicial exception, see MPEP 2106.05(h). Accordingly, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea when considered as an ordered combination and as a whole. Step 2B: This part of the eligibility analysis evaluates whether the claim, as a whole, amounts to significantly more than the recited exception, i.e., whether any additional element, or combination of additional elements, adds an inventive concept to the claim. MPEP 2106.05. First, the additional elements considered as part of the preamble and the additional elements directed to the use of computer technology are deemed insufficient to transform the judicial exception to a patentable invention to a patentable invention because they generally link the judicial exception to the technology environment, see MPEP 2106.05(h). Second, the additional elements directed to mere application of the abstract idea or mere instructions to implement an abstract idea on a computer are deemed insufficient to transform the judicial exception to a patentable invention to a patentable invention because the limitations generally apply the use of a generic computer and/or process with the judicial exception, see MPEP 2106.05(f). Third, the claims are directed to instructions merely indicating a field of use or technological environment in which to apply a judicial exception. The courts have found these types of limitations insufficient to transform the judicial exception to a patentable invention, see MPEP 2106.05(g). Lastly, the claims directed to data gathering activity as noted above, are deemed directed to an insignificant extra-solution activity. The courts have found these types of limitations insufficient to qualify as "significantly more", see MPEP 2106.05(g). Furthermore, when considering evidence in view of Berkheimer v. HP, Inc., 881 F.3d 1360, 1368, 125 USPQ2d 1649, 1654 (Fed. Cir. 2018), see USPTO Berkheimer Memorandum (April 2018). Examiner notes Berkheimer: Option 2 - A citation to one or more of the court decisions discussed in MPEP § 2106.05(d}(II} as noting the well understood, routine, conventional nature of the additional element (s) (e.g., limitations directed to mere data gathering): The courts have recognized the following computer functions as well understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity, see MPEP 2106.05(d). The additional limitations, as analyzed, failed to integrate a judicial exception into a practical application at Step 2A and provide an inventive concept in Step 2B, per the analysis above. Thus, considering the additional elements individually and in combination and the claims as a whole, the additional elements do not provide significantly more than the abstract idea. This claim is not patent eligible. Therefore, in examining elements as recited by the limitations individually and as an ordered combination, as a whole, claims 1, 9 do not recite what the courts have identified as "significantly more". Claim 17: “determining content item feature vectors corresponding to the sampled content items” (mental process of judgement) “determining, using parameters of an agent model and an embedding of the context, an action vector comprising features corresponding to elements in the content item feature vector” (mental process of judgement) “for each sampled content item, determining a dot product of the action vector and the content item feature vector” (mental process of judgement) “sorting the sampled content items based on dot products” (mental process of evaluation) “computing a reward based on a top number of content items having highest dot products” (mental process of evaluation) If the identified limitation(s) falls within at least one of the groupings of abstract ideas, it is reasonable to conclude that the claim(s) recites an abstract idea in Step 2A Prong One. Step 2A Prong Two: This part of the eligibility analysis evaluates whether the claim(s) as a whole integrates the recited judicial exception into a practical application of the exception. As evaluated below: “obtaining a context” “obtaining, from content items corresponding to the context, a number of sampled content items” “wherein a content item feature vector is generated based on a description of a sampled content item, metadata of the sampled content item, one or more past engagement statistics of the sampled content item, and a retrieval strategy used to retrieve the sampled content item” These recitations are deemed insufficient to transform the judicial exception to a patentable invention because the recitation is directed to instructions for mere data gathering or data output, see MPEP 2106.05(g). “and the action vector is a feature representation of weights given to different retrieval strategies” “updating the parameters of the agent model based on the context, the action vector, and the reward” These recitations are deemed insufficient to transform the judicial exception to a patentable invention because the recitation is directed to instructions merely indicating a field of use or technological environment in which to apply a judicial exception, see MPEP 2106.05(h). Accordingly, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea when considered as an ordered combination and as a whole. Step 2B: This part of the eligibility analysis evaluates whether the claim, as a whole, amounts to significantly more than the recited exception, i.e., whether any additional element, or combination of additional elements, adds an inventive concept to the claim. MPEP 2106.05. First, the additional elements considered as part of the preamble and the additional elements directed to the use of computer technology are deemed insufficient to transform the judicial exception to a patentable invention to a patentable invention because they generally link the judicial exception to the technology environment, see MPEP 2106.05(h). Second, the additional elements mere application of the abstract idea or mere instructions to implement an abstract idea on a computer are deemed insufficient to transform the judicial exception to a patentable invention to a patentable invention because the limitations generally apply the use of a generic computer and/or process with the judicial exception, see MPEP 2106.05(f). Lastly, the claims are directed to instructions merely indicating a field of use or technological environment in which to apply a judicial exception. The courts have found these types of limitations insufficient to transform the judicial exception to a patentable invention, see MPEP 2106.05(g). Furthermore, when considering evidence in view of Berkheimer v. HP, Inc., 881 F.3d 1360, 1368, 125 USPQ2d 1649, 1654 (Fed. Cir. 2018), see USPTO Berkheimer Memorandum (April 2018). Examiner notes Berkheimer: Option 2 - A citation to one or more of the court decisions discussed in MPEP § 2106.05(d}(II} as noting the well understood, routine, conventional nature of the additional element (s) (e.g., limitations directed to mere data gathering): The courts have recognized the following computer functions as well understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity, see MPEP 2106.05(d). The additional limitations, as analyzed, failed to integrate a judicial exception into a practical application at Step 2A and provide an inventive concept in Step 2B, per the analysis above. Thus, considering the additional elements individually and in combination and the claims as a whole, the additional elements do not provide significantly more than the abstract idea. This claim is not patent eligible. Therefore, in examining elements as recited by the limitations individually and as an ordered combination, as a whole, claim 17 does not recite what the courts have identified as "significantly more". Furthermore, regarding dependent claims 2-8, which depend from claim 1, claims 10-16, which depend from claim 9, claims 18-20, which depend from claim 17, the claims are directed to a judicial exception (i.e., an abstract idea enumerated in the 2019 PEG, a law of nature, or a natural phenomenon) without significantly more as highlighted below in the claim limitations by evaluating the claim limitations under the Step2A and 2B: Claims 2, 10: Incorporates the rejections of claims 1, 9, respectively. “wherein obtaining the query comprises randomly sampling from historical logs of user activity on a content platform” (mere data gathering by mental process of observation) These recitations are deemed insufficient to transform the judicial exception to a patentable invention because the recitation is directed to instructions for mere data gathering or data output and to mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and are considered to adding the words "apply it" (or an equivalent) with the judicial exception, see MPEP 2106.05(g), MPEP 2106.05(f). Limitations directed to instructions for mere data gathering or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B. Claims 3, 11: Incorporates the rejections of claims 1, 9, respectively. “wherein obtaining the number of sampled content items comprises randomly sampling a first number of positive content items and a second number of positive content items associated with the query using historical logs of user activity on a content platform” (mere data gathering by mental process of observation) These recitations are deemed insufficient to transform the judicial exception to a patentable invention because the recitation is directed to instructions for mere data gathering or data output and to mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and are considered to adding the words "apply it" (or an equivalent) with the judicial exception, see MPEP 2106.05(g), MPEP 2106.05(f). Limitations directed to instructions for mere data gathering or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B. Claims 4, 12: Incorporates the rejections of claims 1, 9, respectively. “wherein computing the reward comprises: determining a proportion of positive content items in the top number of the content items having the highest dot products” (mental process of judgement and evaluation) The recitation is directed to mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and are considered to adding the words "apply it" (or an equivalent) with the judicial exception, See MPEP 2106.05(f). Limitations directed to mere instructions to implement an abstract idea on a computer/using computer as a tool cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B. Claims 5, 13: Incorporates the rejections of claims 1, 9, respectively. “wherein computing the reward comprises: determining a number of positive content items in the top number of the content items having the highest dot products relative to a total number of positive content items in the number of sampled content items” (mental process of judgement and evaluation) The recitation is directed to mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and are considered to adding the words "apply it" (or an equivalent) with the judicial exception, See MPEP 2106.05(f). Limitations directed to mere instructions to implement an abstract idea on a computer/using computer as a tool cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B. Claims 6, 14, 18: Incorporates the rejections of claims 1, 9, 17, respectively. “wherein computing the reward comprises: determining a sum of reciprocal rank(s) of positive content item(s) in the top number of content items having the highest dot products” (mental process of judgement and evaluation) The recitation is directed to mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and are considered to adding the words "apply it" (or an equivalent) with the judicial exception, See MPEP 2106.05(f). Limitations directed to mere instructions to implement an abstract idea on a computer/using computer as a tool cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B. Claims 7, 15, 19: Incorporates the rejections of claims 1, 9, 17, respectively. “wherein computing the reward comprises: determining a key reciprocal rank of a top positive content item having a highest dot product in the top number of content items having the highest dot products” (mental process of judgement and evaluation) The recitation is directed to mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and are considered to adding the words "apply it" (or an equivalent) with the judicial exception, See MPEP 2106.05(f). Limitations directed to mere instructions to implement an abstract idea on a computer/using computer as a tool cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B. Claims 8, 16, 20: Incorporates the rejections of claims 1, 9, 17, respectively. “wherein computing the reward comprises: subtracting a maximum possible reward value of the top number of content items having the highest dot products by a reward value of the top number of content items having the highest dot products” (mental process of judgement and evaluation by mathematics) The recitation is directed to mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and are considered to adding the words "apply it" (or an equivalent) with the judicial exception, See MPEP 2106.05(f). Limitations directed to mere instructions to implement an abstract idea on a computer/using computer as a tool cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B. The dependent claims as analyzed above, do not recite limitations that integrated the judicial exception into a practical application. In addition, the claim limitations do not include additional elements that are sufficient to amount to significantly more than the judicial exception (Step-2B). Therefore, the claims do not recite any limitations, when considered individually or as a whole, that recite what have the courts have identified as "significantly more", see MPEP 2106.05; and therefore, as a whole the claims are not patent eligible. As shown above, the dependent claims do not provide any additional elements that when considered individually or as an ordered combination, amount to significantly more than the abstract idea identified. Therefore, as a whole, the dependent claims do not recite what have the courts have identified as "significantly more" than the recited judicial exception. Therefore, claims 2-8, 10-16, 18-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception and does not recite, when claim elements are examined individually and as a whole, elements that the courts have identified as "significantly more" than the recited judicial exception. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 1, 9, 17 are rejected under 35 U.S.C. 103 as being unpatentable over Bodigutla et al. (U.S. Pre-Grant Publication No. 20220100756, hereinafter ‘Bodigutla'), in view of Kabra et al. (NPL: "Potent Real-Time Recommendations Using Multimodel Contextual Reinforcement Learning", hereinafter 'Kabra'), Zhao et al. (NPL: "Deep Reinforcement Learning for Page-wise Recommendations", hereinafter 'Zhao'), and further in view of Reddy et al. (NPL: "Enhancing Content Personalization at Scale Using Deep Reinforcement Learning and Collaborative Filtering Techniques", hereinafter 'Reddy'). Regarding claim 1 and analogous claim 9, Bodigutla teaches A method, comprising: obtaining a query ([0055] Through computer-based interactions with an end user, search interface 202 extracts user state data 204, user metadata 205, and search query query 206 from the search session.); obtaining, from content items corresponding to the query, a number of sampled content items ([0032] Examples of computer-generated navigation elements that the disclosed technologies may generate and provide to the search interface at any time during a session include but are not limited to computer-generated search re-formulations, such as re-formulations of the user's original query that may refine or expand the user's previous query, computer-generated dynamic re-configurations of search filters and/or facet types, computer-generated conversational query disambiguation elements such as clarifying prompts and informational content elements such as coaching videos and help messages designed to help a new user navigate a search page, presentations of search results retrieved by the search engine, or any combination of any of the foregoing or other forms of navigation elements. The obtaining, from content items corresponding to the query, a number of sampled content items presentation of search results is, for purposes of this disclosure, considered a computer-generated navigation element because the presentation of search results is an option that may be selected by the navigation agent. For example, the navigation agent may determine to both display search results and display search re-formulations and/or re-configured filters.); for each sampled content item, determining a dot product of the action vector and the content item feature vector ([0129] Operation 456 when executed by at least one processor causes one or more computing devices to compute reward scores for candidate filter elements determined by operation 454. In an embodiment, a reward score for a for each sampled content item, determining a dot product of the action vector and the content item feature vector candidate filter element is computed as the dot product between the action weight vector a and the entity embedding e. The candidate filter elements are sampled; that is, each candidate filter element has a different probability of success.); sorting the sampled content items based on dot products; computing a reward based on a top number of sampled content items having highest dot products; and updating the parameters of the agent model based on the query, the action vector, and the reward ([0130] Operation 458 when executed by at least one processor causes one or more computing devices sorting the sampled content items based on dot products to select filters based on the reward scores computed in operation 456. In an embodiment, a computing a reward based on a top number of sampled content items having highest dot products reinforcement learning agent computes a and the reward reward score r, based on the query given a user state s and the action vector a system action a (e.g., presentation of a filter element), based on subsequent user state data indicating user feedback such as click, negate, not click, send message, save the recommended filter, etc. A discount parameter γ measures the present value of future rewards, where a future reward is a reward score computed for a subsequent user state, for example. When γ=0, the reinforcement learning agent only considers immediate rewards, e.g., rewards computed using feedback received only on the current state st and ignores long term rewards, e.g., reward scores computed using feedback received over the course of the entire session. When γ=1, long term rewards are considered as equally important as immediate rewards. Within a search session, rewards may be defined as positive integers indicating the relative significance of various user activities, for instance, r=0 if the recommended entity is not clicked, r=1 if the recommended entity is clicked, r=2 if a positive subsequent action is detected, such as sending a message, viewing a user profile, etc. To select a filter, a deterministic policy gradient algorithm may be used; for example, a deep deterministic policy gradient algorithm (DDPG). FIG. 2F, described above, shows an example of a DDPG algorithm that may be used in operation 458.; [0131] The reinforcement learning-based approach to dynamically generating filter element options enables the updating the parameters of the agent model system 100 to adapt to changes in user behavior and respond differently to different types of queries. For example, once a user selects a filter element, the system 100 dynamically determines one or more additional filters to display and/or an order of arrangement on a display, and automatically refines other candidate filters based on the user state data prior to the presentation of the filter and subsequent user state data (feedback).). Bodigutla fails to teach determining content item feature vectors corresponding to the sampled content items, wherein a content item feature vector is generated based on a description of a sampled content item, metadata of the sampled content item, one or more past engagement statistics of the sampled content item, and a retrieval strategy used to retrieve the sampled content item; determining, using parameters of an agent model and an embedding of the query, an action vector comprising features corresponding to elements in the content item feature vector, and the action vector is a feature representation of weights given to different retrieval strategies; Kabra teaches determining content item feature vectors corresponding to the sampled content items, wherein a content item feature vector is generated based on a description of a sampled content item, metadata of the sampled content item, one or more past engagement statistics of the sampled content item, and a retrieval strategy used to retrieve the sampled content item ([I. INTRODUCTION, pg. 582] In this article, any subsequent mention of the term environment refers to the complete set of users and items available in the dataset. The system refers to our MMCR framework and the term model corresponds to the per-item neural network. State refers to the determining content item feature vectors corresponding to the sampled content items embedding generated using user features, session context features, and user’s item history. The terms arm and action used interchangeably in the text indicate items in the itemset. When a user requests the system for recommendations, his/her specific features along with session context features serve as the input. This input is further processed by the state generator (SG) module to obtain state embedding.; [A. SG Module, pg. 584] wherein a content item feature vector is generated Generating State Embeddings: Auto-encoder is used at this step for generating uniform embeddings. The user and based on a description of a sampled content item context features from the input request are concatenated with the one or more past engagement statistics of the sampled content item interaction history feature vector ih. v = ut +ct +ih. This concatenated vector v is fed as an input to the autoencoder. The output from the autoencoder is the required state embedding st.); Bodigutla and Kabra are considered to be analogous to the claimed invention because they are in the same field of machine learning. In view of the teachings of Bodigutla, it would have been obvious for a person of ordinary skill in the art to apply the teachings of Kabra to Bodigutla before the effective filing date of the claimed invention in order to significantly improve results over various state-of-the-art strategies (cf. Kabra, [Abstract] Our technique ensures that the user item history per item is learned separately; thus, no item is neglected. MMCR has shown an average of 5% increase in CTR rate. Moreover, CCE exploration gives a considerably higher score than state-of-the-art exploration strategies. Thorough experimentation demonstrates that our proposed strategy has shown significantly improved results over various state-of-the-art strategies.). Zhao teaches determining content item feature vectors corresponding to the sampled content items, wherein a content item feature vector is generated based on a description of a sampled content item, metadata of the sampled content item, one or more past engagement statistics of the sampled content item, and a retrieval strategy used to retrieve the sampled content item ([2.1 Framework Overview] As mentioned in Section 1.1, we model the recommendation task as a Markov Decision Process (MDP) and leverage Reinforcement Learning (RL) to automatically learn the determining content item feature vectors corresponding to the sampled content items optimal recommendation strategies, which can continuously update recommendation strategies during the interactions and the optimal strategy is made by maximizing the expected long-term cumulative reward from users. With the above intuitions, we formally define the wherein a content item feature vector is generated based tuple of five elements (S,A,P,R,γ) of MDP– (a) State space S: A state s ∈Sis defined as user’s current preference, which is generated based on user’s browsing history, i.e., the items that a user browsed and her corresponding feedback; (b) Action space A: An action a = {a1,··· ,aM}∈Ais to recommend a page of M items to a user based on current state s; (c) and a retrieval strategy used to retrieve the sampled content item Reward R: After the RA takes an action a at the state s, i.e., recommending a page of items to a user, the user browses these items and provides her feedback. She can skip (not click), click, or purchase these items, and the agent receives immediate reward r(s,a) according to the user’s feedback; (d) Transition P: Transition p(s′|s,a) defines the state transition from s to s′ when RA takes action a; and (e) Discount factor γ: γ ∈[0,1] defines the discount factor when we measure the present value of future reward. In particular, when γ = 0, RA only considers the immediate reward. In other words, when γ = 1, all future rewards can be counted fully into that of the current action.); determining, using parameters of an agent model and an embedding of the query, an action vector comprising features corresponding to elements in the content item feature vector, and the action vector is a feature representation of weights given to different retrieval strategies ([2.1 Framework Overview, pg. 96] As mentioned in Section 1.1, we model the recommendation task as a Markov Decision Process (MDP) and leverage Reinforcement Learning (RL) to automatically learn the optimal recommendation strategies, which can continuously update recommendation strategies during the interactions and the optimal strategy is made by maximizing the expected long-term cumulative reward from users. With the above intuitions, we formally define the tuple of five elements (S,A,P,R,γ) of MDP– (a) using parameters of an agent model and an embedding of the query State space S: A state s ∈Sis defined as user’s current preference, which is generated based on user’s browsing history, i.e., the items that a user browsed and her corresponding feedback; (b) determining an action vector comprising features corresponding to elements in the content item feature vector Action space A: An action a = {a1,··· ,aM}∈Ais to recommend a page of M items to a user based on current state s; (c)Reward R: After the RA takes an action a at the state s, i.e., recommending a page of items to a user, the user browses these items and provides her feedback. She can skip (not click), click, or purchase these items, and the agent receives immediate reward r(s,a) according to the user’s feedback; (d) the action vector is a feature representation of weights given to different retrieval strategies Transition P: Transition p(s′|s,a) defines the state transition from s to s′ when RA takes action a; and (e) Discount factor γ: γ ∈[0,1] defines the discount factor when we measure the present value of future reward. In particular, when γ = 0, RA only considers the immediate reward. In other words, when γ = 1, all future rewards can be counted fully into that of the current action.); Bodigutla, Kabra, and Zhao are considered to be analogous to the claimed invention because they are in the same field of machine learning. In view of the teachings of Bodigutla and Kabra, it would have been obvious for a person of ordinary skill in the art to apply the teachings of Zhao to Bodigutla before the effective filing date of the claimed invention in order to jointly generate a set of complementary items and the corresponding strategy to display them in a 2-D page, optimize a page of items with proper display based on real-time feedback from users (cf. Zhao, [ABSTRACT] In this paper, we study the problem of page-wise recommendations aiming to address aforementioned two challenges simultaneously. In particular, we propose a principled approach to jointly generate a set of complementary items and the corresponding strategy to display them in a 2-D page; and propose a novel page-wise recommendation framework based on deep reinforcement learning, DeepPage, which can optimize a page of items with proper display based on real-time feedback from users. The experimental results based on a real-world e-commerce dataset demonstrate the effectiveness of the proposed framework.). Reddy teaches determining content item feature vectors corresponding to the sampled content items, wherein a content item feature vector is generated based on a description of a sampled content item, metadata of the sampled content item, one or more past engagement statistics of the sampled content item, and a retrieval strategy used to retrieve the sampled content item ([METHODOLOGY, pg. 7-8] 1. Problem Definition and Data Collection: The primary goal is to personalize content recommendations to improve user satisfaction and engagement. We begin by defining the scope of personalization, target metrics, and user demographics. Data collection involves gathering a comprehensive dataset that includes user interactions, metadata of the sampled content item content metadata, and contextual information from a large-scale recommendation system.; 2. Preprocessing and Feature Engineering: Preprocessing is crucial to handle missing values, normalize data, and identify relevant features for both DRL and CF models. Feature engineering involves crafting user profiles, content attributes, and contextual features. wherein a content item feature vector is generated based on Techniques such as one-hot encoding, embedding layers, and principal component analysis (PCA) may be used to reduce dimensionality and enhance feature representation.); Bodigutla, Kabra, Zhao, and Reddy are considered to be analogous to the claimed invention because they are in the same field of machine learning. In view of the teachings of Bodigutla, Kabra, and Zhao, it would have been obvious for a person of ordinary skill in the art to apply the teachings of Reddy to Bodigutla before the effective filing date of the claimed invention in order to leverage DRL's capacity for learning complex user interaction patterns, demonstrating significant improvements in user satisfaction metrics, engagement rates, and system efficiency compared to baseline models (cf. Reddy, [ABSTRACT] To address these challenges, we propose a novel framework that leverages DRL's capacity for learning complex user interaction patterns and CF's proficiency in harnessing user-item relationships. The DRL model is trained to make sequential decisions, adjusting content recommendations in real-time based on user interactions, while the CF component enhances prediction accuracy by analyzing user similarity matrices and latent factors. Our approach is tested on large-scale datasets, demonstrating significant improvements in user satisfaction metrics, engagement rates, and system efficiency compared to baseline models. We further analyze the model's ability to adapt to evolving user preferences and its robustness in handling sparse and noisy data. The findings underscore the potential of combining DRL and CF to push the boundaries of content personalization, offering a scalable solution for con tent providers aiming to deliver more relevant and engaging user experiences. This study not only contributes to the theoretical understanding of personalized content delivery systems but also provides practical insights for deploying such systems in real-world applications.). Regarding claim 17, Bodigutla teaches A method, comprising: obtaining a context ([0055] Through computer-based interactions with an end user, search interface 202 extracts context user state data 204, user metadata 205, and search query 206 from the search session.); obtaining, from content items corresponding to the context, a number of sampled content items ([0032] Examples of computer-generated navigation elements that the disclosed technologies may generate and provide to the search interface at any time during a session include but are not limited to computer-generated search re-formulations, such as re-formulations of the user's original query that may refine or expand the user's previous query, computer-generated dynamic re-configurations of search filters and/or facet types, computer-generated conversational query disambiguation elements such as clarifying prompts and informational content elements such as coaching videos and help messages designed to help a new user navigate a search page, presentations of search results retrieved by the search engine, or any combination of any of the foregoing or other forms of navigation elements. The obtaining, from content items corresponding to the context, a number of sampled content items presentation of search results is, for purposes of this disclosure, considered a computer-generated navigation element because the presentation of search results is an option that may be selected by the navigation agent. For example, the navigation agent may determine to both display search results and display search re-formulations and/or re-configured filters.); for each sampled content item, determining a dot product of the action vector and the content item feature vector ([0129] Operation 456 when executed by at least one processor causes one or more computing devices to compute reward scores for candidate filter elements determined by operation 454. In an embodiment, a reward score for a for each sampled content item, determining a dot product of the action vector and the content item feature vector candidate filter element is computed as the dot product between the action weight vector a and the entity embedding e. The candidate filter elements are sampled; that is, each candidate filter element has a different probability of success.); sorting the sampled content items based on dot products; computing a reward based on a top number of sampled content items having highest dot products; and updating the parameters of the agent model based on the context, the action vector, and the reward ([0130] Operation 458 when executed by at least one processor causes one or more computing devices sorting the sampled content items based on dot products to select filters based on the reward scores computed in operation 456. In an embodiment, a computing a reward based on a top number of sampled content items having highest dot products reinforcement learning agent computes a and the reward reward score r, based on the context given a user state s and the action vector a system action a (e.g., presentation of a filter element), based on subsequent user state data indicating user feedback such as click, negate, not click, send message, save the recommended filter, etc. A discount parameter γ measures the present value of future rewards, where a future reward is a reward score computed for a subsequent user state, for example. When γ=0, the reinforcement learning agent only considers immediate rewards, e.g., rewards computed using feedback received only on the current state st and ignores long term rewards, e.g., reward scores computed using feedback received over the course of the entire session. When γ=1, long term rewards are considered as equally important as immediate rewards. Within a search session, rewards may be defined as positive integers indicating the relative significance of various user activities, for instance, r=0 if the recommended entity is not clicked, r=1 if the recommended entity is clicked, r=2 if a positive subsequent action is detected, such as sending a message, viewing a user profile, etc. To select a filter, a deterministic policy gradient algorithm may be used; for example, a deep deterministic policy gradient algorithm (DDPG). FIG. 2F, described above, shows an example of a DDPG algorithm that may be used in operation 458.; [0131] The reinforcement learning-based approach to dynamically generating filter element options enables the updating the parameters of the agent model system 100 to adapt to changes in user behavior and respond differently to different types of queries. For example, once a user selects a filter element, the system 100 dynamically determines one or more additional filters to display and/or an order of arrangement on a display, and automatically refines other candidate filters based on the user state data prior to the presentation of the filter and subsequent user state data (feedback).). Bodigutla fails to teach determining content item feature vectors corresponding to the sampled content items, wherein a content item feature vector is generated based on a description of a sampled content item, metadata of the sampled content item, one or more past engagement statistics of the sampled content item, and a retrieval strategy used to retrieve the sampled content item; determining, using parameters of an agent model and an embedding of the context, an action vector comprising features corresponding to elements in the content item feature vector, and the action vector is a feature representation of weights given to different retrieval strategies; Kabra teaches determining content item feature vectors corresponding to the sampled content items, wherein a content item feature vector is generated based on a description of a sampled content item, metadata of the sampled content item, one or more past engagement statistics of the sampled content item, and a retrieval strategy used to retrieve the sampled content item ([I. INTRODUCTION, pg. 582] In this article, any subsequent mention of the term environment refers to the complete set of users and items available in the dataset. The system refers to our MMCR framework and the term model corresponds to the per-item neural network. State refers to the determining content item feature vectors corresponding to the sampled content items embedding generated using user features, session context features, and user’s item history. The terms arm and action used interchangeably in the text indicate items in the itemset. When a user requests the system for recommendations, his/her specific features along with session context features serve as the input. This input is further processed by the state generator (SG) module to obtain state embedding.; [A. SG Module, pg. 584] wherein a content item feature vector is generated Generating State Embeddings: Auto-encoder is used at this step for generating uniform embeddings. The user and based on a description of a sampled content item context features from the input request are concatenated with the one or more past engagement statistics of the sampled content item interaction history feature vector ih. v = ut +ct +ih. This concatenated vector v is fed as an input to the autoencoder. The output from the autoencoder is the required state embedding st.); Bodigutla and Kabra are combinable for the same rationale as set forth above with respect to claim 1. Zhao teaches determining content item feature vectors corresponding to the sampled content items, wherein a content item feature vector is generated based on a description of a sampled content item, metadata of the sampled content item, one or more past engagement statistics of the sampled content item, and a retrieval strategy used to retrieve the sampled content item ([2.1 Framework Overview] As mentioned in Section 1.1, we model the recommendation task as a Markov Decision Process (MDP) and leverage Reinforcement Learning (RL) to automatically learn the determining content item feature vectors corresponding to the sampled content items optimal recommendation strategies, which can continuously update recommendation strategies during the interactions and the optimal strategy is made by maximizing the expected long-term cumulative reward from users. With the above intuitions, we formally define the wherein a content item feature vector is generated based tuple of five elements (S,A,P,R,γ) of MDP– (a) State space S: A state s ∈Sis defined as user’s current preference, which is generated based on user’s browsing history, i.e., the items that a user browsed and her corresponding feedback; (b) Action space A: An action a = {a1,··· ,aM}∈Ais to recommend a page of M items to a user based on current state s; (c) and a retrieval strategy used to retrieve the sampled content item Reward R: After the RA takes an action a at the state s, i.e., recommending a page of items to a user, the user browses these items and provides her feedback. She can skip (not click), click, or purchase these items, and the agent receives immediate reward r(s,a) according to the user’s feedback; (d) Transition P: Transition p(s′|s,a) defines the state transition from s to s′ when RA takes action a; and (e) Discount factor γ: γ ∈[0,1] defines the discount factor when we measure the present value of future reward. In particular, when γ = 0, RA only considers the immediate reward. In other words, when γ = 1, all future rewards can be counted fully into that of the current action.); determining, using parameters of an agent model and an embedding of the context, an action vector comprising features corresponding to elements in the content item feature vector, and the action vector is a feature representation of weights given to different retrieval strategies ([2.1 Framework Overview, pg. 96] As mentioned in Section 1.1, we model the recommendation task as a Markov Decision Process (MDP) and leverage Reinforcement Learning (RL) to automatically learn the optimal recommendation strategies, which can continuously update recommendation strategies during the interactions and the optimal strategy is made by maximizing the expected long-term cumulative reward from users. With the above intuitions, we formally define the tuple of five elements (S,A,P,R,γ) of MDP– (a) using parameters of an agent model and an embedding of the context State space S: A state s ∈Sis defined as user’s current preference, which is generated based on user’s browsing history, i.e., the items that a user browsed and her corresponding feedback; (b) determining an action vector comprising features corresponding to elements in the content item feature vector Action space A: An action a = {a1,··· ,aM}∈Ais to recommend a page of M items to a user based on current state s; (c)Reward R: After the RA takes an action a at the state s, i.e., recommending a page of items to a user, the user browses these items and provides her feedback. She can skip (not click), click, or purchase these items, and the agent receives immediate reward r(s,a) according to the user’s feedback; (d) the action vector is a feature representation of weights given to different retrieval strategies Transition P: Transition p(s′|s,a) defines the state transition from s to s′ when RA takes action a; and (e) Discount factor γ: γ ∈[0,1] defines the discount factor when we measure the present value of future reward. In particular, when γ = 0, RA only considers the immediate reward. In other words, when γ = 1, all future rewards can be counted fully into that of the current action.); Bodigutla, Kabra, and Zhao are combinable for the same rationale as set forth above with respect to claim 1. Reddy teaches determining content item feature vectors corresponding to the sampled content items, wherein a content item feature vector is generated based on a description of a sampled content item, metadata of the sampled content item, one or more past engagement statistics of the sampled content item, and a retrieval strategy used to retrieve the sampled content item ([METHODOLOGY, pg. 7-8] 1. Problem Definition and Data Collection: The primary goal is to personalize content recommendations to improve user satisfaction and engagement. We begin by defining the scope of personalization, target metrics, and user demographics. Data collection involves gathering a comprehensive dataset that includes user interactions, metadata of the sampled content item content metadata, and contextual information from a large-scale recommendation system.; 2. Preprocessing and Feature Engineering: Preprocessing is crucial to handle missing values, normalize data, and identify relevant features for both DRL and CF models. Feature engineering involves crafting user profiles, content attributes, and contextual features. wherein a content item feature vector is generated based on Techniques such as one-hot encoding, embedding layers, and principal component analysis (PCA) may be used to reduce dimensionality and enhance feature representation.); Bodigutla, Kabra, Zhao, and Reddy are combinable for the same rationale as set forth above with respect to claim 1. Claims 2, 10 are rejected under 35 U.S.C. 103 as being unpatentable over Bodigutla, in view of Kabra, Zhao, Reddy, and further in view of Yerva et al. (U.S. Pre-Grant Publication No. 20190286656, hereinafter 'Yerva'). Regarding claim 2 and analogous claim 10, Bodigutla, as modified by Kabra, Zhao, and Reddy, teaches The method of claim 1, The one or more non-transitory computer-readable media of claim 9, respectively. Yerva teaches wherein obtaining the query comprises randomly sampling from historical logs of user activity on a content platform ([0076] In at least one embodiment, training triplet inputs <Q, P, N> are generated and/or collected using randomly sampled food search logs, which are stored in the memory 206 (e.g., the operational records 226) and produced by past search activities of users of the health tracking system 100. In one embodiment, the processing circuitry/logic 204 of the server 200 is configured to randomly select a set of past queries Q from the food search logs and retrieve a subset of consumable records 224 and/or food names thereof that have frequently appeared within the top search results (e.g., top 5) for those queries Q, based on the food search logs.). Bodigutla, Kabra, Zhao, Reddy, and Yerva are considered to be analogous to the claimed invention because they are in the same field of machine learning. In view of the teachings of Bodigutla, Kabra, Zhao, and Reddy, it would have been obvious for a person of ordinary skill in the art to apply the teachings of Yerva to Bodigutla before the effective filing date of the claimed invention in order to provide more relevant search results and recommendations (cf. Yerva, [0026] With reference to FIG. 1, an exemplary embodiment of a health tracking system 100 that utilizes deep multi-modal pairwise ranking of consumable records to provide more relevant search results and recommendations is shown. In the illustrated embodiment, the health tracking system 100 includes a plurality of health tracking devices 110 in communication with a system server 200 or other data processing system over a network 120 such as, e.g. the Internet.). Claims 3, 11 are rejected under 35 U.S.C. 103 as being unpatentable over Bodigutla, in view of Kabra, Zhao, Reddy, in view of Lin et al. (NPL: "How to Train Your DRAGON: Diverse Augmentation Towards Generalizable Dense Retrieval", hereinafter 'Lin'), and further in view of Yerva et al. (U.S. Pre-Grant Publication No. 20190286656, hereinafter 'Yerva'). Regarding claim 3 and analogous claim 11, Bodigutla, as modified by Kabra, Zhao, and Reddy, teaches The method of claim 1, The one or more non-transitory computer-readable media of claim 9, respectively. Lin teaches wherein obtaining the number of sampled content items comprises randomly sampling a first number of positive content items and a second number of positive content items associated with the query using historical logs of user activity on a content platform ([A.5 Impacts of Top-k Positive Sampling, pg. 6398] Table 8: Ablation on progressive label augmentation from top-k passages using cropped sentences as queries. top-k positives 1 5 10 MARCO Dev(RR@10) 33.1 36.4 36.6 BEIR-13 (nDCG@10) 42.4 48.0 49.3 ∗ Trajectory: uniCOIL → Contriever → ColBERTv2. In Section 3.2, we mention that our sampling scheme treats top 10 passages from each teacher ranked list as positives and top 45–50 as negatives. We further conduct experiment to study the im pact of the positive sampling scheme. Following the experiment setups in Section 3.3, we use the sentences cropped from MS MARCO corpus as augmented queries and we conduct progressive la bel augmentation using top-k passages as positive. The results are tabulated in Table 8. We observe that treating top-10 passages from each teacher as positives yields the best supervised and zero-shot effectiveness. On the other hand, using only the top passage as positive results in significant effective ness drop. This result indicates that the top passage labeled by a teacher cannot transfer its knowledge well to a student. This result is similar to the obser vation from Chen et al. (2022).). Bodigutla, Kabra, Zhao, Reddy, and Lin are considered to be analogous to the claimed invention because they are in the same field of machine learning. In view of the teachings of Bodigutla, Kabra, Zhao, and Reddy, it would have been obvious for a person of ordinary skill in the art to apply the teachings of Lin to Bodigutla before the effective filing date of the claimed invention in order to progressively train a generalizable dense retrieval (DR) (cf. Lin, [Abstract] Our study shows that commonDA practices such as query augmentation with generative models and pseudo-relevance label creation using a cross-encoder, are often inefficient and sub-optimal. We hence propose a new DA approach with diverse queries and sources of supervision to progressively train a generalizable DR. As a result, DRAGON, 1 our Dense Retriever trained with diverse AuGmentatiON, is the first BERT base-sized DR to achieve state-of-the-art effectiveness in both supervised and zero-shot evaluations and even competes with models using more complex late interaction (ColBERTv2 and SPLADE++).). Yerva teaches wherein obtaining the number of sampled content items comprises randomly sampling a first number of positive content items and a second number of positive content items associated with the query using historical logs of user activity on a content platform ([0076] In at least one embodiment, training triplet inputs <Q, P, N> are generated and/or collected using randomly sampled food search logs, which are stored in the memory 206 (e.g., the operational records 226) and produced by past search activities of users of the health tracking system 100. In one embodiment, the processing circuitry/logic 204 of the server 200 is configured to randomly select a set of past queries Q from the food search logs and retrieve a subset of consumable records 224 and/or food names thereof that have frequently appeared within the top search results (e.g., top 5) for those queries Q, based on the food search logs.). Bodigutla, Kabra, Zhao, Reddy, Lin, and Yerva are considered to be analogous to the claimed invention because they are in the same field of machine learning. In view of the teachings of Bodigutla, Kabra, Zhao, Reddy, and Lin, it would have been obvious for a person of ordinary skill in the art to apply the teachings of Yerva to Bodigutla before the effective filing date of the claimed invention in order to provide more relevant search results and recommendations (cf. Yerva, [0026] With reference to FIG. 1, an exemplary embodiment of a health tracking system 100 that utilizes deep multi-modal pairwise ranking of consumable records to provide more relevant search results and recommendations is shown. In the illustrated embodiment, the health tracking system 100 includes a plurality of health tracking devices 110 in communication with a system server 200 or other data processing system over a network 120 such as, e.g. the Internet.). Claims 4, 12 are rejected under 35 U.S.C. 103 as being unpatentable over Bodigutla, in view of Kabra, Zhao, Reddy, and further in view of Zheng et al. (NPL: "DRN: A Deep Reinforcement Learning Framework for News Recommendation", hereinafter 'Zheng'). Regarding claim 4 and analogous claim 12, Bodigutla, as modified by Kabra, Zhao, and Reddy, teaches The method of claim 1, The one or more non-transitory computer-readable media of claim 9, respectively. Zheng teaches wherein computing the reward comprises: determining a proportion of positive content items in the top number of the content items having the highest dot products ([5.2 Evaluation measures, pg. 173] •Precision@k [10]. Precision at k is calculated as Equation 10 Precision@k = number of clicks in top-k recommended items / k (10)). Bodigutla, Kabra, Zhao, Reddy, and Zheng are considered to be analogous to the claimed invention because they are in the same field of machine learning. In view of the teachings of Bodigutla, Kabra, Zhao, and Reddy, it would have been obvious for a person of ordinary skill in the art to apply the teachings of Zheng to Bodigutla before the effective filing date of the claimed invention in order to propose a Deep Q-Learning based recommendation framework, which can model future reward explicitly (cf. Zheng, [ABSTRACT] Therefore, to address the aforementioned challenges, we propose a Deep Q-Learning based recommendation framework, which can model future reward explicitly. We further consider user return pattern as a supplement to click / no click label in order to capture more user feedback information. In addition, an effective exploration strategy is incorporated to find new attractive news for users. Extensive experiments are conducted on the offline dataset and online production environment of a commercial news recommendation application and have shown the superior performance of our methods.). Claims 5, 13 are rejected under 35 U.S.C. 103 as being unpatentable over Bodigutla, in view of Kabra, Zhao, Reddy, and further in view of Stamenkovic et al. (NPL: "Choosing the Best of Both Worlds: Diverse and Novel Recommendations through Multi-Objective Reinforcement Learning", hereinafter 'Stamenkovic'). Regarding claim 5 and analogous claim 13, Bodigutla, as modified by Kabra, Zhao, and Reddy, teaches The method of claim 1, The one or more non-transitory computer-readable media of claim 9, respectively. Stamenkovic teaches wherein computing the reward comprises: determining a number of positive content items in the top number of the content items having the highest dot products relative to a total number of positive content items in the number of sampled content items ([5.1.2 Quality of Recommendation Metrics., pg. 962] Accuracy metrics. Relevance of the recommended item set is usually measured with two metrics: hit ration (HR) and normalized discounted cumulative gain (NDCG). HR@𝑘 is a recall-based metric, measuring whether the ground-truth item is in the top-𝑘 positions of the recommendation list. We define HR for clicks as: HR(click) = #hits among clicks #clicks On the other hand, NDCG is a rank sensitive metric that assigns higher scores to top positions in the recommendation list [18]). Bodigutla, Kabra, Zhao, Reddy, and Stamenkovic are considered to be analogous to the claimed invention because they are in the same field of machine learning. In view of the teachings of Bodigutla, Kabra, Zhao, and Reddy, it would have been obvious for a person of ordinary skill in the art to apply the teachings of Stamenkovic to Bodigutla before the effective filing date of the claimed invention in order to increase in aggregate diversity, a moderate increase in accuracy, reduced repetitiveness of recommendations, and demonstrate the importance of reinforcing diversity and novelty as complementary objectives (cf. Stamenkovic, [ABSTRACT] In this work, we take on the aforementioned challenge and introduce Scalarized Multi-Objective Reinforcement Learning (SMORL) for the RS setting, a novel Reinforcement Learning (RL) framework that can effectively address multi-objective recommendation tasks. The proposed SMORL agent augments standard recommendation models with additional RL layers that enforce it to simultaneously satisfy three principal objectives: accuracy, diversity, and novelty of recommendations. We integrate this framework with four state of-the-art session-based recommendation models and compare it with a single-objective RL agent that only focuses on accuracy. Our experimental results on two real-world datasets reveal a substantial increase in aggregate diversity, a moderate increase in accuracy, reduced repetitiveness of recommendations, and demonstrate the importance of reinforcing diversity and novelty as complementary objectives.). Claims 6, 14, 18 are rejected under 35 U.S.C. 103 as being unpatentable over Bodigutla, in view of Kabra, Zhao, Reddy, and further in view of Lin et al. (NPL: "How to Train Your DRAGON: Diverse Augmentation Towards Generalizable Dense Retrieval", hereinafter 'Lin'). Regarding claim 6 and analogous claims 14, 18, Bodigutla, as modified by Kabra, Zhao, and Reddy, teaches The method of claim 1, The one or more non-transitory computer-readable media of claim 9, The method of claim 17, respectively. Lin teaches wherein computing the reward comprises: determining a sum of reciprocal rank(s) of positive content item(s) in the top number of content items having the highest dot products ([2.1 Training Dense Retrieval Models] Given a query q, our task is to retrieve a list of documents to maximize some ranking metrics such as nDCG or MRR.Dense retrieval (DR) based on pre-trained transformers (Devlin et al., 2018; Raf- fel et al., 2020) encodes queries and documents as low dimensional vectors with a bi-encoder architec ture and uses the dot product between the encoded vectors as the similarity score: s(q, d) eq[CLS] · ed[CLS] , (1) where eq[CLS] and ed[CLS] are the [CLS] vectors at the last layer of BERT (Devlin et al., 2018).; [A.2 An Intuition Behind Uniform and Progressive Supervisions] For example, at the 3rd iteration of progressive training, given a query, a positive passage is labeled positive by all the three teachers, the probability of the passages being sam pled is 1 3 · (1 k + 1 k + 1 k) = 1 k. In our experiments, each teacher labels the top 10 (k = 10) retrieved passages as positives in our labeling scheme. Note that, in the case where multiple positives have equal probability, we further rank them according to their sum of reciprocal rank. For instance, if the two passages (e.g, p1 and p2) are retrieved by all the three teachers; then, we further rank them according to their scores 1 r11 + 1 r12 + 1 r13 and 1 r21 + 1 r22 + 1 r23 , where rmn denotes the rank of the passage pm by the n-th teacher. In addition, we also estimate the diversity of supervision by computing the number of positive passages in union sets from the sources of supervisions.). Bodigutla, Kabra, Zhao, Reddy, and Lin are considered to be analogous to the claimed invention because they are in the same field of machine learning. In view of the teachings of Bodigutla, Kabra, Zhao, and Reddy, it would have been obvious for a person of ordinary skill in the art to apply the teachings of Lin to Bodigutla before the effective filing date of the claimed invention in order to progressively train a generalizable dense retrieval (DR) (cf. Lin, [Abstract] We hence propose a new DA approach with diverse queries and sources of supervision to progressively train a generalizable DR. As a result, DRAGON, 1 our Dense Retriever trained with diverse AuGmentatiON, is the first BERT base-sized DR to achieve state-of-the-art effectiveness in both supervised and zero-shot evaluations and even competes with models using more complex late interaction (ColBERTv2 and SPLADE++).). Claims 7, 15, 19 are rejected under 35 U.S.C. 103 as being unpatentable over Bodigutla, in view of Kabra, Zhao, Reddy, and further in view of Roy et al. (NPL: "Latent Factor Representations for Cold-Start Video Recommendation", hereinafter 'Roy'). Regarding claim 7 and analogous claims 15, 19, Bodigutla, as modified by Kabra, Zhao, and Reddy, teaches The method of claim 1, The one or more non-transitory computer-readable media of claim 9, The method of claim 17, respectively. Roy teaches wherein computing the reward comprises: determining a key reciprocal rank of a top positive content item having a highest dot product in the top number of content items having the highest dot products PNG media_image1.png 170 368 media_image1.png Greyscale PNG media_image2.png 101 368 media_image2.png Greyscale . Bodigutla, Kabra, Zhao, Reddy, and Roy are considered to be analogous to the claimed invention because they are in the same field of machine learning. In view of the teachings of Bodigutla, Kabra, Zhao, and Reddy, it would have been obvious for a person of ordinary skill in the art to apply the teachings of Roy to Bodigutla before the effective filing date of the claimed invention in order to learn latent factor representations for cold start videos based on implicit feedback, improve MRR performance (cf. Roy, [ABSTRACT] In this work we present a novel method for learning latent factor representation for videos based on modelling the emotional connection between user and item. First of all we present a comparative analysis of state-of-the art emotion modelling approaches that brings out a surprising finding regarding the efficacy of latent factor representations in modelling emotion in video content. Based on this finding we present a method visual-CLiMF for learning latent factor representations for cold start videos based on implicit feedback. Visual-CLiMF is based on the popular collaborative less-is-more approach but demonstrates how emotional aspects of items could be used as auxiliary information to improve MRR performance. Experiments on a new data set and the Amazon products data set demonstrate the effectiveness of visual-CLiMF which outperforms existing CF methods with or without content information.). Claims 8, 16, 20 are rejected under 35 U.S.C. 103 as being unpatentable over Bodigutla, in view of Kabra, Zhao, Reddy, and further in view of Chaudhuri et al. (NPL: "Online Learning to Rank with Top-k Feedback", hereinafter 'Chaudhuri'). Regarding claim 8 and analogous claims 16, 20, Bodigutla, as modified by Kabra, Zhao, and Reddy, teaches The method of claim 1, The one or more non-transitory computer-readable media of claim 9, The method of claim 17, respectively. Chaudhuri teaches wherein computing the reward comprises: subtracting a maximum possible reward value of the top number of content items having the highest dot products by a reward value of the top number of content items having the highest dot products ([1.2 Contextual Setting, pg. 4-5] Our contributions: In this setting, first, we propose a general, efficient algorithm for online learning to rank with top k feedback and show that it works in conjunction with a number of ranking surrogates. We characterize the minimum feedback required, i.e., the value of k, for the algorithm to work with a particular surrogate by formally relating the feedback mechanism with the structure of the surrogates. We then apply our general techniques to three convex ranking surrogates and one non-convex surrogate. The convex surrogates considered are from three major learning to ranking methods: squared loss from a pointwise method (Cossock and Zhang, 2008), hinge loss used in the pairwise RankSVM (Joachims, 2002) method, and (modified) cross-entropy surrogate used in the listwise ListNet (Cao et al., 2007) method.; [2.1 Notation and Preliminaries] The objective of the learner is to minimize the expected regret with respect to best permutation in hindsight: T Eσ1,...,σT t=1 T RL(σt,Rt) −min σ t=1 RL(σ,Rt). (1) When RL is a gain, not loss, we need to negate the quantity above. The worst-case regret of a learner strategy is its maximal regret over all possible choices of R1,...,RT. The minimax regret is the minimal worst-case regret over all learner strategies.). Bodigutla, Kabra, Zhao, Reddy, and Chaudhuri are considered to be analogous to the claimed invention because they are in the same field of machine learning. In view of the teachings of Bodigutla, Kabra, Zhao, and Reddy, it would have been obvious for a person of ordinary skill in the art to apply the teachings of Chaudhuri to Bodigutla before the effective filing date of the claimed invention in order to provide efficient ranking strategies (cf. Chaudhuri, [Abstract] We consider two settings of online learning to rank where feedback is restricted to top ranked items. The problem is cast as an online game between a learner and sequence of users, over T rounds. In both settings, the learners objective is to present ranked list of items to the users. The learner’s performance is judged on the entire ranked list and true relevances of the items. However, the learner receives highly restricted feedback at end of each round, in form of relevances of only the top k ranked items, where k m. The first setting is non-contextual, where the list of items to be ranked is fixed. The second setting is contextual, where lists of items vary, in form of traditional query-document lists. No stochastic assumption is made on the generation process of relevances of items and contexts. We provide efficient ranking strategies for both the settings. The strategies achieve O(T2/3) regret, where regret is based on popular ranking measures in first setting and ranking surrogates in second setting. We also provide impossibility results for certain ranking measures and a certain class of surrogates, when feedback is restricted to the top ranked item, i.e. k = 1. We empirically demonstrate the performance of our algorithms on simulated and real world data sets.). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to MAGGIE MAIDO whose telephone number is (703) 756-1953. The examiner can normally be reached M-Th: 6am - 4pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michael Huntley can be reached on (303) 297-4307. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MM/Examiner, Art Unit 2129 /MICHAEL J HUNTLEY/Supervisory Patent Examiner, Art Unit 2129
Read full office action

Prosecution Timeline

Jan 26, 2024
Application Filed
Aug 13, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737610
ARCHITECTURE FOR UTILIZING KEY-VALUE STORE FOR DISTRIBUTED NEURAL NETWORKS AND DEEP LEARNING
5y 5m to grant Granted Sep 15, 2026
Patent 12725076
ARTIFICIAL INTELLIGENCE MODEL LEARNING INTROSPECTION
4y 10m to grant Granted Sep 01, 2026
Patent 12651159
COMPUTER-IMPLEMENTED METHOD FOR ACCELERATING CONVERGENCE IN THE TRAINING OF GENERATIVE ADVERSARIAL NETWORKS (GAN) TO GENERATE SYNTHETIC NETWORK TRAFFIC, AND COMPUTER PROGRAMS OF SAME
3y 11m to grant Granted Jun 09, 2026
Patent 12651162
METHOD FOR TRAINING FEATURE QUANTIZATION MODEL, FEATURE QUANTIZATION METHOD, DATA QUERY METHODS AND ELECTRONIC DEVICES
3y 9m to grant Granted Jun 09, 2026
Patent 12645922
Machine Learning Systems and Methods for Classification Based Auto-Annotation
4y 11m to grant Granted Jun 02, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
67%
Grant Probability
96%
With Interview (+29.1%)
4y 0m (~1y 4m remaining)
Median Time to Grant
Low
PTA Risk
Based on 55 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month