Prosecution Insights
Last updated: August 17, 2026
Application No. 18/423,825

RETRIEVAL OPTIMIZATION USING REINFORCEMENT LEARNING

Non-Final OA §101§103
Filed
Jan 26, 2024
Priority
Sep 21, 2023 — provisional 63/584,355
Examiner
SHAH, SAYED MUNEER
Art Unit
Tech Center
Assignee
Roku Inc.
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
7 currently pending
Career history
5
Total Applications
across all art units

Statute-Specific Performance

§101
26.7%
-13.3% vs TC avg
§103
43.3%
+3.3% vs TC avg
§102
23.3%
-16.7% vs TC avg
§112
6.7%
-33.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 0 resolved cases

Office Action

§101 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This office action is in response to submission of application on 01/26/2024. Claims 1-20 are presented for examination. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1: Is the claim to a process, machine, manufacture or composition of matter? Claims 1-9 and 19-20 are directed to a method (i.e., a process); and claims 10-18 are directed to an article of manufacture (i.e., a product); therefore, all pending claims are directed to one of the four categories of invention. Independent Claims Step 2A Prong One: Does the claim recite an abstract idea, law of nature, or natural phenomenon? Yes, independent claim 1 recites an abstract idea in the form of mental processes. A mental process is a process that “can be performed in the human mind, or by a human using a pen and paper” (MPEP§ 2106.04(a)(2)(III), paragraph 1). Examples of mental processes include “observations, evaluations, judgments, and opinions” (MPEP § 2106.04(a)(2)(III), paragraph 2). The following limitations of claim 1 are mental processes: randomly sample, from content items corresponding to the query, a number of sampled content items; [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for randomly sample is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.] determining semantic relevance scores corresponding to the sampled content items, wherein a semantic relevance score measures semantic affinity of a content item to the query; [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for semantic relevance scores is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.] determining bucket identifiers corresponding to the sampled content items; [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for determining bucket identifiers is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.] for each sampled content item, scaling the semantic relevance score based on the bucket identifier corresponding to the sampled content item and a weight in the action vector corresponding to the bucket identifier; [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for scaling the semantic relevance score is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.] sorting the sampled content items based on scaled semantic relevance scores; [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for sorting is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.] determining a top number of content items having scaled semantic relevance scores; [This is a mental process that can be performed by observations, evaluations, judgments, and opinions.] computing a reward based on reward values corresponding to a top number of sampled content items having highest scaled semantic relevance scores; and [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for computing a reward is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.] Therefore, the independent claims recite a judicial exception. Step 2A Prong Two: Does the claim recite additional elements that integrate the judicial exception into a practical application? No. The judicial exception recited in the above discussed claims is not integrated into a practical application. obtaining a query; [receiving a query is sending data, which is insignificant, extra-solution activity. See MPEP 2106.05(g). Transmitting data is well-understood, routine, and conventional. See MPEP 2106.05(d)(II)(i).] determining, using parameters of an agent model and an embedding of the query, an action vector comprising weights corresponding to different bucket identifiers; [Determining an action vector are mere instructions to apply the abstract idea. Mere recitation that a judicial exception is to be performed using generic class of computer algorithms in their ordinary capacity, cannot meaningfully integrate the judicial exception into a practical application. See MPEP 2106.05(f). Additionally, this is a description of how the abstract idea is performed, using parameters of an agent model and an embedding of the query. As such, this merely describes a technological environment. See MPEP 2106.05(h).] updating the parameters of the agent model based on the query, the action vector, and the reward. [This additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)). This element merely indicates a field of use or technological environment in which a judicial exception is applied, namely updating parameters of the agent model]. Therefore, under MPEP 2106.04(d), the additional elements of the claims do not integrate the judicial exception into a practical application. Step 2B: Does the claim recite additional elements that amount to significantly more than the judicial exception? No. The claims do not include additional elements that are sufficient for the claims to amount to significantly more than the judicial exception. Additional elements that are mere instructions to apply an exception or merely generally linking or generally linking the use of a judicial exception to a particular technological environment or field of use do not constitute significantly more than a judicial exception under MPEP§2106.05(I)(A). Since the additional elements in the independent claims are all are mere instructions to apply an exception or are merely generally linking or generally linking the use of a judicial exception to a particular technological environment or field of use, they do not constitute significantly more than a judicial exception. Therefore, the additional elements identified in the Step 2A Prong Two analysis do not constitute significantly more than a judicial exception. Independent claim 10 recites the same relevant limitations and a similar analysis applies. Claim 10 recites the additional limitations of " One or more non-transitory computer-readable media having instructions stored thereon, when the instructions are executed by one or more processors, cause the one or more processors to:" – [A non-transitory machine-readable medium are components recited at a high level are construed as generic computer components used to implement the abstract idea. See MPEP 2106.05(f)(2). As such, the limitations do not integrate the abstract idea into a practical application. Nor to do they amount to significantly more.] Claim 19 Step 2A Prong One: Does the claim recite an abstract idea, law of nature, or natural phenomenon? Yes, independent claim 1 recites an abstract idea in the form of mental processes. A mental process is a process that “can be performed in the human mind, or by a human using a pen and paper” (MPEP§ 2106.04(a)(2)(III), paragraph 1). Examples of mental processes include “observations, evaluations, judgments, and opinions” (MPEP § 2106.04(a)(2)(III), paragraph 2). The following limitations of claim 1 are mental processes: randomly sample, from content items corresponding to the context, a number of sampled content items; [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for randomly sample is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.] determining contextual relevance scores corresponding to the sampled content items, wherein a contextual relevance score measures contextual affinity of a content item to the context, [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for contextual relevance scores is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.] determining bucket identifiers corresponding to the sampled content items; [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for determining bucket identifiers is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.] for each sampled content item, scaling the contextual relevance score based on the bucket identifier corresponding to the sampled content item and a weight in the action vector corresponding to the bucket identifier; [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for scaling the contextual relevance score is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.] sorting the sampled content items based on scaled contextual relevance scores; [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for sorting is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.] determining a top number of content items having scaled contextual relevance scores; [This is a mental process that can be performed by observations, evaluations, judgments, and opinions.] computing a reward based on reward values corresponding to a top number of sampled content items having highest scaled contextual relevance scores; and [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for computing a reward is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.] Therefore, the independent claims recite a judicial exception. Step 2A Prong Two: Does the claim recite additional elements that integrate the judicial exception into a practical application? No. The judicial exception recited in the above discussed claims is not integrated into a practical application. obtaining a context; [receiving a context is sending data, which is insignificant, extra-solution activity. See MPEP 2106.05(g). Transmitting data is well-understood, routine, and conventional. See MPEP 2106.05(d)(II)(i).] determining, using parameters of an agent model and an embedding of the context, an action vector comprising weights corresponding to different bucket identifiers; [Determining an action vector are mere instructions to apply the abstract idea. Mere recitation that a judicial exception is to be performed using generic class of computer algorithms in their ordinary capacity, cannot meaningfully integrate the judicial exception into a practical application. See MPEP 2106.05(f). Additionally, this is a description of how the abstract idea is performed, using parameters of an agent model and an embedding of the context. As such, this merely describes a technological environment. See MPEP 2106.05(h).] updating the parameters of the agent model based on the context, the action vector, and the reward. [This additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)). This element merely indicates a field of use or technological environment in which a judicial exception is applied, namely updating parameters of the agent model]. Therefore, under MPEP 2106.04(d), the additional elements of the claims do not integrate the judicial exception into a practical application. Step 2B: Does the claim recite additional elements that amount to significantly more than the judicial exception? No. The claims do not include additional elements that are sufficient for the claims to amount to significantly more than the judicial exception. Additional elements that are mere instructions to apply an exception or merely generally linking or generally linking the use of a judicial exception to a particular technological environment or field of use do not constitute significantly more than a judicial exception under MPEP§2106.05(I)(A). Since the additional elements in the independent claims are all are mere instructions to apply an exception or are merely generally linking or generally linking the use of a judicial exception to a particular technological environment or field of use, they do not constitute significantly more than a judicial exception. Therefore, the additional elements identified in the Step 2A Prong Two analysis do not constitute significantly more than a judicial exception. Therefore, the independent claims are not patent eligible. Dependent Claims The remaining dependent claims being rejected do not recite additional elements, whether considered individually or in combination, that are sufficient to integrate the judicial exception into a practical application or amount to significantly more than the judicial exception. Claims 2 and 11 wherein determining the semantic relevance scores comprises: determining a first feature vector representing the query; [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for determining a feature vector is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.] determining second feature vectors representing metadata of the sampled content items respectively; and [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for determining a vector is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.] determining a dot product of the first feature vector and each one of the second feature vectors, wherein the semantic relevance scores are based on the dot products. [This is a mental process that can be performed by observations, evaluations, judgments, and opinions] Claims 3 and 12 updating a trust parameter of an episode based on the reward and a function; and [Updating a trust parameter are mere instructions to apply the abstract idea. Mere recitation that a judicial exception is to be performed using generic class of computer algorithms in their ordinary capacity, cannot meaningfully integrate the judicial exception into a practical application. See MPEP 2106.05(f)] in response to the trust parameter meeting a criterion and a number of rounds completed in the episode has not reached a maximum number, obtaining a further query to complete a further round of the episode. [receiving a query is sending data, which is insignificant, extra-solution activity. See MPEP 2106.05(g). Transmitting data is well-understood, routine, and conventional. See MPEP 2106.05(d)(II)(i).] Claims 4 and 13 in response to the trust parameter not meeting a criterion and the number of rounds completed in an episode has not reached the maximum number, ending the episode. [Ending the episode are mere instructions to apply the abstract idea. Mere recitation that a judicial exception is to be performed using generic class of computer algorithms in their ordinary capacity, cannot meaningfully integrate the judicial exception into a practical application. See MPEP 2106.05(f)] Claims 5 and 14 in response to the number of rounds completed in an episode has reached the maximum number, ending the episode. [Ending the episode are mere instructions to apply the abstract idea. Mere recitation that a judicial exception is to be performed using generic class of computer algorithms in their ordinary capacity, cannot meaningfully integrate the judicial exception into a practical application. See MPEP 2106.05(f)] Claims 6 and 15 updating the parameters of the agent model is based on an updated value of the trust parameter of the episode. [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for updating a trust parameter is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.] Claims 7 and 16 summing the reward values corresponding to a top number of sampled content items having highest scaled semantic relevance scores. [This is a mental process that can be performed by observations, evaluations, judgments, and opinions.] Claims 8 and 17 determining a binary flag based on whether a sum of the reward values corresponding to a top number of sampled content items having highest scaled semantic relevance scores is positive or negative. [This is a mental process that can be performed by observations, evaluations, judgments, and opinions.] Claims 9 and 18 updating the parameters of the agent model comprises calculating a long-term reward based on a weighted sum of the reward and one or more rewards of future rounds in an episode. [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for calculating a reward is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.] Claims 20 wherein determining the contextual relevance scores comprises: determining a first feature vector representing the context; [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for determining a feature vector is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.] determining second feature vectors representing metadata of the sampled content items respectively; and [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for determining a vector is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.] determining a dot product of the first feature vector and each one of the second feature vectors, wherein the contextual relevance scores are based on the dot products. [This is a mental process that can be performed by observations, evaluations, judgments, and opinions] The prior art used for rejections are provided below: 1. US11544553Bl (January 3, 2023) to He et al. (hereinafter He) 2. US10713263B2 (July 14, 2020) to Lewis et al. (hereinafter Lewis), 3. Reinforcement Learning to Rank in E-Commerce Search Engine: Formalization, Analysis, and Application (May 23, 2018) to Hu et al. (hereinafter Hu), 4. US11709873B2 (July 25, 2023) to Xiao et al. (hereinafter Xiao). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 1- 20 are rejected under 35 USC 103 as being unpatentable over He in view of Lewis, in further view of Hu and Xiao. Per claim 1, He discloses A method, comprising [He, column 2, line 21 “Disclosed are systems, methods, and non-transitory computer-readable media for…”]: obtaining a query [He, column 2, line 52 “Given the query, the classifier is trained to determine the predicted relevance of the retrieved document.”. (note: this shows obtaining a query as the initiating input)]; randomly sample, from content items corresponding to the query, a number of sampled content items [He, column 2, line 32 “When training the classifier module, the reinforcement module gives every candidate a relevance score and sample lower scored documents as irrelevant samples (negative samples)”. (note: this shows sampling from candidate documents (content items) corresponding to a query); computing a reward based on reward values corresponding to a top number of sampled content items having highest scaled semantic relevance scores [He, column 6, line 67 “The classifier 240 then provides (430) a reward signal to the policy gradient function 210 for each data (e.g., document)/query pair”. (note: this shows computing a reward signal based on retrieved documents (content items) corresponding to a query); column 5, line 51 “The reward is a combined reward defined in Equation (4). The goal is to improve the reinforcement ranker's 220 performance by maximizing the objective:”. (note: this shows computing a combined reward corresponding to the quality of the ranked list (top content items))]; and He does not expressly disclose, but He combined with Xiao does teach: determining semantic relevance scores corresponding to the sampled content items, wherein a semantic relevance score measures semantic affinity of a content item to the query [Xiao, column 18, line 42 “the one or more answer retriever engines 320 can determine a dot product (also referred to as cosine similarity) of the query embedding vector (of the input query 304) and the question embedding vectors (of the questions from the QA pairs and/or {Q}A pairs) to determine a semantic similarity between the question defined by the input query 304 and the text of the questions in the QA space and the {Q}A space.”. (note: this directly states semantic similarity, which is semantic affinity); column 18, line 61 “a value close to O indicates independence or orthogonality between two vectors, a value close to 1 indicates a highly positive similarity (e.g., an exact or very close match) between two vectors, and a value close to -1 indicates a highly negative similarity (e.g. opposite semantics).”. (note: this shows that the score is a range measure of semantic affinity))]; determining bucket identifiers corresponding to the sampled content items [Lewis, column 12, line 12 “Token manager 204 can locate the corresponding entry for the content item 140, determine the associated bucketing tokens 244, 246, 248 and provide the tokens to ranking service 120". (note: this shows determining bucket identifiers (bucketing tokens) corresponding to content items); Abstract “…each bucketing token comprises a unique identifier that identifies a plurality of content items as being associated with a group of users of the content sharing platform that have similar interests.”. (note: this shows that each bucket identifier (bucketing token) is a unique identifier associated with content items)]; He and Xiao are analogous art because they are from the same field of endeavor of computing a semantic relevance score. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to apply the bucketing token identification to the ranked content items as taught by He, using the dot-product formulation and embedding approach as taught by Xiao. The suggestion/motivation for doing so is Xiao’s dot-product embedding approach effectively captures semantic meaning, it produces a concrete, computationally efficient measure of semantic affinity [Xiao, column 18, line 61 “…a value close to O indicates independence or orthogonality between two vectors, a value close to 1 indicates a highly positive similarity (e.g., an exact or very close match) between two vectors, and a value close to -1 indicates a highly negative similarity (e.g. opposite semantics).”]. He does not expressly disclose, but He combined with Xiao and Hu does teach: determining, using parameters of an agent model and an embedding of the query, an action vector comprising weights corresponding to different bucket identifiers [He, column 5, line 48 “…the policy gradient function 210 trains the ranker 220. The goal of the ranker 220 is to maximize the future expected cumulative rewards…where τ is a trajectory of state-action sequence s0, a0,…, sT, aT sampled by the reinforcement ranker 20, and P(τ; θ) is the probability of trajectory T under policy πθ”. (note: this shows an agent model (reinforcement ranker) with parameters θ that uses query-based state information to determine action probabilities. The πθ defines an action distribution conditioned on the current sate embedding.); Hu, pg. 7, 6.1 Simulation “…a ranking action of the search engine is a n-dim weight vector μ = (μ1, ..., μn)⊤”. (note: this shows that the action of the agent model is a weight vector (action vector comprising weights). This weight vector is computed by the policy parametrized by θ using the current state, which includes query features.); pg. 9, 6.2 Application “The state of the environment is represented by a 90-dim feature vector, which contains the item page features, user features and query features of the current search session.”. (note: this shows that the state embedding used by the agent to determine its action vector includes query features (embedding of the query)); Lewis, column 12, line 22 “the ranking algorithm can determine the user buckets associated with a user of a given feed and weight content items from those user buckets more highly in the content ranking 160 for that feed.”. (note: this shows that the ranking algorithm uses bucket identifiers with associated weights for different buckets)]; sorting the sampled content items based on scaled semantic relevance scores [He, column 7, line 5 “The ranker 220 can then rank a set of data 230 and enable retrieval of top-ranked data, e.g., 140a, thereby enabling more efficient retrieval of desired data”. (note: this shows sorting (ranking) content items based on computed scores to produce a ranked list); Hu, pg. 1, 1 Introduction “the search engine ranks the items related to the query and displays the top K items (e.g., K = 10) in a page…”. (note: this shows sorting (ranking) items by their computed scores)]; determining a top number of content items having scaled semantic relevance scores [Hu, pg. 1, 1 Introduction “the search engine ranks the items related to the query and displays the top K items (e.g., K = 10) in a page…”. (note: this shows determining a top-K number of content items from ranked results); updating the parameters of the agent model based on the query, the action vector, and the reward [He, column 6, line 37 “…the policy gradient algorithm is then used. The reward signal given by the classifier ϕ* is passed to this reinforcement learning environment. Finally, the gradient of ϕ*is calculated and transferred to the ranker so that the ranker can do one step of parameter update.”. (note: this shows updating the parameters of the agent model (ranker) based on the reward signal. The policy of πθ is conditioned on query derived state and the gradient update uses the reward.); Hu, pg. 7, 5.1 The DPG-FBE Algorithm “…the parameters θ and w will be updated after any search session between the search engine agent and users.”. (note: this shows updating agent model parameters (θ) based on the reward and action taken during a search session initiated by a query)]. He, Xiao, and Hu are analogous art because they are from the same field of endeavor of computing a semantic relevance score. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to apply the bucketing token identification to the ranked content items as taught by He, using the dot-product formulation taught by Hu, and embedding approach as taught by Xiao. The suggestion/motivation for doing so is Hu’s dot-product scoring offers computational efficiency and stronger performance [Hu, pg. 7, 6.1 Simulation “The ranking score of the item x under the ranking action µ is the inner product xTµ of the two vectors”]. Xiao’s dot-product embedding approach effectively captures semantic meaning, it produces a concrete, computationally efficient measure of semantic affinity [Xiao, column 18, line 61 “…a value close to O indicates independence or orthogonality between two vectors, a value close to 1 indicates a highly positive similarity (e.g., an exact or very close match) between two vectors, and a value close to -1 indicates a highly negative similarity (e.g. opposite semantics).”]. He does not expressly disclose, but He combined with Xiao, Hu, and Lewis does teach: for each sampled content item, scaling the semantic relevance score based on the bucket identifier corresponding to the sampled content item and a weight in the action vector corresponding to the bucket identifier [Lewis, Abstract “…the processing device improves a ranking score of each content item from the set of content items that has at least one bucketing token matching a bucketing token associated with the user of the social network platform,”. (note: this shows scaling (improving) the ranking score of each content item based on its bucket identifier (bucketing token)); column 12, line 17 “…content ranking module 124 takes the received bucketing tokens 244, 246, 248 and applies the tokens as inputs to a ranking algorithm. At block 585, method 500 adjusts a ranking of the content item in view of the output of the ranking algorithm.”. (note: this shows that the ranking algorithm receives bucket tokens for a content item and adjusts (scales) the ranking score. This is scaling the relevance score based on bucket identifiers and their corresponding weights.)]; He, Lewis, Hu, and Xiao are analogous art because they are from the same field of endeavor of computing a semantic relevance score. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to apply the bucketing token identification as taught by Lewis to the ranked content items as taught by He, using the dot-product formulation taught by Hu, and embedding approach as taught by Xiao. The suggestion/motivation for doing so is explicitly stated by Lewis as improved ranking quality [Lewis, column 3, line 46 “the use of user bucketing tokens can generate more watchtime and referrers, and for the feed ranking services, the bucketing tokens can deliver more time on site and an improved user experience.”]. Hu’s dot-product scoring offers computational efficiency and stronger performance [Hu, pg. 7, 6.1 Simulation “The ranking score of the item x under the ranking action µ is the inner product xTµ of the two vectors”]. Xiao’s dot-product embedding approach effectively captures semantic meaning, it produces a concrete, computationally efficient measure of semantic affinity [Xiao, column 18, line 61 “…a value close to O indicates independence or orthogonality between two vectors, a value close to 1 indicates a highly positive similarity (e.g., an exact or very close match) between two vectors, and a value close to -1 indicates a highly negative similarity (e.g. opposite semantics).”]. Per Claim 2, He-Lewis-Hu-Xiao discloses claim 1. He does not fully disclose, but with Xiao does teach determining the semantic relevance scores comprises: determining a first feature vector representing the query [Xiao, column 18, line 15 “…the one or more answer retriever engines 320 can generate a vector representation (referred to as an embedding vector or an embedding) for the input query 304”. (note: this shows generating a first feature vector representing the query)]; determining second feature vectors representing metadata of the sampled content items respectively [Xiao, column 18, line 18, “…one or more vector representations (or embedding vectors) for the text passages contained in the QA space and the { Q} A space. In some cases, the embedding vectors can be generated using a neural network encoder, which can convert text to a mathematical vector representation of that text”. (note: this shows generating second feature vectors representing the text (metadata) of each content item respectively)]; and determining a dot product of the first feature vector and each one of the second feature vectors, wherein the semantic relevance scores are based on the dot products [Xiao, column 18, line 42 “the one or more answer retriever engines 320 can determine a dot product (also referred to as cosine similarity) of the query embedding vector (of the input query 304) and the question embedding vectors (of the questions from the QA pairs and/or {Q}A pairs) to determine a semantic similarity…”. (note: this shows computing the dot product of the first feature vector and each second feature vector, with the result being the relevance scores)]. He, Lewis, Hu, and Xiao are analogous art because they are from the same field of endeavor of computing a semantic relevance score. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to apply the bucketing token identification as taught by Lewis to the ranked content items as taught by He, using the dot-product formulation taught by Hu, and embedding approach as taught by Xiao. The suggestion/motivation for doing so is Xiao’s dot-product embedding approach effectively captures semantic meaning, it produces a concrete, computationally efficient measure of semantic affinity [Xiao, column 18, line 61 “…a value close to O indicates independence or orthogonality between two vectors, a value close to 1 indicates a highly positive similarity (e.g., an exact or very close match) between two vectors, and a value close to -1 indicates a highly negative similarity (e.g. opposite semantics).”]. Per Claim 3, He-Lewis-Hu-Xiao discloses claim 1. He does not fully disclose, but with Hu does teach updating a trust parameter of an episode based on the reward and a function [Hu, pg. 7, 5.1 The DPG-FBE Algorithm, Algorithm 1 “ PNG media_image1.png 189 694 media_image1.png Greyscale ”. (note: c is the continuing probability model, the trust parameter. This shows an update step in which the model c is updated using a data sample that is derived directly from the reward rk received at that round. The sample (hk+1, rk) incorporates the reward value rk into the update of these models); pg. 7, 5.1 The DPG-FBE Algorithm, Algorithm 1 “…pretrained conversion probability model b, continuing probability model c, and expected deal price model m of item page histories”. (note: this identifies c as the continuing probability model, the trust parameter that determines whether the episode continues); pg. 7, 5.1 The DPG-FBE Algorithm “These models can be trained using online or offline data by any possible statistical learning method.”. (note: the continuing probability model c is updated using any possible statistical learning method, which is a function. This is applied to reward derived samples as shown in the previous citations. These show that the trust parameter (continuing probability model c) is updated using a function (a statistical learning method) applied to samples derived from the reward.)]; and in response to the trust parameter meeting a criterion [Hu, pg. 4, 3.2 Search Session MDP “ S = H C ​ ∪ H B ​ ∪ H L ​   i s   t h e   s t a t e   s p a c e ,   H C = { C ( h t ) ∣ ∀ h t ∈ H t , 0 ≤ t < T }   i s   t h e   n o n t e r m i n a l   s t a t e   s e t   t h a t   c o n t a i n s   a l l   c o n t i n u a t i o n   e v e n t s ,   H B = { B ( h t ) ∣ ∀ h t ∈ H t , 0 < t ≤ T }   a n d   H L = { L ( h t ) ∣ ∀ h t ∈ H t , 0 < t ≤ T }   a r e   t w o   t e r m i n a l   s t a t e   s e t s   w h i c h   c o n t a i n   a l l   c o n v e r s i o n   e v e n t s   a n d   a l l   a b a n d o n   e v e n t s ,     r e s p e c t i v e l y . ”. (note: this shows a continuation criterion (C(ht)) that determines whether the episode continues to the next round, which is a trust parameter meeting a criterion. The continuation event C(ht) is updated based on the user’s response (reward signal) after each round); , pg. 3, 3.1 Search Session Modeling “terminating the search session by purchasing an item in ht with probability b(ht); (2) leaving the search session from ht with probability l(ht); (3) continuing the search session from ht with probability (1 - b(ht) - l(ht))”. (note: this shows that the session (episode) continues to a further round when the continuation probability (trust parameter) meets the criterion of being non-zero, when the user continues searching. This is the trust parameter meeting a criterion to trigger a further round.)] and a number of rounds completed in the episode has not reached a maximum number [Hu, pg. 4, 3.2 Search Session MDP “ T =   | D | K is the maximal decision step of a search session”. (note: this shows a maximum number of rounds (T = maximal decision step) for each episode (search session), which is a number of rounds completed in the episode has not reached a maximum number)], obtaining a further query to complete a further round of the episode [Hu, pg. 1, 1 Introduction “when a new page is requested, the search engine re-ranks the rest of the items and display the top K items.”. (note: this shows obtaining a further query (new page request) to complete a further round of the episode)]. He and Hu are analogous art because they are from the same field of endeavor of content ranking and recommendation systems. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to apply the bucketing token identification to the ranked content items as taught by He, using the dot-product formulation taught by Hu. The suggestion/motivation for doing so is Hu’s dot-product scoring offers computational efficiency and stronger performance [Hu, pg. 7, 6.1 Simulation “The ranking score of the item x under the ranking action µ is the inner product xTµ of the two vectors”]. Per Claim 4, He-Lewis-Hu-Xiao discloses claim 3. He does not fully disclose, but with Hu does teach in response to the trust parameter not meeting a criterion and the number of rounds completed in an episode has not reached the maximum number, ending the episode [Hu, pg. 7, 5.1 The DPG-FBE Algorithm, Algorithm 1 “ PNG media_image1.png 189 694 media_image1.png Greyscale ”. (note: line 12 in the algorithm, in this branch the trust parameter ( c ) is updated toward a value reflecting non-continuation in that instance.)]. He and Hu are analogous art because they are from the same field of endeavor of content ranking and recommendation systems. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to apply the bucketing token identification to the ranked content items as taught by He, using the dot-product formulation taught by Hu. The suggestion/motivation for doing so is Hu’s dot-product scoring offers computational efficiency and stronger performance [Hu, pg. 7, 6.1 Simulation “The ranking score of the item x under the ranking action µ is the inner product xTµ of the two vectors”]. Per Claim 5, He-Lewis-Hu-Xiao discloses claim 3. He does not fully disclose, but with Hu does teach in response to the number of rounds completed in an episode has reached the maximum number, ending the episode [Hu, pg. 4, 3.2 Search Session MDP “ T =   | D | K is the maximal decision step of a search session”. (note: this shows that the episode (search session) ends when the maximal decision step T is reached, which is ending the episode when the maximum number of rounds is reached); pg. 3, 3.1 Search Session Modeling “Since the item set D is finite, there are at most   | D | K item pages, and correspondingly at most | D | K decision steps in a search session.”. (note: this shows a finite maximum on the number of rounds, after which the session (episode) terminates)]. He and Hu are analogous art because they are from the same field of endeavor of content ranking and recommendation systems. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to apply the bucketing token identification to the ranked content items as taught by He, using the dot-product formulation taught by Hu. The suggestion/motivation for doing so is Hu’s dot-product scoring offers computational efficiency and stronger performance [Hu, pg. 7, 6.1 Simulation “The ranking score of the item x under the ranking action µ is the inner product xTµ of the two vectors”]. Per Claim 6, He-Lewis-Hu-Xiao discloses claim 3. He does not fully disclose, but with Hu does teach the reward used in updating the parameters of the agent model is based on an updated value of the trust parameter of the episode [Hu, pg. 7, 5.1 The DPG-FBE Algorithm, Algorithm 1 ” PNG media_image2.png 127 594 media_image2.png Greyscale ”. (note: this shows that the reward derived term Sk is computed using c(hk+1), which is the continuing probability (trust parameter) value after its update in that round. )]. He and Hu are analogous art because they are from the same field of endeavor of content ranking and recommendation systems. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to apply the bucketing token identification to the ranked content items as taught by He, using the dot-product formulation taught by Hu. The suggestion/motivation for doing so is Hu’s dot-product scoring offers computational efficiency and stronger performance [Hu, pg. 7, 6.1 Simulation “The ranking score of the item x under the ranking action µ is the inner product xTµ of the two vectors”]. Per Claim 7, He-Lewis-Hu-Xiao discloses claim 1. He does not fully disclose, but with Hu does teach summing the reward values corresponding to a top number of sampled content items having highest scaled semantic relevance scores [Hu, pg. 2, 2.1 Reinforcement Learning “The objective of an agent in an MDP is to find an optimal policy which maximizes the expected accumulative rewards starting from any state s (typically under the infinite-horizon discounted setting), which is defined by PNG media_image3.png 55 591 media_image3.png Greyscale ”. (note: this shows that the reward computation involves a summation (Σ) of reward values, the accumulative rewards corresponding to ranked items across the episode. The top K items displayed at each step are the items from which rewards are accumulated.); He, column 6, line 23 “ PNG media_image4.png 87 431 media_image4.png Greyscale ”. (note: this shows the reward Rt as a sum (Σ) of immediate rewards R(sk, ak) across the trajectory, which is summing reward values corresponding to content items)]. He and Hu are analogous art because they are from the same field of endeavor of content ranking and recommendation systems. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to apply the bucketing token identification to the ranked content items as taught by He, using the dot-product formulation taught by Hu. The suggestion/motivation for doing so is Hu’s dot-product scoring offers computational efficiency and stronger performance [Hu, pg. 7, 6.1 Simulation “The ranking score of the item x under the ranking action µ is the inner product xTµ of the two vectors”]. Per Claim 8, He-Lewis-Hu-Xiao discloses claim 1. He does not fully disclose, but with Hu does teach determining a binary flag based on whether a sum of the reward values corresponding to a top number of sampled content items having highest scaled semantic relevance scores is positive or negative [Hu, pg. 5, 4.2 Reward Function “The agent will recieve a positive reward from the environment only when its ranking action leads to a successful transation. In all other cases, the reward is zero”. (note: this shows a binary structure, the reward is positive (non-zero) or zero. A binary flag based on whether the summed rewards are positive or zero (negative/zero) is directly in this binary reward design.)]. He and Hu are analogous art because they are from the same field of endeavor of content ranking and recommendation systems. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to apply the bucketing token identification to the ranked content items as taught by He, using the dot-product formulation taught by Hu. The suggestion/motivation for doing so is Hu’s dot-product scoring offers computational efficiency and stronger performance [Hu, pg. 7, 6.1 Simulation “The ranking score of the item x under the ranking action µ is the inner product xTµ of the two vectors”]. Per Claim 9, He-Lewis-Hu-Xiao discloses claim 3. He does not fully disclose, but with Hu does teach calculating a long-term reward based on a weighted sum of the reward and one or more rewards of future rounds in an episode [Hu, pg. 2, 2.1 Reinforcement Learning “The objective of an agent in an MDP is to find an optimal policy which maximizes the expected accumulative rewards starting from any state s (typically under the infinite-horizon discounted setting), which is defined by PNG media_image3.png 55 591 media_image3.png Greyscale ”. (note: this shows calculating a long-term reward (value function V*) as a weighted sum (yk are the weights) of the reward at the current step and all future rewards (rt+k for k=1, 2, …) across future rounds in the session (episode)); pg. 6, 5 Algorithm “ PNG media_image5.png 98 590 media_image5.png Greyscale ”. (note: this shows the policy objective J(θ) as a sum of rewards across all T rounds of the episode, used to update the parameters θ)]. He and Hu are analogous art because they are from the same field of endeavor of content ranking and recommendation systems. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to apply the bucketing token identification to the ranked content items as taught by He, using the dot-product formulation taught by Hu. The suggestion/motivation for doing so is Hu’s dot-product scoring offers computational efficiency and stronger performance [Hu, pg. 7, 6.1 Simulation “The ranking score of the item x under the ranking action µ is the inner product xTµ of the two vectors”]. Claims 10-18 are directed towards the apparatus performed by the method of claims 1-9, respectively. Therefore, the rejections applied to claims 1-9 also apply to claims 10-18 [He, column 2, line 21 “Disclosed are systems, methods, and non-transitory computer-readable media for…”]. Per claim 19, He discloses A method, comprising [He, column 2, line 21 “Disclosed are systems, methods, and non-transitory computer-readable media for…”]: randomly sample, from content items corresponding to the context, a number of sampled content items [He, column 2, line 32 “When training the classifier module, the reinforcement module gives every candidate a relevance score and sample lower scored documents as irrelevant samples (negative samples)”. (note: this shows sampling from candidate documents (content items) corresponding to a context))]; computing a reward based on reward values corresponding to a top number of content items having highest scaled contextual relevance scores [He, column 6, line 67 “The classifier 240 then provides (430) a reward signal to the policy gradient function 210 for each data (e.g., document)/query pair”. (note: this shows computing a reward signal based on retrieved documents (content items) corresponding to a query); column 5, line 51 “The reward is a combined reward defined in Equation (4). The goal is to improve the reinforcement ranker's 220 performance by maximizing the objective:”. (note: this shows computing a combined reward corresponding to the quality of the ranked list (top content items))]; and He does not expressly disclose, but He combined with Hu does teach: obtaining a context [He, column 2, line 58 “The reinforcement ranker has input from the original feature space, and the other input is transferred from the classifier module.”. (note: this shows that the ranker operates on a broad feature space (context) that goes beyond the query text); column 2, line 52 “Given the query, the classifier is trained to determine the predicted relevance of the retrieved document.”. (note: this shows obtaining a query as the initiating input) Hu, pg. 3, 3.1 Search Session Modeling “For the initial decision step t = 0, the initial item page history h0 = q. For each later decision step t ≥ 1, the item page history up to t is ht = ht−1 ∪ {pt}, where ht−1 is the item page history up to the step (t − 1) and pt is the item page of step t.”. (note: this shows that the context for each ranking round is the item page history ht which is a broader representation that includes the query plus all prior interactions);]; sorting the sampled content items based on scaled contextual relevance scores [He, column 7, line 5 “The ranker 220 can then rank a set of data 230 and enable retrieval of top-ranked data, e.g., 140a, thereby enabling more efficient retrieval of desired data”. (note: this shows sorting (ranking) content items based on computed scores to produce a ranked list); Hu, pg. 1, 1 Introduction “the search engine ranks the items related to the query and displays the top K items (e.g., K = 10) in a page…”. (note: this shows sorting (ranking) items by their computed scores)]; determining a top number of content items having scaled contextual relevance scores [Hu, pg. 1, 1 Introduction “the search engine ranks the items related to the query and displays the top K items (e.g., K = 10) in a page…”. (note: this shows determining a top-K number of content items from ranked results)]; updating the parameters of the agent model based on the context, the action vector, and the reward [He, column 6, line 37 “…the policy gradient algorithm is then used. The reward signal given by the classifier ϕ* is passed to this reinforcement learning environment. Finally, the gradient of ϕ*is calculated and transferred to the ranker so that the ranker can do one step of parameter update.”. (note: this shows updating the parameters of the agent model (ranker) based on the reward signal. The policy of πθ is conditioned on query derived state and the gradient update uses the reward.); Hu, pg. 7, 5.1 The DPG-FBE Algorithm “…the parameters θ and w will be updated after any search session between the search engine agent and users.”. (note: this shows updating agent model parameters (θ) based on the reward and action taken during a search session initiated by a context)]. He and Hu are analogous art because they are from the same field of endeavor of computing a semantic relevance score. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to apply the bucketing token identification to the ranked content items as taught by He, using the dot-product formulation taught by Hu. The suggestion/motivation for doing so is Hu’s dot-product scoring offers computational efficiency and stronger performance [Hu, pg. 7, 6.1 Simulation “The ranking score of the item x under the ranking action µ is the inner product xTµ of the two vectors”]. He does not expressly disclose, but He combined with Hu and Xiao does teach: determining contextual relevance scores corresponding to the sampled content items, wherein a contextual relevance score measures contextual affinity of a content item to the context [Xiao, column 18, line 42 “the one or more answer retriever engines 320 can determine a dot product (also referred to as cosine similarity) of the query embedding vector (of the input query 304) and the question embedding vectors (of the questions from the QA pairs and/or {Q}A pairs) to determine a semantic similarity between the question defined by the input query 304 and the text of the questions in the QA space and the {Q}A space.”. (note: this shows contextual similarity, which is contextual affinity); column 18, line 61 “a value close to O indicates independence or orthogonality between two vectors, a value close to 1 indicates a highly positive similarity (e.g., an exact or very close match) between two vectors, and a value close to -1 indicates a highly negative similarity (e.g. opposite semantics).”. (note: this shows that the score is a range measure of contextual affinity))]; He, Hu, and Xiao are analogous art because they are from the same field of endeavor of computing a semantic relevance score. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to apply the bucketing token identification to the ranked content items as taught by He, using the dot-product formulation taught by Hu, and embedding approach as taught by Xiao. The suggestion/motivation for doing so is Hu’s dot-product scoring offers computational efficiency and stronger performance [Hu, pg. 7, 6.1 Simulation “The ranking score of the item x under the ranking action µ is the inner product xTµ of the two vectors”]. Xiao’s dot-product embedding approach effectively captures semantic meaning, it produces a concrete, computationally efficient measure of semantic affinity [Xiao, column 18, line 61 “…a value close to O indicates independence or orthogonality between two vectors, a value close to 1 indicates a highly positive similarity (e.g., an exact or very close match) between two vectors, and a value close to -1 indicates a highly negative similarity (e.g. opposite semantics).”]. He does not expressly disclose, but He combined with Lewis, Hu, and Xiao does teach: determining bucket identifiers corresponding to the sampled content items [Lewis, column 12, line 12 “Token manager 204 can locate the corresponding entry for the content item 140, determine the associated bucketing tokens 244, 246, 248 and provide the tokens to ranking service 120". (note: this shows determining bucket identifiers (bucketing tokens) corresponding to content items); Abstract “…each bucketing token comprises a unique identifier that identifies a plurality of content items as being associated with a group of users of the content sharing platform that have similar interests.”. (note: this shows that each bucket identifier (bucketing token) is a unique identifier associated with content items)]; determining, using parameters of an agent model and an embedding of the context, an action vector comprising weights corresponding to different bucket identifiers [He, column 5, line 48 “…the policy gradient function 210 trains the ranker 220. The goal of the ranker 220 is to maximize the future expected cumulative rewards…where τ is a trajectory of state-action sequence s0, a0,…, sT, aT sampled by the reinforcement ranker 20, and P(τ; θ) is the probability of trajectory T under policy πθ”. (note: this shows an agent model (reinforcement ranker) with parameters θ that uses context-based state information to determine action probabilities. The πθ defines an action distribution conditioned on the current sate embedding.); Hu, pg. 7, 6.1 Simulation “…a ranking action of the search engine is a n-dim weight vector μ = (μ1, ..., μn)⊤”. (note: this shows that the action of the agent model is a weight vector (action vector comprising weights). This weight vector is computed by the policy parametrized by θ using the current state, which includes context features.); pg. 9, 6.2 Application “The state of the environment is represented by a 90-dim feature vector, which contains the item page features, user features and query features of the current search session.”. (note: this shows that the state embedding used by the agent to determine its action vector includes context features (embedding of the context)); Lewis, column 12, line 22 “the ranking algorithm can determine the user buckets associated with a user of a given feed and weight content items from those user buckets more highly in the content ranking 160 for that feed.”. (note: this shows that the ranking algorithm uses bucket identifiers with associated weights for different buckets)]; for each sampled content item, scaling the contextual relevance score based on the bucket identifier corresponding to the sampled content item and a weight in the action vector corresponding to the bucket identifier [Lewis, Abstract “…the processing device improves a ranking score of each content item from the set of content items that has at least one bucketing token matching a bucketing token associated with the user of the social network platform,”. (note: this shows scaling (improving) the ranking score of each content item based on its bucket identifier (bucketing token)); column 12, line 17 “…content ranking module 124 takes the received bucketing tokens 244, 246, 248 and applies the tokens as inputs to a ranking algorithm. At block 585, method 500 adjusts a ranking of the content item in view of the output of the ranking algorithm.”. (note: this shows that the ranking algorithm receives bucket tokens for a content item and adjusts (scales) the ranking score. This is scaling the relevance score based on bucket identifiers and their corresponding weights.)]; He, Lewis, Hu, and Xiao are analogous art because they are from the same field of endeavor of computing a semantic relevance score. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to apply the bucketing token identification as taught by Lewis to the ranked content items as taught by He, using the dot-product formulation taught by Hu, and embedding approach as taught by Xiao. The suggestion/motivation for doing so is explicitly stated by Lewis as improved ranking quality [Lewis, column 3, line 46 “the use of user bucketing tokens can generate more watchtime and referrers, and for the feed ranking services, the bucketing tokens can deliver more time on site and an improved user experience.”]. Hu’s dot-product scoring offers computational efficiency and stronger performance [Hu, pg. 7, 6.1 Simulation “The ranking score of the item x under the ranking action µ is the inner product xTµ of the two vectors”]. Xiao’s dot-product embedding approach effectively captures semantic meaning, it produces a concrete, computationally efficient measure of semantic affinity [Xiao, column 18, line 61 “…a value close to O indicates independence or orthogonality between two vectors, a value close to 1 indicates a highly positive similarity (e.g., an exact or very close match) between two vectors, and a value close to -1 indicates a highly negative similarity (e.g. opposite semantics).”]. Per Claim 20, He-Lewis-Hu-Xiao discloses claim 19. He does not fully disclose, but with Hu and Xiao does teach determining the contextual relevance scores comprises: determining a first feature vector representing the context [Xiao, column 18, line 15 “…the one or more answer retriever engines 320 can generate a vector representation (referred to as an embedding vector or an embedding) for the input query 304”. (note: this shows generating a first feature vector representing the context. The input query 304 constitutes the context against which semantic similarity is measured.)]; determining second feature vectors representing metadata of the sampled content items respectively [Xiao, column 18, line 18, “…one or more vector representations (or embedding vectors) for the text passages contained in the QA space and the { Q} A space. In some cases, the embedding vectors can be generated using a neural network encoder, which can convert text to a mathematical vector representation of that text”. (note: this shows generating second feature vectors representing the text (metadata) of each content item respectively)]; and determining a dot product of the first feature vector and each one of the second feature vectors, wherein the contextual relevance scores are based on the dot products [Xiao, column 18, line 42 “the one or more answer retriever engines 320 can determine a dot product (also referred to as cosine similarity) of the query embedding vector (of the input query 304) and the question embedding vectors (of the questions from the QA pairs and/or {Q}A pairs) to determine a semantic similarity…”. (note: this shows computing the dot product of the first feature vector and each second feature vector, with the result being the relevance scores)]. He, Lewis, Hu, and Xiao are analogous art because they are from the same field of endeavor of computing a semantic relevance score. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to apply the bucketing token identification as taught by Lewis to the ranked content items as taught by He, using the dot-product formulation taught by Hu, and embedding approach as taught by Xiao. The suggestion/motivation for doing so is Xiao’s dot-product embedding approach effectively captures semantic meaning, it produces a concrete, computationally efficient measure of semantic affinity [Xiao, column 18, line 61 “…a value close to O indicates independence or orthogonality between two vectors, a value close to 1 indicates a highly positive similarity (e.g., an exact or very close match) between two vectors, and a value close to -1 indicates a highly negative similarity (e.g. opposite semantics).”]. /MIRANDA M HUANG/Supervisory Patent Examiner, Art Unit 2124 Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to Sayed M Shah whose telephone number is (571)272-9406. The examiner can normally be reached Monday-Friday 9:00 am - 5:00 pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Miranda Huang can be reached at (571) 270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SAYED MUNEER SHAH/Examiner, Art Unit 2124 /MIRANDA M HUANG/Supervisory Patent Examiner, Art Unit 2124
Read full office action

Prosecution Timeline

Jan 26, 2024
Application Filed
Jul 30, 2026
Non-Final Rejection mailed — §101, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month