DETAILED ACTION
This Office Action is in response to communications filed on March 10th, 2026 for Application No. 18/129,023, in which claims 1-20 are presented for examination. The amendments filed on March 10th, 2026 have been entered, where claims 1-7, 11-17, and 20 are amended.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-7, 9-17, and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Parekh et al. (hereinafter Parekh) (Patent No. US 8,762,364 B2) in view of Zheng et al. (hereinafter Zheng) (“DRN: A Deep Reinforcement Learning Framework for News Recommendation”) and Covington et al. (hereinafter Covington) (“Deep Neural Networks for YouTube Recommendations”).
Regarding Claim 1, Parekh teaches a method comprising, at a computer system comprising a processor and a computer-readable medium (Pg. 1, Abstract, “Embodiments of the invention relate to methods of presenting personalized search results pages to users, and to search engine systems and servers configured to implement such methods”, where “search engine systems and servers” are a computer system, which must have a processor and a computer-readable medium “to implement such methods”; see also Pg. 8, Col. 2, Ln. 2-6, “A system . . . comprises a web search client and a web search server. The web search client is implemented in a computer-readable medium”; see generally Pg. 13, Col. 11, Claim 9, Ln. 3-6, “An apparatus comprising: a processor; a storage medium for tangibly storing thereon program logic for execution by the processor”):
receiving a search query for an item, the search query associated with a user operating a user device (Pg. 9, Col. 3-4, Ln. 66-1, “an SRP [search results page] is generated in response to a search query. The query is first pre-processed”, where the “search query” must be received in order to be “pre-processed”; see also Pg. 3, Fig. 1 and Pg. 9, Col. 3, Ln. 36-42, “As shown in FIG. 1, a typical front page of a commercial search engine, the initial point of user interaction with a search engine comprises a text entry box and a query submission button. In FIG. 1 a user has typed "pizza" in the text entry box. The query submission button is labeled "Search"”, where “pizza” is a search query for an item, which is received once the “user” selects “Search”, and where display and “interaction with” “a commercial search engine” by a “user” requires operation of a user device);
generating a state descriptor based on the search query and the user (Pg. 7, Fig. 5, where the “Group Metadata”, the “Algorithmic Ranked Set”, and the “Advertisement Ranked Set” are, in combination, within the broadest reasonable interpretation of a state descriptor because they are collectively input to “Layout Process” and, in the aggregate, these values provide a context characterization description of the state relevant to the search; see also Pg. 9, Col. 4, Ln. 3-11, “The pre-processed query is provided to two separate sub-process pipelines. The first pipeline ranks all web content . . . in order of relevance to the pre-processed query. This ranking forms the algorithmic search results set for query . . . The second pipeline ranks all available ads, either text-based or graphical, also in order of relevance to the pre-processed query”, where the state descriptor is generated based on the search “query” because its subcomponents, “algorithmic search results set” and “rank[ing set of] all available ads”, are generated based on the search “query”; and Pg. 10, Col. 5-6, Ln. 63-2, “First, an engagement index (EI) is calculated for all users. The users are then partitioned based on EI into engaged and non-engaged users. Then, the engaged users are grouped into User Groups based on a set of Grouping Parameters. Then, the Groups, along with Group Metadata describing the Groups and Grouping Parameters, are supplied to a Layout Process”, where the state descriptor is generated based on the “User” because its subcomponent, “Group Metadata”, is generated based on the “User”);
selecting a candidate content composition to display content items in a particular order on a user interface responsive to the search query (Pg. 10, Col. 6, Ln. 4-7, “The advertising layout on each SRP is based on the Group Metadata. The ads and results in each layout comprise algorithmic results drawn from the two ranked sets” and Pg. 9, Col. 4, Ln. 10-16, “The SRP delivered in response to the query draws on both rankings: of ads and of algorithmic results. Typically, SRP construction involves merging the two rankings and includes a default sub-process to determine the layout of the ads on the page. Typically the rankings determine placement of ads within the layout”, where different combinations of “the rankings” of “results” and “ads”, which can be presented in different proportions, see Pg. 10, Col. 6, Ln. 43-44, “For the heavy and light clickers, the advertising layout is preferably altered as described below” and Pg. 11, Col. 8, Ln. 17-19, “For example, some embodiments decrease the total number of advertisements shown on the page if the user has a low probability of clicking on sponsored listings”, in combination with “advertising layout on each SRP” are a plurality of candidate content compositions for displaying content items in the particular “rank[ed]” order of content on a user interface, see Pg. 9, Col. 3, Ln. 36-42, “As shown in FIG. 1, a typical front page of a commercial search engine, the initial point of user interaction with a search engine comprises a text entry box and a query submission button”, where the operating, “interacting”, “user” utilizes “a typical . . . commercial search engine”, which requires a user interface, to “submi[t]” a search “query”, see also Pg. 4, Fig. 2, and “respons[ive] to the query”, which are selected in regard to layout of “ads” in the “North”, “East”, or “South” of the page, see Pg. 5, Fig. 3 and Pg. 8, Col. 1, Ln. 20-23, “The three locations are labeled as North, East, and South which are situated at the top, right-hand side and bottom (respectively) of the search results page”, and “ranking” order of content “within the layout”, see Pg. 9, Col. 4, Ln. 13-16, “Typically, SRP construction involves merging the two rankings and includes a default sub-process to determine the layout of the ads on the page. Typically the rankings determine placement of ads within the layout”),
wherein selecting the candidate content composition comprises applying . . . [operations] to select the candidate content composition, wherein applying the . . . [operations] comprises (Pg. 10, Col. 6, Ln. 4-7, “The advertising layout on each SRP is based on the Group Metadata. The ads and results in each layout comprise algorithmic results drawn from the two ranked sets” and Pg. 9, Col. 4, Ln. 10-16, “The SRP delivered in response to the query draws on both rankings: of ads and of algorithmic results. Typically, SRP construction involves merging the two rankings and includes a default sub-process to determine the layout of the ads on the page. Typically the rankings determine placement of ads within the layout”, where the selected candidate content composition is the “layout”, “rankings” of content, “ads and results”, and proportions of content, see Pg. 10, Col. 6, Ln. 43-44, “For the heavy and light clickers, the advertising layout is preferably altered as described below” and Pg. 11, Col. 8, Ln. 17-19, “For example, some embodiments decrease the total number of advertisements shown on the page if the user has a low probability of clicking on sponsored listings”, for the “SRP”, which is selected using “construction” operations):
converting the state descriptor [and] prior actions . . . into a state-decision . . . , wherein the state-decision . . . comprises . . . the state descriptor [and] the prior actions . . . (Pg. 9, Col. 4, Ln. 10-16, “The SRP delivered in response to the query draws on both rankings: of ads and of algorithmic results. Typically, SRP construction involves merging the two rankings and includes a default sub-process to determine the layout of the ads on the page. Typically the rankings determine placement of ads within the layout”, where “SRP construction” must be preceded by decisions “SRP construction”, such as layout decisions, see Pg. 10, Col. 6, Ln. 4-7, “The advertising layout on each SRP is based on the Group Metadata” and content proportions decisions, see Pg. 10, Col. 6, Ln. 43-44, “For the heavy and light clickers, the advertising layout is preferably altered as described below” and Pg. 11, Col. 8, Ln. 17-19, “For example, some embodiments decrease the total number of advertisements shown on the page if the user has a low probability of clicking on sponsored listings”, wherein the layout decisions are a conversion of the “Group Metadata” component of the state descriptor and the content proportions decisions are a conversion of prior actions, “heavy and light clickers”, which collectively comprise a decision which, given it is based on the state descriptor, is within the broadest reasonable interpretation of a state-decision);
. . . ; and
selecting content item[s] . . . of the plurality of content item[s] . . . in the particular order based on the similarity scores to present as the candidate content composition (compare Fig. 3, where an “SRP layout” is depicted, see Pg. 9, Col. 4, Ln. 1-21, which, as discussed above, in combination with “rankings” of content, “ads and results”, see Col. 4, Ln. 10-16, “The SRP delivered in response to the query draws on both rankings: of ads and of algorithmic results. Typically, SRP construction involves merging the two rankings and includes a default sub-process to determine the layout of the ads on the page. Typically the rankings determine placement of ads within the layout”, and proportions of content, see Pg. 10, Col. 6, Ln. 43-44, “For the heavy and light clickers, the advertising layout is preferably altered as described below” and Pg. 11, Col. 8, Ln. 17-19, “For example, some embodiments decrease the total number of advertisements shown on the page if the user has a low probability of clicking on sponsored listings”, form the candidate content composition, with Fig. 2, where an “an exemplary search results page (SRP)” is depicted, see Pg. 9, Col. 3, Ln. 43-47, where the transition from a candidate content composition, consistent entirely of layout, ranking, and proportionality metadata, to a “an exemplary search results page (SRP)” requires selecting content items of the plurality of content items contained in the rankings, such as “Pizza Hut” and “New Online Food Delivery Service”, in the manner prescribed by the layout, rankings, and proportions, which is according to the particular order of similarity scores of the ranking, see Pg. 9, Col. 4, Ln. 4-16, “The first pipeline ranks all web content . . . in order of relevance to the pre-processed query. This ranking forms the algorithmic search results set for query. The second pipeline ranks all available ads . . . also in order of relevance to the pre-processed query . . . Typically, SRP construction involves merging the two rankings”, where “ranking . . .web content . . . [and] ads” requires an organization value to determine position for each entity, which are within the broadest reasonable of similarity scores because they are calculated based on “relevance” i.e. similarity to “the pre-processed query”);
populating the candidate content composition with the content items according to a ranking of the similarity scores (compare Fig. 3, where an “SRP layout” is depicted, see Pg. 9, Col. 4, Ln. 1-21, which, as discussed above, in combination with “rankings” of content, “ads and results”, see Col. 4, Ln. 10-16, “The SRP delivered in response to the query draws on both rankings: of ads and of algorithmic results. Typically, SRP construction involves merging the two rankings and includes a default sub-process to determine the layout of the ads on the page. Typically the rankings determine placement of ads within the layout”, and proportions of content, see Pg. 10, Col. 6, Ln. 43-44, “For the heavy and light clickers, the advertising layout is preferably altered as described below” and Pg. 11, Col. 8, Ln. 17-19, “For example, some embodiments decrease the total number of advertisements shown on the page if the user has a low probability of clicking on sponsored listings”, form the candidate content composition, with Fig. 2, where an “an exemplary search results page (SRP)” is depicted, see Pg. 9, Col. 3, Ln. 43-47, where the transition from a candidate content composition, consistent entirely of layout, ranking, and proportionality metadata, to a “an exemplary search results page (SRP)” requires population with content items contained in the rankings, such as “Pizza Hut” and “New Online Food Delivery Service”, in the manner prescribed by the layout, rankings, and proportions, which is according to a ranking of similarity scores, see Pg. 9, Col. 4, Ln. 4-16, “The first pipeline ranks all web content . . . in order of relevance to the pre-processed query. This ranking forms the algorithmic search results set for query. The second pipeline ranks all available ads . . . also in order of relevance to the pre-processed query . . . Typically, SRP construction involves merging the two rankings”, where “ranking . . .web content . . . [and] ads” requires an organization value to determine position for each entity, which are within the broadest reasonable of similarity scores because they are calculated based on “relevance” i.e. similarity to “the pre-processed query”);
providing the candidate content composition in a search result interface (Pg. 8, Col. 2, Ln. 11-13, “generate, in response to the query, a search results page and pass it to the web search client for display, wherein the search results page is personalize”, where the information provided by the “web search client for display” comprises the candidate content composition, the “layout”, content proportions, see Pg. 10, Col. 6, Ln. 43-44, “For the heavy and light clickers, the advertising layout is preferably altered as described below” and Pg. 11, Col. 8, Ln. 17-19, “For example, some embodiments decrease the total number of advertisements shown on the page if the user has a low probability of clicking on sponsored listings”, and ranking that is populated with content of “ads and results”, of the “search results page”, see Pg. 10, Col. 6, Ln. 4-7, “The advertising layout on each SRP is based on the Group Metadata. The ads and results in each layout comprise algorithmic results drawn from the two ranked sets”; Pg. 9, Col. 4, Ln. 10-16, “The SRP delivered in response to the query draws on both rankings: of ads and of algorithmic results. Typically, SRP construction involves merging the two rankings and includes a default sub-process to determine the layout of the ads on the page. Typically the rankings determine placement of ads within the layout”; and Pg. 5, Fig. 3); and
sending the search result interface to the user device, wherein sending causes the user device to present the search result interface for viewing by the user operating the user device (Pg. 6; Fig. 4 and Pg. 9, Col. 3, Ln. 25-27, “FIG. 4 is a diagram outlining backend processing steps required to produce a search results page including a personalization engine consistent with the present invention”, where the search result interface, “search results page”, is “output”; see also Pg. 4, Fig. 2 and Pg. 9, Col. 3, Ln. 43-47, “FIG. 2 illustrates an exemplary search results page (SRP) returned by a commercial search engine in response to the query "pizza" in FIG. 1. The right column of the page comprises advertisements. The left column of the page comprises advertisements and algorithmic search results”, where the “search results page (SRP)” is the search result interface, which is sent, “returned by a commercial search engine”, to the user device for presentation to the operating user, see Pg. 9, Col. 3, Ln. 36-42, “As shown in FIG. 1, a typical front page of a commercial search engine, the initial point of user interaction with a search engine comprises a text entry box and a query submission button”, where the operating, “interacting”, “user” utilizes “a typical . . . commercial search engine”, which requires a user device, to “submi[t]” a search “query”, see Pg. 3, Fig. 1, which results in the presentation of the “search results page” for viewing by the operating user on their user device Pg. 4, Fig. 2; see also Pg. 8, Col. 2, Ln. 11-13, “generate, in response to the query, a search results page and pass it to the web search client for display, wherein the search results page is personalize”).
Parekh does not explicitly disclose . . . a reinforcement learning model . . . reinforcement learning model . . . a predicted reward . . . (where the operations for selecting the candidate content composition do not specifically describe a reinforcement learning model with a predicted reward; redundant recitations of predicted reward omitted)
embedding . . . a vector representation of . . . the vector representation having a same length as a content item embedding . . . comparing the state-decision embedding to a plurality of content item embeddings to obtain similarity scores, wherein the plurality of content item embeddings are associated with the content items . . . embeddings . . . (where the state-decision and content items are not specifically discussed with regard to embeddings and, as a result, the comparing and selecting operations do not specifically discussed with regard to embeddings; redundant recitations of embedding omitted).
However, Zheng teaches . . . [selecting a value by applying] a reinforcement learning model [to a state descriptor] (Pg. 168, Col. 2, Fig. 2; Pg. 168, Col. 2, Para. 1, “Our deep reinforcement recommender system can be shown as Figure 2. We follow the common terminologies in reinforcement learning [37] to describe the system . . . The state is defined as feature representation for users and action is defined as feature representation for news. Each time when a user requests for news, a state representation (i.e., features of users) and a set of action representations (i.e., features of news candidates) are passed to the agent. The agent will select the best action”, where the “agent” comprises of a “reinforcement learning” model, which is applied to the state descriptor, “The state is defined as feature representation for users and action is defined as feature representation for news”, to make a “select[ion”)
. . . the reinforcement learning model . . . [selecting based on] a predicted reward . . . (Pg. 167, Col. 1, Abstract, “, we propose a Deep Q-Learning based recommendation framework, which can model future reward explicitly. . We further consider user return pattern as a supplement to click / no click label in order to capture more user feedback information”, where predicted rewards, “model[led] future reward explicitly” to “consider return pattern” for future sessions, are used for the “Deep Q-Learning based recommendation framework”, which, as discussed above is a reinforcement learning model for selection, see Pg. 168, Col. 2, Para. 1, “Our deep reinforcement recommender system can be shown as Figure 2. We follow the common terminologies in reinforcement learning [37] to describe the system . . . The state is defined as feature representation for users and action is defined as feature representation for news. Each time when a user requests for news, a state representation (i.e., features of users) and a set of action representations (i.e., features of news candidates) are passed to the agent. The agent will select the best action”).
Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to combine the applying operations to select a candidate content composition, wherein the operations comprise converting the state descriptor and prior actions into a state-decision of Parekh with the selecting of a value by applying a reinforcement learning model to a state descriptor, wherein the reinforcement learning model’s selecting is based on a predicted reward of Zheng in order select a candidate content composition by converting a state descriptor, prior actions, and a predicted reward into an informative state-decision, which incorporates valuable information into the decision making process through an operation capable of modeling dynamic advertisement features and user preferences (Zheng, Pg. 175, Col. 2, Para. 2, “our method can effectively model the dynamic news features and user preferences, and plan for future explicitly, in order to achieve higher reward”, where “news features” are comparable to ad features; see also Zheng, Pg. 175, Col. 2, Para. 3, “Our method can be generalized to many other recommendation problems”), which leads to better accuracy, recommendation diversity (Zheng, Pg. 175, Col. 2, Para. 3, “Experiments have shown that our method can improve the recommendation accuracy and recommendation diversity significantly”), and long-term performance (Zheng, Pg. 167, Col. 2, Para. 3, “considering future rewards will help to improve recommendation performance in the long run”) by considering subsequent user activity as a metric for selection success (Zheng, Pg. 167, Col. 2, Para. 4, “how soon one user will return to this service [48] will also indicate how satisfied this user is with the recommendation” and Zheng, Pg. 168, Col. 2, Para. 1, “We consider user activeness to help improve recommendation accuracy, which can provide extra information than simply using user click labels”).
Additionally, Covington teaches . . . [converting a plurality of data into a state-decision] embedding . . . [comprising] a vector representation of [the plurality of data] . . . the vector representation having a same length as a content item embedding . . . (Pg. 2, Col. 2, Para. 2, “where u ∈ RN represents a high-dimensional “embedding” of the user, context pair and the vj ∈ RN represent embeddings of each candidate video. In this setting, an embedding is simply a mapping of sparse entities (individual videos, users etc.) into a dense vector in RN”, where a plurality of data, “user, context pair”, are converted into a state decision embedding, “u ∈ RN”, which comprises a “vector” representation of the plurality of data, “represents a high-dimensional “embedding” of the user, context pair”, and where the state-decision embedding, “u ∈ RN”, and content item embedding, “vj ∈ RN represent embeddings of each candidate video”, have the same vector representation length, “RN”, see also Pg. 3, Col. 1, Para. 2, “dot product space”; see generally Pg. 2, Col. 2, Para. 2, “The task of the deep neural network is to learn user embeddings u as a function of the user’s history and context that are useful for discriminating among videos with a softmax classifier”, where the “user embeddings u” are within the broadest reasonable interpretation of state-decision embeddings because they represent the state of “the user’s history and context” to inform decision making, “useful for discriminating among videos”)
comparing the state-decision embedding to a plurality of content item embeddings to obtain similarity scores (Pg. 3, Col. 1, Para. 2, “we need to compute the most likely N classes (videos) in order to choose the top N to present to the user . . . the scoring problem reduces to a nearest neighbor search in the dot product space” and Pg. 2, Col. 2, Para. 2, “u ∈ RN represents a high-dimensional “embedding” of the user, context pair and the vj ∈ RN represent embeddings of each candidate video. In this setting, an embedding is simply a mapping of sparse entities (individual videos, users etc.) into a dense vector in RN. The task of the deep neural network is to learn user embeddings u as a function of the user’s history and context that are useful for discriminating among videos”, where, the state-decision embedding, “u ∈ RN”, and the plurality of content item embeddings, “vj ∈ RN”, are compared “a nearest neighbor search in the dot product space”, to obtain a “scoring” which is used “to compute the most likely N classes (videos) in order to choose the top N to present to the user”; see also Pg. 3, Col. 1, Para. 3, “we learn high dimensional embeddings for each video in a fixed vocabulary and feed these embeddings into a feedforward neural network”),
wherein the plurality of content item embeddings are associated with the content items (Pg. 2, Col. 2, Para. 2, “the vj ∈ RN represent embeddings of each candidate video”, where the plurality of content item embeddings, “vj ∈ RN represent embeddings”, are associated with the content items, “of each candidate video”); and
[selecting content item] embeddings [of the plurality of content item] embeddings [in the particular order based on the similarity scores] . . . (Pg. 2, Col. 2, Fig. 2, “Recommendation system architecture demonstrating the “funnel” where candidate videos are retrieved and ranked before presenting only a few to the user”, where content items of the plurality of content items, “candidate videos”, are selected, “presenting only a few to the user”; see also Pg. 2, Col. 1, Para. 5, “The ranking network accomplishes this task by assigning a score to each video according to a desired objective function using a rich set of features describing the video and user. The highest scoring videos are presented to the user, ranked by their score”, where the selection is in particular order, “rank”, based on the similarity scores, “by their score”, and is performed by the “ranking network”, which operates on the embeddings, see Pg. 3, Col. 1, Para. 3, “we learn high dimensional embeddings for each video in a fixed vocabulary and feed these embeddings into a feedforward neural network”, thus the selection is of content item embeddings of the plurality of content item embeddings).
Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to combine the applying a reinforcement learning model to select a candidate content composition, wherein a state descriptor, prior actions, and a predicted reward are converted into a state-decision and content items are selected in a particular order, based on similarity scores, to present as the candidate data composition of Parekh in view of Zheng with the converting of a plurality of data into a state-decision embedding, which comprises a vector representation of the plurality of data and has the same length as a content item embedding; comparing the state-decision embedding to a plurality of content item embeddings to obtain similarity scores, wherein the content item embeddings are associated with content items; and selecting content item embeddings of the plurality of content item embeddings in the particular order based on the similarity score of Covington in order to reduce the burdens of data analysis by converting the state-decision information and content item information into a form suitable for model-assisted analysis (compare Covington, Pg. 5, Col. 2, Para. 4, “We typically use hundreds of features in our ranking models, roughly split evenly between categorical and continuous. Despite the promise of deep learning to alleviate the burden of engineering features by hand, the nature of our raw data does not easily lend itself to be input directly into feedforward neural networks” with Covington, Pg. 6, Col. 1, Para. 4, “we use embeddings to map sparse categorical features to dense representations suitable for neural networks” and Covington, Pg. 3, Col. 1, Para. 3, “we learn high dimensional embeddings for each video in a fixed vocabulary and feed these embeddings into a feedforward neural network”) and to utilize similarities between the state-decision and the content items to further customize the rankings (compare Parekh, Pg. 9, Col. 4, Ln. 4-16, “The first pipeline ranks all web content . . . in order of relevance to the pre-processed query. This ranking forms the algorithmic search results set for query. The second pipeline ranks all available ads . . . also in order of relevance to the pre-processed query . . . Typically, SRP construction involves merging the two rankings”, where only the content items, such as “web content” or “ad”, are considered in the similarity ranking, with Covington, Pg. 2, Col. 1, Para. 5, “Presenting a few“best”recommendations in a list requires a fine-level representation . . . The ranking network accomplishes this task by assigning a score to each video according to a desired objective function using a rich set of features describing the video and user. The highest scoring videos are presented to the user, ranked by their score”, where both the content item, “the video”, and the “user” are considered in the similarity ranking, leading to the “presentation” of relevant content that is “best” for the “user”), which will lead to better personalized and more engaging search result pages (Covington, Pg. 2, Col. 1, Para. 6, “The two-stage approach to recommendation allows us to make recommendations from a very large corpus (millions) of videos while still being certain that the small number of videos appearing on the device are personalized and engaging for the user”; Parekh, Abstract, “Embodiments of the invention relate to methods of presenting personalized search results pages to users”).
Regarding Claim 2, Parekh in view of Zheng and Covington teach the method of claim 1, wherein the search query is received in a first interaction session (Parekh, Pg. 9, Col. 3-4, Ln. 66-1, “an SRP [search results page] is generated in response to a search query. The query is first pre-processed”, where the “search query” must be received in order to be “pre-processed”; see also Parekh, Pg. 3, Fig. 1 and Parekh, Pg. 9, Col. 3, Ln. 36-42, “As shown in FIG. 1, a typical front page of a commercial search engine, the initial point of user interaction with a search engine comprises a text entry box and a query submission button. In FIG. 1 a user has typed "pizza" in the text entry box. The query submission button is labeled "Search"”, where “pizza” is a search query for an item, which is received once the “user” selects “Search” during a first interaction session, “interaction with” “a commercial search engine” by a “user”)
and the predicted reward is based on a time until a second interaction session occurring after providing information for the search result interface to the user device (Parekh, Pg. 8, Col. 2, Ln. 11-13, “generate, in response to the query, a search results page and pass it to the web search client for display, wherein the search results page is personalize”, where, as discussed above, the information provided by the “web search client for display”, is based on the selected content composition, which, in view of Zheng is selected based on a reinforcement learning model, see Zheng, Pg. 168, Col. 2, Para. 1, “Our deep reinforcement recommender system”, which is rewarded using a predicted reward, see Zheng, Pg. 167, Col. 1, Abstract, “, we propose a Deep Q-Learning based recommendation framework, which can model future reward explicitly. . We further consider user return pattern as a supplement to click / no click label in order to capture more user feedback information”; Zheng, Pg. 168, Col. 2, Para. 1, “the reward is composed of click labels and estimation of user activeness” and Zheng, Pg. 172, Para. 1, “We use survival models [18, 30] to model user return and user activeness. Survival analysis [18, 30] has been applied in the field of estimating user return time [20]. Suppose T is the time until next event (i.e., user return) happens”, where the “reward” is based on a time until a second interaction session because “the reward is [an] . . . estimation of user activeness” and “activeness” is based on “the time until next event (i.e., user return) happens”; see also Zheng, Pg. 167, Col. 2, Para. 4, “how soon one user will return to this service [48] will also indicate how satisfied this user is with the recommendation”).
The reasons for obviousness were discussed in regard to the rejection of Claim 1 above and remain applicable here.
Regarding Claim 3, Parekh in view of Zheng and Covington teach the method of claim 2, wherein the predicted reward is a penalty that increases as the time until the second interaction session increases (Zheng, Pg. 167, Col. 1, Abstract, “, we propose a Deep Q-Learning based recommendation framework, which can model future reward explicitly. . We further consider user return pattern as a supplement to click / no click label in order to capture more user feedback information”, where, as discussed above, the “model[led] future reward explicitly” to “consider return pattern” is the predicted reward, which is a penalty that increases with time until the second interaction session because, the reward, “rtotal”, increases with user activity, “ractive”, see Zheng, Pg. 172, Col. 1, Para. 4, “rtotal = rclick + βractive”, and user activity decreases, “decays”, as time increases until the second user interaction, see Zheng, Pg. 172, Col. 1, Para. 1-2, “Suppose T is the time until next event (i.e., user return) happens . . . Every time we detect a return of user, we will set S(t) = S(t) + Sa for this particular user. The user activeness score will not exceed 1 . . . Then, the user activeness continues to decay after t1. Similar things happen at t2, t3, t4 and t5”).
The reasons for obviousness were discussed in regard to the rejection of Claim 1 above and remain applicable here.
Regarding Claim 4, Parekh in view of Zheng and Covington teach the method of claim 1, wherein populating the candidate content composition with the content items according to the ranking of the similarity scores (compare Parekh, Fig. 3, where an “SRP layout” is depicted, see Parekh, Pg. 9, Col. 4, Ln. 1-21, which, as discussed above, in combination with “rankings” of content, “ads and results”, see Parekh, Col. 4, Ln. 10-16, “The SRP delivered in response to the query draws on both rankings: of ads and of algorithmic results. Typically, SRP construction involves merging the two rankings and includes a default sub-process to determine the layout of the ads on the page. Typically the rankings determine placement of ads within the layout”, and proportions of content, see Parekh, Pg. 10, Col. 6, Ln. 43-44, “For the heavy and light clickers, the advertising layout is preferably altered as described below” and Parekh, Pg. 11, Col. 8, Ln. 17-19, “For example, some embodiments decrease the total number of advertisements shown on the page if the user has a low probability of clicking on sponsored listings”, form the candidate content composition, with Fig. 2, where an “an exemplary search results page (SRP)” is depicted, see Parekh, Pg. 9, Col. 3, Ln. 43-47, where the transition from a candidate content composition, consistent entirely of layout, ranking, and proportionality metadata, to a “an exemplary search results page (SRP)” requires population with content items contained in the rankings, such as “Pizza Hut” and “New Online Food Delivery Service”, in the manner prescribed by the layout, rankings, and proportions, which is according to a ranking of similarity scores, see Parekh, Pg. 9, Col. 4, Ln. 4-16, “The first pipeline ranks all web content . . . in order of relevance to the pre-processed query. This ranking forms the algorithmic search results set for query. The second pipeline ranks all available ads . . . also in order of relevance to the pre-processed query . . . Typically, SRP construction involves merging the two rankings”, where “ranking . . .web content . . . [and] ads” requires an organization value to determine position for each entity, which are within the broadest reasonable of similarity scores because they are calculated based on “relevance” i.e. similarity to “the pre-processed query”)
is based at least in part on a presentation value of the content items (Parekh, Pg. 5, Fig. 3 and Parekh, Pg. 8, Col. 1, Ln. 20-23, “The three locations are labeled as North, East, and South which are situated at the top, right-hand side and bottom (respectively) of the search results page”, where the selection of which ad organizational content items to populate with selected ad content items is based on the presentation value of the ad content item populated the ad organizational content items, see Parekh, Pg. 8, Col. 1, Ln. 41-48, “the North region can also be a major contributor to a negative user experience. The higher the number of north sponsored listings, the more negative the user experience is likely to be especially if the ads are deemed irrelevant by the user” and Parekh, Pg. 10, Col. 6, Ln. 44-64, “For users identified as heavy clickers, the advertising layout is modified so that a different number of advertisements can be shown in the North region (or other prime location) for a subset of queries. If the advertising layout algorithm determines that N advertisements should be shown in the North (or other prime location), and if the personalization engine estimates the expected revenue (for this user) of moving additional advertisements to the north is above a certain threshold, it modifies the layout so that N+.alpha. ads are displayed in the north, where .alpha. is a positive integer . . . For users identified as light clickers, the advertising layout is modified so that a different number of advertisements can be shown in the North region (or other prime location) for a subset of queries”, where the number of “advertisements” content items shown in the “North (or other prime location)” ad organizational content items is based on the “negative user experience” presentation value).
Regarding Claim 5, Parekh in view of Zheng and Covington teach the method of claim 1, wherein
the candidate content composition is one of a plurality of candidate content compositions (Parekh, Pg. 10, Col. 6, Ln. 4-7, “The advertising layout on each SRP is based on the Group Metadata. The ads and results in each layout comprise algorithmic results drawn from the two ranked sets” and Parekh, Pg. 9, Col. 4, Ln. 10-16, “The SRP delivered in response to the query draws on both rankings: of ads and of algorithmic results. Typically, SRP construction involves merging the two rankings and includes a default sub-process to determine the layout of the ads on the page. Typically the rankings determine placement of ads within the layout”, where different combinations of “results” and “advertising layout on each SRP” are a plurality of candidate content compositions),
wherein the plurality of candidate content compositions describe different arrangements (Parekh, Pg. 10, Col. 6, Ln. 4-7, “The advertising layout on each SRP is based on the Group Metadata. The ads and results in each layout comprise algorithmic results drawn from the two ranked sets” and Parekh, Pg. 9, Col. 4, Ln. 10-16, “The SRP delivered in response to the query draws on both rankings: of ads and of algorithmic results. Typically, SRP construction involves merging the two rankings and includes a default sub-process to determine the layout of the ads on the page. Typically the rankings determine placement of ads within the layout”, where different arrangements of “results” and “advertising layout on each SRP” are a plurality of content compositions for presenting content “respons[ive] to the query”, which are identified in regard to layout of “ad” content items, first set of content items, in the ad organizational content items of “North”, “East”, or “South” regions of the page, second set of content items)
of a first set of content items selected based on relevance to the search query (Parekh, Pg. 9, Col. 4, Ln. 8-15, “The second pipeline ranks all available ads, either text-based or graphical, also in order of relevance to the pre-processed query. The SRP delivered in response to the query draws on both rankings: of ads and of algorithmic results. Typically, SRP construction involves merging the two rankings and includes a default sub-process to determine the layout of the ads on the page. Typically the rankings determine placement of ads within the layout”, where the “ads”, which as discussed above are the first set of content items, are selected based on “rank[ed]” for selection, “determine placement”, based on “relevance to the pre-processed query”)
and a second set of content items selected based at least in part on a presentation value of the second set of content items (Parekh, Pg. 8, Col. 1, Ln. 41-48, “the North region can also be a major contributor to a negative user experience. The higher the number of north sponsored listings, the more negative the user experience is likely to be especially if the ads are deemed irrelevant by the user” and Parekh, Pg. 9, Col. 3, Ln. 60-64, “Typically when advertisements are not shown in the North region, the Algorithmic region subsumes the area of the North region. However, when advertisements are not shown in the East region, that region is typically left blank”, where the “North” ad organizational content item, which as discussed above is part of the second set of content items, is “a major contributor to a negative user experience” and, as a result, the decision against selection “the Algorithmic region subsumes the area of the North region” is made based on the presentation value, “The higher the number of north sponsored listings, the more negative the user experience is likely to be especially if the ads are deemed irrelevant by the user” in relation to the user, see Parekh, Pg. 10, Col. 6, Ln. 44-64, “For users identified as heavy clickers, the advertising layout is modified so that a different number of advertisements can be shown in the North region (or other prime location) for a subset of queries. If the advertising layout algorithm determines that N advertisements should be shown in the North (or other prime location), and if the personalization engine estimates the expected revenue (for this user) of moving additional advertisements to the north is above a certain threshold, it modifies the layout so that N+.alpha. ads are displayed in the north, where .alpha. is a positive integer . . . For users identified as light clickers, the advertising layout is modified so that a different number of advertisements can be shown in the North region (or other prime location) for a subset of queries”).
Regarding Claim 6, Parekh in view of Zheng and Covington teach the method of claim 5, wherein the plurality of candidate content compositions include candidate content compositions having different ordering of the first set of content items and the second set of content items (Parekh, Pg. 10, Col. 6, Ln. 4-7, “The advertising layout on each SRP is based on the Group Metadata. The ads and results in each layout comprise algorithmic results drawn from the two ranked sets” and Parekh, Pg. 9, Col. 4, Ln. 10-16, “The SRP delivered in response to the query draws on both rankings: of ads and of algorithmic results. Typically, SRP construction involves merging the two rankings and includes a default sub-process to determine the layout of the ads on the page. Typically the rankings determine placement of ads within the layout”, where different combinations of “results” and “advertising layout on each SRP” are a plurality of candidate content compositions for presenting content “respons[ive] to the query”, which are identified in regard to layout of “ad” content items, first set, in the ad organizational content items of “North”, “East”, or “South” regions of the page, second set, and where ordering of the first and second set of content items differs across the plurality of candidate content compositions, see Parekh, Fig. 3 and Parekh, Pg. 9, Col. 3, Ln. 60-64, “Typically when advertisements are not shown in the North region, the Algorithmic region subsumes the area of the North region. However, when advertisements are not shown in the East region, that region is typically left blank”, where, in content compositions where the “Algorithmic region subsumes the area of the North region” or “the East region” is “left blank” have different orderings of the first and second sets, such as an ordering without the “North” ad organizational content item or an ordering with a different arrangement of ad content items in the “East” ad organizational content item).
Regarding Claim 7, Parekh in view of Zheng and Covington teach the method of claim 5, wherein the plurality of candidate content compositions include candidate content compositions having different quantities of the second set of content items (Parekh, Pg. 10, Col. 6, Ln. 4-7, “The advertising layout on each SRP is based on the Group Metadata. The ads and results in each layout comprise algorithmic results drawn from the two ranked sets” and Parekh, Pg. 9, Col. 4, Ln. 10-16, “The SRP delivered in response to the query draws on both rankings: of ads and of algorithmic results. Typically, SRP construction involves merging the two rankings and includes a default sub-process to determine the layout of the ads on the page. Typically the rankings determine placement of ads within the layout”, where different combinations of “results” and “advertising layout on each SRP” are a plurality of candidate content compositions for presenting content “respons[ive] to the query”, which are identified in regard to layout of “ad” content items, first set, in the ad organizational content items of “North”, “East”, or “South” regions of the page, second set, and where the quantities of the second set differs across the plurality of content compositions, see Parekh, Fig. 3 and Parekh, Pg. 9, Col. 3, Ln. 60-62, “Typically when advertisements are not shown in the North region, the Algorithmic region subsumes the area of the North region”, where, in content compositions where the “Algorithmic region subsumes the area of the North region” there is one less ad organizational content item).
Regarding Claim 9, Parekh in view of Zheng and Covington teach the method of claim 1, wherein generating the state descriptor is further based on a sequence of previous states for the user (Parekh, Pg. 10, Col. 5-6, Ln. 65-4, “Then, the engaged users are grouped into User Groups based on a set of Grouping Parameters. Then, the . . . Group Metadata describing the Groups and Grouping Parameters, are supplied to a Layout Process. The Layout Process also receives an Advertisement Ranked Set and an Algorithmic Ranked Set, and then constructs an SRP for each Group”, where, as discussed above, the “Group Metadata”, “Advertisement Ranked Set”, and “Algorithmic Ranked Set” are collectively the state descriptor, and as a result, the state descriptor is based on a sequence of previous states for the user because its subcomponent, “Group Metadata”, “describe[es] the Groups and Grouping Parameters”, which in turn is based on the “engage[ment of users]”; see also Parekh, Pg. 10, Col. 6, Ln. 24-28, “personalization of advertising layout is performed for groups of users based on their aggregate user behaviors. The candidate users are as categorized as `heavy`, `light` or `average` clickers by examining the historical click-through rates (CTR) on sponsored ad listings”, where “groups of users [are ] based on their aggregate user behaviors”, which is within the broadest reasonable interpretation of a sequence of previous states because it is the “historical” sequence of “click-through rates (CTR) on sponsored ad listings”).
Regarding Claim 10, Parekh in view of Zheng and Covington teach the method of claim 1, wherein the reinforcement learning model includes a reward based on relevance scores to the search query of content items in a candidate content composition (Parekh, Pg. 9, Col. 4, Ln. 8-18, “The second pipeline ranks all available ads, either text-based or graphical, also in order of relevance to the pre-processed query. The SRP delivered in response to the query draws on both rankings: of ads and of algorithmic results . . . Typically the rankings determine placement of ads within the layout. Though such advertising layout algorithms may formulate advertisement placement in different ways, they use query and individual ad-level features to determine placement”, where the “placement of ads within the layout”, which as discussed above, are a component of the candidate content composition, are selected based on “relevance to the pre-processed query”, which, in view of Zheng, are selected using a reinforcement learning model, see Zheng, Pg. 168, Col. 2, Para. 1, “Our deep reinforcement recommender system”, which is rewarded for selecting relevant items, using “probability” of “user” “click[s]” as a proxy metric, see Zheng, Pg. 170, Col. 2, Para. 4, “Under the setting of reinforcement learning, the probability for a user to click on a piece of news (and future recommended news) is essentially the reward that our agent can get”).
The reasons for obviousness were discussed in regard to the rejection of Claim 1 above and remain applicable here.
Regarding Claim 11, Parekh teaches a computer program product comprising a non-transitory computer readable storage medium having instructions encoded thereon that, when executed by a processor, cause the processor to perform steps comprising: . . . (Pg. 13, Col. 11, Claim 10, Ln. 30-32, “A non-transitory computer readable storage medium tangibly storing computer program instructions capable of being executed by a computer processor, the computer program instructions”, where the “computer program” is a computer program product, see Pg. 12, Col. 12, Claim 15, Ln. 20, “15. The computer program product of claim 10”).
The remaining limitations are substantially the same as limitations of Claim 1, therefore it is rejected under the same rationale.
Regarding Claim 12, the additional elements of the dependent claim are substantially the same as limitations of Claim 2, therefore it is rejected under the same rationale.
Regarding Claim 13, the additional elements of the dependent claim are substantially the same as limitations of Claim 3, therefore it is rejected under the same rationale.
Regarding Claim 14, the additional elements of the dependent claim are substantially the same as limitations of Claim 4, therefore it is rejected under the same rationale.
Regarding Claim 15, the additional elements of the dependent claim are substantially the same as limitations of Claim 5, therefore it is rejected under the same rationale.
Regarding Claim 16, the additional elements of the dependent claim are substantially the same as limitations of Claim 6, therefore it is rejected under the same rationale.
Regarding Claim 17, the additional elements of the dependent claim are substantially the same as limitations of Claim 7, therefore it is rejected under the same rationale.
Regarding Claim 19, the additional elements of the dependent claim are substantially the same as limitations of Claim 9, therefore it is rejected under the same rationale.
Regarding Claim 20, Parekh teaches a system comprising: one or more processors; and a non-transitory computer readable storage medium having instructions encoded thereon that, when executed by a processor, cause the processor to perform steps comprising: . . . (Pg. 1, Abstract, “Embodiments of the invention relate to methods of presenting personalized search results pages to users, and to search engine systems and servers configured to implement such methods”, where the “search engine systems and servers” are a system, which must have a processor and a non-transitory computer-readable medium with encoded instructions “to implement such methods”; see also Pg. 8, Col. 2, Ln. 2-6, “A system . . . comprises a web search client and a web search server. The web search client is implemented in a computer-readable medium”; see generally Pg. 13, Col. 11, Claim 9, Ln. 3-6, “An apparatus comprising: a processor; a storage medium for tangibly storing thereon program logic for execution by the processor” and Pg. 13, Col. 11, Claim 10, Ln. 30-32, “A non-transitory computer readable storage medium tangibly storing computer program instructions capable of being executed by a computer processor, the computer program instructions”).
The remaining limitations are substantially the same as limitations of Claim 1, therefore it is rejected under the same rationale.
Claims 8 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Parekh in view of Zheng, Covington, and Chen et al. (hereinafter Chen) (“Decision Transformer: Reinforcement Learning via Sequence Modeling”).
Regarding Claim 8, Parekh in view of Zheng and Covington teach the method of claim 1, wherein the reinforcement learning model . . . (Zheng, Pg. 168, Col. 2, Para. 1, “Our deep reinforcement recommender system”).
The reasons for obviousness were discussed in regard to the rejection of Claim 1 above and remain applicable here.
Parekh in view of Zheng and Covington do not explicitly disclose . . . is a decision transformer.
However, Chen teaches . . . [a reinforcement learning model that] is a decision transformer (Pg. 1, Abstract, “We introduce a framework that abstracts Reinforcement Learning (RL) as a sequence modeling problem . . . we present Decision Transformer, an architecture that casts the problem of RL as conditional sequence modeling. Unlike prior approaches to RL that fit value functions or compute policy gradients, Decision Transformer simply outputs the optimal actions by leveraging a causally masked Transformer”).
Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to combine the use of a reinforcement learning model to select a candidate content composition of Parekh in view of Zheng and Covington with the use of a reinforcement learning model that is a decision transformer of Chen in order to utilize a simplified and scalable reinforcement learning approach that matches or exceeds more complex and less scalable reinforcement learning approaches (Chen, Pg. 1, Abstract, “We introduce a framework that abstracts Reinforcement Learning (RL) as a sequence modeling problem. This allows us to draw upon the simplicity and scalability of the Transformer architecture . . . we present Decision Transformer, an architecture that casts the problem of RL as conditional sequence modeling. Unlike prior approaches to RL that fit value functions or compute policy gradients, Decision Transformer simply outputs the optimal actions by leveraging a causally masked Transformer . . . Despite its simplicity, Decision Transformer matches or exceeds the performance of state-of-the-art model-free offline RL baselines”).
Regarding Claim 18, the additional elements of the dependent claim are substantially the same as limitations of Claim 8, therefore it is rejected under the same rationale.
Response to Arguments
Applicant's arguments filed on March 10th, 2026 have been fully considered. Each argument is addressed in detail below.
I. Applicant argues the rejections of the claims, under 35 U.S.C. § 101, should be withdrawn (Applicant’s Remarks, 03/10/2026, Pg. 10, Section “Rejections under 35 U.S.C. § 101”).
Applicant’s amendments to the claims have overcome each and every rejection to the claims, under 35 U.S.C. § 101, as previously set forth in the December 11th, 2025 Office Action. As a result, these rejections have been withdrawn.
II. Applicant argues the rejections of the claims, under 35 U.S.C. § 103, should be withdrawn (Applicant’s Remarks, 03/10/2026, Pg. 11-12, Section “Rejections under 35 U.S.C. § 103”).
In response to Applicant’s amendments, the previously communicated rejections under 35 U.S.C. § 103, have been withdrawn. However, Applicants arguments are not persuasive in light of the new grounds for rejection, under 35 U.S.C. § 103, discussed in detail above. The new grounds of rejection rely on new prior art of record to teach the new combination of elements in the amended independent claims, which were not presented in this arrangement in any of the previously presented claims. As a result, Applicant’s arguments are rendered moot.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MATTHEW BRYCE GOLAN whose telephone number is (571)272-5159. The examiner can normally be reached Monday through Friday, 8:00 AM to 5:00 PM ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey Shmatov can be reached at (571) 270-3428. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MATTHEW BRYCE GOLAN/Examiner, Art Unit 2123
/ALEXEY SHMATOV/Supervisory Patent Examiner, Art Unit 2123