DETAILED ACTION
This action is in response to the original filing on February 8, 2023 and the Remarks and Amendments filed on April 27, 2026. Claims 1-20 are pending and have been considered below. Claims 1, 5, and 16 are independent claims. Claims 1, 5, 11, and 16 are amended.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Objections
Claim 5 is objected to because of the following informalities:
“based at least in part on the corresponding predicted actions” should read “based at least in part on the corresponding predicted action”
Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-15 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding claim 1:
Step 2A, Prong 1 – A judicial exception is recited in this claim as it recites mental processes (see MPEP 2106.04(a)(2)(III)):
determining… based at least in part on the sequence of actions, a first user embedding associated with the user that is representative of the user and is configured to predict a plurality of predicted user actions associated with the user, the plurality of predicted user actions including determining images with which the user is predicted to interact… Given a small enough sequence of user actions and a small enough number of images with which the user is predicted to interact, a human can reasonably perform determining a simple, low-dimensional “first user embedding” with the aid of a pen and paper, which is a mental process.
determining… based at least in part on the first user embedding and the contextual information, a second user embedding, the second user embedding providing an updated embedding that is representative of the user in view of the contextual information and is configured to predict a plurality of recommended images for the user… Given a simple, low-dimensional first user embedding, a small enough amount of contextual information, and a small enough number of images to recommend, a human can reasonably perform determining a simple, low-dimensional “second user embedding” with the aid of a pen and paper, which is a mental process.
Step 2A, Prong 2 – The following limitations are additional elements without significantly more than the abstract idea:
providing a first sequence of actions associated with a user to a first trained machine learning model as a first input to the first trained machine learning model, the first sequence of actions associated with the user including information associated with a plurality of images with which the user has interacted… providing a first sequence of actions as input to a model amounts to insignificant extra-solution activity of data gathering or outputting (see MPEP 2106.05(g)).
determining, using the first trained machine learning model… a machine learning model used as a mere tool to apply an exception is a generic element for performing or applying the abstract idea using a generic computing environment (see MPEP 2106.05(f)).
providing the first user embedding to a second trained machine learning model as a first input to the second trained machine learning model… providing the first user embedding as a first input to a model amounts to insignificant extra-solution activity of data gathering or outputting (see MPEP 2106.05(g)).
providing contextual information as a second input to the second trained machine learning model… providing contextual information as a second input to a model amounts to insignificant extra-solution activity of data gathering or outputting (see MPEP 2106.05(g)).
determining, using the second trained machine learning model… a machine learning model used as a mere tool to apply an exception is a generic element for performing or applying the abstract idea using a generic computing environment (see MPEP 2106.05(f)).
providing the plurality of recommended images to a client device associated with the user for presentation… providing recommended images to a device for presentation amounts to insignificant extra-solution activity of data gathering or outputting (see MPEP 2106.05(g)).
Step 2B: These elements are recited at such a high level of generality that they fail to integrate the abstract idea into a practical application, since they provide nothing more than mere instructions to implement an abstract idea on a generic computer (MPEP 2106.05(f)) or only amount to data gathering or outputting without significantly more (MPEP 2106.05(g)). These limitations, taken either alone or in combination, fail to provide an inventive concept. Thus, the claim is not patent eligible.
Claims 2-4 recite limitations which further narrow the abstract idea of claim 1 by specifying more details of the mental processes that occur:
Regarding claim 2, describing the first user embedding as being determined offline, in batch is attempting to limit the field of use without significantly more (see MPEP 2106.05(h)).
Regarding claim 3, this claim further limits the abstract ideas of claim 1 to be based on a mental process: incrementally determining an updated user embedding for the user based at least in part on a subset of the first sequence of user actions and the second sequence of user actions… given a small enough second sequence of user actions and subset of the first sequence of user actions, a human can reasonably perform incrementally determining a simple, low-dimensional updated user embedding with the aid of a pen and paper, which is a mental process. Furthermore, obtaining a second sequence of user actions associated with the user since the first user embedding was determined is still insignificant extra-solution activity (see MPEP 2106.05(g)).
Regarding claim 4, describing the first and second trained machine learning models as a single, end-to-end learned model is attempting to limit the field of use without significantly more (see MPEP 2106.05(h)).
Regarding claim 5:
Step 2A, Prong 1 – A judicial exception is recited in this claim as it recites mental processes:
determine, for one or more user actions of the first sequence of user actions, a corresponding embedding… Given a small enough sequence of user actions, a human can reasonably perform determining a simple, low-dimensional “corresponding embedding” with the aid of a pen and paper, which is a mental process.
determine a plurality of embeddings from the corresponding embeddings determined for each one or more user actions of the first sequence of user actions… Given a small enough number of simple, low-dimensional corresponding embeddings, a human can reasonably perform determining a plurality of simple, low-dimensional embeddings with the aid of a pen and paper, which is a mental process.
determine, for the plurality of embeddings, a corresponding predicted action, wherein a predicted action comprises determining a prediction of a user interaction with a content item… Given a small enough number of simple, low-dimensional embeddings, a human can reasonably perform determining “a corresponding predicted action” within the human mind or with the aid of a pen and paper, which is a mental process.
determine, based at least in part on the corresponding predicted actions, a user embedding that is representative of the user and is configured to predict a plurality of user actions over a defined timeframe… A human can reasonably perform determining a simple, low-dimensional “user embedding” that is configured to predict a small enough plurality of user actions over a short enough defined timeframe with the aid of a pen and paper, which is a mental process.
Step 2A, Prong 2 – The following limitations are additional elements without significantly more than the abstract idea:
one or more processors; and a memory storing program instructions amount to mere tools to apply the judicial exception using a generic computing environment and is not indicative of significantly more (see MPEP 2106.05(f)).
instructions that, when executed by the one or more processors, cause the one or more processors at least: receive a first sequence of user actions associated with a user, the first sequence of actions associated with the user including information associated with a plurality of content items with which the user has interacted… receiving a sequence of user actions amounts to insignificant extra-solution activity of data gathering or outputting (see MPEP 2106.05(g)).
providing for presentation one or more content items to a client device associated with the user over the defined timeframe and based on the predicted user actions… providing content items to a device for presentation amounts to insignificant extra-solution activity of data gathering or outputting (see MPEP 2106.05(g)).
Step 2B: These elements are recited at such a high level of generality that they fail to integrate the abstract idea into a practical application, since they provide nothing more than mere instructions to implement an abstract idea on a generic computer (MPEP 2106.05(f)) or only amount to data gathering or outputting without significantly more (MPEP 2106.05(g)). These limitations, taken either alone or in combination, fail to provide an inventive concept. Thus, the claim is not patent eligible.
Claims 6-15 recite limitations which further narrow the abstract idea of claim 5 by specifying more details of the mental processes that occur:
Regarding claim 6, this claim further limits the abstract ideas of claim 5 to be based on a mental process: incrementally determine an updated embedding for the user based at least in part on a subset of the first sequence of user actions and the second sequence of user actions. Given a small enough second sequence of user actions and subset of the first sequence of user actions, a human can reasonably perform incrementally determining a simple, low-dimensional updated embedding with the aid of a pen and paper, which is a mental process. Furthermore, receiving a second sequence of user actions associated with the user since the user embedding was determined is still insignificant extra-solution activity (see MPEP 2106.05(g)).
Regarding claim 7, specifying wherein the user embedding is further configured to predict a classification associated with the user in this manner does not overcome the rejection of claim 5 as modifying the embedding does not make “determining” to not be a mental process.
Regarding claim 8, this claim further limits the abstract idea of claim 6 to be based on a mathematical concept (see MPEP 2106.04(a)(2)(I)): prior to incrementally determining the updated embedding, determine that a number of actions included in the second sequence of user actions exceeds a threshold value. Determining whether a number of actions exceed a threshold value is a mathematical calculation.
Regarding claim 9, this claim further limits the abstract ideas of claim 5 to be based on a mental process: determine a context aware user embedding based at least in part on the user embedding and the contextual information. Given a small enough amount of context information and a simple, low-dimensional user embedding, a human can reasonably perform determining “a context aware user embedding” with the aid of a pen and paper, which is a mental process. Furthermore, receiving contextual information associated with the user is still insignificant extra-solution activity (see MPEP 2106.05(g)).
Regarding claim 10, specifying wherein the contextual information includes at least one of: a query submitted by the user; an interest associated with the user; or a content item with which the user has interacted in this manner does not overcome the rejection of claim 9 as modifying the contextual information does not make “determining” to not be a mental process.
Regarding claim 11, this claim further limits the abstract idea of claim 9 to be based on a mental process: identify, based at least in part on the context aware user embedding, one or more content items from a corpus of content items to present to the user in response to a request for content items. For example, given a simple enough, low-dimensional context aware user embedding and a small enough corpus of content items, a human can reasonably perform identifying one or more content items to present to the user in response to a request for content items, which is a mental process.
Regarding claim 12, describing wherein the user embedding is generated offline in batch and the context aware user embedding is generated in real-time is attempting to limit the field of use without significantly more (see MPEP 2106.05(h)).
Regarding claim 13, this claim further limits the abstract idea of claim 5 to be based on a mathematical concept: wherein a causal mask is applied to the first sequence of user actions. Applying a causal mask to a first sequence of user actions involves performing matrix operations to form a lower-triangular matrix, which is a mathematical calculation.
Regarding claim 14, specifying wherein the predicted plurality of user actions includes representations of content items with which the user is expected to engage in this manner does not overcome the rejection of claim 5 as modifying the predicted plurality of user actions does not make “determining” to not be a mental process.
Regarding claim 15, specifying wherein the first sequence of user actions includes representations of content items with which the user has engaged in this manner does not overcome the rejection of claim 5 as modifying the first sequence of user actions does not make “determining” to not be a mental process.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-4 are rejected under 35 U.S.C. 103 as being unpatentable over Zhao et al. (US 20210256366 A1, hereinafter Zhao) in view of Pal et al. (“PinnerSage: Multi-Modal User Embedding Framework for Recommendations at Pinterest,” 2020, hereinafter Pal) and further in view of Chen et al. (“Top-K Off-Policy Correction for a REINFORCE Recommender System,” 2019, hereinafter Chen), and further in view of Covington et al. (“Deep Neural Networks for YouTube Recommendations,” 2016, hereinafter Covington).
Regarding claim 1:
Regarding the limitation a computer-implemented method, comprising: providing a first sequence of actions associated with a user to a first trained machine learning model as a first input to the first trained machine learning model, the first sequence of actions associated with the user including information associated with a plurality of images with which the user has interacted, Zhao teaches a computer-implemented method (¶6 “Certain embodiments provide a method”), comprising: providing a first sequence of actions associated with a user to a first trained machine learning model as a first input to the first trained machine learning model (Fig. 1 – 104, 112, ¶28 “The first model 104 and the second model 106 are deep-learning models… that have been trained,” ¶30 “time dependent data 112 stored in the user database 108 is routinely accessed by the first model 104… The time dependent data 112 includes a set of time segments… The time dependent data 112 is generated by identifying user activity for a time segment (e.g., the most recent 24 hours),” ¶33 “For example, the time dependent data 112 can include user search keywords entered by the user in the past 24 hour time segment,” wherein “time dependent data” including “user search keywords entered… in the past 24 hour time segment” encompasses a first sequence of actions associated with a user). However, Zhao fails to teach the first sequence of actions associated with the user including information associated with a plurality of images with which the user has interacted.
Pal, in the same field of endeavor, teaches the first sequence of actions associated with the user including information associated with a plurality of images with which the user has interacted (Page 1, Col. 1, Section 1, ¶1 “Each pin is an image item… Users can save pins on boards,” Page 3, Col. 1, Section 3, ¶1 “Let Au = {a1,a2, . . .} be the sequence of action pins of user u, such that for each a… user either repined, or clicked pin Pa at time Tu[a]”).
Regarding the limitation determining, using the first trained machine learning model and based at least in part on the sequence of actions, a first user embedding associated with the user that is representative of the user and is configured to predict a plurality of predicted user actions associated with the user, the plurality of predicted user actions including determining images with which the user is predicted to interact, Zhao teaches teaches determining, using the first trained machine learning model and based at least in part on the sequence of actions (Fig. 1 – 104, 112, 124, ¶29 “user data includes time dependent data 112,” ¶32 “Using both the user data and the application data, the first model 104 of the recommendation engine 102 generates a relevance score 124 for each available third-party application by transforming the user data and application data into vectors,” wherein generating a “relevance score” using “the user data” which includes “time dependent data” encompasses determining, using the first trained machine learning model and based at least in part on the sequence of actions), a list of applications associated with the user that is representative of the user (Fig. 1 – 104, 126, ¶35 “Once vector representation(s) are generated for each vector, the first model 104 concatenates the vectors generated and passes the concatenated vector through an activation function (e.g., Softmax) to generate a relevance score for each available third-party application. The relevance score indicates whether the third-party application is of relevance or interest to the user… Based on the relevance scores of the third-party applications, the first model 104 determines those third-party applications that are relevant applications 126 to the user,” ¶36 “Once the first model 104 identifies the relevant applications 126, the relevant applications are arranged by row… from highest ranking application of relevance to the lowest ranking application of relevance”). However, Zhao fails to teach a first user embedding associated with the user that is representative of the user and is configured to predict a plurality of predicted user actions associated with the user, the plurality of predicted user actions including determining images with which the user is predicted to interact.
Pal teaches determining images with which the user is predicted to interact (Page 6, Col. 1, Section 5.2, ¶1 “Figure 6 provides an illustration of candidates retrieved by PinnerSage. The recommendation set is a healthy mix of pins that are relevant to the top three interests of the user,” Page 7, Figure 6 depicts the determined images with which the user is predicted to interact). However, Pal fails to teach a first user embedding associated with the user that is representative of the user and is configured to predict a plurality of predicted user actions associated with the user, the plurality of predicted user actions including determining images with which the user is predicted to interact.
Chen, in the same field of endeavor, teaches determining a first user embedding associated with the user… and is configured to predict a plurality of predicted user actions associated with the user (Page 456, Col. 2, ¶3 “user preferences over these items are shifting all the time, resulting in continuously-evolving user states,” Page 457, Col. 1, ¶2 “We offer a novel top-K off-policy correction to account for the fact that our recommender outputs multiple items at a time,” Page 458, Col. 2, Section 4.1, ¶1 “We model our belief on user state at each time t, which capture both evolving user interests using a n-dimensional vector, that is, st… We model the state transition… with a recurrent neural network:
PNG
media_image1.png
31
141
media_image1.png
Greyscale
Page 459, Col. 1, Fig. 1 depicts using “Event 1, 2, . . ., t” to determine a “User state st+1” or “n-dimensional vector,” which encompasses a first user embedding given its broadest reasonable interpretation, ¶2 “Conditioning on a user state s, the policy πθ(a|s) is labeled as a simple softmax,” Page 460, Col. 2, ¶1 “when the desirable item has a small mass in the softmax policy πθ(.|s), the top-K correction more aggressively pushes up its likelihood… Once the softmax policy πθ(.|s) casts a reasonable mass on the desirable item (to ensure it will be likely to appear in the top-K), the correction… no longer tries to push up its likelihood. This in return allows other items of interest to take up some mass in the softmax policy,” Page 461, Col. 1, ¶2 “We consider using a stochastic policy where recommendations are sampled from πθ ,” wherein a “user state” vector used to calculate a “softmax policy” from which “recommendations” are sampled from encompasses configured to predict a plurality of predicted actions associated with the user), the plurality of predicted user actions including videos with which the user is predicted to interact (Page 457, Col. 2, Section 3, ¶2 “we predict the next… videos to recommend”).
Regarding the limitation providing the first user embedding to a second trained machine learning model as a first input to the second trained machine learning model, Zhao teaches providing a list of applications to a second trained (Fig. 1 – 106, ¶28) machine learning model as a first input to the second trained machine learning model (Fig. 1 – 106, 126, ¶36 “The relevant applications 126, arranged by row from highest ranking application of relevance to the lowest ranking application of relevance, is input to the second model 106 to generate a connection score 130 for the relevant applications 126”). However, Zhao fails to teach the first user embedding.
Chen teaches the first user embedding (Page 458, Col. 2, Section 4.1, ¶1 “user state at each time t… using a n-dimensional vector… st,” Page 459, Col. 1, Fig. 1 – “User state st+1”).
Zhao further teaches providing contextual information as a second input to the second trained machine learning model (Fig. 1 – 106, 118, 120, ¶37 “second model 106 accesses unstructured data such as user search history 118, application topic 120, user reviews of an application, tag data, application description, etc.”).
Regarding the limitation determining, using the second trained machine learning model and based at least in part on the first user embedding and the contextual information, a second user embedding, the second user embedding providing an updated embedding that is representative of the user in view of the contextual information and is configured to predict a plurality of recommended images for the user, Zhao teaches determining, using the second trained machine learning model and based at least in part on a list of applications and the contextual information (Fig. 1 – 106, 126, 130, ¶37 “In order to generate the connection scores 130 for each relevant application 126, the second model 106 accesses… unstructured data”), a list of recommended applications… that is representative of the user in view of the contextual information (Fig. 1 – 106, 126, 130-134, ¶38 “The second model 106 generates an engagement score 132 for each relevant third-party application based on combining the connection score 130 of the relevant third-party application with the corresponding relevance score 124 (e.g., multiplying the relevance score and the connection score). Based on the engagement score 132, the recommendation engine 102 determines the top-X third-party applications to recommend to the user”). However, Zhao fails to teach determining, using the second trained machine learning model and based at least in part on the first user embedding and the contextual information, a second user embedding, the second user embedding providing an updated embedding that is representative of the user in view of the contextual information and is configured to predict a plurality of recommended images for the user.
Pal teaches to predict a plurality of recommended images for the user (Page 7, Figure 6 as explained above, Page 6, Col. 1, Section 5.2, ¶1 and Col. 2, Section 5.2, ¶1 “We ran large scale A/B experiments where users are randomly assigned either in control or experiment groups… Users assigned to the experiment group experience PinnerSage recommendations… Users across the two groups are shown equal number of recommendations,” wherein “PinnerSage recommendations” encompass a plurality of recommended images). However, Pal fails to teach determining, using the second trained machine learning model and based at least in part on the first user embedding and the contextual information, a second user embedding, the second user embedding providing an updated embedding that is representative of the user in view of the contextual information and is configured to predict a plurality of recommended images for the user.
Chen teaches the first user embedding (Page 458, Col. 2, Section 4.1, ¶1, Page 459, Col. 1, Fig. 1). However, Chen fails to teach determining, using the second trained machine learning model and based at least in part on the first user embedding and the contextual information, a second user embedding, the second user embedding providing an updated embedding that is representative of the user in view of the contextual information and is configured to predict a plurality of recommended images for the user.
Covington, in the same field of endeavor, teaches determining, based on a user embedding and demographic information, a second user embedding, the second user embedding providing an updated embedding that is representative of the user (Page 3, Col. 1, Section 3.2, ¶1 “we learn high dimensional embeddings for each video… and feed these embeddings into a feedforward neural network. A user’s watch history is represented by a variable-length sequence of sparse video IDs which is mapped to a dense vector representation via the embeddings,” Section 3.3, ¶1 “Demographic features are important for providing priors so that the recommendations behave reasonably for new users,” Page 2, Col. 2, Section 3.1, ¶1 “The task of the deep neural network is to learn user embeddings u as a function of the user’s history and context that are useful for discriminating among videos with a softmax classifier,” Page 4, Fig. 3 – “user vector u”) and is configured to predict recommended videos for the user (Page 3, Col. 1, ¶2 “At serving time we need to compute the most likely N classes (videos) in order to choose the top N to present to the user,” Page 4, Fig. 3, Caption: “At serving, an approximate nearest neighbor lookup is performed to generate hundreds of candidate video recommendations”).
Regarding the limitation providing the plurality of recommended images to a client device associated with the user for presentation, Zhao teaches providing a plurality of applications to a client device associated with the user for presentation (Fig. 1 – 134, ¶38 “The recommended applications 134 are displayed to the user… in the user interface,” Fig. 8 – 800, 808, ¶70 “ Server 800 further includes input/output (I/O) device(s) 808… such as, for example… displays… and other devices that allow for interaction with server 800… server 800 may connect with external I/O devices through physical and wireless connections,” ¶71 “Server 800 further includes network interface 88, which provides server 800 with access to external network 806 and thereby external computing devices”). However, Zhao fails to teach the plurality of recommended images.
Pal teaches the plurality of recommended images (Page 7, Figure 6, Page 6, Col. 1, Section 5.2, ¶1 and Col. 2, Section 5.2, ¶1 all as explained above).
Zhao, Pal, Chen, and Covington are analogous art to the claimed invention as all are from the same field of endeavor of recommender systems. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the image recommendation system of Pal, the first embedding of Chen, and the second, updated embedding of Covington with the methodology of Zhao. The motivation to do so is to design a method that “provides high quality personalized recommendations” (Pal, Abstract), to improve user satisfaction based on their previous actions (Chen, Page 457, Col. 2, Section 3, ¶1 “For each user, we consider a sequence of user historical interactions… Given such a sequence, we predict the next action to take… so that user satisfaction metrics… indicated by clicks or watch time, improve”), and to better predict user actions (Covington, Page 4, Col. 1, ¶2 “We therefore found much better performance predicting the user’s next watch”).
Regarding claim 2, Zhao in view of Pal and further in view of Chen, and further in view of Covington teaches the computer-implemented method of claim 1 (and thus the rejection of claim 1 is incorporated).
Chen teaches wherein the first user embedding is determined offline, in batch (Page 457, Col. 2, Section 3, ¶1 “we consider… user feedback, such as clicks and watch time,” Page 458, Col. 1, ¶1 “We seek a policy… that casts a distribution over the item to recommend… conditional to the user state,” Section 4, ¶1 “we cannot perform online updates to the policy and generate trajectories according to the updated policy immediately. Instead, we receive logged feedback of actions chosen by a historical policy,” Col. 2, ¶1 “we learn from batched feedback collected,” Section 4.1, ¶1 “We model our belief on the user state at each time t, which capture… evolving user interests using a n-dimensional vector… st… The action taken at each time t along the trajectory is embedded using an m-dimensional vector uat… We model the state transition… with a recurrent neural network”).
Zhao and Chen are analogous art to the claimed invention as both are from the same field of endeavor of recommender systems. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the offline batch processing of Chen with the methodology of Zhao. The motivation to do so is to implement a more efficient recommender system (Chen, Page 458, Col. 1, Section 4, ¶1 “our learner does not have real-time interactive control of the recommender due to learning and infrastructure constraints… Instead, we receive logged feedback of actions…”).
Regarding claim 3, Zhao in view of Pal and further in view of Chen, and further in view of Covington teaches the computer-implemented method of claim 1 (and thus the rejection of claim 1 is incorporated).
Chen teaches obtaining a second sequence of user actions associated with the user since the first user embedding was determined (Page 458, Col. 2, Section 4.1, ¶1 “We model our belief on user state at each time t, which capture both evolving user interests using a n-dimensional vector, that is, st… The action taken at each time t along the trajectory is embedded using an m-dimensional vector uat,” wherein at each time t, the model receives a user action uat after a previous user state st, or the first user embedding was determined; hence, user actions taken after the first sequence of actions encompasses a second sequence of user actions associated with the user since the first user embedding was determined).
Chen further teaches and incrementally determining an updated embedding for the user based at least in part on a subset of the first sequence of user actions and the second sequence of user actions (Page 458, Col. 2, Section 4.1, ¶1 “We model the state transition… with a recurrent neural network,” the equation “st+1 = f(st, uat)” describes incrementally determining an updated embedding for the user or user state st+1 based at least in part on a subset of the first sequence of user actions or the original sequence of user actions and the second sequence of user actions or the user actions taken after st was determined).
Zhao and Chen are analogous art to the claimed invention as both are from the same field of endeavor of recommender systems. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the incremental updating of the user actions and embedding of Chen with the methodology of Zhao. The motivation to do so is to improve user satisfaction based on their previous actions (Chen, Page 457, Col. 2, Section 3, ¶1 “For each user, we consider a sequence of user historical interactions… Given such a sequence, we predict the next action to take… so that user satisfaction metrics… indicated by clicks or watch time, improve”).
Regarding claim 4, Zhao in view of Pal and further in view of Chen, and further in view of Covington teaches the computer-implemented method of claim 1 (and thus the rejection of claim 1 is incorporated).
Zhao further teaches wherein the first trained machine learning model and the second trained machine learning model are implemented as a single, end-to-end learned model (Fig. 1 – 102, 104, 106, Fig. 7 – 702-704, 716, ¶28 “The recommendation engine 102 includes a first model 104 and a second model 106. The first model 104 and the second model 106 are deep-learning models (e.g., recurrent neural networks) that have been trained to generate values indicative of a prediction of user interest in and retention of, respectively, third-party applications,” ¶55, ¶58, ¶65, Figure 1 depicts a single RECOMMENDATION ENGINE with a FIRST MODEL linked to a SECOND MODEL; one of ordinary skill in the art would recognize that the RECOMMENDATION ENGINE which maps “user data” and “application data” to “personalized recommendations to the user” encompasses a single, end-to-end learned model given its broadest reasonable interpretation).
Claims 5-7, and 14-15 are rejected under 35 U.S.C. 103 as being unpatentable over Zhao in view of Kang et al. (“Self-Attentive Sequential Recommendation,” 2018, hereinafter Kang) and further in view of Chen.
Regarding claim 5:
Zhao teaches a computing system, comprising: one or more processors (¶7 “a processor of a computing system”).
Zhao further teaches and a memory storing program instructions (¶88 “the functions may be stored or transmitted over as one or more instructions or code on a computer-readable medium… Computer-readable media include… computer storage media”) that, when executed by the one or more processors (¶85 “the various operations of methods described above may be performed by any suitable means capable of performing the corresponding functions. The means may include various hardware… including… a circuit… or processor,” ¶89 “The computer-readable media may comprise a number of software modules. The software modules include instructions that, when executed by an apparatus such as a processor, cause the processing system to perform various functions”), cause the one or more processors at least: receive a first sequence of user actions associated with a user, the first sequence of actions associated with the user including information associated with a plurality of content items with which the user has interacted (Fig. 1 – 102, 104, 112, ¶28 “The recommendation engine 102 includes a first model 104,” ¶30 “time dependent data 112… is routinely accessed by the first model 104 of the recommendation engine 102… time dependent data 112 includes a set of time segments… time dependent data 112 is generated by identifying user activity for a time segment… and aggregating the user activity by the type of user activity… the recommendation engine 102 can track user activity across the time segments,” ¶39 “time dependent data is user activity data collected by an application and aggregated into time segments… in the context of a tax preparation application, such an application can collect information about the user interacting with the application, such as a user searching for a tax form, entry of user information, submitting a tax form, completing a financial transaction,” ¶31 “The first model 104 also accesses application data in the application database 110 regarding applications used previously by the user, indicating the types of applications the user uses… the recommendation engine 102… can access the computing device the application is installed on to identify which applications the user has downloaded on the computing device”).
Zhao fails to each determine, for one or more user actions of the first sequence of user actions, a corresponding embedding. However, Kang, in the same field of endeavor, teaches this limitation (Page 3, Col. 1, Table 1: “Ê = input embedding matrix,” Section III, ¶1 “In the setting of sequential recommendation, we are given a user’s action sequence Su… the model predicts the next item depending on the previous t items,” Section A, ¶1-2 “We transform the training sequence… into a fixed-length sequence s = (s1, s2, . . . sn)… We create an item embedding matrix M… and retrieve the input embedding matrix… where Ei = Msi… we inject a learnable position embedding P… into the input embedding:
PNG
media_image2.png
165
478
media_image2.png
Greyscale
Page 1, Col. 2, Fig. 1 depicts a red, green, and blue block or a corresponding embedding determined at “Embedding “Layer” for every “Training Action Sequence”).
Kang further teaches determine a plurality of embeddings from the corresponding embeddings determined for one or more user actions of the first sequence of user actions (Page 3, Col. 1, Table 1: “S(b) = item embeddings after the b-th self-attention layer, F(b) = item embeddings after the b-th feed-forward network,” Col. 2, Section B, ¶2 “the self-attention operation takes the embedding Ê as input, converts it to three matrices through linear projections, and feeds them into an attention layer:
PNG
media_image3.png
46
591
media_image3.png
Greyscale
Section C, ¶1 “After the first self-attention block, Fi essentially aggregates all previous items’ embeddings (i.e., Êj, j <= i)… it might be useful to learn more complex item transitions via another self-attention block based on F. Specifically, we stack the self-attention block (i.e., a self-attention layer and a feed forward network), and the b-th (b > 1) block is defined as:
PNG
media_image4.png
90
555
media_image4.png
Greyscale
Page 4, Col. 1, ¶2 “if low-layer features are useful, the model can easily propagate them to the final layer… we assume residual connections are also useful in our case… existing sequential recommendation methods have shown that the last visited item plays a key role on predicting the next item… after several self-attention blocks, the embedding of the last visited item is entangled with all previous items; adding residual connections to propagate the last visited item’s embedding to the final layer would make it much easier for the model to leverage low-layer information,” Page 1, Col. 2, Fig. 1 depicts multi-colored blocks or a plurality of embeddings output from “Point-Wise Feed Forward Network” that are determined from the corresponding embeddings determined for one or more user actions of the first sequence of user actions or the red, green, and blue blocks determined at “Embedding Layer”).
Kang further teaches determine, for the plurality of embeddings, a corresponding predicted action, wherein a predicted action comprises determining a prediction of a user interaction with a content item (Abstract: “SASRec seeks to identify which items are ‘relevant’ from a user’s action history, and use them to predict the next item,” Page 3, Col. 1, Table 1: “F(b) = item embeddings after the b-th feed-forward network,” Page 4, Col. 2, Section D, ¶1 “we predict the next item (given the first t items) based on Ft(b),” Page 6, Col. 1, Section A, ¶2 “we treat the presence of a review or rating as implicit feedback (i.e., the user interacted with the item) and use timestamps to determine the sequence order of actions” Page 1, Col. 2, Fig. 1 depicts “Prediction Layer” predicting next item “S4” based on the plurality of embeddings determined after “Point-Wise Feed Forward Network”).
Regarding the limitation determine, based at least in part on the corresponding predicted actions, a user embedding that is representative of the user and is configured to predict a plurality of user actions over a defined timeframe, Kang teaches the corresponding predicted actions (Page 1, Col. 2, Fig. 1 – S4, Page 3, Col. 1, Table 1 – F(b), Page 4, Col. 2, Section D, ¶1, Page 6, Col. 1, Section A, ¶2 all as explained above, wherein the corresponding predicted actions is interpreted to mean “the corresponding predicted action,” as explained under claim objections). However, Kang fails to teach determine, based at least in part on the corresponding predicted actions, a user embedding that is representative of the user and is configured to predict a plurality of user actions over a defined timeframe.
Chen teaches determine, based at least in part on predicted user actions, a user embedding that is representative of the user and is configured to predict a plurality of user actions (Page 456, Col. 2, ¶3, Page 457, Col. 1, ¶2 and Col. 2, Section 3, ¶2, Page 458, Col. 2, Section 4.1, ¶1, Page 459, Col. 1, Fig. 1 and ¶2, Page 460, Col. 2, ¶1, Page 461, Col. 1, ¶2 all as explained above with respect to claim 1, Page 459, Col. 2, Section 4.3, ¶1 to Page 460, ¶1 “As users are going to browse through (the full or partial set of) our recommendations and potentially interact with more than one item, we need to pick a set of relevant items instead of a single one… we seek a policy… each action A is to select a set of k items to maximize the expected cumulative award”) over a defined timeframe (Page 462, Col. 1, Section 6.2, ¶2 “We evaluate these methods on a production RNN candidate generation model in use at YouTube… The model is one of many candidate generators that produce recommendations, which are scored and ranked… before being shown to users on the YouTube Homepage… the immediate reward r is designed to reflect different user activities; videos that are recommended but not clicked receive zero reward. The long term reward R is aggregated over a time horizon of 4-10 hours… Experiments are run for multiple days, during which the model is trained continuously with new events being used as training data,” Page 463, Col. 2, Fig. 5 and ¶1 “We plot the results during a 5-day experiment,” wherein a set of k “recommendations” that the user is expected to click on to maximize a reward which is “aggregated over a time horizon of 4-10 hours” encompasses to predict a plurality of user actions over a defined timeframe).
Regarding the limitation providing for presentation one or more content items to a client device associated with the user over the defined timeframe and based on the predicted user actions, Zhao teaches providing for presentation one or more content items to a client device associated with the user (Fig. 1 – 134, ¶38, Fig. 8 – 800, 808, ¶70-71). However, Zhao fails to teach providing for presentation… over the defined timeframe and based on the predicted user actions.
Chen teaches providing for presentation videos over the defined timeframe and based on the predicted user actions (Page 457, Col. 2, Section 3, ¶2 “For each user, we consider a sequence of user historical interactions with the system, recording the actions taken by the recommender… as well as user feedback, such as clicks and watch time. Given such a sequence, we predict… videos to recommend,” Page 462, Col. 1, Section 6.2, ¶1 “the goal of any recommender systems is ultimately to improve real user experience. We therefore conduct a series of A/B experiments running a live system,” ¶2 “recommendations, which are scored and ranked… before being shown to users on the YouTube Homepage… the immediate reward r is designed to reflect different user activities; videos that are recommended but not clicked receive zero reward. The long term reward R is aggregated over a time horizon of 4-10 hours… Experiments are run for multiple days, during which the model is trained continuously with new events being used as training data,” Page 463, Col. 2, Fig. 5 and ¶1 “We plot the results during a 5-day experiment”).
Zhao, Kang, and Chen are analogous art to the claimed invention as all are from the same field of endeavor of recommender systems. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the SASRec architecture of Kang and the user embedding and recommendations of Chen with the system of Zhao. The motivation to do so is to design a recommender system that “outperforms various state-of-the-art sequential models… on both sparse and dense datasets” (Kang, Abstract) and to improve user satisfaction based on their previous actions (Chen, Page 457, Col. 2, Section 3, ¶1 “For each user, we consider a sequence of user historical interactions… Given such a sequence, we predict the next action to take… so that user satisfaction metrics… indicated by clicks or watch time, improve”).
Regarding claim 6, Zhao in view of Kang and further in view of Chen teaches the computing system of claim 5 (and thus the rejection of claim 5 is incorporated).
Chen teaches wherein the program instructions, when executed by the one or more processors, further cause the one or more processors at least (Abstract: “we present… a production top-K recommender system at YouTube, built with a policy-gradient-based algorithm, i.e. REINFORCE,” wherein a “recommender system” using an “algorithm” implies program instructions which are executed by one or more processors): receive a second sequence of user actions associated with the user since the user embedding was determined (Page 458, Col. 2, Section 4.1, ¶1 as explained above with respect to claim 3).
Chen further teaches and incrementally determine an updated embedding for the user based at least in part on a subset of the first sequence of user actions and the second sequence of user actions (Page 458, Col. 2, Section 4.1, ¶1 as explained above with respect to claim 3).
Zhao and Chen are analogous art to the claimed invention as both are from the same field of endeavor of recommender systems. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the incremental updating of the user actions and embedding of Chen with the system of Zhao. The motivation to do so is to improve user satisfaction based on their previous actions (Chen, Page 457, Col. 2, Section 3, ¶1 “For each user, we consider a sequence of user historical interactions… Given such a sequence, we predict the next action to take… so that user satisfaction metrics… indicated by clicks or watch time, improve”).
Regarding claim 7, Zhao in view of Kang and further in view of Chen teaches the computing system of claim 5 (and thus the rejection of claim 5 is incorporated).
Regarding the limitation wherein the user embedding is further configured to predict a classification associated with the user, Zhao teaches wherein a list of applications is further configured to predict a classification associated with the user (Fig. 1 – 104, 126 ¶35 “The relevance score indicates whether the third-party application is of relevance or interest to the user. For example, if the application data regarding a rarely used application by the user was entered into the first model 104, the relevance score would be low for that third-party application (as well as other third-party applications associated with the same topic), reflecting the user's disinterest in the application. Based on the relevance scores of the third-party applications, the first model 104 determines those third-party applications that are relevant applications 126 to the user,” wherein “interest to the user” or “the user’s disinterest” encompasses a classification associated with the user). However, Zhao fails to teach the user embedding.
Chen teaches the user embedding (Page 458, Col. 2, Section 4.1, ¶1, Page 459, Col. 1, Fig. 1 as explained above with respect to claim 1).
Zhao and Chen are analogous art to the claimed invention as all are from the same field of endeavor of recommender systems. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the user embedding of Chen with the system of Zhao. The motivation to do so is to improve user satisfaction based on their previous actions (Chen, Page 457, Col. 2, Section 3, ¶1 “For each user, we consider a sequence of user historical interactions… Given such a sequence, we predict the next action to take… so that user satisfaction metrics… indicated by clicks or watch time, improve”).
Regarding claim 14, Zhao in view of Kang and further in view of Chen teaches the computing system of claim 5 (and thus the rejection of claim 5 is incorporated).
Zhao teaches wherein the predicted plurality of user actions includes representations of content items with which the user is expected to engage (Fig. 1 – 104, 124-126, ¶29 “The first model 104… determines whether a third-party application is relevant or of interest to the user (e.g., determining whether the user is likely to click on or connect to the third-party application if recommended by the recommendation engine),” ¶32 “the first model 104 of the recommendation engine 102 generates a relevance score 124 for each available third-party application by transforming the user data and application data into vectors,” ¶35 “Once vector representation(s) are generated for each vector, the first model 104 concatenates the vectors generated and passes the concatenated vector through an activation function (e.g., Softmax) to generate a relevance score for each available third-party application”).
Regarding claim 15, Zhao in view of Kang and further in view of Chen teaches the computing system of claim 5 (and thus the rejection of claim 5 is incorporated).
Chen teaches wherein the first sequence of user actions includes representations of content items with which the user has engaged (Page 457, Col. 2, Section 3, ¶2 “a sequence of user historical interactions with the system, recording… user feedback, such as clicks and watch time,” Page 458, Col. 2, Section 4.1, ¶1 “The action taken at each time t along the trajectory is embedded using an m-dimensional vector uat,” Page 459, Col. 1, Fig. 1 depicts multiple “Item embeddings” for “Event 1, 2, . . ., t" or wherein the first sequence of user actions includes representations of content items with which the user has engaged).
Zhao and Chen are analogous art to the claimed invention as all are from the same field of endeavor of recommender systems. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the representations of content items of Chen with the system of Zhao. The motivation to do so is to improve user satisfaction based on their previous actions (Chen, Page 457, Col. 2, Section 3, ¶1 “For each user, we consider a sequence of user historical interactions… Given such a sequence, we predict the next action to take… so that user satisfaction metrics… indicated by clicks or watch time, improve”).
Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Zhao in view of Kang and further in view of Chen, and further in view of Wu et al. (“Rethinking Lifelong Sequential Recommendation with Incremental Multi-Interest Attention,” hereinafter Wu).
Regarding claim 8, Zhao in view of Kang and further in view of Chen teaches the computing system of claim 6 (and thus the rejection of claim 6 is incorporated).
Regarding the limitation wherein the program instructions, when executed by the one or more processors, further cause the one or more processors at least: prior to incrementally determining the updated embedding, determine that a number of actions included in the second sequence of user actions exceeds a threshold value, Chen teaches wherein the program instructions, when executed by the one or more processors, further cause the one or more processors at least (Abstract, as explained above with respect to claim 6): prior to incrementally determining the updated embedding, determine that a number of actions included in the second sequence of user actions is at least one action (Page 458, Col. 2, Section 4.1, ¶1 - equation “st+1 = f(st, uat)” depicts determining the updated embedding after only a single user action uat). However, Chen fails to teach determine that a number of actions included in the second sequence of user actions exceeds a threshold value.
Wu, in the same field of endeavor, teaches determining that a number of actions exceeds a threshold value (Page 8, Col. 1, ¶1 “we extract the subset of users whose behavior sequences has length greater or equal a threshold 𝑙, we compute the performance of different methods only on this subset of users. We then observe how does the performance vary when we increase the threshold 𝑙”).
Zhao, Chen, and Wu are analogous to the claimed invention as all are from the same field of endeavor of recommender systems. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to implement a system that combined the incremental updating of Chen and the threshold value of Wu with the system of Zhao. The motivation to do so is to improve user satisfaction based on their previous actions (Chen, Page 457, Col. 2, Section 3, ¶1 “For each user, we consider a sequence of user historical interactions… Given such a sequence, we predict the next action to take… so that user satisfaction metrics… indicated by clicks or watch time, improve”) and to design “a novel incremental self-attention based method for lifelong sequential recommendation, which goes beyond the limitations of RNN-based memory networks, while also possesses their ability to incrementally update the user representation for online inference” (Wu, Page 2, Col. 2, ¶2).
Claims 9-12 are rejected under 35 U.S.C. 103 as being unpatentable over Zhao in view of Kang and further in view of Chen, and further in view of Covington.
Regarding claim 9, Zhao in view of Kang and further in view of Chen teaches the computing system of claim 5 (and thus the rejection of claim 5 is incorporated).
Zhao teaches wherein the program instructions, when executed by the one or more processors, further cause the one or more processors at least: receive contextual information associated with the user (Fig. 1 – 106, 118, 120, 130, ¶37 “In order to generate the connection scores 130 for each relevant application 126, the second model… accesses additional types of application data… such as user search history 118, application topic 120, user reviews of an application, tag data, application description, etc.”).
Regarding the limitation and determine a context aware user embedding based at least in part on the user embedding and the contextual information, Zhao teaches and determine a context aware list of applications based at least in part on a list of applications and the contextual information (Fig. 1 – 102, 106, 126, 130-134, ¶38 “The second model 106 generates an engagement score 132 for each relevant third-party application based on combining the connection score 130 of the relevant third-party application with the corresponding relevance score 124 (e.g., multiplying the relevance score and the connection score). Based on the engagement score 132, the recommendation engine 102 determines the top-X third-party applications to recommend to the user”). However, Zhao fails to teach and determine a context aware user embedding based at least in part on the user embedding and the contextual information.
Chen teaches the user embedding (Page 458, Col. 2, Section 4.1, ¶1, Page 459, Col. 1, Fig. 1 as explained above with respect to claim 1). However, Chen fails to teach determine a context aware user embedding based at least in part on the user embedding and the contextual information.
Covington teaches determining a context aware user embedding (Page 3, Col. 1, Section 3.2, ¶1, Section 3.3, ¶1, Page 2, Col. 2, Section 3.1, ¶1, Page 4, Fig. 3 – “user vector u”).
Zhao, Chen, and Covington are analogous art to the claimed invention as all are from the same field of endeavor of recommender systems. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the user embedding of Chen and the context-aware embedding of Covington with the system of Zhao. The motivation to do so is to improve user satisfaction based on their previous actions (Chen, Page 457, Col. 2, Section 3, ¶1 “For each user, we consider a sequence of user historical interactions… Given such a sequence, we predict the next action to take… so that user satisfaction metrics… indicated by clicks or watch time, improve”) and to better predict user actions (Covington, Page 4, Col. 1, ¶2 “We therefore found much better performance predicting the user’s next watch”).
Regarding claim 10, Zhao in view of Kang and further in view of Chen, and further in view of Covington teaches the computing system of claim 9 (and thus the rejection of claim 9 is incorporated).
Zhao teaches wherein the contextual information includes at least one of: a query submitted by the user (Fig. 1 – 106, 118, ¶37 “the second model 106 accesses unstructured data such as user search history 118”); an interest associated with the user (¶37 “user reviews of an application”); or a content item with which the user has interacted (Fig. 1 – 120, ¶37 “user reviews” implies a user has already interacted with an application or content item).
Regarding claim 11, Zhao in view of Kang and further in view of Chen, and further in view of Covington teaches the computing system of claim 9 (and thus the rejection of claim 9 is incorporated).
Regarding the limitation wherein the program instructions, when executed by the one or more processors, further cause the one or more processors at least: identify, based at least in part on the context aware user embedding, one or more content items from a corpus of content items to present to the user in response to a request for content items, Zhao teaches wherein the program instructions, when executed by the one or more processors, further cause the one or more processors at least: identify, based at least in part on the context aware list of applications, one or more content items from a corpus of content items to present to the user in response to a request for content items (Fig. 1 – 102, 132, 134, ¶38 “Based on the engagement score 132, the recommendation engine 102 determines the top-X third-party applications to recommend to the user. The recommended applications 134 are displayed to the user. In some cases, the recommended application 134 are displayed upon user request”). However, Zhao fails to teach the context aware user embedding.
Covington teaches the context aware user embedding (Page 3, Col. 1, Section 3.2, ¶1, Section 3.3, ¶1, Page 2, Col. 2, Section 3.1, ¶1, Page 4, Fig. 3 – “user vector u”).
Zhao and Covington are analogous art to the claimed invention as both are from the same field of endeavor of recommender systems. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the context-aware embedding of Covington with the system of Zhao. The motivation to do so is to better predict user actions (Covington, Page 4, Col. 1, ¶2 “We therefore found much better performance predicting the user’s next watch”).
Regarding claim 12, Zhao in view of Kang and further in view of Chen, and further in view of Covington teaches the computing system of claim 9 (and thus the rejection of claim 9 is incorporated).
Regarding the limitation wherein the user embedding is generated offline in batch and the context aware user embedding is generated in real-time, Zhao teaches and the context aware list of applications is generated in real-time (Fig. 1 – 106, 132-134, ¶20 “The third-party application(s) recommended to the user include those third-party applications the recommendation engine has determined to be of relevance to the user and that the user will connect and use. In some cases, the recommendation can be displayed to the user upon request or automatically (e.g., in real time),” ¶38 “The second model 106 generates an engagement score 132 for each relevant third-party application… Based on the engagement score 132, the recommendation engine 102 determines the top-X third-party applications to recommend to the user. The recommended applications 134 are displayed to the user. In some cases, the recommended application 134 are displayed upon user request. In other cases, the recommended application 134 are displayed automatically in the user interface,” wherein the “top-X third party applications to recommend” being displayed “automatically (e.g., in real time)” implies that the recommendations are generated in real-time). However, Zhao fails to teach wherein the user embedding is generated offline in batch and the context aware user embedding is generated in real-time.
Covington teaches the context aware user embedding (Page 3, Col. 1, Section 3.2, ¶1, Section 3.3, ¶1, Page 2, Col. 2, Section 3.1, ¶1, Page 4, Fig. 3 – “user vector u”). However, Covington fails to teach wherein the user embedding is generated offline in batch…
Chen teaches wherein the user embedding is generated offline in batch (Page 457, Col. 2, Section 3, ¶1, Page 458, Col. 1, ¶1, Section 4, ¶1, Col. 2, ¶1, and Section 4.1, ¶1).
Zhao, Chen, and Covington are analogous art to the claimed invention as all are from the same field of endeavor of recommender systems. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the offline batch processing of Chen and the context aware user embedding of Covington with the system of Zhao. The motivation to do so is to implement a more efficient recommender system (Chen, Page 458, Col. 1, Section 4, ¶1 “our learner does not have real-time interactive control of the recommender due to learning and infrastructure constraints… Instead, we receive logged feedback of actions…”) and to better predict user actions (Covington, Page 4, Col. 1, ¶2 “We therefore found much better performance predicting the user’s next watch”).
Claim 13 is rejected under 35 U.S.C. 103 as being unpatentable over Zhao in view of Kang and further in view of Chen, and further in view of Raffel et al. (“Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer,” 2019, hereinafter Raffel).
Regarding claim 13, Zhao in view of Kang and further in view of Chen teaches the computing system of claim 5 (and thus the rejection of claim 5 is incorporated).
Regarding the limitation wherein a causal mask is applied to the first sequence of user actions, Kang teaches a first sequence of user actions (Page 3, Col. 1, Table 1: “Su = historical interaction sequence for a user u”). However, Kang fails to teach wherein a causal mask is applied to the first sequence of user actions.
Raffel, in the same field of endeavor, teaches wherein a causal mask is applied to a sequence of inputs (Fig. 3, Page 16, ¶2 “When producing the ith entry of the output sequence, causal masking prevents the model from attending to the jth entry of the input sequence for j > i. This is used during training so that the model can’t “see into the future” as it produces its output”).
Zhao, Kang, and Raffel are analogous to the claimed invention as all are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the causal mask of Raffel and the user actions of Kang with the system of Zhao. The motivation to do so is to design a recommender system that “outperforms various state-of-the-art sequential models… on both sparse and dense datasets” (Kang, Abstract) and to “achieve state-of-the-art results on many benchmarks covering summarization, question answering, text classification, and more” (Raffel, Abstract).
Claims 16-17 are rejected under 35 U.S.C. 103 as being unpatentable over Grbovic et al. (“E-commerce in Your Inbox: Product Recommendations at Scale,” 2016, hereinafter Grbovic) in view of Kang and further in view of Paraschakis et al. (“Comparative Evaluation of Top-N Recommenders in e-Commerce: An Industrial Perspective,” 2015, hereinafter Paraschakis), and further in view of Chen.
Regarding claim 16:
Grbovic teaches a computer-implemented method for training a sequential machine learning model (Abstract: “we describe a system that leverages user purchase history determined from e-mail receipts to deliver highly personalized product ads… We propose to use a novel neural language-based algorithm specifically tailored for delivering effective product recommendations”), comprising: obtaining a first sequence of user actions, the first sequence of actions associated with the user including information associated with a plurality of content items with which the user has interacted (Page 2, Col. 2, Section 2.1, ¶1 “commercial e-mails in the form of promotions and purchase receipts convey a strong, very direct purchase intent signal that can enable advertisers to reach a high-quality audience,” Page 5, Col. 1, Section 4.1, ¶1 “Our data sets included e-mail receipts sent to users” Page 5, Col. 2, Section 4.1, ¶3 “More formally, data set Dp… was derived by forming e-mail receipt sequences… for each user… along with their timestamps,” wherein “e-mail receipt sequences… along with their timestamps” encompasses a first sequence of user actions associated with the user including information associated with a plurality of content items with which the user has interacted or purchased).
Grbovic fails to teach determining a point in time within the first sequence of user actions. However, Kang teaches this limitation (Page 6, Col. 1, Section A, ¶2 “we… timestamps to determine the sequence order of actions… we split the historical sequence Su for each user u into three parts: (1) the most recent action… (2) the second most recent action… and (3) all remaining actions,” wherein determining a point in time is implicit when using timestamps to partition the data into the first and second most recent actions and “all remaining actions”).
Kang further teaches dividing the first sequence of user actions into a first plurality of user actions that were performed prior to the point in time and a second plurality of user actions that were performed after the point in time (Page 6, Col. 1, Section A, ¶2 “we split the historical sequence Su for each user u into three parts: (1) the most recent action… (2) the second most recent action… and (3) all remaining actions,” wherein the first and second most recent actions encompass a second plurality of user actions that were performed after the implicit point in time and “all remaining actions” encompass a first plurality of user actions that were performed prior…).
Kang further teaches providing the first plurality of user actions to the sequential machine learning model as training inputs (Abstract: “The goal of our work is to balance these two goals, by proposing a self-attention based sequential model (SASRec),” Page 6, Col. 1, Section A, ¶2 “(3) all remaining actions for training… during testing, the input sequences contain training actions,” Section B, ¶1 “we include three groups of recommendation baselines,” Col. 2, ¶3 “The last group contains deep-learning based sequential recommender systems, which consider several (or all) previously visited items,” Page 7, Table III and Col. 2, Section E, ¶2 “Our method SASRec outperforms all baselines”).
Regarding the limitation providing the second plurality of user actions to the sequential machine learning model as labeled positive training data, Kang teaches the second plurality of user actions (Page 6, Col. 1, Section A, ¶2 “(1) the most recent action… (2) the second most recent action”) and the sequential learning model (Abstract: “a self-attention based sequential model (SASRec)”). However, Kang fails to teach providing the second plurality of user actions to the sequential machine learning model as labeled positive training data.
Paraschakis, in the same field of endeavor, teaches providing data after a point in time to a model as labeled positive training data (Page 1026, Col. 1, ¶3-4 “The availability of timestamps allows us to attempt a more realistic setup, where we train on past events to predict future events… we set a split point on the dataset's timeline (indicated with the red vertical line), which acts as our “present”. Past events located to the left of the split point are used for training, whereas future events located on the right side are used for testing… Keeping the split point (and hence the test set) fixed, we train the models using the expanding time window, starting from the most recent history and then going back in time at a fixed rate up to full history… the task is to recommend products matching those of the test set,” Col. 2, Fig. 3).
Regarding the limitation training the sequential machine learning model using the training inputs and the labeled positive training data to generate user embeddings that are representative of corresponding users and are configured to predict a plurality of user actions over a period of time for the corresponding users, Kang teaches training the sequential machine learning model using the training inputs (Page 6, Col. 1, Section A, ¶2)… to generate user embeddings that are representative of corresponding users (Page 1, Col. 2, Fig. 1 – “Embedding Layer,” Page 3, Col. 1, Table 1: “U = user… Su = historical interaction sequence for a user u… Ê = input embedding matrix,” Section III, ¶1, and Section A, ¶1-2 all as explained above with respect to claim 5) and are configured to predict a user action for the corresponding users (Abstract: “recommender systems, which seek to capture the ‘context’ of users’ activities on the basis of actions they have performed recently… SASRec seeks to identify which items are ‘relevant’ from a user’s action history, and use them to predict the next item,” Page 1, Col. 2, Fig. 1 – “Point-Wise Feed Forward Network,” Page 3, Col. 1, Table 1: “S(b) = item embeddings after the b-th self-attention layer, F(b) = item embeddings after the b-th feed-forward network” and Col. 2, Section B, ¶2, Page 4, Col. 1, ¶2 and Col. 2, Section D, ¶1, Page 6, Col. 1, Section A, ¶2 all as explained above with respect to claim 5). However, Kang fails to teach training the sequential machine learning model using the training inputs and the labeled positive training data to generate user embeddings that are representative of corresponding users and are configured to predict a plurality of user actions over a period of time for the corresponding users.
Paraschakis teaches training a model using training inputs prior to a point in time and the labeled positive training data (Page 1026, Col. 1, ¶3-4 and Col. 2, Fig. 3). However, Paraschakis fails to teach user embeddings… configured to predict a plurality of user actions over a period of time.
Chen teaches embeddings configured to predict a plurality of user actions over a period of time (Page 456, Col. 2, ¶3, Page 457, Col. 1, ¶2 and Col. 2, Section 3, ¶2, Page 458, Col. 2, Section 4.1, ¶1, Page 459, Col. 1, Fig. 1 and ¶2, Page 460, Col. 2, ¶1, Page 461, Col. 1, ¶2 all as explained above with respect to claim 1, Page 460, Col. 1, ¶1, Page 462, Col. 1, Section 6.2, ¶2, Page 463, Col. 2, Fig. 5 and ¶1 all as explained above with respect to claim 5).
Grbovic further teaches generating an executable sequential machine learning model from the trained sequential machine learning model (Page 9, Col. 2, Section 6, ¶1 “we employed neural language models capable of learning product embeddings for product-to-product predictions… Several variants of the prediction models were tested offline and the best candidate was chosen for an online bucket test. Following the encouraging bucket test results, we launched the system in production”).
Grbovic further teaches using the executable sequential machine learning model to generate one or more recommendations of content items (Page 7, Col. 1, Section 4.4, ¶1 “we experiment with recommending products to users,” Page 9, Col. 2, ¶3 “Given predictions for a certain user, once the user logs into the e-mail client we show a new recommendation after every user action”).
Grbovic further teaches providing the one or more content items to a user device associated with the user for presentation (Page 9, Col. 2, ¶3 “The product ads are implemented in the so-called “pencil” ad position, just above the first e-mail in the inbox,” Page 2, Fig. 1 depicts providing the one or more content items to an implicit personal user device associated with the user for presentation).
Grbovic, Kang, Paraschakis, and Chen are analogous art to the claimed invention as all are from the same field of endeavor of recommender systems. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the SASRec architecture of Kang, the train-test split method of Paraschakis, and the user embedding of Chen with the methodology of Grbovic. The motivation to do so is to design a recommender system that “outperforms various state-of-the-art sequential models… on both sparse and dense datasets” (Kang, Abstract), to improve the quality of predictions in recommender systems (Paraschakis, Page 1029, Col. 2, ¶2 “the adoption of time-aware recommenders in e-commerce, which have shown improved quality of predictions”), and to improve user satisfaction based on their previous actions (Chen, Page 457, Col. 2, Section 3, ¶1 “For each user, we consider a sequence of user historical interactions… Given such a sequence, we predict the next action to take… so that user satisfaction metrics… indicated by clicks or watch time, improve”).
Regarding claim 17, Grbovic in view Kang and further in view of Paraschakis, and further in view of Chen teaches the computer-implemented method of claim 16 (and thus the rejection of claim 16 is incorporated).
Kang teaches wherein training the sequential machine learning model includes: generating, by the sequential machine learning model, a plurality of embeddings that correspond to the first plurality of user actions provided to the sequential machine learning model (Page 3, Col. 1, Table 1: “M = item embedding matrix… Ê = input embedding matrix,” Section III, ¶1 “we are given a user’s action sequence Su… and seek to predict the next item. During the training process, at time step t, the model predicts the next item depending on the previous t items,” Section A, ¶1-2 “We transform the training sequence… into a fixed-length sequence s = (s1, s2, . . ., sn)… We create an item embedding matrix M… and retrieve the input embedding matrix E… where Ei = Msi… we inject a learnable position embedding… into the input embedding,” (1) depicts the “fixed-length sequence” of “a user’s action sequence” in the “input embedding matrix” as “Ms1, Ms2 …” or a plurality of embeddings that correspond to the first plurality of user actions, Page 6, Col. 1, Section A, ¶2 “we split the historical sequence Su for each user u… and (3) all remaining actions,” Page 1, Col. 1, Fig. 1 – “Training Action Sequence” and “Embedding Layer”).
Kang further teaches determining a subset of the plurality of embeddings (Page 3, Col. 1, Section III, ¶1 “we are given a user’s action sequence S,” Section A, ¶1 “We transform the training sequence… into a fixed length sequence s… where n represents the maximum length that our model can handle. If the sequence length is greater than n, we consider the most recent n actions”).
Kang further teaches and training the sequential machine learning model to predict a respective user action for one or more embeddings of the subset of the plurality of embeddings (Page 3, Col. 1, Section III, ¶1 “During the training process, at time step t, the model predicts the next item depending on the previous t items,” Section A, ¶1 “n represents the maximum length our model can handle”).
Grbovic and Kang are analogous art to the claimed invention as both are from the same field of endeavor of recommender systems. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the SASRec architecture of Kang with the methodology of Grbovic. The motivation to do so is to design a recommender system that “outperforms various state-of-the-art sequential models… on both sparse and dense datasets” (Kang, Abstract).
Claim 18 is rejected under 35 U.S.C. 103 as being unpatentable over Grbovic in view of Kang and further in view of Paraschakis, and further in view of Chen, and further in view Rang et al. (“Data Life Aware Model Updating Strategy for Stream-Based Online Deep Learning,” 2021, hereinafter Rang).
Regarding claim 18, Grbovic in view Kang and further in view of Paraschakis, and further in view of Chen teaches the computer-implemented method of claim 16 (and thus the rejection of claim 16 is incorporated).
Grbovic fails to teach further comprising: updating the sequential machine learning model using a second sequence of user actions by using the second sequence of user actions to re-train an initially trained sequential machine learning model to generate a first updated sequential machine learning model. However, Rang, in the same field of endeavor, teaches this limitation (Page 5, Col. 1, Section 3.2.3, ¶1 “Let m data samples arrive in a sequence,” a machine learning model that receives a sequence of data encompasses a sequential learning model, given its broadest reasonable interpretation, Fig. 6, Page 7, Col. 1, Section 3.4, ¶1 “Fig. 6 depicts the detailed workflow of our proposed architecture. Three training model stages are included: Batch0, Batch1, and Batch2… During the model training procedure with data sample a0, some new data also arrived at the same time. If the current training is finished and its training metrics are profiled, a new data sample a1… for the next training procedure Batch1 is supposed to be built…” Page 3, Col. 2, ¶1 “The next model training Batch1 is triggered with a newly built data sample and model from the previous training procedure,” Fig. 6 depicts the same “Trained Model” from “Batch 0” being trained again or updated in “Batch 1” with a “Data Sample” containing “New Data” or updating the sequential machine learning model using a second sequence of user actions by using the second sequence of user actions to re-train an initially trained sequential machine learning model, wherein to generate a first updated sequential machine learning model is implicit after retraining the “Trained Model”).
Rang further teaches and subsequently updating the first updated sequential machine learning model using a third sequence of user action by using the third sequence of user actions to re-train the initially trained sequential machine learning model to generate a second updated sequential machine learning model (Fig. 6, Page 3, Col. 2, ¶1 “Batch1 trains the model and prepares data sample a2 for the next training stage Batch2. This training loop would continue until no more new data arrives or no improvement in model quality,” Fig. 6 depicts subsequently updating the first updated sequential machine learning model, in this case the “Trained Model” in “Batch 1,” by using a new “Data Sample” in “Batch 2” to train the same “Trained Model” as depicted “Batch 0” and “Batch 1,” which encompasses using a third sequence of actions by using the third sequence of user actions to re-train the initially trained sequential machine learning model, wherein to generate a second updated sequential machine learning model is implicit after retraining the “Trained Model” again on a different “Data Sample” from Batch 2).
Grbovic and Rang are analogous to the claimed invention as both are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the training workflow of Rang with the methodology of Grbovic. The motivation to do so is to design an online learning method that “improves model quality with continually generated data” (Rang, Page 1, Col. 2, Section 2.2, ¶1).
Claim 19 is rejected under 35 U.S.C. 103 as being unpatentable over Grbovic in view of Kang and further in view of Paraschakis, and further in view of Chen, and further in view of Blanco-Mallo et al. (“On the effectiveness of convolutional autoencoders on image-based personalized recommender systems,” 2020, hereinafter Blanco-Mallo).
Regarding claim 19, Grbovic in view Kang and further in view of Paraschakis, and further in view of Chen teaches the computer-implemented method of claim 16 (and thus the rejection of claim 16 is incorporated).
Grbovic teaches further comprising: determining a plurality of parameters associated with a plurality of users associated with the sequence of user actions (Page 5, Col. 2, Section 4.2, ¶1 “To get deeper insights into behavior of online users based on their demographic background and geographic location, we segregated users into cohorts based on their age, gender, and location, and looked at their purchasing habits,” wherein “age, gender, and location” encompass a plurality of parameters associated with a plurality of users associated with the sequence of user actions).
Grbovic fails to teach determining, based at least in part on the plurality of parameters, that the sequence of user actions is unbalanced with respect to at least one parameter of the plurality of parameters. However, Blanco-Mallo, in the same field of endeavor, teaches this limitation (Page 4, Col. 1, Section A, ¶1 “The data used in this work were collected in 2018 and 2019 from the TripAdvisor reviews published by users about restaurants in cities of different sizes,” wherein “TripAdvisor reviews published by users” encompasses sequence of user actions, Table III, Page 5, Col. 1, ¶1-2 “Table III shows the number of images per partition for the three datasets, including the ratios between positive and negative samples. As the datasets are highly unbalanced, a strategy must be applied to reduce its impact on the model performance,” wherein “Positive samples” or “Negative samples” for “cities of different sizes” encompasses based at least in part on the plurality of parameters).
Blanco-Mallo further teaches and at least one of up-sampling or down-sampling user actions of at least some of the plurality of users based at least in part on the at least one parameter, so as to balance the sequence of user actions with respect to the at least one parameter (Page 5, Col. 1, ¶2-3 “image data augmentation involves expanding the size of the train set by creating modified versions of the original images. The objective of this technique is not only to increase the amount of data available, but also their variability, thus improving the robustness of the learning models. Data augmentation can be applied to all the samples in the train set but… we only over-sampled the minority class…” Page 5, Col. 1, ¶1 “As a result, the imbalance problem is alleviated and the ratios between the positive and the negative classes are very close to 1:1 (1.2:1 for Santiago de Compostela, 1.02:1 for Barcelona, and 1.39:1 for New York)”).
Grbovic and Blanco-Mallo are analogous to the claimed invention as both are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to implement a method that combined the plurality of parameters of Grbovic with the up-sampling or down-sampling of Blanco-Mallo. The motivation to do so, as stated by Blanco-Mallo, is to design an image recommender system that “makes use of the context of the problem… works better than standard approaches that use a pre-trained convolutional neural network (CNN), and… is less computationally expensive that integrating and fine-tuning a CNN” (Blanco-Mallo, Page 2, Col. 1, ¶1).
Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Grbovic in view of Kang and further in view of Paraschakis, and further in view of Chen, and further in view of Solano Gomez (US 20220224683 A1, hereinafter Solano Gomez).
Regarding claim 20, Grbovic in view Kang and further in view of Paraschakis, and further in view of Chen teaches the computer-implemented method of claim 16 (and thus the rejection of claim 16 is incorporated).
Grbovic fails to teach further comprising: obtaining a plurality of labeled negative training data. However, Solano Gomez, in the same field of endeavor, teaches this limitation (Fig. 4 – 404, ¶71 “a plurality of positive and a plurality of negative pairs of biometric data are obtained from the set of biometric data associated with a plurality of users performing login sessions”).
Solano Gomez further teaches and providing the plurality of labeled negative training data to the sequential machine learning model (Fig. 3 – 306, ¶69 “the Siamese network may be trained with pairs of behavior data (e.g., to train for pairwise similarity) of prior users performing logins to a computer system,” ¶71 “training data is obtained from a set of biometric data associated with a plurality of users performing login sessions… a plurality of negative pairs of biometric data are obtained”).
Solano Gomez further teaches wherein: training the sequential machine learning model is further based on the plurality of labeled negative training data (Fig. 3 – 306, ¶60 “the machine learning model includes a Siamese neural network… the Siamese network may be trained by positive pairs of training data and negative pairs of training data”).
Regarding the limitation and the plurality of negative training data includes a portion of the second plurality of user actions that were a positive engagement for a different respective user, Kang teaches the second plurality of user actions (Page 6, Col. 1, Section A, ¶2 “we split the historical sequence Su for each user u… (1) the most recent action… (2) the second most recent action). However, Kang fails to teach and the plurality of negative training data includes a portion of the second plurality of user actions that were a positive engagement for a different respective user.
Solano Gomez teaches the plurality of negative training data includes a portion of user actions that were a positive engagement for a different respective user (¶60 “a positive pair of biometric data may be a pair of two login behaviors associated with the same user during different sessions to login an account. A negative pair of biometric data may be a pair of two login behaviors associated with two different users during different sessions to login accounts,” wherein “login behaviors associated with two different users” encompasses positive engagement, given its broadest reasonable interpretation of user engagement labeled as “positive” training data for a different user).
Grbovic, Kang, and Solano Gomez are analogous to the claimed invention as both are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the second plurality of user actions of Kang and the negative training data of Solano Gomez with the methodology of Grbovic. The motivation to do so is to design a recommender system that “outperforms various state-of-the-art sequential models… on both sparse and dense datasets” (Kang, Abstract) and is more efficient to train (Solano Gomez, ¶21 “Trained with a large set of prior other users' behaviors associated with logging into computer systems to learn similarity based on the features specific to the domain of user's login behaviors, the machine learning model does away with the conventionally required training with a particular user's historic login behavior to authenticate new login behaviors”).
Response to Amendment
In review of Applicant’s amendments, filed April 27, 2026, the objection to the drawing made in the previous office action has been withdrawn.
The rejection of claim 11 under 35 U.S.C. 112(b) set forth in the previous office action is withdrawn in view of the amendment to the claim.
The rejections of claims 1 and 4 under 35 U.S.C. 102(a)(2) set forth in the previous office action are withdrawn in view of the amendments to claim 1.
Response to Arguments
Applicant’s amendments and arguments, see page 1 filed April 27, 2026, regarding the abstract idea rejections from the previous office action made under 35 U.S.C. 101 have been considered but are not persuasive.
On page 2 of the Remarks, under section I, Applicant asserts that “while the claim elements may recite actions that involve mathematical concepts such as embedding vectors, the applicant submits that the claim does not recite operations or formulas.” Applicant cites as reasoning an excerpt from MPEP 2106.04 II(A)(1) cautioning examiners to “distinguish claims that recite an exception… and claims that merely involve an exception.” Examiner agrees with the assertion. However, claims 1 and 5 still recite “determining,” which can reasonably be interpreted as a mental process. Thus, the claim is not subject matter eligible.
On page 2 of the Remarks, under section II, Applicant argues that “Even if the independent claims are determined to recite a judicial exception… The applicant submits that the additional elements in the claims as a whole integrate the purported judicial exception into a practical application of the exception.” Applicant cites as reasoning excerpts from MPEP 2106.04(d) and “Reminders on evaluating subject matter eligibility of claims under 35 USC 101.” Applicant submits that the amended independent claims are patent eligible because they “recite at least the practical application of providing recommended content to a device associated with a user for presentation… This provides an improvement by using both the embedding information and contextual information to predict recommended images.” Examiner respectfully disagrees. Providing, to a device, recommended content that “is determined based on determining an updated user embedding that is representative of the user in view of contextual information and is configured to predict a plurality of recommended images” is still insignificant extra-solution activity (see MPEP 2106.05(g) “all uses of the recited judicial exception require such data gathering or data output”) of necessary data gathering (inputting contextual information and an embedding into a model that performs the judicial exception) and outputting (using the model to apply the judicial exception and output an updated embedding configured to predict recommended images that are then provided to a user device) that does not integrate the judicial exception into a practical application (see MPEP 2106.05(a)(II) “Examples that the courts have indicated may not be sufficient to show an improvement to technology include… iii. Gathering and analyzing information using conventional techniques and displaying the result, TLI Communications, 823 F.3d at 612-13, 118 USPQ2d at 1747-48”) or amount to significantly more (see MPEP 2106.05(d)(II) “Courts have held computer‐implemented processes not to be significantly more than an abstract idea (and thus ineligible) where the claim as a whole amounts to nothing more than generic computer functions merely used to implement an abstract idea, such as an idea that could be done by a human analog (i.e., by hand or by merely thinking)… Below are examples of other types of activity that the courts have found to be well-understood, routine, conventional activity when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity… iv. Presenting offers and gathering statistics, OIP Techs., 788 F.3d at 1362-63, 115 USPQ2d at 1092-93; v. Determining an estimated outcome and setting a price”).
In consideration of this conclusion, independent claims 1 and 5 and their associated dependent claims 2-4 and 6-15, respectively, are subject matter ineligible, and thus, the rejections under 35 U.S.C. 101 stand.
Applicant’s arguments, filed April 27, 2026 regarding the rejections from the previous office action made under 35 U.S.C. 102(a)(2) have been fully considered but are moot as they do not apply to the new references Pal, Chen, and Covington being used in the current rejection of claim 1 and its associated dependent claims 2-4, Kang, Chen, and Covington being used in the current rejection of claim 5 and its associated dependent claims 5-15, and Kang, Paraschakis, and Chen being used in the current rejection of claim 16 and its associated dependent claims 17-20, to teach the amended claim limitations directed to recommending images based on user interactions and providing said recommendations to a user device.
On page 4 of the Remarks, Applicant asserts that “a set of application scores does not disclose an embedding that is representative of the user and is configured to predict a plurality of user actions associated with the user where the plurality of predicted user actions including determining images with which the user is predicted to interact.” Examiner agrees with the assertion. However, the combination of the new references Chen and Pal disclose a user embedding that captures “evolving user interests using a n-dimensional vector” (Chen, Page 458, Col. 2, Section 4.1, ¶1) which is configured to determine videos that a user is predicted to interact with given a softmax policy and an image recommender system “that powers personalized recommendation at Pinterest” (Pal, Page 9, Col. 1, Section 7, ¶1), respectively. The combination of the new references explicitly describes an embedding that is representative of the user and is configured to determine images with which the user is predicted to interact.
Additionally, the new reference Covington explicitly teaches the limitation of claim 1 directed to determining “the second user embedding” as opposed to the list of top-K applications of Zhao cited in the previous office action.
On page 6 of the Remarks, Applicant asserts that “three different embeddings are generated: i) embeddings corresponding to one or more user actions, ii) a plurality of embeddings from the corresponding embeddings determined for the one or more user actions, and iii) a user embedding based at least in part on the correspond predicted action (determined for the plurality of embeddings).” Applicant argues that “a list of relevant applications does not correspond to a user embedding,” and that even if it does, it at most “corresponds to the user embedding already identified in Ulanov. Consequently, even in combination the portions of Ulanov and Zhao fail to disclose or suggest the combination of the corresponding embeddings for one or more user actions of the first sequence of user actions, the plurality of embeddings determined from the corresponding embeddings, and the user embedding that is representative of the user.” Examiner agrees with the assertion. However, the combination of the new references Kang and Chen disclose generating (i) embeddings corresponding to one or more user actions, or the “input embedding matrix” (Kang, Page 3, Col. 1, Table 3), (ii) a plurality of embeddings from the corresponding embeddings, or “item embeddings after the b-th feed forward network” (Kang, Page 3, Col. 1, Table 3), and (iii) a user embedding based at least in part on corresponding predicted actions, or “a n-dimensional vector” (Chen, Page 458, Col. 2, Section 4.1, ¶1) that continuously updates when users interact with recommended videos. The combination of the new references explicitly describes the three different embeddings, as opposed to the content and user embeddings of Ulanov and the list of relevant applications of Zhao.
Additionally, the new reference Covington explicitly teaches the limitation of claim 9 directed to determining “the second user embedding” as opposed to the list of top-K applications of Zhao cited in the previous office action.
On page 7 of the Remarks, Applicant argues that “Grbovic simply describes evaluating predictions based on a month of user purchases that was… excluded from the training. Thus, the predictions can be compared to the held out purchase data. The portion does not disclose or suggest dividing the first sequence of user actions into a first plurality of user actions ‘that were performed prior to the point in time and a second plurality of user actions that were performed after the point in time.’” Examiner agrees with the assertion. However, the new reference, Kang, uses “timestamps to determine the sequence order of actions… For partitioning, we split the historical sequence… for each user… into three parts: (1) the most recent action… for testing, (2) the second most recent action… for validation, and (3) all remaining actions for training” (Kang, Page 6, Col. 1, Section A, ¶2), which explicitly describes dividing a sequence of user actions into pluralities of actions that happen before and after a determined point in time.
Applicant further argues that the “chart shown in Fig. 7 illustrating how popular products can change frequently. See Grbovic page 6, section 4.3. This portion does not relate to the training data input to a sequential machine learning model and moreover does not relate to the separate user sequences divided into labeled and unlabeled training according to a point in time.” Examiner agrees with the assertion. However, the combination of the new references, Kang and Paraschakis, explicitly describe separating a user sequence into “all remaining actions for training” (Kang, Page 6, Col. 1, Section A, ¶2) input to a sequential learning model and dividing sequential data into labeled and unlabeled training according to a train-test split around “a split point on the dataset’s timeline” (Paraschakis, Page 1026, Col. 1, ¶3), as opposed to the frequently changing popular products of Grbovic.
Additionally, the new reference Chen replaces Lomada and explicitly teaches the limitation of claim 16 directed to generating user embeddings configured to predicting a plurality of user actions over a period of time for the corresponding users.
With the addition of the references Pal, Chen, Covington, Kang, and Paraschakis teaching the subject matter introduced in the amendments and the existing claims, claims 1 and 4 are now rejected under 35 U.S.C. 103, while the previous rejections under 35 U.S.C. 103 still stand.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to WILLIAM M LEE whose telephone number is (571)272-4761. The examiner can normally be reached Mon-Fri. 8am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Cesar Paula can be reached at (571)272-4128. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/WILLIAM MICHAEL LEE/
Examiner, Art Unit 2145
/CESAR B PAULA/Supervisory Patent Examiner, Art Unit 2145