Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 1-3, 6-15, 18, and 21-26 are presented for examination.
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on April 2, 2026 has been entered.
Claim Objections
Claims 11, 12, and 26 objected to because of the following informalities: “the group” should read “a group”.
Appropriate correction is required.
Claim Rejections - 35 USC § 112
Claims 1-3, 6-15, 18, and 21-26 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claims 1, 13, and 18 recite the limitation "the corresponding episodes,” which has insufficient antecedent basis in the claims. As currently drafted, it is unclear if “the corresponding episodes” used to generate augmented episode embeddings are the “first plurality of episodes,” or a different set of episodes. Claims 2-3, 6-12, 14-15, and 21-26 are rejected due to dependency on a rejected base claim.
Claim 12 recites the limitation “the plurality of features of the episode,” which has insufficient antecedent basis in the claim. Claim 1 from which claim 12 depends recites both “a plurality of features of the corresponding episodes” and “the plurality of features of the corresponding episode,” thus it is unclear what singular episode “the episode” is meant to refer to.
Claim Rejections - 35 USC § 101
Claims 1-3, 6-15, 18 and 21-26 are rejected under 35 U.S.C. 101 because the claimed invention is
directed to an abstract idea without significantly more. The analysis of the claims will follow the 2019 Revised Patent Subject Matter Eligibility Guidance (“2019 PEG”).
Regarding claim 1:
Subject Matter Eligibility Analysis Step 1:
Claim 1 recites "A method of recommending content to a user " thus it is a process,
one of the four statutory categories of patentable subject matter.
Subject Matter Eligibility Analysis Step 2A Prong 1:
Claim 1 recites the steps of:
“generating episode embeddings from features of at least a first plurality of episodes of content items available through the media-providing service”: This could encompass a human utilizing features of a first plurality of episodes of content items to generate episode embeddings. Thus, this is a mental process.
"using the feature-level augmentation to generate augmented
episode embeddings by masking subsets of the plurality of features
of the corresponding episodes": This could encompass a human masking subsets of features to generate augmented episode embeddings as a feature-level augmentation operation. Thus, this is a mental process.
"using the instance-level augmentation to identify a correlated
episode for each episode of at least a second plurality of episodes of content items available through the media-providing service and generating a correlated episode embedding for the correlated episode using user-interaction data for the correlated episode": This could encompass a human identifying a correlated episode for a given episode of a plurality of episodes, then generating an embedding for the correlated episode using user-interaction data. Thus, this is a mental process.
"generating… a user embedding based on a plurality of features of the user": This could encompass a human generating a user embedding using the features of a user via pen and paper. Thus, this is a mental process.
"generating… a respective episode embedding for each episode of a third plurality of episodes, each respective episode embedding based on the plurality of features of the corresponding episode": This could encompass a human taking each episode of a plurality of episodes and embedding each episode based on its features, thus this is a mental process.
"generating a respective similarity score for each episode of the
plurality of episodes, the respective similarity score corresponding to
a latent similarity between the user embedding and each respective
episode embedding": This could encompass a human mentally computing a similarity score/metric between the user embedding and each respective episode embedding. Thus, this is a mental process.
"ranking the third plurality of episodes in accordance with the respective similarity scores": This could encompass a human using the similarity scores for the third plurality of episodes to mentally rank the episodes, thus this is a mental process.
"recommending a highest ranked episode of the third plurality of episodes to the user": This could encompass a human taking the rankings and then recommending the top ranked episode to a user, therefore this is a mental process.
Claim 1 therefore recites abstract ideas.
Subject Matter Eligibility Analysis Step 2A Prong 2:
The judicial exception is not integrated into a practical application.
Claim 1 recites the additional elements:
"A method of recommending content to a user": This element does not integrate the abstract ideas above into a practical application because it merely recites a field of use (content recommendation) in which to apply a judicial exception (MPEP 2106.05(h)).
"training a recommender model for a media-providing service
using contrastive learning with feature-level augmentation and instance-level augmentation": This element does
not integrate the abstract ideas into a practical application because it amounts to mere instructions to apply the judicial exception (feature-level augmentation and instance-level augmentation) to a generically recited model training step (MPEP 2105.06(f)).
“training the recommender model using at least the episode embeddings, the augmented episode embeddings, and the correlated episode embeddings”: This element does not integrate the abstract ideas into a practical application because it amounts to mere instructions to apply the judicial exception (e.g. generating episode embeddings, augmented episode embeddings, and correlated episode embeddings) to a generically recited model training step (MPEP 2105.06(f)).
that the “generating… a user embedding” and “generating… a respective episode embedding” steps are performed “via the trained recommender model”: This element does not integrate the abstract ideas into a practical application because it recites mere instructions to apply a judicial exception on a generic computer programmed with a generic class of computer algorithms (MPEP 2106.05(f)).
Thus, claim 1 is directed to the abstract ideas.
Subject Matter Eligibility Analysis Step 2B:
The claim does not contain significantly more than the judicial exception. The analysis at this step mirrors that of Step 2A Prong 1 above. As an ordered whole, the claim is directed to a mentally performable process of generating episode embeddings, augmented episode embeddings, and correlated episode embeddings, generating a user embedding, generating a similarity score between a user embedding and each episode embedding, ranking the episodes, and recommending a highest ranked episode to the user. Nothing in the claim provides significantly more than this. As such, the claim is not patent eligible.
Regarding claim 2:
Subject Matter Eligibility Analysis Step 1:
Claim 2 is a process as in claim 1.
Subject Matter Eligibility Analysis Step 2A Prong 1:
In addition to the mental processes in claim 1, claim 2, recites:
"the third plurality of episodes consists of episodes with which the user has not previously interacted": This further expands on the mental processes in claim 1, by defining "the plurality of episodes", therefore this is a mental process.
Claim 2 therefore recites abstract ideas.
Subject Matter Eligibility Analysis Step 2A Prong 2:
This judicial exception is not integrated into a practical application. No further additional elements are recited, see analysis of claim 1.
Subject Matter Eligibility Analysis Step 2B:
The claim does not contain significantly more than the judicial exception. No further additional elements are recited, see analysis of claim 1.
Regarding claim 3:
Subject Matter Eligibility Analysis Step 1:
Claim 3 is a process as in claim 1
Subject Matter Eligibility Analysis Step 2A Prong 1:
Claim 3 recites the same mental processes as claim 1, therefore it recites abstract ideas.
Subject Matter Eligibility Analysis Step 2A Prong 2:
This judicial exception is not integrated into a practical application. In addition to the additional elements in claim 1, claim 3 recites:
"the trained recommender model is a two-tower model having a user
function and an episode function": This element does not integrate the
abstract ideas above into a practical application because it amounts to mere instructions to apply a judicial exception using a generic computer programmed with a generic class of computer algorithms (MPEP 2106.05(f)).
Therefore claim 3 is directed to the abstract ideas.
Subject Matter Eligibility Analysis Step 2B:
The claim does not contain significantly more than the judicial exception. The analysis at this step mirrors that of Step 2A Prong 2 above.
Regarding claim 6:
Subject Matter Eligibility Analysis Step 1:
Claim 6 is a process as in claim 1.
Subject Matter Eligibility Analysis Step 2A Prong 1:
In addition to the mental processes in claim 1, claim 6, recites:
"generating the correlated episode embedding comprises applying a
second feature-level augmentation to the features of the correlated
episode": This could encompass a human performing another feature-level augmentation using the features of the correlated episode to
generate an embedding for that correlated episode. Thus, this is a
mental process.
Claim 6 therefore recites abstract ideas.
Subject Matter Eligibility Analysis Step 2A Prong 2:
This judicial exception is not integrated into a practical application. No further additional elements are recited, see analysis of claim 1.
Subject Matter Eligibility Analysis Step 2B:
The claim does not contain significantly more than the judicial exception. No further additional elements are recited, see analysis of claim 1.
Regarding claim 7:
Subject Matter Eligibility Analysis Step 1:
Claim 7 is a process as in claim 1.
Subject Matter Eligibility Analysis Step 2A Prong 1:
In addition to the mental processes in claim 1, claim 7 recites:
"the correlated episode is identified using a semantic similarity
approach": This could encompass a human identifying a correlated
episode by parsing through the second plurality of episodes mentioned in claim 1 and performing a semantic similarity operation/algorithm. Thus, this is a mental process.
Claim 7 therefore recites abstract ideas.
Subject Matter Eligibility Analysis Step 2A Prong 2:
This judicial exception is not integrated into a practical application. No further additional elements are recited, see analysis of claim 1.
Subject Matter Eligibility Analysis Step 2B:
The claim does not contain significantly more than the judicial exception. No further additional elements are recited, see analysis of claim 1.
Regarding claim 8:
Subject Matter Eligibility Analysis Step 1:
Claim 8 is a process as in claim 7.
Subject Matter Eligibility Analysis Step 2A Prong 1:
In addition to the mental processes in claim 7, claim 8, recites:
"the semantic similarity approach comprises using a nearest
neighbor search": This could encompass a human identifying a correlated episode by mentally performing a nearest neighbor search algorithm. Thus, this is a mental process.
Claim 8 therefore recites abstract ideas.
Subject Matter Eligibility Analysis Step 2A Prong 2:
This judicial exception is not integrated into a practical application. No further additional elements are recited, see analysis of claim 7.
Subject Matter Eligibility Analysis Step 2B:
The claim does not contain significantly more than the judicial exception. No further additional elements are recited, see analysis of claim 7.
Regarding claim 9:
Subject Matter Eligibility Analysis Step 1:
Claim 9 is a process as in claim 1.
Subject Matter Eligibility Analysis Step 2A Prong 1:
In addition to the mental processes in claim 1, claim 9 recites:
"the correlated episode is identified using a knowledge graph
similarity approach": This could encompass a human identifying a
correlated episode by observing a knowledge graph and identifying a similar episode. Thus, this is a mental process.
Claim 9 therefore recites abstract ideas.
Subject Matter Eligibility Analysis Step 2A Prong 2:
This judicial exception is not integrated into a practical application. No further additional elements are recited, see analysis of claim 1.
Subject Matter Eligibility Analysis Step 2B:
The claim does not contain significantly more than the judicial exception. No further additional elements are recited, see analysis of claim 1.
Regarding claim 10:
Subject Matter Eligibility Analysis Step 1:
Claim 10 is a process as in claim 1.
Subject Matter Eligibility Analysis Step 2A Prong 1:
In addition to the mental processes in claim 1, claim 10 recites:
"the correlated episode is identified using a cosine similarity
approach": This encompasses using a cosine similarity formula to identify a correlated episode; thus, this is a mathematical calculation.
Claim 10 therefore recites abstract ideas.
Subject Matter Eligibility Analysis Step 2A Prong 2:
This judicial exception is not integrated into a practical application. No further additional elements are recited, see analysis of claim 1.
Subject Matter Eligibility Analysis Step 2B:
The claim does not contain significantly more than the judicial exception. No further additional elements are recited, see analysis of claim 1.
Regarding claim 11:
Subject Matter Eligibility Analysis Step 1:
Claim 11 is a process as in claim 1.
Subject Matter Eligibility Analysis Step 2A Prong 1:
In addition to the mental processes in claim 1, claim 11 recites:
"the plurality of features of the user include one or more of the
group consisting of: a gender, an age, a country, a language, a
recent topic liked, a streaming statistic, and a collaborative filtering
vector": This further expands on the mental processes in claim 1, by
defining "plurality of features of the user,” therefore this is a mental
process.
Claim 11 therefore recites abstract ideas.
Subject Matter Eligibility Analysis Step 2A Prong 2:
This judicial exception is not integrated into a practical application. No further additional elements are recited, see analysis of claim 1.
Subject Matter Eligibility Analysis Step 2B:
The claim does not contain significantly more than the judicial exception. No further additional elements are recited, see analysis of claim 1.
Regarding claim 12:
Subject Matter Eligibility Analysis Step 1:
Claim 12 is a process as in claim 1.
Subject Matter Eligibility Analysis Step 2A Prong 1:
In addition to the mental processes in claim 1, claim 12 recites:
"the plurality of features of the episode include one or more of the
group consisting of: a topic, a country, a language, a licensor, a publisher, a collaborative filtering vector, and a semantic
embedding": This further expands on the mental processes in claim 1, by defining "plurality of features of the episode,” therefore this is a mental process.
Claim 12 therefore recites abstract ideas.
Subject Matter Eligibility Analysis Step 2A Prong 2:
This judicial exception is not integrated into a practical application. No further additional elements are recited, see analysis of claim 1.
Subject Matter Eligibility Analysis Step 2B:
The claim does not contain significantly more than the judicial exception. No further additional elements are recited, see analysis of claim 1.
Regarding claim 13:
Subject Matter Eligibility Analysis Step 1:
Claim 13 recites "A computing device comprising: one or more processors…" thus it is a machine, one of the four statutory categories of patentable subject matter.
Subject Matter Eligibility Analysis Step 2A Prong 1:
Claim 13 recites the steps of:
“generating episode embeddings from features of at least a first plurality of episodes of content items available through the media-providing service”: This could encompass a human utilizing features of a first plurality of episodes of content items to generate episode embeddings. Thus, this is a mental process.
"using the feature-level augmentation to generate augmented
episode embeddings by masking subsets of the plurality of features
of the corresponding episodes": This could encompass a human masking subsets of features to generate augmented episode embeddings as a feature-level augmentation operation. Thus, this is a mental process.
"using the instance-level augmentation to identify a correlated
episode for each episode of at least a second plurality of episodes of content items available through the media-providing service and generating a correlated episode embedding for the correlated episode using user-interaction data for the correlated episode": This could encompass a human identifying a correlated episode for a given episode of a plurality of episodes, then generating an embedding for the correlated episode using user-interaction data. Thus, this is a mental process.
"generating… a user embedding based on a plurality of features of the user": This could encompass a human generating a user embedding using the features of a user via pen and paper. Thus, this is a mental process.
"generating… a respective episode embedding for each episode of a third plurality of episodes, each respective episode embedding based on the plurality of features of the corresponding episode": This could encompass a human taking each episode of a plurality of episodes and embedding each episode based on its features, thus this is a mental process.
"generating a respective similarity score for each episode of the
plurality of episodes, the respective similarity score corresponding to
a latent similarity between the user embedding and each respective
episode embedding": This could encompass a human mentally computing a similarity score/metric between the user embedding and each respective episode embedding. Thus, this is a mental process.
"ranking the third plurality of episodes in accordance with the respective similarity scores": This could encompass a human using the similarity scores for the third plurality of episodes to mentally rank the episodes, thus this is a mental process.
"recommending a highest ranked episode of the third plurality of episodes to the user": This could encompass a human taking the rankings and then recommending the top ranked episode to a user, therefore this is a mental process.
Claim 13 therefore recites abstract ideas.
Subject Matter Eligibility Analysis Step 2A Prong 2:
The judicial exception is not integrated into a practical application.
Claim 13 recites the additional elements:
" A computing device, comprising: one or more processors; memory; and one or more programs stored in the memory and configured for execution by the one or more processors, the one or more programs comprising instructions for…": This element does not integrate the abstract ideas above into a practical application because it amounts to mere instructions to apply a judicial exception using a generic computer (MPEP 2106.05(f)).
"training a recommender model for a media-providing service
using contrastive learning with feature-level augmentation and instance-level augmentation": This element does
not integrate the abstract ideas into a practical application because it amounts to mere instructions to apply the judicial exception (feature-level augmentation and instance-level augmentation) to a generically recited model training step (MPEP 2105.06(f)).
“training the recommender model using at least the episode embeddings, the augmented episode embeddings, and the correlated episode embeddings”: This element does not integrate the abstract ideas into a practical application because it amounts to mere instructions to apply the judicial exception (e.g. generating episode embeddings, augmented episode embeddings, and correlated episode embeddings) to a generically recited model training step (MPEP 2105.06(f)).
that the “generating… a user embedding” and “generating… a respective episode embedding” steps are performed “via the trained recommender model”: This element does not integrate the abstract ideas into a practical application because it recites mere instructions to apply a judicial exception on a generic computer programmed with a generic class of computer algorithms (MPEP 2106.05(f)).
Thus, claim 13 is directed to the abstract ideas.
Subject Matter Eligibility Analysis Step 2B:
The claim does not contain significantly more than the judicial exception. The analysis at this step mirrors that of Step 2A Prong 1 above. As an ordered whole, the claim is directed to a mentally performable process of generating episode embeddings, augmented episode embeddings, and correlated episode embeddings, generating a user embedding, generating a similarity score between a user embedding and each episode embedding, ranking the episodes, and recommending a highest ranked episode to the user. Nothing in the claim provides significantly more than this. As such, the claim is not patent eligible.
Regarding claim 14:
Claim 14 is a machine as in claim 13 and is subject-matter ineligible for the same reasons as
claim 2.
Regarding claim 15:
Claim 15 is a machine as in claim 13 and is subject-matter ineligible for the same reasons as
claim 3.
Regarding claim 18:
Subject Matter Eligibility Analysis Step 1:
Claim 18 recites "A non-transitory computer-readable storage medium…" thus it is an article of manufacture, one of the four statutory categories of patentable subject matter.
Subject Matter Eligibility Analysis Step 2A Prong 1:
Claim 18 recites the steps of:
“generating episode embeddings from features of at least a first plurality of episodes of content items available through the media-providing service”: This could encompass a human utilizing features of a first plurality of episodes of content items to generate episode embeddings. Thus, this is a mental process.
"using the feature-level augmentation to generate augmented
episode embeddings by masking subsets of the plurality of features
of the corresponding episodes": This could encompass a human masking subsets of features to generate augmented episode embeddings as a feature-level augmentation operation. Thus, this is a mental process.
"using the instance-level augmentation to identify a correlated
episode for each episode of at least a second plurality of episodes of content items available through the media-providing service and generating a correlated episode embedding for the correlated episode using user-interaction data for the correlated episode": This could encompass a human identifying a correlated episode for a given episode of a plurality of episodes, then generating an embedding for the correlated episode using user-interaction data. Thus, this is a mental process.
"generating… a user embedding based on a plurality of features of the user": This could encompass a human generating a user embedding using the features of a user via pen and paper. Thus, this is a mental process.
"generating… a respective episode embedding for each episode of a third plurality of episodes, each respective episode embedding based on the plurality of features of the corresponding episode": This could encompass a human taking each episode of a plurality of episodes and embedding each episode based on its features, thus this is a mental process.
"generating a respective similarity score for each episode of the
plurality of episodes, the respective similarity score corresponding to
a latent similarity between the user embedding and each respective
episode embedding": This could encompass a human mentally computing a similarity score/metric between the user embedding and each respective episode embedding. Thus, this is a mental process.
"ranking the third plurality of episodes in accordance with the respective similarity scores": This could encompass a human using the similarity scores for the third plurality of episodes to mentally rank the episodes, thus this is a mental process.
"recommending a highest ranked episode of the third plurality of episodes to the user": This could encompass a human taking the rankings and then recommending the top ranked episode to a user, therefore this is a mental process.
Claim 18 therefore recites abstract ideas.
Subject Matter Eligibility Analysis Step 2A Prong 2:
The judicial exception is not integrated into a practical application.
Claim 18 recites the additional elements:
"A non-transitory computer-readable storage medium storing one or more programs configured for execution by a computing device having one or more processors and memory, the one or more programs comprising instructions for…": This element does not integrate the abstract ideas above into a practical application because it amounts to mere instructions to apply a judicial exception using a generic computer (MPEP 2106.05(f)).
"training a recommender model for a media-providing service
using contrastive learning with feature-level augmentation and instance-level augmentation": This element does
not integrate the abstract ideas into a practical application because it amounts to mere instructions to apply the judicial exception (feature-level augmentation and instance-level augmentation) to a generically recited model training step (MPEP 2105.06(f)).
“training the recommender model using at least the episode embeddings, the augmented episode embeddings, and the correlated episode embeddings”: This element does not integrate the abstract ideas into a practical application because it amounts to mere instructions to apply the judicial exception (e.g. generating episode embeddings, augmented episode embeddings, and correlated episode embeddings) to a generically recited model training step (MPEP 2105.06(f)).
that the “generating… a user embedding” and “generating… a respective episode embedding” steps are performed “via the trained recommender model”: This element does not integrate the abstract ideas into a practical application because it recites mere instructions to apply a judicial exception on a generic computer programmed with a generic class of computer algorithms (MPEP 2106.05(f)).
Thus, claim 18 is directed to the abstract ideas.
Subject Matter Eligibility Analysis Step 2B:
The claim does not contain significantly more than the judicial exception. The analysis at this step mirrors that of Step 2A Prong 1 above. As an ordered whole, the claim is directed to a mentally performable process of generating episode embeddings, augmented episode embeddings, and correlated episode embeddings, generating a user embedding, generating a similarity score between a user embedding and each episode embedding, ranking the episodes, and recommending a highest ranked episode to the user. Nothing in the claim provides significantly more than this. As such, the claim is not patent eligible.
Regarding claim 21:
Claim 21 is a machine as in claim 13 and is subject-matter ineligible for the same reasons as claim 6.
Regarding claim 22:
Claim 22 is a machine as in claim 13 and is subject-matter ineligible for the same reasons as
claim 7.
Regarding claim 23:
Claim 23 is machine as in claim 22 and is subject-matter ineligible for the same reasons as claim 8.
Regarding claim 24:
Claim 24 is a machine as in claim 13 and is subject-matter ineligible for the same reasons as
claim 9.
Regarding claim 25:
Claim 25 is a machine as in claim 13 and is subject-matter ineligible for the same reasons as claim 10.
Regarding claim 26:
Claim 26 is a machine as in claim 13 and is subject-matter ineligible for the same reasons as
claim 11.
Claim Rejections - 35 USC § 103
Claims 1-3, 6-8, 10-15, 18, 21-23 and 25-26 are rejected under 35 U.S.C. 103 as being unpatentable over Wang et al. (“Exploring Heterogeneous Metadata for Video Recommendation with Two-tower Model”) (hereinafter “Wang”) in view of Yao et al. (“Self-supervised Learning for Large-scale Item Recommendations”) (hereinafter “Yao”), further in view of Dwibedi et al. (“With a Little Help from My Friends: Nearest-Neighbor Contrastive Learning of Visual Representations”) (hereinafter “Dwibedi”), and further in view of Tamborrino (“Introducing Natural Language Search for Podcast Episodes”).
Regarding claim 1, Wang discloses “A method of recommending content to a user, the method comprising…
training a recommender model for a media-providing service… (Wang, 1: “Online video services like Amazon Prime Video add cold-start contents, i.e. titles that do not have any watch history, on a regular basis to their catalog. Newly added titles are typically the most interesting for users due to the novelty factor. In fact, many users scroll the page endlessly to find new titles that pique their interest. Thus, it is important to match these highly desired titles to the right users as soon as possible” and Wang, 3.1, Architecture: “As the name implies, the model has two components: user and item towers” and Preference Prediction: “In the training process, we use cross-entropy to calculate the loss. In the prediction process, we will predict users’ preference scores on the set of videos and recommend top scoring titles”)
…generating a[n]… [item] embedding for… [an item] using user-interaction data for the… [item] (Wang, 3.2 Metadata: In the following section, we will elaborate on three types of metadata used in the item tower: categorical features, synopsis and cover art. Categorical features is a subset of metadata listed on the product page to characterize the videos. They include categorical features represented with a high-dimensional binary vector… • Popularity is based on the number of views in a certain period… By concatenating the vectors generated from each of the features above, we obtain a high dimensional vector for each of the title; Examiner notes that number of views corresponds to “user-interaction data”) …
generating, via the trained recommender model, a user embedding based on a plurality of features of the user (Wang 3.1, User Tower: “Users are represented by their watch histories in the training time period and some additional user-level features such as their country… With the user input, the watch history is firstly encoded with a position-aware attention layer to extract the sequential patterns. The user-level features are directly embedded with an MLP layer. Then the encoder concatenates watch history embeddings and user feature embeddings, passing them through the three residual blocks” and Fig. 2: “User Embedding”);
generating, via the trained recommender model, a respective… [item] embedding for each… [item] of a third plurality of… [items], each respective… [item] embedding based on the plurality of features of the corresponding… [item] (Wang 3.1, Item Tower: “For each item appearing in the training period, we can use so-called ID feature that is in essence the one-hot encoding for all known items. The embedding layer of the ID feature is uniquely linked to the ID of the video and is updated during the training process. Meanwhile, each item also has other metadata available (genres, actors, directors, etc; synopsis, cover art image), which can be used to generate a dense item representation” and Fig. 2: “Item Embedding” and 4.1, Data: For the test set, we collect user streaming history in 2-year as input features to predict what user is going to watch in the next 7-day (Time Period 𝑌′). To evaluate how the proposed model works for cold-start title recommendation we did a series of offline experiments. Here we have a set of items that have no watch history in time period 𝑋, 𝑌 or 𝑋′ but have at least one watch in time period 𝑌′. There are 360 movies and 75 TV series in total in this category. In the following experiments, during the scoring/evaluation, we only calculate preference scores for this set of cold-start titles and rank among them; Examiner notes that the test set of items corresponds to “a third plurality of items”);
generating a respective similarity score for each… [item] of the third plurality of… [items], the respective similarity score corresponding to a latent similarity between the user embedding and each respective… [item] embedding (Wang 3.1, Preference Prediction: “Given a user-item pair (𝑢,𝑖), we can generate the user embedding u and item embedding i with the corresponding towers. Then following the idea of matrix factorization [10], we can use the dot product u · i to approximate 𝑢’s preference on item 𝑖. Note that we also apply the Sigmoid function 𝜎(·) on the dot product u · i as activation... In the prediction process, we will predict users’ preference scores on the set of videos and recommend top scoring titles; Examiner notes that the dot product of the embeddings corresponds to a latent similarity which determines the preference score);
ranking the third plurality of… [items] in accordance with the respective similarity scores (Wang 4.1, Evaluation methodology and metrics: “For each user, we calculate preference scores for a list of candidate movies or TV shows. We rank these titles based on the predicted scores…”); and
recommending a highest ranked… [item] of the third plurality of… [items] to the user (Wang 4.1, Evaluation methodology and metrics: “For each user, we calculate preference scores for a list of candidate movies or TV shows. We rank these titles based on the predicted scores and select the Top-K titles for different categories (i.e.,movies and series). We pick 𝐾 = 6 to simulate the use cases in Amazon Prime Video. With the top-k titles and the ground-truth, we adopt 4 different metrics to evaluate the recommendation performance”).
Wang does not appear to explicitly disclose the further limitations of the claim.
However, Yao discloses “training a recommender model… using contrastive learning with feature-level augmentation (Yao, 1 Introduction, Paragraph 7: “In this paper, we propose to leverage self-supervised learning based auxiliary tasks to improve item representations, especially with long-tail distributions and sparse data. Different from CV or NLU applications, input space of recommendation model is highly sparse and represented by a set of categorical features (e.g. item ids) with large cardinality. For such sparse models, we propose a new SSL framework, where the key idea is to: (i) augment data by masking input information; (ii) encode each pair of augmented examples by a two-tower DNN; and (iii) apply a contrastive loss to learn representations of augmented data”) …including:
generating… [item] embeddings from features of at least a first plurality of… [items] (Yao, 3.3 Multi-task Training, Loss for Main Task: “In detail, let q𝑖, x𝑖 be the embeddings of query and item examples (𝑞𝑖,𝑥𝑖) after being encoded by two neural networks, then for a batch of pairs
{
(
q
i
,
x
i
)
}
i
=
1
N
and temperature 𝜏, the batch softmax cross entropy loss is…” and Yao, Figure 3: item features-> embedding);
using the feature-level augmentation to generate augmented… [item] embeddings by masking subsets of a plurality of features of the corresponding… [items] (Yao, 3.1 Framework, Paragraph 2: “We consider a batch of 𝑁 item examples 𝑥1,...,𝑥𝑁, where 𝑥𝑖 ∈ X represents a set of features for example 𝑖. In the context of recommenders, an example indicates a query, an item or a query-item pair” and 3.2 A Two-Stage Data Augmentation: “We introduce the data augmentation, i.e., ℎ and 𝑔 in Figure 2. Given a set of item features, the key idea is to create two augmented examples by masking part of the information… The two-stage augmentation includes: • Masking. Apply a masking pattern on the set of item features. We use a default embedding in the input layer to represent the features that are masked” and Yao, Figure 3: augmented item features-> embedding) … and
training the recommender model using at least the… [item] embeddings, the augmented… [item] embeddings…” (Yao, Figure 3: item features-> embedding-> supervised loss and augmented item features-> embedding-> self-supervised loss and Yao 3.3 Multi-task Training: “To enable SSL learned representations to help improve the learning for the main supervised task such as regression or classification, we leverage a multi-task training strategy where the main supervised task and the auxiliary SSL task are jointly optimized. Precisely, let {(𝑞𝑖,𝑥𝑖)} be a batch of query-item pairs sampled from the training data distribution D𝑡𝑟𝑎𝑖𝑛, and let {𝑥𝑖} be a batch of items sampled from an item distribution D𝑖𝑡𝑒𝑚. Then the joint loss is: [see eq (5)]”; Examiner notes that the item embeddings contribute to the supervised loss, the augmented item embeddings contribute to the self-supervised loss, and the model is trained with a joint loss that combines the two losses).
Yao and the instant application both relate to machine learning for item recommendation and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified Wang with the teachings of Yao to include training the recommender model for a media-providing service using contrastive learning with feature-level augmentation, generating item embeddings from features of at least a first plurality of items, using the feature-level augmentation to generate augmented item embeddings by masking subsets of a plurality of features of the corresponding items, and training the recommender model using at least the item embeddings and the augmented item embeddings, and one would have been motivated to do so, as doing so would allow for learning better latent relationships of item features, improving item representation learning and improving generalization (see Yao, Abstract, Paragraph 2).
Neither Wang nor Yao appear to explicitly disclose the further limitations of the claim.
However, Dwibedi discloses “at a computing device having one or more processors and memory (Dwibedi, 4.4, Compute overhead: “We find increasing the size of the queue results in improved performance but this improvement comes at a cost of additional memory and compute required while training. In Table 9 we show how queue scaling with d = 256 affects memory required during training and number of training steps per second. With a support size of about 98k elements we require a modest 100 MB more in memory”): training a… model… using contrastive learning with… instance-level augmentation (Dwibedi, Abstract: “Self-supervised learning algorithms based on instance discrimination train encoders to be invariant to pre-defined transformations of the same instance. While most methods treat different views of the same image as positives for a contrastive loss, we are interested in using positives from other instances in the dataset. Our method, Nearest Neighbor Contrastive Learning of visual Representations (NNCLR), samples the nearest neighbors from the dataset in the latent space, and treats them as positives”; Examiner notes that using other instances in the dataset as positives for contrastive learning corresponds to “instance-level augmentation”), including:
using the instance-level augmentation to identify a correlated… [item] for each… [item] of at least a second plurality of… [items] and generating a correlated… [item] embedding for the correlated… [item] (Dwibedi, 3.2: Nearest Neighbor CLR (NNCLR): “In order to increase the richness of our latent representation and go beyond single instance positives, we propose using nearest-neighbours to obtain more diverse positive pairs. This requires keeping a support set of embeddings which is representative of the full data distribution. SimCLR uses two augmentations (zi, zi+) to form the positive pair. Instead, we propose using zi’s nearest neighbor in the support set Q to form the positive pair. In Figure 2 we visualize this process schematically…” and Figure 2; Examiner notes that, in Figure 2, the mini-batch of images corresponds to “a second plurality of items,” the embedding in support set Q that is zi’s nearest neighbor corresponds to a generated “correlated item embedding,” and the item that the embedding corresponds to (e.g. the turtle image outlined in red) corresponds to an identified “correlated item”); and
training the… model using at least the… correlated… [item] embeddings” (Dwibedi, 3.2: Nearest Neighbor CLR (NNCLR): “Similar to SimCLR we obtain the negative pairs from the mini-batch and utilize a variant of the InfoNCE loss (1) for contrastive learning. Building upon the SimCLR objective (2) we define NNCLR loss as below: [see eq (3)] where NN(z,Q) is the nearest neighbor operator as defined below: [see eq(4)]”).
Dwibedi and the instant application both relate to contrastive learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Wang/Yao with the teachings to Dwibedi to include performing the method at a computing device having one or more processors and memory, training the recommender model for a media-providing service using contrastive learning with instance-level augmentation, using the instance-level augmentation to identify a correlated item for each item of at least a second plurality of items and generating a correlated item embedding for the correlated item using user-interaction data for the correlated item (note that generating an embedding using user-interaction data was taught by Wang as shown above), and training the recommender model using at least the correlated item embeddings, and one would have been motivated to do so, as doing so would provide more semantic variations than pre-defined transformations, leading to improved performance of self-supervised representation learning (see Dwibedi, Abstract).
Neither Wang, Yao, nor Dwibedi appear to explicitly disclose the further limitations of the claim.
However, Tamborrino discloses “generating episode embeddings from features of at least a first plurality of episodes of content items available through… [a] media-providing service” (Tamborrino, Technical solution: “We are using a machine learning technique called Dense Retrieval, which consists of training a model that produces query and episode vectors in a shared embedding space… for episodes, we use a concatenation of textual metadata fields of the episode such as its title, description, its parent podcast show’s title and description, and so on” and Preparing the data: “Once we have this powerful pre-trained Transformer model, we need to fine-tune it on our target task of performing Natural Language Search on Spotify’s podcast episodes”).
Tamborrino and the instant application both relate to machine learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Wang/Yao/Dwibedi with Tamborrino such that the “items” are instead “episodes of content items available through the media-providing service” and the “item embeddings” are instead “episode embeddings,” and one would have been motivated to do so, as doing so would result in an increase in podcast engagement (see Tamborrino, Conclusion and future works).
Regarding claim 2, the rejection of claim 1 is incorporated. Wang as modified by Yao, Dwibedi, and Tamborrino further discloses “wherein the third plurality of episodes consists of episodes with which the user has not previously interacted” (Yang, 4.1, Data: “For the test set, we collect user streaming history in 2-year as input features to predict what user is going to watch in the next 7-day (Time Period 𝑌′). To evaluate how the proposed model works for cold-start title recommendation we did a series of offline experiments. Here we have a set of items that have no watch history in time period 𝑋, 𝑌 or 𝑋′ but have at least one watch in time period 𝑌′. There are 360 movies and 75 TV series in total in this category. In the following experiments, during the scoring/evaluation, we only calculate preference scores for this set of cold-start titles and rank among them”).
Regarding claim 3, the rejection of claim 1 is incorporated. Wang as modified by Yao, Dwibedi, and Tamborrino further discloses “wherein the trained recommender model is a two-tower model having a user function and an episode function” (Yang, 3 Two-Tower Model, 3.1 Architecture: “As the name implies, the model has two components: user and item towers, each producing the corresponding embeddings, culminating in a dot product between the two and passing through a sigmoid activation function”).
Regarding claim 6, the rejection of claim 1 is incorporated. Wang as modified by Yao, Dwibedi, and Tamborrino further discloses “wherein generating the correlated episode embedding comprises applying a second feature-level augmentation to the features of the correlated episode” (Dwibedi, 3.2, Support set: “We update it [the support set] at the end of each training step by taking the n (batch size) embeddings from the current training step and concatenating them at the end of the queue. We discard the oldest n elements from the queue. We only use embeddings from one view to update the support set” and 3.1, paragraph 2: “Formally, given a mini-batch of images {x1,x2..,xn}, two different random augmentations (or views) are generated for each image xi”; Examiner notes that the correlated episode embeddings come from the support set (see rejection of claim 1), and the random augmentation used to generate the view that is then embedded corresponds to “a second feature-level augmentation”).
Dwibedi and the instant application both relate to contrastive learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Wang/Yao/Tamborrino with the teachings of Dwibedi such that generating the correlated episode embedding comprises applying a second feature-level augmentation to the features of the correlated episode, and one would have been motivated to do so, as doing so would provide more semantic variations, leading to improved performance of self-supervised representation learning (see Dwibedi, Abstract).
Regarding claim 7, the rejection of claim 1 is incorporated. Wang as modified by Yao, Dwibedi, and Tamborrino further discloses “wherein the correlated episode is identified using a semantic similarity approach” (Dwibedi, Figure 1: “NNCLR Training. We propose a simple self-supervised learning method that uses similar examples from a support set as positives in a contrastive loss” and 3.2, Nearest Neighbor CLR (NNCLR): “Instead, we propose using zi’s nearest neighbor in the support set Q to form the positive pair”).
Dwibedi and the instant application both relate to contrastive learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Wang/Yao/Tamborrino with the teachings to Dwibedi such that the correlated episode is identified using a semantic similarity approach, and one would have been motivated to do so, as doing so would provide more semantic variations than pre-defined transformations, leading to improved performance of self-supervised representation learning (see Dwibedi, Abstract).
Regarding claim 8, the rejection of claim 7 is incorporated. Wang as modified by Yao, Dwibedi, and Tamborrino further discloses “wherein the semantic similarity approach comprises using a nearest neighbor search” (Dwibedi, 3.2: Nearest Neighbor CLR (NNCLR): “Instead, we propose using zi’s nearest neighbor in the support set Q to form the positive pair”).
Dwibedi and the instant application both relate to contrastive learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Wang/Yao/Tamborrino with the teachings to Dwibedi such that the semantic similarity approach comprises using a nearest neighbor search, and one would have been motivated to do so, as doing so would provide more semantic variations than pre-defined transformations, leading to improved performance of self-supervised representation learning (see Dwibedi, Abstract).
Regarding claim 10, the rejection of claim 1 is incorporated. Wang as modified by Yao, Dwibedi, and Tamborrino further discloses “wherein the correlated episode is identified using a cosine similarity approach” (Tamborrino, Training: “Cosine similarity is used as the similarity measure between a query vector and an episode vector”).
Tamborrino and the instant application both relate to machine learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Wang/Yao/Dwibedi with Tamborrino such that the correlated episode is identified using a cosine similarity approach, and one would be motivated to do so, as doing so would allow for efficiently retrieving episodes whose vectors are closest to the query vector (see Tamborrino, Technical solution).
Regarding claim 11, the rejection of claim 1 is incorporated. Wang as modified by Yao, Dwibedi, and Tamborrino further discloses “wherein the plurality of features of the user include one or more of the group consisting of: a gender, an age, a country, a language, a recent topic liked, a streaming statistic, and a collaborative filtering vector” (Wang, 3.1, User Tower: “Users are represented by their watch histories in the training time period and some additional user-level features such as their country”).
Regarding claim 12, the rejection of claim 1 is incorporated. Wang as modified by Yao, Dwibedi, and Tamborrino further discloses “wherein the plurality of features of the episode include one or more of the group consisting of: a topic, a country, a language, a licensor, a publisher, a collaborative filtering vector, and a semantic embedding” (Wang, 3.2, Metadata: “In the following section, we will elaborate on three types of metadata used in the item tower: categorical features, synopsis and cover art. Categorical features is a subset of metadata listed on the product page to characterize the videos. They include categorical features represented with a high-dimensional binary vector • Genre is used to indicate the topic of a title… • Country of Origin containing 159 different values, indicating the country/region title is created in”).
Regarding claim 13, Wang discloses
…training a recommender model for a media-providing service… (Wang, 1: “Online video services like Amazon Prime Video add cold-start contents, i.e. titles that do not have any watch history, on a regular basis to their catalog. Newly added titles are typically the most interesting for users due to the novelty factor. In fact, many users scroll the page endlessly to find new titles that pique their interest. Thus, it is important to match these highly desired titles to the right users as soon as possible” and Wang, 3.1, Architecture: “As the name implies, the model has two components: user and item towers” and Preference Prediction: “In the training process, we use cross-entropy to calculate the loss. In the prediction process, we will predict users’ preference scores on the set of videos and recommend top scoring titles”)
…generating a[n]… [item] embedding for… [an item] using user-interaction data for the… [item] (Wang, 3.2 Metadata: In the following section, we will elaborate on three types of metadata used in the item tower: categorical features, synopsis and cover art. Categorical features is a subset of metadata listed on the product page to characterize the videos. They include categorical features represented with a high-dimensional binary vector… • Popularity is based on the number of views in a certain period… By concatenating the vectors generated from each of the features above, we obtain a high dimensional vector for each of the title; Examiner notes that number of views corresponds to “user-interaction data”) …
generating, via the trained recommender model, a user embedding based on a plurality of features of the user (Wang 3.1, User Tower: “Users are represented by their watch histories in the training time period and some additional user-level features such as their country… With the user input, the watch history is firstly encoded with a position-aware attention layer to extract the sequential patterns. The user-level features are directly embedded with an MLP layer. Then the encoder concatenates watch history embeddings and user feature embeddings, passing them through the three residual blocks” and Fig. 2: “User Embedding”);
generating, via the trained recommender model, a respective… [item] embedding for each… [item] of a third plurality of… [items], each respective… [item] embedding based on the plurality of features of the corresponding… [item] (Wang 3.1, Item Tower: “For each item appearing in the training period, we can use so-called ID feature that is in essence the one-hot encoding for all known items. The embedding layer of the ID feature is uniquely linked to the ID of the video and is updated during the training process. Meanwhile, each item also has other metadata available (genres, actors, directors, etc; synopsis, cover art image), which can be used to generate a dense item representation” and Fig. 2: “Item Embedding” and 4.1, Data: For the test set, we collect user streaming history in 2-year as input features to predict what user is going to watch in the next 7-day (Time Period 𝑌′). To evaluate how the proposed model works for cold-start title recommendation we did a series of offline experiments. Here we have a set of items that have no watch history in time period 𝑋, 𝑌 or 𝑋′ but have at least one watch in time period 𝑌′. There are 360 movies and 75 TV series in total in this category. In the following experiments, during the scoring/evaluation, we only calculate preference scores for this set of cold-start titles and rank among them; Examiner notes that the test set of items corresponds to “a third plurality of items”);
generating a respective similarity score for each… [item] of the third plurality of… [items], the respective similarity score corresponding to a latent similarity between the user embedding and each respective… [item] embedding (Wang 3.1, Preference Prediction: “Given a user-item pair (𝑢,𝑖), we can generate the user embedding u and item embedding i with the corresponding towers. Then following the idea of matrix factorization [10], we can use the dot product u · i to approximate 𝑢’s preference on item 𝑖. Note that we also apply the Sigmoid function 𝜎(·) on the dot product u · i as activation... In the prediction process, we will predict users’ preference scores on the set of videos and recommend top scoring titles; Examiner notes that the dot product of the embeddings corresponds to a latent similarity which determines the preference score);
ranking the third plurality of… [items] in accordance with the respective similarity scores (Wang 4.1, Evaluation methodology and metrics: “For each user, we calculate preference scores for a list of candidate movies or TV shows. We rank these titles based on the predicted scores…”); and
recommending a highest ranked… [item] of the third plurality of… [items] to the user (Wang 4.1, Evaluation methodology and metrics: “For each user, we calculate preference scores for a list of candidate movies or TV shows. We rank these titles based on the predicted scores and select the Top-K titles for different categories (i.e.,movies and series). We pick 𝐾 = 6 to simulate the use cases in Amazon Prime Video. With the top-k titles and the ground-truth, we adopt 4 different metrics to evaluate the recommendation performance”).
Wang does not appear to explicitly disclose the further limitations of the claim.
However, Yao discloses “training a recommender model… using contrastive learning with feature-level augmentation (Yao, 1 Introduction, Paragraph 7: “In this paper, we propose to leverage self-supervised learning based auxiliary tasks to improve item representations, especially with long-tail distributions and sparse data. Different from CV or NLU applications, input space of recommendation model is highly sparse and represented by a set of categorical features (e.g. item ids) with large cardinality. For such sparse models, we propose a new SSL framework, where the key idea is to: (i) augment data by masking input information; (ii) encode each pair of augmented examples by a two-tower DNN; and (iii) apply a contrastive loss to learn representations of augmented data”) …including:
generating… [item] embeddings from features of at least a first plurality of… [items] (Yao, 3.3 Multi-task Training, Loss for Main Task: “In detail, let q𝑖, x𝑖 be the embeddings of query and item examples (𝑞𝑖,𝑥𝑖) after being encoded by two neural networks, then for a batch of pairs
{
(
q
i
,
x
i
)
}
i
=
1
N
and temperature 𝜏, the batch softmax cross entropy loss is…” and Yao, Figure 3: item features-> embedding);
using the feature-level augmentation to generate augmented… [item] embeddings by masking subsets of a plurality of features of the corresponding… [items] (Yao, 3.1 Framework, Paragraph 2: “We consider a batch of 𝑁 item examples 𝑥1,...,𝑥𝑁, where 𝑥𝑖 ∈ X represents a set of features for example 𝑖. In the context of recommenders, an example indicates a query, an item or a query-item pair” and 3.2 A Two-Stage Data Augmentation: “We introduce the data augmentation, i.e., ℎ and 𝑔 in Figure 2. Given a set of item features, the key idea is to create two augmented examples by masking part of the information… The two-stage augmentation includes: • Masking. Apply a masking pattern on the set of item features. We use a default embedding in the input layer to represent the features that are masked” and Yao, Figure 3: augmented item features-> embedding) … and
training the recommender model using at least the… [item] embeddings, the augmented… [item] embeddings…” (Yao, Figure 3: item features-> embedding-> supervised loss and augmented item features-> embedding-> self-supervised loss and Yao 3.3 Multi-task Training: “To enable SSL learned representations to help improve the learning for the main supervised task such as regression or classification, we leverage a multi-task training strategy where the main supervised task and the auxiliary SSL task are jointly optimized. Precisely, let {(𝑞𝑖,𝑥𝑖)} be a batch of query-item pairs sampled from the training data distribution D𝑡𝑟𝑎𝑖𝑛, and let {𝑥𝑖} be a batch of items sampled from an item distribution D𝑖𝑡𝑒𝑚. Then the joint loss is: [see eq (5)]”; Examiner notes that the item embeddings contribute to the supervised loss, the augmented item embeddings contribute to the self-supervised loss, and the model is trained with a joint loss that combines the two losses).
Yao and the instant application both relate to machine learning for item recommendation and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified Wang with the teachings of Yao to include training the recommender model for a media-providing service using contrastive learning with feature-level augmentation, generating item embeddings from features of at least a first plurality of items, using the feature-level augmentation to generate augmented item embeddings by masking subsets of a plurality of features of the corresponding items, and training the recommender model using at least the item embeddings and the augmented item embeddings, and one would have been motivated to do so, as doing so would allow for learning better latent relationships of item features, improving item representation learning and improving generalization (see Yao, Abstract, Paragraph 2).
Neither Wang nor Yao appear to explicitly disclose the further limitations of the claim.
However, Dwibedi discloses “A computing device, comprising: one or more processors; memory; and one or more programs stored in the memory and configured for execution by the one or more processors, the one or more programs comprising instructions for (Dwibedi, 4.4, Compute overhead: “We find increasing the size of the queue results in improved performance but this improvement comes at a cost of additional memory and compute required while training. In Table 9 we show how queue scaling with d = 256 affects memory required during training and number of training steps per second. With a support size of about 98k elements we require a modest 100 MB more in memory”): training a… model… using contrastive learning with… instance-level augmentation (Dwibedi, Abstract: “Self-supervised learning algorithms based on instance discrimination train encoders to be invariant to pre-defined transformations of the same instance. While most methods treat different views of the same image as positives for a contrastive loss, we are interested in using positives from other instances in the dataset. Our method, Nearest Neighbor Contrastive Learning of visual Representations (NNCLR), samples the nearest neighbors from the dataset in the latent space, and treats them as positives”; Examiner notes that using other instances in the dataset as positives for contrastive learning corresponds to “instance-level augmentation”), including:
using the instance-level augmentation to identify a correlated… [item] for each… [item] of at least a second plurality of… [items] and generating a correlated… [item] embedding for the correlated… [item] (Dwibedi, 3.2: Nearest Neighbor CLR (NNCLR): “In order to increase the richness of our latent representation and go beyond single instance positives, we propose using nearest-neighbours to obtain more diverse positive pairs. This requires keeping a support set of embeddings which is representative of the full data distribution. SimCLR uses two augmentations (zi, zi+) to form the positive pair. Instead, we propose using zi’s nearest neighbor in the support set Q to form the positive pair. In Figure 2 we visualize this process schematically…” and Figure 2; Examiner notes that, in Figure 2, the mini-batch of images corresponds to “a second plurality of items,” the embedding in support set Q that is zi’s nearest neighbor corresponds to a generated “correlated item embedding,” and the item that the embedding corresponds to (e.g. the turtle image outlined in red) corresponds to an identified “correlated item”); and
training the… model using at least the… correlated… [item] embeddings” (Dwibedi, 3.2: Nearest Neighbor CLR (NNCLR): “Similar to SimCLR we obtain the negative pairs from the mini-batch and utilize a variant of the InfoNCE loss (1) for contrastive learning. Building upon the SimCLR objective (2) we define NNCLR loss as below: [see eq (3)] where NN(z,Q) is the nearest neighbor operator as defined below: [see eq(4)]”).
Dwibedi and the instant application both relate to contrastive learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Wang/Yao with the teachings to Dwibedi to include performing the method at a computing device comprising: one or more processors; memory; and one or more programs stored in the memory and configured for execution by the one or more processors, the one or more programs comprising instructions, training the recommender model for a media-providing service using contrastive learning with instance-level augmentation, using the instance-level augmentation to identify a correlated item for each item of at least a second plurality of items and generating a correlated item embedding for the correlated item using user-interaction data for the correlated item (note that generating an embedding using user-interaction data was taught by Wang as shown above), and training the recommender model using at least the correlated item embeddings, and one would have been motivated to do so, as doing so would provide more semantic variations than pre-defined transformations, leading to improved performance of self-supervised representation learning (see Dwibedi, Abstract).
Neither Wang, Yao, nor Dwibedi appear to explicitly disclose the further limitations of the claim.
However, Tamborrino discloses “generating episode embeddings from features of at least a first plurality of episodes of content items available through… [a] media-providing service” (Tamborrino, Technical solution: “We are using a machine learning technique called Dense Retrieval, which consists of training a model that produces query and episode vectors in a shared embedding space… for episodes, we use a concatenation of textual metadata fields of the episode such as its title, description, its parent podcast show’s title and description, and so on” and Preparing the data: “Once we have this powerful pre-trained Transformer model, we need to fine-tune it on our target task of performing Natural Language Search on Spotify’s podcast episodes”).
Tamborrino and the instant application both relate to machine learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Wang/Yao/Dwibedi with Tamborrino such that the “items” are instead “episodes of content items available through the media-providing service” and the “item embeddings” are instead “episode embeddings,” and one would have been motivated to do so, as doing so would result in an increase in podcast engagement (see Tamborrino, Conclusion and future works).
Regarding claim 18, Wang discloses
…training a recommender model for a media-providing service… (Wang, 1: “Online video services like Amazon Prime Video add cold-start contents, i.e. titles that do not have any watch history, on a regular basis to their catalog. Newly added titles are typically the most interesting for users due to the novelty factor. In fact, many users scroll the page endlessly to find new titles that pique their interest. Thus, it is important to match these highly desired titles to the right users as soon as possible” and Wang, 3.1, Architecture: “As the name implies, the model has two components: user and item towers” and Preference Prediction: “In the training process, we use cross-entropy to calculate the loss. In the prediction process, we will predict users’ preference scores on the set of videos and recommend top scoring titles”)
…generating a[n]… [item] embedding for… [an item] using user-interaction data for the… [item] (Wang, 3.2 Metadata: In the following section, we will elaborate on three types of metadata used in the item tower: categorical features, synopsis and cover art. Categorical features is a subset of metadata listed on the product page to characterize the videos. They include categorical features represented with a high-dimensional binary vector… • Popularity is based on the number of views in a certain period… By concatenating the vectors generated from each of the features above, we obtain a high dimensional vector for each of the title; Examiner notes that number of views corresponds to “user-interaction data”) …
generating, via the trained recommender model, a user embedding based on a plurality of features of the user (Wang 3.1, User Tower: “Users are represented by their watch histories in the training time period and some additional user-level features such as their country… With the user input, the watch history is firstly encoded with a position-aware attention layer to extract the sequential patterns. The user-level features are directly embedded with an MLP layer. Then the encoder concatenates watch history embeddings and user feature embeddings, passing them through the three residual blocks” and Fig. 2: “User Embedding”);
generating, via the trained recommender model, a respective… [item] embedding for each… [item] of a third plurality of… [items], each respective… [item] embedding based on the plurality of features of the corresponding… [item] (Wang 3.1, Item Tower: “For each item appearing in the training period, we can use so-called ID feature that is in essence the one-hot encoding for all known items. The embedding layer of the ID feature is uniquely linked to the ID of the video and is updated during the training process. Meanwhile, each item also has other metadata available (genres, actors, directors, etc; synopsis, cover art image), which can be used to generate a dense item representation” and Fig. 2: “Item Embedding” and 4.1, Data: For the test set, we collect user streaming history in 2-year as input features to predict what user is going to watch in the next 7-day (Time Period 𝑌′). To evaluate how the proposed model works for cold-start title recommendation we did a series of offline experiments. Here we have a set of items that have no watch history in time period 𝑋, 𝑌 or 𝑋′ but have at least one watch in time period 𝑌′. There are 360 movies and 75 TV series in total in this category. In the following experiments, during the scoring/evaluation, we only calculate preference scores for this set of cold-start titles and rank among them; Examiner notes that the test set of items corresponds to “a third plurality of items”);
generating a respective similarity score for each… [item] of the third plurality of… [items], the respective similarity score corresponding to a latent similarity between the user embedding and each respective… [item] embedding (Wang 3.1, Preference Prediction: “Given a user-item pair (𝑢,𝑖), we can generate the user embedding u and item embedding i with the corresponding towers. Then following the idea of matrix factorization [10], we can use the dot product u · i to approximate 𝑢’s preference on item 𝑖. Note that we also apply the Sigmoid function 𝜎(·) on the dot product u · i as activation... In the prediction process, we will predict users’ preference scores on the set of videos and recommend top scoring titles; Examiner notes that the dot product of the embeddings corresponds to a latent similarity which determines the preference score);
ranking the third plurality of… [items] in accordance with the respective similarity scores (Wang 4.1, Evaluation methodology and metrics: “For each user, we calculate preference scores for a list of candidate movies or TV shows. We rank these titles based on the predicted scores…”); and
recommending a highest ranked… [item] of the third plurality of… [items] to the user (Wang 4.1, Evaluation methodology and metrics: “For each user, we calculate preference scores for a list of candidate movies or TV shows. We rank these titles based on the predicted scores and select the Top-K titles for different categories (i.e.,movies and series). We pick 𝐾 = 6 to simulate the use cases in Amazon Prime Video. With the top-k titles and the ground-truth, we adopt 4 different metrics to evaluate the recommendation performance”).
Wang does not appear to explicitly disclose the further limitations of the claim.
However, Yao discloses “training a recommender model… using contrastive learning with feature-level augmentation (Yao, 1 Introduction, Paragraph 7: “In this paper, we propose to leverage self-supervised learning based auxiliary tasks to improve item representations, especially with long-tail distributions and sparse data. Different from CV or NLU applications, input space of recommendation model is highly sparse and represented by a set of categorical features (e.g. item ids) with large cardinality. For such sparse models, we propose a new SSL framework, where the key idea is to: (i) augment data by masking input information; (ii) encode each pair of augmented examples by a two-tower DNN; and (iii) apply a contrastive loss to learn representations of augmented data”) …including:
generating… [item] embeddings from features of at least a first plurality of… [items] (Yao, 3.3 Multi-task Training, Loss for Main Task: “In detail, let q𝑖, x𝑖 be the embeddings of query and item examples (𝑞𝑖,𝑥𝑖) after being encoded by two neural networks, then for a batch of pairs
{
(
q
i
,
x
i
)
}
i
=
1
N
and temperature 𝜏, the batch softmax cross entropy loss is…” and Yao, Figure 3: item features-> embedding);
using the feature-level augmentation to generate augmented… [item] embeddings by masking subsets of a plurality of features of the corresponding… [items] (Yao, 3.1 Framework, Paragraph 2: “We consider a batch of 𝑁 item examples 𝑥1,...,𝑥𝑁, where 𝑥𝑖 ∈ X represents a set of features for example 𝑖. In the context of recommenders, an example indicates a query, an item or a query-item pair” and 3.2 A Two-Stage Data Augmentation: “We introduce the data augmentation, i.e., ℎ and 𝑔 in Figure 2. Given a set of item features, the key idea is to create two augmented examples by masking part of the information… The two-stage augmentation includes: • Masking. Apply a masking pattern on the set of item features. We use a default embedding in the input layer to represent the features that are masked” and Yao, Figure 3: augmented item features-> embedding) … and
training the recommender model using at least the… [item] embeddings, the augmented… [item] embeddings…” (Yao, Figure 3: item features-> embedding-> supervised loss and augmented item features-> embedding-> self-supervised loss and Yao 3.3 Multi-task Training: “To enable SSL learned representations to help improve the learning for the main supervised task such as regression or classification, we leverage a multi-task training strategy where the main supervised task and the auxiliary SSL task are jointly optimized. Precisely, let {(𝑞𝑖,𝑥𝑖)} be a batch of query-item pairs sampled from the training data distribution D𝑡𝑟𝑎𝑖𝑛, and let {𝑥𝑖} be a batch of items sampled from an item distribution D𝑖𝑡𝑒𝑚. Then the joint loss is: [see eq (5)]”; Examiner notes that the item embeddings contribute to the supervised loss, the augmented item embeddings contribute to the self-supervised loss, and the model is trained with a joint loss that combines the two losses).
Yao and the instant application both relate to machine learning for item recommendation and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified Wang with the teachings of Yao to include training the recommender model for a media-providing service using contrastive learning with feature-level augmentation, generating item embeddings from features of at least a first plurality of items, using the feature-level augmentation to generate augmented item embeddings by masking subsets of a plurality of features of the corresponding items, and training the recommender model using at least the item embeddings and the augmented item embeddings, and one would have been motivated to do so, as doing so would allow for learning better latent relationships of item features, improving item representation learning and improving generalization (see Yao, Abstract, Paragraph 2).
Neither Wang nor Yao appear to explicitly disclose the further limitations of the claim.
However, Dwibedi discloses “A non-transitory computer-readable storage medium storing one or more programs configured for execution by a computing device having one or more processors and memory, the one or more programs comprising instructions for (Dwibedi, 4.4, Compute overhead: “We find increasing the size of the queue results in improved performance but this improvement comes at a cost of additional memory and compute required while training. In Table 9 we show how queue scaling with d = 256 affects memory required during training and number of training steps per second. With a support size of about 98k elements we require a modest 100 MB more in memory”): training a… model… using contrastive learning with… instance-level augmentation (Dwibedi, Abstract: “Self-supervised learning algorithms based on instance discrimination train encoders to be invariant to pre-defined transformations of the same instance. While most methods treat different views of the same image as positives for a contrastive loss, we are interested in using positives from other instances in the dataset. Our method, Nearest Neighbor Contrastive Learning of visual Representations (NNCLR), samples the nearest neighbors from the dataset in the latent space, and treats them as positives”; Examiner notes that using other instances in the dataset as positives for contrastive learning corresponds to “instance-level augmentation”), including:
using the instance-level augmentation to identify a correlated… [item] for each… [item] of at least a second plurality of… [items] and generating a correlated… [item] embedding for the correlated… [item] (Dwibedi, 3.2: Nearest Neighbor CLR (NNCLR): “In order to increase the richness of our latent representation and go beyond single instance positives, we propose using nearest-neighbours to obtain more diverse positive pairs. This requires keeping a support set of embeddings which is representative of the full data distribution. SimCLR uses two augmentations (zi, zi+) to form the positive pair. Instead, we propose using zi’s nearest neighbor in the support set Q to form the positive pair. In Figure 2 we visualize this process schematically…” and Figure 2; Examiner notes that, in Figure 2, the mini-batch of images corresponds to “a second plurality of items,” the embedding in support set Q that is zi’s nearest neighbor corresponds to a generated “correlated item embedding,” and the item that the embedding corresponds to (e.g. the turtle image outlined in red) corresponds to an identified “correlated item”); and
training the… model using at least the… correlated… [item] embeddings” (Dwibedi, 3.2: Nearest Neighbor CLR (NNCLR): “Similar to SimCLR we obtain the negative pairs from the mini-batch and utilize a variant of the InfoNCE loss (1) for contrastive learning. Building upon the SimCLR objective (2) we define NNCLR loss as below: [see eq (3)] where NN(z,Q) is the nearest neighbor operator as defined below: [see eq(4)]”).
Dwibedi and the instant application both relate to contrastive learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Wang/Yao with the teachings to Dwibedi to include a non-transitory computer-readable storage medium storing one or more programs configured for execution by a computing device having one or more processors and memory, the one or more programs comprising instructions, training the recommender model for a media-providing service using contrastive learning with instance-level augmentation, using the instance-level augmentation to identify a correlated item for each item of at least a second plurality of items and generating a correlated item embedding for the correlated item using user-interaction data for the correlated item (note that generating an embedding using user-interaction data was taught by Wang as shown above), and training the recommender model using at least the correlated item embeddings, and one would have been motivated to do so, as doing so would provide more semantic variations than pre-defined transformations, leading to improved performance of self-supervised representation learning (see Dwibedi, Abstract).
Neither Wang, Yao, nor Dwibedi appear to explicitly disclose the further limitations of the claim.
However, Tamborrino discloses “generating episode embeddings from features of at least a first plurality of episodes of content items available through… [a] media-providing service” (Tamborrino, Technical solution: “We are using a machine learning technique called Dense Retrieval, which consists of training a model that produces query and episode vectors in a shared embedding space… for episodes, we use a concatenation of textual metadata fields of the episode such as its title, description, its parent podcast show’s title and description, and so on” and Preparing the data: “Once we have this powerful pre-trained Transformer model, we need to fine-tune it on our target task of performing Natural Language Search on Spotify’s podcast episodes”).
Tamborrino and the instant application both relate to machine learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Wang/Yao/Dwibedi with Tamborrino such that the “items” are instead “episodes of content items available through the media-providing service” and the “item embeddings” are instead “episode embeddings,” and one would have been motivated to do so, as doing so would result in an increase in podcast engagement (see Tamborrino, Conclusion and future works).
Regarding claim 21, the rejection of claim 13 is incorporated. Claim 21 is a device claim corresponding to method claim 6, thus the rejection of claim 21 follows the same rationale given for the rejection of claim 6 above.
Regarding claim 22, the rejection of claim 13 is incorporated. Claim 22 is a device claim corresponding to method claim 7, thus the rejection of claim 22 follows the same rationale given for the rejection of claim 7 above.
Regarding claim 23, the rejection of claim 22 is incorporated. Claim 23 is a device claim corresponding to method claim 8, thus the rejection of claim 23 follows the same rationale given for the rejection of claim 8 above.
Regarding claim 25, the rejection of claim 13 is incorporated. Claim 25 is a device claim corresponding to method claim 10, thus the rejection of claim 25 follows the same rationale given for the rejection of claim 10 above.
Regarding claim 26, the rejection of claim 13 is incorporated. Claim 26 is a device claim corresponding to method claim 11, thus the rejection of claim 26 follows the same rationale given for the rejection of claim 11 above.
Claims 9 and 24 are rejected under 35 U.S.C 103 as being unpatentable over Wang in view of Yao, Dwibedi, and Tamborrino, and further in view of Bhutani et al. (US11307881) (hereinafter “Bhutani”).
Regarding claim 9, the rejection of claim 1 is incorporated. Neither Wang, Yao, Dwibedi, nor Tamborrino appear to explicitly disclose the further limitations of the claim.
However, Bhutani discloses "[a] correlated… [item] is identified using a knowledge graph similarity approach" (Bhutani, Col 5, lines 40-57: “To generate the recommendation 126, the suggestion module 110 computes a dot product between the input embedding vectors and the knowledge graph embedding vectors… The suggestion module 110 applies the weight to the result of the dot product to determine scores for items that indicate similarity between the input embedding vectors and the knowledge graph embedding vectors. The recommendation 126 corresponds to an item having a highest score”).
Bhutani and the instant application both relate to machine learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Wang/Yao/Dwibedi/Tamborrino with Bhutani such that the correlated episode is identified using a knowledge graph similarity approach, and one would have been motivated to do so, as doing so would improve computational efficiency (see Bhutani, Col 3, lines 33-36).
Regarding claim 24, the rejection of claim 13 is incorporated. Claim 24 is a device claim corresponding to method claim 9, thus the rejection of claim 24 follows the same rationale given for the rejection of claim 9 above.
Response to Arguments
Applicant's arguments filed March 6, 2026 have been fully considered but they are, except insofar as rendered moot by the new grounds of rejection, not persuasive.
Regarding the remarks concerning rejections under 35 U.S.C. 101, Applicant argues that “amended claim 1 integrates any alleged mental process into a practical application, namely an improvement to computer technology by providing an improved way to train a model, using episode embeddings, augmented episode embeddings, and correlated episode embeddings, in circumstances in which training data is sparse,” thus reciting “a specific improvement to the process for training the claimed model” (Remarks, pages 9-10). Examiner respectfully disagrees. The claim language merely recites “training the recommender model using at least the episode embeddings, the augmented episode embeddings, and the correlated episode embeddings,” which amounts to mere instructions to apply an exception (i.e. the generating of the episode embeddings, augmented episode embeddings, and the correlated episode embeddings) to a generically recited training operation, as there are no specific steps or detail as to how this training is accomplished (see MPEP 2106.05(f)). Further, the increased size of the training dataset which is purported as the improvement comes from the mentally performable augmentation operations, and the judicial exception alone cannot provide the improvement (MPEP 2106.05(a)). Applicant further argues: (1) the claimed method is not how a human would go about providing recommendations in the absence of relevant data; and (2) a human is not capable of training a machine learning model. Regarding (1), the analysis as to whether a claim recites a mental process is whether it contains limitations that can be practically performed in the human mind, thus whether or not a human would perform this method is irrelevant to the analysis. Regarding (2), Examiner agrees that training a machine learning model is not mentally performable, however the training limitation amounts to mere instructions to apply an exception as analyzed above.
Regarding the remarks concerning rejections under 35 U.S.C. 103, the new grounds of rejection rely on Yao rather than Tan to teach the contested limitations, thus the arguments are rendered moot.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to GWYNEVERE A DETERDING whose telephone number is (571) 272-7657. The examiner can normally be reached Mon-Fri. 9am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kamran Afshar can be reached at (571) 272-7796. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000
/G.A.D./Examiner, Art Unit 2125
/KAMRAN AFSHAR/Supervisory Patent Examiner, Art Unit 2125