DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This office action is in responsive to RCE filed on 07/01/2026. Claims remain pending in the application. Claims 1, 9, and 17 are independent.
Claim Objections
Applicant's amendment to claims corrects previous objections; therefore, the previous objections are withdrawn.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 3-4, 7, 9, 11-12, 15, 17, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Flanagan (US 2022/0083911 A1, pub. date: 03/17/2022, filed on 01/18/2019), hereinafter Flanagan in view of Wu et al. ("Intent-aware Multi-source Contrastive Alignment for Tag-enhanced Recommendation", ARXIV ID: 2211.06370, Nov. 12, 2022), hereinafter Wu and Shen et al. ("From Distributed Machine Learning To Federated Learning: In The View Of Data Privacy And Security", arXiv:2010.09258v1, Oct 19, 2020, pp. 1-19), hereinafter Shen.
Independent Claims 1, 9, and 17
Flanagan discloses a method of training a model (Flanagan, ¶¶ [0040] and [0047] with FIG. 2: the application of Differential Privacy encoding applied to the model updates ΔXi, sent from the client or user equipment 100 back to the server 200 and the training of a machine learning model in federated mode with the application of differential privacy to the model updates ΔXi, as applied to the Federated Collaborative Filter (FCF)), comprising:
obtaining, at a client device (Flanagan, ¶ [0030] with 100 in FIG. 1: user equipment or device 100; ¶ [0039] with FIG. 2: one or more user equipment or device 100a-100m, also referred to as client devices) from a server (Flanagan, ¶ [0030] with 200 in FIG. 1: backend server 200; ¶ [0039] with FIG. 2: the backend server 200), a first version of a machine learning model for media content recommendation (Flanagan, ¶¶ [0006]-[0009], [0013]-[0014], and [0017]: download a master machine learning model for generating a user recommendation related to one or more of a use or interaction with an application of the user equipment; generate the user recommendation related to the use of the application (e.g., a video service running on the user equipment) based on the downloaded master machine learning model; the user uses the personalized recommendations that propose video choices to the user based on video preference selections, user demographic and/or gender data, or videos they have previously selected and/or watched through the service; the master machine learning model is downloaded from a backend server associated with an application service to a user equipment; ¶¶ [0031]-[0032] with FIG. 1: download a master machine learning model for generating a user recommendation related to use of an application of the user equipment 100; the master machine learning model can be downloaded from the backend server 200 to the user equipment 100; the user recommendation can provide one or more different options, or recommendations, to the user related to the use of the application or service, wherein the application is a video service; ¶¶ [0040]-[0041] with FIG. 2: the Master model Y for generating a recommendation in the Federated Learning (FL) mode is distributed to all of the user devices 100a-100m from the backend server 200; the Master model Y will be stored locally on the user equipment 100a-100m as Xi; ¶ [0048] with FIGS. 1-2 and 302 in FIG.3: a master machine learning model is downloaded 302 to a particular user device, such as user equipment 100 shown in FIG. 1; the master machine learning model can be downloaded from the server 200; as described above with respect to FIG. 2, this can be the master model Y);
recommending, based on local information of the client device, a first set of media contents according to the first version of the machine learning model (Flanagan, ¶¶ [0008]-[0009] and [0017]: generate the user recommendation related to the use of the application based on the downloaded master machine learning model and the data related to one or more of the user of the user equipment or the user interaction with the user equipment; minimize the risk of exposing user data by generating the recommendations on the user equipment; provide a high level of user privacy when the user uses the personalized recommendations that propose video choices to the user based on video preference selections, user demographic and/or gender data, or videos they have previously selected and/or watched through the service; ¶¶ [0031]-[0033] with FIG. 1: generating a user recommendation related to use of an application of the user equipment 100; the user recommendation can provide one or more different options, or recommendations, to the user related to the use of the application or service; when the application is a video service, the data can include information pertaining to a video watched by the user in the video service; another form of data can include information about the user; e.g., the data can include any form of user demographic data; the data can include meta data such as location of the user and the user equipment, a type of the user equipment, user gender, or user age, or any combination thereof; this data is obtained by and stored locally in the user equipment 100; ¶ [0042] with FIG. 2: using a combination of the locally stored master model Xi, and local user data, such as for example videos the user has previously watched, a set of personalized recommendations for the user of the user equipment 100a-100m can be generated; ¶ [0051]: Huawei video service provides an application to users to run on their mobile device that allows them to watch videos through the service; the service backend is hosted in a cloud service; the video service would like to offer users a personalized recommendation service to propose video choices to users based on videos they have previously watched through the service, as well as other user specific preferences and demographics; the video service would like to provide the highest level of user privacy they can when the user uses the personalized recommendations; the video service decides to use a Collaborative Filter (CF) recommendation algorithm/model and use a Federated Learning mode to build and update the CF model);
determining an update to the machine learning model based on respective interactions of a user with the first set of media contents (Flanagan, ¶¶ [0006], [0013], and [0031]-[0034] with FIG. 1: calculate a model update for the master machine learning model based on the master machine learning model and data related to one or more of the user of the user equipment or the user's interaction with the user equipment; the data, also referred to as user data, can have different types; e.g., the data can include data obtained or recorded from the user's interaction with the application or service; this can include data recorded based on a user's selection of an item or option of the application, or selection of one or more items being recommended; e.g., when the application is a video service, the data can include information pertaining to a video watched by the user in the video service; the data can include user behavioral data and/or user meta data, or any combination thereof; the calculated model update is encoded using an c-differential privacy mechanism, and the c-differential privacy encoded model update is then transmitted; ¶¶ [0039] and [0042]-[0043] with FIG. 2: model updates ΔXi are generated from the model Xi, and user data; the model updates ΔXi, to "learn" the model Y are then calculated in the user equipment 100a-100m for each user or client, such as Client 1-Client M, respectively, from the master model Xi, stored locally on a specific user equipment 100a-100m, and the corresponding local user data; Differential Privacy encoding is applied to the model update ΔXi, of a particular user equipment 100a-100m to give E(ΔXi) the DP encoded updates; ¶¶ [0048]-[0049] with FIG. 2 and 304 and 306 in FIG. 3: the machine learning model update is calculated 304, such as for example the model update ΔXi, described above with reference to FIG. 2; the model update ΔXi, is encoded 306 by applying ε-Differential Privacy);
; and
providing the (Flanagan, ¶¶ [0006], [0010], [0012]-[0013], [0015]-[0016], and [0018]-[0019]: the model updates uploaded from the user equipment to the backend server; the server apparatus receives a plurality of ε-differential privacy encoded model updates for a master machine learning model; use ε-Differential Privacy to encode the model updates sent from a user equipment to the backend in such a way that it is impossible or very difficult for any agent (including the backend itself) to intercept or view the encoded updates to reverse engineer the encoded updates to extract any useful information about the user data; ¶¶ [0034]-[0035] with FIG. 1: the encoded model update is transmitted to the apparatus 200, referred to herein as the server, or backend server; use ε-Differential Privacy (DP) to encode a model update on the user equipment or device 100 and decode an aggregation of user model updates on the backend server 200; hashing-randomization process is applied to the model updates (which are numbers) and instead of transferring the plain model updates from the user device 100 to the backend server 200, the encoded version is transferred from the user device 100 to the backend server 200; ¶¶ [0044]-[0047] with FIG. 2: the encoded model updates E(ΔXi) are transferred back to the back end server 200 and are aggregated on the server 200 as E(ΔY)=ΣiΔXi; a decoding is applied to E(ΔY) to give an approximation to ΔY; the master model Y is updated as Y=Y+η ΔŶ; the process can continue with the distribution of the updated master model Y, as described above; ¶¶ [0049]-[0050] and [0052] with FIGS. 1-2, 308 in FIG. 3, and FIG.4: the encoded model update, referred to as E(ΔXi) is then sent 308 to the backend server, such as the backend server 200 of FIGS. 1 and 2; a plurality of encoded model updates E(ΔXi) are received 310 at the backend server, such as backend server 200 illustrated in FIGS. 1 and 2; the plurality of encoded model updates E(ΔXi) will be for a given master model Y; the plurality of encoded model updates E(ΔXi) will be aggregated 312 and decoded 314, generally as described with respect to FIG. 2; the given master model Y will be updated 316; the video service applies ε-Differential Privacy to encode the model updates; the encoded model updates are then sent, via the cloud service, to the service backend; the service backend aggregates the encoded model updates and decodes the resulting aggregate to calculate an estimate of the actual model updates; in this manner, the privacy of the user is enhanced since the model updates cannot be decoded individually to learn anything about the user),
wherein determining the update to the machine learning model comprises: (Flanagan, ¶¶ [0006], [0013], and [0031]-[0034] with FIG. 1: calculate a model update for the master machine learning model based on the master machine learning model and data related to one or more of the user of the user equipment or the user's interaction with the user equipment; the data, also referred to as user data, can have different types; e.g., the data can include data obtained or recorded from the user's interaction with the application or service; this can include data recorded based on a user's selection of an item or option of the application, or selection of one or more items being recommended; e.g., when the application is a video service, the data can include information pertaining to a video watched by the user in the video service; the data can include user behavioral data and/or user meta data, or any combination thereof; the calculated model update is encoded using an c-differential privacy mechanism, and the c-differential privacy encoded model update is then transmitted; ¶¶ [0039] and [0042]-[0043] with FIG. 2: model updates ΔXi are generated from the model Xi, and user data; the model updates ΔXi, to "learn" the model Y are then calculated in the user equipment 100a-100m for each user or client, such as Client 1-Client M, respectively, from the master model Xi, stored locally on a specific user equipment 100a-100m, and the corresponding local user data; Differential Privacy encoding is applied to the model update ΔXi, of a particular user equipment 100a-100m to give E(ΔXi) the DP encoded updates; ¶¶ [0048]-[0049] with FIG. 2 and 304 and 306 in FIG. 3: the machine learning model update is calculated 304, such as for example the model update ΔXi, described above with reference to FIG. 2; the model update ΔXi, is encoded 306 by applying ε-Differential Privacy).
Flanagan further discloses an electronic device (Flanagan, ¶ [0030] with 100 in FIG. 1: the user equipment or device 100; ¶ [0053] with 100/200 in FIG. 1 and 1000 in FIG. 5: the apparatus 1000 is appropriate for use in a wireless network and can be implemented in one or more of the user equipment apparatus 100 or the backend server apparatus 200), comprising a computer processor (Flanagan, ¶ [0030] with 102 in FIG. 1: user equipment 100, includes one or more processors 102; ¶¶ [0054]-[0055] with 1002 in FIG. 5: processor or computing hardware 1002) coupled to a computer-readable memory unit (Flanagan, ¶ [0030] with 108 in FIG. 1: connected or coupled to one or more memory devices 108; ¶¶ [0054] and [0056] with 1004 in FIG. 5: coupled to a memory 1004), the memory unit comprising instructions that when executed by the computer processor implements a method for the media content recommendation described above (Flanagan, ¶ [0030] with FIG. 1: the processor 102 is configured to execute non-transitory machine readable program instructions; ¶¶ [0056]-[0057] with FIG. 5: the memory 1004 is configured to store computer program instructions that may be accessed and executed by the processor 1002 to cause the processor 1002 to perform a variety of desirable computer implemented processes or methods such as the methods as described herein).
Flanagan also discloses a computer program product, the computer program product comprising a non- transitory computer readable storage medium (Flanagan, ¶ [0030] with 108 in FIG. 1: one or more memory devices 108; ¶¶ [0054] and [0056] with 1004 in FIG. 5: a memory 1004) having program instructions embodied therewith (Flanagan, ¶ [0057]: the program instructions stored in memory 1004), the program instructions executable by an electronic device (Flanagan, ¶ [0030] with 100 in FIG. 1: the user equipment or device 100; ¶ [0053] with 100/200 in FIG. 1 and 1000 in FIG. 5: the apparatus 1000 is appropriate for use in a wireless network and can be implemented in one or more of the user equipment apparatus 100 or the backend server apparatus 200) to cause the electronic device to perform a method for media content recommendation described above (Flanagan, ¶¶ [0056]-[0057] with FIG. 5: the memory 1004 is configured to store computer program instructions that may be accessed and executed by the processor 1002 to cause the processor 1002 to perform a variety of desirable computer implemented processes or methods such as the methods as described herein).
Flanagan fails to explicitly disclose(1) encrypting the update with a private key of the client device; (2) providing the encrypted update to the server; and (3) for a given media content in the first set of media contents, assigning, to the given media content, a label corresponding to an interaction of the user with the given media content, the label indicating a degree of interest of the user in the given media content; determining a difference between the label and a prediction of the given media content.
Wu teaches a system and a method related to recommendation service (Wu, Abstract), wherein for a given media content in the first set of media contents, assigning, to the given media content, a label corresponding to an interaction of the user with the given media content, the label indicating a degree of interest of the user in the given media content; determining a difference between the label and a prediction of the given media content (Wu, Abstract in Page 1: use a self-supervision signal to pair users with the auxiliary information (tags) associated with the items they have interacted with before; for a given item, the model predicts which is the correct pairing between the representations obtained from the users that have interacted with this item and the tags assigned to it; provide an efficient solution, using the auxiliary information (tags) directly to enhance the quality of user and item embeddings; user behavior in recommendation systems is driven by the complex interactions of many factors behind the users’ decision-making processes; to make the pairing process more fine-grained and avoid embedding collapse, propose a user intent-aware self-supervised pairing process where split the user embeddings into multiple sub-embedding vectors; each sub-embedding vector captures a specific user intent via self-supervised alignment with a particular cluster of tags; integrate designed framework with various recommendation models, demonstrating its flexibility and compatibility; Section I in Pages 1-2 with FIG. 1 in Page 1: recommendation systems are primarily interested in using the user-item interaction history to predict the users’ interests and thereby recommend potential satisfactory items to users; to alleviate the cold-start problem and improve the recommendation quality, auxiliary information (e.g., tags of items, reviews of items, profiles of users) is usually introduced into the item recommendation process to enrich the modeling of user-item interactions; the intent here means the motivation behind a user’s interaction with an item; e.g., a user’s intent behind visiting a restaurant may be to experience good "service" or to "taste" good food, and thus, "service" and "taste" are related to two kinds of user intents; a user may like a restaurant because she likes certain attributes (tags) of it (e.g., good service, wonderful taste, etc.), and it is unnecessary for her to like all the attributes of the restaurant; i.e., a review to a restaurant received from a reviewer with words only related to "good service" indicates that the reviewer has higher degree of interest in the "service" of the restaurant than the degree of interest in the food "taste" of the restaurant; similarly, a review to a restaurant received from a reviewer with words only related to "good taste" indicates that the reviewer has higher degree of interest in the food "taste" of the restaurant than the degree of interest in the "service" of the restaurant; focusing on tag-enhanced recommendation due to the ubiquity and accessibility of tags; propose a method to efficiently bridge the collaborative filtering (CF) signal and the auxiliary semantic information; the core idea is to refine the learned representations through contrastive objectives; in addition to modeling the user-item interaction using Bayesian Personalized Ranking (BPR), construct a self-supervised learning (SSL) task to conduct alignment between multiple sources (i.e., users, items, and tags); employ Intent-aware Representation Modeling (IRM) to decompose user and item embeddings into multiple components, where each component captures a specific intent whose semantic meaning is identified by a corresponding tag cluster, derived using a self-supervised end-to-end clustering method; meanwhile, enforce independence of different intents, ensuring that intents are effectively disentangled; introduce an Intent-aware Multi-source Contrastive Alignment (IMCA) module; for each intent and item, first aggregate associated users and tags; then, employ contrastive learning to optimize the alignment of the aggregated representations; also align user and item representations; propose an Intent-aware Set-to-set Alignment (ISA) module to improve the performance of IMCA on cold-start users and long-tail items; for each intent (tag cluster), identify whether two items are similar by evaluating the Jaccard index between the items’ tag sets, limiting our attention to tags in the cluster; then extend the intent-aware contrastive alignment in IMCA to optimize the alignment of aggregated user and tag representations derived from the sets of similar items, rather than individual items; interpret this as an intent-aware augmentation of the user-item interactions; it leads to an augmented interaction graph with a more uniform degree distribution, mitigating the problem of high-degree nodes exerting too much influence and improving learning significantly for low-degree nodes; proposed method are called Intent-aware Multi-source Contrastive Alignment for Tag-enhanced recommendation (IMCAT); Section II in Pages 3-4 and Section V.C in Pages 8-9: summarize the related works from three perspectives in the field of recommender systems: (i) the use of auxiliary tags, (ii) the use of knowledge graph, and (iii) the use of self-supervised learning technique; CFA represents the users by the tags they have interacted with, and then uses a deep neural network to extract the features layer-by-layer to predict the final score; DSPR makes use of a Multilayer Perceptron (MLP) to translate tag-based user and item profiles into an abstract embedding space, and then maximizes the similarity between the user representations and the relevant items; HDLPR leverages an autoencoder to compress tag-based user and item profiles into a low-rank feature space; TGCN builds a unified graph containing user nodes, item nodes, and tag nodes; it allows the model to leverage the contextual semantics of multi-hop neighbors in the user-tag-item graph through a message passing paradigm; focus on using tags since they are more accessible in practice; the KG-enhanced recommendation methods can be straightforwardly adapted and used in the tag-enhanced recommendation scenario; for applying these methods, treat the tags and items as entities, and each connection to a specific tag as a unique relation; most works for Self-supervised learning (SSL) use the contrastive learning method, which maximizes the similarity of the representation of a target sample with the representations of corresponding positive samples (mutations of the target sample) and minimizes the similarity with representations of negative samples (samples known to be different); SQN combines SSL with reinforcement learning to capture long-term user interests in sequential recommendation; incorporate auxiliary information via GNNs, construct self-supervised objectives from multiple sources to refine the representations for collaborative filtering, naturally bringing the tag information into training (i.e., using tags as training data); this reduces the time complexity and makes it compatible with a wide variety of recommendation models; make the positive sampling pairing process more fine-grained to avoid embedding collapse; the method is not restricted to a specific type of recommendation task (e.g., sequential recommendation), recommendation model (e.g., GNNs), or form of auxiliary knowledge (e.g., knowledge graph); the only auxiliary information required is tags, which are usually easy to obtain; this provides the strategy with a high level of compatibility; Sections III-IV in Pages 4-7 with FIG. 2 in Page 6, FIG. 3 in Page 7, and FIG. 4 in Page 8: denote the sets of all users, items, and tags as U, V, and T , respectively; for each user u [Symbol font/0xCE] U, the user preference data is represented by a set of items she has interacted with as Iu+ := { i [Symbol font/0xCE] I | Yu,i = 1 } where Yu,i [Symbol font/0xCE] R|U|[Symbol font/0xB4]|T| is the binary user-item rating matrix (i.e., indicating degree of user's interest in item); analogously, use Y'u,i [Symbol font/0xCE] R|I|[Symbol font/0xB4]|T| to represent the labelling history between items and tags; then split Iu+ into a training set Su+ and a test set Tu+; then the tag-enhanced top-N recommendation task is formulated as: given the training item set Su+, and the non-empty test item set Tu+ for user u, train a model to recommend an ordered set of N items Xu such that Xu ∩ Su+ = 0; and |Xu| = N; the model should learn from the collaborative filtering signal Y and the auxiliary tag information Y'; the recommendation quality is evaluated by a matching metric between Xu and Tu+ such as Recall@N; Bayesian Personalized Ranking (BPR) is one of the most widely studied methods in recommendation systems for learning the user preference from the implicit user-item interaction history; the core idea of BPR is to maximize the ranking of an item that the user has accessed (treated as a positive sample; i.e., a set of items she has interacted in Iu+) relative to a randomly sampled item (treated as a negative sample) (i.e., a user has higher degree of interest for items in positive samples than items in negative samples); this goal is achieved via a carefully designed loss function LUV as in eqn. (1), where (u, v+, v–) is a training triplet with a positive item v+ and a negative item v– for user u; ỹuv+ refers to the relevance score between u and v+; adopt a similar formulation for learning the relations between items and tags; this can be viewed as recommending tags to items based on the previous item-tag pairing history; the loss function LVT for this task can be formulated as eqn. (2), where (v, t+, t–) is a training triplet with a positive tag t+ and a negative tag t– for item v; use u, v, t to represent the individual embeddings of a user, item, and tag, respectively; to model the user-item interaction behavior in a more fine-grained manner, decompose the representation of each user and item into K components; each component (sub-embedding) aims to represent a distinct user preference; conduct the intent-aware initialization for users and items as represented in eqn. (3), where K denotes the number of user intents; multiple tags with similar semantic meanings can be regarded as a common factor that biases a user to interact with items with similar traits; e.g., consider a restaurant recommendation scenario; a tag cluster "delicious food, yummy, amazing dessert" may correspond to the same factor that the users like the restaurant due to the "taste" of the food, while a tag cluster "feel at home, friendly waiter" can be used to explain another intent corresponding to a desire for good "service"; focus on how to cluster tags so that the kth tag cluster can be properly aligned with the kth intent embedding for users and items; iteratively apply the K-means algorithm on the learned tag embeddings as the training procedure proceeds; the tag embeddings can be trained through the objective LUV + α LVT, where α is a scaling factor; employ an end-to-end self-supervised clustering approach to adaptively obtain the tag clusters; use a Student’s t-distribution to model the probability of assigning the tag tl to the kth cluster as eqn. (4); construct a target distribution which strives to push the representations closer to cluster centers, strengthening the cohesion of the clusters; the target distribution is defined as eqn. (5); construct a self-supervised loss objective LKL for the end-to-end clustering as the Kullback–Leibler (KL)-divergence between the above two matrices as eqn. (6); use a hard allocation to assign each item to one tag cluster; the assigned tag cluster index is determined by argmaxk(Qlk) for tag tl; introduce a new formulation by combining the two modalities of information (user-item collaborative signals Y and item-tag auxiliary information Y') into a common user-item-tag space using a contrastive learning paradigm; treat the items as a middle ground to link the information coming from the other two sources because items are present in both user-item interactions and item-tag labels; for a given item, conduct an aggregation on those users who previously interacted with the item, but design this aggregation to be intent-aware; use an arithmetic average on each intent component over the user embeddings; the tag clusters are obtained through self-supervised training, where each cluster is related to a user intent; now, for a given item vj, compute its tag cluster embedding for each cluster through aggregating only over those tags assigned to vj; conduct this across all K clusters and all |V| items; different items may have varying degrees of relatedness to distinct tag clusters and user intents; e.g., if an item has 10 tags related to intent 1 and only 1 tag related to intent 2, the item is more closely related to intent 1 than intent 2 (i.e., degree of user's interest on item for intent 1 (e.g., "taste") is higher than degree of user's interest on item for intent 2 (e.g., "service")); use a vector mj [Symbol font/0xCE] RK to store the relatedness of vj with respect to all intents; mj is computed based on the number of vj’s tags in each cluster, and the kth entry can be written as eqn. (9); define M = [m1, m2, …, m|V|]T [Symbol font/0xCE] R|V|[Symbol font/0xB4]K, where each row contains the relatedness of an item to all intents (i.e., matrix M indicating degree of user's interest on item for different intents obtained from user-item interactions Yu,i and item-tag labels Y'u,i of the training data); use this matrix for re-weighting the contrastive loss; aim to maximize the alignment of the pairs of positive samples from the sources of users and tags; first use a linear layer to transform the tag aggregation to make it share the same dimension as the user aggregation; propose to merge the user-tag and user-item intent-aware alignments into a single unified alignment task; maximize the alignment between the aggregated user representation for intent k and the sum of the item embedding and its corresponding aggregated tag embedding for intent k when j = j', and minimize it when j ≠ j'; adopt the commonly used InfoNCE loss formulation to maximize the cosine similarity of the correct pairings of user representations and item/tag representations in each training batch while minimizing the cosine similarity of the embeddings of the incorrect pairings; adopt a bidirectional contrastive alignment loss formulation LCA to ensure the pairing process across the multiple sources can be jointly exploited as eqn. (11); the user to item-tag (u2it) alignment under the kth intent is formalized as eqn. (12); use the predefined matrix M here to capture the degree of alignment for each item with respect to each intent based on the corresponding relatedness; Mj,k refers to the entry of M located at the jth row and the kth column, which denotes the relatedness of item vj to the kth cluster and intent (i.e., indicating degree of user's interest on item vj for kth intent); N(vj) is the set of negative samples of vj; treat all items other than j as candidate negatives; analogously, the item-tag to user (it2u) alignment under the kth intent can be formalized as eqn. (13); introduce a learnable nonlinear transformation between the representations and the contrastive loss, which further improves the quality of the learned representations; design more diverse positive sample pairs by aligning users with the tags not only from the items they have interacted with but also the tags from other similar items; this serves to enrich the representations of the cold start users and items; compute the similarity metric as shown in eqn. (15) based on the Jaccard index for items j and j' for the kth intent; then treat any pair of items larger than a predefined threshold as similar items and regard them within the same set under the kth user intent factor; obtain updated loss function LCA* as eqns. (16)-(17) from eqn. (11)-(13); adapt the model to allow forward and backward propagation for mini-batches of data; overall training objective can be formulated as L = LUV + α LVT + β LCA* + γ LKL in eqn. (18), where α, β, and γ are scaling factors; Section V.A in Pages 7-8: evaluate our proposed method on seven real-world datasets with different domains and sparsity, where the first three datasets are all released in HetRec 2011; HetRec-MV is a movie recommendation dataset; it links movies in the MovieLens dataset with their corresponding Internet Movie Database (IMDb) web pages and Rotten Tomatoes movie reviews, where each movie is assigned with tags provided by users; HetRec-FM is an artist recommendation dataset obtained from Last.fm; it contains social networks, music tags, and music-artist listening histories of users; HetRec-Del is gathered from the Delicious social bookmarking system, which contains social relations, bookmarks, and tags from users; AMZBook-Tag is a real-world online product recommendation dataset derived from the Amazon review datasets; to be consistent with the implicit feedback setting, for the datasets with explicit ratings, retain any ratings no less than four (out of five) as positive feedback and treat all other ratings as missing entries).
Flanagan and Wu are analogous art because they are from the same field of endeavor, a system and a method related to recommendation service. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to apply the teaching of Wu to Flanagan. Motivation for doing so would (1) i.
Flanagan in view of Wu fails to explicitly disclose encrypting the update with a private key of the client device; and providing the encrypted update to the server.
Shen teaches a system and a method relating to federated learning (Shen, Summary), wherein encrypting the update with a private key of the client device; and providing the encrypted update to the server (Shen, Summary in Page 1: federated learning is an improved version of distributed machine learning that further offloads operations which would usually be performed by a central server; the server becomes more like an assistant coordinating clients to work together rather than micro-managing the workforce as in traditional DML; one of the greatest advantages of federated learning is the additional privacy and security guarantees it affords; federated learning architecture relies on smart devices, such as smartphones and IoT sensors, that collect and process their own data, so sensitive information never has to leave the client device; rather, clients train a sub-model locally and send an encrypted update to the central server for aggregation into the global model; these strong privacy guarantees make federated learning an attractive choice in a world where data breaches and information theft are common and serious threats; identify the different mechanisms used to provide privacy and security, such as differential privacy, secure multi-party computation and secure aggregation; Section 1 with FIG. 2 of Pages 1-3: FIGURE 2 illustrates a typical federated learning system. First, a central server publishes a machine learning task and selects clients to participate in each epoch of the training process; then it sends the model and relevant sources to the clients and waits for their training results; clients train the model with the data on their device and return an update of the model parameters or gradients to the server; the server then aggregates those details and updates the ’master’ model for the next training epoch; there two key advantages with this type of learning scheme: reduced computational and communications overhead and better privacy; in fact, federated learning can incorporate many privacy preserving and security mechanisms across the entire system – from the collaborative training process to aggregating updates at the server; e.g., differential privacy (DP) and local DP can guarantee that both the training data and the updates remain private at the numeric level; secure aggregation protocols on the server side, consisting of secure multi-party computation, secret sharing and homomorphic encryption, can perturb the updates to guarantee model security during transfer and aggregation; Section 2.2 with FIG. 4 in Pages 3-6: federated Learning is a specific type of DML, designed to overcome some of the privacy and security issues with classic DML architecture; the basic architecture of federated learning including its data flows is illustrated in FIGURE 4; like traditional DML, there is a central server, which is responsible for the overall control and management of training a global model and some clients who receive training subtasks from the central server; there are two key differences however: a) Instead of each client working individually on their own piece of the model, in federated learning, all selected clients work on the same training task in each epoch; and b) The clients in a federated learning system are typically devices like smart phones, tablets, and sensors that are able to capture or collect information as opposed to desktop computers or routers; therefore, because the training data is gathered, stored and used at the client level, the only information that ever needs to be transmitted is the model updates; the learning procedure is relatively straightforward; in each training epoch, the server allocates a training task and computing resources to any client that is ready to learn, then it transmits the current model; the client trains the model with its own local data and sends the updated parameters as encrypted training results back to an aggregator for compilation; as such, there is greater data privacy because there is no need to transmit sensitive information, and encrypting the updated parameters before sending them to aggregators increases security over the models; the aggregators, also controlled by the central server, average the parameter updates; there are two types of aggregators: master and temporary; master aggregators manage the number of training epochs and generate an appropriate number of “temporary” aggregators for each epoch to consolidate the training results; these temporary aggregators do not store any permanent information, and all aggregators follow what is called a “secure aggregation protocol”, which means encrypted data can be processed and compiled without knowing the true data; the master aggregators then fully aggregate the results from the temporary aggregators and deliver the results to the central server that updates the model; the server then schedules the next training task and starts a new training epoch; Section 3.2 in Pages 8-11; in federated learning, most of the communications surround model aggregation because all devices must upload their training updates to the aggregator for averaging; to prevent leaks of any individual’s training results, a specific protocol called "secure aggregation" encrypts the client updates at the device level before they are uploaded for aggregation; the protocol guarantees that all updates are aggregated in a secure way and that any other party can only access the cipher-text of a client’s updates – even the server; these protocols involve secret sharing schemes, secure multi-party computation and homomorphic encryption; the secret sharing scheme for the collection distributes shares of the secret to these parties by a dealer according to two requirements: any subset in the collection can reconstruct the secret from its shares of the secret, and any subset not in the collection cannot reveal any partial information about the secret, separately; secure multi-party computation addresses the problem of having a set of parties calculate an agreed-on function over their private inputs such that all parties can reveal the intended output without obtaining other parties’ inputs; the idea is that all parties’ private inputs are protected by an encryption scheme that guarantees the utility of the data for accurately answering a query function; in this sense, multi-party computation is more like a general notion of secure computation comprising a set of techniques as opposed to being a single method; homomorphic encryption is an encryption scheme that allows complex mathematical operations to be performed on cipher-text without changing the nature of the encryption; the two different types of homomorphic encryption are fully homomorphic encryption and partially homomorphic encryption; fully homomorphic encryption supports both additive and multiplicative operations, while partially homomorphic encryption only supports one or the other; fully homomorphic encryption is strongly recommended in federated learning, even though the cost of computation is much greater because the aggregation process involves both addition and multiplication; also, because the central server should not be able to decrypt the client updates, a trusted third party must be involved to hold a key, and the central server must be able to sum the client updates using only cipher-text; homomorphic encryption exactly meets all these requirements; secure aggregation is a subclass of multi-party computation algorithms where a group of parties that do not trust each other each hold sensitive information and must collaborate to calculate an aggregated value; the aggregated value should not reveal any party’s information (except what it can learn from its own information); like homomorphic encryption and secret sharing schemes, each client’s outputs are encrypted before they are shared, which guarantees a secure transit process; in late 2016, Bonawitz et al. propose the first secure aggregation suggested secure aggregation protocol for federated learning to protect the privacy of clients’ model gradients and to guarantee that the server only learns the sum of the clients’ inputs while the users learn nothing; later, in early 2017, Bonawitz et al. further developed a full version of the protocol for practical applications; a random number masks each client’s raw input to prevent direct disclosure to the central server, and each client generates a private-public key pair for each epoch of the aggregation process; each client is allowed to combine its private key and every other client’s public key, to generate a private shared key with a hash function; the hash function involves Pseudo Random Generator and Decisional Diffie-Hellman assumption to guarantee each pair of clients’ private shared keys are additive inverse; because the sum of a pair of private shared keys is zero, all clients’ masks are offset during the aggregation process, and the server can offset the effect of the masks to calculate an accurate aggregation result without needing to know any of the clients’ true inputs; Mandal et al. propose the non-interactive key establishment protocol (NIKE) and a secure aggregation protocol based on NIKE; NIKE addresses the cost of key sharing; it comprises two non-colluding cryptographic secret service providers who independently calculate pairwise polynomial functions for each client; to generate a shared private key, each client generates a private polynomial function as a private key by multiplying the two polynomial functions; further, each client has a unique order number assigned by the server, which is public information, and any client is allowed to generate a shared private key by placing the targeted client’s order number into their private polynomial function; thus, there is no communication cost for generating a shared key, and the protocol guarantees that each pair of client calculations with its own private polynomial function will have the same results; the NIKE-based secure aggregation protocol reduces the communication costs associated with the secret sharing scheme; each client can only calculate private shared keys with their neighbors via the NIKE protocol; the communications costs for reconstructing a disconnected client’s mask is reduced ; again, each client generates a double mask to protect its inputs for the same reasons as outlined above; an additive homomorphic encryption scheme and a secure aggregation protocol with practical crypto-primitives imposed at the beginning of each learning epoch guarantee a safe environment for the training process client-side and the aggregation process server-side)
Flanagan in view of Wu, and Shen are analogous art because they are from the same field of endeavor, a system and a method relating to Federated learning. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to apply the teaching of Shen to Flanagan in view of Wu. Motivation for doing so would provide additional privacy and security guarantees, and guarantee model security during transfer and aggregation (Shen, Summary in Page 1; Section 1 in Pages 1-3; Section 3.2.2 in Pages 10-11).
Claims 3 and 11
Flanagan in view of Wu and Shen discloses all the elements as stated in Claims 1 and 9 respectively and further discloses wherein the label is selected from one or more of following: a first label corresponding to a positive user interaction, and a second label corresponding to a negative user interaction (Wu, Sections III-IV in Pages 4-7 with FIG. 2 in Page 6, FIG. 3 in Page 7, and FIG. 4 in Page 8: denote the sets of all users, items, and tags as U, V, and T , respectively; for each user u [Symbol font/0xCE] U, the user preference data is represented by a set of items she has interacted with as Iu+ := { i [Symbol font/0xCE] I | Yu,i = 1 } where Yu,i [Symbol font/0xCE] R|U|[Symbol font/0xB4]|T| is the binary user-item rating matrix; analogously, use Y'u,i [Symbol font/0xCE] R|I|[Symbol font/0xB4]|T| to represent the labelling history between items and tags; then split Iu+ into a training set Su+ and a test set Tu+; then the tag-enhanced top-N recommendation task is formulated as: given the training item set Su+, and the non-empty test item set Tu+ for user u, train a model to recommend an ordered set of N items Xu such that Xu ∩ Su+ = 0; and |Xu| = N; the model should learn from the collaborative filtering signal Y and the auxiliary tag information Y'; the recommendation quality is evaluated by a matching metric between Xu and Tu+ such as Recall@N; Bayesian Personalized Ranking (BPR) is one of the most widely studied methods in recommendation systems for learning the user preference from the implicit user-item interaction history; the core idea of BPR is to maximize the ranking of an item that the user has accessed (treated as a positive sample) relative to a randomly sampled item (treated as a negative sample); this goal is achieved via a carefully designed loss function LUV as in eqn. (1), where (u, v+, v–) is a training triplet with a positive item v+ and a negative item v– for user u; ỹuv+ refers to the relevance score between u and v+; adopt a similar formulation for learning the relations between items and tags; this can be viewed as recommending tags to items based on the previous item-tag pairing history; the loss function LVT for this task can be formulated as eqn. (2), where (v, t+, t–) is a training triplet with a positive tag t+ and a negative tag t– for item v; use u, v, t to represent the individual embeddings of a user, item, and tag, respectively; to model the user-item interaction behavior in a more fine-grained manner, decompose the representation of each user and item into K components; each component (sub-embedding) aims to represent a distinct user preference; conduct the intent-aware initialization for users and items as represented in eqn. (3), where K denotes the number of user intents; focus on how to cluster tags so that the kth tag cluster can be properly aligned with the kth intent embedding for users and items; iteratively apply the K-means algorithm on the learned tag embeddings as the training procedure proceeds; the tag embeddings can be trained through the objective LUV + α LVT, where α is a scaling factor; employ an end-to-end self-supervised clustering approach to adaptively obtain the tag clusters; use a Student’s t-distribution to model the probability of assigning the tag tl to the kth cluster as eqn. (4) construct a target distribution which strives to push the representations closer to cluster centers, strengthening the cohesion of the clusters; the target distribution is defined as eqn. (5); construct a self-supervised loss objective LKL for the end-to-end clustering as the Kullback–Leibler (KL)-divergence between the above two matrices as eqn. (6); use a hard allocation to assign each item to one tag cluster; the assigned tag cluster index is determined by argmaxk(Qlk) for tag tl; introduce a new formulation by combining the two modalities of information (user-item collaborative signals Y and item-tag auxiliary information Y') into a common user-item-tag space using a contrastive learning paradigm; treat the items as a middle ground to link the information coming from the other two sources because items are present in both user-item interactions and item-tag labels; for a given item, conduct an aggregation on those users who previously interacted with the item, but design this aggregation to be intent-aware; use an arithmetic average on each intent component over the user embeddings; the tag clusters are obtained through self-supervised training, where each cluster is related to a user intent; now, for a given item vj, compute its tag cluster embedding for each cluster through aggregating only over those tags assigned to vj; conduct this across all K clusters and all |V| items; different items may have varying degrees of relatedness to distinct tag clusters and user intents; e.g., if an item has 10 tags related to intent 1 and only 1 tag related to intent 2, the item is more closely related to intent 1 than intent 2; use a vector mj [Symbol font/0xCE] RK to store the relatedness of vj with respect to all intents; mj is computed based on the number of vj’s tags in each cluster, and the kth entry can be written as eqn. (9); define M = [m1, m2, …, m|V|]T [Symbol font/0xCE] R|V|[Symbol font/0xB4]K, where each row contains the relatedness of an item to all intents; use this matrix for re-weighting the contrastive loss; aim to maximize the alignment of the pairs of positive samples from the sources of users and tags; first use a linear layer to transform the tag aggregation to make it share the same dimension as the user aggregation; propose to merge the user-tag and user-item intent-aware alignments into a single unified alignment task; maximize the alignment between the aggregated user representation for intent k and the sum of the item embedding and its corresponding aggregated tag embedding for intent k when j = j', and minimize it when j ≠ j'; adopt the commonly used InfoNCE loss formulation to maximize the cosine similarity of the correct pairings of user representations and item/tag representations in each training batch while minimizing the cosine similarity of the embeddings of the incorrect pairings; adopt a bidirectional contrastive alignment loss formulation LCA to ensure the pairing process across the multiple sources can be jointly exploited as eqn. (11); the user to item-tag (u2it) alignment under the kth intent is formalized as eqn. (12); use the predefined matrix M here to capture the degree of alignment for each item with respect to each intent based on the corresponding relatedness; Mj,k refers to the entry of M located at the jth row and the kth column, which denotes the relatedness of item vj to the kth cluster and intent; N(vj) is the set of negative samples of vj; analogously, the item-tag to user (it2u) alignment under the kth intent can be formalized as eqn. (13); introduce a learnable nonlinear transformation between the representations and the contrastive loss, which further improves the quality of the learned representations; design more diverse positive sample pairs by aligning users with the tags not only from the items they have interacted with but also the tags from other similar items; this serves to enrich the representations of the cold start users and items; compute the similarity metric as shown in eqn. (15) based on the Jaccard index for items j and j' for the kth intent; then treat any pair of items larger than a predefined threshold as similar items and regard them within the same set under the kth user intent factor; obtain updated loss function LCA* as eqns. (16)-(18) from eqn. (11)-(13); adapt the model to allow forward and backward propagation for mini-batches of data; overall training objective can be formulated as L = LUV + α LVT + β LCA* + γ LKL in eqn. (18), where α, β, and γ are scaling factors).
It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to apply the teaching of Wu to Flanagan. Motivation for doing so would (1) i.
Claims 4 and 12
Flanagan in view of Wu and Shen discloses all the elements as stated in Claims 1 and 9 respectively and further discloses wherein the label is selected from two or more of following: a third label corresponding to tagging the given media content positively, a fourth label corresponding to spreading of the given media content, a fifth label corresponding to commenting the given media content, a sixth label corresponding to ignoring the given media content, and a seventh label corresponding to tagging the given media content negatively (Wu, Sections III-IV in Pages 4-7 with FIG. 2 in Page 6, FIG. 3 in Page 7, and FIG. 4 in Page 8: denote the sets of all users, items, and tags as U, V, and T , respectively; for each user u [Symbol font/0xCE] U, the user preference data is represented by a set of items she has interacted with as Iu+ := { i [Symbol font/0xCE] I | Yu,i = 1 } where Yu,i [Symbol font/0xCE] R|U|[Symbol font/0xB4]|T| is the binary user-item rating matrix; analogously, use Y'u,i [Symbol font/0xCE] R|I|[Symbol font/0xB4]|T| to represent the labelling history between items and tags; then split Iu+ into a training set Su+ and a test set Tu+; then the tag-enhanced top-N recommendation task is formulated as: given the training item set Su+, and the non-empty test item set Tu+ for user u, train a model to recommend an ordered set of N items Xu such that Xu ∩ Su+ = 0; and |Xu| = N; the model should learn from the collaborative filtering signal Y and the auxiliary tag information Y'; the recommendation quality is evaluated by a matching metric between Xu and Tu+ such as Recall@N; Bayesian Personalized Ranking (BPR) is one of the most widely studied methods in recommendation systems for learning the user preference from the implicit user-item interaction history; the core idea of BPR is to maximize the ranking of an item that the user has accessed (treated as a positive sample) relative to a randomly sampled item (treated as a negative sample); this goal is achieved via a carefully designed loss function LUV as in eqn. (1), where (u, v+, v–) is a training triplet with a positive item v+ and a negative item v– for user u; ỹuv+ refers to the relevance score between u and v+; adopt a similar formulation for learning the relations between items and tags; this can be viewed as recommending tags to items based on the previous item-tag pairing history; the loss function LVT for this task can be formulated as eqn. (2), where (v, t+, t–) is a training triplet with a positive tag t+ and a negative tag t– for item v; use u, v, t to represent the individual embeddings of a user, item, and tag, respectively; to model the user-item interaction behavior in a more fine-grained manner, decompose the representation of each user and item into K components; each component (sub-embedding) aims to represent a distinct user preference; conduct the intent-aware initialization for users and items as represented in eqn. (3), where K denotes the number of user intents; focus on how to cluster tags so that the kth tag cluster can be properly aligned with the kth intent embedding for users and items; iteratively apply the K-means algorithm on the learned tag embeddings as the training procedure proceeds; the tag embeddings can be trained through the objective LUV + α LVT, where α is a scaling factor; employ an end-to-end self-supervised clustering approach to adaptively obtain the tag clusters; use a Student’s t-distribution to model the probability of assigning the tag tl to the kth cluster as eqn. (4) construct a target distribution which strives to push the representations closer to cluster centers, strengthening the cohesion of the clusters; the target distribution is defined as eqn. (5); construct a self-supervised loss objective LKL for the end-to-end clustering as the Kullback–Leibler (KL)-divergence between the above two matrices as eqn. (6); use a hard allocation to assign each item to one tag cluster; the assigned tag cluster index is determined by argmaxk(Qlk) for tag tl; introduce a new formulation by combining the two modalities of information (user-item collaborative signals Y and item-tag auxiliary information Y') into a common user-item-tag space using a contrastive learning paradigm; treat the items as a middle ground to link the information coming from the other two sources because items are present in both user-item interactions and item-tag labels; for a given item, conduct an aggregation on those users who previously interacted with the item, but design this aggregation to be intent-aware; use an arithmetic average on each intent component over the user embeddings; the tag clusters are obtained through self-supervised training, where each cluster is related to a user intent; now, for a given item vj, compute its tag cluster embedding for each cluster through aggregating only over those tags assigned to vj; conduct this across all K clusters and all |V| items; different items may have varying degrees of relatedness to distinct tag clusters and user intents; e.g., if an item has 10 tags related to intent 1 and only 1 tag related to intent 2, the item is more closely related to intent 1 than intent 2; use a vector mj [Symbol font/0xCE] RK to store the relatedness of vj with respect to all intents; mj is computed based on the number of vj’s tags in each cluster, and the kth entry can be written as eqn. (9); define M = [m1, m2, …, m|V|]T [Symbol font/0xCE] R|V|[Symbol font/0xB4]K, where each row contains the relatedness of an item to all intents; use this matrix for re-weighting the contrastive loss; aim to maximize the alignment of the pairs of positive samples from the sources of users and tags; first use a linear layer to transform the tag aggregation to make it share the same dimension as the user aggregation; propose to merge the user-tag and user-item intent-aware alignments into a single unified alignment task; maximize the alignment between the aggregated user representation for intent k and the sum of the item embedding and its corresponding aggregated tag embedding for intent k when j = j', and minimize it when j ≠ j'; adopt the commonly used InfoNCE loss formulation to maximize the cosine similarity of the correct pairings of user representations and item/tag representations in each training batch while minimizing the cosine similarity of the embeddings of the incorrect pairings; adopt a bidirectional contrastive alignment loss formulation LCA to ensure the pairing process across the multiple sources can be jointly exploited as eqn. (11); the user to item-tag (u2it) alignment under the kth intent is formalized as eqn. (12); use the predefined matrix M here to capture the degree of alignment for each item with respect to each intent based on the corresponding relatedness; Mj,k refers to the entry of M located at the jth row and the kth column, which denotes the relatedness of item vj to the kth cluster and intent; N(vj) is the set of negative samples of vj; analogously, the item-tag to user (it2u) alignment under the kth intent can be formalized as eqn. (13); introduce a learnable nonlinear transformation between the representations and the contrastive loss, which further improves the quality of the learned representations; design more diverse positive sample pairs by aligning users with the tags not only from the items they have interacted with but also the tags from other similar items; this serves to enrich the representations of the cold start users and items; compute the similarity metric as shown in eqn. (15) based on the Jaccard index for items j and j' for the kth intent; then treat any pair of items larger than a predefined threshold as similar items and regard them within the same set under the kth user intent factor; obtain updated loss function LCA* as eqns. (16)-(18) from eqn. (11)-(13); adapt the model to allow forward and backward propagation for mini-batches of data; overall training objective can be formulated as L = LUV + α LVT + β LCA* + γ LKL in eqn. (18), where α, β, and γ are scaling factors; Section V.A in Pages 7-8: evaluate our proposed method on seven real-world datasets with different domains and sparsity, where the first three datasets are all released in HetRec 2011; HetRec-MV is a movie recommendation dataset; it links movies in the MovieLens dataset with their corresponding Internet Movie Database (IMDb) web pages and Rotten Tomatoes movie reviews, where each movie is assigned with tags provided by users; HetRec-FM is an artist recommendation dataset obtained from Last.fm; it contains social networks, music tags, and music-artist listening histories of users; HetRec-Del is gathered from the Delicious social bookmarking system, which contains social relations, bookmarks, and tags from users).
It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to apply the teaching of Wu to Flanagan. Motivation for doing so would (1) i.
Claims 7, 15, and 20
Flanagan in view of Wu and Shen discloses all the elements as stated in Claims 1, 9, and 17 respectively and further discloses obtaining, from the server, a second version of the machine learning model, the second version being updated from the first version at least based on the update and a further update provided by a further client device (Flanagan, ¶¶ [0039]-[0047] with FIG. 2: the Master model Yin the Federated Learning (FL) mode is distributed to all of the user devices 100a-100m from the backend server 200; model updates ΔXi are generated from the model Xi, and user data; the model updates ΔXi, to "learn" the model Y are then calculated in the user equipment 100a-100m for each user or client, such as Client 1-Client M, respectively, from the master model Xi, stored locally on a specific user equipment 100a-100m, and the corresponding local user data; Differential Privacy encoding is applied to the model update ΔXi, of a particular user equipment 100a-100m to give E(ΔXi) the DP encoded updates; the encoded model updates E(ΔXi) are transferred back to the back end server 200 and are aggregated on the server 200 as E(ΔY)=ΣiΔXi; a decoding is applied to E(ΔY) to give an approximation to ΔY; the master model Y is updated as Y=Y+η ΔŶ; the process can continue with the distribution of the updated master model Y, as described above; i.e., the distribution of the updated master model Y to all of the user devices 100a-100m from the backend server 200).
Claims 6, 8, 14, 16, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Flanagan in view of Wu and Shen as applied to Claims 1, 9, and 17 respectively above, and further in view of AMAD-UD-DIN et al. (US 2022/0012601 A1, pub. date: 01/13/2022; filed on 09/24/2021), hereinafter AMAD-UD-DIN'601.
Claims 6, 14, and 19
Flanagan in view of Wu and Shen discloses all the elements as stated in Claims 1, 9, and 17 respectively and further discloses wherein recommending the first set of media contents comprises: obtaining, from the server, content information concerning a first number of candidate media contentscontents by feeding the local information to the machine learning model; and presenting an indication of the first set of media contents (Flanagan, ¶ [0051]: Huawei video service provides an application to users to run on their mobile device that allows them to watch videos through the service; the service backend is hosted in a cloud service; the video service would like to offer users a personalized recommendation service to propose video choices to users based on videos they have previously watched through the service, as well as other user specific preferences and demographics; the video service would like to provide the highest level of user privacy they can when the user uses the personalized recommendations; the video service decides to use a Collaborative Filter (CF) recommendation algorithm/model and use a Federated Learning mode to build and update the CF model; ¶¶ [0008] and [0031]-[0033] with FIG. 1: generate the user recommendation related to the use of the application based on the downloaded master machine learning model and the data related to one or more of the user of the user equipment or the user interaction with the user equipment; the user recommendation can provide one or more different options, or recommendations, to the user related to the use of the application or service; the data, also referred to as user data, can have different types; e.g. the data can include data obtained or recorded from the user's interaction with the application or service; this can include data recorded based on a user's selection of an item or option of the application, or selection of one or more items being recommended; e.g., when the application is a video service, the data can include information pertaining to a video watched by the user in the video service; another form of data can include information about the user; e.g. the data can include any form of user demographic data; the data can include meta data such as location of the user and the user equipment, a type of the user equipment, user gender, or user age, or any combination thereof; the data can include user behavioral data and/or user meta data, or any combination thereof; this data is obtained by and stored locally in the user equipment 100; this type of data is obtained in any suitable manner and stored on any suitable storage medium accessible by the user equipment 100; ¶¶ [0009] and [0017]: provide a high level of user privacy when the user uses the personalized recommendations that propose video choices to the user based on video preference selections, user demographic and/or gender data, or videos they have previously selected and/or watched through the service; ¶ [0042] with FIG. 1: using a combination of the locally stored master model Xi, and local user data, such as for example videos the user has previously watched, a set of personalized recommendations for the user of the user equipment 100a-100m can be generated).
Flanagan in view of Wu and Shen fails to explicitly disclose wherein the first number of candidate media contents being selected from a second number of candidate media contents.
AMAD-UD-DIN'601 teaches a system and a method relating to Federated Recommendation (AMAD-UD-DIN'601, ¶ [0002]), wherein the first number of candidate media contents being selected from a second number of candidate media contents (AMAD-UD-DIN'601, ¶¶ [0041]-[0044] and [0048]-[0049] with FIGS. 2-3: the server side 202 of the recommendation system 300 is composed of one or more processors running two algorithms operating in Federated Learning mode; a Collaborative Filter (CF) 312 is used to generate a user specific candidate set of video recommendations; a Predictive Model (PM) 314 is used to score each video in the candidate set and to generate the final video recommendations; a client on the client side 200, also referred to herein as a client side device or client side devices, is also composed of one or more processors running two algorithms operating in Federated Learning mode; a Collaborative Filter (CF) 322 on the client side 200 is used to receive and generate a user specific candidate set 325 of video recommendations; a Predictive Model (PM) 324 on the client side 200 is used to score 326 each video in the candidate set and to generate the final set 327 of video recommendations; each client 200a-200n on the client side will generate a final set 327 of video recommendations; the Collaborative Filter 322 generates the candidate set 325 based on a user's video watch event or behavioral data; the Predictive Model 324 re-scores 326 the candidate set 325 based on the user's personal data; the candidate set 325 is seen as a sub-set of the total number of videos, filtered based on the user's watching behavior; this filtered set is then re-scored 326 such that the videos which have high probability of being liked by the user get a high score and are recommended; each of the master models and metrics described above are distributed to each of Huawei Video services user devices on the client side 200; the master models along with the metrics from the server side 202, referred to as the local master models on the client side 200, now reside on the user devices, such as the user devices 200a-200n shown in FIG. 2, and have the same hyper-parameter configurations as the master models on the servers 202; the local master models that now reside on the client side 200 are generally configured to generate recommendations, update and train and evaluate; the local master model of the collaborative filter 322 is used to generate a candidate set 325 of videos for the user using the local user data which can include, but is not limited to, the videos watched by the user on that device; the generated candidate set 325 of videos is scored 326 by the local predictive module 324 based on user personal data which can include for example, but is not limited to other applications used by the user, date of birth stored on the user device, location of the device etc.; the result of the scoring 326 is the final list or set 327 of videos, which is generated or provided as a personal set of video recommendations to the user; the locally generated video recommendations can then be shown or otherwise presented to the user on the device; in this manner, the user of a particular client device 200 is encouraged to select one or more of the video recommendations from this personalized set 327 for watching).
Flanagan in view of Wu and Shen, and AMAD-UD-DIN'601 are analogous art because they are from the same field of endeavor, a system and a method relating to Federated Recommendation. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to apply the teaching of AMAD-UD-DIN'601 to Flanagan in view of Wu and Shen. Motivation for doing so would maximize the client's privacy.
Claims 8 and 16
Flanagan in view of Wu and Shen discloses all the elements as stated in Claims 7 and 15 respectively and further discloses recommending, based on the local information, a second set of media contents according to the second version of the machine learning model (Flanagan, ¶¶ [0039]-[0047] with FIG. 2: the Master model Yin the Federated Learning (FL) mode is distributed to all of the user devices 100a-100m from the backend server 200; using a combination of the locally stored master model Xi, and local user data, such as for example videos the user has previously watched, a set of personalized recommendations for the user of the user equipment 100a-100m can be generated; the process can continue with the distribution of the updated master model Y, as described above; i.e., the distribution of the updated master model Y to all of the user devices 100a-100m from the backend server 200, and generate new recommendations based on the updated master model Y).
Flanagan in view of Wu and Shen fails to explicitly disclose wherein determining a metric for evaluating the machine learning model based on respective interactions of the user with the second set of media contents; and providing the metric to the server.
AMAD-UD-DIN'601 teaches a system and a method relating to Federated Recommendation (AMAD-UD-DIN'601, ¶ [0002]), wherein determining a metric for evaluating the machine learning model based on respective interactions of the user with the second set of media contents; and providing the metric to the server (AMAD-UD-DIN'601, ¶ [0017] and [0035]: the only information that is required from the clients, without knowing their identities, is the validation set performances, also referred as accuracy metrics; ¶ [0038] with FIG. 2: the hyper-parameter optimizer 204 receives as an input 212 from the Federated Learning Server master model 202, current hyper-parameter configuration and performance metrics; the Federated Learning Server master model 202 collects and stores the current configuration and performance metrics as part of model updates sent by clients; ¶¶ [0041]-[0057] with FIGS. 2-4: a Collaborative Filter (CF) 322 on the client side 200 is used to receive and generate a user specific candidate set 325 of video recommendations; a Predictive Model (PM) 324 on the client side 200 is used to score 326 each video in the candidate set and to generate the final set 327 of video recommendations; the Collaborative Filter 322 generates the candidate set 325 based on a user's video watch event or behavioral data; the Predictive Model 324 re-scores 326 the candidate set 325 based on the user's personal data; the candidate set 325 is seen as a sub-set of the total number of videos, filtered based on the user's watching behavior. This filtered set is then re-scored 326 such that the videos which have high probability of being liked by the user get a high score and are recommended; Huawei Video Service initializes validation set performance metrics namely Root Mean Squared Error (RMSE) and log-loss on its server 202, one for the Collaborative Filter 312 and one for the Predictive Model 314, respectively; Huawei Video Service creates two master models on its server 202, one for the Collaborative Filter 312 and one for the Predictive Model 314; the two master models are initialized with the respective hyper-parameters suggested by the hyper-parameter optimizer 204; each of the master models and metrics described above are distributed to each of Huawei Video services user devices on the client side 200; the master models along with the metrics from the server side 202, referred to as the local master models on the client side 200, now reside on the user devices, such as the user devices 200a-200n shown in FIG. 2, and have the same hyper-parameter configurations as the master models on the servers 202; the local master models that now reside on the client side 200 are generally configured to generate recommendations, update and train and evaluate; based on the user's viewing of different videos in the Huawei video service, the different videos are randomly divided into training, validation and test sets; using the training set, the local master model of the respective collaborative filter 322 is updated; the local master model updates for each user, or client side device 200, are different and independent; using the training set and based on the user's personal data, such as for example, the user's uses of other services on the device, the user's age and gender, the local predictive model 324 is updated; on the client side 200, using the local data, the validation set 325 and training set 327 video recommendations are generated for each user independently; the training set 327 is used to update the local model; the validation set 325 is used to evaluate the local model and compute the validation set performance metrics; the validation set recommendations are evaluated to update the validation set performance metrics; the validation set performance metrics updates for the local collaborative filter 322 and predictive model 324 models are transferred back to the Federated Learning Server, or in this example, the Huawei video service server 202, where the Federated Learning Master Model is residing; the collaborative filter model updates received from the client side devices 200 are aggregated 402 to update the collaborative filter 312 master model; the collaborative filter validation set performance metric updates received from each client side device 200 are averaged to create a new updated collaborative filter metric, referred to as RMSE; the server 202 is also configured to aggregate the predictive model updates obtained from each client 200 and update the master model of the predictive model 314 of the server 202. The validation set performance metric updates received from each client for the predictive model 324 are averaged to create a new updated predicted model metric, generally referred to herein as log-loss).
Flanagan in view of Wu and Shen, and AMAD-UD-DIN'601 are analogous art because they are from the same field of endeavor, a system and a method relating to Federated Recommendation. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to apply the teaching of AMAD-UD-DIN'601 to Flanagan in view of Wu and Shen. Motivation for doing so would maximize the client's privacy.
Response to Arguments
Applicant’s arguments filed 07/01/2026 with respect to Claims 1, 9, and 17 have been fully considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to HWEI-MIN LU whose telephone number is (313)446-4913. The examiner can normally be reached Mon - Fri: 9:00 AM - 6:00 PM EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Mariela D. Reyes can be reached at (571) 270-1006. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/HWEI-MIN LU/Primary Examiner, Art Unit 2142