Prosecution Insights
Last updated: October 01, 2026
Application No. 18/579,089

Systems and Methods for Federated Learning of Machine-Learned Models with Sampled Softmax

Non-Final OA §102§103
Filed
Jan 12, 2024
Priority
Jul 12, 2021 — nonprovisional of PCTUS2021041225
Examiner
HWANG, MEGAN ELIZABETH
Art Unit
Tech Center
Assignee
Google LLC
OA Round
1 (Non-Final)
54%
Grant Probability
Moderate
1-2
OA Rounds
1y 3m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 54% of resolved cases
54%
Career Allowance Rate
18 granted / 33 resolved
-5.5% vs TC avg
Strong +58% interview lift
Without
With
+57.5%
Interview Lift
resolved cases with interview
Typical timeline
4y 0m
Avg Prosecution
9 currently pending
Career history
46
Total Applications
across all art units

Statute-Specific Performance

§101
31.1%
-8.9% vs TC avg
§103
43.9%
+3.9% vs TC avg
§102
8.3%
-31.7% vs TC avg
§112
15.2%
-24.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 33 resolved cases

Office Action

§102 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claims 1-20 are presented for examination. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless –(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 13-18 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Niu et al. (“Secure Federated Submodel Learning”, published 11/11/2019), hereinafter Niu. Niu was cited in the IDS dated 01/12/2024. Regarding Claim 13, Niu teaches A computer-implemented method for federated learning of a machine- learned model with reduced computing resource usage (Niu: “our new framework further decouples the ability to accomplish federated learning from the need to use the prohibitively large full model, which can dramatically improve efficiency. For example, in our evaluation, the size of a client’s desired submodel is only 1.99% of the full model’s size. Thus, our framework is more practical for resource constrained clients and deep learning tasks.” [Section I. Introduction]), the method comprising: receiving, by a server computing system comprising one or more computing devices, one or more class sets from one or more client computing systems, the one or more class sets comprising, for each client computing system of the one or more client computing systems, a plurality of local positive class labels respective to a plurality of training examples of a local training dataset at the client computing system and one or more negative class labels sampled by the client computing system (Niu: “First, the real index is extracted from a client’s private data and is kept secret from the other system participants, including the cloud server and any other chosen client. Second, the perturbed index set is used to interact with others in the download and upload phases. It is generated by applying randomized response twice with one memoization step between. Such a design, together with secure aggregation, allows the client to hold a self-controllable deniability against whether she really intends or does not intend to download some row and to upload the modification of this row, even if the client may be chosen to participate in multiple communication rounds.” [Section I.E. Our Solution Overview and Major Contributions]; [Algorithm 1, Lines 3-4]; “Given the questionnaire, client i basically uses two probability parameters p1(i), p2(i) in randomized response to fine-tune the tension among effectiveness, efficiency, and privacy (Lines 3–6). In particular, p1(i) denotes the probability that an index in client i’s real index set will return a “Yes” answer and controls the factual size of a client’s user data contributed to federated submodel learning. Thus, a larger p1(i) implies better effectiveness in terms of convergency rate. In addition, among privacy, effectiveness, and efficiency. More specifically, p2(i) denotes the probability that an index outside client i’s real index set will return a “Yes” answer and determines the number of redundant rows to be downloaded and the number of padded zero vectors to be uploaded through the secure aggregation protocol.” [Section IV.B.2. Index Set Perturbation]); communicating, by the server computing system, data descriptive of a respective classification submodel to each client computing system of the one or more client computing systems, wherein the respective classification submodel is configured to provide a classification output limited to the client class set received from the client computing system (Niu: “The union result will be further delivered to live clients, based on which each client can perturb her real index set with a customized local differential privacy guarantee (Line 12). In addition, each client will use the perturbed index set, rather than the real index set, to download her submodel and upload the submodel update (Lines 13 and 19). In other words, when interacting with the cloud server, a client’s real index set is replaced with her perturbed index set, which provides deniability of her real index set and thus obscures her training data. Upon receiving the perturbed index set from a client, the cloud server stores it for later usage and returns the corresponding submodel and the training hyperparameters to the client (Line 6).” [Section IV.B.1. Secure Federated Submodel Learning]; [Algorithm 1, Lines 5-6]); receiving, by the server computing system, one or more model updates from the one or more client computing systems (Niu: “each chosen client trains the global model on her data and uploads the update of the global model together with the size of her training data to the cloud server.” [Section IV.B.1. Secure Federated Submodel Learning]; [Algorithm 1, Line 16-19]); aggregating, by the server computing system, the one or more model updates to produce an aggregate model update (Niu: “Finally, the weighted submodel updates and the count vectors from live clients are securely aggregated under the guidance of the cloud server (Lines 7–9 and 19). Specifically, the cloud server guides the secure aggregation by enumerating every index in the union of the chosen clients’ real index sets. For each index, the cloud server first determines the set of live clients whose perturbed index sets contain this index and then lets these clients submit the materials for securely adding up the weighted updates and the count numbers with respect to this index (Line 8). The cloud server finally applies the update to the global model in this row by adding the quotient of the sum of the weighted updates and the total count number, namely the weighted average (Line 9).” [Section IV.B.1. Secure Federated Submodel Learning]; [Algorithm 1, Lines 7-9]); and updating, by the server computing system, a machine-learned global classification model based at least in part on the aggregate model update (Niu: “The cloud server takes a weighted average of all updates, where one client’s weight is proportional to the size of her local data, and finally adds the aggregate update to the global model.” [Section IV.B.1. Secure Federated Submodel Learning]). Regarding Claim 14, Niu teaches the method of Claim 13, wherein the respective classification submodel comprises a feature extractor (Niu: “In our scheme, each chosen client generates three types of index sets locally: real, perturbed, and succinct. First, the real index is extracted from a client’s private data and is kept secret from the other system participants, including the cloud server and any other chosen client.” [Section I.E. Our Solution Overview and Major Contributions]; “Similarly, when federated submodel learning is applied to the natural language scenario(e.g., next word prediction in Gboard), a client’s real index set to locate her wanted parameters of word embedding is actually the vocabulary extracted from her typed texts.” [Section III.A. Details on Privacy Risks and Security Requirements]). Regarding Claim 15, Niu teaches the method of Claim 14, wherein the global classification model comprises the feature extractor (Niu: “We use a two-dimensional matrix with m rows and d columns to represent the global/full model, denoted as W. Such a matrix-based representation not only suffices for the recommendation models used in Alibaba but also can easily degenerate to a widely used vector-based representation [18], [30], by setting the number of columns d to 1. Additionally, we let S = {1, 2, … , m} denote the entire row index set of W. Moreover, we let C denote those clients who are selected by the cloud server to participate in one communication round of federated submodel learning. For a chose client i ϵ C, we let S(i) ⊂ S denote her real index set, which implies that the user data of client i involves the rows in W with indicies S(i).” [Section III. Preliminaries]). Regarding Claim 16, Niu teaches the method of Claim 14, wherein the global classifier model comprises a classifier weight matrix comprising class representations for the plurality of classes (Niu: “Additionally, to facilitate the cloud server in averaging submodel updates according to the sizes of relevant local training data, each client also needs to count the number of her samples involving every index in the perturbed index set (Line 17). In particular, the numbers of samples involving the indices outside the succinct index set are all zeros. Furthermore, each client prepares the submodel update to be uploaded by multiplying each row with the corresponding count number, namely the weight, in advance (Line18).” [Section IV.B.1. Secure Federated Submodel Learning]; “We use a two-dimesional matrix with m rows and d columns to represent the global/full model, denoted as W. Such a matrix-based representation not only suffices for the recommendation models used in Alibaba but also can easily degenerate to a widely used vector-based representation,” [Section III. Preliminaries]; “For example, each row of the embedding matrix for goods in the recommendation model is linked with a certain goods ID, which indicates that a client’s real index set, specifying her required rows of the embedding matrix, is in fact the goods IDs in her private data.” [Section III.A. Details on Privacy Risks and Security Requirements). Regarding Claim 17, Niu teaches the method of Claim 16, wherein the respective classification submodel comprises a sub classifier model comprising a sub classifier weight matrix, the sub classifier weight matrix comprising class representations only for classes in the client class set of a respective client computing system (Niu: “Additionally, to facilitate the cloud server in averaging submodel updates according to the sizes of relevant local training data, each client also needs to count the number of her samples involving every index in the perturbed index set (Line 17). In particular, the numbers of samples involving the indices outside the succinct index set are all zeros. Furthermore, each client prepares the submodel update to be uploaded by multiplying each row with the corresponding count number, namely the weight, in advance (Line18).” [Section IV.B.1. Secure Federated Submodel Learning]; “We use a two-dimesional matrix with m rows and d columns to represent the global/full model, denoted as W. Such a matrix-based representation not only suffices for the recommendation models used in Alibaba but also can easily degenerate to a widely used vector-based representation,” [Section III. Preliminaries]; “For example, each row of the embedding matrix for goods in the recommendation model is linked with a certain goods ID, which indicates that a client’s real index set, specifying her required rows of the embedding matrix, is in fact the goods IDs in her private data.” [Section III.A. Details on Privacy Risks and Security Requirements). Regarding Claim 18, Niu teaches the method of Claim 14, wherein the feature extractor comprises a neural network (Niu: “We start with the first problem. One trivial method is that the client downloads the full matrix, as in conventional federated learning, and then extracts the required row locally. Although this method perfectly hides the fetched row index, it incurs significant communication cost, which can be unaffordable for resource-constrained mobile devices, especially when the matrix is huge, e.g., representing a deep neural network.” [Section I.D. Fundamental Problems and Challenges]). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-12 and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Niu et al. (“Secure Federated Submodel Learning”, published 11/11/2019), hereinafter Niu; in view of Yu et al. (US 20210019654 A1, published 01/21/2021), hereinafter X. Yu. Regarding Claim 1, Niu teaches A computer-implemented method for federated learning of a machine-learned model with reduced computing resource usage (Niu: “our new framework further decouples the ability to accomplish federated learning from the need to use the prohibitively large full model, which can dramatically improve efficiency. For example, in our evaluation, the size of a client’s desired submodel is only 1.99% of the full model’s size. Thus, our framework is more practical for resource constrained clients and deep learning tasks.” [Section I. Introduction]), the method comprising: sampling, at a client computing system comprising one or more computing devices, one or more negative class labels from a negative sampling distribution, wherein the client computing system comprises a local training dataset, the local training dataset comprising a plurality of training examples each associated with one of a plurality of local positive class labels (Niu: “When the memoization technique is applied to federated submodel learning, we let client i maintain two index sets with “Yes” and “No” answers in the permanent randomized response, respectively (Input). Here, the permanent randomized response mechanism is parameterized by two probabilities p1(i), p2(i) to tune privacy and utility (Lines 3–6)” [Section IV.B.2. Index Set Perturbation]; [Algorithm 1, Lines 10-12]; “First, the real index is extracted from a client’s private data and is kept secret from the other system participants, including the cloud server and any other chosen client. Second, the perturbed index set is used to interact with others in the download and upload phases. It is generated by applying randomized response twice with one memoization step between. Such a design, together with secure aggregation, allows the client to hold a self-controllable deniability against whether she really intends or does not intend to download some row and to upload the modification of this row, even if the client may be chosen to participate in multiple communication rounds. The strength of deniability is rigorously quantified using local differential privacy. Further, rather than trivially using the prohibitively large-scale full index set as the questionnaire of randomize response in every communication round, we identify a necessary and sufficient index set, namely the union of the chosen clients’ real index sets. Considering the secrecy of each client’s real index set, we propose an efficient and scalable Private Set Union (PSU) protocol based on Bloom filter, secure aggregation, and randomization, allowing clients to obtain the union under the mediation of an untrusted cloud server without revealing any individual real index set.” [Section I.E. Out Solution Overview and Major Contributions]); communicating, by the client computing system, a client class set comprising a union of the one or more negative class labels with the plurality of local positive class labels to a server computing system, the server computing system comprising a current version of a machine- learned global classification model configured to provide class labels associated with a plurality of candidate classes (Niu: “First, the real index is extracted from a client’s private data and is kept secret from the other system participants, including the cloud server and any other chosen client. Second, the perturbed index set is used to interact with others in the download and upload phases. It is generated by applying randomized response twice with one memoization step between. Such a design, together with secure aggregation, allows the client to hold a self-controllable deniability against whether she really intends or does not intend to download some row and to upload the modification of this row, even if the client may be chosen to participate in multiple communication rounds.” [Section I.E. Our Solution Overview and Major Contributions]; [Algorithm 1, Lines 10-12]; “Given the questionnaire, client i basically uses two probability parameters p1(i), p2(i) in randomized response to fine-tune the tension among effectiveness, efficiency, and privacy (Lines 3–6). In particular, p1(i) denotes the probability that an index in client i’s real index set will return a “Yes” answer and controls the factual size of a client’s user data contributed to federated submodel learning. Thus, a larger p1(i) implies better effectiveness in terms of convergency rate. In addition, among privacy, effectiveness, and efficiency. More specifically, p2(i) denotes the probability that an index outside client i’s real index set will return a “Yes” answer and determines the number of redundant rows to be downloaded and the number of padded zero vectors to be uploaded through the secure aggregation protocol.” [Section IV.B.2. Index Set Perturbation]; “At the beginning of one communication round, the cloud server sends the up-to-date parameters of the global model and the training hyperparameters to some clients.” [Section IV.B.1. Secure Federated Submodel Learning]); receiving, by the client computing system, data descriptive of a classification submodel, the classification submodel configured to provide a classification output limited to classes of the client class set (Niu: “Depending on the intersection of the real index set and the perturbed index set, called the succinct index set, the client extracts a succinct submodel and prepares involved data as the succinct training set (Line 14).” [Section IV.B.1. Secure Federated Submodel Learning]; [Algorithm 1, Lines 13-14]); determining, by the client computing system, a model update based at least in part on use of the classification submodel with the local training dataset at the client computing system (Niu: “each chosen client trains the global model on her data and uploads the update of the global model together with the size of her training data to the cloud server.” [Section IV.B.1. Secure Federated Submodel Learning]; [Algorithm 1, Line 15]); and communicating, by the client computing system, the model update to the server computing system, wherein the current version of the machine-learned global classification model is updated based at least in part on the model update (Niu: “each chosen client trains the global model on her data and uploads the update of the global model together with the size of her training data to the cloud server. The cloud server takes a weighted average of all updates, where one client’s weight is proportional to the size of her local data, and finally adds the aggregate update to the global model.” [Section IV.B.1. Secure Federated Submodel Learning]; [Algorithm 1, Line 16-19]). However, Niu fails to expressly disclose determining a model update based at least in part on a sampled SoftMax loss, the sampled SoftMax loss determined at least in part on use of the classification model. In the same field of endeavor, X. Yu teaches determining a model update based at least in part on a sampled SoftMax loss, the sampled SoftMax loss determined at least in part on use of the classification model (X. Yu: “The training computing system 150 can include a model trainer 160 that trains the machine-learned models 120 and/or 140 stored at the user computing device 102 and/or the server computing system 130 using various training or learning techniques, such as, for example, backwards propagation of errors. For example, a loss function can be backpropagated through the model(s) to update one or more parameters of the model(s) (e.g., based on a gradient of the loss function).” [0058]; “the sampled softmax distribution can be defined as p i ' = e o i ' Z ' , where Z ' = ∑ j = 1 m + 1 e o j ' . The sampled softmax loss can correspond to the cross entropy loss with respect to the sampled softmax distribution: L = log ⁡ p t ' = - o t + log ⁡ Z ' .” [0025]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have incorporated determining a model update based at least in part on a sampled SoftMax loss, the sampled SoftMax loss determined at least in part on use of the classification model, as taught by X. Yu to the method of Niu because both of these methods are directed toward machine learning training optimization with a reduced sampled class set. X. Yu discusses training submodels on sets of local client data but is agnostic to what loss function is employed to do so. In making this combination and employing a sampled softmax method for model updating, it would allow the method of X. Yu to “speed up training” “for cases where a large number of classes are involved”, such as those applicable to the large set of clients in X. Yu’s federated learning architecture (X. Yu: [0023]). Regarding Claim 2, Niu and X. Yu teach the method of Claim 1, wherein the negative sampling distribution comprises a uniform distribution over the plurality of candidate classes not included in the plurality of local positive class labels (Niu: “Given the questionnaire, client i basically uses two probability parameters p1(i), p2(i) in randomized response to fine-tune the tension among effectiveness, efficiency, and privacy (Lines 3–6). In particular, p1(i) denotes the probability that an index in client i’s real index set will return a “Yes” answer and controls the factual size of a client’s user data contributed to federated submodel learning. Thus, a larger p1(i) implies better effectiveness in terms of convergency rate. In addition, among privacy, effectiveness, and efficiency. More specifically, p2(i) denotes the probability that an index outside client i’s real index set will return a “Yes” answer and determines the number of redundant rows to be downloaded and the number of padded zero vectors to be uploaded through the secure aggregation protocol.” [Section IV.B.2. Index Set Perturbation]; “Intuitively, the above definition says that the output distribution of the randomized mechanism does not change too much, given distinct inputs from the client. Thus, local differential privacy formalizes a sort of plausible deniability: no matter what output is revealed, it is approximately equally as likely to have come from one input as any other input.” [Section III.A. Details on Privacy Risks and Security Requirements]). Regarding Claim 3, Niu and X. Yu teach the method of Claim 1, wherein the client class set is anonymized such that the server computing system cannot discern the plurality of local positive class labels from the one or more negative class labels (Niu: “First, the real index is extracted from a client’s private data and is kept secret from the other system participants, including the cloud server and any other chosen client. Second, the perturbed index set is used to interact with others in the download and upload phases. It is generated by applying randomized response twice with one memoization step between. Such a design, together with secure aggregation, allows the client to hold a self-controllable deniability against whether she really intends or does not intend to download some row and to upload the modification of this row, even if the client may be chosen to participate in multiple communication rounds. The strength of deniability is rigorously quantified using local differential privacy. Further, rather than trivially using the prohibitively large-scale full index set as the questionnaire of randomize response in every communication round, we identify a necessary and sufficient index set, namely the union of the chosen clients’ real index sets. Considering the secrecy of each client’s real index set, we propose an efficient and scalable Private Set Union (PSU) protocol based on Bloom filter, secure aggregation, and randomization, allowing clients to obtain the union under the mediation of an untrusted cloud server without revealing any individual real index set.” [Section I.E. Out Solution Overview and Major Contributions]). Regarding Claim 4, Niu and X. Yu teach the method of Claim 1, wherein the global classification model comprises a feature extractor and a classifier model (Yu: “In some embodiments, the machine-learned model can be configured for image classification. For example, inputs associated with image classification can include raw and/or processed image data. For instance, the image data can be input into a feature extraction algorithm to extract input features from the image data.” [0075]). Regarding Claim 5, Niu and X. Yu teach the method of Claim 4, wherein the classification submodel comprises the feature extractor (Niu: “each chosen client downloads part of the global model as she requires, namely a submodel, from the cloud server. For example, in the e-commerce scenario above, a client’s submodel mainly consists of the embedding parameters for the displayed and clicked goods in her historical data, as well as the parameters of the other network layers.” [Section I.B. Framework of Federated Submodel Learning]). Regarding Claim 6, Niu and X. Yu teach the method of Claim 4, wherein the classifier model comprises a classifier weight matrix comprising class representations for the plurality of classes (Niu: “Additionally, to facilitate the cloud server in averaging submodel updates according to the sizes of relevant local training data, each client also needs to count the number of her samples involving every index in the perturbed index set (Line 17). In particular, the numbers of samples involving the indices outside the succinct index set are all zeros. Furthermore, each client prepares the submodel update to be uploaded by multiplying each row with the corresponding count number, namely the weight, in advance (Line18).” [Section IV.B.1. Secure Federated Submodel Learning]; “We use a two-dimesional matrix with m rows and d columns to represent the global/full model, denoted as W. Such a matrix-based representation not only suffices for the recommendation models used in Alibaba but also can easily degenerate to a widely used vector-based representation,” [Section III. Preliminaries]; “For example, each row of the embedding matrix for goods in the recommendation model is linked with a certain goods ID, which indicates that a client’s real index set, specifying her required rows of the embedding matrix, is in fact the goods IDs in her private data.” [Section III.A. Details on Privacy Risks and Security Requirements). Regarding Claim 7, Niu and X. Yu teach the method of Claim 6, wherein the classification submodel comprises a sub classifier model comprising a sub classifier weight matrix, the sub classifier weight matrix comprising class representations only for classes in the client class set (Niu: “Additionally, to facilitate the cloud server in averaging submodel updates according to the sizes of relevant local training data, each client also needs to count the number of her samples involving every index in the perturbed index set (Line 17). In particular, the numbers of samples involving the indices outside the succinct index set are all zeros. Furthermore, each client prepares the submodel update to be uploaded by multiplying each row with the corresponding count number, namely the weight, in advance (Line18).” [Section IV.B.1. Secure Federated Submodel Learning]; “We use a two-dimesional matrix with m rows and d columns to represent the global/full model, denoted as W. Such a matrix-based representation not only suffices for the recommendation models used in Alibaba but also can easily degenerate to a widely used vector-based representation,” [Section III. Preliminaries]; “For example, each row of the embedding matrix for goods in the recommendation model is linked with a certain goods ID, which indicates that a client’s real index set, specifying her required rows of the embedding matrix, is in fact the goods IDs in her private data.” [Section III.A. Details on Privacy Risks and Security Requirements). Regarding Claim 8, Niu and X. Yu teach the method of Claim 4, wherein the feature extractor comprises a neural network (Niu: “We start with the first problem. One trivial method is that the client downloads the full matrix, as in conventional federated learning, and then extracts the required row locally. Although this method perfectly hides the fetched row index, it incurs significant communication cost, which can be unaffordable for resource-constrained mobile devices, especially when the matrix is huge, e.g., representing a deep neural network.” [Section I.D. Fundamental Problems and Challenges]). Regarding Claim 9, Niu and X. Yu teach the method of Claim 1, wherein the local training dataset comprises image data (X. Yu: “For instance, cross-entropy loss based on the softmax function can be used in multi-class classification tasks such as natural language processing, image classification, and recommendation systems.” [0020]; “For example, the input features may include features used for natural language processing, such as raw or processed linguistic information. As another example, the input features may include image classification features, such as raw or processed images. As another example, the input features may include features used for content recommendation services, such as web usage or other suitable information.” [0060]). Regarding Claim 10, Niu and X. Yu teach the method of Claim 1, wherein the sampled SoftMax loss comprises a sum of an adjusted logit for a true class label and a logarithm of a sum of the exponents of adjusted logits for all classes in the client class set (X. Yu: “The distribution in the full softmax function is referred to as the softmax distribution. Given a collection of inputs and their true labels, the objective is to identify the model parameters by minimizing the cross-entropy loss based on the softmax function or the full softmax loss L = - log ⁡ p t = - o t + log ⁡ Z , where tϵ[n] denotes the true label or true class for the input x.” [0021]; “Formally, let the number of sampled classes during each iteration be m, which class I being picked with probability qi, Let N t ≜ [ n ] { t } be the set of negative classes. Assuming that s1, … , sm ∈ N t denote the sampled class indices, adjusted logits o ' = ( o ' 1 ,   o ' 2 ,   . . . ,   o ' m + 1 } can be defined such that o ' 1 = o t and for i ∈ m ,   o ' i + 1 = o s i - log ⁡ ( m q s i ) . ” [0024]). Regarding Claim 11, Niu and X. Yu teach the method of Claim 9, wherein the adjusted logits comprise the sum of a full SoftMax logit and the logarithm of a number of negative classes multiplied by a sampling probability (X. Yu: “Formally, let the number of sampled classes during each iteration be m, which class I being picked with probability qi, Let N t ≜ [ n ] { t } be the set of negative classes. Assuming that s1, … , sm ∈ N t denote the sampled class indices, adjusted logits o ' = ( o ' 1 ,   o ' 2 ,   . . . ,   o ' m + 1 } can be defined such that o ' 1 = o t and for i ∈ m ,   o ' i + 1 = o s i - log ⁡ ( m q s i ) . Accordingly, the sampled softmax distribution can be defined as p i ' = e o i ' Z ' , where Z ' = ∑ j = 1 m + 1 e o j ' . The sampled softmax loss can correspond to the cross entropy loss with respect to the sampled softmax distribution: L = log ⁡ p t ' = - o t + log ⁡ Z ' .” [0024]-[0025]). Regarding Claim 12, Niu and X. Yu teach the method of Claim 1, wherein the client computing system comprises a mobile device (Niu: “We can see that this protocol is quite efficient for large-scale data vectors, especially from communication overhead, and thus can apply to mobile applications.” [Section III.B.2. Secure Aggregation]). Regarding Claim 19, Niu teaches the method of Claim 13. However, it fails to expressly disclose wherein the one or more model updates are determined with respect to a sampled SoftMax loss, wherein the sampled SoftMax loss comprises a sum of an adjusted logit for a true class label and a logarithm of a sum of the exponents of adjusted logits for all classes in the client class set of a respective client computing system. In the same field of endeavor, X. Yu teaches wherein the one or more model updates are determined with respect to a sampled SoftMax loss, wherein the sampled SoftMax loss comprises a sum of an adjusted logit for a true class label and a logarithm of a sum of the exponents of adjusted logits for all classes in the client class set of a respective client computing system (X. Yu: “The distribution in the full softmax function is referred to as the softmax distribution. Given a collection of inputs and their true labels, the objective is to identify the model parameters by minimizing the cross-entropy loss based on the softmax function or the full softmax loss L = - log ⁡ p t = - o t + log ⁡ Z , where tϵ[n] denotes the true label or true class for the input x.” [0021]; “Formally, let the number of sampled classes during each iteration be m, which class I being picked with probability qi, Let N t ≜ [ n ] { t } be the set of negative classes. Assuming that s1, … , sm ∈ N t denote the sampled class indices, adjusted logits o ' = ( o ' 1 ,   o ' 2 ,   . . . ,   o ' m + 1 } can be defined such that o ' 1 = o t and for i ∈ m ,   o ' i + 1 = o s i - log ⁡ ( m q s i ) . ” [0024]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have incorporated wherein the one or more model updates are determined with respect to a sampled SoftMax loss, wherein the sampled SoftMax loss comprises a sum of an adjusted logit for a true class label and a logarithm of a sum of the exponents of adjusted logits for all classes in the client class set of a respective client computing system, as taught by X. Yu to the method of Niu because both of these methods are directed toward machine learning training optimization with a reduced sampled class set. X. Yu discusses training submodels on sets of local client data but is agnostic to what loss function is employed to do so. In making this combination and employing a sampled softmax method for model updating, it would allow the method of X. Yu to “speed up training” “for cases where a large number of classes are involved”, such as those applicable to the large set of clients in X. Yu’s federated learning architecture (X. Yu: [0023]). Regarding Claim 20, Niu and X. Yu teach the method of Claim 19, wherein the adjusted logits comprise the sum of a full SoftMax logit and the logarithm of a number of negative classes multiplied by a sampling probability (X. Yu: “Formally, let the number of sampled classes during each iteration be m, which class I being picked with probability qi, Let N t ≜ [ n ] { t } be the set of negative classes. Assuming that s1, … , sm ∈ N t denote the sampled class indices, adjusted logits o ' = ( o ' 1 ,   o ' 2 ,   . . . ,   o ' m + 1 } can be defined such that o ' 1 = o t and for i ∈ m ,   o ' i + 1 = o s i - log ⁡ ( m q s i ) . Accordingly, the sampled softmax distribution can be defined as p i ' = e o i ' Z ' , where Z ' = ∑ j = 1 m + 1 e o j ' . The sampled softmax loss can correspond to the cross entropy loss with respect to the sampled softmax distribution: L = log ⁡ p t ' = - o t + log ⁡ Z ' .” [0024]-[0025]). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Li et al. (“Sample-level Data Selection for Federated Learning”) discusses constructing a global learning model through federated learning without sharing local training data with the remote server, aiming to solve the exigent optimization problem that selects high quality training samples for a given FL task. Any inquiry concerning this communication or earlier communications from the examiner should be directed to MEGAN E HWANG whose telephone number is (703)756-1377. The examiner can normally be reached Monday-Thursday 10:00AM-7:30PM ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jennifer Welch can be reached at (571) 272-7212. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /M.E.H./Examiner, Art Unit 2143 /JENNIFER N WELCH/Supervisory Patent Examiner, Art Unit 2143
Read full office action

Prosecution Timeline

Jan 12, 2024
Application Filed
Aug 21, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737670
TRANSFORMATION OF DATA FROM LEGACY ARCHITECTURE TO UPDATED ARCHITECTURE
4y 11m to grant Granted Sep 15, 2026
Patent 12737693
LEARNING DEVICE, LEARNING METHOD, AND COMPUTER-READABLE STORAGE MEDIUM
4y 4m to grant Granted Sep 15, 2026
Patent 12718153
SURROGATE MODEL FOR TIME-SERIES MODEL INTERPRETATION
4y 4m to grant Granted Aug 25, 2026
Patent 12699918
SYSTEMS AND METHODS FOR PHOTOVOLTAIC FAULT DETECTION USING A FEEDBACK-ENHANCED POSITIVE UNLABELED LEARNING
4y 9m to grant Granted Aug 04, 2026
Patent 12682604
A GENERIC MODULAR SPARSE THREE-DIMENSIONAL (3D) CONVOLUTION DESIGN UTILIZING SPARSE 3D GROUP CONVOLUTION
4y 10m to grant Granted Jul 14, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
54%
Grant Probability
99%
With Interview (+57.5%)
4y 0m (~1y 3m remaining)
Median Time to Grant
Low
PTA Risk
Based on 33 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month