Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This action is in response to the amendment and remarks filed May 11th, 2026. In the
amendment, claims 1, 3-6, 9-15, and 17-20 were amended and no claims were cancelled or
added. As such, claims 1-20 are pending.
Response to Arguments
Applicant’s arguments, see Page 12, filed March May 11th 2026, with respect to the claim objections have been considered and are persuasive. Amendments to the claims obviate the
objections of record. The objections of claims have been withdrawn. See updated objections below.
Applicant’s arguments with respect to the rejections of claims 1-20 under 35 U.S.C § 101 and 103 are not persuasive for the following reasons:
35 U.S.C. 101:
Applicant argues that amended claim 1 is directed to a transfer learning pipeline in which an initial loss distribution is computed prior to fine-tuning and a regularization term is applied during fine-tuning to penalize divergence, and therefore does not recite a mental process or a mathematical concept (Pages 12-13 of Remarks). The Examiner respectfully disagrees. As set forth in the updated rejection, the recited steps of computing an initial loss distribution, computing a batch loss and its loss distribution, computing a divergence metric, multiplying the divergence by a hyperparameter, and adding the result as a regularizer are each mathematical calculations that fall within the mathematical concepts grouping under MPEP 2106.04(a). The recited objective of preventing data leakage from membership inference attacks states an intended result of protecting data privacy and falls within the mental process grouping as previously set forth. Describing the ordered performance of these mathematical operations before and during fine-tuning as a "pipeline" does not remove the underlying limitations from the abstract idea groupings. Therefore, claim 1 recites an abstract idea under Step 2A Prong One.
Applicant argues that even if the claim recites an abstract idea, the alleged abstract idea is integrated into a practical application because the invention provides an improvement in preventing data leakage from membership inference attacks and enhances robustness against attacks citing paragraphs 4, 39, 41-42, 49, and 59 from the specification (Pages 13-15 of Remarks). The Examiner respectfully disagrees. Making member and non-member loss distributions less distinguishable so that an adversary cannot infer membership is an improvement to the result produced by the recited abstract ideas. The claim does not recite any specific technical mechanism that improves how the computer or the model operates such as improving the machine learning model, memory system, or processor. Instead, the claim applies the recited regularization math steps to training model data. The additional elements of training and updating the model amount to mere instructions to apply the judicial exception on a computer under MPEP 2106.05(f) and do not integrate the exception into a practical application.
Applicant argues that the claimed method improves computational efficiency because it achieves the same accuracy in fewer epochs, saving computer resources and time (Page 14 of Remarks). The Examiner respectfully disagrees. The claim does not recite achieving the same accuracy in fewer epochs, reducing the number of epochs, or any other mechanism for conserving computer resources. An alleged advantage that appears only in the specification but is not reflected in the claim cannot integrate the exception into a practical application. The reduction in training time is a result of performing the regularization math steps and represents an improvement to the abstract idea itself rather than a specific technical improvement to the computer. The claim should recite how the recited steps as a whole lead to an improved computational efficiency rather than pointing out to an alleged advantage that only appears in the specification and is not reflected in the claims.
Applicant argues that the claim addresses a technical problem because trained models are vulnerable to membership inference attacks due to the loss signal, and that adding the regularizer removes the statistical signal that attackers rely on, establishing a nexus between the claim language and a practical implementation (Page 15 of Remarks). The Examiner respectfully disagrees. Reducing the distinguishability of member and non-member losses is an improvement to the privacy of the abstract idea and does not impose a meaningful limit that integrates the exception into a practical application. The applicant also has not shown that the claim is directed to significantly more than the abstract idea.
Applicant argues that the features of amended claim 1 add specific limitations that are not well-understood, routine, or conventional, and that the claim elements taken individually and in combination amount to significantly more than the alleged abstract idea (Page 16 of Remarks). The Examiner respectfully disagrees. See the updated 101 rejections in light of the amendments.
Applicant argues that the specification at paragraph 64 expressly discloses that the computer-readable storage medium is not to be construed as transitory signals per se, and that claims 19 and 20 therefore recite statutory subject matter (Pages 16-17 of Remarks). The Examiner respectfully disagrees. Claim 19 has not been amended and still recites "one or more computer-readable storage devices" rather than "non-transitory computer-readable storage media". The cited paragraph describes "computer-readable storage media" as excluding transitory signals, but the claim still recites "storage devices". Because the claim term still encompasses subject matter that does not fall within one of the four statutory categories, claims 19 and 20 remain non-statutory. The applicant may overcome this rejection by amending "computer-readable storage devices" to "non-transitory computer-readable storage media device" or by changing “computer-readable storage devices” to “computer-readable storage media”.
35 U.S.C. 103:
Applicant argues that Carlini in view of Chen does not teach computing, based on querying the pre-trained model with private data, an initial loss distribution LINIT of a plurality of loss values that is computed prior to beginning of a fine-tuning operation (Pages 17-19 of Remarks). The Examiner respectfully disagrees. Carlini establishes Q_out as the baseline distribution of losses for data points on which the model has not been trained, which corresponds to the claimed LINIT (Page 4 Section IV of Carlini). Chen teaches querying the pre-trained model with private data by performing a forward pass on member and non-member query samples (Page 3 Algorithm 1 of Chen). Newly cited Leidner further teaches that an operation is computed prior to beginning the fine-tuning operation (Claim 11 of Leidner, reciting pre-training performed prior to fine-tuning). The combination therefore teaches computing LINIT from querying the pre-trained model with private data before fine-tuning begins.
Applicant argues that Carlini does not teach computing a batch loss of a minibatch from the private data after fine-tuning begins for the same model, computing LBATCH, and further does not teach a divergence metric between LINIT and LBATCH (Page 19 of Remarks). The Examiner respectfully disagrees. Carlini defines Q_in as the distribution of losses for a model trained on datasets containing the example (x,y), which characterizes the loss behavior of the same model after its parameters are updated with the private data, which corresponds to the claimed LBATCH (Page 4 Section IV of Carlini). Zhang teaches computing the batch loss of a minibatch under mini-batch SGD by averaging losses across selected samples (Paragraph 111 of Zhang), and Carlini teaches the divergence metric through the likelihood-ratio test that quantifies the statistical distance between Q_in and Q_out (Page 4 Equation 2 of Carlini).
Applicant argues that Chen and Zhang do not remedy the alleged deficiencies of Carlini (Page 19 of Remarks). The Examiner respectfully disagrees. Chen supplies the features Carlini does not expressly recite, including receiving a pre-trained model and hyperparameter λ, querying the pre-trained model with private data via a forward pass, and applying a hyperparameter-scaled divergence penalty as a regularizer during backpropagation (Section 7 Page 9 and Algorithm 1 of Chen). Zhang supplies the minibatch batch loss computation (Paragraph 111 of Zhang), and newly cited Leidner supplies the "prior to fine-tuning" timing (Claim 11 of Leidner). Together, the references teach each limitation and an alleged deficiency of any single reference does not establish non-obviousness of the combination.
Applicant argues that Carlini cannot be combined with Chen and Zhang because Carlini evaluates membership inference using loss-based comparisons while Chen applies a divergence-based regularizer and Zhang merely teaches minibatch loss (Pages 19-20 of Remarks). The Examiner respectfully disagrees. A reference is analogous art if it is in the same field of endeavor as the claimed invention, or is reasonably pertinent to the particular problem with which the inventor was concerned. Carlini is in the same field of endeavor as the claimed method of training a machine learning model to prevent data leakage from membership inference attacks, being directed to measuring the membership privacy leakage of trained models (Section 1 of Carlini). Chen is likewise in the same field, being directed to a training framework that defends membership inference attacks (Introduction of Chen). Zhang is in the same field of endeavor of training machine learning models, disclosing training a neural network by mini-batch stochastic gradient descent on a loss function including a hyperparameter-scaled regularization term (Paragraphs 96-97 of Zhang). Each reference is analogous to the claimed invention, and does not render the combination as improper.
Applicant argues that even in combination the references do not teach computing LINIT from the pre-trained model with private data prior to fine-tuning, computing LBATCH from minibatch loss values during fine-tuning, and computing a divergence metric between them (Page 20 of Remarks). The Examiner respectfully disagrees. As mapped in the rejection, Carlini teaches LINIT (Q_out) and the divergence metric (likelihood-ratio test), Chen teaches querying the pre-trained model with private data and the regularizer, Zhang teaches the minibatch batch loss, and Leidner teaches performing the initial computation prior to fine-tuning. Each recited limitation is therefore accounted for by the combined teachings. Applicant again separates the references rather than addressing what the combination teaches as a whole.
Applicant argues that the dependent claims are patentable by virtue of their dependence on claim 1 and that each separately recites subject matter not taught (Pages 20-21 of Remarks). The Examiner respectfully disagrees. Since claim 1 remains properly rejected over Carlini, Chen, Zhang, and Leidner for the reasons above, the dependency argument does not overcome the rejection. Claims 1-20 remain rejected under 35 U.S.C. 103.
Specification
The disclosure is objected to because of the following informalities:
Paragraph 2: “using transfer learning.. .” contains multiple periods and should read “using transfer learning.”
Paragraph 6: the two sentences “…a computer-implemented method of training a machine learning model to prevent data leakage from membership inference attacks. A pre-trained model and a pre-defined hyperparameter X are received as an input.” are duplicated and should be recited only once.
Paragraph 7: “+In one or more embodiments” should read “In one or more embodiments”.
Paragraph 42: “These changes provide for a more robust defense against MIAs, and provides for a more efficient user of computer resources,” should read “provides for a more efficient use of computer resources”
Claim Objections
Claims 1, 10, 11, and 19 and their respective dependent claims are objected to because of the following informalities:
Claim 1 recites “wherein the initial loss distribution LINIT is computed prior to beginning of a fine-tuning operation,” which should read “LINIT is computed prior to a beginning of a fine-tuning operation”.
Claim 10 recites “to change a privacy-robustness values of the pre-trained model,” which should read “to change privacy-robustness values of the pre-trained model”.
Claims 11 and 19 each recite “computing a batch loss of a minibatch from the private data after the beginning the fine-tuning operation,” which should read “after the beginning of the fine-tuning operation”
Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 19 and 20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. The claims do not fall within at least one of the four categories of patent eligible subject matter because claim 19 recites a computer program product comprising “one or more computer-readable storage devices” storing instructions executable by a processor. While non-transitory computer-readable storage media are considered a statutory manufacture, the claim as written uses the term “devices,” which could encompass transitory signals, and therefore doesn’t fall under any of the four statutory categories of machine, manufacture, composition of matter, or process. The specification at paragraph 64 only describes “computer-readable storage media” as non-transitory media. Claim 19 and 20 should be rejected as non-statutory. The applicant may overcome this issue by amending “devices” to “non-transitory computer-readable storage media” to clearly claim a statutory manufacture.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an
abstract idea without significantly more.
Claim 1
Step 1: The claim recites a method; therefore, it is directed to the statutory category of a process.
Step2A Prong 1: The claim recites, inter alia:
computing, based on the querying of the pre-trained model with the private data, an initial loss distribution LINIT of a plurality of loss values…: This limitation is a mathematical concept because it deals with performing mathematical calculations to derive LINIT. See Paragraph 65 which states “The modeling of the LINIT is performed by computing the mean and variance of the logit-scaled loss values. The mean and the variance aid in providing a Gaussian variable. The logit scaling also aids in obtaining a Gaussian random variable of the minibatch losses to make an MIA attack more difficult for an adversary.”.
computing a batch loss of a minibatch from the private data after the beginning of the fine-tuning operation: This is a mathematical concept because the batch loss is computed by applying a loss function over the minibatch which applies math equations in order to compute the loss.
computing a loss distribution LBATCH of the batch loss: This limitation is a mathematical concept because it deals with calculating the mean and variance between the batches. See Paragraph 9 that states, “…the modeling of the LBATCH as a second Gaussian Random Variable is performed by computing the mean and variance of the logit-scaled batch loss.”.
computing a divergence metric between the initial loss distribution LINIT and the loss distribution LBATCH: This limitation is viewed as a mathematical concept because it deals with applying Kullback-Leibler (KL) divergence method to measure the distance between two probability distributions dealing with math calculations. See Paragraph 10 which states, “In one or more embodiments, the computing of the divergence metric between LNIT and LBATCH includes using a Kullback-Leibler divergence to measure a distance between LINIT and LBATCH. Kullback-Leibler (KL) divergence is a method useful to measure a distance between two probability distributions.”.
multiplying an output of the divergence metric with the pre-defined hyperparameter X to obtain a result: This limitation is viewed as a mathematical concept since it deals with multiplying two different numbers to get a result.
adding the result to the batch loss as a regularizer to penalize a distance between the initial loss distribution LINIT and the loss distribution LBATCH: This limitation is a mathematical concept dealing with the addition of the result from the previous step and the batch loss to obtain a regularizer.
…computing backpropagation on the batch loss to which the result is added as the regularizer: This limitation is a mathematical concept because it involves computing derivatives and loss functions during the backpropagation process.
Step2A Prong 2: This judicial exception is not integrated into a practical application because the
additional elements are as follows:
A computer-implemented method of training a machine learning model to prevent data leakage from membership inference attacks, the computer-implemented method comprising… updating model parameters of the pre-trained model by…: Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
receiving a pre-trained model and a pre-defined hyperparameter X as an input for machine learning: Mere data gathering recited at a high level of generality, and thus is an insignificant extra-solution activity (MPEP 2106.05(g)).
applying a forward pass by querying the pre-trained model with private data: The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)).
wherein the initial loss distribution LINIT is computed prior to beginning of a fine-tuning operation: The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)).
wherein the fine-tuning operation is performed to transform the pre-trained model into a fine-tuned model: Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
and outputting the fine-tuned model based on the updating: Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly
more than the judicial exception because the additional elements are as follows:
…training a machine learning model… updating model parameters of the pre-trained model by…: Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and cannot provide inventive concept (MPEP 2106.05(f)).
receiving a pre-trained model and a pre-defined hyperparameter X as an input for machine learning: The additional element of “receiving” does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of receiving steps amounts to no more than mere data gathering. This element amounts to receiving data over a network and are well-understood, routine, conventional activity. See MPEP 2106.05(d), subsection II (i). This cannot provide an inventive concept.
applying a forward pass by querying the pre-trained model with private data: The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself which cannot provide inventive concept (MPEP 2106.05(h)).
wherein the initial loss distribution LINIT is computed prior to beginning of a fine-tuning operation: The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself which cannot provide inventive concept (MPEP 2106.05(h)).
wherein the fine-tuning operation is performed to transform the pre-trained model into a fine-tuned model: Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and cannot provide inventive concept (MPEP 2106.05(f)).
and outputting the fine-tuned model based on the updating: Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and cannot provide inventive concept (MPEP 2106.05(f)).
The elements in combination as an ordered whole still do not amount to significantly more than the judicial exception (i.e., the abstract ideas of mental processes and mathematical concepts for preventing data leakage from membership inference attacks). The claim merely describes applying known mathematical operations to machine learning training data, including computing loss distributions, calculating statistical parameters such as mean and variance, measuring divergence between probability distributions, and performing backpropagation to update model parameters. The recitation of training and fine-tuning a machine learning model, applying a forward pass, receiving a pre-trained model and hyperparameters, and outputting a fine-tuned model merely describes generic computer implementation and routine data gathering and output, without improving the functioning of a computer or machine learning model itself.
Claim 2
Step 1: A process, as above.
Step2A Prong 1: The claim recites, inter alia:
applying logit-scaling to the plurality of loss values: This limitation is seen as a mathematical concept because logit-scaling is a math operation applied to numerical values.
Step 2A Prong Two and Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B. Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2. The claim is ineligible.
Even when considered in combination, these additional elements represent mere instructions to apply an exception and therefore does not provide an inventive concept. The claim is ineligible.
Claim 3
Step 1: A process, as above.
Step2A Prong 1: The claim recites, inter alia:
performing a modeling of the initial loss distribution LINIT as a first Gaussian Random Variable by computing a mean and a variance of the plurality of loss values, wherein the modeling is performed based on the applied logit-scaling to the plurality of loss values: This limitation is seen as a mathematical concept because it recites the application of statistical formulas to numeric data including calculating the mean and variance. The limitation additionally recites applying logit-scaling to the plurality of loss values, which involves performing a math operation on numerical values.
Step 2A Prong Two and Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B. Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2. The claim is ineligible.
Even when considered in combination, these additional elements represent mere instructions to apply an exception and therefore does not provide an inventive concept. The claim is ineligible.
Claim 4
Step 1: A process, as above.
Step2A Prong 1: The claim recites, inter alia:
applying logit-scaling to the batch loss and performing a modeling of the loss distribution LBATCH as a second Gaussian Random Variable by computing a mean and a variance of the batch loss, wherein the modeling is performed based on the applied logit-scaling to the batch loss: This limitation is seen as a mathematical concept because it recites performing logit-scaling, computing a mean and variance, which involve math operations and statistical formulas.
Step 2A Prong Two and Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B. Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2. The claim is ineligible.
Even when considered in combination, these additional elements represent mere instructions to apply an exception and therefore does not provide an inventive concept. The claim is ineligible.
Claim 5
Step 1: A process, as above.
Step2A Prong 1: The claim recites, inter alia:
the computing of the divergence metric between the initial loss distribution LINIT and the loss distribution LBATCH comprises using a Kullback-Leibler divergence to measure the distance between the initial loss distribution LINIT and the loss distribution LBATCH: This limitation is seen as a mathematical concept because it recites performing a statistical comparison between two probability distributions.
Step 2A Prong Two and Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B. Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2. The claim is ineligible.
Even when considered in combination, these additional elements represent mere instructions to apply an exception and therefore does not provide an inventive concept. The claim is ineligible.
Claim 6
Step 1: A process, as above.
Step2A Prong 1: The claim recites, inter alia:
determining, prior to outputting the fine-tuned model, that the fine-tuned model meets a termination criterion: This limitation is seen as a mental process because it involves the determination of that a model meets a certain criterion.
Step2A Prong 2: This judicial exception is not integrated into a practical application because the
additional elements are as follows:
the updating of the model parameters is performed by using an update rule, and the computer-implemented method further comprising: Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly
more than the judicial exception because the additional elements are as follows:
the updating of the model parameters is performed by using an update rule, and the computer-implemented method further comprising: Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and cannot provide inventive concept (MPEP 2106.05(f)).
Even when considered in combination, these additional elements represent mere instructions to apply an exception and therefore do not provide an inventive concept. The claim is ineligible.
Claim 7
Step 1: A process, as above.
Step2A Prong 1: The claim recites, inter alia:
computing of the backpropagation to update the model parameters is performed using a Stochastic Gradient Descent (SGD): This limitation is seen as a mathematical concept because it recites an algorithmic procedure for adjusting numeric model parameters based on computed gradients.
Step 2A Prong Two and Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B. Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2. The claim is ineligible.
Even when considered in combination, these additional elements represent mere instructions to apply an exception and therefore does not provide an inventive concept. The claim is ineligible.
Claim 8
Step 1: A process, as above.
Step2A Prong 1: The claim recites, inter alia:
the computing of the backpropagation to update the model parameters is performed using an adaptive learning rate method (Adam): This limitation is a mathematical concept because it recites an algorithmic procedure for updating model parameters using gradients and other mathematical operations.
Step 2A Prong Two and Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B. Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2. The claim is ineligible.
Even when considered in combination, these additional elements represent mere instructions to apply an exception and therefore does not provide an inventive concept. The claim is ineligible.
Claim 9
Step 1: A process, as above.
Step2A Prong 1: This claim does not recite an additional abstract idea, but the claim depends on
claim 1, which recites an abstract idea.
Step2A Prong 2: This judicial exception is not integrated into a practical application because the
additional elements are as follows:
the fine-tuning operation of the pre-trained model occurs in a federated learning setting, and the computer-implemented method further comprising: The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)).
providing an output of the fine-tuned model to an aggregation server: Insignificant extra-solution as the limitation amounts to necessary data outputting (MPEP 2106.05(g)(3)).
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly
more than the judicial exception because the additional elements are as follows:
the fine-tuning operation of the pre-trained model occurs in a federated learning setting, and the computer-implemented method further comprising: The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself which cannot provide inventive concept (MPEP 2106.05(h)).
providing an output of the fine-tuned model to an aggregation server: Insignificant extra-solution as the limitation amounts to necessary data outputting (MPEP 2106.05(g)(3)). This falls under Well-Understood, Routine, Conventional activity -see MPEP 2106.05(d)(II)(vi).
Even when considered in combination, these additional elements represent mere instructions to apply an exception and therefore does not provide an inventive concept. The claim is ineligible.
Claim 10
Step 1: A process, as above.
Step2A Prong 1: The claim recites, inter alia:
Step2A Prong 2: This judicial exception is not integrated into a practical application because the
additional elements are as follows:
varying a pre-defined value of the pre-defined hyperparameter X to change a privacy-robustness values of the pre-trained model: This limitation is seen as a mathematical concept because it recites adjusting a numeric parameter that influences the outcome of a mathematical training algorithm.
Step 2A Prong Two and Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B. Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2. The claim is ineligible.
Even when considered in combination, these additional elements represent mere instructions to apply an exception and therefore does not provide an inventive concept. The claim is ineligible.
Claim 11
Step 1: The claim recites a computer device; therefore, it is directed to the statutory category of a
machine.
Step2A Prong 1: The claim recites, inter alia:
and prevent data leakage from membership inference attacks…: This limitation is seen as a mental process because it recites an intended result or objective of protecting data privacy.
computing, based on the querying of the pre-trained model with the private data, an initial loss distribution LINIT of a plurality of loss values: This limitation is a mathematical concept because it deals with performing mathematical calculations to derive LINIT. See Paragraph 65 which states “The modeling of the LINIT is performed by computing the mean and variance of the logit-scaled loss values. The mean and the variance aid in providing a Gaussian variable. The logit scaling also aids in obtaining a Gaussian random variable of the minibatch losses to make an MIA attack more difficult for an adversary.”.
computing a batch loss of a minibatch from the private data after the beginning of the fine-tuning operation…: This is a mathematical concept because the batch loss is computed by applying a loss function over the minibatch which applies math equations in order to compute the loss.
computing a loss distribution LBATCH of the batch loss: This limitation is a mathematical concept because it deals with calculating the mean and variance between the batches. See Paragraph 9 that states, “…the modeling of the LBATCH as a second Gaussian Random Variable is performed by computing the mean and variance of the logit-scaled batch loss.”.
computing a divergence metric between the initial loss distribution LINIT and the loss distribution LBATCH: This limitation is viewed as a mathematical concept because it deals with applying Kullback-Leibler (KL) divergence method to measure the distance between two probability distributions dealing with math calculations. See Paragraph 10 which states, “In one or more embodiments, the computing of the divergence metric between LNIT and LBATCH includes using a Kullback-Leibler divergence to measure a distance between LINIT and LBATCH. Kullback-Leibler (KL) divergence is a method useful to measure a distance between two probability distributions.”.
multiplying an output of the divergence metric with the pre-defined hyperparameter a to obtain a result: This limitation is viewed as a mathematical concept since it deals with multiplying two different numbers to get a result.
adding the result to the batch loss as a regularizer to penalize a distance between the initial loss distribution LINIT and the loss distribution LBATCH: This limitation is a mathematical concept dealing with the addition of the result from the previous step and the batch loss to obtain a regularizer.
…computing backpropagation on the batch loss to which the result is added as the regularizer: This limitation is a mathematical concept because it involves computing derivatives and loss functions during the backpropagation process.
Step2A Prong 2: This judicial exception is not integrated into a practical application because the additional elements are as follows:
A computing device configured to train a machine learning model… the computing device comprising: a processor; and a memory coupled to the processor, the memory storing instructions to cause the processor to perform operations comprising… updating model parameters by…: Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
receiving a pre-trained model and a pre-defined hyperparameter X as an input: Mere data gathering recited at a high level of generality, and thus are insignificant extra-solution activity (MPEP 2106.05(g)).
applying a forward pass by querying the pre-trained model with private data: The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)).
wherein the initial loss distribution LINIT is computed prior to beginning of a fine-tuning operation: The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)).
wherein the fine-tuning operation is performed to transform the pre-trained model into a fine-tuned model: Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
and outputting the fine-tuned model based on the updating: Insignificant extra-solution as the limitation amounts to necessary data outputting (MPEP 2106.05(g)(3)).
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because the additional elements are as follows:
A computing device configured to train a machine learning model… the computing device comprising: a processor; and a memory coupled to the processor, the memory storing instructions to cause the processor to perform operations comprising… updating model parameters by…: Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and cannot provide inventive concept (MPEP 2106.05(f)).
receiving a pre-trained model and a pre-defined hyperparameter X as an input: The additional element of “receiving” does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of receiving steps amounts to no more than mere data gathering. This element amounts to receiving data over a network and are well-understood, routine, conventional activity. See MPEP 2106.05(d), subsection II (i). This cannot provide an inventive concept.
applying a forward pass by querying the pre-trained model with private data: The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself which cannot provide inventive concept (MPEP 2106.05(h)).
wherein the initial loss distribution LINIT is computed prior to beginning of a fine-tuning operation: The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself which cannot provide inventive concept (MPEP 2106.05(h)).
wherein the fine-tuning operation is performed to transform the pre-trained model into a fine-tuned model: Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and cannot provide inventive concept (MPEP 2106.05(f)).
and outputting the fine-tuned model based on the updating: Insignificant extra-solution as the limitation amounts to necessary data outputting (MPEP 2106.05(g)(3)). This falls under Well-Understood, Routine, Conventional activity -see MPEP 2106.05(d)(II)(vi).
The elements in combination as an ordered whole still do not amount to significantly more than the judicial exception (i.e., the abstract ideas of mental processes and mathematical concepts for preventing data leakage from membership inference attacks). The claim merely describes applying known mathematical operations to machine learning training data, including computing loss distributions, calculating statistical parameters such as mean and variance, measuring divergence between probability distributions, and performing backpropagation to update model parameters. The recitation of training and fine-tuning a machine learning model, applying a forward pass, receiving a pre-trained model and hyperparameters, and outputting a fine-tuned model merely describes generic computer implementation and routine data gathering and output, without improving the functioning of a computer or machine learning model itself.
Claim 12 is a machine claim that recites similar limitations to claim 2. Therefore, claim 12 is rejected using the same rationale as claim 2.
Claim 13
Step 1: A machine, as above.
Step2A Prong 1: The claim recites, inter alia:
performing a modeling of the initial loss distribution LINIT as a first Gaussian Random Variable by computing a mean and a variance of the plurality of loss values, wherein the modeling is performed based on the applied logit-scaling to the plurality of loss values: This limitation is seen as a mathematical concept because it recites the application of statistical formulas to numeric data including calculating the mean and variance. The limitation additionally recites applying logit-scaling to the plurality of loss values, which involves performing a math operation on numerical values.
performing a modeling of the loss distribution LBATCH as a second Gaussian Random Variable by computing a mean and a variance of the batch loss, wherein the modeling is performed based on the applied logit-scaling to the batch loss values: This limitation is seen as a mathematical concept because it recites the application of statistical formulas to numeric data including calculating the mean and variance. The limitation additionally recites applying logit-scaling to the plurality of loss values, which involves performing a math operation on numerical values.
Step 2A Prong Two and Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B. Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2. The claim is ineligible.
Even when considered in combination, these additional elements represent mere instructions to apply an exception and therefore do not provide an inventive concept. The claim is ineligible.
Claim 14 is a machine claim that recites similar limitations to claim 5. Therefore, claim 14 is rejected using the same rationale as claim 5.
Claim 15 is a machine claim that recites similar limitations to claim 6. Therefore, claim 15 is rejected using the same rationale as claim 6.
Claim 16
Step 1: A machine, as above.
Step2A Prong 1: The claim recites, inter alia:
the update rule for the computing of the backpropagation to update the model parameters is performed by a Stochastic Gradient Descent (SGD) or an adaptive learning rate method (Adam): This limitation is seen as a mathematical concept because it recites an algorithmic procedure for adjusting numeric model parameters based on computed gradients.
Step 2A Prong Two and Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B. Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2. The claim is ineligible.
Even when considered in combination, these additional elements represent mere instructions to apply an exception and therefore do not provide an inventive concept. The claim is ineligible.
Claim 17
Step 1: A machine, as above.
Step2A Prong 1: This claim does not recite an additional abstract idea, but the claim depends on
claim 16, which recites an abstract idea.
Step2A Prong 2: This judicial exception is not integrated into a practical application because the
additional elements are as follows:
performing the fine-tuning operation of the pre-trained model in a federated learning setting: The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)).
and providing an output of the fine-tuned model to an aggregation server that shares the update of the fine-tuned model: Insignificant extra-solution as the limitation amounts to necessary data outputting (MPEP 2106.05(g)(3)).
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because the additional elements are as follows:
performing the fine-tuning operation of the pre-trained model in a federated learning setting: The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself which cannot provide inventive concept (MPEP 2106.05(h)).
and providing an output of the fine-tuned model to an aggregation server that shares the update of the fine-tuned model: Insignificant extra-solution as the limitation amounts to necessary data outputting (MPEP 2106.05(g)(3)). This falls under Well-Understood, Routine, Conventional activity -see MPEP 2106.05(d)(II)(vi).
Even when considered in combination, these additional elements represent mere instructions to apply an exception and therefore do not provide an inventive concept. The claim is ineligible.
Claim 18 is a machine claim that recites similar limitations to claim 10. Therefore, claim 18 is rejected using the same rationale as claim 10.
Claim 19
Step 1. The claim is assumed to be directed to a statutory category of an article of manufacture for the purposes of the abstract idea analysis.
Step2A Prong 1: The claim recites, inter alia:
…to compute an initial loss distribution LINIT of a plurality of the loss values: This limitation is a mathematical concept because it deals with performing mathematical calculations to derive LINIT. See Paragraph 65 which states “The modeling of the LINIT is performed by computing the mean and variance of the logit-scaled loss values. The mean and the variance aid in providing a Gaussian variable. The logit scaling also aids in obtaining a Gaussian random variable of the minibatch losses to make an MIA attack more difficult for an adversary.”.
…to compute a batch loss of a minibatch from the private data after the beginning of the a fine-tuning operation…: This is a mathematical concept because the batch loss is computed by applying a loss function over the minibatch which applies math equations in order to compute the loss.
…to compute a loss distribution LBATCH of the batch loss: This limitation is a mathematical concept because it deals with calculating the mean and variance between the batches. See Paragraph 9 that states, “…the modeling of the LBATCH as a second Gaussian Random Variable is performed by computing the mean and variance of the logit-scaled batch loss.”.
…to compute a divergence metric between the initial loss distribution LINIT and the loss distribution LBATCH: This limitation is viewed as a mathematical concept because it deals with applying Kullback-Leibler (KL) divergence method to measure the distance between two probability distributions dealing with math calculations. See Paragraph 10 which states, “In one or more embodiments, the computing of the divergence metric between LNIT and LBATCH includes using a Kullback-Leibler divergence to measure a distance between LINIT and LBATCH. Kullback-Leibler (KL) divergence is a method useful to measure a distance between two probability distributions.”.
multiply an output of the divergence metric with the pre-defined hyperparameter X to obtain a result: This limitation is viewed as a mathematical concept since it deals with multiplying two different numbers to get a result.
to add the result to the batch loss as a regularizer to penalize a distance between the initial loss distribution LINIT and the loss distribution LBATCH: This limitation is a mathematical concept dealing with the addition of the result from the previous step and the batch loss to obtain the regularizer.
…by computing backpropagation on the batch loss to which the result is added as the regularizer: This limitation is a mathematical concept because it involves computing derivatives and loss functions during the backpropagation process.
Step2A Prong 2: This judicial exception is not integrated into a practical application because the additional elements are as follows:
A computer program product comprising: one or more computer-readable storage devices and program instructions stored on at least one of the one or more computer-readable storage devices, the program instructions executable by a processor, the program instructions comprising: program instructions… to update model parameters of the pre-trained model: Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
…to receive a pre-trained model and a pre-defined hyperparameter X as an input: Mere data gathering recited at a high level of generality, and thus is an insignificant extra-solution activity (MPEP 2106.05(g)).
…to apply a forward pass by querying the pre-trained model with a private data: The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)).
wherein the initial loss distribution LINIT is computed prior to beginning of a fine-tuning operation: The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)).
wherein the fine-tuning operation is performed to transform the pre-trained model into a fine-tuned model: Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
…to output the fine-tuned model based on the updating: Insignificant extra-solution as the limitation amounts to necessary data outputting (MPEP 2106.05(g)(3)).
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because the additional elements are as follows:
A computer program product comprising: one or more computer-readable storage devices and program instructions stored on at least one of the one or more computer-readable storage devices, the program instructions executable by a processor, the program instructions comprising: program instructions… to update model parameters of the pre-trained model: Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and cannot provide inventive concept (MPEP 2106.05(f)).
…to receive a pre-trained model and a pre-defined hyperparameter X as an input: The additional element of “receiving” does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of receiving steps amounts to no more than mere data gathering. This element amounts to receiving data over a network and are well-understood, routine, conventional activity. See MPEP 2106.05(d), subsection II (i). This cannot provide an inventive concept.
…to apply a forward pass by querying the pre-trained model with private data: The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself which cannot provide inventive concept (MPEP 2106.05(h)).
wherein the initial loss distribution LINIT is computed prior to beginning of a fine-tuning operation: The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself which cannot provide inventive concept (MPEP 2106.05(h)).
wherein the fine-tuning operation is performed to transform the pre-trained model into a fine-tuned model: Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and cannot provide inventive concept (MPEP 2106.05(f)).
…to output the fine-tuned model based on the updating: Insignificant extra-solution as the limitation amounts to necessary data outputting (MPEP 2106.05(g)(3)). This falls under Well-Understood, Routine, Conventional activity -see MPEP 2106.05(d)(II)(vi).
The elements in combination as an ordered whole still do not amount to significantly more than the judicial exception (i.e., the abstract ideas of mental processes and mathematical concepts for preventing data leakage from membership inference attacks). The claim merely describes applying known mathematical operations to machine learning training data, including computing loss distributions, calculating statistical parameters such as mean and variance, measuring divergence between probability distributions, and performing backpropagation to update model parameters. The recitation of training and fine-tuning a machine learning model, applying a forward pass, receiving a pre-trained model and hyperparameters, and outputting a fine-tuned model merely describes generic computer implementation and routine data gathering and output, without improving the functioning of a computer or machine learning model itself.
Claim 20
Step 1: An article of manufacture, as above.
Step2A Prong 1: The claim recites, inter alia:
apply logit-scaling to the plurality of loss values obtained from the forward pass: This limitation is seen as a mathematical concept because logit-scaling is a math operation applied to numerical values.
perform a modeling of the initial loss distribution LINIT as a first Gaussian Random Variable by computing a mean and a variance of the plurality of loss values, wherein the modeling is performed based on the applied logit- scaling to the plurality of loss values: This limitation is seen as a mathematical concept because it recites the application of statistical formulas to numeric data including calculating the mean and variance. The limitation additionally recites applying logit-scaling to the plurality of loss values, which involves performing a math operation on numerical values.
apply logit-scaling to the batch loss: This limitation is seen as a mathematical concept because it recites performing logit-scaling, which involves math operations.
perform a modeling of the loss distribution LBATCH as a second Gaussian Random Variable by computing a mean and variance of the batch loss, wherein the modeling is performed based on the applied logit-scaling to the batch loss: This limitation is seen as a mathematical concept because it recites the application of statistical formulas to numeric data including calculating the mean and variance. The limitation additionally recites applying logit-scaling to the plurality of loss values, which involves performing a math operation on numerical values.
compute the backpropagation to update the model parameters by using an update rule: This limitation recites a mathematical concept because it involves computing the backpropagation, which involves using math calculations to update the model parameters.
and determine, prior to outputting the fine-tuned model, that the fine-tuned model meets a termination criterion: This limitation is seen as a mental process because it involves the determination of that a model meets a certain criterion.
Step 2A Prong Two and Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B. Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2. The claim is ineligible.
Even when considered in combination, these additional elements represent mere instructions to apply an exception and therefore do not provide an inventive concept. The claim is ineligible.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1-5, 7-8, 10-14, 16, and 18-19 are rejected under 35 U.S.C. 103 as being unpatentable over Carlini (“Membership Inference Attacks From First Principles”, 2022) in view of Chen (“RELAXLOSS: DEFENDING MEMBERSHIP INFERENCE ATTACKS WITHOUT LOSING UTILITY”, 2022), in view of Zhang (US 20240064160 A1), and in further view of Leidner (US 20210043211 A1).
Regarding claim 1,
Carlini teaches [a] computer-implemented method of training a machine learning model to
prevent data leakage from membership inference attacks, the computer-implemented method comprising (Section 2 Page 2 of Carlini, “The field of training data privacy constructs attacks that leak data, develops techniques to prevent memorization, and measures the privacy of proposed defenses.”, Section 3 Page 2, “The objective of a membership inference attack (MIA) [60] is to predict if a specific training example was, or was not, used as training data in a particular model. This makes MIAs the simplest and most widely deployed attack for auditing training data privacy… This section formalizes the membership inference attack security game (§III-A), and introduces our membership inference evaluation methodology”):
computing… an initial loss distribution LINIT of a plurality of loss values, wherein the initial loss distribution LINIT is computed… (Section IV Page 4, “We formalize this by considering two distributions over models… is the distribution of models trained on datasets containing (x,y), and then Qout(x,y)={f←T(D∖{(x,y)})∣D←D}… To simplify the situation, we instead define Qin and Qout as the distributions of losses on (x,y) for models either trained, or not trained, on this example.”
Carlini establishes Q_out as the baseline distribution of losses for data points that the model has
not yet been trained on which corresponds to the initial loss distribution LINIT.);
computing a loss distribution LBATCH of the … loss (Section IV Page 4, “We formalize this by considering two distributions over models: Qin(x,y)={f←T(D∪{(x,y)})∣D←D} is the distribution of models trained on datasets containing (x,y)… To simplify the situation, we instead define Qin and Qout as the distributions of losses on (x,y) for models either trained, or not trained, on this example.”
Carlini defines Q_in as a distribution of losses for models trained on datasets containing specific
training example (x,y), thereby characterizing how the inclusion of data changes the trained
model’s loss behavior.);
computing a divergence metric between the initial loss distribution LINIT and the loss distribution LBATCH (Section IV Page 4, “…the best hypothesis test… is obtained by thresholding the Likelihood-ratio Test between the two hypotheses:
PNG
media_image1.png
79
625
media_image1.png
Greyscale
”
The Likelihood-ratio test maps onto the divergence metric because both are mathematical methods used to quantify the “distance” or statistical difference between the two distributions Q_in (LBatch) and Q_out (Linit). The divergence metric determines how much the model’s behavior has shifted toward the trained state. If the result of the likelihood ratio is a high value, then it means that the model’s behavior shifted away from the un-trained baseline where it now belongs to the trained distribution.);
Carlini does not teach receiving a pre-trained model and a pre-defined hyperparameter X as an input for machine learning; applying a forward pass by querying the pre-trained model with a private data; and multiplying an output of the divergence metric with a pre-defined hyperparameter a to obtain a result; adding the result to the batch loss as a regularizer; updating model parameters by computing backpropagation on the regularized batch loss; and outputting the fine-tuned model.
Chen, in the same field of endeavor, teaches receiving a pre-trained model and a pre-defined hyperparameter λ as an input for machine learning (Section 7 Page 9 of Chen, “Our method involves a single hyperparameter α that controls the trade-off between privacy and utility. A fine-grained grid search on a validation set (i.e., first estimating the privacy-utility trade-off with varying value of α, and subsequently selecting the α corresponding to the desired privacy/utility level) allows precise control over the expected privacy/utility level of the target model.”, Adaptive Attack Section C.4 Page 20, “And for the NN-based attack, we use the complete logits prediction from the pre-trained shadow models as features to train the adaptive attack models (modeled as a NN).”);
applying a forward pass by querying the pre-trained model with private data (Page 3 Preliminaries Section, “We consider the standard setting of MIA: the attacker has access to a query set S = {(zi , mi)} N i=1 containing both member (training) and non-member (testing) samples drawn from the same data distribution Pdata, where mi is the membership attribute (mi = 1 if zi is a member). The task is to infer the value of the membership attribute mi associated with each query sample zi… which predicts mi for a given query sample zi and a target model parametrized by θ.”, Page 17 Model Architectures, “…we adopt the same architecture as used in Nasr et al. (2018) 12: a 4-layer fully-connected neural network”, See Algorithm 1, “for epoch in {1, ..., E} do for batch_index in {1, ..., K} do Get sample batch {(xi , yi)} B i=1 Perform forward pass: pi = f(xi ; θ)”, Section IV Page 4, “We formalize this by considering two distributions over models: Qin(x,y)={f←T(D∪{(x,y)})∣D←D} is the distribution of models trained on datasets containing (x,y)… To simplify the situation, we instead define Qin and Qout as the distributions of losses on (x,y) for models either trained, or not trained, on this example.”
The forward pass in Algorithm 1 simulates the attacker’s query by processing private training data to generate output predictions. The algorithm then applies backpropagation to fine tune the model’s weights. The membership attribute m_i identifies whether a queried sample z_i is private training data, and Algorithm 1 explicitly performs a forward pass on such samples.);
…the … data after the beginning of the fine-tuning operation, wherein the fine-tuning operation is performed to transform the pre-trained model into a fine-tuned model (Page 18 Memguard, “Memguard modifies the output predictions of pre-trained target models during test-time, i.e., output predictions are perturbed by adversarial noise to fool a surrogate attack model… Each surrogate attack model is trained on the target model’s predictions when inputting the target model’s training data (used as member samples) and a separate hold-out set data (used as non-member samples).”
Chen performs iterative optimization on a model by repeatedly processing batches through forward passes and updating model parameters during training. These operations occur after training has begun and progressively transform the model into an updated model, corresponding to processing the claimed data after the beginning of the fine-tuning operation.)
multiplying an output of the divergence metric with the pre-defined hyperparameter λ to obtain a result (Defense Methods Section B.5 Page 17, “Label-smoothing prevents overconfident predictions by incorporating a regularization term into the training objective that penalizes the distance (measured by the KL-divergence) between the model predictions and the uniform distribution. The objective is formularized as follows
PNG
media_image2.png
44
516
media_image2.png
Greyscale
”
The equation defines a regularized loss by scaling the KL-divergence (divergence metric term) by the hyperparameter a to penalize the overconfident predictions.);
adding the result to the… loss as a regularizer to penalize a distance between the [first] distribution … and the [second value] … (Defense Methods Section B.5 Page 17, “Label-smoothing prevents overconfident predictions by incorporating a regularization term into the training objective that penalizes the distance (measured by the KL-divergence) between the model predictions and the uniform distribution. The objective is formularized as follows
PNG
media_image2.png
44
516
media_image2.png
Greyscale
”
This mathematical result is added to the standard cross-entropy loss to force the model toward a uniform distribution.);
updating model parameters of the pre-trained model by computing backpropagation on the … loss to which the result is added as the regularizer (Pages 17-18 Label-smoothing, “The objective is formularized as follows
PNG
media_image2.png
44
516
media_image2.png
Greyscale
…DKL is the KL-divergence; U denotes the uniform distribution; pθ(y|x) denotes the output prediction. α is a hyper-parameter with range [0,1] that balances the cross-entropy loss LCE and the regularization term.”, See Algorithm 1 on Page 4,
PNG
media_image3.png
761
375
media_image3.png
Greyscale
The algorithm uses a forward pass to query the model with private data, then calculates a regularized loss where hyperparameter a (from Equation 14) scales the divergence metric that penalizes overconfident predictions. Backpropagation is performed on the CE loss to update the weights that fine-tunes the model into a privacy-preserving state that balances classification utility.);
and outputting the fine-tuned model based on the updating (See Algorithm 1 on Page 4, “
Output: Model f(·;θ) with parameters θ… Initialize model parameter θ; for epoch in {1,...,E} do… Update model parameters: θ ←θ−τ∇L(θ)… return model f(·;θ)”
Chen’s algorithm updates the parameters of the model and then outputs/returns the fine-tuned model.).
based on the querying of the pre-trained model with the private data (Page 2 Introduction, “(iii) Extensive evaluations on five datasets with diverse modalities demonstrate that our method outperforms state-of-the-art approaches…”, Page 4 Algorithm 1, “Input: Dataset {(xi,yi)}N i=1… Get sample batch {(xi,yi)}B i=1 Perform forward pass: pi = f(xi;θ)”, Page 16 Section B.1, “CH-MNIST… contains 5000 greyscale images of 8 different types of tissues from patients with colorectal cancer… Texas100 contains medical records of 67,330 patients published by the Texas Department of State Health Services 8.”, Page 18 Section B.5, “Memguard modifies the output predictions of pre-trained target models during test-time, i.e., output predictions are perturbed by adversarial noise to fool a surrogate attack model… Each surrogate attack model is trained on the target model’s predictions when inputting the target model’s training data (used as member samples) and a separate hold-out set data (used as non-member samples).”)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to combine Carlini’s teaching that membership inference vulnerability is measured by the divergence between loss distributions of member and non-member data with Chen’s teaching of a training algorithm that applies a hyperparameter-scaled divergence penalty as a regularizer during backpropagation in order to defend against attacks and minimize the measurable privacy risk (Introduction of Chen).
Carlini in view of Chen does not teach computing a batch loss of a minibatch from the private data after the beginning of the fine-tuning operation, wherein the fine-tuning operation is performed to transform the pre-trained model into a fine-tuned model.
Zhang, in the same field of endeavor, teaches computing a batch loss of a minibatch from the … data… (Paragraph 111 of Zhang, “Under mini-batch SGD, the mini-batch loss of the l.sup.th selected vehicle can be calculated by averaging losses across all selected message graphs:”
Zhang discloses computing a mini-batch loss by averaging losses of multiple training samples under mini-batch SGD, which corresponds to computing a batch loss of a minibatch.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to combine Carlini in view of Chen’s teaching with Zhang’s teaching of computing a mini-batch loss in order to fine tune a model in a privacy-preserving manner while leveraging mini-batch SGD techniques used in deep learning (Paragraph 97 of Zhang).
Carlini in view of Chen in further view of Zhang does not teach wherein the [operation] is computed prior to beginning of a fine-tuning operation.
Leidner, in the same field of endeavor, teaches wherein the [operation] is computed prior to beginning of a fine-tuning operation (Claim 11, “The computer system of claim 8, wherein the processor is further programmed to: prior to fine-tuning, further pre-train the general language model based on a pre-training transcript corpus to learn a domain-specific language model that is transcript-specific.”
Leidner teaches that the processor computes the process of pre-training the language model based on a pre-trained transcript corpus before fine-tuning.)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to combine Carlini in view of Chen in further view of Zhang’s LINIT distribution computation with Leidner’s computing of an operation prior to fine-tuning in order to fine tune the model that is specific to the target data (Paragraph 7 of Leidner).
Regarding claim 2,
Carlini teaches further comprising applying logit-scaling to the plurality of loss values (Page 6 Under Algorithm 1,“We first train N shadow models [60] on random samples from the data distribution D, so that half of these models are trained on the target point (x, y), and half are not (we call these respectively IN and OUT models for (x, y)).”, Page 10 of Logit scaling the loss function, “The first step of our attack projects the model’s confidences to a logit scale to ensure that the distributions that we work with are approximately normal… we find that using the model’s confidence f(x)y∈[0,1], or its logarithm (the cross-entropy loss), leads to poor performance of the attack since these statistics do not behave like Gaussians (recall from Figure 4).Our logit rescaling performs best, but the exact numerical computation of the logit function ϕ(p)=log(p1−p) matters. ”
Logit-scaling is performed on the plurality of model confidences to convert them into Gaussian distributions.).
Regarding claim 3,
Carlini teaches performing a modeling of the initial loss distribution LINIT as a first Gaussian Random Variable by computing a mean and a variance of the plurality of loss values, wherein the modeling is performed based on the applied logit-scaling to the plurality of loss values (Page 4 Under Equation 3, “To minimize the number of shadow models necessary, we assume Q˜ in/out is a Gaussian distribution, reducing our attack to estimating just four parameters: the mean and variance of each distribution.”, Page 6 Under Algorithm 1, “In this case, we fit m dimensional spherical Gaussians
PNG
media_image4.png
34
250
media_image4.png
Greyscale
to the losses collected from querying the shadow models m times per example, and compute a standard likelihood-ratio test between two multivariate normal distributions.”
Carlini teaches this claim by instructing the user to collect a plurality of scores from shadow models, transform them to a logit scale, and calculate the mean and variance of the Q_out (Linit) to fit the Gaussian model.).
Regarding claim 4,
Carlini teaches applying logit-scaling to the batch loss (Page 6 Under Algorithm 1, “We then fit two Gaussians to the confidences of the IN and OUT models on (x, y) (in logit scale).”, Page 10 Logit scaling the loss function section, “The first step of our attack projects the model’s confidences to a logit scale to ensure that the distributions that we work with are approximately normal.”),
and performing a modeling of the loss distribution LBATCH as a second Gaussian Random Variable by computing a mean and a variance of the logit-scaled batch loss wherein the modeling is performed based on the applied logit-scaling to the batch loss (Page 4 Under Equation 3, “we assume Q~in/out is a Gaussian distribution, reducing our attack to estimating just four parameters: the mean and variance of each distribution.”, Algorithm 1 (lines 10-15) explicitly calculate mean and variance to define the Gaussian parameters, “µin ← mean(confsin)… σ 2 in ← var(confsin)…
PNG
media_image5.png
40
250
media_image5.png
Greyscale
”
Carlini’s method instructs the user to model the (Q_in / L_Batch) distribution as a Gaussian by calculating the mean (µin) and variance (σ^2in) from the logit-scaled confidence scores. Carlini teaches using these Gaussian parameters to perform a statistical test to determine membership).
Regarding claim 5,
Carlini teaches the computing of the divergence metric between the initial loss distribution LINIT and the loss distribution LBATCH… to measure the distance between the initial loss distribution LINIT and the loss distribution LBATCH (See Equation 2 on Page 4 of Carlini, “
PNG
media_image6.png
43
263
media_image6.png
Greyscale
”
The Likelihood-ratio test maps onto the divergence metric because both are mathematical methods used to quantify the “distance” or statistical difference between the two distributions Q_in (LBatch) and Q_out (Linit). The divergence metric determines how much the model’s behavior has shifted toward the trained state.)
Carlini does not teach comprises using a Kullback-Leibler divergence.
Chen, in the same field of endeavor, teaches the computing of the divergence metric between the [first] loss distribution LINIT and the [second] loss distribution comprises using a Kullback-Leibler divergence… (Page 17 Under Section B.5, “Label-smoothing prevents overconfident predictions by incorporating a regularization term into the training objective that penalizes the distance (measured by the KL-divergence
PNG
media_image7.png
38
517
media_image7.png
Greyscale
”)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to combine Carlini’s teaching of measuring membership inference vulnerability by computing a divergence metric between loss distributions with Chen’s teaching of using Kullback-Leibler (KL) divergence metric within a training regularization objective in order to implement a known statistical divergence method between loss distributions to defend against privacy attacks (Page 17 Section B.5).
Regarding claim 7,
Carlini teaches the computing of the backpropagation to update the model parameters is performed using a Stochastic Gradient Descent (SGD) (Page 2 Section A Machine learning notation, “Neural networks are trained via stochastic gradient descent [32] to minimize some loss function :
PNG
media_image8.png
43
280
media_image8.png
Greyscale
”, Page 12 Mismatched training procedures section, “In Figure 11b we fix the architecture to a WRN28-10, and vary the training optimizer: SGD, SGDM (SGD with momentum) or Adam.”
Equation 1 explicitly represents the parameter update step of the backpropagation algorithm, where gradients are used to minimize the loss via Stochastic Gradient Descent.).
Regarding claim 8,
Carlini teaches the computing of the backpropagation to update the model parameters is performed using an adaptive learning rate method (Adam) (Page 2 Section A Machine learning notation, “to minimize some loss function:
PNG
media_image8.png
43
280
media_image8.png
Greyscale
”, Page 12 Mismatched training procedures section, “In Figure 11b we fix the architecture to a WRN28-10, and vary the training optimizer: SGD, SGDM (SGD with momentum) or Adam.”
Carlini explicitly identifies Adam as a training optimizer used to minimize the loss function by dynamically adjusting learning rates, which serves as a functional implementation of the parameter update phase in the backpropagation process.).
Regarding claim 10,
Carlini does not teach varying a pre-defined value of the hyperparameter X to change a privacy-robustness values of the training model.
Chen, in the same field of endeavor, teaches varying a pre-defined value of the pre-defined hyperparameter X to change a privacy-robustness values of the pre-trained model (Page 9 Under Practicality, “Our method involves a single hyperparameter α that controls the trade-off between privacy and utility. A fine-grained grid search on a validation set (i.e., first estimating the privacy-utility trade-off with varying value of α, and subsequently selecting the α corresponding to the desired privacy/utility level) allows precise control over the expected privacy/utility level of the target model.”, Page 18 Under Distillation, “To determine the hyper-parameter that best describes the privacy-utility trade-off, we conduct preliminary experiments and investigate the effect of α and T independently.”
Chen discloses a predefined hyperparameter that is intentionally varied to control the privacy-utility (privacy-robustness) trade-off of the trained model.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to combine Carlini’s teaching of a framework for measuring and defending against membership inference attacks with Chen’s teaching of using a hyperparameter tuned to control the trade-off between the privacy robustness and utility in order to enable control over the final model’s privacy robustness level for optimization (Page 9, Practicality of Chen).
Regarding claim 11,
Carlini teaches …train a machine learning model and prevent data leakage from membership inference attacks (Section 1 Introduction Page 1, “Neural networks are now trained on increasingly sensitive datasets, and so it is necessary to ensure that trained models are privacy-preserving. In order to empirically verify if a model is in fact private, membership inference attacks [60] have become the de facto standard [42], [63] because of their simplicity. A membership inference attack receives as input a trained model and an example from the data distribution, and predicts if that example was used to train the model.”, Section IV Page 4, “We formalize this by considering two distributions over models: Qin(x,y)={f←T(D∪{(x,y)})∣D←D} is the distribution of models trained on datasets containing (x,y)… To simplify the situation, we instead define Qin and Qout as the distributions of losses on (x,y) for models either trained, or not trained, on this example.”
Carlini’s definition of Q_in as the distribution for models trained on datasets containing (x,y) maps to the “private data” limitation because it describes how a model’s behavior is fundamentally altered by the inclusion of specific, sensitive training samples. Since a fine-tuned model’s loss distribution (L_batch) is derived from its interaction with this private data, it functions the same as Q_in.);
computing… an initial loss distribution LINIT of a plurality of loss values, wherein the initial loss distribution LINIT is computed… (Section IV Page 4, “We formalize this by considering two distributions over models… is the distribution of models trained on datasets containing (x,y), and then Qout(x,y)={f←T(D∖{(x,y)})∣D←D}… To simplify the situation, we instead define Qin and Qout as the distributions of losses on (x,y) for models either trained, or not trained, on this example.”
Carlini establishes Q_out as the baseline distribution of losses for data points that the model has not yet been trained on which corresponds to the initial loss distribution LINIT.);
and computing a loss distribution LBATCH of the … loss (Section 1 Introduction Page 1, “Neural networks are now trained on increasingly sensitive datasets, and so it is necessary to ensure that trained models are privacy-preserving. In order to empirically verify if a model is in fact private, membership inference attacks [60] have become the de facto standard [42], [63] because of their simplicity. A membership inference attack receives as input a trained model and an example from the data distribution, and predicts if that example was used to train the model.”, Section IV Page 4, “We formalize this by considering two distributions over models: Qin(x,y)={f←T(D∪{(x,y)})∣D←D} is the distribution of models trained on datasets containing (x,y)… To simplify the situation, we instead define Qin and Qout as the distributions of losses on (x,y) for models either trained, or not trained, on this example.”
Carlini’s definition of Q_in as the distribution for models trained on datasets containing (x,y) maps to the “private data” limitation because it describes how a model’s behavior is fundamentally altered by the inclusion of specific, sensitive training samples. Since a fine-tuned model’s loss distribution (L_batch) is derived from its interaction with this private data, it functions the same as Q_in.);
computing a divergence metric between the initial loss distribution LINIT and the loss distribution LBATCH (Section IV Page 4, “…the best hypothesis test… is obtained by thresholding the Likelihood-ratio Test between the two hypotheses:
PNG
media_image1.png
79
625
media_image1.png
Greyscale
”
The Likelihood-ratio test maps onto the divergence metric because both are mathematical methods used to quantify the “distance” or statistical difference between the two distributions Q_in (LBatch) and Q_out (Linit). The divergence metric determines how much the model’s behavior has shifted toward the trained state. If the result of the likelihood ratio is a high value, then it means that the model’s behavior shifted away from the un-trained baseline where it now belongs to the trained distribution.);
Carlini does not teach receiving a pre-trained model and a pre-defined hyperparameter X as an input for machine learning; applying a forward pass by querying the pre-trained model with a private data; and multiplying an output of the divergence metric with a pre-defined hyperparameter a to obtain a result; adding the result to the batch loss as a regularizer; updating a model parameters by computing backpropagation on the regularized batch loss; and outputting the fine-tuned model.
Chen, in the same field of endeavor, teaches receiving a pre-trained model and a pre-defined hyperparameter X as an input for machine learning (Section 7 Page 9, “Our method involves a single hyperparameter α that controls the trade-off between privacy and utility. A fine-grained grid search on a validation set (i.e., first estimating the privacy-utility trade-off with varying value of α, and subsequently selecting the α corresponding to the desired privacy/utility level) allows precise control over the expected privacy/utility level of the target model.”, Adaptive Attack Section C.4 Page 20, “And for the NN-based attack, we use the complete logits prediction from the pre-trained shadow models as features to train the adaptive attack models (modeled as a NN).”);
applying a forward pass by querying the pre-trained model with private data (Page 3 Preliminaries Section, “We consider the standard setting of MIA: the attacker has access to a query set S = {(zi , mi)} N i=1 containing both member (training) and non-member (testing) samples drawn from the same data distribution Pdata, where mi is the membership attribute (mi = 1 if zi is a member). The task is to infer the value of the membership attribute mi associated with each query sample zi… which predicts mi for a given query sample zi and a target model parametrized by θ.”, Page 17 Model Architectures, “…we adopt the same architecture as used in Nasr et al. (2018) 12: a 4-layer fully-connected neural network”, See Algorithm 1, “for epoch in {1, ..., E} do for batch_index in {1, ..., K} do Get sample batch {(xi , yi)} B i=1 Perform forward pass: pi = f(xi ; θ), Section IV Page 4, “We formalize this by considering two distributions over models: Qin(x,y)={f←T(D∪{(x,y)})∣D←D} is the distribution of models trained on datasets containing (x,y)… To simplify the situation, we instead define Qin and Qout as the distributions of losses on (x,y) for models either trained, or not trained, on this example.”
The forward pass in Algorithm 1 simulates the attacker’s query by processing private training data to generate output predictions. The algorithm then applies backpropagation to fine tune the model’s weights. The membership attribute m_i identifies whether a queried sample z_i is private training data, and Algorithm 1 explicitly performs a forward pass on such samples.
…the private data after the beginning of the fine-tuning operation, wherein the fine-tuning operation is performed to transform the pre-trained model into a fine-tuned model (Page 18 Memguard, “Memguard modifies the output predictions of pre-trained target models during test-time, i.e., output predictions are perturbed by adversarial noise to fool a surrogate attack model… Each surrogate attack model is trained on the target model’s predictions when inputting the target model’s training data (used as member samples) and a separate hold-out set data (used as non-member samples).”
Chen performs iterative optimization on a model by repeatedly processing batches through forward passes and updating model parameters during training. These operations occur after training has begun and progressively transform the model into an updated model, corresponding to processing the claimed data after the beginning of the fine-tuning operation.),
multiplying an output of the divergence metric with the pre-defined hyperparameter λ to obtain a result (Defense Methods Section B.5 Page 17, “Label-smoothing prevents overconfident predictions by incorporating a regularization term into the training objective that penalizes the distance (measured by the KL-divergence) between the model predictions and the uniform distribution. The objective is formularized as follows
PNG
media_image2.png
44
516
media_image2.png
Greyscale
”
The equation defines a regularized loss by scaling the KL-divergence (divergence metric term) by the hyperparameter a to penalize the overconfident predictions.);
adding the result to the… loss as a regularizer to penalize a distance between the [first] loss distribution … and the [second] loss distribution… (Defense Methods Section B.5 Page 17, “Label-smoothing prevents overconfident predictions by incorporating a regularization term into the training objective that penalizes the distance (measured by the KL-divergence) between the model predictions and the uniform distribution. The objective is formularized as follows
PNG
media_image2.png
44
516
media_image2.png
Greyscale
”
This mathematical result is added to the standard cross-entropy loss to force the model toward a uniform distribution.);
updating model parameters of the pre-trained model by computing backpropagation of the … loss to which the result is added as the regularizer (Pages 17-18 Label-smoothing, “The objective is formularized as follows
PNG
media_image2.png
44
516
media_image2.png
Greyscale
…DKL is the KL-divergence; U denotes the uniform distribution; pθ(y|x) denotes the output prediction. α is a hyper-parameter with range [0,1] that balances the cross-entropy loss LCE and the regularization term.”, See Algorithm 1 on Page 4,
PNG
media_image3.png
761
375
media_image3.png
Greyscale
The algorithm uses a forward pass to query the model with private data, then calculates a regularized loss where hyperparameter a (from Equation 14) scales the divergence metric that penalizes overconfident predictions. Backpropagation is performed on the CE loss to update the weights that fine-tunes the model into a privacy-preserving state that balances classification utility.);
and outputting the fine-tuned model based on the updating (See Algorithm 1 on Page 4, “
Output: Model f(·;θ) with parameters θ… Initialize model parameter θ; for epoch in {1,...,E} do… Update model parameters: θ ←θ−τ∇L(θ)… return model f(·;θ)”
Chen’s algorithm updates the parameters of the model and then outputs/returns the fine-tuned model.).
based on the querying of the pre-trained model with the private data (Page 2 Introduction, “(iii) Extensive evaluations on five datasets with diverse modalities demonstrate that our method outperforms state-of-the-art approaches…”, Page 4 Algorithm 1, “Input: Dataset {(xi,yi)}N i=1… Get sample batch {(xi,yi)}B i=1 Perform forward pass: pi = f(xi;θ)”, Page 16 Section B.1, “CH-MNIST… contains 5000 greyscale images of 8 different types of tissues from patients with colorectal cancer… Texas100 contains medical records of 67,330 patients published by the Texas Department of State Health Services 8.”, Page 18 Section B.5, “Memguard modifies the output predictions of pre-trained target models during test-time, i.e., output predictions are perturbed by adversarial noise to fool a surrogate attack model… Each surrogate attack model is trained on the target model’s predictions when inputting the target model’s training data (used as member samples) and a separate hold-out set data (used as non-member samples).”)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to combine Carlini’s teaching that membership inference vulnerability is measured by the divergence between loss distributions of member and non-member data with Chen’s teaching of a training algorithm that applies a hyperparameter-scaled divergence penalty as a regularizer during backpropagation in order to defend against attacks and minimize the measurable privacy risk (Introduction of Chen).
Carlini in view of Chen do not teach [a] computing device… the computing device comprising: a processor; and a memory coupled to the processor, the memory storing instructions to cause the processor to perform acts and computing a batch loss of a minibatch.
Zhang, in the same field of endeavor, teaches [a] computing device… the computing device comprising: a processor; and a memory coupled to the processor, the memory storing instructions to cause the processor to perform operations comprising (Paragraph 157, “Various embodiments described herein are implemented using dedicated hardware, configurable hardware or programmed processors executing programming instructions that are broadly described in flow chart form that can be stored on any suitable electronic storage medium or transmitted over any suitable electronic communication medium.”)
computing a batch loss of a minibatch from the … data… (Paragraph 111 of Zhang, “Under mini-batch SGD, the mini-batch loss of the l.sup.th selected vehicle can be calculated by averaging losses across all selected message graphs:”
Zhang discloses computing a mini-batch loss by averaging losses of multiple training samples under mini-batch SGD, which corresponds to computing a batch loss of a minibatch.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to combine Carlini and Chen’s teaching with Zhang’s teaching of computing a mini-batch loss in order to fine tune a model in a privacy-preserving manner while leveraging mini-batch SGD techniques used in deep learning (Paragraph 97 of Zhang).
Carlini in view of Chen in further view of Zhang does not teach wherein the [operation] is computed prior to beginning of a fine-tuning operation.
Leidner, in the same field of endeavor, teaches wherein the [operation] is computed prior to beginning of a fine-tuning operation (Claim 11, “The computer system of claim 8, wherein the processor is further programmed to: prior to fine-tuning, further pre-train the general language model based on a pre-training transcript corpus to learn a domain-specific language model that is transcript-specific.”
Leidner teaches that the processor computes the process of pre-training the language model based on a pre-trained transcript corpus before fine-tuning.)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to combine Carlini in view of Chen in further view of Zhang’s LINIT distribution computation with Leidner’s computing of an operation prior to fine-tuning in order to fine tune the model that is specific to the target data (Paragraph 7 of Leidner).
Claim 12 is a machine claim that recites similar limitations to claim 2. Therefore, claim 12 is rejected using the same rationale as claim 2.
Regarding claim 13,
Carlini teaches performing a modeling of the initial loss distribution LINIT as a Gaussian Random Variable by computing a mean and a variance of plurality of loss values, wherein the modeling is performed based on the applied logit-scaling to the plurality of loss values (Page 4 Under Equation 3, “To minimize the number of shadow models necessary, we assume Q˜ in/out is a Gaussian distribution, reducing our attack to estimating just four parameters: the mean and variance of each distribution.”, Page 6 Under Algorithm 1, “In this case, we fit m dimensional spherical Gaussians
PNG
media_image4.png
34
250
media_image4.png
Greyscale
to the losses collected from querying the shadow models m times per example, and compute a standard likelihood-ratio test between two multivariate normal distributions.”
Carlini teaches this claim by instructing the user to collect a plurality of scores from shadow models, transform them to a logit scale, and calculate the mean and variance of the Q_out (Linit) to fit the Gaussian model.);
performing a modeling of the LINIT as a first Gaussian Random Variable by computing a mean and a variance of the logit-scaled loss values Page 4 Under Equation 3, “To minimize the number of shadow models necessary, we assume Q˜ in/out is a Gaussian distribution, reducing our attack to estimating just four parameters: the mean and variance of each distribution.”, Page 6 Under Algorithm 1, “In this case, we fit m dimensional spherical Gaussians
PNG
media_image4.png
34
250
media_image4.png
Greyscale
to the losses collected from querying the shadow models m times per example, and compute a standard likelihood-ratio test between two multivariate normal distributions.”
Carlini teaches this claim by instructing the user to collect a plurality of scores from shadow models, transform them to a logit scale, and calculate the mean and variance of the Q_out (Linit) to fit the Gaussian model.);
and performing a modeling of the LBATCH as a second Gaussian Random Variable by computing a mean and a variance of the logit-scaled batch loss (Page 4 Under Equation 3, “we assume Q~in/out is a Gaussian distribution, reducing our attack to estimating just four parameters: the mean and variance of each distribution.”, Algorithm 1 (lines 10-13) explicitly calculate mean and variance to define the Gaussian parameters.).
Carlini does not teach the instructions cause the processor to perform additional acts.
Chen, in the same field of endeavor, teaches the instructions further cause the processor to perform the operations comprising (Paragraph 157 of Chen, “Various embodiments described herein are implemented using dedicated hardware, configurable hardware or programmed processors executing programming instructions that are broadly described in flow chart form that can be stored on any suitable electronic storage medium or transmitted over any suitable electronic communication medium.”).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to combine Carlini’s teaching of training a model to prevent data leakage using a divergence-based regularizer with Chen’s teaching of a generic computing device in order to implement a system configured to perform Carlini’s method to prevent data leakage from membership inference attacks (Paragraph 157 of Chen).
Claim 14 is a machine claim that recites similar limitations to claim 5. Therefore, claim 14 is rejected using the same rationale as claim 5.
Regarding claim 16,
Carlini teaches the update rule for the computing of the backpropagation to update the model parameters is performed by a Stochastic Gradient Descent (SGD) or an adaptive learning rate method (Adam) (Page 2 Section A Machine learning notation, “Neural networks are trained via stochastic gradient descent [32] to minimize some loss function :
PNG
media_image8.png
43
280
media_image8.png
Greyscale
”, Page 12 Mismatched training procedures section, “In Figure 11b we fix the architecture to a WRN28-10, and vary the training optimizer: SGD, SGDM (SGD with momentum) or Adam.”
Equation 1 explicitly represents the parameter update step of the backpropagation algorithm, where gradients are used to minimize the loss via Stochastic Gradient Descent.).
Claim 18 is a machine claim that recites similar limitations to claim 10. Therefore, claim 18 is rejected using the same rationale as claim 10.
Regarding claim 19,
Carlini teaches compute… an initial loss distribution LINIT of a plurality of the loss values, wherein the initial loss distribution LINIT is computed… (Section IV Page 4, “We formalize this by considering two distributions over models… is the distribution of models trained on datasets containing (x,y), and then Qout(x,y)={f←T(D∖{(x,y)})∣D←D}… To simplify the situation, we instead define Qin and Qout as the distributions of losses on (x,y) for models either trained, or not trained, on this example.”
Carlini establishes Q_out as the baseline distribution of losses for data points that the model has not yet been trained on which corresponds to the initial loss distribution LINIT.);
and to compute a loss distribution LBATCH of the … loss (Section IV Page 4, “We formalize this by considering two distributions over models: Qin(x,y)={f←T(D∪{(x,y)})∣D←D} is the distribution of models trained on datasets containing (x,y)… To simplify the situation, we instead define Qin and Qout as the distributions of losses on (x,y) for models either trained, or not trained, on this example.”
Carlini defines Q_in as a distribution of losses for models trained on datasets containing specific
training example (x,y), thereby characterizing how the inclusion of data changes the trained
model’s loss behavior.);
compute a divergence metric between the initial loss distribution LINIT and the loss distribution LBATCH… (Section IV Page 4, “…the best hypothesis test… is obtained by thresholding the Likelihood-ratio Test between the two hypotheses:
PNG
media_image1.png
79
625
media_image1.png
Greyscale
”
The Likelihood-ratio test maps onto the divergence metric because both are mathematical methods used to quantify the “distance” or statistical difference between the two distributions Q_in (LBatch) and Q_out (Linit). The divergence metric determines how much the model’s behavior has shifted toward the trained state. If the result of the likelihood ratio is a high value, then it means that the model’s behavior shifted away from the un-trained baseline where it now belongs to the trained distribution.)
Carlini does not teach …receive a pre-trained model and a pre-defined hyperparameter X as an input; …apply a forward pass by querying the pre-trained model with a private data……and multiply an output of the divergence metric with a pre-defined hyperparameter A. to obtain a result, and to add the result to the batch loss as a regularizer; …to update a model parameters by computing backpropagation on the regularized batch loss; and… to output the fine-tuned model.
Chen, in the same field of endeavor, teaches receive a pre-trained model and a pre-defined hyperparameter λ as an input (Section 7 Page 9, “Our method involves a single hyperparameter α that controls the trade-off between privacy and utility. A fine-grained grid search on a validation set (i.e., first estimating the privacy-utility trade-off with varying value of α, and subsequently selecting the α corresponding to the desired privacy/utility level) allows precise control over the expected privacy/utility level of the target model.”, Adaptive Attack Section C.4 Page 20, “And for the NN-based attack, we use the complete logits prediction from the pre-trained shadow models as features to train the adaptive attack models (modeled as a NN).”);
apply a forward pass by querying the pre-trained model with private data (Page 3 Preliminaries Section, “We consider the standard setting of MIA: the attacker has access to a query set S = {(zi , mi)} N i=1 containing both member (training) and non-member (testing) samples drawn from the same data distribution Pdata, where mi is the membership attribute (mi = 1 if zi is a member). The task is to infer the value of the membership attribute mi associated with each query sample zi… which predicts mi for a given query sample zi and a target model parametrized by θ.”, Page 17 Model Architectures, “…we adopt the same architecture as used in Nasr et al. (2018) 12: a 4-layer fully-connected neural network”, See Algorithm 1, “for epoch in {1, ..., E} do for batch_index in {1, ..., K} do Get sample batch {(xi , yi)} B i=1 Perform forward pass: pi = f(xi ; θ)”, Section IV Page 4, “We formalize this by considering two distributions over models: Qin(x,y)={f←T(D∪{(x,y)})∣D←D} is the distribution of models trained on datasets containing (x,y)… To simplify the situation, we instead define Qin and Qout as the distributions of losses on (x,y) for models either trained, or not trained, on this example.”
The forward pass in Algorithm 1 simulates the attacker’s query by processing private training data to generate output predictions. The algorithm then applies backpropagation to fine tune the model’s weights. The membership attribute m_i identifies whether a queried sample z_i is private training data, and Algorithm 1 explicitly performs a forward pass on such samples.);
…the private data after the beginning of the fine-tuning operation, wherein the fine-tuning operation is performed to transform the pre-trained model into a fine-tuned model (Page 18 Memguard, “Memguard modifies the output predictions of pre-trained target models during test-time, i.e., output predictions are perturbed by adversarial noise to fool a surrogate attack model… Each surrogate attack model is trained on the target model’s predictions when inputting the target model’s training data (used as member samples) and a separate hold-out set data (used as non-member samples).”
Chen performs iterative optimization on a model by repeatedly processing batches through forward passes and updating model parameters during training. These operations occur after training has begun and progressively transform the model into an updated model, corresponding to processing the claimed data after the beginning of the fine-tuning operation.)
multiply an output of the divergence metric with the pre-defined hyperparameter λ to obtain a result (Defense Methods Section B.5 Page 17, “Label-smoothing prevents overconfident predictions by incorporating a regularization term into the training objective that penalizes the distance (measured by the KL-divergence) between the model predictions and the uniform distribution. The objective is formularized as follows
PNG
media_image2.png
44
516
media_image2.png
Greyscale
”
The equation defines a regularized loss by scaling the KL-divergence (divergence metric term) by the hyperparameter a to penalize the overconfident predictions.),
add the result to the … loss as a regularizer to penalize a distance between the [first] loss distribution … and the [second] loss distribution … (Defense Methods Section B.5 Page 17, “Label-smoothing prevents overconfident predictions by incorporating a regularization term into the training objective that penalizes the distance (measured by the KL-divergence) between the model predictions and the uniform distribution. The objective is formularized as follows
PNG
media_image2.png
44
516
media_image2.png
Greyscale
”
The equation defines a regularized loss by scaling the KL-divergence (divergence metric) by the hyperparameter a to penalize the overconfident predictions. This mathematical result is added to the standard cross-entropy loss to force the model toward a uniform distribution.);
update model parameters of the pre-trained model by computing backpropagation on the … loss to which the result is added as the regularizer (Pages 17-18 Label-smoothing, “The objective is formularized as follows
PNG
media_image2.png
44
516
media_image2.png
Greyscale
…DKL is the KL-divergence; U denotes the uniform distribution; pθ(y|x) denotes the output prediction. α is a hyper-parameter with range [0,1] that balances the cross-entropy loss LCE and the regularization term.”, See Algorithm 1 on Page 4,
PNG
media_image3.png
761
375
media_image3.png
Greyscale
The algorithm uses a forward pass to query the model with private data, then calculates a regularized loss where hyperparameter a (from Equation 14) scales the divergence metric that penalizes overconfident predictions. Backpropagation is performed on the CE loss to update the weights that fine-tunes the model into a privacy-preserving state that balances classification utility.);
and to output the fine-tuned model based on the updating (See Algorithm 1 on Page 4, “
Output: Model f(·;θ) with parameters θ… Initialize model parameter θ; for epoch in {1,...,E} do… Update model parameters: θ ←θ−τ∇L(θ)… return model f(·;θ)”
Chen’s algorithm updates the parameters of the model and then outputs/returns the fine-tuned model.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to combine Carlini’s teaching that membership inference vulnerability is measured by the divergence between loss distributions of member and non-member data with Chen’s teaching of a training algorithm that applies a hyperparameter-scaled divergence penalty as a regularizer during backpropagation in order to defend against attacks and minimize the measurable privacy risk (Introduction of Chen).
Carlini in view of Chen does not teach a computer program product comprising: one or more computer-readable storage devices and program instructions stored on at least one of the one or more computer-readable storage devices, the program instructions executable by a processor, the program instructions comprising.
Zhang, in the same field of endeavor, teaches a computer program product comprising: one or more computer-readable storage devices and program instructions stored on at least one of the one or more computer-readable storage devices, the program instructions executable by a processor, the program instructions cause the processor to (Paragraph 157, “Various embodiments described herein are implemented using dedicated hardware, configurable hardware or programmed processors executing programming instructions that are broadly described in flow chart form that can be stored on any suitable electronic storage medium or transmitted over any suitable electronic communication medium.”)
computing a batch loss of a minibatch from the … data… (Paragraph 111 of Zhang, “Under mini-batch SGD, the mini-batch loss of the l.sup.th selected vehicle can be calculated by averaging losses across all selected message graphs:”
Zhang discloses computing a mini-batch loss by averaging losses of multiple training samples under mini-batch SGD, which corresponds to computing a batch loss of a minibatch.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to combine Carlini in view of Chen’s teaching with Zhang’s teaching of computing a mini-batch loss in order to fine tune a model in a privacy-preserving manner while leveraging mini-batch SGD techniques used in deep learning (Paragraph 97 of Zhang).
Carlini in view of Chen in further view of Zhang does not teach wherein the [operation] is computed prior to beginning of a fine-tuning operation.
Leidner, in the same field of endeavor, teaches the [operation] is computed prior to beginning of a fine-tuning operation (Claim 11, “The computer system of claim 8, wherein the processor is further programmed to: prior to fine-tuning, further pre-train the general language model based on a pre-training transcript corpus to learn a domain-specific language model that is transcript-specific.”
Leidner teaches that the processor computes the process of pre-training the language model based on a pre-trained transcript corpus before fine-tuning.)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to combine Carlini in view of Chen in further view of Zhang’s LINIT distribution computation with Leidner’s computing of an operation prior to fine-tuning in order to fine tune the model that is specific to the target data (Paragraph 7 of Leidner).
Claims 6, 15, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Carlini (“Membership Inference Attacks From First Principles”, 2022) in view of Chen (“RELAXLOSS: DEFENDING MEMBERSHIP INFERENCE ATTACKS WITHOUT LOSING UTILITY”, 2022), in view of Zhang (US 20240064160 A1), in view of Leidner (US 20210043211 A1), and in view of Sadeghi (US 12537178 B2).
Regarding claim 6,
Carlini teaches the updating of the model parameters is performed by using an update rule, and the computer-implemented method further comprising (“Page 2 Section A Machine learning notation, “Neural networks are trained via stochastic gradient descent [32] to minimize some loss function :
PNG
media_image8.png
43
280
media_image8.png
Greyscale
”):
Carlini does not teach determining, prior to outputting the fine-tuned , that the fine-tuned model meets a termination criterion.
Chen, in the same field of endeavor, teaches outputting the fine-tuned model (See Algorithm 1 on Page 4, “
PNG
media_image3.png
761
375
media_image3.png
Greyscale
”
The algorithm uses a forward pass to query the model with private data, then calculates a regularized loss where hyperparameter a (from Equation 14) scales the divergence metric that penalizes overconfident predictions. Performing backpropagation on this result to update the weights fine-tunes the model into a privacy-preserving state that balances classification utility.).
Carlini in view of Chen does not teach determining, prior to outputting the… model, that the … model meets a termination criterion.
Sadeghi, in the same field of endeavor, teaches determining, prior to outputting the… model, that the … model meets a termination criterion (Paragraph 223, “Control returns to the 2952 if the model does not meet the predetermined training criteria. Control ends if the model meets the predetermined training criteria.”).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to combine Carlini in view of Chen’s teaching of updating model parameters via stochastic gradient descent (SGD) update rule during training with Sadeghi’s teaching of applying a predetermined termination criterion to control the training of the model in order to provide an automated mechanism that would terminate training based on a threshold (Paragraph 223 of Sadeghi).
Claim 15 is a machine claim that recites similar limitations to claim 6. Therefore, claim 15 is rejected using the same rationale as claim 6.
Regarding claim 20,
Carlini teaches apply logit-scaling to the plurality of loss values obtained from the forward pass (Page 6 Under Algorithm 1,“We first train N shadow models [60] on random samples from the data distribution D, so that half of these models are trained on the target point (x, y), and half are not (we call these respectively IN and OUT models for (x, y)).”, Page 10 of Logit scaling the loss function, “The first step of our attack projects the model’s confidences to a logit scale to ensure that the distributions that we work with are approximately normal… we find that using the model’s confidence f(x)y∈[0,1], or its logarithm (the cross-entropy loss), leads to poor performance of the attack since these statistics do not behave like Gaussians (recall from Figure 4).Our logit rescaling performs best, but the exact numerical computation of the logit function ϕ(p)=log(p1−p) matters. ”);
perform a modeling of the initial loss distribution LINIT as a first Gaussian Random Variable by computing a mean and a variance of the plurality of loss values, wherein the modeling is performed based on the applied logit-scaling to the plurality of loss values (Page 4 Under Equation 3, “To minimize the number of shadow models necessary, we assume Q˜ in/out is a Gaussian distribution, reducing our attack to estimating just four parameters: the mean and variance of each distribution.”, Page 6 Under Algorithm 1, “In this case, we fit m dimensional spherical Gaussians
PNG
media_image4.png
34
250
media_image4.png
Greyscale
to the losses collected from querying the shadow models m times per example, and compute a standard likelihood-ratio test between two multivariate normal distributions.”
Carlini teaches this claim by instructing the user to collect a plurality of scores from shadow models, transform them to a logit scale, and calculate the mean and variance of the Q_out (Linit) to fit the Gaussian model.);
apply logit-scaling to the batch loss (Page 6 Under Algorithm 1, “We then fit two Gaussians to the confidences of the IN and OUT models on (x, y) (in logit scale).”, Page 10 Logit scaling the loss function section, “The first step of our attack projects the model’s confidences to a logit scale to ensure that the distributions that we work with are approximately normal.”),
and perform a modeling of the loss distribution LBATCH as a second Gaussian Random Variable by computing a mean and variance of the batch loss, wherein the modeling is performed based on the applied logit-scaling to the batch loss (Page 4 Under Equation 3, “we assume Q~in/out is a Gaussian distribution, reducing our attack to estimating just four parameters: the mean and variance of each distribution.”, Algorithm 1 (lines 10-13) explicitly calculate mean and variance to define the Gaussian parameters.);
compute the backpropagation to update the model parameters by using an update rule (“Page 2 Section A Machine learning notation, “Neural networks are trained via stochastic gradient descent [32] to minimize some loss function :
PNG
media_image8.png
43
280
media_image8.png
Greyscale
”):
Carlini does not teach the program instructions further cause the processor to… outputting the fine-tuned model.
Chen, in the same field of endeavor, teaches outputting the fine-tuned model (See Algorithm 1 on Page 4, “
PNG
media_image3.png
761
375
media_image3.png
Greyscale
”
Performing backpropagation on the private data to update the weights fine-tunes the model into a privacy-preserving state that balances classification utility.)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to combine Carlini’s teaching that membership inference vulnerability is measured by the divergence between loss distributions of member and non-member data with Chen’s teaching of a training algorithm that applies a hyperparameter-scaled divergence penalty as a regularizer during backpropagation in order to defend against attacks and minimize the measurable privacy risk (Introduction of Chen).
Carlini in view of Chen does not teach the program instructions further cause the processor to…
Zhang, in the same field of endeavor, teaches the program instructions further cause the processor to… (Paragraph 157, “Various embodiments described herein are implemented using dedicated hardware, configurable hardware or programmed processors executing programming instructions that are broadly described in flow chart form that can be stored on any suitable electronic storage medium or transmitted over any suitable electronic communication medium.”)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to combine Carlini and Chen’s teaching with Zhang’s teaching of computing a mini-batch loss in order to fine tune a model in a privacy-preserving manner while leveraging mini-batch SGD techniques used in deep learning (Paragraph 97 of Zhang).
Carlini in view of Chen in view of Zhang in further view of Leidner does not teach …prior to output of the fine-tuned model to determine that the fine-tuned model meets a termination criterion and program instructions.
Sadeghi teaches determine, prior to outputting the … model, that the … model meets a termination criterion (Paragraph 223, “Control returns to the 2952 if the model does not meet the predetermined training criteria. Control ends if the model meets the predetermined training criteria.”).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to combine Carlini’s teaching of updating model parameters via stochastic gradient descent (SGD) update rule during training with Sadeghi’s teaching of applying a predetermined termination criterion to control the training of the model in order to provide an automated mechanism that would terminate training based on a threshold (Paragraph 223 of Sadeghi).
Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Carlini (“Membership Inference Attacks From First Principles”, 2022) in view of Chen (“RELAXLOSS: DEFENDING MEMBERSHIP INFERENCE ATTACKS WITHOUT LOSING UTILITY”, 2022), in view of Zhang (US 20240064160 A1), in view of Leidner (US 20210043211 A1), and in further view of Nasr (“Comprehensive Privacy Analysis of Deep Learning: Passive and Active White-box Inference Attacks against Centralized and Federated Learning”, 2019).
Regarding claim 9,
Carlini does not teach the fine-tuning operation of the pre-trained model occurs in a federated learning setting, and the method further comprising: providing the output of the fine-tuned model to an aggregation server.
Nasr, in the same field of endeavor, teaches the fine-tuning operation of the pre-trained model occurs in a federated learning setting (Page 5 Stand-alone fine-tunning section, “At a later stage it is updated to f△ after being fine-tuned using a new dataset D△…. The model for inference attacks against fine-tunned models is a special case of our membership inference model for at-tacking federated learning.”),
and the computer-implemented method further comprising: providing an output of the fine-tuned model to an aggregation server (Page 5 Federated Learning Section, “A central server keeps the latest version of the parameters W for the global model… In each epoch of training, each participant downloads the global parameters, updates them locally using SGD algorithm on their local training data, and uploads them back to the server.”
Nasr teaches fine-tuning a pre-trained model in a federated learning setting where participants locally update the model and provide the resulting fine-tuned model outputs (parameter updates) to a central aggregation server for global model aggregation.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to combine Carlini’s teaching of a method for fine-tuning a model to prevent membership inference attacks by measuring and minimizing loss distribution divergence with Nasr’s teaching of fine-tuning operations within an federated learning environment in order to apply the privacy-preserving fine-tuning method within a federated learning framework to protect sensitive data across distributed participants (Page 5, Federated Learning section of Nasr)
Claims 17 is rejected under 35 U.S.C. 103 as being unpatentable over Carlini (“Membership Inference Attacks From First Principles”, 2022) in view of Chen (“RELAXLOSS: DEFENDING MEMBERSHIP INFERENCE ATTACKS WITHOUT LOSING UTILITY”, 2022), in view of Zhang (US 20240064160 A1), in view of Leidner (US 20210043211 A1), in view of Sadeghi (US 12537178 B2), and in further view of Nasr (“Comprehensive Privacy Analysis of Deep Learning: Passive and Active White-box Inference Attacks against Centralized and Federated Learning”, 2019).
Regarding claim 17,
Carlini in view of Chen does not teach the instructions cause the processor to perform additional acts comprising: training the fine-tuned model in a federated learning setting, and providing the output of the fine-tuned model to an aggregation server that shares an update of fine-tuned model.
Zhang, in the same field of endeavor, teaches the instructions further cause the processor to perform the operations comprising (Paragraph 157, “Various embodiments described herein are implemented using dedicated hardware, configurable hardware or programmed processors executing programming instructions that are broadly described in flow chart form that can be stored on any suitable electronic storage medium or transmitted over any suitable electronic communication medium.”)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to combine Carlini and Chen’s teaching with Zhang’s processor computing a mini-batch loss in order to fine tune a model in a privacy-preserving manner while leveraging mini-batch SGD techniques used in deep learning (Paragraph 97 of Zhang).
Carlini in view of Chen, in view of Zhang, in view of Leidner, and in further view of Sadeghi does not teach teaches performing the fine-tuning operation of the pre-trained model in a federated learning setting, and providing an output of the fine-tuned model to an aggregation server that shares the update of the fine-tuned model.
Nasr, in the same field of endeavor, teaches performing the fine-tuning operation of the pre-trained model in a federated learning setting (Page 5 Stand-alone fine-tunning section, “At a later stage it is updated to f△ after being fine-tuned using a new dataset D△…. The model for inference attacks against fine-tunned models is a special case of our membership inference model for attacking federated learning.”);
and providing an output of the fine-tuned model to an aggregation server that shares the update of the fine-tuned model (Page 5 Federated Learning Section, “A central server keeps the latest version of the parameters W for the global model… In each epoch of training, each participant downloads the global parameters, updates them locally using SGD algorithm on their local training data, and uploads them back to the server.”, Page 5 Federated Learning, “In this setting, N participants… collaborate to train a global model”
Nasr teaches fine-tuning a pre-trained model in a federated learning setting where participants locally update the model and provide the resulting fine-tuned model outputs (parameter updates) to a central aggregation server for global model aggregation.)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to combine Carlini’s method of fine-tuning a model to prevent data leakage with Nasr’s teaching of performing fine-tuning in a federated learning setting in order to extend the privacy-preserving training method to a distributed, federated environment (Page 5, Federated Learning Section of Nasr.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MAJD MAHER HADDAD whose telephone number is (571)272-2265. The examiner can normally be reached Mon-Friday 8-5 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kamran Afshar, can be reached at (571) 272-7796. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/M.M.H./Examiner, Art Unit 2125
/KAMRAN AFSHAR/Supervisory Patent Examiner, Art Unit 2125