Prosecution Insights
Last updated: October 02, 2026
Application No. 18/344,419

UN-LEARNING OF TRAINING DATA FOR MACHINE LEARNING MODELS

Final Rejection §103§112
Filed
Jun 29, 2023
Priority
Dec 15, 2022 — provisional 63/387,590
Examiner
NYE, LOUIS CHRISTOPHER
Art Unit
2141
Tech Center
2100 — Computer Architecture & Software
Assignee
Amazon Technologies Inc.
OA Round
2 (Final)
29%
Grant Probability
At Risk
3-4
OA Rounds
11m
Est. Remaining
59%
With Interview

Examiner Intelligence

Grants only 29% of cases
29%
Career Allowance Rate
4 granted / 14 resolved
-26.4% vs TC avg
Strong +30% interview lift
Without
With
+30.0%
Interview Lift
resolved cases with interview
Typical timeline
4y 2m
Avg Prosecution
23 currently pending
Career history
37
Total Applications
across all art units

Statute-Specific Performance

§101
26.2%
-13.8% vs TC avg
§103
58.7%
+18.7% vs TC avg
§102
8.4%
-31.6% vs TC avg
§112
6.7%
-33.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 14 resolved cases

Office Action

§103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 6-15 is/are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 6 recites the limitation "the set of adapter weights" in line 10. There is insufficient antecedent basis for this limitation in the claim. It is unclear what set of adapter weights the limitation refers to. Examiner suggests amending the limitation to “a set of adapter weights”. Claim 14 recites “a set of adapter weights” in line 2; it is unclear if this recitation of adapter weights is the same as the set of adapter weights recited in claim 6, or if this recitation of adapter weights is intended to distinguish from the adapter weights of claim 6. Claim 15 recites “a respective subset of the adapter weights” in line 2; it is unclear if this recitation of the adapter weights refers to the adapter weights of claim 14 or if this recitation of the adapter weights refers to the adapter weights of claim 6. Claims 7-15 are rejected under 35 U.S.C. 112(b) for reasons discussed above. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Bourtoule et al. (NPL from IDS: Machine Unlearning, published 2021, hereinafter “Bourtoule”) in view of Houlsby et al. (NPL from IDS: Parameter-Efficient Transfer Learning for NLP, published June 2019, hereinafter “Houlsby”). Regarding claim 1, Bourtoule teaches a computer-implemented method, comprising: receiving a request to remove a data sample from a dataset (Bourtoule, Pg. 2, Col. 1, Paragraph 1 – “Finally, when a request to unlearn a training point arrives, we need to retrain only the affected model.” – teaches receiving a request to remove a data samples from a dataset), the dataset comprising a plurality of shards each corresponding to a portion of the dataset, training data in the plurality of shards used to train a respective plurality of instances of a model (Bourtoule, Pg. 2, Col. 1, Paragraph 1 – “First, we divide the training data into multiple disjoint shards such that a training point is included in one shard only; shards partition the data. Then, we train models in isolation on each of these shards, which limits the influence of a point to the model that was trained on the shard containing the point.” – teaches the dataset comprising a plurality of shards each corresponding to a portion of the dataset (divides training data into multiple disjoint shards), training data in the plurality of shards used to train a respective plurality of instances of a model (train models in isolation on each of these shards)); identifying an instance of the model that is trained using a shard that contains the data sample (Bourtoule, Fig. 2 and Pg. 2, Col. 1, Paragraph 1 – “Then, we train models in isolation on each of these shards, which limits the influence of a point to the model that was trained on the shard containing the point. Finally, when a request to unlearn a training point arrives, we need to retrain only the affected model.” – teaches identifying an instance of the model that is trained using a shard that contains the data sample (trains models in isolation on each shard, limits influence of a point to the model that was trained on the shard containing the point, thus identifying an instance of the model that is trained using a shard that contains the sample)); identifying a slice of the shard that contains the data sample, the shard comprising a plurality of slices of the training data each corresponding to [[a]] one of a plurality of checkpoints set during training of the model (Bourtoule, Fig. 2 and Pg. 2, Col. 1, Paragraph 1 – “In addition, rather than training each model on the entire shard directly, we can divide each shard’s data into slices and present slices incrementally during training. We save the state of model parameters before introducing each new slice, allowing us to start retraining the model from the last known parameter state that does not include the point to be unlearned—rather than a random initialization.” – teaches identifying a slice of the shard that contains the data sample, the shard comprising a plurality of slices of the training data each corresponding to one of a plurality of checkpoints set during training (can divide each shard’s data into slices and present slices incrementally during training, can save state of model parameters before introducing each slice allowing to retrain from one of a plurality of checkpoints set during training)); removing the data sample from the identified slice of the dataset (Bourtoule, Fig. 1 and Pg. 2, Col. 1, Paragraph 1 – “We save the state of model parameters before introducing each new slice, allowing us to start retraining the model from the last known parameter state that does not include the point to be unlearned—rather than a random initialization.” – teaches removing the data sample from the identified slice of the dataset (retrains with last known parameter state that does not include the point to be unlearned, thus removing the data sample from the identified slice of the dataset, and further in Fig. 1, shows removing data sample from identified slice of the dataset)); retraining the identified instance of the model using a set of weights stored at a checkpoint that was most recently set before the identified instance was trained using the training data in the slice (Bourtoule, Pg. 2, Col. 1, Paragraph 1 – “We save the state of model parameters before introducing each new slice, allowing us to start retraining the model from the last known parameter state that does not include the point to be unlearned—rather than a random initialization.” – teaches retraining the identified instance of the model using a set of weights stored at a checkpoint (saved state of model parameters) that most recently set before the identified instance was trained using the training data in the slice (retraining from last known parameter state that does not include the point to be unlearned)); and providing the retrained instance with the other instances, of the plurality of instances of the model, to generate a plurality of inferences to be used to generate a consensus inference output (Bourtoule, Fig. 2 and Pg. 2, Col. 1, Paragraph 1 – “At inference, we use different strategies to aggregate the predictions of models trained on each shard: the simplest one is a majority vote over predicted labels.” – teaches instances of the plurality of instances (models trained on each shard) to generate a plurality of inferences (predictions) to be used to generate a consensus inference output (majority vote over predicted labels)). Bourtoule fails to explicitly teach a [[the]] language model; retraining the identified instance of the language model using a set of adapter weights. However, analogous to the field of the claimed invention, Houlsby teaches: a [[the]] language model (Houlsby, Fig. 2 and Pg. 2, Col. 2, Last Paragraph – We instantiate adapter-based tuning for text Transformers. These models attain state-of-the-art performance in many NLP tasks, including translation, extractive QA, and text classification problems” – teaches a language model); retraining the identified instance of the language model using a set of adapter weights (Houlsby, Pg. 2, Col. 2, Paragraph 3 – “Adapter modules perform more general architectural modifications to re-purpose a pre trained network for a downstream task. In particular, the adapter tuning strategy involves injecting new layers into the original network. The weights of the original network are untouched, whilst the new adapter layers are initialized at random. In standard fine-tuning, the new top-layer and the original weights are co-trained. In contrast, in adapter tuning, the parameters of the original network are frozen and therefore may be shared by many tasks.” and in Fig. 2 Description – “During adapter tuning, the green layers are trained on the downstream data, this includes the adapter, the layer normalization parameters, and the final classification layer” – teaches training the identified instance of the language model using a set of adapter weights (adapter modules initialize adapter layers with adapter weights, trains identified instance of language model using adapter weights). In addition to the previously cited passages, Houlsby further teaches in Pg. 2, Col. 2, Paragraph 4 – “Adapter modules have two main features: a small number of parameters, and a near-identity initialization…A near-identity initialization is required for stable training of the adapted model; we investigate this empirically in Section 3.6. By initializing the adapters to a near-identity function, original network is unaffected when training starts. During training, the adapters may then be activated to change the distribution of activations throughout the network.” – teaches wherein the adapter modules are initialized with a near-identity initialization) Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the language model and adapter weights of Houlsby to the shards, slices, checkpoints, retraining, and unlearning of Bourtoule in order to unlearn a data sample from a language model. Doing so would meet the right to be forgotten requirements, thus achieving erasure of data from ML models (Bourtoule, Introduction), would allow the model to be extended to new tasks without affecting previous ones, and would provide parameter-efficient tuning for NLP by reducing the number of parameters to be tuned while maintaining performance (Houlsby, Introduction). Regarding claim 6, Bourtoule teaches a computer-implemented method, comprising: receiving a request to remove a data sample from a dataset used to train a plurality of instances of a model (Bourtoule, Pg. 2, Col. 1, Paragraph 1 – “Finally, when a request to unlearn a training point arrives, we need to retrain only the affected model.” – teaches receiving a request to remove a data sample from a dataset used to train a plurality of instances of a model), wherein the plurality of instances are trained using respective portions of the dataset (Bourtoule, Pg. 2, Col. 1, Paragraph 1 – “First, we divide the training data into multiple disjoint shards such that a training point is included in one shard only; shards partition the data. Then, we train models in isolation on each of these shards, which limits the influence of a point to the model that was trained on the shard containing the point.” – teaches wherein the plurality of instances are trained using respective portions of the dataset (divides training data into multiple disjoint shards, trains models in isolation on each of these shards, thus the plurality of instances are trained using respective portions, or shards, of the dataset)); identifying an instance of the model that was trained using a portion of the dataset including the data sample, the portion of the data set corresponding to one of a plurality of checkpoints set during training of the machine learning model (Bourtoule, Fig. 2 and Pg. 2, Col. 1, Paragraph 1 – “Then, we train models in isolation on each of these shards, which limits the influence of a point to the model that was trained on the shard containing the point. Finally, when a request to unlearn a training point arrives, we need to retrain only the affected model… We save the state of model parameters before introducing each new slice, allowing us to start retraining the model from the last known parameter state that does not include the point to be unlearned—rather than a random initialization.” – teaches identifying an instance of the model that is trained using a shard that contains the data sample (trains models in isolation on each shard, limits influence of a point to the model that was trained on the shard containing the point, thus identifying an instance of the model that is trained using a shard that contains the sample), the portion of the data set corresponding to one of a plurality of checkpoints set during training of the machine learning model (saves state of model parameters before introducing each slice, or portion of the data set. Thus the portion, or slice, of the dataset corresponds to one of a plurality of checkpoints, or saved state of model parameters, set during training of the machine learning model)); removing the data sample from the dataset (Bourtoule, Fig. 1 and Pg. 2, Col. 1, Paragraph 1 – “We save the state of model parameters before introducing each new slice, allowing us to start retraining the model from the last known parameter state that does not include the point to be unlearned—rather than a random initialization.” – teaches removing the data sample from the dataset (retrains with last known parameter state that does not include the point to be unlearned, thus removing the data sample from the dataset, and further in Fig. 1, shows removing data sample from the dataset)); retraining the identified instance of the model using the portion of the dataset with the data sample removed, wherein the set of weights correspond to a checkpoint that was most recently set before the identified instance was trained using the data sample (Bourtoule, Pg. 2, Col. 1, Paragraph 1 – “We save the state of model parameters before introducing each new slice, allowing us to start retraining the model from the last known parameter state that does not include the point to be unlearned—rather than a random initialization.” – teaches retraining the identified instance of the model using the portion of the dataset with the data sample removed (retrains model from last known parameter state that does not include the point), wherein the set of weights correspond to a checkpoint (saved state of model parameters) that was most recently set before the identified instance was trained using the data sample (retrains from last known parameter state that does not include the point to be unlearned)); and providing the retrained instance for use in the plurality of instances to generate inferences to be aggregated into a single inference output (Bourtoule, Fig. 2 and Pg. 2, Col. 1, Paragraph 1 – “At inference, we use different strategies to aggregate the predictions of models trained on each shard: the simplest one is a majority vote over predicted labels.” – teaches providing the retrained instance for use in the plurality of instances (models trained on each shard) to generate a plurality of inferences (predictions) to be aggregated into a single inference output (majority vote over predicted labels to aggregate predictions of models, thus aggregating inferences into a single inference output)). Bourtoule fails to explicitly teach a [[the]] language model; and the [[a]] set of adapter weights. However, analogous to the field of the claimed invention, Houlsby teaches: a [[the]] language model (Houlsby, Fig. 2 and Pg. 2, Col. 2, Last Paragraph – We instantiate adapter-based tuning for text Transformers. These models attain state-of-the-art performance in many NLP tasks, including translation, extractive QA, and text classification problems” – teaches a language model); the [[a]] set of adapter weights (Houlsby, Pg. 2, Col. 2, Paragraph 3 – “Adapter modules perform more general architectural modifications to re-purpose a pre trained network for a downstream task. In particular, the adapter tuning strategy involves injecting new layers into the original network. The weights of the original network are untouched, whilst the new adapter layers are initialized at random. In standard fine-tuning, the new top-layer and the original weights are co-trained. In contrast, in adapter tuning, the parameters of the original network are frozen and therefore may be shared by many tasks.” and in Fig. 2 Description – “During adapter tuning, the green layers are trained on the downstream data, this includes the adapter, the layer normalization parameters, and the final classification layer” – teaches training the the language model using a set of adapter weights (adapter modules initialize adapter layers with adapter weights, trains identified instance of language model using adapter weights). In addition to the previously cited passages, Houlsby further teaches in Pg. 2, Col. 2, Paragraph 4 – “Adapter modules have two main features: a small number of parameters, and a near-identity initialization…A near-identity initialization is required for stable training of the adapted model; we investigate this empirically in Section 3.6. By initializing the adapters to a near-identity function, original network is unaffected when training starts. During training, the adapters may then be activated to change the distribution of activations throughout the network.” – teaches wherein the adapter modules are initialized with a near-identity initialization) Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the language model and adapter weights of Houlsby to the shards, slices, checkpoints, retraining, and unlearning of Bourtoule in order to unlearn a data sample from a language model. Doing so would meet the right to be forgotten requirements, thus achieving erasure of data from ML models (Bourtoule, Introduction), would allow the model to be extended to new tasks without affecting previous ones, and would provide parameter-efficient tuning for NLP by reducing the number of parameters to be tuned while maintaining performance (Houlsby, Introduction). Claim 16 incorporates substantively all the limitations of claim 6 in a system and is rejected on similar grounds as above. Bourtoule teaches the processor and memory of this claim at Pg. 10, Col. 1, Paragraph 2 – “We run our experiments using P100 and T4 Nvidia GPUs, with 12 and 16 GB of dedicated memory, respectively. We use Intel Xeon Silver 4110 CPUs with 8 cores each and 192GB of Ram”. Regarding claim 2, the combination of Bourtoule and Houlsby teaches the computer-implemented method of claim 1, wherein generating the updated inference output further comprising: determining the consensus inference output based on a majority vote based on the plurality of inferences (Bourtoule, Pg. 2, Col. 1, Paragraph 1 – “At inference, we use different strategies to aggregate the predictions of models trained on each shard: the simplest one is a majority vote over predicted labels.” – teaches determining the consensus inference output based on a majority vote based on the plurality of inferences (performs majority vote over predicted labels produced by the plurality of models trained on each shard)). Claims 8 and 18 are similar to claim 2, hence similarly rejected. Regarding claim 3, the combination of Bourtoule and Houlsby teaches the computer-implemented method of claim 1, wherein data samples are positioned in the slices of a shard of the dataset based at least in part on a likelihood that a request will be received to remove the data samples from the dataset (Bourtoule, Pg. 13, Col. 2, Paragraph 3 – “We discuss one such approach in Algorithm 1, under the following assumptions: (a) the distribution of unlearning requests is known precisely, and (b) this distribution is relatively constant over a time interval. Recall that each data point du ∈ D has an associated probability p(u) with which it may be erased. We first sort the data points in the order of their erasure probability, and points to a shard Di till the desired value of E(Di) is reached.” – teaches wherein data samples are position in the slices of a shard of the dataset based at least in part on a likelihood that a request will be received to remove the data samples from the dataset). Regarding claim 4, the combination of Bourtoule and Houlsby teaches the computer-implemented method of claim 1, wherein the plurality of instances of the language model are trained using a set of the adapter weights and a set of base weights (Houlsby, Pg. 2, Col. 2, Paragraph 1 – “We present a strategy for tuning a large text model on several downstream tasks.”, Fig. 2, and in Pg. 2, Col. 2, Paragraph 3 – “Adapter modules perform more general architectural modifications to re-purpose a pre trained network for a downstream task. In particular, the adapter tuning strategy involves injecting new layers into the original network. The weights of the original network are untouched, whilst the new adapter layers are initialized at random. In standard fine-tuning, the new top-layer and the original weights are co-trained. In contrast, in adapter tuning, the parameters of the original network are frozen and therefore may be shared by many tasks.” – teaches wherein the language model (large text model) is trained (trains by only tuning adapter weights and keeping base weights frozen) using a set of adapter weights (adapter weights of adapter layers injected into network) and a set of base weights (original weights)). Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the language model, adapter weights, and base weights of Houlsby to the plurality of instances of Bourtoule in order to train a plurality of instances of the language model on a set of adapter weights and base weights. Doing so would provide a method for parameter-efficient tuning for NLP and permit sequential training of a plurality of models (Houlsby, Introduction). Claim 14 is similar to claim 4, hence similarly rejected. Regarding claim 5, the combination of Bourtoule and Houlsby teaches the computer-implemented method of claim 4, wherein a subset of the weights is stored for each slice (Bourtoule, Fig. 2 Description – “One constituent model is trained on each shard by presenting it with incrementally many slices and saving its parameters before the training set is augmented with a new slice. When data needs to be unlearned, only one of the constituent models whose shards contains the point to be unlearned needs to be retrained — retraining can start from the last parameter values saved before including the slice containing the data point to be unlearned.” – teaches wherein a subset of weights is stored for each slice (parameters are saved before training is augmented with a new slice of the shard)) Bourtoule fails to explicitly teach the adapter weights and the instance of the language model is retrained using a respective set of the adapter weights without modifying the base weights. However, analogous to the field of the claimed invention, Houlsby teaches: the adapter weights (Houlsby, Pg. 2, Col. 2, Paragraph 3 – “Adapter modules perform more general architectural modifications to re-purpose a pre trained network for a downstream task. In particular, the adapter tuning strategy involves injecting new layers into the original network. The weights of the original network are untouched, whilst the new adapter layers are initialized at random. In standard fine-tuning, the new top-layer and the original weights are co-trained. In contrast, in adapter tuning, the parameters of the original network are frozen and therefore may be shared by many tasks.” and in Fig. 2 Description – “During adapter tuning, the green layers are trained on the downstream data, this includes the adapter, the layer normalization parameters, and the final classification layer” – teaches the adapter weights) and the instance of the language model is retrained using a respective set of the adapter weights without modifying the base weights (Houlsby, Pg. 2, Col. 2, Paragraph 1 – “We present a strategy for tuning a large text model on several downstream tasks.”, Fig. 2, and in Pg. 2, Col. 2, Paragraph 3 – “Adapter modules perform more general architectural modifications to re-purpose a pre trained network for a downstream task. In particular, the adapter tuning strategy involves injecting new layers into the original network. The weights of the original network are untouched, whilst the new adapter layers are initialized at random. In standard fine-tuning, the new top-layer and the original weights are co-trained. In contrast, in adapter tuning, the parameters of the original network are frozen and therefore may be shared by many tasks.” – teaches wherein the language model (large text model) is trained (trains by only tuning adapter weights and keeping base weights frozen) using a set of adapter weights (adapter weights of adapter layers injected into network) without modifying the base weights (original weights)). Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the language model and adapter weights of Houlsby to the shards, slices, checkpoints, and retraining of Bourtoule in order to retrain the language model using a set of adapter weights and base weights. Doing so would allow the model to be extended to new tasks without affecting previous ones, and would provide parameter-efficient tuning for NLP (Houlsby, Introduction). Claim 15 is similar to claim 5, hence similarly rejected. Claim 20 is similar to claims 4 and 5, hence similarly rejected. Regarding claim 7, the combination of Bourtoule and Houlsby teaches the computer-implemented method of claim 6, wherein each portion of the dataset corresponds to a shard and each shard comprises a plurality of slices, each slice comprising a portion of data samples in a respective shard (Bourtoule, Pg. 2, Col. 1, Paragraph 1 – “First, we divide the training data into multiple disjoint shards such that a training point is included in one shard only; shards partition the data. Then, we train models in isolation on each of these shards, which limits the influence of a point to the model that was trained on the shard containing the point… In addition, rather than training each model on the entire shard directly, we can divide each shard’s data into slices and present slices incrementally during training. We save the state of model parameters before introducing each new slice, allowing us to start retraining the model from the last known parameter state that does not include the point to be unlearned—rather than a random initialization.” – teaches wherein each portion of the dataset corresponds to a shard (divide training data into multiple disjoint shards) and each shard comprises a plurality of slices (divide each shard’s data into slices), each slice comprising a portion of data samples in a respective shard (divides each shard’s data into slices, thus each slice comprises a portion of data samples in a respective shard)). Claim 17 is similar to claim 7, hence similarly rejected. Regarding claim 9, the combination of Bourtoule and Houlsby teaches the computer-implemented method of claim 7, wherein each slice corresponds to a checkpoint set after training of a respective instance of the model (Bourtoule, Pg. 2, Col. 1, Paragraph 1 – “In addition, rather than training each model on the entire shard directly, we can divide each shard’s data into slices and present slices incrementally during training. We save the state of model parameters before introducing each new slice, allowing us to start retraining the model from the last known parameter state that does not include the point to be unlearned—rather than a random initialization.” – teaches the shard comprising a plurality of slices of the training data each corresponding to a checkpoint set during training a respective instance of the model (can divide each shard’s data into slices and present slices incrementally during training, can save state of model parameters before introducing each slice allowing to retrain from a checkpoint set during training)). Bourtoule fails to explicitly teach the language model. However, analogous to the field of the claimed invention, Houlsby teaches: the language model (Houlsby, Fig. 2 and Pg. 2, Col. 2, Last Paragraph – We instantiate adapter-based tuning for text Transformers. These models attain state-of-the-art performance in many NLP tasks, including translation, extractive QA, and text classification problems” – teaches a language model); Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the language model of Houlsby to the shards, slices, and training checkpoints of Bourtoule in order to set checkpoints in training a language model. Doing so would provide parameter-efficient tuning for NLP and enable the model to have perfect memory of previous tasks using a small number of task-specific parameters (Houlsby, Introduction). Regarding claim 10, the combination of Bourtoule and Houlsby teaches the computer-implemented method of claim 9, further comprising: determining a slice that contains the data sample to be removed, the slice corresponding to a checkpoint (Bourtoule, Fig. 1 and Pg. 2, Col. 1, Paragraph 1 – “We save the state of model parameters before introducing each new slice, allowing us to start retraining the model from the last known parameter state that does not include the point to be unlearned—rather than a random initialization.” – teaches determining a slice that contains the data sample to be removed, the slice corresponding to a checkpoint (saves state of model parameters before introducing each slice, thus each slice corresponds to a checkpoint, and retrains from last known state that does not include point to be unlearned, thus determining the slice that contains the data sample to be removed)) ; removing the data sample from the slice (Bourtoule, Fig. 1 and Pg. 2, Col. 1, Paragraph 1 – “We save the state of model parameters before introducing each new slice, allowing us to start retraining the model from the last known parameter state that does not include the point to be unlearned—rather than a random initialization.” – teaches removing the data sample from the slice (retrains with last known parameter state that does not include the point to be unlearned, thus removing the data sample from the slice, and further in Fig. 1, shows removing data sample from slice)); and retraining the instance of the model from a checkpoint that was most recently set before the slice was used to train the identified instance (Bourtoule, Fig. 1 and Pg. 2, Col. 1, Paragraph 1 – “We save the state of model parameters before introducing each new slice, allowing us to start retraining the model from the last known parameter state that does not include the point to be unlearned—rather than a random initialization.” – teaches retraining the instance of the model from a checkpoint that was most recently set before the slice was used to train the identified instance). Bourtoule fails to explicitly teach the language model. However, analogous to the field of the claimed invention, Houlsby teaches: the language model (Houlsby, Fig. 2 and Pg. 2, Col. 2, Last Paragraph – We instantiate adapter-based tuning for text Transformers. These models attain state-of-the-art performance in many NLP tasks, including translation, extractive QA, and text classification problems” – teaches a language model); Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the language model of Houlsby to the shards, slices, and training checkpoints of Bourtoule in order to set checkpoints in training a language model. Doing so would provide parameter-efficient tuning for NLP and enable the model to have perfect memory of previous tasks using a small number of task-specific parameters (Houlsby, Introduction). Claim 19 is similar to claims 9 and 10, hence similarly rejected. Regarding claim 11, the combination of Bourtoule and Houlsby teaches the computer-implemented method of claim 7, wherein the data samples are positioned in the slices of a shard of the dataset based on a determined ranking of the data samples (Bourtoule, Pg. 13, Col. 2, Paragraph 3 – “We discuss one such approach in Algorithm 1, under the following assumptions: (a) the distribution of unlearning requests is known precisely, and (b) this distribution is relatively constant over a time interval. Recall that each data point du ∈ D has an associated probability p(u) with which it may be erased. We first sort the data points in the order of their erasure probability, and points to a shard Di till the desired value of E(Di) is reached.” – teaches wherein data samples are positioned in the slices of a shard of the dataset based on a determined ranking of the data samples (sorts the data points in order of their erasure probability)). Regarding claim 12, the combination of Bourtoule and Houlsby teaches the computer-implemented method of claim 11 wherein data samples with a higher likelihood of being removed from the dataset are placed in slices used for training after data samples with a lower likelihood of being removed (Bourtoule, Algorithm 1 and Pg. 13, Col. 2, Paragraph 3 – “We first sort the data points in the order of their erasure probability, and points to a shard Di till the desired value of E(Di) is reached. Once this value is exceeded, we create a new shard Di+1 and restart the procedure with the residual data D\Di8. By enforcing a uniform cumulative probability of unlearning a across shards, Algorithm 1 naturally aggregates the training points that are likely to require unlearning into a fewer shards that are also smaller in size.” – teaches wherein data samples with a higher likelihood of being removed from the dataset are placed in slices used for training after data samples with a lower likelihood of being removed (creates new, empty shard and populates shard with samples of highest erasure probability by removing the lowest probability samples, thus the samples of higher erasure probability are placed in slices after samples with lower likelihood)). Regarding claim 13, the combination of Bourtoule and Houlsby teaches the computer-implemented method of claim 11 wherein data samples associated with a higher determined importance are placed in slices used for training before data samples associated with a lower determined importance (Bourtoule, Algorithm 1 and Pg. 13, Col. 2, Paragraph 3 – “We first sort the data points in the order of their erasure probability, and points to a shard Di till the desired value of E(Di) is reached. Once this value is exceeded, we create a new shard Di+1 and restart the procedure with the residual data D\Di8. By enforcing a uniform cumulative probability of unlearning a across shards, Algorithm 1 naturally aggregates the training points that are likely to require unlearning into a fewer shards that are also smaller in size.” – teaches wherein data samples with a higher importance are placed in slices used for training before data samples with a lower importance (creates new, empty shard and populates shard with samples of highest probability by removing the lowest probability samples, thus the samples of higher importance are placed in slices before samples with lower importance)). Response to Arguments Applicant's arguments filed 4 June 2026 have been fully considered but they are not persuasive. Applicant argues on pp. 3-4 of Remarks that the cited references fails to teach or suggest “identifying a slice of the shard that contains the data sample, the shard comprising…” and “retraining the identified instance of the language model using a set of adapter weights…”. Applicant further argues on pp. 3 of Remarks that Houlsby’s adapter weights are initialized at random. Examiner respectfully disagrees. Bourtoule teaches “identifying a slice of the shard that contains the data sample, the shard comprising…” at Fig. 2 and Pg. 2, Col. 1, Paragraph 1 – “In addition, rather than training each model on the entire shard directly, we can divide each shard’s data into slices and present slices incrementally during training. We save the state of model parameters before introducing each new slice, allowing us to start retraining the model from the last known parameter state that does not include the point to be unlearned—rather than a random initialization” – which describes dividing each shard’s data into slices that are presented incrementally during training. Each slice corresponds to one of a plurality of states of saved model parameters, as the model’s parameters are saved as a state before introducing each slice. Thus, each slice corresponds to one of a plurality of checkpoints set during training of the model. Bourtoule teaches “retraining the identified instance of the model using a set of weights stored at a checkpoint that was most recently set before the identified instance was trained…” at Pg. 2, Col. 1, Paragraph 1 – “We save the state of model parameters before introducing each new slice, allowing us to start retraining the model from the last known parameter state that does not include the point to be unlearned—rather than a random initialization” – which describes retraining the model from the last known parameter state that does not include the point to be unlearned, rather than random initialization. Thus Bourtoule teaches retraining the identified instance of the model using a set of weights (in Bourtoule, parameters) stored (in Bourtoule, saved) at a checkpoint that was most recently set before the identified instance was trained using the training data in the slice (in Bourtoule, retrains at the saved state most recently set before training using the point to be unlearned). Bourtoule fails to explicitly teach language models and adapter weights. However, Houlsby teaches at Pg. 2, Col. 2, Paragraphs 2-4 – tuning a large text model using a set of adapter weights that are initialized using a near-identity initialization (See Houlsby at Pg. 2, Col. 2, Paragraph 4 – “Adapter modules have two main features: a small number of parameters, and a near-identity initialization… A near-identity initialization is required for stable training of the adapted model;”). Thus, Houlsby teaches a language model tuned using adapter weights of adapter modules that are initialized using near-identity initialization. A person of ordinary skill in the art given the language model and adapter weights with near-identity initialization of Houlsby, would have found it obvious to apply the parameter saving, retraining, and unlearning techniques of Bourtoule in order to save the adapter weights at a plurality of saved states, or checkpoints, for later retraining. Doing so would meet the right to be forgotten requirements, thus achieving erasure of data from ML models (Bourtoule, Introduction), would allow the model to be extended to new tasks without affecting previous ones, and would provide parameter-efficient tuning for NLP by reducing the number of parameters to be tuned while maintaining performance (Houlsby, Introduction). See MPEP 2143(I)(G). In response to applicant's arguments against the references individually, one cannot show nonobviousness by attacking references individually where the rejections are based on combinations of references. See In re Keller, 642 F.2d 413, 208 USPQ 871 (CCPA 1981); In re Merck & Co., 800 F.2d 1091, 231 USPQ 375 (Fed. Cir. 1986). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Yu et al. (NPL: LegoNet: A Fast and Exact Unlearning Architecture, published Oct. 2022) teaches methods for machine unlearning with a framework of a fixed encoder and multiple adapters. Teaches wherein each sample can only affect very few adapters, and thus during unlearning, parameters and samples that need to be re-trained are both reduced. Teaches storing weights of the adapters, and removing a sample’s impact on the adapters during unlearning. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to LOUIS C NYE whose telephone number is 571-272-0636. The examiner can normally be reached Monday - Friday 9:00AM - 5:00PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, MATT ELL can be reached at 571-270-3264. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /LOUIS CHRISTOPHER NYE/Examiner, Art Unit 2141 /MATTHEW ELL/Supervisory Patent Examiner, Art Unit 2141
Read full office action

Prosecution Timeline

Jun 29, 2023
Application Filed
Mar 05, 2026
Non-Final Rejection mailed — §103, §112
Apr 21, 2026
Examiner Interview Summary
Apr 21, 2026
Applicant Interview (Telephonic)
Jun 04, 2026
Response Filed
Aug 25, 2026
Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12725092
SYSTEM AND METHOD FOR DECENTRALIZED FEDERATED LEARNING
5y 2m to grant Granted Sep 01, 2026
Patent 12639577
SYSTEMS AND METHODS FOR SELF SUPERVISED MULTI-VIEW REPRESENTATION LEARNING FOR TIME SERIES
4y 8m to grant Granted May 26, 2026
Patent 12524683
METHOD FOR PREDICTING REMAINING USEFUL LIFE (RUL) OF AERO-ENGINE BASED ON AUTOMATIC DIFFERENTIAL LEARNING DEEP NEURAL NETWORK (ADLDNN)
3y 2m to grant Granted Jan 13, 2026
Study what changed to get past this examiner. Based on 3 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
29%
Grant Probability
59%
With Interview (+30.0%)
4y 2m (~11m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 14 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month