Prosecution Insights
Last updated: August 17, 2026
Application No. 18/493,962

SYSTEMS AND METHODS FOR EFFICIENT MACHINE UNLEARNING

Non-Final OA §101§103
Filed
Oct 25, 2023
Examiner
FITCH, GRANT FREDERICK
Art Unit
2124
Tech Center
2100 — Computer Architecture & Software
Assignee
JPMorgan Chase Bank, N.A.
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-55.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
4 currently pending
Career history
4
Total Applications
across all art units

Statute-Specific Performance

§101
25.0%
-15.0% vs TC avg
§103
50.0%
+10.0% vs TC avg
§102
8.3%
-31.7% vs TC avg
§112
16.7%
-23.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 0 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This office action is in response to submission of application on 10/25/2023. Claims 1-20 are presented for examination. Drawings The Drawings filed on 10/25/2023 are acceptable for examination purposes. Specification The Specification filed on 10/25/2023 is acceptable for examination purposes. The disclosure is objected to because of the following informalities: Typographical error ¶0055, “IN” should be “In”. Appropriate correction is required. Claim Objections Claims 3-5, 10-12, and 17-19 are objected to because of the following informalities: The use of the abbreviation MSA layer that is not explicitly defined within the claims or the specification. For the purposes of examination, MSA layer will be interpreted as Multi-head Self-Attention layer. Appropriate correction is required. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The analysis of the claims will follow the 2019 Revised Patent Subject Matter Eligibility Guidelines (“2019 PEG”). Step 1: Is the claim directed at one of the four statutory categories? Claims 1-7 are directed to a method, which falls within the statutory category of a process. Claims 8-14 are directed to a system comprising at least one computer including a processor and a memory, which falls within the statutory category of a machine. Claims 15-20 are directed to A non-transitory computer readable storage medium, including instructions stored thereon, which falls within the statutory category of an article of manufacture. Therefore, claims 1-20 are directed to one of the four statutory categories of invention, i.e., process, machine, manufacture, or composition of matter. (MPEP § 2106.03) Independent Claims Step 2A Prong One: Does the claim recite an abstract idea, law of nature, or natural phenomenon? Yes, independent claim 1 recites an abstract idea in the form of mental processes. A mental process is a process that “can be performed in the human mind, or by a human using a pen and paper” (MPEP § 2106.04(a)(2)(III), paragraph 1). Examples of mental processes include “observations, evaluations, judgments, and opinions” (MPEP § 2106.04(a)(2)(III), paragraph 2). The following limitations of claim 1 are mental processes: wherein the NTK-based machine unlearning algorithm is configured to: approximate a final training state of model parameters trained with an unfiltered dataset; approximate a final training state of model parameters trained with a retain dataset; [The approximation of a final training state is a mental process that can be performed by observations, evaluations, and judgements. While a specific formula is provided in the specification for approximation, no specific details of the dataset size or complexity are recited in the claim; therefore, it broadly encompasses datasets that can be approximated by the human mind with aid of pen and paper.] compute a vector for shifting parameter weights from the final training state of model parameters trained with the unfiltered dataset to the final training state of model parameters trained with the retain dataset; [Comparing two final states and computing a vector that represents the difference is a mental process that can be performed by the human mind with aid of pen and paper.] tuning a batch normalization layer of a convolutional neural network included in a machine learning model with the NTK-based machine unlearning algorithm, wherein parameters of a convolution layer of the convolutional neural network remain fixed; [Tuning a layer while certain parameters remain fixed is a mental process that can be performed by observations, evaluations, judgements, and opinions.] tuning prompt parameters of a transformer model included in the machine learning model with the NTK-based machine unlearning algorithm, wherein other parameters of the transformer model remain fixed. [Tuning parameters of a model while other parameters remain fixed is a mental process that can be performed by observations, evaluations, judgements, and opinions.] Therefore, the independent claims recite a judicial exception Step 2A Prong Two: Does the claim recite additional elements that integrate the judicial exception into a practical application? No. The judicial exception recited in the above discussed claims is not integrated into a practical application. providing a neural-tangent-kernel-based (NTK-based) machine unlearning algorithm, [This additional element is mere recitation that a judicial exception is to be performed using generic computer equipment running general class of computer algorithms in their ordinary capacity (MPEP § 2106.05(f)). This element merely provides additional instructions for providing an algorithm for applying the judicial exception, and therefore does not integrate the judicial exception into a practical application or amount to significantly more than the abstract idea itself.] Therefore, under MPEP § 2106.04(d), the additional elements of the claims do not integrate the judicial exception into a practical application. Step 2B: Does the claim recite additional elements that amount to significantly more than the judicial exception? No. The claims do not recite additional elements that are sufficient for the claims to amount to significantly more than the judicial exception. Additional elements that merely further define the judicial exception or are mere instructions to apply an exception using a generic class of computer algorithms do not constitute significantly more than a judicial exception under MPEP § 2106.05(f). Since the additional elements in the independent claims are all mere instructions to apply an exception, they do not constitute significantly more than a judicial exception. Therefore, the additional elements identified in the Step 2A Prong Two analysis do not constitute significantly more than a judicial exception. Claims 8 and 15 are substantially similar in scope and spirit to claim 1. Therefore, it would be rejected under similar analysis. Dependent Claims The remaining dependent claims being rejected do not recite additional elements, whether considered individually or in combination, that are sufficient to integrate the judicial exception into a practical application or amount to significantly more than the judicial exception. Claims 2, 9, & 16: Step 2A, Prong 1: partitioning the unfiltered dataset into a forget dataset and the retain dataset; [These further limitations merely further define the mental process recited in the parent claim and are therefore considered to be part of the judicial exception.] Step 2A, Prong 2 and Step 2B: There are no additional elements recited, as such the claim does not provide a practical application and is not considered to be significantly more. Claims 3, 10, & 17: Step 2A, Prong 1: There are no additional abstract idea limitations. Step 2A, Prong 2 and Step 2B: wherein the other parameters of the transformer model include parameters of an attention layer and parameters of an MSA layer. [This additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use, and therefore rejected under Step 2A Prong 2 and Step 2B(MPEP § 2106.05(h)). This element merely limits the judicial exception to a particular computing environment, namely a transformer model including attention and MSA layers, and therefore does not integrate the judicial exception into a practical application or amount to significantly more than the abstract idea itself.] Claims 4, 11, & 18: Step 2A, Prong 1: There are no additional abstract idea limitations. Step 2A, Prong 2 and Step 2B: wherein the MSA layer includes an input query, a key, and values. [This additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use, and therefore rejected under Step 2A Prong 2 and Step 2B(MPEP § 2106.05(h)). This additional element merely describes the details within technological environment, specifically an MSA layer, and therefore does not integrate the judicial exception into a practical application or amount to significantly more than the abstract idea itself.] Claim 5, 12, 19: Step 2A, Prong 1: wherein the prompt parameters of the transformer model are divided into key prompts and value prompts, and wherein the key prompts are prepended to the key of the MSA layer and the value prompts are prepended to the values of the MSA layer. [These further limitations merely further define the mental process recited in the parent claim and are therefore considered to be part of the judicial exception.] Step 2A, Prong 2 and Step 2B: There are no additional elements recited, as such the claim does not provide a practical application and is not considered to be significantly more. Claims 6 & 13: Step 2A, Prong 1: There are no additional abstract idea limitations. Step 2A, Prong 2 and Step 2B: wherein the NTK-based machine unlearning algorithm includes a matrix between the retain dataset and the forget dataset. [This additional element is mere recitation that a judicial exception is to be performed using generic computer equipment running general class of computer algorithms in their ordinary capacity, and therefore rejected under Step 2A Prong 2 and Step 2B (MPEP § 2106.05(f)). This element merely provides additional instructions for applying the judicial exception by specifying a matrix used in the NTK-based machine unlearning algorithm, and therefore does not integrate the judicial exception into a practical application or amount to significantly more than the abstract idea itself.] Claims 7 & 14: Step 2A, Prong 1: There are no additional abstract idea limitations. Step 2A, Prong 2 and Step 2B: Wherein the matrix between the retain dataset and the forget dataset includes a matrix whose columns are gradients of a sample from the forget dataset. [This additional element is mere recitation that a judicial exception is to be performed using generic computer equipment running general class of computer algorithms in their ordinary capacity, and therefore rejected under Step 2A Prong 2 and Step 2B (MPEP § 2106.05(f)). This element merely provides additional instructions for applying the judicial exception by further specifying details of the matrix used in the NTK-based machine unlearning algorithm, and therefore does not integrate the judicial exception into a practical application or amount to significantly more than the abstract idea itself.] Claim 20 is substantially similar in scope and spirit to claims 13 and 14. Therefore, it would be rejected under similar analysis. The following references are relied upon for the art rejection set forth below Golatkar et al. Forgetting outside the box: Scrubbing deep networks of information accessible from input-output observations, Oct 2020, hereinafter Golatkar Kanavati et al. Partial transfusion: on the expressive influence of trainable batch norm parameters for transfer learning, Feb 2021,hereinafter Kanavati Zhang et al Towards Adaptive Prefix Tuning for Parameter-Efficient Language Model Fine-tuning, July 2023, hereinafter Zhang Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Golatkar, in view of Kanavati, and Zhang. Regarding Claim 1, Golatkar discloses A method comprising: providing a neural-tangent-kernel-based (NTK-based) machine unlearning algorithm, [Golatkar: §1 Pg 3] “The forgetting procedure we propose is obtained using the Neural Tangent Kernel (NTK).” wherein the NTK-based machine unlearning algorithm is configured to: approximate a final training state of model parameters trained with an unfiltered dataset; approximate a final training state of model parameters trained with a retain dataset; and [Golatkar: §1 Pg 2] “… let a dataset D be partitioned into a subset 𝒟f to be forgotten and its complement 𝒟𝑟 to be retained.” [Golatkar: §4 Pg 8] “Using this dynamics, we can approximate in closed form the final training point when training with 𝒟 and 𝒟𝑟 …”. compute a vector for shifting parameter weights from the final training state of model parameters trained with the unfiltered dataset to the final training state of model parameters trained with the retain dataset; [Golatkar: §4 Pg 8] “… compute the optimal ‘one-shot forgetting’ vector to jump from the weights 𝓌𝒟 that have been obtained by training on 𝒟 to the weights 𝓌𝒟𝑟 that would have been obtained training on 𝒟𝑟 alone.” The ‘one-shot forgetting’ vector for jumping weights reasonably corresponds to the vector described in the claim and is calculated through a similar process. Although Golatkar discloses a machine learning model with the NTK-based machine unlearning algorithm, Golatkar does not specifically teach tuning a batch normalization layer of a convolutional neural network […] wherein parameters of a convolution layer of the convolutional neural network remain fixed; However, Kanavati is reasonably pertinent to the problem of selectively updating a subset of a network’s parameters and discloses the above limitation. [Kanavati: §6] “Our results demonstrate that fine-tuning only the batch norm affine parameters leads to similar performance as to fine-tuning all of the model parameters… overall results in faster convergence from… the fine-tuning of a smaller number of parameters without loss in performance. We observed this result with four different model architectures, and in particular with DenseNet121, highlighting the expressive power of simply scaling and offsetting outputs of pretrained convolutional layers…” Kanavati teaches tuning the trainable parameters of a batch-normalization layer in model architectures employing pretrained convolutional layers, this reasonably corresponds to tuning a batch-normalization layer of a convolutional neural network. Because only the batch-normalization parameters are fine-tuned, the remaining model parameters, including the convolution-layer parameters, remain fixed. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Kanavati’s batch-normalization fine-tuning with the NTK-based machine unlearning algorithm of Golatkar in order to achieve similar performance while reducing computation cost. [Kanavati: §6] “Our results demonstrate that fine-tuning only the batch norm affine parameters leads to similar performance as to fine-tuning all of the model parameters… This overall results in faster convergence from the use of a higher learning rate and the fine-tuning of a smaller number of parameters without loss in performance.” The combination of Golatkar and Kanavati does not specifically teach tuning prompt parameters of a transformer model included in the machine learning model with the NTK-based machine unlearning algorithm, wherein other parameters of the transformer model remain fixed. However, Zhang in the same field of endeavor discloses the above limitation. [Zhang: Appendix A] “We use the pre-trained model BERT-base …, BERT-large … RoBERTa-large … and DeBERTa-xlarge”. The experiment is performed on BERT base models which are transformer models, corresponding to a transformer model included in the machine learning model. [Zhang: Abstract] “Parameter-efficient fine-tuning has attracted attention that only optimizes a few task-specific parameters with the frozen pre-trained model. In this work, we focus on prefix tuning, which only optimizes continuous prefix vectors (i.e. pseudo tokens) inserted into Transformer layers.” The task-specific parameters reasonably correspond to the prompt parameters and it is disclosed that only task-specific parameters are fine-tuned, therefore the other parameters are is frozen, or remain fixed. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Adaptive Prefix Tuning (APT) introduced by Zhang with the NTK-based machine unlearning algorithm of Golatkar because [Zhang: Abstract] “Fine-tuning large pre-trained language models on various downstream tasks with whole parameters is prohibitively expensive… the adaptive prefix will be further tailored to each layer… enabling the fine-tuning more effective and efficient.” Regarding Claim 2, the combination of Golatkar, Kanavati, and Zhang discloses the method according to claim 1, comprising: partitioning the unfiltered dataset into a forget dataset and the retain dataset; [Golatkar: §1 Pg 2] “let a data set 𝒟 be partitioned into a subset 𝒟𝑓 to be forgotten, and its complement 𝒟𝑟 to be retained…” Regarding Claim 3, the combination of Golatkar, Kanavati, and Zhang discloses the method according to claim 1, wherein the other parameters of the transformer model include parameters of an attention layer and parameters of an MSA layer. [Zhang: §3.1] “Transformer is the block consisting of multi-head attention concatenated by multiple single self-attention functions…” Teaching that the Transformer model includes an attention layer, specifically a multi-head self-attention (MSA) layer. Further, as disclosed by Zhang above, only the prompt/prefix parameters are optimized while the pretrained model remains frozen, the remaining parameters include the parameters of the attention and MSA layers. Regarding Claim 4, the combination of Golatkar, Kanavati, and Zhang discloses the method according to claim 3, wherein the MSA layer includes an input query, a key, and values. [Zhang: §3.1] “Transformer block is calculated as follows: Attn(Q, K, V )…”. Figure 1 identifies Q, K and V as the inputs to the multi-head attention structure, and the text further identifies the K and V as the keys and values [Zhang: §3.1]. Therefore, it is broadly taught that the attention/MSA layer includes an input query, key, and value. Regarding Claim 5, the combination of Golatkar, Kanavati, and Zhang discloses the method according to claim 4, wherein the prompt parameters of the transformer model are divided into key prompts and value prompts, [Zhang: §3.1] “… by concatenating inserted Specifically, let 𝑃𝑘, 𝑃𝑣 ∈ ℝ 𝑙 x 𝑑 be the keys and values of the engaged prefix separately,”. Zhang discloses prefix parameters that reasonably correspond to prompt parameters as they are used as the prompt during prefix tuning. Further, prefix parameters comprise separate key parameters(𝑃𝑘) and value parameters (𝑃𝑣), showing that the equivalent prompt parameters are divided into key and value prompts. and wherein the key prompts are prepended to the key of the MSA layer and the value prompts are prepended to the values of the MSA layer. [Zhang: §3.1] “Attn(Q, K′, V′ )… where K′ = [Pk ;K ], V′ =[Pv ;V ] Here [ ; ] donates concatenation function.” This teaches that the key and value prefix parameters (𝑃𝑘, 𝑃v) are prepended to the original key and values (K, V ) of the MSA layer by concatenation. As noted above, the prefix parameters reasonably correspond to the prompt parameters in prefix training. Regarding Claim 6, the combination of Golatkar, Kanavati, and Zhang discloses the method according to claim 2, wherein the NTK-based machine unlearning algorithm includes a matrix between the retain dataset and the forget dataset. [Golatkar: §4 Proposition 2] “…the optimal scrubbing procedure under the NTK approximation is given by ℎNTK(𝑤)= 𝑤 +P ∇𝑓0 (𝒟-𝑓)T𝑀𝑉 where … P = I −∇𝑓0 (𝒟-𝑟)T Θ𝑟𝑟-1∇𝑓0 (𝒟-𝑟) is a projection matrix, that projects the gradients of the samples to forget ∇𝑓0 (𝒟-𝑓) onto the orthogonal space to the space spanned by the gradients of all samples to retain.” Golatkar teaches a projection matrix defined using gradients of samples from both the retain and forget dataset. This matrix corresponds to the claimed matrix between the retain and forget datasets. Regarding Claim 7, the combination of Golatkar, Kanavati, and Zhang discloses the method according to claim 6, wherein the matrix between the retain dataset and the forget dataset includes a matrix whose columns are gradients of a sample from the forget dataset. [Golatkar: §4 Proposition 2] “…the optimal scrubbing procedure under the NTK approximation is given by ℎNTK(𝑤)= 𝑤 +P ∇𝑓0 (𝒟-𝑓)T𝑀𝑉 where ∇𝑓0 (𝒟-𝑓)T is the matrix whose columns are the gradients of the sample to forget…”. Golatkar’s disclosed NTK approximation contains a matrix that directly corresponds to the claimed matrix with columns. Claims 8-14 disclose a system that implements the method of claims 1-7 respectively, with substantially the same limitations. Therefore, the rejection applied to claims 1-7 also applies. In addition, Zhang discloses A system comprising at least one computer including a processor and a memory, wherein the at least one computer is configured to: execute a neural-tangent-kernel-based (NTK-based) machine unlearning algorithm, [Zhang: Appendix A] “We conduct experiments on NVIDIA V100 or A100 GPUs for each task” Because a GPU is a processor, Zhang teaches execution of the disclosed algorithm by a processor. Further, a person of ordinary skill in the art would understand that execution of the disclosed algorithm on the GPU requires memory storing the model and associated parameters. Claims 15-20 disclose a computer program product that implement the method of claims 1-7 respectively, with substantially the same limitations. Therefore, the rejection applied to claims 1-7 also applies. In addition, Zhang discloses A non-transitory computer readable storage medium, including instructions stored thereon, which instructions, when read and executed by one or more computer processors, cause the one or more computer processors to perform steps comprising: Zhang discloses that experiments involved “…an implementation of the prefix tuning, configured with hyper-parameters public in the released code” [Zhang: §4.1] and these experiments were conducted on “NVIDIA V100 or A100 GPUs for each task.” [Zhang: Appendix A]. The released code corresponds to computer readable instructions, and execution of the released code on the GPUs corresponds to execution by one or more computer processors. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Jia et al. “Visual Prompt Tuning”Includes methods of fine-tuning prompt parameters while keeping a transformer backbone fixed. Chen et al. “Inducer-tuning: Connecting Prefix-tuning and Adapter-tuning” Prefix/prompt tuning in transformer models including key/value parameters in the attention layer/MSA while keeping portions of the model fixed. Any inquiry concerning this communication or earlier communications from the examiner should be directed to GRANT F FITCH whose telephone number is (571)270-0621. The examiner can normally be reached Bi-Week M-F 6-3 Friday Flex. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Miranda Huang can be reached at (571) 270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /G.F.F./Examiner, Art Unit 2124 /MIRANDA M HUANG/ Supervisory Patent Examiner, Art Unit 2124
Read full office action

Prosecution Timeline

Oct 25, 2023
Application Filed
Jul 28, 2026
Non-Final Rejection mailed — §101, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month