Prosecution Insights
Last updated: October 02, 2026
Application No. 18/772,615

METHOD AND DEVICE FOR RETRAINING A MACHINE LEARNING SYSTEM

Non-Final OA §103§112
Filed
Jul 15, 2024
Priority
Jul 24, 2023 — DE 10 2023 207 010.3
Examiner
CADY, MATTHEW ALAN
Art Unit
Tech Center
Assignee
Robert Bosch GmbH
OA Round
1 (Non-Final)
0%
Grant Probability
At Risk
1-2
OA Rounds
1y 2m
Est. Remaining
0%
With Interview

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 1 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
3y 4m
Avg Prosecution
26 currently pending
Career history
19
Total Applications
across all art units

Statute-Specific Performance

§101
10.4%
-29.6% vs TC avg
§103
68.7%
+28.7% vs TC avg
§102
11.3%
-28.7% vs TC avg
§112
9.6%
-30.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1 resolved cases

Office Action

§103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claim 15 recites the limitation "the hyperparameter r," without reciting “a hyperparameter r” in the claim or in the parent claim 10. There is insufficient antecedent basis for this limitation in the claim. Claim 13 recites the limitation "a prefix of length at least” without saying what the prefix length must at least be. It is unclear if the prefix length must be at least a specific number or at least the value of another parameter, rendering the scope of the claim unclear. For examination, this limitation is being interpreted to mean “a prefix of length K,” as supported by the applicant’s specification ([pg. 9, line 10]). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 10, 12, 14-17 is/are rejected under 35 U.S.C.103 as being anticipated by Da Yu et al. (Hereinafter Yu) (“Differentially Private Fine-tuning of Language Models,” 2022-07-14) in view of Edward Hu et al. (hereinafter Hu) (“LORA: LOW-RANK ADAPTATION OF LARGE LAN GUAGE MODELS,” 2021-10-16). Regarding claim 10, Yu teaches; A computer-implemented method for gradient-based retraining of a machine learning system ([pg. 5] Fine-tuning is done by running DPSGD [Differential Privacy Stochastic Gradient Descent]) with non-public training data with regard to a target task ([pg. 2] model is … privately fine-tuned on a smaller, private, task-specific dataset), wherein parameters of the machine learning system have been adjusted to a basic task in a pre-training ([pg. 2] First, the model is pre-trained on a large, public dataset), the method comprising the following steps: adding further parameters to the pre-trained machine learning system ([pg. 2] First, the model is pre-trained … Next, new parameters are introduced), wherein (i) the further parameters are added by inserting at least one additional layer, parameterized with the further parameters, into an architecture of the learning system ([pg. 6] adapter-based fine-tuning, in which we modify the architecture of the pre trained model by adding new “adapter” layers … parameters of the adapter layers), and/or (ii) the further parameters are added by splitting at least one weight matrix to be adjusted in the retraining in a layer of the machine learning system, into a sum of a pre-trained weight matrix and a further summand added in the retraining, wherein the further summand is given by the matrix product of two further matrices, wherein the two further matrices are parameterized with the further parameters ([pg. 5-6] For each dense weight matrix WiPT of size a×b in the pre-trained network … Wi = WiPT + LiRi … where Li ∈ ℝ^axr, Ri ∈ ℝ^rxb are the new trainable parameters … apply this reparameterization only to the Transformer attention weights … [pg. 8] attention layers), wherein parameters of the pre-trained weight matrix have been adjusted in the pre-training ([pg. 5] weight matrix WiPT of size a×b in the pre-trained network) and are retained in the retraining ([pg. 7] pre-trained weight matrix WPT is frozen during training), and adjusting the added further parameters with the non-public training data ([pg. 2] new parameters are introduced and privately fine-tuned on a smaller, private, task-specific dataset) using a differentially private backpropagation method ([pg. 4] Differential Privacy (DP) … DP stochastic gradient descent (DPSGD) … [pg. 5] Fine-tuning is done by running DPSGD on the additional parameters θ, while freezing the weights of pre-trained model WPT), wherein the added parameters are adjusted with regard to the target task ([pg. 2] new parameters are introduced and privately fine-tuned on a smaller, private, task-specific dataset) Yu fails to explicitly disclose but Hu teaches; wherein ranks of the two further matrices are each lower than the rank of the pre-trained weight matrix; ([pg. 4] dense layers …The weight matrices in these layers typically have full-rank … For a pre-trained weight matrix W0 ∈ Rd× k, we constrain its update by representing the latter with a low-rank decomposition W0 + ∆W = W0 +BA, where B ∈ R^d×r, A ∈ R^r×k, and the rank r << min(d,k)) OBVIOUSNESS TO COMBINE HU WITH YU: Yu and Hu are analogous art to the present disclosure as they pertain parameter efficient fine-tuning of pre-trained machine learning models. It would have been obvious to one of ordinary skill in the art to implement Yu’s low-rank matrices Li and Ri according to the LoRA configuration taught by Hu, such that the ranks of Li and Ri are each lower than the rank of the corresponding pre-trained weight matrix WiPT , because Yu expressly identified Hu’s LoRA technique as the basis for Yu’s low-rank reparameterization ([Yu pg. 5] Low-Rank Adaptation (LoRA) (Hu et al., 2021)). Hu teaches selecting a rank r corresponding to the low rank matrices such that r << min(d,k) where a dense layer W0 ∈ Rd× k, thereby reducing the number of trainable parameters and the memory/computational requirements associated with fine-tuning the model ([Hu, pg. 5] we reduce that VRAM usage by up to 2/3 if r << dmodel). Thus, applying Hu’s expressly taught low-rank configuration to Yu’s LoRA-based private fine-tuning would have amounted to using the known implementation details of the very technique adopted by Yu to obtain the known benefits of parameter efficient fine tuning. Regarding claim 12, Yu teaches; wherein two adapter layers are in each case inserted at least in last L transformer blocks into the architecture of the pre-trained machine learning system ([pg. 6] adapter-based fine-tuning, in which we modify the architecture of the pre trained model by adding new “adapter” layers after each attention and feed-forward layer), and wherein the parameters added by inserting the adapter layers are adjusted ([pg. 6] When fine-tuning, … only parameters of the adapter layers, … are modified.) using the differentially private backpropagation method ([pg. 4] Differential Privacy (DP) … DP stochastic gradient descent (DPSGD) … [pg. 5] Fine-tuning is done by running DPSGD on the additional parameters θ). Regarding claim 14, Yu teaches; wherein the further parameters are added by splitting a weight matrix Wi to be adjusted in the retraining, in a layer of the machine learning system, into a sum Wi=Wi,0 + Wi,A * Wi,B of a pre-trained weight matrix Wi,0 and a product of two further matrices, … ([pg. 5] For each dense weight matrix WiPT of size a×b in the pre-trained network … Wi = WiPT + LiRi) wherein the elements of the two further matrices are each added parameters to be adjusted in the retraining ([pg. 6] where Li ∈ ℝ^axr, Ri ∈ ℝ^rxb are the new trainable parameters), wherein: Wi,0 denotes an AxB weight matrix of the machine learning system, which corresponds to the layer and the entries of which have been adjusted in the pre-training ([pg. 5] dense weight matrix WiPT of size a×b in the pre-trained network) and are not changed ([pg. 7] pre-trained weight matrix WPT is frozen during training), Wi,A denotes an Axr matrix and Wi,B an rxB matrix, the entries of which are added parameters, which are adjusted ([pg. 5-6] LoRA is an additive fine-tuning scheme … where Li ∈ ℝ^axr, Ri ∈ ℝ^rxb are the new trainable parameters) with regard to the target task ([pg. 2] model is … fine-tuned on a … task-specific dataset) by means of the differentially private backpropagation method ([pg. 4] Differential Privacy (DP) … DP stochastic gradient descent (DPSGD) … [pg. 5] Fine-tuning is done by running DPSGD on the additional parameters θ), r is a freely selectable hyperparameter determining a rank of the matrices Wi,A,Wi,B ([pg. 6] Li ∈ ℝ^axr, Ri ∈ ℝ^rxb … [pg. 8] we choose the best-performing rank r from the set {4,16,48,64}) Yu fails to explicitly teach but Hu teaches; each with a lower rank than the rank of the weight matrix Wi,0, ([pg. 4] dense layers …The weight matrices in these layers typically have full-rank … For a pre-trained weight matrix W0 ∈ Rd× k, we constrain its update by representing the latter with a low-rank decomposition W0 + ∆W = W0 +BA, where B ∈ R^d×r, A ∈ R^r×k, and the rank r << min(d,k)) OBVIOUSNESS: Using the same reasoning from claim 10. Regarding claim 15, Yu teaches; adding the further parameters by splitting a weight matrix Wi to be adjusted in the retraining, in a layer of the machine learning system, into a sum Wi=Wi,0 + Wi,A * Wi,B of a pre-trained weight matrix Wi,0 and a product of two further matrices, … ([pg. 5] For each dense weight matrix WiPT of size a×b in the pre-trained network … Wi = WiPT + LiRi) wherein the elements of the two further matrices are each added parameters to be adjusted in the retraining ([pg. 6] where Li ∈ ℝ^axr, Ri ∈ ℝ^rxb are the new trainable parameters), wherein: Wi,0 denotes an AxB weight matrix of the machine learning system, which corresponds to the layer and the entries of which have been adjusted in the pre-training ([pg. 5] dense weight matrix WiPT of size a×b in the pre-trained network) and are not changed ([pg. 7] pre-trained weight matrix WPT is frozen during training), Wi,A denotes an Axr matrix and Wi,B an rxB matrix, the entries of which are added parameters, which are adjusted ([pg. 5-6] LoRA is an additive fine-tuning scheme … where Li ∈ ℝ^axr, Ri ∈ ℝ^rxb are the new trainable parameters) with regard to the target task ([pg. 2] model is … fine-tuned on a … task-specific dataset) by means of the differentially private backpropagation method ([pg. 4] Differential Privacy (DP) … DP stochastic gradient descent (DPSGD) … [pg. 5] Fine-tuning is done by running DPSGD on the additional parameters θ), r is a freely selectable hyperparameter determining a rank of the matrices Wi,A,Wi,B ([pg. 6] Li ∈ ℝ^axr, Ri ∈ ℝ^rxb … [pg. 8] we choose the best-performing rank r from the set {4,16,48,64}) selecting the hyperparameter r for the obtained learning system with a best performance metric ([pg. 8] we choose the best-performing rank r from the set {4, 16, 48, 64}); performing the adding of the parameters ([pg. 5] we add … Li Ri) by splitting ([pg. 5] Wi = WiPT + LiRi) for adjusting the added parameters of the machine learning system ([pg. 5-6] additive fine-tuning scheme … Li Ri … are new trainable parameters) with non-public training data ([pg. 2] new parameters are introduced and privately fine-tuned on a … private … dataset) with the selected hyperparameter r ([pg. 6] Li ∈ ℝ^axr, Ri ∈ ℝ^rxb … [pg. 8] we choose the best-performing rank r from the set {4,16,48,64}). Yu does not expressly disclose, but Hu teaches; wherein the hyperparameter r is first ascertained according to the following method steps: performing multiple times (See table below) [pg. 10] PNG media_image1.png 343 917 media_image1.png Greyscale each with a lower rank than the rank of the weight matrix Wi,0, ([pg. 4] dense layers …The weight matrices in these layers typically have full-rank … For a pre-trained weight matrix W0 ∈ Rd× k, we constrain its update by representing the latter with a low-rank decomposition W0 + ∆W = W0 +BA, where B ∈ R^d×r, A ∈ R^r×k, and the rank r << min(d,k)) with in each case different specified values of the hyperparameter r ([pg. 10] with different rank r) and with public training data ([pg. 18] WikiSQL … is release under the BSD 3-Clause License … WebNLG … released under Creative Commons BY-NC-SA 4.0) so that a machine learning system with adjusted added parameters is obtained in each case ([pg. 4] weight matrix W0 ∈ Rd× k, … W0 + ∆W = W0 +BA, where B ∈ R^d×r, A ∈ R^r×k [pg. 10] r = 1, r = 2, r = 4, etc.); validating the obtained machine learning systems on public validation data in each case by ascertaining an associated performance metric ([pg. 10] Table 6: Validation accuracy on WikiSQL and MultiNLI with different rank r); OBVIOUSNESS: It would have been obvious to one of ordinary skill in the art, before the effective filing date, to modify Yu’s DP-LoRA method according to Hu by evaluating a plurality of candidate ranks r using public training data and comparing validation performance for the resulting models, because Hu expressly teaches that LoRA performance varies as a function of rank and evaluates different ranks to determine an optimal rank [see Hu section 7.2]. One of ordinary skill would have been motivated to perform such rank evaluation using Yu’s disclosed DP-LoRA training procedure so that the selected rank reflects performance under the same differentially private training regime subsequently employed on the non-public target task data. Such a modification would provide the predictable benefit of selecting a LoRA rank suited to the target task without using the private training data for hyperparameter selection. Regarding claim 16, Yu teaches; A device Yu recites a computer implemented process. The remaining limitations are substantially similar to claim 10 and are taught using the same reasoning. Regarding claim 17, Yu teaches; A non-transitory machine-readable medium on which is stored a computer program … the computer program, when executed by a processor, causing the processor to … Yu recites a computer implemented process. The remaining limitations are substantially similar to claim 10 and are taught using the same reasoning. Claim(s) 11 is/are rejected under 35 U.S.C.103 as being anticipated by Yu (“Differentially Private Fine-tuning of Language Models,” 2022-07-14) in view of Hu (“LORA: LOW-RANK ADAPTATION OF LARGE LAN GUAGE MODELS,” 2021-10-16) as applied to claim 10 above, further in view of Martín Abadi et al. (hereinafter Abadi) (“Deep Learning with Differential Privacy,” 2016-10-25). Regarding claim 11, Yu teaches; wherein the differentially private backpropagation method ([pg. 4] Differential Privacy (DP) … DP stochastic gradient descent (DPSGD)) … noisy gradients ([pg. 4] per-example gradient clipping and Gaussian noise addition steps) … non-public training data ([pg. 2] fine-tuned on a … private, … dataset [pg. 5] Fine-tuning is done by running DPSGD) Yu does not expressly disclose, but Abadi teaches; wherein the differentially private backpropagation method ([pg. 3] Differentially Private SGD) includes a step-by-step minimization of a cost function ([pg. 3] minimizing the empirical loss function L(θ)), wherein, in one step of the backpropagation method, an averaged and noisy gradient of the cost function is in each case ascertained ([pg. 3] At each step of the SGD, we compute the gradient ∇θL(θ,xi) … compute the average, add noise in order to protect privacy), wherein the averaged and noisy gradient ([pg. 3] g̃t): (i) includes a weighted sum ([pg. 3] g̃t ← 1/L (∑i ḡt(xi) + N(0,σ2C2I)) of the contributions of limited magnitude ([pg. 3] Clip gradient: ḡt(xi) ← gt(xi)/max(1, (||gt(xi)||2)/C)) of the gradients of individual … training data to the gradient of the cost function ([pg. 3] compute gt(xi) ← ∇θt L(θt,xi)), and (ii) is subjected to an additional noise term ([pg. 3] compute the gradient … add noise … N(0,σ2C2I)). OBVIOUSNESS TO COMBINE ABADI: Abadi is analogous art to the present disclosure as it pertains to computing gradients of neural network loss with respect to the model parameters for individual training examples, and using those gradients in iterative training. It would have been obvious to one of ordinary skill in the art, before the effective filing date, to implement Yu’s differentially private fine-tuning method using the DP-SGD gradient computation procedure taught by Abadi. Yu expressly teaches fine-tuning the additional parameters of a pretrained model using DP-SGD, including per-example gradient clipping and Gaussian noise addition, but does not detail the precise computation of the privatized batch gradient. Abadi (which is directly cited by Yu) teaches a known implementation of DP-SGD in which individual example gradients are clipped to limit their magnitude, the clipped gradients are averaged, Gaussian noise is added to the aggregate, and the resulting noisy averaged gradient is used to iteratively minimize the loss function. A person of ordinary skill would have been motivated to employ Abadi’s disclosed DP-SGD procedure in Yu because it provides the known gradient clipping, aggregation, and noise operations for carrying out the same differentially private SGD expressly contemplated by Yu. The combination would therefore have amounted to use of a known technique to implement Yu’s disclosed DP-SGD fine-tuning method, with the predictable result of limiting the influence of individual private training examples while providing differential privacy during gradient based retraining. Claim(s) 13 is/are rejected under 35 U.S.C.103 as being anticipated by Yu (“Differentially Private Fine-tuning of Language Models,” 2022-07-14) in view of Hu (“LORA: LOW-RANK ADAPTATION OF LARGE LAN GUAGE MODELS,” 2021-10-16) as applied to claim 12 above, further in view of Renrui Zhang et al. (hereinafter Zhang) (“LLAMA-ADAPTER: EFFICIENT FINE-TUNING OF LARGE LANGUAGE MOD ELS WITH ZERO-INITIALIZED ATTENTION,” 2023-03-28). Regarding claim 13, Yu teaches; added parameters which are adjusted ([pg. 2] new parameters are introduced and privately fine-tuned) with the differentially private backpropagation method ([pg. 5] Fine-tuning is done by running DPSGD on the additional parameters). Yu fails to explicitly teach but Zhang teaches; wherein the machine learning system is in each case prepended by a prefix ([pg. 3] the adaption prompt is concatenated with Tl along the token dimension as prefix, formulated as [Pl; Tl] ∈ R^(K+M)× C) of length at least ([pg. 3] learnable adaption prompt … Pl ∈ R^K× C with K denoting the prompt length) in the last L transformer blocks ([pg. 3] we insert the prompts into the topmost L layers of the trans former (L ≤ N)) wherein a key vector ([pg. 4] keys … Kl) and a value vector ([pg. 4] values … Vl) in a self-attention layer (NOTE: queries, keys, and values are derived from the token sequences of the same transformer layer, thereby providing self-attention, see pg. 3-4) of a transformer block can in each case be modified ([pg. 3] we modify the vanilla attention mechanisms at the last L transformer layers to be zero-init attention … l-th inserted layer) by a prefix of the associated transformer block ([pg. 3-4] adaption prompt is concatenated with Tl … as prefix, formulated as [Pl; Tl] … In the attention mechanism, … transform the input tokens into … keys, and values as … Kl = Lineark ( [Pl; Tl; tl] ); Vl = Linearv ( [Pl; Tl; tl] )), wherein an additional gating mechanism with a scalar parameter is introduced ([pg. 4] we adopt a learnable gating factor, denoted as gl), wherein parameters associated with the prefix ([pg. 3] learnable adaption prompts … Pl ∈ R^K× C) and the scalar parameter of the gating mechanism ([pg. 4] learnable gating factor, denoted as gl) are the added parameters ([pg. 8] LLaMA-Adapter only introduces a few learnable parameters and keep the pre trained 7B LLaMA frozen) OBVIOUSNESS TO COMBINE ZHANG: Zhang is analogous art to the present disclosure as it pertains to a parameter efficient adaptation of a frozen pretrained transformer. It would have been obvious to one of ordinary skill in the art, before the effective filing date, to modify Yu’s differentially private parameter efficient fine-tuning method to include the learnable prefix and gating mechanism taught by Zhang. Yu teaches adapting a pretrained transformer while freezing its pretrained parameters and training only newly introduced lightweight parameters using DP-SGD. Zhang similarly teaches parameter efficient adaptation of a frozen pretrained transformer by introducing learnable prompts in the last L transformer layers together with a learnable gating factor that controls the contribution of the prompts to the attention mechanism. One of ordinary skill would have been motivated to employ Zhang’s prefix and gating mechanism as additional trainable parameters within Yu’s parameter efficient fine-tuning framework because Zhang teaches that their mechanism permits efficient adaptation of the pretrained model while leaving the underlying pretrained parameters frozen ([Zhang, pg. 1] LLaMA-Adapter only introduces 1.2M learnable parameters upon the frozen LLaMA 7B model, and costs less than one hour for fine-tuning on 8 A100 GPUs.). Applying Yu’s disclosed DP-SGD procedure to Zhang’s newly introduced learnable prompt and gating parameters would have been a predictable use of a known parameter-efficient adaptation technique in Yu’s expressly contemplated private fine-tuning framework. CONCLUSION Any inquiry concerning this communication or earlier communications from the examiner should be directed to Matthew Alan Cady whose telephone number is (571) 272-7229. The examiner can normally be reached Monday - Friday, 7:30 am - 5:00 pm ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Cesar Paula can be reached on (571)272-4128. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MATTHEW ALAN CADY/ Examiner, Art Unit 2145 /CESAR B PAULA/ Supervisory Patent Examiner, Art Unit 2145
Read full office action

Prosecution Timeline

Jul 15, 2024
Application Filed
Aug 12, 2024
Response after Non-Final Action
Sep 02, 2026
Non-Final Rejection mailed — §103, §112 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
0%
Grant Probability
0%
With Interview (+0.0%)
3y 4m (~1y 2m remaining)
Median Time to Grant
Low
PTA Risk
Based on 1 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month