DETAILED ACTION
Status of Claims
Claim(s) 1-7 and 10-20 are pending and are examined herein.
Claim(s) 1, 7, 11, and 16 have been Amended. Claim(s) 8-9 are Canceled.
Claim(s) 1-7 and 10-20 remain rejected under 35 U.S.C. § 103.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
The amendment filed on August 06, 2026, has been entered. Claims 1-7 and 10-20 are pending in the application. Applicant’s amendments to the claims have been fully considered and are addressed in the rejections below.
Response to Arguments
Applicant's arguments, with respect to the rejection under 35 U.S.C. § 103, filed on 08/06/2026 (see Remarks Pp. 8-9), have been fully considered but are moot in view of the new grounds of rejection necessitated by amendments.
Applicant specifically argues that the cited references do not independent claims 1, 11 and 16, as currently presented. However, Applicant’s arguments are moot in view of the new grounds of rejection necessitated by amendments. The examiner notes that at least independent claims 1 and 11 are currently rejected as being unpatentable over Nayak in view of Stutz.
The examiner refers to the updated rejection under 35 U.S.C. § 103 for more details.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1-2, 7, 11-12, and 16-17are rejected under 35 U.S.C. 103 as being unpatentable over Nayak et al., (NPL: "DE-CROP: Data-efficient Certified Robustness for Pretrained Classifiers." (2022)) in view of Stutz et al., (NPL: "Random and Adversarial Bit Error Robustness: Energy-Efficient and Secure DNN Accelerators." (2022)). Hereinafter, Nayak in view of Stutz.
Regarding Currently Amended Claim 1,
Nayak discloses the following:
A method, comprising: (Nayak, [P. 2, Section: 1] “Fig. 1, our method DE-CROP significantly improves certified performance across different sample budgets (100%, 20%, 10%, 5%, 1%) on CIFAR-10 compared to [27]. Our contributions are summarized as follows: • Given only limited training data, we provide robust ness guarantees for a non-robust pretrained classifier against l2 perturbations on both white-box and black box setups. To the best of our knowledge, we are the first to provide certified adversarial defense using only the few training samples. …etc.”)
adding a first machine learning model on an input side of a deployed deep neural network model, (Nayak, [P. 4, Section: 4] “We aim to provide certified robustness to the given pre trained base classifier Bc. However, obtaining certification using randomized smoothing expects the model Bc to be robust against the input perturbations with the random Gaussian noise, which may not be the case with the model Bc supplied by the API provider. In order to make the base model Bc appropriate for randomized smoothing based certification without modifying/retraining Bc, a denoiser network Dn is prepended to Bc.”) [Examiner’s Note: the denoiser network Dn prepended to the pretrained base classifier Bc directly correspond to the “first machine learning model” added on the input side of the deployed DNN.]) wherein the deployed deep neural network comprises a restricted access architecture in which backpropagating through the model is not possible; (Nayak, [Abstract] “Existing works use this technique to provably secure a pretrained non-robust model by training a custom denoiser network on entire training data. However, access to the training set may be restricted to a handful of data samples due to constraints such as high transmission cost and the proprietary nature of the data. Thus, we formulate a novel problem of “how to certify the robustness of pretrained models using only a few training samples”.” [P. 2, Section: 1] “Given only limited training data, we provide robust ness guarantees for a non-robust pretrained classifier against l2 perturbations on both white-box and black box setups.” [P. 8, Section: 8] “In previous sections we assumed white-box access to the pre-trained base classifier Bc i.e. we can backpropagate the gradients through Bc to optimize the denoiser (Dn). However, this may not always be the case as the API provider can limit access to only Bc’s predictions (i.e. black-box) due to proprietary reasons. Since the black-box setup restricts the gradient information of Bc, we first use a black-box model stealing technique [2] to train a surrogate model: Sm.”) [Examiner’s Note: the black-box setting where the API provider can restrict gradient information. Thus, the frozen pretrained classifier in the back-box setting represents the restricted access architecture in which backpropagating through the model is not possible.])
training the first machine learning model to perform input transformation for reducing low-voltage bit errors for the deployed deep neural network …, wherein the training comprises: (Nayak, [Abstract] “We train the denoiser by maximizing the similarity between the denoised output of the generated sample and the original training sample in the classifier’s logit space.” [P. 4, Section: 4] “Thus,
B
c
∘
D
n
is the newbase classifier using which the prediction on smoothed classifier
S
c
on an ith training sample is defined as follows:
l
a
b
e
l
(
S
c
(
x
o
i
)
)
=
argmaxProb
c
∈
C
(
l
a
b
e
l
B
c
D
n
x
o
i
+
ϵ
=
c
)
where
ϵ
∼
N
(
0
,
σ
2
I
)
” [Pp. 4-5, Section: 4.1] “For any ith training sample of
D
t
r
a
i
n
l
i
m
(
i
.
e
.
,
x
o
i
)
, we obtain boundary sample (
x
b
i
) by computing adversarial noise δ for which the following relation holds:
x
b
i
=
x
o
i
+
δ
,
∥
δ
∥
∞
<
ϵ
,
ϵ
>
0
,
l
a
b
e
l
(
B
c
(
x
b
i
)
)
≠
l
a
b
e
l
(
B
c
(
x
o
i
)
)
(5) Next, we obtain interpolated features by performing a mixup between the features of generated boundary sample xi b and original training sample xi o as follows:
L
o
g
i
t
i
n
t
i
=
α
×
B
c
(
x
o
i
)
+
(
1
-
α
)
×
B
c
(
x
b
i
)
(6)”) [Examiner’s Note: the Dn network is trained as a preprocessing/denoising (input transformation) for the frozen pretrained classifier Bc(deep neural network).]
inputting training data into the first machine learning model such that, in response, the first machine learning model produces transformed training data; (Nayak, [Abstract] “We train the denoiser by maximizing the similarity between the denoised output of the generated sample and the original training sample in the classifier’s logit space.” [P. 4, Section: 4] “Thus,
B
c
∘
D
n
is the newbase classifier using which the prediction on smoothed classifier
S
c
on an ith training sample is defined as follows:
l
a
b
e
l
(
S
c
(
x
o
i
)
)
=
argmaxProb
c
∈
C
(
l
a
b
e
l
B
c
D
n
x
o
i
+
ϵ
=
c
)
where
ϵ
∼
N
(
0
,
σ
2
I
)
” [Pp. 4-5, Section: 4.1] “For any ith training sample of
D
t
r
a
i
n
l
i
m
(
i
.
e
.
,
x
o
i
)
, we obtain boundary sample (
x
b
i
) by computing adversarial noise δ for which the following relation holds:
x
b
i
=
x
o
i
+
δ
,
∥
δ
∥
∞
<
ϵ
,
ϵ
>
0
,
l
a
b
e
l
(
B
c
(
x
b
i
)
)
≠
l
a
b
e
l
(
B
c
(
x
o
i
)
)
(5) Next, we obtain interpolated features by performing a mixup between the features of generated boundary sample xi b and original training sample xi o as follows:
L
o
g
i
t
i
n
t
i
=
α
×
B
c
(
x
o
i
)
+
(
1
-
α
)
×
B
c
(
x
b
i
)
(6)”) [Examiner Notes: the training data samples (
x
o
i
) are input to Dn, and the Dn produces the transformed output (
D
n
(
x
-
o
i
), which is the denoised representation of the training sample.]
inputting the transformed training data into a surrogate model comprising a clean machine learning model in which backpropagating through the model is possible and into perturbed machine learning models, (Nayak, [Pp. 5-6, Section: 4.2] “The denoiser network Dn is attached before the pre trained base classifier Bc to make it suitable for randomized smoothing. Apart from this, we also add a domain discriminator network Dd whose input is the normalized features of the penultimate layer of Bc. The discriminator Dd learns to distinguish between distribution of clean samples and distribution of denoised outputs of gaussian perturbed input samples. Motivated from domain adaptation literature [12], we use gradient reversal layer (GRL) before feeding the normalized penultimate layer features to the discriminator that allows normal forward pass but reverses the direction of gradient in the backward pass. As a consequence, these negative gradients backpropagates to the denoiser network Dn that helps it to produce denoised output which yields in distinguishable domain-invariant features on the pretrained Bc classifier. Refer to supplementary for architectural details of the networks Dn and Dd. The overall steps involved in the proposed framework (‘DE-CROP’) are also shown in Fig. 3. The network is trained with our generated data and limited training data
D
t
r
a
i
n
l
i
m
by using different losses aimed at different objectives- label consistency to ensure correct predictions on Bc and matching of high-level feature similarity obtained on Bc for denoised output of gaussian perturbed input and clean original training samples both at the sample level and distribution level. The respective losses are described below: … we use cross entropy loss (
L
c
e
) to ensure that the label predicted by the pretrained network Bc on original clean data and denoised output of its gaussian perturbed counterpart are same.
L
l
c
=
1
N
k
∑
i
=
1
N
k
L
c
e
(
s
o
f
t
m
a
x
(
B
c
(
D
n
(
x
-
o
i
)
)
)
,
l
a
b
e
l
(
B
c
(
x
o
i
)
)
)
)
(8). ….” [P. 8, Section: 5.5] “Since the black-box setup restricts the gradient information of Bc, we first use a black-box model stealing technique [2] to train a surrogate model: Sm. We use Sm which allows gradient backpropagation, to train Dn using our proposed approach DE-CROP (refer Fig. 3). Finally, for evaluation we use the denoiser (trained via Sm) to certify robustness of the black-box classifier Bc.” [Examiner Notes: the surrogate model Sm is used in place of the target DNN (‘black-box access’) to train the denoiser network model, which allows gradient backpropagation. In other words, in the black-box setting, the pretrained deep neural network Bc, is frozen and the Bc DNN is replaced with surrogate model to train the Dn network. The Examiner notes that the recitation of “a surrogate model comprising a clean machine learning model … and into perturbed machine learning models,” is interpreted as suggesting that the perturbed models represents multiple perturbed versions of the clean/base model.]) wherein the surrogate model is selected based on a predicted architecture and performance for the deployed deep neural network comprising the restricted access architecture; (Nayak, [P. 8, Section: 5.5] “In Fig. 5, we compare the certification performance (of Bc) using denoisers trained via Sm (‘black-box access’) to the denoiser trained directly on Bc (‘white-box access’). We take Bc as Alexnet and two different choices for Sm, namely ResNet-18 and Half-Alexnet. Our method DE CROP yields very similar performance in the black box setting across different architectures of Sm. Also, the performance drop compared to white box setting is marginal, highlighting the suitability of our technique even when the pretrained classifier weights are not shared.”)
optimizing the first machine learning model based on a comparison of output of the clean machine learning model … compared to groundtruth labels for the training data. (Nayak, [P. 5, Section: 4.1] “We craft the input sample corresponding to the interpolated feature
L
o
g
i
t
i
n
t
i
(
i
.
e
.
,
x
i
_
i
n
t
)
by perturbing the original sample
x
o
i
to match its feature response to
L
o
g
i
t
i
n
t
i
using mean square error loss (Lmse). Mathematically, we obtain interpolated sample
x
i
n
t
i
as follows:
x
i
n
t
i
←
min
x
L
m
s
e
(
B
c
(
x
)
,
L
o
g
i
t
i
n
t
i
)
(7) where x is initialized with
x
o
i
and kept trainable with ground truth as
L
o
g
i
t
i
n
t
i
. The model Bc is non-trainable, but the gradients are allowed to backpropagate from the model to update the input x. Thus, we obtain the boundary and the interpolated samples corresponding to each training sample. Refer supplementary for visualization of both the boundary and interpolated samples in the input space. All of them preserve the class semantics. Our generated boundary samples
x
b
and interpolated samples
x
i
n
t
i
are used along with limited training samples xo to train the denoiser network
D
n
, which we discuss in the subsequent subsection.” [Pp. 5-6, Section: 4.2] “we use gradient reversal layer (GRL) before feeding the normalized penultimate layer features to the discriminator that allows normal forward pass but reverses the direction of gradient in the backward pass. As a consequence, these negative gradients backpropagates to the denoiser network Dn that helps it to produce denoised output which yields in distinguishable domain-invariant features on the pretrained Bc classifier. Refer to supplementary for architectural details of the networks Dn and Dd. The overall steps involved in the proposed framework (‘DE-CROP’) are also shown in Fig. 3. The network is trained with our generated data and limited training data
D
t
r
a
i
n
l
i
m
by using different losses aimed at different objectives- label consistency to ensure correct predictions on Bc and matching of high-level feature similarity obtained on Bc for denoised output of gaussian perturbed input and clean original training samples both at the sample level and distribution level. The respective losses are described below: … we use cross entropy loss (
L
c
e
) to ensure that the label predicted by the pretrained network Bc on original clean data and denoised output of its gaussian perturbed counterpart are same.
L
l
c
=
1
N
k
∑
i
=
1
N
k
L
c
e
(
s
o
f
t
m
a
x
(
B
c
(
D
n
(
x
-
o
i
)
)
)
,
l
a
b
e
l
(
B
c
(
x
o
i
)
)
)
)
(8). … we also train the domain discriminator network Dd (parameterized by ϕ) using binary cross entropy loss (
L
b
c
e
) to distinguish the distribution of gaussian perturbed samples and clean samples… The negative gradients are backpropagated via GRL [11] (multiplies calculated gradient by −1), to update the parameters θ of denoiser network Dn such that features of limited training data
D
t
r
a
i
n
l
i
m
and its corresponding denoised output of gaussian corrupted data (
D
n
(
D
-
t
r
a
i
n
l
i
m
)
) on network Bc are domain invariant. Hence, the total loss can be written as follows:
L
(
D
n
θ
,
D
d
ϕ
)
=
β
1
L
l
c
-
β
2
L
c
s
+
β
3
L
m
m
d
+
β
4
L
d
d
(12) At test time, the trained denoiser network
D
n
with optimal parameters
θ
*
prepended with base classifier
B
c
is used for evaluation.) [Examiner’s Note: The Dn network is optimized via cross-entropy loss between the output of the classifier (i.e., Sm surrogate model) output on the denoised data (transformed input) and the reference labels (i.e., pseudo-label from the limited training sample), which serves as the ground truth. See Fig. 3 and Eq. 8.]
As outlined above, Nayak discloses the proposed method (DE-CROP) that provides robust guarantees on pretrained classifier (DNN) by training denoiser network (i.e., first machine learning model) using surrogate model (
S
m
) during a black-box access setting, which allows gradient backpropagation to train the denoiser network, and prepending the trained denoiser network with optimal parameters with the pretrained classifier (deployed DNN). However, Nayak does not appear to explicitly teach:
Salient on whether the pretrained classifier (DNN) operating in a low-voltage regime;
the perturbed machine learning models being generated by applying random bit errors to the clean machine learning model; and
optimizing the first machine learning model based on a comparison of output of the clean machine learning model and of the perturbed machine learning models compared to groundtruth labels for the training data.
However, Nayak in view of Stutz teaches the limitations:
training … for reducing low-voltage bit errors for a deep neural network operating in a low-voltage regime (Stutz, [Abstract] “Low-voltage operation of DNN accelerators allows to further reduce energy consumption, however, causes bit-level failures in the memory storing the quantized weights. Furthermore, DNN accelerators are vulnerable to adversarial attacks on voltage controllers or individual bits. In this paper, we show that a combination of robust fixed-point quantization, weight clipping, as well as random bit error training (RANDBET) or adversarial bit error training (ADVBET) improves robustness against random or adversarial bit errors in quantized DNN weights significantly.” [P. 1, Section: 1] “In this paper, we aim to enable very low voltage operation of DNN accelerators by developing DNNs robust to random bit errors in their (quantized) weights. This also improves security against manipulation of voltage settings [8]. Furthermore, we address robustness against a limited number of adversarial bit errors, similar to [11], [12], [13].” [P. 4, Section: 3.1, Col. 2] “the probability of memory bit cell failures increases exponentially as operating voltage is scaled below Vmin, i.e., the minimal voltage required for reliable operation, see Fig. A.” [P. 6, Section: 4, Col. 2] “we perform random bit error training (RANDBET) (Sec. 4.3) or adversarial bit error training (ADVBET) (Sec. 4.4). For RANDBET, in contrast to the fixed bit error patterns in [4], [19], we train on completely random bit errors and, thus, generalize across chips and voltages. Regarding ADVBET, we train on adversarial bit errors, computed as outlined in Sec. 3.2. Generalization of bit error robustness is measured using robust test error (RErr), the test error after injecting bit errors (lower is more robust).”)
inputting the transformed training data into a surrogate model comprising a clean machine learning model in which backpropagating through the model is possible and into perturbed machine learning models, the perturbed machine learning models being generated by applying random bit errors to the clean machine learning model, (Stutz, [pp. 1-2, Section: 1] “We demonstrate that the obtained DNNs do not generalize across voltages or to unseen bit error patterns, e.g., from other memory arrays, and propose random bit error training (RANDBET), in combination with weight clipping and robust quantization, to obtain robustness against completely random bit error patterns, see Fig. b (violet). Thereby, it generalizes across chips and voltages, without any profiling, hardware-specific data mapping or other circuit-level mitigation strategies. Finally, in contrast to [4], [19], we also consider bit errors in activations and inputs, as both are temporally stored on the chip’s memory and thus subject to bit errors.” [p. 4, Section: 3.1, Col. 2] “Random Bit Error Model: The probability of a bit error is p (in %) for all weight values and bits. For a fixed memory array, bit errors are persistent across supply voltages, i.e., bit errors at probability p≤p also occur at probability p. A bit error flips the currently stored bit. Random bit error injection is denoted BErrp.” [Pp. 5-6, Sections: 3.1-3.2] “Fig. e: Random Bit Error Training (RANDBET). We illustrate the data-flow for RANDBET as in Alg. 2. Here, BErrp injects random bit errors in the quantized weights v(t) = Q(w(t)), resulting in ˜v(t), while the forward pass is performed on the de-quantized perturbed weights ˜ w(t) q =Q−1(˜v(t))...” [Algorithm 2: # perturbed forward and backward pass:
ω
~
q
t
=
Q
-
1
B
E
r
r
p
v
t
o
r
A
D
V
B
I
T
E
R
R
O
R
S
v
t
,
ϵ
]. … Algorithm 2 Random Bit Error Training (RANDBET). The forward passes are performed using de-quantized weights (blue). Perturbed weights are obtained by injecting bit errors in the quantized weights (in red). The update, averaging gradients from both forward passes, is performed in floating-point (magenta). Also see Fig. e. [Forward and backward pass for unperturbed & perturbed models]” [P. 5, Section: 3.2] “Overall, our white-box threat model is defined as follows: Adversarial Bit Error Model: An adversary can flip up to bits, at most one bit per (quantized) weight, to reduce accuracy and has full access to the DNN, its weights and gradients. Note that we do not consider adversarial bit errors in inputs or activations. We also emphasize that this assumes a white-box settings where the adversary can not only access the DNNs weights, but also knows about the used quantization scheme.”) [Examiner’s Note: The proposed RANDBET training uses a clean (unperturbed version model) and perturbed models. The perturbed models by injecting random bit errors with probability p to the clean quantized weight (clean ML model). The BErrp operator is a stochastic bit-flip function applied directly to the quantized weight representation with random bit errors and provides a perturbed copy of the clean model version. The Examiner notes that the recitation of “a surrogate model comprising a clean machine learning model … and into perturbed machine learning models,” is interpreted as suggesting that the perturbed models represents multiple perturbed versions of the clean/base model.]
optimizing the first machine learning model based on a comparison of output of the clean machine learning model and of the perturbed machine learning models compared to ground-truth labels for the training data. (Stutz, [Pp. 5-6, Sections: 3.1-3.2] “Algorithm 2 Random Bit Error Training (RANDBET). The forward passes are performed using de-quantized weights (blue). Perturbed weights are obtained by injecting bit errors in the quantized weights (in red). The update, averaging gradients from both forward passes, is performed in floating-point (magenta). Also see Fig. e. Algorithm 2: # clean forward and backward pass:
∆
(
t
)
=
∇
w
∑
b
=
1
B
L
(
f
x
b
;
w
q
t
,
y
b
)
# perturbed forward and backward pass:
ω
~
q
t
=
Q
-
1
B
E
r
r
p
v
t
o
r
A
D
V
B
I
T
E
R
R
O
R
S
v
t
,
ϵ
∆
~
(
t
)
=
∇
w
∑
b
=
1
B
L
(
f
x
b
;
w
~
q
t
,
y
b
)
# average gradients and weight update:
w
(
t
+
1
)
=
w
(
t
)
-
γ
(
∆
(
t
)
+
∆
~
(
t
)
).” ) [Pp. 8-9, Section: 4.3-4.4] “In addition to weight clipping and robust quantization, we inject random bit errors with probability p during training to further improve robustness. This results in the following learning problem, which we optimize as illustrated in Fig. e:
mi
n
ω
E
L
f
x
;
ω
~
,
y
+
L
f
x
;
ω
,
y
s
.
t
.
v
=
Q
ω
,
v
~
=
B
E
r
r
p
v
,
ω
~
=
Q
-
1
v
~
.
(5) where (x,y) are labeled examples, L is the cross-entropy loss and v = Q(w) denotes the (element-wise) quantized weights w which are to be learned. BErrp(v) injects random bit errors with rate p in v. Note that we consider both the loss on clean weights and weights with bit errors.” Further see Fig. e.) [Examiner’s Note: Stutz teaches a combined optimization objective using prediction losses of both the clean and perturbed models and compared to the ground-truth label yb. Updating the model iteratively using both losses to optimized and improve training robustness directly read on the optimization step.]
Accordingly, at the effective filing date, it would have been prima facie obvious to one ordinarily skilled in the art of machine learning to modify the DE-CROP framework of Nayak to incorporate the proposed random bit error training (RANDBET) as taught by Stutz. One would have been motivated to make such a combination in order to obtain high robustness against low voltage induced, random bit errors or maliciously crafted, adversarial bit errors. Doing so would improve the accuracy and security of the DNN against adversarial attacks (Stutz [Abstract]).
Regarding Original Claim 2, Nayak in view of Stutz teaches the elements of claim 1 as outlined above, and further teaches:
calculating respective losses for the clean machine learning model and for the perturbed machine learning models; calculating a gradient based on the calculated losses; and updating parameters of the first machine learning model based on the calculated gradient. (Stutz, [Pp. 5-6, Sections: 3.1-3.2] “Algorithm 2 Random Bit Error Training (RANDBET). The forward passes are performed using de-quantized weights (blue). Perturbed weights are obtained by injecting bit errors in the quantized weights (in red). The update, averaging gradients from both forward passes, is performed in floating-point (magenta). Also see Fig. e. Algorithm 2, Lines 8-12: # clean forward and backward pass:
∆
(
t
)
=
∇
w
∑
b
=
1
B
L
(
f
x
b
;
w
q
t
,
y
b
)
. # perturbed forward and backward pass:
ω
~
q
t
=
Q
-
1
B
E
r
r
p
v
t
o
r
A
D
V
B
I
T
E
R
R
O
R
S
v
t
,
ϵ
∆
~
(
t
)
=
∇
w
∑
b
=
1
B
L
(
f
x
b
;
w
~
q
t
,
y
b
)
... Note that yb are the ground truth labels. We also consider a targeted version of Eq. (1), similar to [12], where we minimize the cross-entropy loss between predictions and an arbitrary but fixed target label:
min
v
ˇ
∑
b
=
1
B
L
f
x
b
;
Q
-
1
v
~
)
,
y
t
where yt is the same target label across all examples xb.” [Pp. 8-9, Section: 4.3-4.4] “In addition to weight clipping and robust quantization, we inject random bit errors with probability p during training to further improve robustness. This results in the following learning problem, which we optimize as illustrated in Fig. e:
mi
n
ω
E
L
f
x
;
ω
~
,
y
+
L
f
x
;
ω
,
y
s
.
t
.
v
=
Q
ω
,
v
~
=
B
E
r
r
p
v
,
ω
~
=
Q
-
1
v
~
.
(5) where (x,y) are labeled examples, L is the cross-entropy loss and v = Q(w) denotes the (element-wise) quantized weights w which are to be learned. BErrp(v) injects random bit errors with rate p in v. Note that we consider both the loss on clean weights and weights with bit errors... Following Alg. 2, we use stochastic gradient descent to optimize Eq. (5), by performing the gradient computation using the perturbed weights
w
~
=
Q
-
1
(
v
~
)
with
v
~
=
B
E
r
r
p
(
v
)
, while applying the gradient update on the (floating-point) clean weights w.” Further see Fig. e.)
Regarding Currently Amended Claim 7, Nayak in view of Stutz teaches the elements of claim 1 as outlined above:
inputting additional data into the trained first machine learning model such that the trained first machine learning model transforms the additional data; inputting the transformed additional data into the deployed deep neural network; performing, via the deployed deep neural network, an inference based on the transformed additional data. (Nayak, [P. 3, Section: 3] “The complete original training dataset is de noted by Do = {Dtrain,Dtest}, where Dtrain and Dtest are the training set and the test set respectively.” [0076] “In order to make the base model Bc appropriate for randomized smoothing based certification without modifying/retraining Bc, a denoiser network Dn is prepended to Bc. Thus, Bc◦Dn is the new base classifier using which the prediction on smoothed classifier Sc on an ith training sample…” [P. 6, Section: 4.2] “At test time, the trained denoiser network Dn with optimal parameters θ∗ prepended with base classifier Bc is used for evaluation.” [P. 8, Section: 5.5] “Finally, for evaluation we use the denoiser (trained via Sm) to certify robustness of the black-box classifier Bc.”)
Regarding Currently Amended Claim 11,
The claim recites substantially similar limitations as corresponding claim 1 and is rejected for similar reasons as claim 1 using similar teachings and rationale. Claim 1 is directed to a method, and claim 11 is directed to a system.
Nayak in view of Stutz also discloses the implementation using CPU/GPU and memory.
Regarding Original Claim 12,
The claim recites substantially similar limitations as corresponding claim 2 and is rejected for similar reasons as claim 2 using similar teachings and rationale.
Regarding Currently Amended Claim 16,
The claim recites substantially similar limitations as corresponding claim 1 and is rejected for similar reasons as claim 1 using similar teachings and rationale. Claim 1 is directed to a method, and claim 16 is directed to a computer program product.
Nayak in view of Stutz also discloses a computer program product comprising program instructions (i.e., PyTorch) executable by one or more processors (CPU/GPU) perform operations.
Regarding Original Claim 17,
The claim recites substantially similar limitations as corresponding claim 2 and is rejected for similar reasons as claim 2 using similar teachings and rationale.
Claim(s) 3, 13, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Nayak in view of Stutz as outlined above and further in view of Hariharan et al., (Pub. No.: US 20230409673 A1).
Regarding Original Claim 3, Nayak in view of Stutz teaches the elements of claim 1 as outlined above.
While Nayak in view of Stutz defines multiple perturbed models (i.e., perturbed weights), Nayak in view of Stutz does not appear to explicitly suggest:
wherein the perturbed machine learning models consist of from five to ten perturbed machine learning models.
However, Hariharan, in combination with Nayak and Stutz, teaches the limitation:
wherein the perturbed machine learning models consist of from five to ten perturbed machine learning models. (Hariharan, [0039] “the perturbation component of the computerized tool can electronically generate a set of perturbed instantiations of the trained neural network. More specifically, the perturbation component can copy and/or replicate the trained neural network any suitable number of times, thereby yielding a set of network copies each of which can be identical to the trained neural network.” [0077] “the perturbation component 114 can electronically create a set of network copies 302 based on the trained neural network 104. In various aspects, the set of network copies 302 can include n network copies, for any suitable positive integer n...”) [Examiner’s Note: the defined range falls within the scope disclosed by the prior art.]
Accordingly, it would have been obvious to a person having ordinary skill in the art, before the effective filing date of the claimed invention, having the combination of Nayak, Stutz, and Hariharan, to incorporate the method that facilitates improved uncertainty scoring for neural networks via stochastic weight perturbations as taught by Hariharan. One would have been motivated to make such a combination in order to facilitate uncertainty scoring of an already-trained neural network by computing a standard deviation of predictions outputted by a set of perturbed versions of the already-trained neural network. Doing so would provide improvement in the field of neural networks (Hariharan [0140]).
Regarding Original Claim 13,
The claim recites substantially similar limitations as corresponding claim 3 and is rejected for similar reasons as claim 3 using similar teachings and rationale.
Regarding Original Claim 18,
The claim recites substantially similar limitations as corresponding claim 3 and is rejected for similar reasons as claim 3 using similar teachings and rationale.
Claim(s) 4, 6, 10, 14, and 19 is rejected under 35 U.S.C. 103 as being unpatentable over Nayak in view of Stutz as outlined above and further in view of Meng et al., (NPL: “MagNet: a Two-Pronged Defense against Adversarial Examples.” (2017)).
Regarding Original Claim 4, Nayak in view of Stutz teaches the elements of claim 1 as outlined above:
Stutz in view of Chen does not appear to explicitly define the model architecture as an encoder-decoder structure. However, it would have been obvious in view of Meng.
Hereinafter, Meng, in combination with Nayak and Stutz, teaches:
wherein the first machine learning model comprises an encoder-decoder structure. (Zhang, [P. 2, Section: 2.1] “MagNet uses a reformer to reform adversarial examples. For this we use autoencoders, which are neural networks trained to attempt to copy its input to its output. Autoencoders leverage simpler hidden representation to introduce regularization to uncover useful properties of the data [6, 35, 36]. We train an autoencoder with adequate normal examples for it to learn an approximate manifold of the data. Given an adversarial example x close to the boundary of the manifold, we expect the autoencoder to output an example y on the manifold where y is close to x.” [P. 7, Section: 4.2.2] “Autoencoder-based reformer. We propose to use autoencoders as the reformer. We train the autoencoder to minimize the reconstruction error on the training set and ensures that it generalizes well on the validation set.”)
Accordingly, it would have been obvious to a person having ordinary skill in the art, before the effective filing date of the claimed invention, having the combination of Nayak, Stutz, and Meng, to incorporate the MagNet framework for defending neural network classifiers against adversarial examples as taught by Meng. One would have been motivated to make such a combination in order to improves the classification accuracy of adversarial examples while keeping the classification accuracy of normal examples unchanged (Meng [Section: 4.2.2]).
Regarding Original Claim 6, Nayak in view of Stutz teaches the elements of claim 1 as outlined above:
Nayak in view of Stutz does not appear to explicitly suggest:
wherein the transformed training data lies within a valid range based on a type of the training data.
However, it would have been obvious in view of Meng. Hereinafter, Meng, in combination with Nayak and Stutz, teaches the limitation:
wherein the transformed training data lies within a valid range based on a type of the training data. (Meng, [P. 7, Section: 4.2] “The reformer is a function r : S → Nt that tries to reconstruct the test input. The output of the reformer is then fed to the target classifier. Note that we do not use the reformer when training the target classifier, but use the reformer only when deploying the target classifier. … A naive reformer is a function that adds random noise to the input. If we use Gaussian noise, we get the following reformer r(x) = clip(x +ϵ · y) where y~N(y;0,I) is the normal distribution with zero mean and identity covariance matrix, ϵ scales the noise, and clip is a function that clips each element of its input vector to be in the valid range.” [P. 8, Section: 5.1] “The accuracy of both these classifiers is near the state of the art on these datasets. Table 1 and Table 2 show the architecture and training parameters of these classifiers. We used a scaled range of [0,1] instead of [0,255] for simplicity.”)
The same motivation that was utilized for combining Nayak, Stutz, and Meng as set forth in claim 4 is equally applicable to claim 6.
Regarding Original Claim 10, Nayak in view of Stutz teaches the elements of claim 7 as outlined above:
Nayak in view of Stutz does not appear to explicitly suggest:
wherein the first machine learning model comprises a DNN model having fewer layers than the deployed DNN model has.
However, it would have been obvious in view of Meng. Hereinafter, Meng, in combination with Nayak and Stutz, teaches the limitation:
wherein the first machine learning model comprises a DNN model having fewer layers than the deployed DNN model has. (Meng, [P. 8, Tables 1, 3, and 4] Table 1 defines the architecture of the target deep neural network classifier to be protected. Table 3 defines the defensive models architecture including detector/reformer. It shows that the prepended defensive DNN (reformer/autoencoder) has fewer layers than the targeted/deployed DNN classifier. [P. 7, Section: 4.2] “The reformer is a function r : S → Nt that tries to reconstruct the test input. The output of the reformer is then fed to the target classifier. Note that we do not use the reformer when training the target classifier, but use the reformer only when deploying the target classifier.” Figure 2: MagNet workflow in test phase. MagNet includes one or more detectors. It considers a test example x adversarial if any detector considers x adversarial. If x is not considered adversarial, MagNet reforms it before feeding it to the target classifier. )
The same motivation that was utilized for combining Nayak, Stutz, and Meng as set forth in claim 4 is equally applicable to claim 10.
Regarding Original Claim 14,
The claim recites substantially similar limitations as corresponding claim 4 and is rejected for similar reasons as claim 4 using similar teachings and rationale.
Regarding Original Claim 19,
The claim recites substantially similar limitations as corresponding claim 4 and is rejected for similar reasons as claim 4 using similar teachings and rationale.
Claim(s) 5, 15, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Nayak in view of Stutz as outlined above and further in view of Zhang et al., (Pub. No.: WO 2022193077 A1).
Regarding Original Claim 5, Nayak in view of Stutz teaches the elements of claim 1 as outlined above:
Nayak in view of Stutz does not appear to explicitly define:
the first machine learning model is selected from a group consisting of a convolution-based model, a deconvolution-based model, and a U-Net-based model.
However, it would have been obvious in view of Zhang. Hereinafter, Zhang, in combination with Nayak and Stutz, teaches:
wherein the first machine learning model is selected from a group consisting of a convolution-based model, a deconvolution-based model, and a U-Net-based model. (Zhang, [0088] “In some embodiments, the first model may be a process or an algorithm that is configured to transform non-image data into an image format. For example, the first model may include a convolutional neural network (CNN) model, a deep convolution-deconvolution network (e.g., an encoder-decoder) , a U-shaped convolutional neural network (U-Net) , a V-shaped convolutional neural network (V-Net) , a residual network (Res-Net) , a residual dense network (Red-Net) , a deep insight-feature selection algorithm, or the like, or any combination thereof.”)
Accordingly, it would have been obvious to a person having ordinary skill in the art, before the effective filing date of the claimed invention, having the combination of Nayak, Stutz, and Zhang, to incorporate the image segmentation model architecture as taught by Zhang. One would have been motivated to make such a combination in order to provide efficient and accurate systems and methods for image segmentation (Zhang [0002]).
Regarding Original Claim 15,
The claim recites substantially similar limitations as corresponding claim 5 and is rejected for similar reasons as claim 5 using similar teachings and rationale.
Regarding Original Claim 20,
The claim recites substantially similar limitations as corresponding claim 5 and is rejected for similar reasons as claim 5 using similar teachings and rationale.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
NPL: Zheng et al. " Regularizing Neural Networks via Adversarial Model Perturbation." (2021). See Fig. 1 and Algorithm 1.
PNG
media_image1.png
390
899
media_image1.png
Greyscale
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SADIK ALSHAHARI whose telephone number is (703)756-4749. The examiner can normally be reached Monday Friday, 9 A.M - 6 P.M. ET..
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Li Zhen can be reached on (571) 272-3768. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/S.A.A./Examiner, Art Unit 2121
/Li B. Zhen/Supervisory Patent Examiner, Art Unit 2121