DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
NOTE: The following rejections are interpreting noise and perturbation to have equivalent meaning based on page 16, lines 16-17 of the applicants spec: "Within the meaning of the invention, adversarial perturbations can also be understood as noise."
Claim(s) 16-17, 27-29 is/are rejected under 35 U.S.C. 103 as being unpatentable over Nitin Sharma et al. (Hereinafter Sharma) (US 20220094709 A1, 2022-03-24) in view of Sturlaugson Liessman E. (Hereinafter Liessman) (US 20180346151 A1, 2018-12-06) further in view of Edan Habler et al. (Hereinafter Habler) (“Using LSTM encoder-decoder algorithm for detecting anomalous ADS-B messages”, 2018) further in view of Angelo Sotigiu et al. (Hereinafter Angelo) (“Deep neural rejection against adversarial examples”, 2020-04-07) further in view of Almog Lahav et al. (Hereinafter Lahav) (“Mahalanonbis Distance Informed by Clustering”, 2017).
Regarding claim 16, Sharma teaches;
A computer-implemented method for training a machine learning system, ([Abstract] train a machine learning model.)
training the machine learning system ([Abstract] train a machine learning model), the training including: a. ascertaining a first training ([Abstract] training examples) ([Abstract] training examples … whose classifications satisfy a classification difference threshold) ([Abstract] classifications satisfy classification difference threshold)
b. ascertaining a first adversarial example, wherein the first adversarial example is an overlap of the first training ([Abstract] In some embodiments, a computer system perturbs, using a set of adversarial attack methods, a set of training examples…)
NOTE: The first adversarial example is a perturbed version of the original training example and is therefore an overlap.
wherein a first noise value of the first adversarial perturbation is not greater than a specifiable threshold, ([Abstract] In some embodiments, the computer system identifies … [a] set of sparse perturbed training examples includ[ing] examples whose perturbations are below a perturbation threshold)
NOTE: The set of sparse perturbed training includes examples whose perturbations / noise are below a threshold.
c. ascertaining a training output signal for the first adversarial example using the machine learning system; ([0031] Trained machine learning classifier … generates classifications 202 for the perturbed examples)
NOTE: Teaches ascertaining a training output signal for the first adversarial example using the machine learning system (generates classifications for the perturbed examples of the set of training examples)
and d. adapting at least one parameter of the machine learning system ([0061] retraining further includes updating, based on the identified error, one or more weights of the respective nodes of the classifier) according to a gradient of a loss value ([0061]retraining includes backpropagating the set of sparse perturbed training examples through the classifier to identify error), the loss value characterizing a deviation of the desired training output signal from the ascertained training output signal ([0061] retraining includes inputting the set of sparse perturbed training examples and the set of training examples into the classifier)
NOTE: Sharma teaches adapting the weights of the machine learning system according to a gradient of a loss value (backpropagation is gradient based loss optimization), the loss value characterizing a deviation of the desired training output from the ascertained training output (error based on classifier outputs).
Sharma fails to teach but Liessman teaches;
time series of input signals
the machine learning system being configured to ascertain an output signal based on a time series of input signals of a technical system,
([Abstract] machine learning models … applied to the feature data [0081] The feature extraction module 62 may be configured to determine a statistic of sensor values and/or control input values during a time window, a difference of sensor values and/or control input values during a time window, a difference between sensor values and/or control input values measured at different locations and/or different points in time, and/or a statistic of derived sensor values and/or control input values (e.g., an average difference, a moving average etc.)
the output signal characterizing a classification and/or a regression result of at least one first operating state and/or at least one first operating variable of the technical system, the method comprising the following steps:
([0049] machine learning algorithms that identify (e.g., classify) the category (one of at least two states) to which a new observation (set of extracted features) belongs)
OBVIOUSNESS TO COMBINE LIESSMAN WITH SHARMA:
Liessman and Sharma are analogous art to each other and to the present disclosure as they all pertain to data analysis using machine learning. Specifically, Sharma pertains to a method for training a machine learning model to handle adversarial attacks while Liessman pertains to a machine learning method for determining status of components in an aircraft using time-series sensor data.
Additionally, Sharma further states;
([Abstract] The disclosed techniques may advantageously enable a machine learning model to correctly classify data associated with adversarial attacks.)
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to implement the methods of training a machine learning system to detect adversarial attacks taught by Sharma using the time-series data from the technical system disclosed by Liessman as the training time-series to enable the system to accurately determine adversarial attacks in the time series input signals of the technical system.
Sharma and Liessman Fail to teach but Habler teaches;
wherein the specifiable threshold is based on ascertained noise values of the training time series
([pg. 161] Injected anomalies. In order to evaluate the performance of the learned model, we injected three types of anomalies (in a segment of 70 sequential messages, from message 180 to message 250) into the flights included in the test sets: Random noise (RND)- anomalies are generated by adding random noise. We multiplied the original values of the message attributes of the ADS-B messages with a randomly generated floating number between zero and two.)
NOTE: The anomalies of the disclosure of Habler include random noise.
([pg. 162] In order to set the threshold value for an anomalous window, we performed a 5-fold cross-validation evaluation on the training dataset... We computed the anomaly scores for the training set (according to Eq. (2)) and defined the threshold as the value for which 95% of the anomaly scores are smaller than the value.)
PNG
media_image1.png
714
969
media_image1.png
Greyscale
NOTE: Teaches a specifiable threshold being based on ascertained noise (the threshold is determined based on the anomaly scores of the training set, where the anomalies are generated by adding random noise) from the time-series training data (the data from the training dataset is in sequential windows representing the data at different points in time, and is therefore time-series training data)
OBVIOUSNESS TO COMBINE HABLER WITH SHARMA AND LIESSMAN:
Habler is analogous art to Sharma, Liessman, and the present disclosure as they all pertain to analyzing data by utilizing machine learning. Specifically, Sharma pertains to a machine learning based security solution for detecting anomalous messages, which are altered or spoofed messages that could be sent by an attacker or compromised system.
Additionally, Habler states;
([pg. 162] We computed the anomaly scores for the training set (according to Eq. (2)) and defined the threshold as the value for which 95% of the anomaly scores are smaller than the value. … From the results we can infer from the results that the proposed model can efficiently predict an ongoing anomaly)
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to use Habler’s method of determining the threshold for the training system of claim Sharma as modified by Liessman, to efficiently recognize adversarial examples having an exceedingly high perturbation magnitude.
Sharma, Liessman, and Habler fail to teach but Angelo teaches;
wherein the first adversarial perturbation is ascertained according to the following steps:
NOTE: Ascertaining an adversarial example includes ascertaining an adversarial perturbation coupled with the original sample. From this, when ascertaining an adversarial example, you are also ascertaining an adversarial perturbation. Therefore, ascertaining an adversarial example can be used interchangeably with ascertaining an adversarial perturbation.
h. providing a second adversarial perturbation;
[AltContent: rect]
PNG
media_image2.png
41
717
media_image2.png
Greyscale
[pg. 4] ^
NOTE: x within the loop is provided as a second adversarial perturbation.
i. ascertaining a third adversarial perturbation, wherein with respect to the first training time series, the third adversarial perturbation is stronger than the second adversarial perturbation;
[AltContent: rect]
PNG
media_image2.png
41
717
media_image2.png
Greyscale
[pg. 4] ^
NOTE: x' within the loop is considered the ascertained 3rd perturbation. Each iteration of the ascertained third perturbation (x' within the loop) results in a more adverse model result than the second adversarial perturbation (x within the loop) therefore making the ascertained third perturbation stronger than the second perturbation.
wherein the third adversarial perturbation ([Algorithm 1] x’) is ascertained based on the second adversarial perturbation ([Algorithm 1] x), a specifiable step-width value ([pg. 4] variable step size η),
a gradient ([algorithm 1] ∇) for gradient ascent,
[pg. 4]
PNG
media_image3.png
60
679
media_image3.png
Greyscale
[Algorithm 1]
PNG
media_image2.png
41
717
media_image2.png
Greyscale
NOTE: The paper defines the attack objective in a minimization form, but the same objective can be rewritten as a maximization of its negative. Therefore, performing projected gradient descent on omega can then be equivalently considered gradient ascent on negative omega. See further reasoning below:
PNG
media_image4.png
930
1430
media_image4.png
Greyscale
j. providing the third adversarial perturbation as the first adversarial perturbation when a distance of the third adversarial perturbation from the second adversarial perturbation is less than or equal to a specifiable threshold;
[AltContent: rect][AltContent: rect]
PNG
media_image2.png
41
717
media_image2.png
Greyscale
[pg. 4] ^
NOTE: The third perturbation (x' within the loop) is provided as the first perturbation (returned x') when the distance of the third adversarial perturbation from the second adversarial perturbation (distance between omega x' and omega x) is less than or equal to a specifiable threshold (t)
k. otherwise, when a noise value of the third adversarial perturbation is less than or equal to an expected noise value, performing step i., wherein, in the performance of step i., the third adversarial perturbation is used as the second adversarial perturbation;
[AltContent: rect][AltContent: rect][AltContent: rect]
PNG
media_image2.png
41
717
media_image2.png
Greyscale
[pg. 4] ^
NOTE: Otherwise, when a noise value of the third adversarial perturbation is less than or equal to an expected noise value, performing step i. (when the noise value [perturbation] of the third adversarial perturbation x’ within the loop is less than or equal to epsilon [epsilon indicating the allowed perturbation, which is considered an expected noise value] x’ is not a projected value, and after each iteration where the ‘until’ condition is not met, the loop restarts, thereby performing step i again), wherein, in the performance of step i., the third adversarial perturbation is used as the second adversarial perturbation (in each iteration, the previous third adversarial perturbation x' is used as the second adversarial perturbation x);
1. otherwise, ascertaining a projected perturbation and performing step j., wherein, in the performance of step j., the projected perturbation is used as the third adversarial perturbation, and wherein the projected perturbation is ascertained by an optimization such that a distance of the projected perturbation from the second adversarial perturbation is as small as possible and the noise value of the projected perturbation is equal to the expected noise value.
[AltContent: rect][AltContent: rect][AltContent: rect]
PNG
media_image2.png
41
717
media_image2.png
Greyscale
[pg. 4] ^
[AltContent: rect][AltContent: rect]
PNG
media_image5.png
234
799
media_image5.png
Greyscale
[pg. 4] ^
NOTE: Teaches otherwise, ascertaining a projected perturbation (when the candidate within the projection operator is greater than epsilon [epsilon being the aforementioned expected noise value], a projected perturbation is ascertained using the projection operator) and performing step j., wherein, in the performance of step j., the projected perturbation is used as the third adversarial perturbation (the projected perturbation indicated by the projection operator is used as the third perturbation, x', and the loop restarts, performing step j again), and wherein the projected perturbation is ascertained by an optimization (omega(x) within the projection operation indicates an optimization) such that a distance of the projected perturbation from the second adversarial perturbation is as small as possible (the projection operator is bounded by a constraint which keeps the projected perturbation as close to the second adversarial perturbation (x) as possible [the distance of the perturbation from the second adversarial perturbation x must be less than or equal to epsilon]) and the noise value of the projected perturbation is equal to the expected noise value (when the projected perturbation exceeds the expected noise value epsilon, the perturbation is projected back onto the expected noise value defined by the bounds of epsilon)
OBVIOUSNESS TO COMBINE ANGELO:
Angelo is analogous art to Sharma, Liessman, and Habler as well as the present invention as they all pertain to Machine Learning. Specifically, Angelo pertains to making Machine Learning algorithms more robust to adversarial attacks.
Additionally, Angelo states;
[pg. 3-4]
PNG
media_image6.png
46
312
media_image6.png
Greyscale
PNG
media_image7.png
146
308
media_image7.png
Greyscale
PNG
media_image8.png
58
316
media_image8.png
Greyscale
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to use the time-series training data of Liessman in the gradient based optimization process taught by Angelo to create optimized adversarial perturbations for training the system of Sharma as modified by Liessman and Habler, to properly evaluate the adversarial robustness of the system.
Sharma, Liessman, Habler, and Angelo fail to teach but Lahav teaches;
and a first matrix ([pg. 7, algorithm 1] Ũk = Uk) characterizing a specifiable number k of greatest eigenvalues and corresponding eigenvectors ([pg. 7] Uk … taking the K eigenvectors of ΣD with the largest eigenvalues) of a covariance matrix ([pg. 3] sample covariance ΣD) of at least a subset of the plurality of ([pg. 7] input: Data matrix D)
OBVIOUSNESS TO COMBINE LAHAV:
Lahav is analogous art to the present disclosure as it pertains to determining the k greatest eigenvalues and corresponding eigenvectors of a covariance matrix.
Sharma teaches generating adversarial perturbations to train a machine learning model, Liessman teaches time-series training data of a technical system, Angelo teaches an iterative projected gradient-based optimization for determining an effective third adversarial perturbation based on a second adversarial perturbation, and Lahav teaches obtaining the greatest eigenvalues and corresponding eigenvectors of a covariance matrix, identifying the principal directions of variation represented by the covariance structure of the data.
One of ordinary skill in the art, before the effective filing date, would have been motivated to use Lahav’s covariance-derived principal directions to inform the direction of the gradient-based adversarial perturbation optimization of Angelo, in the system of Sharma as modified by Liessman, Habler, and Angelo, in order to account for the covariance structure of the training data when determining the direction of the adversarial perturbation, predictably directing the iterative perturbation toward principal directions of variation actually represented in the data, thereby providing a data informed manner of searching for effective adversarial perturbations.
Thus, the combination of Liessman, Angelo, and Lahav reasonably teaches;
wherein the third adversarial perturbation is ascertained based on the second adversarial perturbation, a specifiable step-width value, a gradient for gradient ascent (Angelo), and a first matrix characterizing a specifiable number k of greatest eigenvalues and corresponding eigenvectors of a covariance matrix of at least a subset of the plurality of (Lahav) training time series (Liessman);
Regarding claim 17, Sharma and Liessman fail to teach but Habler teaches;
wherein the specifiable threshold corresponds to an average noise value of the first training time series of the plurality of training time series.
PNG
media_image1.png
714
969
media_image1.png
Greyscale
([pg. 162] In order to set the threshold value for an anomalous window, we performed a 5-fold cross-validation evaluation on the training dataset... We computed the anomaly scores for the training set (according to Eq. (2)) and defined the threshold as the value for which 95% of the anomaly scores are smaller than the value.)
NOTE: Teaches the specifiable threshold corresponding to an average noise value of the first training time series (the threshold is set to the typical or average noise value of the training time series, as shown in the above image and excerpt) of the plurality of training time series (there are multiple different training time-series in the above image).
OBVIOUSNESS:
Using the same reasoning from claim 16, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to use Habler’s method of determining the threshold for the training system of claim Sharma as modified by Liessman, to efficiently recognize adversarial examples having an exceedingly high perturbation magnitude.
Claims 27 and 28 are apparatus claims directly corresponding to method claim 16, and are therefore rejected using the same reasoning.
Regarding claim 29, Sharma teaches;
A non-transitory machine-readable storage medium on which is stored a computer program for training a machine learning system, … the computer program, when executed by a processor, causing the processor to perform the following steps:
([0126] Various articles of manufacture that store instructions (and, optionally, data) executable by a computing system to implement techniques disclosed herein are also contemplated. The computing system may execute the instructions using one or more processing elements. The articles of manufacture include non-transitory computer-readable memory media. The contemplated non-transitory computer-readable memory media include portions of a memory subsystem of a computing device as well as storage media or memory media such as magnetic media (e.g., disk) or optical media (e.g., CD, DVD, and related technologies, etc.). The non-transitory computer-readable media may be either volatile or nonvolatile memory.)
NOTE: Teaches a non-transitory machine-readable storage medium on which is stored a computer program for training a machine learning system and the computer program, when executed by a processor, causing the processor to perform the methods of the disclosure of Sharma.
The remaining limitations directly correspond to claim 16 and are therefore rejected using the same reasoning.
Claim(s) 18-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Sharma (US 20220094709 A1, 2022-03-24) in view of Liessman (US 20180346151 A1, 2018-12-06) further in view of Habler (“Using LSTM encoder-decoder algorithm for detecting anomalous ADS-B messages”, 2018) further in view of Angelo (“Deep neural rejection against adversarial examples”, 2020-04-07) further in view of Lahav (“Mahalanonbis Distance Informed by Clustering”, 2017) further in view of Kuniaki (JP 2011090382 A, 2011-05-06).
Regarding claim 18, Sharma, Liessman, Habler, Angelo and Lahav fail to teach but Kuniaki teaches;
wherein a noise value of each training time series or adversarial perturbation or adversarial example is ascertained according to a Mahalanobis distance.
([pg. 5] acquires these monitoring target data in real time at a predetermined measurement interval, and stores them in the storage unit 24 together with information on the measurement time.)
NOTE: Teaches that the monitoring targets are time-series (a set of data measured at time intervals)
([pg. 8] the Mahalanobis distance D changes gradually due to … noise … the environmental change is determined as an abnormality of the monitoring target 10.)
NOTE: Teaches a noise value ascertained according to Mahalanobis distance of each time-series monitoring target.
OBVIOUSNESS TO COMBINE KUNIAKI:
Kuniaki is analogous art to the present disclosure as it uses Mahalanobis distance to measure a noise related value for time-series training data.
Additionally, Kuniaki states;
([pg. 4] Here, when an abnormality occurs in the operating state of the monitoring target 10, the monitoring target data at that time appears at a position outside the reference data group, and the Mahalanobis distance D increases. Therefore, by calculating the Mahalanobis distance D between the monitoring target data and the reference data group and comparing the Mahalanobis distance D with a predetermined threshold k, it can be determined whether or not the operating state of the monitoring target 10 is abnormal.)
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to use Kuniaki’s Mahalnobis distance derived noise value to quantify the noise for the training time series in the system of Sharma as modified by Liessman, Habler, and Lahav, beneficially providing a measure of noise for distinguishing ordinary noise / variation of the data from meaningful adversarial perturbations.
Regarding claim 19, Sharma teaches;
Adversarial perturbation, adversarial example
(Using the same reasoning from the claim 16 rejection)
Sharma, Liessman, Habler, and Angelo fail to teach but Lahav teaches;
wherein the
PNG
media_image9.png
26
20
media_image9.png
Greyscale
wherein s is the
[pg. 6]
PNG
media_image10.png
22
27
media_image10.png
Greyscale
NOTE: Teaches Mahalanobis distance between columns (samples) being ascertained according to the pictured formula d. By taking the square root of both sides of the pictured formula d, where (c1 – c2) can be the adversarial perturbation s of the system of claim 16 (for example, if c1 is an adversarial example and c2 is the overlapped unperturbed example, (c1 – c2) would represent the adversarial perturbation of the adversarial example c1), and
PNG
media_image10.png
22
27
media_image10.png
Greyscale
is the pseudo-inverse covariance matrix
PNG
media_image9.png
26
20
media_image9.png
Greyscale
, the pictured formula d is equivalent to the claimed formula
PNG
media_image9.png
26
20
media_image9.png
Greyscale
(see below).
PNG
media_image11.png
323
1430
media_image11.png
Greyscale
pseudo-inverse covariance matrix characterizing a specifiable number k of greatest eigenvalues and corresponding eigenvectors of at least a subset of the plurality of
[pg. 7]
PNG
media_image12.png
389
612
media_image12.png
Greyscale
NOTE: Teaches Ũk = Uk characterizing the k greatest eigenvalues (lambda) and corresponding eigenvectors (u) of at least a subset of the data, which can be the training time-series taught in claim 16 (reasoning for why this combination would be obvious will be provided later).
[pg. 7]
PNG
media_image13.png
1
1
media_image13.png
Greyscale
NOTE: Teaches the pseudo-inverse covariance matrix being ascertained using Ũk, and therefore the pseudo-inverse covariance matrix characterizes a specifiable number k of greatest eigenvalues and corresponding eigenvectors of at least a subset of the data, which can be the training time-series taught in claim 16 (reasoning for why this combination would be obvious will be provided later).
OBVIOUSNESS:
As discussed above with respect to claim 18, Kuniaki teaches using a Mahalanobis distance to provide a quantitative measure of noise-related deviation of time-series training data, while Lahav further teaches calculation of a Mahalanobis distance using a pseudo-inverse of the covariance matrix and k greatest eigenvectors and corresponding eigenvalues.
Additionally, Lahav states;
[pg. 3]
PNG
media_image14.png
102
1027
media_image14.png
Greyscale
NOTE: Lahav explains that the pseudo-inverse of the covariance matrix is used when the data is not full rank, which can make the covariance matrix singular or not invertible.
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to implement Kuniaki’s Mahalanobis distance based noise determination using Lahav’s pseudo-inverse covariance matrix technique, including the k eigenvectors corresponding to the k greatest eigenvalues, beneficially allowing the Mahalanobis distance based noise value to be determined even when the covariance matrix of the training data is not full rank and therefore cannot be conventionally inverted.
Sharma, Liessman, Habler, Angelo and Lahav fail to teach but Kuniaki teaches;
noise value ascertained by Mahalanobis distance
(Using the same reasoning from the claim 18 rejection)
Regarding claim 20, Sharma fails to teach but Liessman teaches;
plurality of training time series
(Using the same reasoning from the claim 16 rejection)
Sharma, Liessman, Habler, and Angelo fail to teach but Lahav teaches;
wherein the pseudo-inverse covariance matrix is ascertained by the following steps:
e. ascertaining a covariance matrix of the at least subset of the
[pg.3-4]
PNG
media_image15.png
37
688
media_image15.png
Greyscale
PNG
media_image16.png
17
20
media_image16.png
Greyscale
NOTE: The covariance matrix (
PNG
media_image16.png
17
20
media_image16.png
Greyscale
) is ascertained using at least a subset of the data (the data being the random vector c).
f. ascertaining a predefined plurality of greatest eigenvalue of the covariance matrix and eigenvectors corresponding to the eigenvalues;
[pg. 7]
PNG
media_image17.png
19
33
media_image17.png
Greyscale
NOTE: Teaches ascertaining a predefined plurality of greatest eigenvalues (k largest eigenvalues lambda) of the covariance matrix (
PNG
media_image17.png
19
33
media_image17.png
Greyscale
) and eigenvectors corresponding to the eigenvalue (k eigenvectors u);
g. ascertaining the pseudo-inverse covariance matrix according to the formula
PNG
media_image18.png
59
132
media_image18.png
Greyscale
[pg. 6]
PNG
media_image13.png
1
1
media_image13.png
Greyscale
NOTE: The pseudo-inverse covariance (
PNG
media_image13.png
1
1
media_image13.png
Greyscale
) is ascertained according to the pictured formula (
PNG
media_image13.png
1
1
media_image13.png
Greyscale
).
[pg. 7]
PNG
media_image19.png
45
642
media_image19.png
Greyscale
PNG
media_image20.png
21
144
media_image20.png
Greyscale
PNG
media_image21.png
59
264
media_image21.png
Greyscale
NOTE: From these definitions, the pseudo-inverse formula given by Lahav (
PNG
media_image13.png
1
1
media_image13.png
Greyscale
) is equivalent to the formula provided in the claims (
PNG
media_image18.png
59
132
media_image18.png
Greyscale
). See reasoning below:
PNG
media_image22.png
1190
1429
media_image22.png
Greyscale
wherein Lambda i is the i-th eigenvalue of the plurality of greatest eigenvalues, vi is the eigenvector corresponding to the eigenvalue and k is the specifiable number of greatest eigenvalues.
OBVIOUSNESS:
Using the same reasoning from claim 19, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to implement Kuniaki’s Mahalanobis distance based noise determination as taught in claim 18 using Lahav’s pseudo-inverse covariance matrix technique, including the k eigenvectors corresponding to the k greatest eigenvalues, beneficially allowing the Mahalanobis distance based noise value to be determined even when the covariance matrix of the training data is not full rank and therefore cannot be conventionally inverted.
Claim(s) 22 is/are rejected under 35 U.S.C. 103 as being unpatentable over Sharma (US 20220094709 A1, 2022-03-24) in view of Liessman (US 20180346151 A1, 2018-12-06) further in view of Habler (“Using LSTM encoder-decoder algorithm for detecting anomalous ADS-B messages”, 2018) further in view of Angelo (“Deep neural rejection against adversarial examples”, 2020-04-07) further in view of Lahav (“Mahalanonbis Distance Informed by Clustering”, 2017) further in view of N. Gopalan Nair (hereinafter Nair) (“Gradient Eigenspace Projections for Adaptive Filtering”, 1995-08-16).
Regarding claim 22, Sharma teaches;
Desired training output signal
(Using the same reasoning as the claim 16 rejection)
Sharma fails to teach but Liessman teaches;
first training time series
(Using the same reasoning as the claim 16 rejection)
Sharma, Liessman, and Habler fail to teach but Angelo teaches teaches;
wherein, in step i., the third adversarial perturbation is ascertained using the gradient ascent
[pg. 4]
[AltContent: textbox (Attack objective)][AltContent: rect]
PNG
media_image3.png
60
679
media_image3.png
Greyscale
[pg. 4]
[AltContent: rect][AltContent: rect]
PNG
media_image2.png
41
717
media_image2.png
Greyscale
NOTE: The paper defines the attack objective in a minimization form, but the same objective can be rewritten as a maximization of its negative. Therefore, performing projected gradient descent on omega can then be equivalently considered gradient ascent on negative omega. See further reasoning below:
PNG
media_image23.png
830
1316
media_image23.png
Greyscale
based on an output of the machine learning system with respect to the [example]
[pg. 4]
[AltContent: rect][AltContent: textbox (A)][AltContent: rect]
PNG
media_image3.png
60
679
media_image3.png
Greyscale
NOTE: The aforementioned gradient ascent is based on an output of a machine learning system with respect to a first source sample overlapped with the second adversarial perturbation (the aforementioned second adversarial example / perturbation is a source example overlapped with a perturbation, as detailed in ‘A’) and with respect to a desired training output signal (minimizing the output of the true class while maximizing the output of the competing class).
OBVIOUSNESS:
Using the same reasoning from claim 16, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to use the time-series training data of Liessman in the gradient based optimization process taught by Angelo to create optimized adversarial perturbations for training the system of Sharma as modified by Liessman and Habler, to properly evaluate the adversarial robustness of the system.
Sharma, Liessman, Habler and Angelo fail to teach but Nair teaches;
wherein the gradient for the gradient ascent is adapted according to the eigenvalues and eigenvectors.
[pg. 1]
PNG
media_image24.png
346
374
media_image24.png
Greyscale
PNG
media_image25.png
469
398
media_image25.png
Greyscale
NOTE: Nair teaches a method for gradient projections according to subspaces defined by eigenvalues and eigenvectors (projecting the gradient vector on the high subspace, which is a matrix of the highest eigenvectors defined by their corresponding eigenvalues). This therefore teaches adapting a gradient according to the eigenvalues and eigenvectors.
OBVIOUSNESS TO COMBINE NAIR:
Nair is analogous art to the present disclosure as it pertains to a method for adapting gradients according to eigenvalues and eigenvectors.
Angelo teaches a gradient based optimization algorithm for determining effective adversarial perturbations, Lahav teaches determining top k eigenvalues and corresponding eigenvectors of a covariance matrix, while Nair teaches adapting a gradient according to top eigenvectors and corresponding eigenvalues.
Additionally, Nair states;
PNG
media_image24.png
346
374
media_image24.png
Greyscale
NOTE: This excerpt details that the disclosed method for gradient projection onto eigen-subspaces improves convergence properties for gradient algorithms when data is highly correlated.
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, for the gradient based optimization algorithm taught by Angelo to be adapted according to eigenvalues and eigenvectors using the method disclosed by Nair, to improve the convergence properties for the gradient based optimization algorithm of the system of Sharma as modified by Liessman, Habler, Angelo, and Lahav when data is highly correlated.
Claim(s) 23 is/are rejected under 35 U.S.C. 103 as being unpatentable over Sharma (US 20220094709 A1, 2022-03-24) in view of Liessman (US 20180346151 A1, 2018-12-06) further in view of Habler (“Using LSTM encoder-decoder algorithm for detecting anomalous ADS-B messages”, 2018) further in view of Angelo (“Deep neural rejection against adversarial examples”, 2020-04-07) further in view of Lahav (“Mahalanonbis Distance Informed by Clustering”, 2017) further in view of Bai Li et al. (Hereinafter Li) (“Certified Adversarial Robustness with Additive Noise”, 2019-11-10).
Regarding claim 23, Sharma, Liessman, Habler, Angelo, and Lahav fail to teach but Li teaches;
wherein the first adversarial example is ascertained using certifiable robustness training.
([Abstract] Defensive methods that provide theoretical robustness guarantees have been studied intensively, yet most fail to obtain non-trivial robustness when a large-scale model and data are present. To address these limitations, we introduce a framework that is scalable and provides certified bounds on the norm of the input manipulation for constructing adversarial examples. We establish a connection between robustness against adversarial perturbation and additive random noise, and propose a training strategy that can significantly improve the certified bounds.)
NOTE: Teaches ascertaining an adversarial example (provides certifiable bounds on the norm for constructing adversarial examples) using certifiable robustness training (Li teaches a framework providing certified bounds on input manipulation for constructing adversarial examples, and further teaches a training strategy that improves such certified bounds. Thus, Li teaches ascertaining an adversarial example within a certifiable robustness training framework).
OBVIOUSNESS TO COMBINE LI:
Li is analogous art to the present disclosure as it pertains to certifiable robustness training.
Additionally, Li states;
([Abstract] We establish a connection between robustness against adversarial perturbation and additive random noise, and propose a training strategy that can significantly improve the certified bounds. Our evaluation on MNIST, CIFAR-10 and ImageNet suggests that the proposed method is scalable to complicated models and large data sets, while providing competitive robustness to state-of-the-art provable defense methods.)
NOTE: This excerpt explains that the method disclosed by Li gives stronger and more practically provable robustness guarantees, while being scalable to complicated models and large datasets.
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to ascertain the first adversarial example of the system of claim 16 using the certifiable robustness training disclosed by Li, to give stronger and more practically provable robustness guarantees for the system.
Claim(s) 24 is/are rejected under 35 U.S.C. 103 as being unpatentable over Sharma (US 20220094709 A1, 2022-03-24) in view of Liessman (US 20180346151 A1, 2018-12-06) further in view of Habler (“Using LSTM encoder-decoder algorithm for detecting anomalous ADS-B messages”, 2018) further in view of Angelo (“Deep neural rejection against adversarial examples”, 2020-04-07) further in view of Lahav (“Mahalanonbis Distance Informed by Clustering”, 2017) further in view of Patel Shwetak N et al. (Hereinafter Patel) (CN 102460104 B, 2015-05-20).
Regarding claim 24, Sharma teaches;
Desired training output signal
(Using the same reasoning from claim 16 rejection)
Sharma fails to teach but Liessman teaches;
Training time series
(Using the same reasoning from claim 16 rejection)
Sharma, Liessman, Habler, Angelo, and Lahav fail to teach but Patel teaches;
wherein the technical system dispenses a liquid via a valve,
([pg. 10] For each residential construction, first base line quiescent hydraulic pressure is measured, and then suitable pressure transducer (that is: scope is at the pressure transducer of 0psi to 50psi or 0psi to 100psi) is arranged on operational hose adapter, kitchen sink (utlity sink) water swivel, or on the water discharging valve of water heater.)
NOTE: Patel teaches the technical system dispensing liquid from a valve (the water discharging valve).
wherein each time series
([pg. 10] These pressure characteristics use a kind of log recording instrument of drawing to record, and this also provides the Real-time Feedback of pressure data by a kind of time series line chart of rolling.)
NOTE: Teaches each time series (the pressure data is represented by a time-series line chart) characterizing a sequence of pressure values of the technical system (teaches a time series [which is considered a sequence] of pressure data of the system).
It would be obvious for each training time series of the system of claim 16 to characterizes a sequence of pressure values of the technical system disclosed by Patel, further explained later.
and the output signal and
([pg. 4] This illustrative methods comprises the following steps, and namely monitors the fluid pressure at first place in distribution system, and in response to this, produces the output signal that illustrates pressure in distribution system.)
NOTE: The output signal characterizes an amount of liquid dispensed by the valve. (an output signal indicating the fluid pressure of the aforementioned water discharging valve would indicate the rate or amount of fluid dispensed by said valve)
It would also be obvious for the desired training output signal of the system of claim 16 to also each characterize an amount of liquid dispensed by the valve, further explained later.
OBVIOUSNESS TO COMBINE PATEL:
Patel is analogous art to the present disclosure as it pertains to monitoring and processing sensor values in a liquid distribution system.
Liessman teaches ascertaining output signals characterizing a first operating variable of a technical system based on time-series input signals, while Patel teaches time-series data characterizing a sequence of pressure values.
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, for the aforementioned technical system in the system of Sharma as modified by Liessman, Habler, Angelo, and Lahav to be the liquid distribution system disclosed by Patel, to improve monitoring of operating variables in a liquid distribution system.
Claim(s) 25 is/are rejected under 35 U.S.C. 103 as being unpatentable over Sharma (US 20220094709 A1, 2022-03-24) in view of Liessman (US 20180346151 A1, 2018-12-06) further in view of Habler (“Using LSTM encoder-decoder algorithm for detecting anomalous ADS-B messages”, 2018) further in view of Angelo (“Deep neural rejection against adversarial examples”, 2020-04-07) further in view of Lahav (“Mahalanonbis Distance Informed by Clustering”, 2017) further in view of Taehyun Kim et al. (Hereinafter Kim) (“KR 20190103084 A”, 2019-09-04).
Regarding claim 25, Sharma fails to teach but Liessman teaches;
Training time series
(Using the same reasoning from claim 16 rejection)
Sharma, Liessman, Habler, Angelo, and Lahav fail to teach but Kim teaches;
wherein the technical system is a robot
([Abstract] The intelligent electronic device of the present invention can be connected to an artificial intelligence module, a drone (unmanned aerial vehicle, UAV), a robot...)
NOTE: Teaches the technical system of the disclosure being a robot.
and each time series and
([pg. 25] The position sensor may sense the current position of the intelligent electronic device 100 in real time. The position sensor may use a GPS sensor to detect the position information, and the sensors may be triggered through the position information and the time information based on the GPS sensor.)
NOTE: Teaches each time series (the position sensor is recorded with time information, and is therefore time series) characterizes position data of the robot using a corresponding sensor (position sensor characterizes the position of the device, which can be a robot).
and the output signal or the desired training output signal characterizes a position and/or an acceleration and/or a center of gravity and/or a zero moment point of the robot.
([pg.26] The processor 180 may extract a feature value from each of the illuminance information, the sound information, and the position information. The feature value is determined to recognize the surrounding environment for the current location among at least one feature that can be extracted from the illumination information, the sound information, and the location information, and to specifically indicate whether the feature is a specific place.)
NOTE: The output signal (extracted feature values) characterizes a position of the robot (the extracted feature values can be position information of the electronic device, which can be a robot)
OBVIOUSNESS TO COMBINE KIM:
Kim is analogous art to the present disclosure as it pertains to monitoring and processing sensor values in an intelligent electronic system such as a robot.
Sharma teaches input and output signals of a machine learning system, Liessman teaches ascertaining output signals characterizing a first operating variable of a technical system based on time-series input signals, and Kim teaches time-series data characterizing position data of a robot.
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, for the aforementioned technical system of the system of Sharma as modified by Liessman, Habler, Angelo, and Lahav to be the robot disclosed by Kim, beneficially allowing the machine learning model to provide improved monitoring of operating variables for controlling a robot.
Claim(s) 26 is/are rejected under 35 U.S.C. 103 as being unpatentable over Sharma (US 20220094709 A1, 2022-03-24) in view of Liessman (US 20180346151 A1, 2018-12-06) further in view of Habler (“Using LSTM encoder-decoder algorithm for detecting anomalous ADS-B messages”, 2018) further in view of Angelo (“Deep neural rejection against adversarial examples”, 2020-04-07) further in view of Lahav (“Mahalanonbis Distance Informed by Clustering”, 2017), further in view of Imanari Hiroyuki et al. (Hereinafter Hiroyuki) (“WO 2020194534 A1”, 2020-10-01).
Regarding claim 26, Sharma, Liessman, Habler, Angelo, and Lahav fail to teach but Hiroyuki teaches;
wherein the technical system is a production machine that produces at least one part,
([Abstract] This abnormality determination assistance device is provided with a to-be-analyzed data creation unit, a primary determination unit, and a secondary determination unit. The to-be-analyzed data creation unit acquires, from a data extraction device of a production facility, a time-series signal representing at least one of the status of the production facility and a product quality, and extracts to-be-analyzed data from the time-series signal.)
NOTE: Teaches the technical system being a production machine (production facility) that produces at least one part (product quality indicates that the production facility produces at least one part).
wherein the input signals of each of the time series characterize a force and/or a torque of the production machine,
([pg. 5] FIG. 3 is a diagram illustrating an example of the processing flow of the analysis target data creation unit 3. In step S101, when the rolling of the rolled material to be analyzed is completed in the manufacturing facility 20, a time series signal including before and after rolling is acquired from the data collecting device 1. This time-series signal includes data representing the state of the manufacturing equipment 20 and sensor data representing the product quality.
NOTE: Teaches time-series input signals representing the state of the manufacturing equipment.
([pg. 5] In step S102, data such as rolling load, rolling torque, electric machine current, and speed of rotating equipment are being rolled (load) for each rolling facility (two rolling mills and seven rolling stands constituting the finishing rolling mill).))
NOTE: The aforementioned states of the manufacturing equipment include torque (rolling torque). This therefore teaches the input signals of each the time series each characterize a force and/or a torque of the production machine.
and the output signal characterizes a classification as to whether or not the part was produced correctly.
([pg. 5] In step S103, sensor data such as plate thickness and plate width indicating product quality are classified into measurement state data and non-measurement state data.)
NOTE: Teaches that the output signal characterizes a classification as to whether the part was produced correctly (sensor data indicating product quality are classified).
OBVIOUSNESS TO COMBINE HIROYUKI:
Hiroyuki is analogous art to the present disclosure as it pertains to abnormality detection of a production machine using sensor values.
Sharma teaches a machine learning system for making predictions based on input data, Liessman teaches time-series inputs of a technical machine, and Hiroyuki teaches input and output data relating to a production machine.
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, for the technical system of Sharma as modified by Liessman, Habler, Angelo, and Lahav to be the production machine disclosed by Hiroyuki to provide monitoring and state information of the production machine, predictably improving monitoring and control of the machine.
Response to Arguments
Applicant’s arguments, starting page 1, filed 06/22/2026, with respect to 35 U.S.C. 112b have been fully considered and are persuasive. The 35 U.S.C. 112b rejection of claim 22 has been withdrawn.
Applicant’s arguments, starting page 1, filed 06/22/2026, with respect to 35 U.S.C. 101 have been fully considered and are persuasive. The 35 U.S.C. 101 abstract idea rejections of claims 16-20 and 22-29 have been withdrawn.
Applicant’s arguments, starting page 2, filed 06/22/2026, with respect to 35 U.S.C. 103 have been fully considered and are not persuasive.
Claims 16, 27-29 have been amended to incorporate the subject matter of claim 21 and additional features, claims 20, 22, and 26 have been amended to fix minor errors, and claim 21 has been cancelled. The applicant states that “amended claim 16 will be treated as if amended claim 16 was rejected over Sharma, Liessman, Habler, and Angelo for prosecution efficiency … The cited references, individually or in combination, fail to disclose or suggest the specific combination of features of amended claim 16.” Specifically, the applicant identifies the following portions of amended claim 16 as being allegedly not taught by the combination of Sharma, Liessman, Habler, and Angelo: “ascertaining a third adversarial perturbation, wherein with respect to the first training time series, the third adversarial perturbation is stronger than the second adversarial perturbation, wherein the third adversarial perturbation is ascertained based on the second adversarial perturbation, a specifiable step-width value, a gradient for gradient ascent, and a first matrix characterizing a specifiable number k of greatest eigenvalues and corresponding eigenvectors of a covariance matrix of at least a subset of the plurality of training time series… [the references] do not disclose ascertaining a third adversarial perturbation based on a first matrix that characterizes a specifiable number k of greatest eigenvalues and corresponding eigenvectors of a covariance matrix of at least a subset of the plurality of training time series… Thus, claim 16 is patentable. Independent claims 27-29 have been amended to recite similar limitations as amended claim 16 and are also patentable for the same reasons set forth above.” The applicant further remarks that claims 17-20, and 22-26 depend from claim 16, and are therefore allegedly patentable for the same reasons set forth for claim 16.
However, as reflected by the current office action, these limitations are taught by the combination of Sharma, Liessman, Habler, Angelo, and Lahav, as necessitated by the amendments:
“ascertaining a third adversarial perturbation, wherein with respect to the first training time series, the third adversarial perturbation is stronger than the second adversarial perturbation, wherein the third adversarial perturbation is ascertained based on the second adversarial perturbation, a specifiable step-width value, a gradient for gradient ascent (Angelo), and a first matrix characterizing a specifiable number k of greatest eigenvalues and corresponding eigenvectors of a covariance matrix of at least a subset of (Lahav) the plurality of training time series. (Liessman)”
From this reasoning, the rejections of amended independent claims 16, 27-29 and associated dependent claims 17-20, and 22-26 under 35 U.S.C. 103 stand in the current office action.
CONCLUSION
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Matthew Alan Cady whose telephone number is (571) 272-7229. The examiner can normally be reached Monday - Friday, 7:30 am - 5:00 pm ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Cesar Paula can be reached on (571)272-4128. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC)
at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MATTHEW ALAN CADY/ Examiner, Art Unit 2145
/CESAR B PAULA/Supervisory Patent Examiner, Art Unit 2145