Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
This action is in response to the application filed 30 May 2024. Claims 1-20 are pending and have been examined.
Claim Objections
Applicant is advised that should Claim 9 be found allowable, Claim 15 will be objected to under 37 CFR 1.75 as being a substantial duplicate thereof. When two claims in an application are duplicates or else are so close in content that they both cover the same thing, despite a slight difference in wording, it is proper after allowing one claim to object to the other as being a substantial duplicate of the allowed claim. See MPEP § 608.01(m).
Claim 15 recites "fitting at least one of a parametric distribution or a conditionally informed probability confidence estimation to a histogram ..." (emphasis added) while Claim 9 recites "fitting a parametric distribution to a histogram ...." As Claim 15 recites a non-required element, Claim 15 is a duplicate of Claim 9 when Claim 15 is interpreted to recite the step of fitting only a parametric distribution to a histogram.
Claim 11 is objected to because of the following informalities. Claim 11 recites "wherein a confident that the probabilities are corrected is determined according to Bayes' Rule" (emphasis added). In light of the instant specification at [0084] and [0085], for the purposes of examination, Examiner has interpreted the claim to read: "wherein a confidence that the probabilities are correct is determined according to Bayes' Rule." Appropriate confirmation or correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1, 9, 12, 15, and 20, and dependent Claims 2-8, 10-14, and 16-20, are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 1 recites the limitation "splitting a validation data set from the training data set, the validation data set and the training data set being mutually exclusive" (emphasis added). Later, Claim 1 recites "one truth datum contained within the training data set." There is insufficient antecedent basis for the later recitation of the element "the training data set," as "training data set" appears to have been re-defined in the previously recited limitation. Claims 9 and 15 are rejected according to the same rationale. Appropriate confirmation or correction is required.
Claim 1 recites the limitation "using the probabilities to provide a human or a machine decision maker with the likelihood information" (emphasis added). Claim 1 has previously recited probabilities in the prior separate steps of "estimating probabilities" and "computing probabilities." It is unclear whether one or both of the estimated and computed probabilities are referenced by the step of using the probabilities. Thus there is insufficient antecedent basis for "the probabilities" in the step of "using the probabilities." Claims 9 and 15 are rejected according to the same rationale. Appropriate confirmation or correction is required.
Claim 9 recites the limitation "computing distribution parameters from the parametric function." Claim 9 has previously recited "fitting a parametric distribution" but has not previously recited a parametric function. Thus there is insufficient antecedent basis for this limitation in the claim. Claim 15 is rejected according to the same rationale. For the purposes of examination, the "parametric function" of Claims 9 and 15 are interpreted to read "parametric distribution." Appropriate correction is required.
Claim 12 recites the limitation "wherein the parametric function is determined for each combination." Claim 9, upon which Claim 12 depends, has previously recited "fitting a parametric distribution" but has not previously recited a parametric function, nor has Claim 10, upon which Claim 12 depends. Thus there is insufficient antecedent basis for this limitation in the claim. For the purposes of examination, the "parametric function" of Claim 12 is interpreted to read "parametric distribution." Appropriate correction is required.
The term "statistically significant" in Claim 20 is a relative term which renders the claim indefinite. The term "statistically significant" is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. The limitation "computing the estimated probabilities of a ... number of samples" in the claim has been rendered indefinite by the use of the term "statistically significant." For the purposes of examination, the element "a statistically significant number of samples" has been interpreted to read "a number of samples." Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to and abstract idea without significantly more.
Regarding Claim 1
Step 1
Claim 1 recites a method of estimating the confidence in a neural network, and thus the claimed process falls within a statutory category of invention.
Step 2A Prong 1
The claim recites defining a problem to be solved using a neural network, which is a mental process step of defining that recites an intended use of a neural network with no patentable weight. The claim recites splitting the data into a training data set and a test data set, the test data set and the training data set being mutually exclusive, which is a mental process. The claim recites splitting a validation data set from the training data set, the validation data set and the training data set being mutually exclusive, which is a mental process. The claim recites to minimize plural aggregate differences between at least one prediction from the neural network and at least one truth datum contained within the training data set, which is a mental process. The claim recites determining a decision plurality of decision vectors and a weight plurality of weight vectors, which is a mental process. The claim recites pairing individual decision vectors from the decision plurality of decision vectors with corresponding individual weight vectors from the weight plurality of weight vectors, which is a mental process. The claim recites computing angle distributions between the individual decision vectors and the corresponding individual weight vectors, which is a mental process and/or a mathematical calculation. The claim recites computing a combination of labelled class parameters and predicted class parameters from the angle distributions, which is a mental process. The claim recites fitting a parametric function to a histogram of the angle distributions for each combination of labeled class parameters and predicted class parameters, which is a mental process and/or a mathematical calculation. The claim recites computing distribution parameters from the parametric function, which is a mental process and/or a mathematical calculation. The claim recites estimating probabilities from the distribution parameters that the neural network predictions are correct, which is a mental process. The claim recites computing probabilities that the values from the distribution parameters are correct, which is a mental process and/or a mathematical calculation.
Thus, the claim recites an abstract idea.
Step 2A Prong 2
The additional element providing data to be input to the neural network amounts to insignificant extra-solution activity (see MPEP 2106.05(g), "mere data gathering or outputting"). The additional element training the neural network to have weight parameters and bias parameters invokes a computer or other machinery merely as a tool to perform an existing process (see MPEP 2106.05(f), "apply it"). The additional element using the probabilities to provide a human or a machine decision maker with the likelihood information about the neural network's prediction needed to make a risk informed decision amounts to insignificant extra-solution activity (see MPEP 2106.05(g), "mere data gathering or outputting," as the claim recites no positive step of decision making).
Step 2B
The additional element providing data to be input to the neural network is well-understood, routine, conventional activity (see MPEP 2106.05(d), "receiving or transmitting data over a network"). The additional element training the neural network to have weight parameters and bias parameters invokes a computer or other machinery merely as a tool to perform an existing process (see MPEP 2106.05(f), "apply it"). The additional element using the probabilities to provide a human or a machine decision maker with the likelihood information about the neural network's prediction needed to make a risk informed decision is well-understood, routine, conventional activity (see MPEP 2106.05(d), "receiving or transmitting data over a network," as the claim recites no positive step of decision making).
The claim lacks additional elements that integrate it into a practical application or provide significantly more, so it is directed to an abstract idea and is ineligible.
Regarding Claim 2
Step 1
Regarding Claim 2, the rejection of Claim 1 is incorporated.
Step 2A Prong 1
The claim recites splitting the data into a training data set and a test data set, the test data set and the training data set being mutually exclusive (as recited by Claim 1), wherein the step of splitting the data into a training data set and a test data set comprises randomly splitting the data which is a mental process.
Thus, the claim recites an abstract idea.
Step 2A Prong 2, Step 2B
The claim lacks additional elements that integrate it into a practical application or provide significantly more, so it is directed to an abstract idea and is ineligible.
Regarding Claim 3
Step 1
Regarding Claim 3, the rejection of Claim 1 is incorporated.
Step 2A Prong 1
The claim recites sequestering the test data set while splitting the data into the training data set and the test data set, which is a mental process.
Thus, the claim recites an abstract idea.
Step 2A Prong 2, Step 2B
The claim lacks additional elements that integrate it into a practical application or provide significantly more, so it is directed to an abstract idea and is ineligible.
Regarding Claim 4
Step 1
Regarding Claim 4, the rejection of Claim 1 is incorporated.
Step 2A Prong 1
The claim recites determining a decision plurality of decision vectors and a weight plurality of weight vectors (as recited by Claim 1), wherein the step of determining the decision plurality of decision vectors and the weight plurality of weight vectors comprises making the determinations from data in the validation data set which is a mental process.
Thus, the claim recites an abstract idea.
Step 2A Prong 2, Step 2B
The claim lacks additional elements that integrate it into a practical application or provide significantly more, so it is directed to an abstract idea and is ineligible.
Regarding Claim 5
Step 1
Regarding Claim 5, the rejection of Claim 1 is incorporated.
Step 2A Prong 1
Claim 5 recites the abstract ideas recited by Claim 1.
Step 2A Prong 2
The additional element the predicted class parameters are saved in a data structure for later reference to estimate probabilities of data being true amounts to insignificant extra-solution activity (see MPEP 2106.05(g), "mere data gathering").
Step 2B
The additional element the predicted class parameters are saved in a data structure for later reference to estimate probabilities of data being true is well-understood, routine, conventional activity (see MPEP 2106.05(d), "storing and retrieving information in memory").
The claim lacks additional elements that integrate it into a practical application or provide significantly more, so it is directed to an abstract idea and is ineligible.
Regarding Claim 6
Step 1
Regarding Claim 6, the rejection of Claim 1 is incorporated.
Step 2A Prong 1
The claim recites defining a problem to be solved using a neural network (as recited by Claim 1), wherein the problem is defined using a neural network classifier model, which is a mental process step of defining that recites an intended use of a neural network classifier with no patentable weight.
Thus, the claim recites an abstract idea.
Step 2A Prong 2, Step 2B
The claim lacks additional elements that integrate it into a practical application or provide significantly more, so it is directed to an abstract idea and is ineligible.
Regarding Claim 7
Step 1
Regarding Claim 7, the rejection of Claim 6 is incorporated.
Step 2A Prong 1
The claim recites defining a problem to be solved using a neural network (as recited by Claim 1), wherein the problem is defined using a neural network classifier model (as recited by Claim 6), wherein the neural network classifier model defines a type of data input and a plurality of classes of the data input, which is a mental process.
Thus, the claim recites an abstract idea.
Step 2A Prong 2, Step 2B
The claim lacks additional elements that integrate it into a practical application or provide significantly more, so it is directed to an abstract idea and is ineligible.
Regarding Claim 8
Step 1
Regarding Claim 8, the rejection of Claim 7 is incorporated.
Step 2A Prong 1
Claim 8 recites the abstract ideas recited by Claim 7.
Step 2A Prong 2, Step 2B
The additional element wherein the type of data comprises visual images does not amount to more than generally linking the use of a judicial exception to a particular field of use (see MPEP 2106.05(h), "limit the use of the abstract idea to a particular technological environment").
The claim lacks additional elements that integrate it into a practical application or provide significantly more, so it is directed to an abstract idea and is ineligible.
Regarding Claim 9
Step 1
Claim 9 recites a method of estimating the probabilities of correct values from a neural network, and thus the claimed process falls within a statutory category of invention.
Step 2A Prong 1
The claim recites defining a problem to be solved using a neural network, which is a mental process. The claim recites splitting the data into a training data set and a test data set, the test data set and the training data set being mutually exclusive, which is a mental process. The claim recites splitting a validation data set from the training data set, the validation data set and the training data set being mutually exclusive, which is a mental process. The claim recites to minimize plural aggregate differences between at least one prediction from the neural network and at least one truth datum contained within the training data set, which is a mental process. The claim recites determining a decision plurality of decision vectors and a weight plurality of weight vectors, which is a mental process. The claim recites pairing individual decision vectors from the decision plurality of decision vectors with corresponding individual weight vectors from the weight plurality of weight vectors, which is a mental process. The claim recites computing angle distributions between the individual decision vectors and the corresponding individual weight vectors, which is a mental process and/or a mathematical calculation. The claim recites computing a combination of labelled class parameters and predicted class parameters from the angle distributions, which is a mental process. The claim recites fitting a parametric function to a histogram of the angle distributions for each combination of labeled class parameters and predicted class parameters, which is a mental process and/or a mathematical calculation. The claim recites computing distribution parameters from the parametric function, which is a mental process and/or a mathematical calculation. The claim recites estimating probabilities from the distribution parameters that the neural network predictions are correct, which is a mental process. The claim recites computing probabilities that the values from the distribution parameters are correct, which is a mental process and/or a mathematical calculation.
Thus, the claim recites an abstract idea.
Step 2A Prong 2
The additional element providing data to be input to the neural network amounts to insignificant extra-solution activity (see MPEP 2106.05(g), "mere data gathering or outputting"). The additional element training the neural network to have weight parameters and bias parameters invokes a computer or other machinery merely as a tool to perform an existing process (see MPEP 2106.05(f), "apply it"). The additional element using the probabilities to provide a human or a machine decision maker with the likelihood information about the neural network's prediction needed to make a risk informed decision amounts to insignificant extra-solution activity (see MPEP 2106.05(g), "mere data gathering or outputting," as the claim recites no positive step of decision making).
Step 2B
The additional element providing data to be input to the neural network is well-understood, routine, conventional activity (see MPEP 2106.05(d), "receiving or transmitting data over a network"). The additional element training the neural network to have weight parameters and bias parameters invokes a computer or other machinery merely as a tool to perform an existing process (see MPEP 2106.05(f), "apply it"). The additional element using the probabilities to provide a human or a machine decision maker with the likelihood information about the neural network's prediction needed to make a risk informed decision is well-understood, routine, conventional activity (see MPEP 2106.05(d), "receiving or transmitting data over a network," as the claim recites no positive step of decision making).
The claim lacks additional elements that integrate it into a practical application or provide significantly more, so it is directed to an abstract idea and is ineligible.
Regarding Claim 10
Step 1
Regarding Claim 10, the rejection of Claim 9 is incorporated.
Step 2A Prong 1
The claim recites computing angle distributions between the individual decision vectors and the corresponding individual weight vectors (as recited by Claim 9), wherein the distribution of angles between the decision vectors and the weight vectors are modeled using a Gaussian distribution, which is a mathematical calculation.
Thus, the claim recites an abstract idea.
Step 2A Prong 2, Step 2B
The claim lacks additional elements that integrate it into a practical application or provide significantly more, so it is directed to an abstract idea and is ineligible.
Regarding Claim 11
Step 1
Regarding Claim 11, the rejection of Claim 10 is incorporated.
Step 2A Prong 1
The claim recites estimating probabilities from the distribution parameters that the neural network predictions are correct (as recited by Claim 9), wherein a confident that the probabilities are corrected is determined according to Bayes' Rule, which is a mathematical calculation.
Thus, the claim recites an abstract idea.
Step 2A Prong 2, Step 2B
The claim lacks additional elements that integrate it into a practical application or provide significantly more, so it is directed to an abstract idea and is ineligible.
Regarding Claim 12
Step 1
Regarding Claim 12, the rejection of Claim 10 is incorporated.
Step 2A Prong 1
The claim recites fitting a parametric function to a histogram of the angle distributions for each combination of labeled class parameters and predicted class parameters (as recited by Claim 9), wherein the parametric function is determined for each combination of labeled class parameters and predicted class parameters, which is a mental process.
Thus, the claim recites an abstract idea.
Step 2A Prong 2, Step 2B
The claim lacks additional elements that integrate it into a practical application or provide significantly more, so it is directed to an abstract idea and is ineligible.
Regarding Claim 13
Step 1
Regarding Claim 13, the rejection of Claim 9 is incorporated.
Step 2A Prong 1
The claim recites to minimize plural aggregate differences between at least one prediction from the neural network and at least one truth datum contained within the training data set (as recited by Claim 9), wherein the plural aggregate differences are expressed as a cross-entropy loss, which is a mental process.
Thus, the claim recites an abstract idea.
Step 2A Prong 2, Step 2B
The claim lacks additional elements that integrate it into a practical application or provide significantly more, so it is directed to an abstract idea and is ineligible.
Claim 14, dependent on Claim 9, incorporates the rejection of Claim 9. Claim 14 incorporates substantively all the limitations of Claim 8 in method form and is rejected under the same rationales.
Regarding Claim 15
Step 1
Claim 15 recites a method of estimating the probabilities of correct values from a neural network, and thus the claimed process falls within a statutory category of invention.
Step 2A Prong 1
The claim recites defining a problem to be solved using a neural network, which is a mental process. The claim recites splitting the data into a training data set and a test data set, the test data set and the training data set being mutually exclusive, which is a mental process. The claim recites splitting a validation data set from the training data set, the validation data set and the training data set being mutually exclusive, which is a mental process. The claim recites to minimize plural aggregate differences between at least one prediction from the neural network and at least one truth datum contained within the training data set, which is a mental process. The claim recites determining a decision plurality of decision vectors and a weight plurality of weight vectors, which is a mental process. The claim recites pairing individual decision vectors from the decision plurality of decision vectors with corresponding individual weight vectors from the weight plurality of weight vectors, which is a mental process. The claim recites computing angle distributions between the individual decision vectors and the corresponding individual weight vectors, which is a mental process and/or a mathematical calculation. The claim recites computing a combination of labelled class parameters and predicted class parameters from the angle distributions, which is a mental process. The claim recites fitting at least one of a parametric distribution or a conditionally informed probability confidence estimation to a histogram of the angle distributions for each combination of labeled class parameters and predicted class parameters, which is a mental process and/or a mathematical calculation. The claim recites computing distribution parameters from the parametric function, which is a mental process and/or a mathematical calculation. The claim recites estimating probabilities from the distribution parameters that the neural network predictions are correct, which is a mental process. The claim recites computing probabilities that the values from the distribution parameters are correct, which is a mental process and/or a mathematical calculation.
Thus, the claim recites an abstract idea.
Step 2A Prong 2
The additional element providing data to be input to the neural network amounts to insignificant extra-solution activity (see MPEP 2106.05(g), "mere data gathering or outputting"). The additional element training the neural network to have weight parameters and bias parameters invokes a computer or other machinery merely as a tool to perform an existing process (see MPEP 2106.05(f), "apply it"). The additional element using the probabilities to provide a human or a machine decision maker with the likelihood information about the neural network's prediction needed to make a risk informed decision amounts to insignificant extra-solution activity (see MPEP 2106.05(g), "mere data gathering or outputting," as the claim recites no positive step of decision making).
Step 2B
The additional element providing data to be input to the neural network is well-understood, routine, conventional activity (see MPEP 2106.05(d), "receiving or transmitting data over a network"). The additional element training the neural network to have weight parameters and bias parameters invokes a computer or other machinery merely as a tool to perform an existing process (see MPEP 2106.05(f), "apply it"). The additional element using the probabilities to provide a human or a machine decision maker with the likelihood information about the neural network's prediction needed to make a risk informed decision is well-understood, routine, conventional activity (see MPEP 2106.05(d), "receiving or transmitting data over a network," as the claim recites no positive step of decision making).
The claim lacks additional elements that integrate it into a practical application or provide significantly more, so it is directed to an abstract idea and is ineligible.
Regarding Claim 16
Step 1
Regarding Claim 16, the rejection of Claim 15 is incorporated.
Step 2A Prong 1
The claim recites computing a calibration error for at least one estimated probability, which is a mental process and/or a mathematical calculation.
Thus, the claim recites an abstract idea.
Step 2A Prong 2, Step 2B
The claim lacks additional elements that integrate it into a practical application or provide significantly more, so it is directed to an abstract idea and is ineligible.
Regarding Claim 17
Step 1
Regarding Claim 17, the rejection of Claim 16 is incorporated.
Step 2A Prong 1
The claim recites optimizing the weight parameters and bias parameters using a gradient descent, which is a mathematical calculation.
Thus, the claim recites an abstract idea.
Step 2A Prong 2, Step 2B
The claim lacks additional elements that integrate it into a practical application or provide significantly more, so it is directed to an abstract idea and is ineligible.
Regarding Claim 18
Step 1
Regarding Claim 18, the rejection of Claim 17 is incorporated.
Step 2A Prong 1
The claim recites optimizing the weight parameters and bias parameters using a gradient descent, (as recited by Claim 17), wherein the gradient descent is a stochastic gradient descent, which is a mathematical calculation.
Thus, the claim recites an abstract idea.
Step 2A Prong 2, Step 2B
The claim lacks additional elements that integrate it into a practical application or provide significantly more, so it is directed to an abstract idea and is ineligible.
Regarding Claim 19
Step 1
Regarding Claim 19, the rejection of Claim 18 is incorporated.
Step 2A Prong 1
Claim 19 recites the abstract ideas recited by Claim 18.
Step 2A Prong 2, Step 2B
The additional element training the neural network to have weight parameters and bias parameters (as recited by Claim 15), wherein the step of training the neural network is terminated when the aggregate differences for the validation data set reach a minimum value, in order to avoid overfitting the model invokes a computer or other machinery merely as a tool to perform an existing process (see MPEP 2106.05(f), "apply it").
The claim lacks additional elements that integrate it into a practical application or provide significantly more, so it is directed to an abstract idea and is ineligible.
Regarding Claim 20
Step 1
Regarding Claim 20, the rejection of Claim 9 is incorporated.
Step 2A Prong 1
The claim recites computing the estimated probabilities of a statistically significant number of samples, which is a mental process and/or a mathematical calculation.
Thus, the claim recites an abstract idea.
Step 2A Prong 2, Step 2B
The claim lacks additional elements that integrate it into a practical application or provide significantly more, so it is directed to an abstract idea and is ineligible.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1, 3-9, and 11-20 are rejected under 35 U.S.C. 103 as being unpatentable over Waagen, et al., "A Geometric Statistic for Deep Learning Model Confidence and Adversarial Defense" (hereinafter "Waagen") in view of Naeini, et al., "Obtaining Well Calibrated Probabilities Using Bayesian Binning" (hereinafter "Naeini").
Regarding Claim 1, Waagen teaches:
A method of estimating the confidence in a neural network (Waagen, p. 1, Abstract: " We propose a simple statistic, which we call Angular Margin, to characterize the 'confidence' of the model given a new input"), the method comprising the steps of:
defining a problem to be solved using a neural network (Waagen, p. 3, 4.1 Datasets and Models: "For the study of the
Δ
ϕ
statistic in DL classifiers under nominal conditions, we used two significantly different model architectures applied to different image classification problems");
providing data to be input to the neural network (Waagen, p. 3, 4.1 Datasets and Models: "One architecture is a very simple convolutional neural network (CNN) trained on the Fashion-MNIST dataset. The second is a residual neural network (ResNet) model trained for 10-class RGB-image classification using the CIFAR-10 dataset");
splitting the data into a training data set and a test data set, the test data set and the training data set being mutually exclusive (Waagen, p. 4, 4.1 Datasets and Models, Fashion-MNIST Trained Simple CNN: "Fashion-MNIST contains a 60,000 28x28 pixel image training set, with 6,000 images per class, and a separate 10,000 image test set, with 1,000 image per class. ... Before using them, we lightly pre-processed the images in the dataset by linearly rescaling their pixels from [0,255] integer values to [0,1] floating point values");
splitting a validation data set from the training data set, the validation data set and the training data set being mutually exclusive (Waagen, p. 4, 4.1 Datasets and Models: "The rescaled Fashion-MNIST training images were used for model fitting and the rescaled test images for model validation");
training the neural network to have weight parameters and bias parameters (Waagen, p. 2, 2. Geometry Of Deep Learning Classification: "The output layer
L
O
is the layer that defines the final mapping of the data to the set of class labels.... ¶ Given
c
classes, the output layer
L
O
consists of
c
neurons, with each neuron having a vector of input weights
w
∈
R
d
and a bias term
b
") to minimize plural aggregate differences between at least one prediction from the neural network and at least one truth datum contained within the training data set (Waagen, p. 4, 4.1 Datasets and Models: "The rescaled Fashion-MNIST training images were used for model fitting and the rescaled test images for model validation. For model fitting the loss function was categorical cross entropy, and empirical risk minimization was performed using the Adam16 variant of stochastic gradient descent (SGD). ... The training was run for 30 epochs and the model state with the highest accuracy on the validation data was selected," where accuracy is based on labeled test data, per p. 4: "Fashion-MNIST contains ... a separate 10,000 image test set, with 1,000 image per class," and where Waagen's per-class label corresponds to the instant truth datum, per p. 3, 3. Angular Statistics: "Note that to compute this statistic, one must have a priori knowledge of the correct class label of the image under evaluation. Therefore, its applicability is limited to model analysis with labeled data");
determining a decision plurality of decision vectors and a weight plurality of weight vectors (Waagen, p. 2, 2. Geometry Of Deep Learning Classification: "The output layer
L
O
is the layer that defines the final mapping of the data to the set of class labels, i.e.,
L
O
:
R
d
→
R
c
. This layer typically consists of c neurons (one neuron per class). Softmax can be viewed as the activation function of
L
O
or as a separate transformation
L
O
+
1
. Either way, it is usually applied across the
L
O
neurons. Let
X
O
-
1
denote the
d
-dimensional input vector space to
L
O
, which we have previously named the decision space (DS). ¶ Given
c
classes, the output layer
L
O
consists of
c
neurons, with each neuron having a vector of input weights
w
∈
R
d
and a bias term
b
," where Waagen's weights in
R
d
and class labels in
R
c
correspond to the instant weight and decision vectors, respectively);
pairing individual decision vectors from the decision plurality of decision vectors with corresponding individual weight vectors from the weight plurality of weight vectors (Waagen, p. 3, 3. Angular Statistics: "As previously referenced, Chen et al. define a geometric statistic called angular visual hardness, AVH. Given an image's decision space datum
x
i
and associated class label
y
, AVH is given by
A
V
H
≝
A
x
i
,
w
y
∑
j
A
x
i
,
w
j
=
ϕ
y
,
i
∑
j
ϕ
x
i
,
w
j
(7)
where
w
y
is the weight vector of the output neuron associated with the correct label
y
, and
A
x
i
,
w
j
is the angular distance eq. (3) between un-augmented vectors
x
i
and
w
j
," where Waagen's weight vector
w
y
for label
y
corresponds to the instant pairing);
computing angle distributions between the individual decision vectors and the corresponding individual weight vectors (Waagen, p. 7, 6.1 Original (Non-Adversarial) Image Results: "we wish to understand if there is a relationship between the
Δ
ϕ
values and image classification accuracy, that is, identify the distributions of
Δ
ϕ
values conditioned on being correctly (or incorrectly) classified. To this end, Figure 6(a) and (b) display the histograms of the
Δ
ϕ
values computed for the simple CNN and ResNet- 20 models' respective training (top plots) and validation (bottom plots) images," where Waagen's histograms of correctly and incorrectly classified images based on
Δ
ϕ
angle values correspond to the instant distribution);
computing a combination of labelled class parameters and predicted class parameters from the angle distributions (Waagen, p. 7, 6.1 Original (Non-Adversarial) Image Results: "we wish to understand if there is a relationship between the
Δ
ϕ
values and image classification accuracy, that is, identify the distributions of
Δ
ϕ
values conditioned on being correctly (or incorrectly) classified," where Waagen's angle distribution of correctly labeled images or angle distribution of incorrectly labeled images corresponds to the instant combination of labeled and predicted class parameters);
fitting a parametric function to a histogram of the angle distributions (Waagen, p. 6, 5. Quantifying Confidence: "To map the
Δ
ϕ
values to classification confidence levels for a given model and dataset, we developed a simple, greedy partitioning approach to group the
Δ
ϕ
values into 'statistically distinctive' confidence regions. ... Using the model's validation data, it bins the
Δ
ϕ
values of the correctly and incorrectly classified images into two separate histograms. The histograms are smoothed to reduce noise using the averaged shifted histogram (ASH) method. ¶ Next, it iteratively applies a null hypothesis test (NHT) of two-independent proportions to the smoothed number of correct classification / number of classifications made between adjacent bins. ... The final
Δ
ϕ
bins are the assigned
Δ
ϕ
regions and the estimated accuracies with the regions are the confidence estimates. The initial bin sizes, the ASH smoothing kernel size, and the p-value threshold are the algorithm's only hyperparameters. The results provide a coarse calibration of the
Δ
ϕ
statistic with model confidence," where Waagen's iterative algorithm corresponds to the instant fitting a function, and Waagen's algorithm's hyperparameters correspond to the instant parametric) for each combination of labeled class parameters and predicted class parameters (Waagen, p. 8, Figure 6, "Histograms of the
Δ
ϕ
statistic of the validation sets for (a) the simple CNN model and (b) the ResNet-20 model," depicting both the correct pairing and the incorrect pairing of predicted and actual labels);
computing distribution parameters from the parametric function (Waagen, p. 6, 5. Quantifying Confidence: "Using the model's validation data, it bins the
Δ
ϕ
values of the correctly and incorrectly classified images into two separate histograms. The histograms are smoothed to reduce noise using the averaged shifted histogram (ASH) method. ¶ Next, it iteratively applies a null hypothesis test (NHT) of two-independent proportions to the smoothed number of correct classification / number of classifications made between adjacent bins. ... The bin merging iteration is only stopped when a stability configuration is reached, i.e., no more bins can be merged with the given p-value threshold. The final
Δ
ϕ
bins are the assigned
Δ
ϕ
regions and the estimated accuracies with the regions are the confidence estimates," where Waagen's final bins correspond to the instant distribution parameters);
estimating probabilities from the distribution parameters that the neural network predictions are correct (Waagen, p. 9, Figure 7, "Estimated confidence regions using validation data for (a) the simple CNN model, and (b) the ResNet-20 model," depicting confidence regions corresponding to the confidence estimate values 0.538, 0.693, 0.826, etc., relating confidence to
Δ
ϕ
angle values, where the instant specification refers to confidence estimates as probabilities at [00109]: "providing improved confidence estimates (AKA 'probabilities' or 'likelihoods) in the classification decisions");
computing probabilities that the values from the distribution parameters are correct (Waagen, p. 9, Figure 7, "Estimated confidence regions using validation data for (a) the simple CNN model, and (b) the ResNet-20 model," depicting confidence regions corresponding to the confidence estimate values 0.538, 0.693, 0.826, etc., relating confidence to
Δ
ϕ
angle values, where the instant specification refers to confidence estimates as probabilities at [00109]: "providing improved confidence estimates (AKA 'probabilities' or 'likelihoods) in the classification decisions"); and
using the probabilities to provide ... likelihood information about the neural network's prediction ... (Waagen, p. 6, 5. Quantifying Confidence: "To map the
Δ
ϕ
values to classification confidence levels for a given model and dataset, we developed a simple, greedy partitioning approach to group the
Δ
ϕ
values into 'statistically distinctive' confidence regions. ... ¶ ... The results provide a coarse calibration of the
Δ
ϕ
statistic with model confidence," where Waagen's calibrated confidence corresponds to the instant provided likelihood in light of the specification at [00109]: "providing improved confidence estimates (AKA 'probabilities' or 'likelihoods) in the classification decisions made by neural networks").
Waagen teaches using probabilities to provide likelihood information about the neural network's prediction.
Waagen does not explicitly teach provide a human or a machine decision maker with the likelihood information ... needed to make a risk informed decision.
However, Naeini teaches:
provide a human or a machine decision maker with the likelihood information ... needed to make a risk informed decision (Naeini, p. 2901, Abstract: "Learning probabilistic predictive models that are well calibrated is critical for many prediction and decision-making tasks in artificial intelligence. In this paper we present a new non-parametric calibration method called Bayesian Binning into Quantiles (BBQ)" and p. 2901, Introduction: "Producing well-calibrated probabilistic predictions is critical in many areas of science (e.g., determining which experiments to perform), medicine (e.g., deciding which therapy to give a patient), business (e.g., making investment decisions), and others. ... An ... approach is to construct well-calibrated models by relying on the existing machine learning methods and by modifying their outputs in a post-processing step to obtain the desired model," where Naeini's calibrated confidence corresponds to the instant provided likelihood in light of the specification at [00109]: "providing improved confidence estimates (AKA 'probabilities' or 'likelihoods) in the classification decisions made by neural networks").
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Waagen regarding estimating probabilities from the distribution parameters that the neural network predictions are correct with those of Naeini regarding provide a human or a machine decision maker with the likelihood information needed to make a risk informed decision.
The motivation to do so would be to facilitate providing a model with well calibrated confidence without requiring additional effort during the model training stage (Naeini, p. 2901, Introduction: "An ... approach is to construct well-calibrated models by relying on the existing machine learning methods and by modifying their outputs in a post-processing step to obtain the desired model. This approach is often preferred because of its generality, flexibility, and the fact that it frees the designer of the machine learning model from the need to add additional calibration measures into the objective function used to learn the model").
Regarding Claim 9, Waagen teaches:
A method of estimating the probabilities of correct values from a neural network (Waagen, p. 1, Abstract: " We propose a simple statistic, which we call Angular Margin, to characterize the 'confidence' of the model given a new input," where Waagen's estimated confidence corresponds to the instant probability of correct value per the instant specification at [00109]: 'providing improved confidence estimates (AKA 'probabilities' or 'likelihoods) in the classification decisions made by neural networks'), the method comprising the steps of:
defining a problem to be solved using a neural network (Waagen, p. 3, 4.1 Datasets and Models: "For the study of the
Δ
ϕ
statistic in DL classifiers under nominal conditions, we used two significantly different model architectures applied to different image classification problems");
providing data to be input to the neural network (Waagen, p. 3, 4.1 Datasets and Models: "One architecture is a very simple convolutional neural network (CNN) trained on the Fashion-MNIST dataset. The second is a residual neural network (ResNet) model trained for 10-class RGB-image classification using the CIFAR-10 dataset");
splitting the data into a training data set and a test data set, the test data set and the training data set being mutually exclusive (Waagen, p. 4, 4.1 Datasets and Models, Fashion-MNIST Trained Simple CNN: "Fashion-MNIST contains a 60,000 28x28 pixel image training set, with 6,000 images per class, and a separate 10,000 image test set, with 1,000 image per class. ... Before using them, we lightly pre-processed the images in the dataset by linearly rescaling their pixels from [0,255] integer values to [0,1] floating point values");
splitting a validation data set from the training data set, the validation data set and the training data set being mutually exclusive (Waagen, p. 4, 4.1 Datasets and Models: "The rescaled Fashion-MNIST training images were used for model fitting and the rescaled test images for model validation");
training the neural network to have weight parameters and bias parameters (Waagen, p. 2, 2. Geometry Of Deep Learning Classification: "The output layer
L
O
is the layer that defines the final mapping of the data to the set of class labels.... ¶ Given
c
classes, the output layer
L
O
consists of
c
neurons, with each neuron having a vector of input weights
w
∈
R
d
and a bias term
b
") to minimize plural aggregate differences between at least one prediction from the neural network and at least one truth datum contained within the training data set (Waagen, p. 4, 4.1 Datasets and Models: "The rescaled Fashion-MNIST training images were used for model fitting and the rescaled test images for model validation. For model fitting the loss function was categorical cross entropy, and empirical risk minimization was performed using the Adam16 variant of stochastic gradient descent (SGD). ... The training was run for 30 epochs and the model state with the highest accuracy on the validation data was selected," where accuracy is based on labeled test data, per p. 4: "Fashion-MNIST contains ... a separate 10,000 image test set, with 1,000 image per class," and where Waagen's per-class label corresponds to the instant truth datum, per p. 3, 3. Angular Statistics: "Note that to compute this statistic, one must have a priori knowledge of the correct class label of the image under evaluation. Therefore, its applicability is limited to model analysis with labeled data");
determining a decision plurality of decision vectors and a weight plurality of weight vectors
(Waagen, p. 2, 2. Geometry Of Deep Learning Classification: "The output layer
L
O
is the layer that defines the final mapping of the data to the set of class labels, i.e.,
L
O
:
R
d
→
R
c
. This layer typically consists of c neurons (one neuron per class). Softmax can be viewed as the activation function of
L
O
or as a separate transformation
L
O
+
1
. Either way, it is usually applied across the
L
O
neurons. Let
X
O
-
1
denote the
d
-dimensional input vector space to
L
O
, which we have previously named the decision space (DS). ¶ Given
c
classes, the output layer
L
O
consists of
c
neurons, with each neuron having a vector of input weights
w
∈
R
d
and a bias term
b
," where Waagen's weights in
R
d
and class labels in
R
c
correspond to the instant weight and decision vectors, respectively);
pairing individual decision vectors from the decision plurality of decision vectors with corresponding individual weight vectors from the weight plurality of weight vectors (Waagen, p. 3, 3. Angular Statistics: "As previously referenced, Chen et al. define a geometric statistic called angular visual hardness, AVH. Given an image's decision space datum
x
i
and associated class label
y
, AVH is given by
A
V
H
≝
A
x
i
,
w
y
∑
j
A
x
i
,
w
j
=
ϕ
y
,
i
∑
j
ϕ
x
i
,
w
j
(7)
where
w
y
is the weight vector of the output neuron associated with the correct label
y
, and
A
x
i
,
w
j
is the angular distance eq. (3) between un-augmented vectors
x
i
and
w
j
," where Waagen's weight vector
w
y
for label
y
corresponds to the instant pairing);
computing angle distributions between the individual decision vectors and the corresponding individual weight vectors (Waagen, p. 7, 6.1 Original (Non-Adversarial) Image Results: "we wish to understand if there is a relationship between the
Δ
ϕ
values and image classification accuracy, that is, identify the distributions of
Δ
ϕ
values conditioned on being correctly (or incorrectly) classified. To this end, Figure 6(a) and (b) display the histograms of the
Δ
ϕ
values computed for the simple CNN and ResNet- 20 models' respective training (top plots) and validation (bottom plots) images," where Waagen's histograms of correctly and incorrectly classified images based on
Δ
ϕ
angle values correspond to the instant distribution);
computing a combination of labelled class parameters and predicted class parameters from the angle distributions (Waagen, p. 7, 6.1 Original (Non-Adversarial) Image Results: "we wish to understand if there is a relationship between the
Δ
ϕ
values and image classification accuracy, that is, identify the distributions of
Δ
ϕ
values conditioned on being correctly (or incorrectly) classified," where Waagen's angle distribution of correctly labeled images or angle distribution of incorrectly labeled images corresponds to the instant combination of labeled and predicted class parameters);
fitting a parametric distribution to a histogram of the angle distributions (Waagen, p. 6, 5. Quantifying Confidence: "To map the
Δ
ϕ
values to classification confidence levels for a given model and dataset, we developed a simple, greedy partitioning approach to group the
Δ
ϕ
values into 'statistically distinctive' confidence regions. ... Using the model's validation data, it bins the
Δ
ϕ
values of the correctly and incorrectly classified images into two separate histograms. The histograms are smoothed to reduce noise using the averaged shifted histogram (ASH) method. ¶ Next, it iteratively applies a null hypothesis test (NHT) of two-independent proportions to the smoothed number of correct classification / number of classifications made between adjacent bins. ... The final
Δ
ϕ
bins are the assigned
Δ
ϕ
regions and the estimated accuracies with the regions are the confidence estimates. The initial bin sizes, the ASH smoothing kernel size, and the p-value threshold are the algorithm's only hyperparameters. The results provide a coarse calibration of the
Δ
ϕ
statistic with model confidence," where Waagen's iterative algorithm corresponds to the instant fitting, and Waagen's algorithm's hyperparameters correspond to the instant parametric) for each combination of labeled class parameters and predicted class parameters (Waagen, p. 8, Figure 6, "Histograms of the
Δ
ϕ
statistic of the validation sets for (a) the simple CNN model and (b) the ResNet-20 model," depicting both the correct pairing and the incorrect pairing of predicted and actual labels);
computing distribution parameters from the parametric function (Waagen, p. 6, 5. Quantifying Confidence: "Using the model's validation data, it bins the
Δ
ϕ
values of the correctly and incorrectly classified images into two separate histograms. The histograms are smoothed to reduce noise using the averaged shifted histogram (ASH) method. ¶ Next, it iteratively applies a null hypothesis test (NHT) of two-independent proportions to the smoothed number of correct classification / number of classifications made between adjacent bins. ... The bin merging iteration is only stopped when a stability configuration is reached, i.e., no more bins can be merged with the given p-value threshold. The final
Δ
ϕ
bins are the assigned
Δ
ϕ
regions and the estimated accuracies with the regions are the confidence estimates," where Waagen's final bins correspond to the instant distribution parameters);
estimating probabilities from the distribution parameters that the neural network predictions are correct (Waagen, p. 9, Figure 7, "Estimated confidence regions using validation data for (a) the simple CNN model, and (b) the ResNet-20 model," depicting confidence regions corresponding to the confidence estimate values 0.538, 0.693, 0.826, etc., relating confidence to
Δ
ϕ
angle values, where the instant specification refers to confidence estimates as probabilities at [00109]: "providing improved confidence estimates (AKA 'probabilities' or 'likelihoods) in the classification decisions");
computing probabilities that the values from the distribution parameters are correct (Waagen, p. 9, Figure 7, "Estimated confidence regions using validation data for (a) the simple CNN model, and (b) the ResNet-20 model," depicting confidence regions corresponding to the confidence estimate values 0.538, 0.693, 0.826, etc., relating confidence to
Δ
ϕ
angle values, where the instant specification refers to confidence estimates as probabilities at [00109]: "providing improved confidence estimates (AKA 'probabilities' or 'likelihoods) in the classification decisions"); and
using the probabilities to provide ... likelihood information about the neural network's prediction (Waagen, p. 6, 5. Quantifying Confidence: "To map the
Δ
ϕ
values to classification confidence levels for a given model and dataset, we developed a simple, greedy partitioning approach to group the
Δ
ϕ
values into 'statistically distinctive' confidence regions. ... ¶ ... The results provide a coarse calibration of the
Δ
ϕ
statistic with model confidence," where Waagen's calibrated confidence corresponds to the instant provided likelihood in light of the specification at [00109]: "providing improved confidence estimates (AKA 'probabilities' or 'likelihoods) in the classification decisions made by neural networks").
Waagen teaches using probabilities to provide likelihood information about the neural network's prediction.
Waagen does not explicitly teach provide a human or a machine decision maker with the likelihood information ... needed to make a risk informed decision.
However, Naeini teaches:
provide a human or a machine decision maker with the likelihood information ... needed to make a risk informed decision (Naeini, p. 2901, Abstract: "Learning probabilistic predictive models that are well calibrated is critical for many prediction and decision-making tasks in artificial intelligence. In this paper we present a new non-parametric calibration method called Bayesian Binning into Quantiles (BBQ)" and p. 2901, Introduction: "Producing well-calibrated probabilistic predictions is critical in many areas of science (e.g., determining which experiments to perform), medicine (e.g., deciding which therapy to give a patient), business (e.g., making investment decisions), and others. ... An ... approach is to construct well-calibrated models by relying on the existing machine learning methods and by modifying their outputs in a post-processing step to obtain the desired model," where Naeini's calibrated confidence corresponds to the instant provided likelihood in light of the specification at [00109]: "providing improved confidence estimates (AKA 'probabilities' or 'likelihoods) in the classification decisions made by neural networks").
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Waagen regarding estimating probabilities from the distribution parameters that the neural network predictions are correct with those of Naeini regarding provide a human or a machine decision maker with the likelihood information needed to make a risk informed decision.
The motivation to do so would be to facilitate providing a model with well calibrated confidence without requiring additional effort during the model training stage (Naeini, p. 2901, Introduction: "An ... approach is to construct well-calibrated models by relying on the existing machine learning methods and by modifying their outputs in a post-processing step to obtain the desired model. This approach is often preferred because of its generality, flexibility, and the fact that it frees the designer of the machine learning model from the need to add additional calibration measures into the objective function used to learn the model").
Regarding Claim 15, Waagen teaches:
A method of estimating the probabilities of correct values from a neural network (Waagen, p. 1, Abstract: " We propose a simple statistic, which we call Angular Margin, to characterize the 'confidence' of the model given a new input," where Waagen's estimated confidence corresponds to the instant probability of correct value per the instant specification at [00109]: 'providing improved confidence estimates (AKA 'probabilities' or 'likelihoods) in the classification decisions made by neural networks'), the method comprising the steps recited in the rejection of Claim 9. Claim 15 is rejected under the same rationale as Claim 9.
Claim 15 recites "fitting at least one of a parametric distribution or a conditionally informed probability confidence estimation to a histogram," which contains the non-essential element "a conditionally informed probability confidence estimation." This limitation of Claim 15, when interpreted not to recite the non-essential element, is obvious in light of Waagen for the reason given in the rejection of the corresponding limitation of Claim 9.
Regarding Claim 3, the rejection of Claim 1 is incorporated. The Waagen/Naeini combination teaches:
further comprising the step of sequestering the test data set while splitting the data into the training data set and the test data set (Waagen, p. 4, 4. Experiments, 4.1 Datasets and Models: "Fashion-MNIST contains a 60,000 28x28 pixel image training set, with 6,000 images per class, and a separate 10,000 image test set, with 1,000 image per class" and "The rescaled Fashion-MNIST training images were used for model fitting and the rescaled test images for model validation," where Waagen's training and test images being used for different phases corresponds to the instant sequestering, per the instant specification, [00102]: "The test data set 45 is sequestered until it is time to evaluate model performance").
Regarding Claim 4, the rejection of Claim 1 is incorporated. The Waagen/Naeini combination teaches:
wherein the step of determining the decision plurality of decision vectors and the weight plurality of weight vectors comprises making the determinations from data in the validation data set (Waagen, p. 6, 5. Quantifying Confidence: "To map the
Δ
ϕ
values to classification confidence levels for a given model and dataset, we developed a simple, greedy partitioning approach to group the
Δ
ϕ
values into 'statistically distinctive' confidence regions. ... Using the model's validation data, it bins the
Δ
ϕ
values of the correctly and incorrectly classified images into two separate histograms").
Regarding Claim 5, the rejection of Claim 1 is incorporated. The Waagen/Naeini combination teaches:
wherein the predicted class parameters are saved in a data structure for later reference to estimate probabilities of data being true (Waagen, p. 6, 5. Quantifying Confidence: "Once all adjacent bins are either left alone or merged, the process is repeated with the updated bins. The bin merging iteration is only stopped when a stability configuration is reached, i.e., no more bins can be merged with the given p-value threshold. The final
Δ
ϕ
bins are the assigned
Δ
ϕ
regions and the estimated accuracies with the regions are the confidence estimates," where Waagen's iteratively merged bins correspond to the instant data structure, and where Waagen's confidence estimates correspond to the instant probabilities, per the specification at [00109]: "providing improved confidence estimates (AKA 'probabilities' or 'likelihoods) in the classification decisions made by neural networks").
Regarding Claim 6, the rejection of Claim 1 is incorporated. The Waagen/Naeini combination teaches:
wherein the problem is defined using a neural network classifier model (Waagen, p. 3, 4.1 Datasets and Models: "For the study of the
Δ
ϕ
statistic in DL [deep learning] classifiers under nominal conditions, we used two significantly different model architectures applied to different image classification problems").
Regarding Claim 7, the rejection of Claim 6 is incorporated. The Waagen/Naeini combination teaches:
wherein the neural network classifier model defines a type of data input and a plurality of classes of the data input (Waagen, p. 3, 4. Experiments, 4.1 Datasets and Models, Fashion-MNIST Trained Simple CNN: "In our first experiments, we selected the Fashion-MNIST classification problem. The Fashion-MNIST dataset is a collection of grayscale images of stylized clothing and other fashion items for ten classes (t-shirt/top, trouser, pullover, dress,...)").
Regarding Claim 8, the rejection of Claim 7 is incorporated. The Waagen/Naeini combination teaches:
wherein the type of data comprises visual images (Waagen, p. 3, 4. Experiments, 4.1 Datasets and Models, Fashion-MNIST Trained Simple CNN: "In our first experiments, we selected the Fashion-MNIST classification problem. The Fashion-MNIST dataset is a collection of grayscale images of stylized clothing and other fashion items for ten classes (t-shirt/top, trouser, pullover, dress,...)").
Regarding Claim 11, the rejection of Claim 10 is incorporated. Naeini further teaches:
wherein a confident that the probabilities are corrected is determined according to Bayes' Rule (Naeini, p. 2902, Methods: "the marginal likelihood can be expressed as ...
P
D
M
... The term
P
M
in Equation 1 specifies the prior probability of the binning model
M
. In our experiments we use a uniform prior for modeling
P
M
. BBQ uses the above Bayesian score to perform model averaging over the space of all possible equal frequency binnings. ... [A] calibrated prediction in our BBQ framework is defined as:
P
z
=
1
y
... where ...
P
z
=
1
y
,
M
i
is the probability estimate obtained using the binning model
M
i
, for the (uncalibrated) classifier output
y
," where Naeini's
P
z
=
1
y
is calculated according to Bayes Rule applied to Naeini's Eq. 1).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of the Waagen/Naeini combination regarding computing probabilities that the values from the distribution parameters are correct with the further teachings of Naeini regarding wherein a confident that the probabilities are corrected is determined according to Bayes' Rule.
The motivation to do so would be to facilitate use of Bayesian model averaging to determine a superior combination of binning models (Naeini, p. 2, Methods: "BBQ extends the simple histogram-binning calibration method (Zadrozny and Elkan 2001) by considering multiple binning models and their combination. The main challenge here is to decide on how to pick the models and how to combine them. BBQ considers multiple equal-frequency binning models that distribute the data-points in the training set equally across all bins. ... We combine them using a Bayesian score derived from the BDeu ... score used for learning Bayesian network structures" and "BBQ uses the above Bayesian score to perform model averaging over the space of all possible equal frequency binnings. We could have also used the above Bayesian score to perform the model selection, which in our case would yield a single binning model. However, model averaging is typically superior to model selection").
Regarding Claim 12, the rejection of Claim 10 is incorporated. The Waagen/Naeini combination teaches:
wherein the parametric function is determined for each combination of labeled class parameters and predicted class parameters (Waagen, p. 8, Figure 6, "Histograms of the
Δ
ϕ
statistic of the validation sets for (a) the simple CNN model and (b) the ResNet-20 model," depicting both the correct pairing and the incorrect pairing of predicted and actual labels).
Regarding Claim 13, the rejection of Claim 9 is incorporated. The Waagen/Naeini combination teaches:
wherein the plural aggregate differences are expressed as a cross-entropy loss (Waagen, p. 4, 4.1 Datasets and Models: "The rescaled Fashion-MNIST training images were used for model fitting and the rescaled test images for model validation. For model fitting the loss function was categorical cross entropy, and empirical risk minimization was performed using the Adam16 variant of stochastic gradient descent (SGD)").
Regarding Claim 14, the rejection of Claim 9 is incorporated. The Waagen/Naeini combination teaches:
wherein the type of data comprises visual images (Waagen, p. 4, 4.1 Datasets and Models, Fashion-MNIST Trained Simple CNN: "Fashion-MNIST contains a 60,000 28x28 pixel image training set, with 6,000 images per class, and a separate 10,000 image test set, with 1,000 image per class").
Regarding Claim 16, the rejection of Claim 15 is incorporated. The Waagen/Naeini combination teaches:
further comprising the step of computing a calibration error for at least one estimated probability (Waagen, p. 9, Figure 7, "Estimated confidence regions using validation data for (a) the simple CNN model, and (b) the ResNet-20 model," depicting confidence as estimated accuracy and thus reasonably suggesting estimated error, where the confidence is referred to as a calibrated confidence at p. 5. Quantifying Confidence: "The final
Δ
ϕ
bins are the assigned
Δ
ϕ
regions and the estimated accuracies with the regions are the confidence estimates. ... The results provide a coarse calibration of the
Δ
ϕ
statistic with model confidence").
Regarding Claim 17, the rejection of Claim 16 is incorporated. The Waagen/Naeini combination teaches:
wherein the step of training the neural network to minimize plural aggregate differences comprises optimizing the weight parameters and bias parameters using a gradient descent (Waagen, p. 4, 4.1 Datasets and Models: "The rescaled Fashion-MNIST training images were used for model fitting and the rescaled test images for model validation. For model fitting the loss function was categorical cross entropy, and empirical risk minimization was performed using the Adam16 variant of stochastic gradient descent (SGD)").
Regarding Claim 18, the rejection of Claim 17 is incorporated. The Waagen/Naeini combination teaches:
wherein the gradient descent is a stochastic gradient descent (Waagen, p. 4, 4.1 Datasets and Models: "The rescaled Fashion-MNIST training images were used for model fitting and the rescaled test images for model validation. For model fitting the loss function was categorical cross entropy, and empirical risk minimization was performed using the Adam16 variant of stochastic gradient descent (SGD)").
Regarding Claim 19, the rejection of Claim 18 is incorporated. The Waagen/Naeini combination teaches:
wherein the step of training the neural network is terminated when the aggregate differences for the validation data set reach a minimum value (Waagen, p. 6, 5. Quantifying Confidence: "Using the model’s validation data, it bins the
Δ
ϕ
values of the correctly and incorrectly classified images into two separate histograms. ... ¶ ... If the p-value is greater than the threshold, the two bins are merged. Once all adjacent bins are either left alone or merged, the process is repeated with the updated bins. The bin merging iteration is only stopped when a stability configuration is reached, i.e., no more bins can be merged with the given p-value threshold," where Waagen's bins correspond to the aggregate differences, and Waagen's threshold corresponds to the instant minimum value), in order to avoid overfitting the model (Waagen, p. 6, 5. Quantifying Confidence: "The bin merging iteration is only stopped when a stability configuration is reached, i.e., no more bins can be merged with the given p-value threshold. ... The results provide a coarse calibration of the
Δ
ϕ
statistic with model confidence," where Waagen's stopping according to the p-value threshold for a coarse calibration reasonably suggests the instant avoiding overfitting).
Regarding Claim 20, the rejection of Claim 9 is incorporated. The Waagen/Naeini combination teaches:
further comprising the step of computing the estimated probabilities of a statistically significant number of samples (Waagen, p. 4, 4.1 Datasets and Models, Fashion-MNIST Trained Simple CNN: "Fashion-MNIST contains a 60,000 28x28 pixel image training set, with 6,000 images per class, and a separate 10,000 image test set, with 1,000 image per class. ... ¶ The rescaled Fashion-MNIST training images were used for model fitting and the rescaled test images for model validation. For model fitting the loss function was categorical cross entropy, and empirical risk minimization was performed using the Adam variant of stochastic gradient descent (SGD)").
Claim 2 is rejected under 35 U.S.C. 103 as being unpatentable over Waagen, et al., "A Geometric Statistic for Deep Learning Model Confidence and Adversarial Defense" (hereinafter "Waagen") in view of Naeini, et al., "Obtaining Well Calibrated Probabilities Using Bayesian Binning" (hereinafter "Naeini") in view of Gong, et al., "Confidence Calibration for Domain Generalization under Covariate Shift" (hereinafter "Gong").
Regarding Claim 2, the rejection of Claim 1 is incorporated.
The Waagen/Naeini combination teaches splitting the data into a training data set and a test data set.
The Waagen/Naeini combination does not explicitly teach wherein the step of splitting the data into a training data set and a test data set comprises randomly splitting the data.
However, Gong teaches:
wherein the step of splitting the data into a training data set and a test data set comprises randomly splitting the data (Gong, p. 5, 5.1. Datasets: "Office-Home [30] contains images of 65 classes across four domains.... We split these four domains into three subsets: one domain as the source for training the classifier, two domains for post-hoc calibration of the classifier, and one holdout domain as the target for evaluating the calibrated classifier. ... We randomly divide data from each domain into a Large subset (80%) and a Small subset (20%). We use the Large subset for either training the classifier or evaluating the calibration performance").
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of the Waagen/Naeini combination regarding splitting the data into a training data set and a test data set with those of Gong regarding wherein the step of splitting the data into a training data set and a test data set comprises randomly splitting the data.
The motivation to do so would be to facilitate evaluation of a trained classifier according to a definition of calibration relying on sampling from the distribution of predicted data (Gong, p. 1, 1. Introduction: "A classifier is calibrated with respect to a distribution (or a dataset sampled from that distribution) if its predicted probability of being correct matches its true probability" and p. 12, Table A.4, "Quantitative evaluation of calibration on the DomainNet dataset [23] for experiments using Quickdraw (Q), Infograph (I) or Sketch (S) as the target domain. For each experiment, we evaluate each of the three proposed algorithms over 1000 experiments with randomly selected test data from the target domain. Results with statistically significant improvement against source-only method are highlighted with asterisks").
Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Waagen, et al., "A Geometric Statistic for Deep Learning Model Confidence and Adversarial Defense" (hereinafter "Waagen") in view of Naeini, et al., "Obtaining Well Calibrated Probabilities Using Bayesian Binning" (hereinafter "Naeini") in view of Waagen, et al., "Adventures in Deep Learning Geometry" (hereinafter "Waagen-b").
Regarding Claim 10, the rejection of Claim 9 is incorporated.
The Waagen/Naeini combination teaches computing angle distributions between the individual decision vectors and the corresponding individual weight vectors.
The Waagen/Naeini combination does not explicitly teach wherein the distribution of angles between the decision vectors and the weight vectors are modeled using a Gaussian distribution.
However, Waagen-b teaches:
wherein the distribution of angles between the decision vectors and the weight vectors are modeled using a Gaussian distribution (Waagen-b, p. 11, Figure 6, "Angular test data distributions with respect to each output layer weight vector for ... decision space model instances," depicting normalized distributions of per-class angles, and p. 8, 4.2 Decision space geometry and dimensionality: "For each output layer neuron, each class' angular distribution is empirically estimated by a normalized histogram
h
^
θ
k
,
j
,
k
=
1
,
…
,
c
. These histograms are color coded to allow visual estimation of class separation (or lack thereof) of these angular distributions. ¶ The histograms of angular distributions
h
^
θ
k
,
j
between the output layer weight vectors and the training set are shown in Figure 5 for four model instances").
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of the Waagen/Naeini combination regarding computing angle distributions between the individual decision vectors and the corresponding individual weight vectors with those of Waagen-b regarding wherein the distribution of angles between the decision vectors and the weight vectors are modeled using a Gaussian distribution.
The motivation to do so would be to facilitate performing comparisons between training and validation result histograms (Waagen-b, p. 8, 4.2 Decision space geometry and dimensionality: "The previous section provides a sense (in terms of class mean vectors in the decision space) of how the class data are angularly related to the output layer neurons. In this section, given the trained simple CNN models, we examine the normalized histograms of class-conditioned angular relationships ... between the data ... and the output layer neuron weight vectors.... ¶ The histograms of angular distributions
h
^
θ
k
,
j
between the output layer weight vectors and the training set are shown in Figure 5 for four model instances. Figure 6 provides the angular histograms of the validation set. ... As can be seen, the training and validation data histograms are quite similar to each other. As expected, the angular relationships of the same-labeled data ... are closer in angle to the output layer weight than the angular distributions of differing-labeled classes .... There is a right-ward shift of the same-labled
h
^
θ
k
,
j
histograms as the dimension of the decision space increases, which was also noted with respect to the (epoch 20) means in Figures 3 and 4").
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ROBERT N DAY whose telephone number is (703)756-1519. The examiner can normally be reached M-F 9-5.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Kakali Chaki can be reached at (571) 272-3719. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/R.N.D./Examiner, Art Unit 2122
/MICHAEL H HOANG/ PRIMARY EXAMINER, Art Unit 2122